Tuesday, 31 March 2015

Apache Tajos Big Data warehouse comes to Hadoop


The Apache Software Foundation (ASF) has just announced the relatively unknown Apache Tajo open source data warehouse software is ready for commercial use, almost two years after it first became a top-level project, and five years after its development began

Tajo is an SQL-on-Hadoop platform thats designed to help organizations extract more intelligence from their Hadoop deployments, and its first official release comes with updates that provide more connectivity to third-party databases like Oracle and PostGreSQL, plus Java programs

The software helps analyze data stored on the Hadoop Distributed File System (HDFS), as well as other data sources like Amazon S3, Openstack Swift and local file systems

Although less well-known, its somewhat similar to solutions like Apache Hive and Clouderas Impala, providing an extract, transform and load (ETL) feature set and an extensible query re-write system that lets users and external programs query data through SQL

The software could be a good fit for organizations that have outgrown their commercial data warehouses, writes Joab Jackson in Computerworld

On top of that, it could also be a useful solution for companies looking to analyze Hadoop data using more familiar commercial businesses intelligence tools, rather than the MapReduce framework.

No comments:

Post a Comment