EMC Introduces World’s Most Powerful Hadoop Distribution: Pivotal HD
February 25, 2013 No CommentsEMC announced a new distribution of Apache Hadoop: Pivotal HD. Pivotal HD features native integration of EMC’s industry leading Greenplum® massively parallel processing (MPP) database with Apache Hadoop—the most cost-effective and flexible open source Big Data platform ever developed. The new EMC Greenplum-developed HAWQ™ technology brings 10 years of large scale data management research and development to Hadoop and delivers more than 100X performance improvements when compared to existing SQL-like services on top of Hadoop , making Pivotal HD the single most powerful Hadoop distribution in the industry .
Hadoop has rapidly emerged as the preferred solution for Big Data analytics applications that grapple with vast repositories of unstructured data. It is flexible, scalable, inexpensive, fault-tolerant, and enjoys rapid adoption rates and a rich ecosystem surrounded by massive investment. However, customers face high hurdles to broadly adopting Hadoop as their singular data repository due to a lack of useful interfaces and high-level tooling for Business Intelligence and datamining—components that are critical to data analytics and building a data-driven enterprise. As the world’s first true SQL processing for Hadoop, Pivotal HD addresses these challenges.
By offering the full spectrum of the SQL interface, and by extension, the entire ecosystem of products that support SQL, customers no longer need an army of developers to build a dashboard or run a report. Unlike competitive Hadoop distributions, Pivotal HD does this without moving data between systems or using connectors that require users to store the data twice. Pivotal HD cuts out the complexity of using Hadoop, thus expanding the platform’s potential and productivity, and allowing customers to enjoy the benefits of the most cost-effective and flexible data processing platform ever developed.
ABOUT HAWQ
HAWQ (pronounced hawk) represents the EMC Greenplum engineering effort that brings 10 years of large-scale data management research and development to the Apache Hadoop framework. Leveraging the feature richness and maturity of the industry leading Greenplum MPP analytical database, this innovation has resulted in the world’s first true SQL parallel database on top of the Hadoop Distributed File System (HDFS). HAWQ is the key differentiating technology in making Pivotal HD the world’s most powerful Hadoop distribution. Capabilities of note include Dynamic Pipelining, a world-class query optimizer, horizontal scaling, SQL compliant, interactive query, deep analytics, and support for common Hadoop formats.
PIVOTAL HD AND HAWQ DELIVER:
- True SQL Query Capabilities – With Pivotal HD’s advanced database services (HAWQ) enterprises can now unlock the potential of Hadoop’s scalable, fault-tolerant storage capabilities by bringing to bear the vast pool of “data worker” tools and languages into the Hadoop ecosystems. With Pivotal HD’s support for true, SQL-standards compliant query interfaces data mining tools, SQL-trained data analysts, and standard BI tools can now easily connect to, query, and analyze data sets stored in the Hadoop file system (HDFS).
- Unprecedented Query Performance – Bringing over 10 years of parallel database processing technology to Hadoop, Pivotal HD delivers query response time improvements that are up to 600x faster than current SQL-like interfaces for Hadoop.
- Robust Operational Support – Command Center enables administrators and developers to easily install and manage large clusters from interactive web user interfaces. Command Center also exposes Command Line Interface for scripting and programmer friendly web services API for complex automation tasks. Using Command Center administrators can deploy large cluster, configure services/roles, manage services and monitor HDFS jobs and tasks.
HADOOP—THE FOUNDATION FOR CHANGE
EMC believes that Hadoop has the potential to reach beyond Big Data to catalyze new levels of business productivity and transformation. As the foundation for change in business, Hadoop represents an unprecedented opportunity to improve how organizations can get the most value from large amounts of data. Businesses that rely on Hadoop as the core of their infrastructure can not only do analytics on top of vast amounts of data, but can also go beyond analytics and the foundation for that data layer to build applications that are meaningful, and that have a very tightly coupled relationship with the data. Consumer Internet companies have reaped the benefits of this approach, and EMC believes more traditional enterprises will adopt the same model as they evolve and transform their businesses.
Pivotal HD is expected to be available at the end of the first quarter of this year as a software-only or appliance-based solution, backed by EMC’s global 24×7 support infrastructure.
ADDITIONAL RESOURCES
- Learn more about Pivotal HD
- Read more about Pivotal HD on EMC Pulse , EMC’s news and technology blog
- Read the blog from Greenplum Solutions Architect Donald Miner: Introducing Pivotal HD
- Interview: Former CTO for the Obama 2012 Campaign Harper Reed on the Power of Data, at Hadoop: the Foundation for Change
- For the latest industry news, research, webcasts and use cases covering Big Data, analytics, and the data scientist community visit the Greenplum Media Center
- Connect with EMC via Twitter , Facebook , YouTube , LinkedIn and Greenplum
ABOUT GREENPLUM, A DIVISION OF EMC
Greenplum, a division of EMC, is driving the future of Big Data analytics with breakthrough products that harness the skills of data science teams to help global organizations realize the full promise of business agility and become data-driven, predictive enterprises. The division’s products include Greenplum® Unified Analytics Platform, Greenplum® Data Computing Appliance, Greenplum® Database, Greenplum® Analytics Lab, Greenplum® HD and Greenplum® Chorus™. They embody the power of open systems, cloud computing, virtualization and social collaboration, enabling global organizations to gain greater insight and value from their data than ever before possible. Learn more at www.greenplum.com
ABOUT EMC
EMC Corporation is a global leader in enabling businesses and service providers to transform their operations and deliver IT as a service. Fundamental to this transformation is cloud computing. Through innovative products and services, EMC accelerates the journey to cloud computing, helping IT departments to store, manage, protect and analyze their most valuable asset — information — in a more agile, trusted and cost-efficient way. Additional information about EMC can be found at www.EMC.com.