The Enterprise Data Hub: A place to store all your data with enterprise grade data management, integrations

Hadoop World/Strata has been full of activity and announcements. We will be providing more through the week. One of the more important announcements was Cloudera’s articulation of an Enterprise Data hub. The significance of this is huge for enterprise data. Imagine a place where you can store all data, structured and unstructured, for a very economical cost. This alone is a fantastic, highly desired capability. The enterprise data hub construct has far more capability and features you would expect from a well engineered solution. This includes enhancements to simplify storage, processing, analyzing and managing data. It also includes enhanced security and auditing. And very tight integration into your existing legacy infrastructure and applications.  This is going to be big.

The press release from Cloudera is below:

Cloudera Enterprise 5 Sets New Standard for Data Management; Lays Foundation for The Enterprise Data Hub

Oct/29/2013

Company Extends Category Leadership With Public Beta Release of CDH 5 and Cloudera Enterprise 5; Unveils Industry’s First Enterprise Data Hub and Analysis Platform

PALO ALTO, CA and NEW YORK, NY–(Marketwired – Oct 29, 2013) - From Strata + Hadoop World: Cloudera, the leader in enterprise analytic data management powered by Apache Hadoop™, today unveiled the fifth generation of its Platform for Big Data, Cloudera Enterprise, which is now available for public beta. The new product release, powered by Apache Hadoop 2, offers unique features and advancements that simplify storing, processing, analyzing and managing large structured and unstructured datasets, while offering increased security, robust data management and tight integration with third-party applications. The combination of innovative updates to CDH (Cloudera’s Distribution Including Apache Hadoop) at the core — plus enhancements to Cloudera Manager for Hadoop system administration and Cloudera Navigator for Hadoop audit and access control, data discovery and lineage analysis — together deliver the industry’s first Enterprise Data Hub.

“With Cloudera Enterprise 5, Cloudera has taken several important steps toward realizing its vision to transform Hadoop into an enterprise data hub for analytics,” said Tony Baer, Principal Analyst for Ovum. “Adding support for in-memory data tiering and user-defined functions are essential for delivering the kind of performance that enterprises expect from their analytic data platforms.”

Rethink Data: Introducing the Enterprise Data Hub
Organizations currently employ a variety of systems to support their diverse data hub goals: data warehouses for operational reporting; storage systems to keep data available and safe; specialized massively-parallel databases for large-scale analytics; and search systems for finding and exploring documents. While these systems are suitable for traditional data and workloads, they are not equipped to handle today’s exponential growth in data volume and variety, or the range of users who seek insights from that data. And because each system is purpose-built for a particular class of data and workload, no single system can provide unified access to all relevant information to diverse business users. A new hybrid approach is required, which pragmatically extends the value of existing investments while enabling fundamentally new ways of delivering value from data.

The objective is simple: Acquire and combine any amount or type of data in its original fidelity, in one place, for as long as is necessary, and deliver insights to all kinds of users, as fast as possible. And do so with maximum efficiency of capital and resources.

The solution? The Enterprise Data Hub. One place to store and work with all data, with the flexibility to run a variety of enterprise workloads — including batch processing, interactive SQL, enterprise search and advanced analytics — together with the integrations to existing systems, robust security, governance, data protection, and management that enterprises require. The Enterprise Data Hub is the emerging and necessary center of enterprise data management, complementing existing infrastructure.

Cloudera Enterprise 5: Next Generation Platform for Big Data powered by Apache Hadoop
Built for the demanding requirements of enterprise customers, Cloudera Enterprise enables companies to store, process and analyze unlimited amounts of data and applications from a single system. The newest innovations in Cloudera Enterprise 5 offer customers a significant leap forward in the evolution of the platform, which can now be used to efficiently address an even wider range of business problems. Customers can now use Cloudera to easily handle the rapidly increasing data volume and variety they face, absorbing a growing share of data and workloads from legacy infrastructure while optimizing the efficiency of those existing systems.

Cloudera Enterprise 5 offers a single platform from which organizations can tackle diverse critical business problems:

  • Automatically archiving the complete set of enterprise data to meet compliance requirements while retaining queryable access;
  • Complementing data warehouses to offload data and workloads to help customers increase efficiency and manage costs, while delivering faster ETL/ELT data processing at scale;
  • Supporting business intelligence, through familiar tools, on more data and more kinds of data than ever before possible;
  • Enabling and consolidating enterprise search on data and documents in-place within the single environment; and
  • Accelerating a diverse array of advanced analytics solutions, like recommendation engines, fraud detection or image processing.

Increasingly, strategic partners like Informatica are certifying reference architectures to bring these benefits to joint customers. For example, Informatica and Cloudera together provide a “Data Warehouse Optimization” solution to address the challenges facing traditional data warehouse infrastructures, where capacity is too quickly consumed by increasing data volumes, leading to performance bottlenecks and costly upgrades.

Key advances in Cloudera Enterprise 5 include:

Accelerated Time-to-Value

  • In-Memory HDFS Caching: Datasets from HDFS can now be cached in-memory, boosting MapReduce data processing performance and Cloudera Impala’s analytic query response times for even faster time to insight.
  • User-Defined Functions (UDFs): Customers can now use the custom query functions they depend on in conjunction with Cloudera Impala to deliver the business insights they require. They can also take advantage of the popular open source MADlib library of pre-built statistical and analytic functions to enable scalable in-database analytics.

Improved Efficiency

  • Resource Management: Cloudera Enterprise now delivers advanced resource management for running multiple frameworks for data processing and analysis on a single cluster through the powerful combination of Hadoop YARN (Yet Another Resource Negotiator) and Cloudera Manager. For the first time, administrators can allocate resources not only by workload, but by workgroup, ensuring the best combination of performance and utilization. For example, customers can dedicate 50% of capacity for IT to run mission critical data processing jobs, 30% to the marketing team for ad-hoc BI queries, and so on.
  • Unified Management of Third Party Applications. Cloudera Manager now provides extensibility to enable customers to deploy, manage and monitor products from Cloudera partners such as SAS, Revolution Analytics, Syncsort and many more. Now, customers can manage complex clustered environments from within a single, intuitive interface.

Comprehensive Data Management

  • Manage and Explore Big Data. In addition to enabling centralized data auditing for Hadoop, Cloudera Navigator now provides:
    • Data Discovery: Analysts and data modelers can search, explore, define and tag datasets through the Cloudera Navigator interface, to help identify relevant information for downstream analysis or processing.
    • Data Lineage: As the amount of data in Cloudera Enterprise grows, so does the importance of understanding how that data is used across the organization. Cloudera Navigator delivers the industry’s first data lineage solution for Hadoop, enabling customers to meet regulatory requirements, find associated datasets, and satisfy data governance and retention policies.
  • Data Protection: HDFS and HBase now support snapshots to help prevent data loss.
  • NFS-based Data and Application Access: Easily integrate Cloudera Enterprise with data in and applications running on existing filesystems with native support for NFSv3.

“Over the last five years, we have worked closely with enterprises around the world to help them capture the value in the data they have. Resoundingly, they have asked for a more secure, more reliable real-time data platform that streamlines their existing architectures and speeds up time to insight,” said Mike Olson, chairman and chief strategy officer, Cloudera. “The market has spoken and we are listening. The new capabilities introduced in Cloudera Enterprise 5 deliver the industry’s first Enterprise Data Hub.”

Product Availability and Documentation
Public beta releases of Cloudera Enterprise 5 and CDH 5 are now available. To learn more about Cloudera Enterprise 5, visit http://cloudera.com/CE5. To learn more about CDH 5, or to download it for free, visithttp://www.cloudera.com/content/cloudera/en/products/cdh.html.

The Cloudera Enterprise Data Hub is available today on Cloudera Enterprise 4, for more information contact Cloudera on info@cloudera.com.

This information is not a commitment, promise or legal obligation to deliver any material, code, or functionality. Cloudera does not guarantee that the beta software will be made generally available or that any individual feature in the beta version will be made generally available. Cloudera may make the beta software generally available, or not, in its sole discretion and without obligation to make any communication of any kind with regard to such availability.

About Cloudera
Cloudera is revolutionizing enterprise data management by offering the first unified Platform for Big Data: The Enterprise Data Hub. Cloudera offers enterprises one place to store, process and analyze all their data, empowering them to extend the value of existing investments, while enabling fundamental new ways to derive value from their data. Founded in 2008, Cloudera was the first, and is still today, the leading provider and supporter of Hadoop for the enterprise. Cloudera also offers software for business critical data challenges, including storage, access, management, analysis, security and search. With over 15,000 individuals trained, Cloudera is a leading educator of data professionals, offering the industry’s broadest array of Hadoop training and certification programs. Cloudera works with over 700 hardware, software and services partners to meet customers’ big data goals. Leading organizations in every industry run Cloudera in production, including finance, telecommunications, retail, internet, utilities, oil and gas, healthcare, biopharmaceuticals, networking and media, plus top public sector organizations globally. www.cloudera.com

Connect with Cloudera
Read our blog: http://blog.cloudera.com/blog/
Follow us on Twitter: https://twitter.com/cloudera
Visit us on Facebook: https://www.facebook.com/cloudera

Cloudera, Cloudera Manager, Cloudera Navigator, CDH, Cloudera Enterprise, Cloudera Standard and Cloudera Enterprise Data Hub are trademarks or registered trademarks of Cloudera in the United States and in jurisdictions throughout the world. All other company and product names may be trade names or trademarks of their respective owners.

CTOvision Pro Special Technology Assessments

We produce special technology reviews continuously updated for CTOvision Pro members. Categories we cover include:

  • Analytical Tools - With a special focus on technologies that can make dramatic positive improvements for enterprise analysts.
  • Big Data - We cover the technologies that help organizations deal with massive quantities of data.
  • Cloud Computing - We curate information on the technologies enabling enterprise use of the cloud.
  • Communications - Advances in communications are revolutionizing how data gets moved.
  • GreenIT - A great and virtuous reason to modernize!
  • Infrastructure  - Modernizing Infrastructure can have dramatic benefits on functionality while reducing operating costs.
  • Mobile - This revolution is empowering the workforce in ways few of us ever dreamed of.
  • Security  -  There are real needs for enhancements to security systems.
  • Visualization  - Connecting computers with humans.
  • Hot Technologies - Firms we believe warrant special attention.

 

Recent Research

Finding The Elusive Data Scientist In The Federal Space

DoD Public And Private Cloud Mandates: And insights from a deployed communications professional on why it matters

Intel CEO Brian Krzanich and Cloudera CSO Mike Olson on Intel and Cloudera’s Technology Collaboration

Watch For More Product Feature Enhancements for Actifio Following $100M Funding Round

Navy Information Dominance Corps: IT still searching for the right governance model

DISA Provides A milCloud Overview: Looks like progress, but watch for two big risks

Innovators, Integrators and Tech Vendors: Here is what the government hopes they will buy from you in 2015

Navy continues to invest in innovation: Review their S&T efforts here

MSPA Unified Certification Standard For Cloud Service Providers: Is This A Commercial Version of FedRamp?

Watch Ben Fry And His Visualizations: Multiple use-cases come to mind, including national security efforts

Agenda And More Details for 4-5 March NIST Data Science Symposium

Actionable Insights From AFCEA Western Conference and Exposition 2014

solid
About Bob Gourley

Bob Gourley is the publisher of CTOvision.com and DelphiBrief.com and the new analysis focused Analyst One Bob's background is as an all source intelligence analyst and an enterprise CTO. Find him on Twitter at @BobGourley