Dell, EMC, Dell Technologies, Cisco,

Showing posts with label MapR. Show all posts
Showing posts with label MapR. Show all posts

Saturday, March 24, 2018

Elephant in the room: Hortonworks CEO thinks Hadoop software will keep driving big data

Since its founding in 2011, Hortonworks Inc. has fought battles on two fronts: Persuade corporations to adopt an entirely new, open-source data platform called Hadoop for a novel type of analytic processing and build a business selling software that customers can get elsewhere for free. By most accounts, @Hortonworks is making good progress. The company recorded its first-ever positive cash flow in its most recent earnings report, with revenues up 44 percent from a year ago, to $75 million. Competitors @MapR Technologies Inc. and @Cloudera Inc. have been gradually edging away in recent years from their roots in #Hadoop, which was named after the beloved toy elephant of one of the creators’ children. But Hortonworks remains steadfastly committed to the big-data platform as the core of its business. Chief Executive @Rob Bearden (pictured) recently discussed that strategy in an interview with SiliconANGLE. Your early competitors in the Hadoop market have largely deemphasized Hadoop as a focus, while Hortonworks hasn’t. Why?   Customers spend more than $100 million a year with us, and the vast majority of that is Hadoop subscriptions. We’re managing the entire lifecycle of the data from the point of origination to data at rest, and we do the at-rest part with Apache Hadoop. We’re in the platform business with three principal platforms: Hortonworks DataFlow brings data off the edge and processes it at any point as an event, condition or action. You can to take a proscriptive and surgical approach to that data. When we bring data to rest, we manage and process it with Hadoop. Our third platform is DataPlane. It allows us to deploy those workloads wherever it’s best from an economics or physics standpoint. Many analysts are saying the future of big data is in the cloud. Do you agree? The math says they’re wrong. If you look at the revenue stream coming in to our company, two-thirds is on-prem. How have customer attitudes evolved since you started the company? It’s now assumed that they’re going to have Hadoop in an architecture for managing significant volumes of new-paradigm data. The math says data volumes are doubling every 14 to 16 months in the enterprise, and 80 percent of that is coming from new data sources. There’s nowhere else to go except Hadoop. [Enterprise resource planning] was intended to be a central hub for the entire order-to-cash lifecycle with one common view for the customer, but now that’s highly fractured with multiple ERPs and best-of-breed apps. Customers have gone further away from having a common central view. They can’t increase their velocity with having an aggregate view of the data. Hadoop is the most economical platform to do that. Looking back, which trends in big data took off faster than you expected and which didn’t? What surprised me is the acceleration of horizontal uses of Hadoop. Enterprise have learned the value of bringing all this data together on Hadoop for predictive insights that allowed them to engage earlier with customers and the supply chain. In the old world, everything was post-event. With Hadoop, they can bring all the data together about progress up to and through the event. What was surprising was how fast customers wanted to go all the way to the edge and grab that data from the point of origin. They want to engage with data as far upstream as they can. We bet on that when we bought Onyara’s data stream platform.

https://siliconangle.com/blog/2018/03/23/elephant-room-hortonworks-ceo-thinks-hadoop-software-will-keep-driving-big-data/

Thursday, March 15, 2018

Cloudera founder Mike Olson: ‘We’re moving from automating processes to automating decisions’

The growing momentum of big data in the cloud has been described as a threat to @Cloudera Inc., which is ironic given that the company’s original business plan was to sell big data as a cloud service. The market wasn’t ready in 2008, so Cloudera shifted to selling an integrated platform that combines various open-source projects with proprietary extensions. Of the three prominent startups to emerge from the @Hadoop ecosystem, the other two being @MapR Technologies Inc. and @Hortonworks Inc., Cloudera was the most prominent, in no small part because of the $740 million investment it received from @Intel Corp. in 2014. Its stock-market performance since going public last April has been underwhelming, but few would deny that it’s a market leader. Chief Strategy Officer Mike Olson, one of Cloudera’s four founders, joined SiliconANGLE recently for a telephone interview on where big data’s going next. Gartner has estimated that 85 percent of enterprise big data projects failed. Does that surprise you? I don’t understand who they’re talking to. We’re a high-growth company in the $300 million-plus forecast range and most of our business is still on-prem. People aren’t shutting down large football fields of Teradata [Corp.], but the opportunity was never to displace data warehouses. It was to capture more data than we could before and see what we could do if we had better tools to derive value. I think Gartner is comparing the wrong things. The growth rate of machine learning is zero if you compare it against traditional markets because there is no traditional market. But it has created value in huge new ways. It’s true that the biggest growth area has been in cloud databases like Cloudera Altus and Redshift on Amazon Web Services. We’ve believed since the early days that much of the action would be in cloud services, and that’s why we named the company Cloudera. We still believe that, but I think there’s enormous potential and success in on-prem deployments. Three years ago, Cloudera defined its mission around a unified platform built on Hadoop, the pioneering framework for managing big data. Today, you don’t even mention Hadoop in your description. What changed? We talked about Hadoop in the early days because people needed to know what the platform was and what we had. Today it’s a much richer environment. We’re seeing native machine learning using Spark integrated with AWS storage buckets. There’s no Hadoop in there. This platform is doing more than just what Hadoop did. Our original platform was Hadoop and MapReduce. Today, we ship 26 different open-source projects, 18 of which were created by Clouderans. Hadoop is always going to be part of the foundation of the company — and we’re proud that we spotted it so early — but it’s only a small part of what we do today.

https://siliconangle.com/blog/2018/03/13/cloudera-founder-mike-olson-moving-automating-processes-automating-decisions/

Tuesday, December 19, 2017

Microsoft's cloud Big Data service cuts prices up to 52 percent

@Microsoft has decided to get down to business with @HDInsight (HDI), its #Azure cloud-hosted #BigData offering, based on @Apache #Hadoop, @HBase, @Spark, @Storm, @Kafka, @Hive LLAP, and #MicrosoftRServer. Ostensibly, Microsoft previously considered its competition to be on-premises Hadoop implementations. But it now offers pricing that is far more competitive with @Amazon Web Services' (AWS') @Elastic @MapReduce (EMR), while still offering a 3-nines service level agreement (SLA) as a differentiator. Details, details The pricing changes, highlighted in a blog post by Microsoft's Rimma Nehme and detailed on a separate page, offer varying price cuts depending on the virtual machine type used for the head and worker nodes in the HDInsight cluster. Price cuts are up to 52 percent, Microsoft says, while the service itself remains largely the same. In addition, for those customers wishing to run data science workloads with code written in R, the surcharge for running R Server in a distributed fashion on an HDI cluster has been cut by 80 percent, down to just $0.016 (i.e. 1.6 US cents) per CPU core, per hour. Microsoft points out that because of Azure's numerous global data centers (regions), HDI is available at more points of presence than any other cloud Hadoop service. In addition to Azure's mainstream cloud, the service is also available on its US government cloud and on so-called sovereign clouds, including those in Germany and China. Per various regulatory requirements, the sovereign clouds run in facilities operated by local partners, rather than Microsoft itself. In other news Microsoft has a few other announcements to go with the price change: The introduction, available in preview, of HDInsight Enterprise Security Package, which integrates Microsoft Active Directory with Apache Ranger. This is essentially a re-branding of HDInsight's Premium cluster tier General availability of its Apache Kafka cluster type, which had been in preview until recently General availability of HDInsight's integration with Azure Log Analytics Public preview of support for Power BI DirectQuery, specifically against Hive LLAP, used in HDI "Interactive Query" cluster types New HDInsight add-in developer tools for IntelliJ, Eclipse and Visual Studio Code (Microsoft's cross-platform code editor for MacOS, Linux and Windows). The IntelliJ and Eclipse tooling include the ability to submit and debug distributed Spark code right from those development environments. The VS Code tooling allows for interactive execution of PySpark (a Python library for Apache Spark) code.

http://www.zdnet.com/article/microsofts-cloud-big-data-service-cuts-prices-up-to-52/

Tuesday, December 12, 2017

Cisco & MapR set a Software Defined Storage World Record

It sounds like a long time, but first we had to wait for a few milestones to precede us: UNIVAC I 1951 – UNIVAC pioneers use of magnetic tape for storage 1993 – Severe Tire Damage is the first band to live stream 2007 – Netflix launches their streaming media business and finally, 2017 – @Cisco & @MapR achieve a #SPEC #SFS world record for streaming workloads   The #DataAge That’s some pretty impressive company.  Why is this particular benchmark result so significant? Because data is the lifeblood of business.  According to IDC, we’ve entered the Data Age: The world’s data will double every two years for the next decade.  How companies store, manage, access, and deliver hypercritical data will determine winners and losers. But such volume of data cannot be easily managed with conventional systems.  As data-driven businesses gain prominence, demand is growing for integrated solutions encompassing high-density storage, powerful compute, and improved scalability: Enterprises are collecting and storing more data for longer periods to support predictive analytics. Data is classified by “temperatures” (hot/warm/cold) for dynamic data tiering. Decision makers want to access and analyze data for real time insights. But traditional storage arrays are expensive and don’t easily support large scale-out landscapes.   A New Approach to Software Defined Storage That’s why Cisco developed the UCS S3260 Storage Server, a modular architecture designed to deliver efficient, industry-leading storage for data-intensive workloads.   And that’s why we’re partnering with MapR to provide a complete end-to-end software defined storage solution. Our S3260 Storage Server offers the industry’s best performance in a compact 4RU form factor.  It supports either Scale Up (expandable to 28 drives/node), or Scale Out (easily adding multiple nodes to your integrated infrastructure), and is a proven Hadoop platform.   MapR-XD provides a highly reliable distributed data fabric with enterprise-grade features such as Global Namespace and Multi-temperature store.   We Blew It Away Our SPEC SFS world record demonstrates the performance of the S3260 with MapR-XD. SPEC’s benchmark suite is the standardized method for evaluating performance using file server throughput and response time.  We ran the Video Data Acquisition (VDA) streaming workload because it simulates applications that store data acquired from temporally volatile sources such as surveillance cameras. Cisco and MapR have been setting records since we published the industry’s first Big Data benchmark in 2015.    But this was the first SPEC SFS benchmark on a HDFS (Hadoop file system) compatible platform.  Together, Cisco and MapR blew away all previously published SPEC SFS benchmarks: 2070 streams with an overall response time of just 12.94 msecs! You can read the complete, audited benchmark report here.   Back in the Real World But companies don’t run benchmarks for a living.   What’s this mean for our customers? Referring back to our data delivery milestones, the UCS S3260 Storage Server is perfect for modern data-intensive workloads like big data and video steaming.  Analytics and content distribution applications require both simultaneous scalability and high performance.  This need is ubiquitous in industries as different as video surveillance and entertainment. Remember the frustrating mid-movie buffering we put up with back in 2007?  Yesterday I watched Sunday Night Football’s live stream on my laptop connected to the Cisco network without so much as a hiccup. Nor did I expect one.  We’ve become accustomed to a right-now world.  We generate and consume more data in real time today than ever before in history.  As our demand for storage capacity and performance in this Data Age continues to increase, we’ll want greater density and higher performance at lower costs. Cisco and MapR deliver on those expectations.   The innovative, modular design of our UCS S3260 Storage Server allows independent refresh of compute, storage, and network.  And it scales to over ½ PB.   Like the rest of Cisco’s Intent-based Data Center portfolio, the S3260 benefits from profile-based management, 40gE speed, and best-in-class Security to create the most flexible, programmable, cost effective infrastructure you can deploy. Click here to learn more about how the UCS 3260 Storage Server can benefit your Data Center.

https://blogs.cisco.com/datacenter/cisco-mapr-sds-world-record

Monday, December 11, 2017

Apache Bigtop Adds OpenJDK 8 Support

@Apache has released #Bigtop 1.2.1 with support for #OpenJDK 8, and a new sandbox feature that lets you run big data pseudo clusters on @Docker. Bigtop is an Apache Foundation project that you can use for packaging, testing, and configuration of the big name open source big data components that make up the #Hadoop infrastructure. Bigtop supports a wide range of components and projects, including Hadoop, @HBase and @Spark. The primary goal of Bigtop is to build a community around the packaging, deployment and interoperability testing of Hadoop-related projects. This includes testing at various levels, including packaging, platform, runtime, and upgrade, and focussing on the system as a whole, rather than individual projects. While Hadoop is generally used to refer to the central collection of tools, Bigtop looks at the wider selection that makes up the Hadoop-related projects, including Hbase, Pig, Hive MapReduce, Zookeeper and Avro among others.  Bigtop packages Hadoop RPMs and DEBs, so that you can manage and maintain your Hadoop cluster, and it provides an integrated smoke testing framework, alongside a suite of 50 test files. It also helps with virtualization testing, with vagrant recipes, raw images, and (work-in-progress) docker recipes for deploying Hadoop from zero, and you can use Bigtop Provisioner to spin-up a virtual cluster with a single command. The new release of Bigtop adds a Sandbox feature that lets you use Docker to run pseudo clusters. Creating it is a single line command, and for HDFS the process takes around 30 seconds. You then get a local WebUI to play around it. You can run HDFS and Spark standalong, or HDFS, Yarn, Hive and Pig - there are simple instructions on the Bigtop site. This presentation from DataWorks Summit 2017 on using the Sandbox comes from Apache Bigtop Project Committer and PMC member, Evans Ye:   The new release also includes a faster Docker Provisioner which has been rewritten to fully embrace the Docker ecosystem. The OpenJDK support in the new release means that all the components are now built on JDK8. The major components have been updated to recent versions, including Hadoop 2.7.3, Spark 2.1.1, HBase 1.1.9, and Zeppelin 0.72, and most of the ecosystem projects have also been updated, including Apex, Crunch, Flume, Ignite, Mahout, Oozie, and Phoenix, among others.

http://www.i-programmer.info/news/197-data-mining/11374-apache-bigtop-adds-openjdk-8-support.html

Wednesday, November 29, 2017

Find Out Current Status of Hadoop-as-a-Service Market by Manufacturers, Type and Application, Forecast to 2022

Hadoop is a provisioning model offered to organizations seeking to incorporate a hosted implementation of the #Hadoop platform. @Apache Hadoop is an open-source software platform that uses the @MapReduce technology to perform distributed computations on various hardware servers. Hadoop-as-a-service ( #HDaaS ) providers offer #HadoopPaaS, which enables technical experts of enterprises to perform various operations including #bigdata #analytics, #bigdata management, and big data storage in a cloud. top players in global market, like @Amazon Web Services @Microsoft @IBM @EMC Corp @Altiscale   Vendors provide the HDaaS platform as a web-based subscription service on a pay-per-use basis. The platform eliminates the need for any on-premise hardware; it abstracts the Hadoop architecture and supporting applications into a single cloud-based offering. The HDaaS platform enables enterprises to use the Hadoop technology in a cost-effective manner, while ensuring minimal time consumption. Market Highlights: Hadoop has become a leading platform for big data analytics today. Hadoop-based applications are used by enterprises which require real-time analytics from data such as video, audio, email, machine generated data from a multitude of sensors and data from external sources such as social media and the internet. Hadoop-as-a-service enables technical experts of organizations to perform several operations which include big data management, big data analytics and big data storage in a cloud. The Hadoop-as-a-Service platform enables organizations to use Hadoop technology in a highly cost-effective manner, along with ensuring minimal consumption of time. Hadoop-as-a-service is being widely accepted across various industries including IT, banking, manufacturing and telecommunication among others. One of the emerging trends in this market is the increased adoption of Hadoop-as-a-Service by small and medium enterprises (SMEs). In fact, SMEs have been among the earliest adopters of this technology and cloud computing, as this end-user segment is already conversant with the benefits associated with cloud computing. Owing to this, the Hadoop-as-a-Service providers are looking to capitalize on the increase in demand of this technology from the SME segment. Download Sample Copy @ https://www.precisemarketreports.com/report/sample/pmr-98890 Key Players: This report studies the global Hadoop-as-a-Service (HaaS) market, analyzes and researches the Hadoop-as-a-Service (HaaS) development status and forecast in United States, EU, Japan, China, India and Southeast Asia. This report focuses on the Highlights of the report: A complete backdrop analysis, which includes an assessment of the market Important changes in market dynamics Market segmentation up to the second or third level Historical, current, and projected size of the market from the standpoint of both value and volume Reporting and evaluation of recent industry developments Market shares and strategies of key players Emerging niche segments and regional markets An objective assessment of the trajectory of the market Recommendations to companies for strengthening their foothold in the market

http://www.satprnews.com/2017/11/29/find-out-current-status-of-hadoop-as-a-service-market-by-manufacturers-type-and-application-forecast-to-2022/

Tuesday, November 28, 2017

Cloudera Bringing Impala to AWS Cloud

#Apache #Impala, the #SQL-based #analytical #database that originated at #Cloudera, will soon be available as a managed service on the @Amazon Web Services cloud, the #Hadoop software distributor announced today at #AWS Re:Invent. Cloudera Altus Analytic DB, as the hosted cloud version of Impala will be known, will be available as a beta service on AWS by the end of the year, with general availability expected in 2018. The company says support for the #Microsoft #Azure cloud will follow, but no timeline was given. Impala is one of the most popular engines in Cloudera’s Distribution of Hadoop (CDH), and the open source software is also offered in other Hadoop distributions. The software essentially allows customers to run a host of standard SQL queries against massive stores of relational data stored in Parquet, an optimized Hadoop file format. It’s not the only parallel SQL data warehouse designed to run atop Hadoop, but it is one of the most mature. Getting customers the capability to run SQL workloads against data hosted in cloud-based object storage repositories was a big priority for Cloudera, says Alex Gutow, a product manager with the Palo Alto, California-based company. “When we look around at our customers and the types of workloads that are really best suited to take advantage of the agility and cost efficiencies of the cloud, BI and analytics is one of the key workloads,” Gutow tells Datanami. Specifically, BI workloads running on Impala will benefit tremendously from features customers can find on AWS cloud, such as multi-tenant isolation and workload elasticity, Gutow says. Instead of leaning on in-house administrators to procure computational resources and then manage it for the user on an on-going basis, the Altus Analytic DB service allows customers to lean on Amazon and Cloudera to do that heavy lifting for them, she says. “So you can have a very specific cluster to run reporting workloads, and you can have another to run ad hoc queries or self-service BI,” Gutow continues. “It allows for much more of that agility, giving all different types of analysts access to shared data very quickly, giving them much more flexibility, and being able to elastically scale up those resources as you need to meet different performance requirements, or to make sure there’s predicable performance for those workloads.” Altus Analytic DB will access data stored in customers’ Simple Storage Service (S3) accounts. Impala has been able to access data stored in the S3 object store via an HDFS API for about a year, says Greg Rahn, a Cloudera product manager for Impala and Altus. “HDFS provides an API to S3 known as an S3A connector. Impala uses this,” Rahn says. “It looks to Impala as if it’s kind of the HDFS file system, or the abstraction thereof.” There’s a similar API that exposes data stored in Microsoft’s cloud object store, ALDS, through HDFS, and the company will use that connector when Altus Analytics DB is supported on the Azure cloud in the future. “So at the end of the day, whether the data is in S3 or ALDS or HDFS itself, it all kind of looks the same in terms of the visibly it to impala,” Rahn says. Cloudera has pre-selected certain Elastic Cloud Compute (EC2) instances that Altus Analytic DB will be allowed to run upon. Customers will be able to spin up Altus Analytic DB clusters with just a few clicks of the mouse, Rahn says.  Cloudera is positioning its SDX as a key product uniting customers’ cloud and on-premise Hadoop deployments “The Altus deployment makes it quite trivial to start up these things, probably on the order of three to four clicks to provision a cluster,” he says. “You log in, name the cluster, pick the size of the instance, the number you want, then you hit ‘create cluster.’ So it’s very simple.” Customers will be able to quickly spin up analytic clusters on AWS, run a workload, and then quickly dispose of it. None of the data, metadata, or state information for these jobs will be lost when the cloud cluster is deleted because it’s all managed centrally under Cloudera’s Shared Data Experience (SDX), which the company announced at the Strata Data Conference in September. The SDX provides a way to manage data access and permissions for on-premise and cloud environments from a central console. The software sports hooks into core management tools, including Cloudera Navigator, Cloudera Manager, and Sentry for on-premise implementations and Altus controls for cloud-based environments. “Not only can you provide these different isolated resources for each of the different workloads,” Gutow explains, “but from the management side of things, these all benefit from having shared security, shared governance, and shared metadata as running actors shared data layer in the cloud, the shared object storage. So each time any of these different workloads are provisioned or run for different self-service workloads, you don’t have to go and redefine the different security policies. You can easily manage them from an enterprise standpoint.” Altus Analytic DB will be the second hosted offering under the Altus banner since Cloudera announced its new platform as a service (PaaS) in May. The first offering, Altus Data Engineering, was focused on data ingest and transformation tasks, and includes Spark, Hive, Hive on Spark, and MapReduce2 engines. Cloudera was mum on what engines will come next for Altus. Kudu, its fast-data layer, is one obvious candidate. Cloudera is currently in a quiet period before it announces financial results on December 7.

https://www.datanami.com/2017/11/28/cloudera-bringing-impala-aws-cloud/

Monday, November 27, 2017

Key Insights on the Global Hadoop Market | Technavio

LONDON--(BUSINESS WIRE)--The latest market research report by Technavio on the global Hadoop market predicts a CAGR of more than 39% during the period 2017-2021. Global #Hadoop market is set to grow at a CAGR of more than 39% during the period 2017-2021. @Technavio Tweet this The report has further categorized the global Hadoop market into different segments by end-user (banking, financial services, and insurance sector, government sector, communications sector, healthcare sector, and others) and by geography (the Americas, EMEA, and APAC). It provides a detailed illustration of the major factors influencing the market, including drivers, opportunities, trends, and industry-specific challenges. Here are some key findings of the global Hadoop market, according to Technavio ICT researchers: Growing structured and unstructured data: a major market driver The BFSI sector was the largest end-user of Hadoop in 2016 In 2016, the Americas dominated the global Hadoop market with a share of more than 55% The major players in the market include #Amazon, #Cloudera, #Hortonworks, #IBM, #MapR Technologies, #Microsoft, #Pivotal Software and #Teradata This report is available at a USD 1,000 discount for a limited time only: View market snapshot before purchasing Buy 1 Technavio report and get the second for 50% off. Buy 2 Technavio reports and get the third for free. Growing structured and unstructured data: a major market driver Growing structured and unstructured data is one of the key factors driving the global Hadoop market. Enterprise data, including structured data and unstructured data, is generated from various sources such as enterprise applications, web-based search, social networks, and cloud-based applications. The data coming from embedded systems and metadata are some of the fastest-growing data segments. Hadoop is necessary for organizations to process the huge volumes of big data generated and to use the data effectively. Hadoop and big data analytics assist enterprises to optimize their business decisions and innovate new business models, products, and services offerings. According to Amrita Choudhury, a lead analyst at Technavio for research on enterprise application, “Accessibility to Hadoop, as a part of or as an extension to the corporate information framework, is necessary. It should be made available for analysis and decision-making. Companies use Hadoop and Spark for real-time analytics. They are also used for fraud detection, product design and development, and process automation. Therefore, the growth of structured and unstructured data is expected to fuel the demand for Hadoop during the forecast period.” Looking for more information on this market? Request a free sample report Technavio’s sample reports are free of charge and contain multiple sections of the report including the market size and forecast, drivers, challenges, trends, and more. BFSI sector: largest end-user segment Hadoop is utilized across the BFSI sector for various applications such as fraud detection, data security, customer intelligence, data modeling, neural network scoring, social media management, and customer analysis. Big data on Hadoop assists to pick up odd patterns and alerts the bank of the same. Sigorta Bilgi ve Gözetim Merkezi (SBM), also known as Insurance Information and Monitoring Center, is a non-commercial legal entity in the body of the Union of Insurance and Reinsurance Companies of Turkey. SBM uses SAS Fraud Framework and analytics to improve the fraud detection rates and to focus more on organized fraud cases. Competitive vendor landscape The global Hadoop market is not intensely fragmented. The dominating players in the market are Hortonworks, Cloudera, MapR Technologies, and Microsoft. There is intense competition among these vendors. Owing to the increased competition, consolidation is being observed in the industry wherein smaller players are being acquired by or merged with the major players. Moreover, the changing technological environment is a major challenge for the global vendors. To survive and succeed in this intensely competitive environment, it is essential that the vendors differentiate their products and services through clear and unique value propositions.

http://www.businesswire.com/news/home/20171126005027/en/Key-Insights-Global-Hadoop-Market-Technavio

Sunday, November 12, 2017

Hadoop – Global Market to witness astonishing growth of 39% in next years

HTF’s analysts forecast the global hadoop market to grow at a CAGR of 39.49% during the period 2017-2021. HTF recognizes the following companies as the key players in the global #Hadoop market: @Amazon, @Cloudera, @Hortonworks, @IBM, @MapR Technologies, @Microsoft, @Pivotal Software, and @Teradata. Commenting on the report, an analyst from HTF’s team said: “One trend in the market is increasing market consolidation and partnerships. The consolidation in the global Hadoop market is increasing. Many large enterprise computing vendors are increasingly acquiring companies to attain new big data technologies. Larger vendors are targeting smaller companies to expand their business portfolio in the global big data and Hadoop market.”

http://www.satprnews.com/2017/11/10/hadoop-global-market-to-witness-astonishing-growth-of-39-in-next-years/

Monday, November 6, 2017

Should Spark In-Memory Run Natively On IBM i?

There’s a revolution happening in the field of data analytics, and an #opensource computing framework called Apache Spark is right smack in the middle of it. Spark is such a powerful tool that IBM elected to create a distribution of it that runs natively on its System z mainframe. Will it do the same for its baby mainframe, the IBM i? So, what is Apache Spark, and why should you care? Great questions! Let’s introduce you to Spark. Spark came out of UC Berkeley’s AMPLab about five years ago to provide a faster and easier-to-use alternative to MapReduce, which at that point was the primary computational engine for running big data processing jobs on Apache Hadoop. While Spark has a learning curve of its own, the Scala-based framework has not only replaced Java-based MapReduce, but also eclipsed Hadoop in importance in the emerging big data ecosystem. Spark is useful for developing and running all sorts of data-intensive applications, including familiar programs like ETL jobs and SQL analytics, as well as more advanced approaches like real-time stream processing, machine learning, and graph analytics. This versatility, as well as well-documented APIs for developers working in Java, Scala, Python, and R languages and its familiar DataFrame construct, have fueled Spark’s meteoritic rise in the emerging field of big data analytics. IBM took notice of Spark several years ago, and has since worked on several fronts to help accelerate the maturation of Spark on the one hand, and to embed Spark within its various products on the other, including: ML for z/OS, which executes Watson machine learning functions in a Spark runtime in the mainframe’s Linux-based System z Integrated Information Processor (zIIP). Integrated Analytics System, which combines Spark, Db2 Warehouse, and its Data Science Experience, a Jupyter-based data science “notebook” for data scientists to quickly iterate with Spark scripts. Project DataWorks, which brings Spark and Watson analytics together on the Bluemix cloud. Open Data Analytics for z/OS, a runtime that combines Spark, Python, and the Anaconda package of (mostly) Python-based data science libraries from Anaconda. And Spark running directly on its Bluemix cloud. And considering that IBM opened a Spark Technology Center in 2015, it’s safe to say that IBM is quite bullish on Spark. (That’s a major understatement, actually.) But perhaps the most interesting data point for this discussion came in 2016, when Big Blue launched its z/OS Platform for Apache Spark, which is a native distribution of Spark for the System z mainframe. Native Spark On The Mainframe IBM received kudos for the work from various industry insiders who participated in this video on the z/OS Platform for Apache Spark webpage. Among those singing IBM’s praise was Bryan Smith, the former CTO and VP of R&D at Rocket Software. “IBM did a really good job in porting Apache Spark to z/OS,” Smith says. “They could have just done a very simple port. But they didn’t. They didn’t cut any corners. They really exploited the underlying hardware architecture. They’re using specialty engines. They’re using the hardware compression facilities. They’re able to leverage the 10 TB of memory that you have on a z13 machine and the . . . processors, so you can actually run those Apache Spark clusters on z/OS.” Another software vendor that appreciates having Spark running natively on z/OS is Jack Henry & Associates, the Missouri banking software developer that also has a fairly big IBM i business. “The pain point to us is getting the data out to our customers,” Todd Hill, Jack Henry’s direct of card processing, says in the video. “Currently we have data on the mainframe. We have a distributed stack for across many types of applications. What Apache Spark does for us is to keep your data centralized in the one location. So instead of moving all that data off from multiple platforms into other applications, I can run Apache Spark directly on the mainframe, at low cost, and get it built out, and get the data to the people that need it.” Mike Rohrbaugh, zSystem lead for Accenture, says having Spark on the mainframe helps by automating the generation of intelligence and reducing the complexity. “It’s just so simple to bring the analytics engine back to the data to do intelligent automation,” he says in the video. IBM i Versus The Mainframe So how does this relate to the question in the headline of this story? For starters, let’s compare the similarities and differences between the IBM i and the z/OS mainframe platforms. First, the similarities. Both the IBM i server and the z/OS mainframe are relied upon to run transactional applications that are core to the businesses that use them. Both of them are used to store structured data that’s arguably the most critical data for the businesses that use them. They both store data in the EBCDIC format, and are heralded for best-in-class reliability and security. The also both run proprietary operating systems as well as open OSes like Linux, mostly utilize older languages (RPG and Cobol, respectively), and sport text-based interfaces that use the 5250 and 3270 datastreams, respectively. Now, the differences. Mainframes have their own processor type, while IBM i runs on the more popular Power processor. The mainframe stores data many different data stores (Db2 for z, copy books, etc.), while most IBM i data is stored in Db2 or IFS. Demographically, mainframe customers tend to be the largest companies in the world, whereas IBM i has a bigger installed base among small and midsized business. There’s also a large concentration of mainframes in banking, insurance, and healthcare, whereas IBM i has a stronger foothold in manufacturing, distribution, and retail. IBM i and mainframes are strong transactional systems, and are less known for their analytical prowess. However, data analytics are becoming increasingly important in this day and age, especially as part of a company’s digital transformation strategy. The pundits often say that all companies will need data analytics strategies to effectively compete in the coming decades. That’s probably a bit of an exaggeration, but only for the timing. The question, then, becomes the places where this analytical processing is going to take happen. Today, most mainframe and IBM i shops offload it to another system. It’s fairly common for users of both mainframes and IBM i servers to set up elaborate workflows to move data from the “big iron” transactional systems to dedicated analytical systems, including massively parallel processing (MPP) column-oriented systems like Teradata, Netezza, or Vertica. With the advent of Apache Hadoop clusters running on commodity X86 processors, many companies started experimenting with Hadoop computing, which invariably introduced them to the in-memory Spark framework. IBM wants to keep those analytic workloads on the mainframe if at all possible, which is why it made Spark run natively. This not only keeps costs down for its customers, but it also make the mainframe more “sticky” and lessens the urgency to migrate data and workloads off its biggest cash cow. The question, then, is whether IBM sees similar dynamics at play for the average IBM i user. Mainframe customers, owing to their size and tendency to be in financial services, are early adopters of new technologies, like Spark. They’re arguably closer to the cutting edge than the average IBM i shop, and the dollars at stake for each mainframe client are much larger. It’s safe to say that IBM i members of the Large User Group (LUG) probably are more closely resemble their mainframe brethren, and could benefit from having a powerful, cutting-edge tool like Spark running natively on the IBM i. They’re more apt to have a bigger investment in separate analytical environments, be it a Teradata machine or a Hadoop cluster. They’re also more likely to have some data science Skunk Works project running somewhere in their shop, and are more likely to already be running Spark in Linux, which is where it was originally developed to run. Spark On IBM i While Spark may not be on the radar of the average IBM i shop yet, folks within IBM are starting to ask questions about whether Spark will impact the IBM i installed base, and if it’s going to be important to them, how it ought to be introduced. If the company is planning to support Spark natively on IBM i, the company isn’t saying publically, which is not surprising. What we do know, however, is that IBM executives are at least talking about the prospect of bringing Spark to IBM i in some way, shape, or form. “It’s part of some discussions,” IBM’s product development manager for Db2 Web Query Robert Bestgen recently told IT Jungle. There are two general options for bringing Spark to the platform: porting Spark to run natively on IBM i or running in a Linux partition running on Power Systems. Spark was written in Scala, and therefore can run within a Java virtual machine (JVM), which the IBM i platform obviously runs. It may not be a stretch to get it running there, but there could be other factors that come into play, such as IBM i’s single level storage architecture, and how that maps to how Spark tries to keep everything in RAM (but will spill out to disk if needed). Should the Spark port be native? “Depends on who you talk to,” Bestgen said. The widely held thinking within IBM is that the Linux route makes more practical sense – if Spark is to come to IBM i at all (which, as far as we know, hasn’t been decided). “If you back up [and look at it] from an IBM i perspective, IBM would say that IBM i is part of the Power Systems portfolio, or what we call Cognitive Systems now,” Bestgen says. “For Power Systems, those platforms [like Hadoop and Spark] tend to run best on a… Linux kind of environment. That’s what folks think about it.” Few IBM i shops today are even running Linux partitions. According to HelpSystems‘ 2017 IBM i Marketplace study, fewer than 8 percent of organizations are running Linux next to IBM i on a Power Systems box, while about 9 percent are running Linux on other Power boxes. AIX’s penetration is about 50 percent higher, for what it’s worth. There’s a case to be made that IBM i shops are lousy at figuring out how to leverage the wealth of available tools for Linux, even after IBM went through the trouble of supporting little endian, X86-style Linux to go along with its existing support for big endian Linux within Power. “One of the areas that IBM could do a better job selling is saying, you seem to be willing to run Linux on a different platform. Why not run it on the platform that you have in your system now?” Bestgen says. At the end of the day, there are a lot of unanswered questions, including whether the IBM i installed base needs or wants such a powerful tool as Spark, let alone how it should run. So the question to the answer in the headline is no. “I don’t think we’re there yet in terms of running those things natively on i,” Bestgen says.

https://www.itjungle.com/2017/11/06/spark-memory-run-natively-ibm/

Wednesday, October 25, 2017

Global Hadoop And Big Data Analytics Market 2017: Cloudera, Karmasphere, Greenplum, Hstreaming, Zettaset

QY Market Research include new Hadoop And Big Data Analytics market research report "2016-2022 Report on Global Hadoop And Big Data Analytics Market Competition, Status and Forecast, Market Size by Players, Regions, Type, Application" to its huge collection of research reports. This report on the global Hadoop And Big Data Analytics market is highly helpful because it covers all the aspects which are important in deciding the future of this industry. The Hadoop And Big Data Analytics report is collated by full-fledged analysts who have created use of their market intelligence to explain all basic and crucial knowledge regarding the worldwide Hadoop And Big Data Analytics industry.  Get Free Sample of Report Here: bit.ly/2lclWf8 The Hadoop And Big Data Analytics market report makes use of tables, charts, graphs, maps, and statistics to present the data in the easiest way to understand. The Hadoop And Big Data Analytics market report is a comprehensive analysis of the key factors impacting the global Hadoop And Big Data Analytics industry. This Hadoop And Big Data Analytics market report includes both the driving factors as well as the restraining factors that are influencing the market's performance positively and negatively, respectively. Top Manufacturers Analysis Of global #Hadoop And #BigData #Analytics market @Cloudera Inc. @Hortonworks @Hadapt @AmazonWebServices LLC @Outerthought @MapR Technologies Platform Computing @Karmasphere @Greenplum @Hstreaming LLC @Pentaho Corporation @Zettaset The present trends shaping the worldwide Hadoop And Big Data Analytics market and the way it will influence the Hadoop And Big Data Analytics market within the future are studied. Additionally to this, the future opportunities within the Hadoop And Big Data Analytics market that have the potential to assist the market to expand are given within the Hadoop And Big Data Analytics report. Enquire Here: bit.ly/2z5gTDE The segments inside the global Hadoop And Big Data Analytics industry and their sub-segments also are studied in detail. This ensures that the whole Hadoop And Big Data Analytics market is covered. The leading segment along with the declining segment and also the most promising segment has been given during this Hadoop And Big Data Analytics report. This helps the newcomers, shareholders, readers, stockholders and leading manufacturers to choose which segment or sub-segment to speculate on therefore on acquire most profits. The Hadoop And Big Data Analytics market report discusses the degree of competition, bargaining power of suppliers, bargaining power of buyers, and a threat of substitutes, both internal and external substitute of Hadoop And Big Data Analytics market. The threat of recent entrants or the barriers experienced by new and prime players in getting into the worldwide Hadoop And Big Data Analytics market has been mentioned. This offers new players a plan on whether or not they will exist within the competitive Hadoop And Big Data Analytics industry. About Us: Business Worldwide is a trusted brand in the research industry with a capability of commissioning complex projects within a short span of time with high level of accuracy. At Business Worldwide, we believe in building long-term relations with our clients. Our services cover a broad spectrum of industries including Energy, Chemicals and Materials, Automotive, Software and Aerospace.

https://www.openpr.com/news/784960/Global-Hadoop-And-Big-Data-Analytics-Market-2017-Cloudera-Karmasphere-Greenplum-Hstreaming-Zettaset.html

Thursday, October 19, 2017

Attunity Relishes Role as Connective Tissue for Big Data

In the #bigdata world, there are a few patterns that repeat themselves over and over again. One of those is the need to move data from one place to another. For #dataintegration tool maker #Attunity, that pattern has provided ample opportunities for its changed data capture (CDC) technology. CDC emerged just after the turn of the century, at least five years before @Yahoo started filling the first @Hadoop cluster. Back in those days, enterprises built #datawarehouses and stocked them with the latest data from core business applications running atop @IBM #DB2, @Microsoft #SQLServer, @Sybase (now owned by @SAP), and @Oracle ‘s #eponymousdatabase. At first, bulk extracts were the norm. But as CDC technology filtered into the marketplace, enterprises realized that automatically detecting changes recorded by the database’s log, and then replicating the changed parts – not the whole table –constituted a better way. Over the years, the data ecosystem has evolved in dramatic fashion. While relational data warehouses are still widely used, enterprises have rushed to embrace Apache Hadoop-style storage and processing, whereby raw data is loaded into a distributed file system and processed at run-time. This “schema on read” approach has flipped the script, so to speak, on the highly structured nature of the old data warehousing world. But it didn’t eliminate the need to move data in carefully orchestrated time slices. Apache Sqoop may be widely used to load bulk data into Hadoop, but CDC technology has retained a foothold when it comes to loading data from relational databases, which remain the go-to platforms for operational systems in the Fortune 500. Getting data out of those databases and into analytic systems is a booming business at the moment. “We’re best known for real-time movement. CDC is the heart of what we do,” says Attunity’s vice president of product management and marketing Dan Potter. “We make our money by unlocking the data from those enterprise systems.”  Accessing data from source systems remains a key part of analytic projects (TechnoVectors/Shutterstock) Hadoop remains a popular target for data scraped off operational systems with CDC, Potter says. But cloud data repositories and their associated analytic systems, such as Google Cloud Platform‘s BigQuery and Amazon Web Services‘ Redshift, are quickly becoming popular targets for the data generated from traditional line-of-business apps. “Snowflake is hot,” Potter says, referring to another cloud-hosted MPP-style database. “In the last six months, we’ve seen more data movement to the cloud and we’ve seen a lot more interest in S3 as the repository. We’ll continue to make investments around that.” Real-time streaming analytics is also providing a fertile bed for CDC technology. While one might think that modern data message busses like Apache Kafka and Amazon Kinesis might in some ways compete with 15-year-old CDC technology, the reality is that they’re complementary to each other, says Itamar Ankorion, the chief marketing officer for the Burlington, Vermont company. “What we do with CDC is turn a database into a live stream. It broadcasts from the database whenever something changes,” he tells Datanami at the recent Strata Data Conference. “If you’re adopting streaming architectures, it would only make sense that you’re feeding data into it as a stream.” Potter concurs. ” Kafka is perfect for CDC,” he says. “People are starting to think about real-time data movement. The best way to get real-time data movement is through CDC. So we generate these events in real time and we’re allowing enterprise transaction systems to particulate in real-time streams.” The company recently shipped a new version of its flagship CDC product, Attunity Replicate version 6.0, which brings several new features that open up the product to more big data use cases. That includes performance optimizations for Hadoop; expanded cloud data integration for AWS, Azure, and Snowflake; better integration with Kafka, Kinesis, Azure Event Hub, and MapR Streams. It also gets new streaming metadata integration support for JSON and Avro formats, a new central repository for metadata storage and management, and support for microservices via a new REST API.  Attunity, which trades on the NASDAQ National Market under the ticker symbol ATTU, also unveiled Attunity Compose for Hive, an ETL tool that enables users to better align relational data with the data formats that the Hadoop-based SQL analytics engine expects to see. The update takes advantage of the new ACID merge feature that Hortonworks built into Hive with the latest version 2.6 release of its Hortonworks Data Platform (HDP) offering. Potter says that feature will help give a big data customer more confidence that the data they’re loading from a relational database into Hive is clean and ready for processing. “If I’m going from an Oracle system, I know the schema and structure of that Oracle system. I can move that over and automatically create that same schema and structure in Hive,” Potter says. “There’s no coding at all involved. We pick up any of the changes on the source system, like table changes or other things, and we automatically accommodate that.” Attunity likes its neutral position when it comes to accessing data. While the targets may change, for Attunity the goal is all about getting data out of host systems, including older ones that are harder to access, like IBM mainframes, IBM i servers, NonStop, VAX, and HP3000 systems, not to mention newer ones that nevertheless store data in unique ways. “We’re Switzerland,” Potter says. “We’re independent. It’s easy for us to work with Oracle’s competitors, like SAP.” Lately, the architectural discussions have trended away from using Hadoop and toward using cloud systems. Attunity execs monitor those discussions and aren’t afraid to formulate a position. “I don’t think [Hadoop] is dead. I think there’s lots of evolution,” Potter says. “Hadoop is the best thing that happened to us because those projects are much larger. They’re transformative initiatives.” Hadoop was important for the industry because it spurred a wave of innovation. “When Attunity started doing replications, it was ‘Let’s move from Oracle to SQL Server, point to point.’ It’s not that interesting,” Potter says. “Now it’s about ‘Let’s move a big set of our enterprise data into this data lake and start doing meaningful things with it.’ So there’s a lot more complexity, a lot larger data set, and it’s a lot more strategic to organizations.”

https://www.datanami.com/2017/10/18/attunity-relishes-role-connective-tissue-big-data/

Wednesday, October 18, 2017

PSSC Labs Named a 10 Best Hadoop Solution Provider Companies to Watch

LAKE FOREST, CALIF. (PRWEB) OCTOBER 18, 2017 @PSSCLabs, a developer of custom #HPC and #BigData #computing solutions, today announced it has been named to Insights Success Magazine’s 10 Best @Hadoop Solution Provider Companies to Watch. Providing industry leading server platforms engineered specifically for Hadoop, PSSC Labs is a leader in providing turn-key solutions for enterprise #BigData needs. Hadoop is a constantly evolving and expanding ecosystem that is disrupting the traditional storage and analytics platforms with more flexibility, speed, and reliability – creating an ideal, low cost, and scalable solution for today’s enterprise market. PSSC Labs’ exceptional record of offering the most powerful Hadoop solutions to its customers around the world has earned the company a place on this year’s Best Hadoop Solution Providers list. Editors at Insight Success Magazine analyze companies large and small to find the most unique and disruptive Hadoop solution providers on the market, choosing PSSC Labs as one of its standouts. To view the complete profile visit http://www.insightssuccess.com/pssc-labs-delivering-hand-crafted-hpc-extraordinary-big-data-computing-solutions/. Designed for Hadoop PSSC Labs has deployed over 100 petabytes of Hadoop infrastructure, including small POC clusters to production environments exceeding 10 PB. PSSC Labs’ CloudOOP Big Data Server Line is ideal for organizations looking for Hadoop solutions, and ideal for numerous industries including Ad Tech, Healthcare, Telecom, Media & Entertainment, and Cybersecurity. Designed specifically for Hadoop, Kafka, Big Data and edge computing with IOT devices and sensors, the CloudOOP line is engineered for faster data ingestion and processing. Proven compatible with leading data platforms including Hortonworks, Cloudera and MapR, the line offers enterprises a platform that will significantly lower both CapEx and OpEx costs while achieving higher data throughput performance. Utilizing PSSC Labs’ specialized server design and unique, customized solutions, the CloudOOP line delivers high performance, high density, flexibility and scalability while reducing overall footprint and power usage. Offering 2x the density, 35% reduction in power use and up to 50% increase in data throughput, the CloudOOP line allows companies to scale their infrastructure while reducing the need for additional network, rack and power components – reducing total cost of ownership.
http://www.prweb.com/releases/2017/10/prweb14808617.htm

Monday, October 16, 2017

GLOBAL HADOOP MARKET 2017- MAPR TECHNOLOGIES , PENTAHO, HORTON WORKS AND CLOUDERA .

Global Hadoop Market research report starts with basic introduction of Hadoop market, basic definitions, end-user applications, classifications and industry chain structure. The report also predicts upcoming Hadoop market tendencies by analyzing past market values and current Hadoop market needs. Comparison of past, current and future data is also accomplished in the global Hadoop industry report so as to envision drastic conversion of Hadoop market. The report keeps a complete picture of Hadoop market size and growth in front of our clients and enhances them to make right Hadoop business decisions. It entails bureaucratic outlook of the Hadoop market both regionally and internationally. Clients face major challenges and disputes while analyzing the Hadoop market. Hence, Hadoop market demand and supply analysis mentioned in the Hadoop study would help clients to face those challenges conveniently. Different strategies used to retrieve the relevant and crucial data are also mentioned in the Hadoop research report. The Hadoop report also acknowledges remarkable data including sales margin, business deceits, Hadoop company profiles along with their company information, and scale of demand to supply. It also throws a light on cropping up products of Hadoop industry, details of product price/cost and various market forces. Towards the end, the Hadoop report highlights the critical process analysis carried out by Hadoop experts and professionals. For more details about Global Hadoop Market report enquire here: https://market.biz/report/global-hadoop-market-icrw/42628/#inquiry The global Hadoop market research report is isolated according to key manufacturers, different categories of products & Hadoop applications. According to the research information, the Hadoop market is highly diverse and competing because of a large number of local and global Hadoop vendors. The Hadoop players focusing on the development of new Hadoop technologies and feedstock to strengthen the technological expertise in Hadoop industry. Manufacturers based Segmentation of Global Hadoop Market gives detailed infomation about leading players of @Hadoop which includes @HStreaming LLC, @CiscoSystems, @HortonWorks, @Cloudera ., @Karmasphere ., @EMC – @Greenplum, @Pentaho, @Teradata Corp., @MapR Technologies ., @IBM Corp. and .. Types based Segmentation of Global Hadoop Market categories into SQL Layer, Analytics and Visualization, Searching and Indexing, Machine Learning, Hadoop Application Software and Hadoop Performance Monitoring Software. Application analysis is also included along with type analysis which divides the Hadoop market into Manufacturing, Media and Entertainment, Banking, Financial services and Insurance (BFSI), Telecommunications, Healthcare and Life Sciences and Retail. Regions based Segmentation of Global Hadoop Market mainly focuses on Hadoop specific regions of the world including North America, South America, Europe, Asia-Pacific and the Middle East. Market growth of Hadoop industry is expected to grow with significant CAGR during the Hadoop forecast period from 2017 to 2022. Click here for request a global Hadoop market report: https://market.biz/report/global-hadoop-market-icrw/42628/#requestforsample Major highlights of the Global Hadoop market Report: The Hadoop report chiefly evaluates the in-depth groundwork of the Hadoop market and covers major geographical Hadoop regions. It provides clear tolerant about the Hadoop market along with different opportunities, constraints, Hadoop growth, and practicality. It also displays various Hadoop plans and policies, industrial chain, rules and regulations of the Hadoop industry. At last, the overall global Hadoop market report will assist the new aspirants in making right Hadoop business choices.

http://publicistreport.com/market-research-news/global-hadoop-market-2017.html

Sunday, October 8, 2017

Global Hadoop Market Report 2017: Analysis & Trends 2014-2016 & Industry Forecasts 2017-2025

Dublin, Oct. 06, 2017 (GLOBE NEWSWIRE) -- The "Global Hadoop Market Analysis & Trends - Industry Forecast to 2025" report has been added to Research and Markets' offering. The Global Hadoop Market is poised to grow at a CAGR of around 42.7% over the next decade to reach approximately $54.2 billion by 2025. This industry report analyzes the market estimates and forecasts of all the given segments on global as well as regional levels presented in the research scope. The study provides historical market data for 2014, 2015 revenue estimations are presented for 2016 and forecasts from 2017 till 2025. The study focuses on market trends, leading players, supply chain trends, technological innovations, key developments, and future strategies for the existing players, new entrants and the future investors. Some of the prominent trends that the market is witnessing include rising demand for data analytics and increasing number of partnerships and fundings taking place in hadoop Market.

10 Leading @Hadoop Companies 

@Cloudera, Inc
@AmazonWebServices
@Datameer,
@Mapr
@TeradataCorporation
@MarkLogic
@Karmasphere,Inc
@Pentaho
@Hortonworks
@Cisco Systems, Inc
@IBM Corporation
@Oracle Corporation, Inc
@OpenX
@Vmware
@Pivotal

http://www.globenewswire.com/news-release/2017/10/06/1142270/0/en/Global-Hadoop-Market-Report-2017-Analysis-Trends-2014-2016-Industry-Forecasts-2017-2025.html

Running Hadoop on a Raspberry Pi 2 cluster

I've been involved with cluster computing ever since #DEC introduced #VAXcluster in 1984. In those days, a three node VAXcluster cost about $1 million. Today you can build a much more powerful cluster for under $1,000, including much more storage than anyone could afford back then. @Hadoop is the open-source version of @Google 's #MapReduce and #GoogleFileSystem ( #GFS ), widely used for large data-crunching applications. It is a shared-nothing cluster, which means that as you add cluster nodes, performance scales up smoothly.

In the paper, Performance of a Low Cost Hadoop Cluster for Image Analysis, researchers Basit Qureshia, Yasir Javeda, Anis Kouba, Mohamed-Foued Sritic, and Maram Alajlan, built a 20 node RPi Model 2 cluster, brought up Hadoop on it, and used it for surveillance drone image analysis. They also benchmarked the RPi cluster against a 4-node PC cluster based on 3GHz Intel i7 CPUs, each with 4GB of RAM.

CONFIGURATION

The 20 node cluster was divided into four, 5-node subnets, each attached to 16 port switches that are, in turn, networked to a managed 24 port core switch. The extra switch ports enable easy cluster expansion.

Each 700MHz RPi B runs Raspbian, an ARM-optimized version of Debian Linux. Each RPi has a Class 10, 16 GB SD card capable of up to 80MB/s read/write speeds. An image of the OS with Hadoop 2.6.2 was copied onto the SD cards. The Hadoop Master node, which implements the name-node only, was installed on a PC running Ubuntu 14.4 and Hadoop.

In the paper, Performance of a Low Cost Hadoop Cluster for Image Analysis, researchers Basit Qureshia, Yasir Javeda, Anis Kouba, Mohamed-Foued Sritic, and Maram Alajlan, built a 20 node RPi Model 2 cluster, brought up Hadoop on it, and used it for surveillance drone image analysis. They also benchmarked the RPi cluster against a 4-node PC cluster based on 3GHz Intel i7 CPUs, each with 4GB of RAM.

http://www.zdnet.com/article/running-hadoop-on-a-raspberry-pi-2-cluster/

Thursday, September 21, 2017

Syncsort quality manager aims to purify Hadoop data lakes

#Syncsort Inc. is extending the data quality features of the #TrilliumSoftware Inc. subsidiary it acquired last November to native #Hadoop environments with #TrilliumQuality for #BigData. The offering combines Trillium’s data quality features with its Intelligent Execution data integration platform to enable information technology organizations to normalize and integrate data at the same time. The Trillium platform was previously available in native format only on #Linux, #Unix and Windows operating systems. The Hadoop support is the first time Syncsort has applied its data quality features to applications. Data quality is about identifying inconsistencies, errors or duplication. Examples include a ZIP code entered in a date field or duplicate customer records that appear to be different because of misspellings. Normalizing data is a tricky process. For example, different countries have different address and date formats and two people with the same name in the same ZIP Code may or may not be the same person. Users are rushing to extract data from production systems and load it into analytics engines, but are discovering that quality problems limit their effectiveness. “Everybody is trying to govern the data once it’s in the data lake so it doesn’t turn into a data swamp,” said Tendü Yoğurtçu, Syncsort’s chief technology officer. “The volume and variety of data makes it complex.” Trillium has hundreds of matching algorithms to identify such problems, and can be configured to automatically apply corrective algorithms, Yoğurtçu said. The offering includes address- and name-matching data for 150 countries as well as postal directories and geocoding. Intelligent Execution examines the topology of a data flow and optimizes resources for the job without changes to the application. It supports both new and existing Trillium data quality projects across Hadoop, MapReduce and Apache Spark on-premises or in the cloud. “Once you understand the data you can create the rules to cleanse that data,” Yoğurtçu said. “For example, if you have duplicates you can specify a process to flag them or get rid of them.” Trillium Quality for Big Data is available on all Hadoop distributions including Cloudera Inc.’s CDH, Hortonworks Inc.’s HDP and MapR Technologies Inc.’s Converged Data Platform. It deploys and installs via Cloudera Manager and Apache Ambari. Pricing is on a per-node basis or cloud subscription, but Sync

https://siliconangle.com/blog/2017/09/20/syncsort-quality-manager-aims-purify-hadoop-data-lakes/

Tuesday, September 5, 2017

Hadoop distributor MapR raises $56M following major sales jump

A year after closing its last funding round, MapR Technologies Inc. has secured another cash infusion to establish a stronger position in the global analytics market. The company, which sells a commercial version of the popular Hadoop data crunching framework, today announced the completion of a $56 million investment. Lightspeed Venture Partners led the round with participation from several other returning backers. The funding follows a quarter in which MapR claims to have doubled new subscription billings year-over-year. This growth puts the company on track to close 2017 with total sales up 70 percent on an annualized basis, according to a press release. #MapR generates most of its revenue from its #ConvergedDataPlatform, which augments the core components of #Hadoop with other open-source tools and proprietary enhancements. MapR recently bolstered its portfolio with a homegrown data store designed to ease large analytics projects. Dubbed #MapRXD, the system enables companies to aggregate records from disparate systems into a unified environment so that they can be processed centrally. MapR mainly sells its software to large organizations. Notable clients include American Express Co., Cisco Systems Inc. and Audi AG. The company’s software also powers Aadhaar, the database that India has built to store the biometric identification information of its more than 1 billion residents. The new funding will enable MapR to further expand its global presence. Chief Executive Officer Matt Mills told Forbes that the company’s expansion plans places a particular focus on Australia, Japan and South Korea. Mills also shared some backroom information about the investment. He said the funding came at a higher valuation than the company’s previous raise last year, which was in turn a down round according to Forbes. It’s a sign that MapR’s investors are optimistic even as competing Hadoop distributors expand their own turf. Hortonworks Inc. last month reported a 42 percent revenue increase for the second quarter, while Cloudera Inc. recorded a 41 percent sale jump during the same period. However, both companies are trading well below their initial public offering prices because of their large losses. MapR isn’t profitable yet either, according to CEO Mills, but he said it’s hoping to become cash-neutral by the end of 2018. The company will likely wait until then before going public given its competitors’ somewhat lackluster stock market performance. Mills was quoted as saying “we don’t have a timetable, we want to do it when things are right” regarding a potential MapR IPO.

https://siliconangle.com/blog/2017/09/05/hadoop-distributor-mapr-raises-56m-following-major-sales-jump/

Sunday, August 27, 2017

Global Hadoop Market is expected to grow at a CAGR of +52% over the forecast period 2015 to 2022

This report gives an in-depth research about the overall state of Hadoop Systems market and projects an overview of its growth market. It also gives the crucial elements of the market and across major global regions in detail. Number on primary and secondary research has been carried out in order to collect required data for completing this particular report. Sever industry based analytical techniques has been narrowed down for a better understanding of this market. Hadoop is a data storage processing system that enables data storage, file sharing, data analytics etc. The technology is scalable & enables effective analysis from large unstructured data therefore adding value. With increasing role of social media and internet communication Hadoop is being largely used by various spectrum of companies ranging from Facebook to Yahoo. Other big users of #Hadoop include #Cloudera, #Hortonworks, #IBM, #Amazon, #Intel, #Mapr, #Microsoft. This technology facilitates its users to handle more data through enhanced storage capacity also enables data retrieval in case of hardware failure.

Some of the key players in the global hadoop market include: #IBM, #DellTechnologies, #CiscoSystems, Inc., #HewlettPackard, #Zettaset, Inc., #AmazonWebServices, #Datastax, Inc., #MapRTechnologies, Inc., #FujitsuLtd., #HitachiDataSystems, #Datameer, Inc., #TeradataCorporation, #Rainstor, #Cloudera, Inc. and #Hortonworks Inc.

Hadoop as a solution is increasingly offering data retrieval and data security features. These features are getting better with time. This is leading to enhanced solution to database management systems (DBMS). Hadoop software is the highest growing market in comparison to hardware and services. In terms of geographical segmentation, North America expected to lead the revenues in this market due to higher rate of technology adoption.

Service segment is valued to account for the largest share across the global market and software segment is expected to register the highest growth due to increasing penetration of Hadoop software in various industries such as BFSI, government and telecommunication. Hadoop application software has the dominant share in global market and is expected to continue its dominance throughout forecast period owing to increasing operation by developers to make real time applications.

http://www.military-technologies.net/2017/08/25/global-hadoop-market-is-expected-to-grow-at-a-cagr-of-52-over-the-forecast-period-2015-to-2022/

Sunday, August 20, 2017

Big Data Analytics & Hadoop Market is expected to reach at a CAGR of +41% from 2017 to 2021

The global #Hadoop #BigData Analytics market is explained in detail in this report, starting with a basic overview, which includes definitions and various specifics related to the raw materials used in manufacturing Hadoop Big Data Analytics products. It includes a categorized distinction of major and minor factors that influence this global industry. The overview also includes a description of the value chain structure of the global industry and a status update for the different major regional segments of this industry.

Big Data Analytics & Hadoop Market is expected to reach at a CAGR of +41% from 2017 to 2021. Rise of big data & growing need for big data analytics and rapid growth in consumer data are some of the factors fueling the market growth. Lack of skilled workers and Lack of security features in the Hadoop framework are restraining the market growth. Venture capital funding is the major opportunity for vendors in big data analytics and hadoop market.

Some of the key players in the  The global #Hadoop #BigData #Analytics market  include: #DellTechnologies  #Karmasphere Inc., #Talend, Inc., #DataDirectNetworks, Inc., #AmazonWebService LLC, #HORTONWORKS , INC., #Appistry, Inc., #NetApp, Inc., #TeradataCorporation, #ClouderaInc., The #HewlettPackard Company, #Greenplum, Inc., #Datameer, Inc., #Zettaset, Inc., #Fujitsu Ltd., #PentahoCorporation, #DataStax, Inc., #PlatformComputing, #HStreamingLLC , #MapRTechnologies, Inc., #IBM and #HadaptInc

http://www.satprnews.com/2017/08/19/big-data-analytics-hadoop-market-is-expected-to-reach-at-a-cagr-of-41-from-2017-to-2021/