Dell, EMC, Dell Technologies, Cisco,

Showing posts with label Bigdata. Show all posts
Showing posts with label Bigdata. Show all posts

Thursday, September 21, 2017

Syncsort quality manager aims to purify Hadoop data lakes

#Syncsort Inc. is extending the data quality features of the #TrilliumSoftware Inc. subsidiary it acquired last November to native #Hadoop environments with #TrilliumQuality for #BigData. The offering combines Trillium’s data quality features with its Intelligent Execution data integration platform to enable information technology organizations to normalize and integrate data at the same time. The Trillium platform was previously available in native format only on #Linux, #Unix and Windows operating systems. The Hadoop support is the first time Syncsort has applied its data quality features to applications. Data quality is about identifying inconsistencies, errors or duplication. Examples include a ZIP code entered in a date field or duplicate customer records that appear to be different because of misspellings. Normalizing data is a tricky process. For example, different countries have different address and date formats and two people with the same name in the same ZIP Code may or may not be the same person. Users are rushing to extract data from production systems and load it into analytics engines, but are discovering that quality problems limit their effectiveness. “Everybody is trying to govern the data once it’s in the data lake so it doesn’t turn into a data swamp,” said Tendü Yoğurtçu, Syncsort’s chief technology officer. “The volume and variety of data makes it complex.” Trillium has hundreds of matching algorithms to identify such problems, and can be configured to automatically apply corrective algorithms, Yoğurtçu said. The offering includes address- and name-matching data for 150 countries as well as postal directories and geocoding. Intelligent Execution examines the topology of a data flow and optimizes resources for the job without changes to the application. It supports both new and existing Trillium data quality projects across Hadoop, MapReduce and Apache Spark on-premises or in the cloud. “Once you understand the data you can create the rules to cleanse that data,” Yoğurtçu said. “For example, if you have duplicates you can specify a process to flag them or get rid of them.” Trillium Quality for Big Data is available on all Hadoop distributions including Cloudera Inc.’s CDH, Hortonworks Inc.’s HDP and MapR Technologies Inc.’s Converged Data Platform. It deploys and installs via Cloudera Manager and Apache Ambari. Pricing is on a per-node basis or cloud subscription, but Sync

https://siliconangle.com/blog/2017/09/20/syncsort-quality-manager-aims-purify-hadoop-data-lakes/

Wednesday, December 14, 2016

Syncsort's Third Annual Hadoop Survey Uncovers Big Iron to Big Data Trends to Watch in 2017

PEARL RIVER, N.Y., Dec. 13, 2016 /PRNewswire/ -- Syncsort, a global leader in #BigIron to #BigData solutions, today announced the results from its third annual #Hadoop survey, showing that as users gain more experience with Hadoop, they are building on their early success and expanding the size and scope of Hadoop projects. Download the Hadoop survey results for 2017. The number one role of Hadoop continues to be increasing data warehouse capacity and reducing costs (cited by 62% of respondents) but the number of organizations using it for better, faster analytics is on the rise (49.4% of respondents, up from 45.5% last year). As the ways in which they are using Hadoop and the benefits they expect to achieve are evolving, there are several key trends to watch in 2017. Traditional data sources are more popular for filling the data lake, but newer sources are a sizable part of the mix. As Hadoop projects multiply, so too, will the sources and volume of data required to support them – from mainframes to NoSQL databases and beyond. Legacy systems remain essential for populating the data lake, with the top three sources including: enterprise data warehouses (EDW) (69%), relational database management systems (RDBMs) (61%) and mainframes (40%). Newer sources are gaining importance, especially those that generate streaming data, such as smart devices and sensors (53.5%). Therefore, having a tool that allows organizations to easily access and integrate all data – legacy and newer sources, batch and streaming – is vital. Big Data insights will become a mainstay in running all business operations. Respondents expect significant benefits from Hadoop, with increased business/IT agility (66.9%) and greater operational efficiency and reduced costs (66.5%) in a virtual tie at the top of the list. When looking at how respondents are turning data into business insights, the majority are interested in using Hadoop for advanced/predictive analytics (62%), followed by data discovery and visualization (57%) and ETL (53%). As Big Data projects mature, businesses will become more confident in and dependent on their results. Big data won't just be fodder for data scientists and analysts, but will feed the KPIs and other important metrics needed to run the business. Data governance grows in importance – mainframe data continues to be key. Over two-thirds of respondents from organizations with mainframes find accessing and integrating these critical data assets into Hadoop for real-time analytics valuable or very valuable (70%). Highly-regulated industries such as financial services, banking, insurance and healthcare have the challenge of making the critical legacy data assets available for real-time analytics while also maintaining cross-platform data lineage for compliance purposes. As Hadoop implementations spread across organizations, data governance, and the data quality needed to support it, will become more critical to meet the regulatory and compliance mandates. There will be a move toward simplicity. For the third year in a row, the need to learn a new set of skills and tools and finding affordable Hadoop programmers is among the top challenges (58%), followed by difficulty keeping up with evolving compute frameworks (53%) and integration with existing data sources and applications (48%). As more technology designed to support specific use-cases comes online, tools that remove the complexity and streamline processes will be in high demand. Organizations will increasingly implement strategies and tools that easily adapt to take advantage of a rapidly evolving technology stack. As the Hadoop market matures, MapReduce remains the more popular framework (62%) and half of respondents use at least one other compute framework along with MapReduce (51%). Spark adoption is on the rise – 55% currently use Spark, but that number is expected to grow to 76% in the future. 47% of those currently using MapReduce plan to move away from the framework. In addition, a third of respondents deploy Hadoop both on premise and in the cloud. To achieve the efficiencies enterprises seek with Hadoop, users will demand tools that handle different environments with ease, without burdening their teams with continuous application maintenance or requiring complex coding skills. "Insights from Big Data can offer new revenue streams, new customer insights, improved decision-making, better quality products, enhanced customer experiences, and more, but organizations are still struggling to have easy access to all their enterprise data," said Tendü Yoğurtçu, General Manager of Syncsort's Big Data business. "From our survey results, it's clear businesses are realizing the importance of tapping into a full range of data sources, and the benefits of using products and solutions that can simplify and bridge the gap between legacy and emerging platforms, Big Iron to Big Data solutions to streamline and accelerate the ROI on their Big Data initiatives." For more information on the Hadoop survey: http://bit.ly/2hgvgLz. Methodology: Syncsort polled over 250 respondents including data architects, IT managers, developers, business intelligence/data analysts and data scientists, with 86% coming from organizations with revenues of more than $100 million. Participants represent a range of industries including financial services/insurance, healthcare, government, telecommunications, retail and more.

http://www.prnewswire.com/news-releases/syncsorts-third-annual-hadoop-survey-uncovers-big-iron-to-big-data-trends-to-watch-in-2017-300377138.html

Monday, November 7, 2016

Cloudera launches new office in Dubai; expands big data management services to the Middle East

DUBAI, UAE – November 3, 2016 #Cloudera, the global provider of the fastest, easiest, and most secure data management and analytics platform built on #Apache #Hadoop and the latest open source technologies, today [date] announced its formal entry to the Middle Eastern market. The data management technology provider made the announcement at its very first Cloudera Sessions in Dubai – held at Sofitel the Palm. The inaugural session gathered innovators, data professionals and IT decision-makers to explore Cloudera technologies and creative ways of harnessing big data for competitive advantage. “Companies are realizing that understanding, analysing and utilising big data isn’t a passing fad. It’s a matter of survival. There’s been an absolute explosion in the rate of data creation. 90% of all data in existence has been created in the past 2 years, and only 0.5% of it has currently been analysed. For businesses, big data is a very exciting opportunity to gain customer insight to drive new revenue streams,” says Cloudera co-founder and chief technology officer, Dr. Amr Awadallah. Cloudera, which helps enterprises around the world solve their most challenging business problems with data, is planning to expand its global 2,500 strong partner ecosystem to include companies in the UAE and region-wide. “We are very excited about serving enterprise customers in the Middle East with the big data management technologies that enable incisive analysis and better decision-making. The UAE has always been the pace setter for innovation region-wide, and we are delighted to be part of a market that is prioritising innovation and beginning to recognise the power of big data,” says Cloudera vice president of EMEA David Pieterse. Big data capabilities are crucial to Dubai’s transformation towards a smart city, with analytical data management platforms such as Cloudera’s helping turn vast amounts of data into insights and intelligent real-time management.

http://www.zawya.com/mena/en/story/Cloudera_launches_new_office_in_Dubai_expands_big_data_management_services_to_the_Middle_East-ZAWYA20161103112129/

Tuesday, November 1, 2016

10 Popular Big Data Software and Tools We Found for you.

Data processing is a do-or-die requirement for businesses today. There are many #BigData software out there. Many of them promise to save you money, time and help you find never-before-seen insights. Although that may be true, navigating the internet for the possible best software can be difficult when there are numerous choices to select from. The following are some of the most popular Big Data software we found for you: #Apache #Hadoop Originally developed by Mike Cafarella and Doug Cutting in 2006, Apache Hadoop is an open-source software framework. It is made to handle numerous data sets. It consists of two main parts: #MapReduce and Hadoop Distributed File System ( #HDFS ). HDFS is a storage part while MapReduce is a processing part. The software is scalable, cost effective, flexible, fast and resilient to failure. Apache Spark Apache #Spark is fast developing as a data processing software built around ease of use, sophisticated analytics and speed. Built in 2009, Spark has over 250 developers who are contributing to it already. Frequently, this software is used as an alternative to Hadoop because of its ability to analyse data faster for some applications. Apache #Storm Apache Storm is used to processes various streaming data including logs or social data. It is a real-time computation that enables users to process limitless streams of data. You can use it for any programming language. Apache Storm advantages include fault tolerance, ease of use and scalability. Apache #Flink Apache Flink is known for processing data fast with high fault tolerance and low data latency. The software processes streaming data in real time. Apache Flink developed quickly as big data processing software and within months, it had captured the attention of a wider audience. Attivio #Attivio software empowers users to get the right data, work with it and quickly use it to make decisions. Many of the world’s top brands rely on this software to gain visibility into their information. #Splunk Splunk software helps users to search, monitor, and process machine-generated data, through web-style interface. The software collects, indexes, and correlates data in a searchable source from which it can create visualizations, dashboards, alerts, reports and graphs. #Samza Samza is a popular data processing software. It is built on #YARN for cluster resource management and Apache Kafka for messaging. Companies that use Samza include LinkedIn, Intuit, MobileAware, Fortscale and Project Florida. DataPlay DataPlay is a cloud based software suite that automates data processing and reporting. It significantly cuts time spent on processing and presentation of data. #DataPlay has a lot of success stories with clients. #SAP #HANA SAP HANA provides users with the ability to store, search, process, explore, and analyse interlinked and textual data with relationships. SAP HANA benefits and capabilities include advanced analytics processing, openness, data access and database services. #Bottlenose Bottlenose makes data processing easy. The software analyses business data to detect trends for brands. It detects trends in huge amounts of fast changing data. The software is used by many successful enterprises to make decisions. For more information about any listed big data software, please visit their respective websites. If you know any software that should be on our list, send an email to us or make a comment below.

http://www.techbullion.com/10-popular-big-data-software-recommended-for-you/

Sunday, October 30, 2016

Containers and microservices find home in Hadoop ecosystem

Much of the recent #bigdata experience has been a bare-metal affair, meaning #Hadoop has happened largely on non-virtualized servers. That could change as containers and microservices gain traction in application development circles.

Both containers and microservices break up monolithic application code into more finely grained pieces.  That streamlines development and makes for easier testing, which is one of the keys to more flexible application deployment and code reuse.

It is early on for such techniques to be applied to big data, but, for new jobs like data streaming, microservices shows promise.  For a technology manager at a leading European e-commerce company, the microservices approach simplifies development and enables code reuse.

With microservices, "you can very much economize on what you're doing," according to Rupert Steffner, chief platform architect for business intelligence systems at Otto GmbH, a multichannel retailer based in Hamburg, Germany. He goes further: For some types of applications, not using microservices "is stupid. You're building the same functionality over and over again."

The types of applications Steffner is talking about are multiple artificial intelligence (AI) bots that run various real-time analytics jobs on the company's online retail site. Otto uses a combination of microservices, #Docker containers and stream processing technologies to power these AI bots.
http://searchdatamanagement.techtarget.com/news/450401973/Containers-and-microservices-find-home-in-Hadoop-ecosystem

Monday, October 10, 2016

Why Java in Big Data? What about Scala?

Why #Java ? Why not? What about #Scala ? Or #Python. I use all three for various parts of #BigData projects. Use the best tool for the job. A lot of things can be orchestrated and managed without any coding through #Apache NiFi 1.0. Some things like #TensorFlow are best done in Python, while #Spark and #Flink jobs could be Scala, Python, or Java. #ApacheBeam is Java only (Spotify added a Scala interface, but it's not official yet. If you are a really strong Java 8 developer and code clean, you can write #Hadoop #MapReduce, #Kafka, Spark, Flink, #Apex. Apache NiFi is written in Java and so is most of Hadoop, so it's Big Data scale. Spark and others are written mostly in Scala. Ecosystem Scala and Java share a ton of libraries, as they run on the JVM. Python has its own huge ecosystem, but for many Hadoop things the JVM languages have a bit of an advantage. You can run JPython on the JVM, but I really haven't seen that used for Big Data, Spark, or Machine Learning. I am wondering if anyone is doing this? Please comment here. Python has TensorFlow and some nice Deep Learning and Machine Learning libraries. They are also starting to get more Universities teaching Python instead of Java. Not too many Universities are teaching Scala.

https://dzone.com/articles/why-java-in-big-data

Monday, October 3, 2016

Big data technology market speeding up – with NoSQL and Hadoop at forefront

The #bigdata technology landscape is varied and vast – and a new report from Forrester examines the key strands, with non-relational databases taking the bulk of the honours. The report, claimed to be a first of its kind from the analyst firm titled Big Data Management Solutions Forecast 2016 to 2021, argues #NoSQL and #Hadoop will see the biggest growth during the five year period, with the markets growing 25.0% and 32.9% per year respectively. The analysts also claim that the big data technology space will grow at three times the rate of the overall technology market. Forrester defines big data technology in six buckets; enterprise data warehousing, NoSQL, Hdaoop, big data integration, data virtualisation, and in-memory data fabric. The latter, a product which usually offers data, compute and service grids as well as an in-memory database, is predicted to grow at 29.2% annually over the coming five years. Usage varies by industry, but the report notes that while telecommunications, professional services, finance and government are the largest users of these technologies now, the pharmaceutical, transport and primary production industries will be the quickest growing. Almost 40% of firms polled by Forrester say they are implementing and expanding their big data technology adoption, with another 30% planning to up their usage in the next 12 months.
http://www.forbes.com/sites/bernardmarr/2016/10/03/big-data-news-the-top-insights-from-strata-hadoop-world-2016-in-new-york/#504e5878620e

Big Data News: The Top Insights From Strata + Hadoop World 2016 In New York

Just about every player in the #BigData and analytics game was in New York last week at the #Strata + #Hadoop World conference, to showcase their latest technologies. Over 7,000 people attended the event where keynote speakers, including White House chief data scientist DJ Patil, laid out their visions for where machine learning, analytics, the Internet of Things, autonomous vehicles and smart cities will be taking us in the near future.

http://www.forbes.com/sites/bernardmarr/2016/10/03/big-data-news-the-top-insights-from-strata-hadoop-world-2016-in-new-york/#504e5878620e

Sunday, October 2, 2016

Maturing of Hadoop in the Enterprise

As the #bigdata conference season winds down for 2016, I believe that one of the key trends we can see prevalent across the board is that the #Apache #Hadoop platform has made long strides in maturing in the enterprise. Early on, Hadoop was adopted principally by those organizations that had “bleeding edge” use cases like sentiment analysis or needed predictive/prescriptive analytics on a massive scale. Today’s landscape looks notably different, with the average Hadoop cluster more likely to be augmenting an enterprise data warehouse than to be locked in a lab.
http://www.cio.com/article/3126457/data-center/maturing-of-hadoop-in-the-enterprise.html

Sunday, September 25, 2016

AWS Redshift Feels the Heat

#Amazon Web Service’s ( #AWS ) #Redshift data warehouse service is taking a beating this week, with database analytics competitors #Oracle Corp. and #Cloudera claiming their platforms run much faster for less money ( #Bigdata ) back up its claims, Cloudera released benchmark results said to show advanced capabilities for cloud-native workloads running on its analytics platform as well as improved price performance compared to Amazon Redshift. Cloudera’s analytics database platform also runs on Apache Impala, an open source incubator project designed to develop a distributed SQL query engine for Hadoop. In benchmark results released on Thursday (Sept. 22), the Palo Alto, Calif., company reported that Impala was 22 percent cheaper than Redshift when querying data stored in AWS (NASDAQ: AMZN) Simple Cloud Storage Service (S3) using cloud native tools. Impala cost 60 percent less and ran 40 percent faster than Redshift over local EBS storage. It also registered major cost savings and faster time-to-result, Cloudera asserted. The added horsepower attributed to Impala comes as more business intelligence and analytic workloads move to the cloud. Cloudera argues that the benchmark testing shows that these workloads can tap into the flexibility and cost savings of public cloud services without sacrificing the performance of on-premise analytical databases. “This comparison [with Amazon Redshift] is clear evidence that Impala is unmatched for these BI and analytic workloads in the cloud,” claimed Charles Zedlewski, Cloudera’s vice president for products. Impala is touted as decoupling data and computing to provide comparable performance for SQL analytics whether running cloud-natively over data in S3 or across other on-premise and cloud storage options. The open source tools works natively with data stored on a range of storage engines, including Amazon S3 object store. That, Cloudera said, eliminates the need to move or load data specifically into Impala clusters. “Especially for cloud deployments, this translates to cost-savings and efficiencies as transient clusters can be spun up as needed for BI and reporting workloads and, with cost-effective storage from S3, more data is quickly and readily available for analysis,” the company added in releasing in the benchmark results. Earlier this week, cofounder and CTO Larry Ellison claimed its Oracle Cloud was more than 100 times faster for database analytics than Amazon Redshift. “Why is Amazon Redshift so slow? Because it’s 20 years behind Oracle,” (NYSE: ORCL) Ellison boasted during a company event this week in San Francisco. The broadside was part of a larger Oracle campaign aimed at challenging public cloud leader AWS with an expanded analytics package that includes applications, infrastructure and database analytics. Skeptics called Ellison’s challenge to the the public cloud leader “ridiculous.” Added Justin Moore, CEO of data security specialist Axcient: “At the end of the day Oracle will likely be a niche player for certain database and application workloads” running in the cloud.
https://www.datanami.com/2016/09/22/aws-redshift-feels-heat/

Cloudera’s Analytic Database Enables Unrivaled Elastic Scale, Agility, and Performance for BI and Analytics in the Cloud

#Bigdata #Cloudera, the global provider of the fastest, easiest, and most secure data management and analytics platform built on #Apache #Hadoop and the latest open source technologies, today released benchmark results that validate Cloudera’s modern analytic database solution, powered by #ApacheImpala (incubating), not only delivers unprecedented capabilities for cloud-native workloads but does so at better cost performance compared to alternatives. Impala uniquely offers elastic scalability, better flexibility, and direct #Amazon S3 query ability unavailable from traditionally architected systems such as Redshift. With a modern design, Impala decouples data and compute to provide the same high-performance SQL analytics whether cloud-natively over data in S3 or across a wide range of on-premise and cloud storage options. Furthermore, Impala enables all these capabilities while also delivering up to 275% more cost-efficiency and up to 10x greater performance compared to Amazon’s analytic database Redshift, equating to more value all within an open platform. Using queries from the TPC-DS industry standard benchmark, Cloudera compared Impala running on the cloud (both cloud-natively over S3 and over local EBS storage) to Amazon Redshift (only able to run over its own storage on dedicated AWS instances). Results from the benchmark show: Impala is over 200% less costly and over 10x faster on S3 compared to a general purpose tuned Redshift Impala is still 8% less costly and 90% faster on S3 compared to a pre-tuned Redshift for specific fixed reporting queries Impala is 28-275% less costly and 42-400% faster on EBS compared to either pre-tuned or general purpose tuned Redshift “Increasingly our customers are looking to move BI and analytic workloads to cloud environments to tap into the cost-effectiveness of elastic scale and greater flexibility. But they still require the high-performance analytics and big data agility they’re used to on-premises,” said Charles Zedlewski, Vice President, Products, at Cloudera. “Impala brings all its advantages it has over traditional, on-premise analytic databases to the cloud with a modern architecture that enables unprecedented agility no matter where the data lives. This comparison is clear evidence that Impala is unmatched for these BI and analytic workloads in the cloud.” As businesses look to bring in more data from new sources, actively adjust models based on changing needs, and iteratively design for a variety of use cases, they need a modern analytic database that is built to address these requirements, without hindering business productivity. The rigid design and inelastic scale of traditionally architected, monolithic systems, whether on-premise or in the cloud, simply are not able to keep up with today’s ever-changing business needs. Cloudera’s analytic database, powered by Impala as the interactive SQL engine, is purpose-built to bring high-performance SQL analytics to big data, with elastic scalability for cloud and on-premise deployments, as and when it is needed. Impala works natively with data stored on a number of storage engines, including Amazon S3 object store, eliminating the need to move or load data specifically into Impala clusters. Especially for cloud deployments, this translates to cost-savings and efficiencies as transient clusters can be spun up as needed for BI and reporting workloads and, with cost-effective storage from S3, more data is quickly and readily available for analysis. Advancing Impala’s performance, concurrency, and scalability is a consistent area of focus for Cloudera. The company has widened the performance gap between Impala’s analytic database architecture and other alternatives for both single and multi-user workloads. The latest release delivers 12x better performance on secure workloads compared to its two prior versions. Cloudera plans to continue expanding Impala’s value and price performance benefits by adding support in the future for other object stores in the public cloud.

http://globenewswire.com/news-release/2016/09/22/873814/0/en/Cloudera-s-Analytic-Database-Enables-Unrivaled-Elastic-Scale-Agility-and-Performance-for-BI-and-Analytics-in-the-Cloud.html

Wednesday, September 21, 2016

Automate and Accelerate Big Data Initiatives with BMC at Strata + Hadoop World New York

HOUSTON, Sept. 21, 2016 /PRNewswire/ -- BMC is showcasing solutions to optimize and automate #BigData initiatives for #Hadoop environments at O'Reilly Strata + Hadoop World, September 27-29 at the Javits Center in New York City. As enterprises fast track Big Data initiatives to support their digital business strategies, BMC is enabling its customers to quickly gain a competitive advantage from new business insights secured from Hadoop data. WHO: BMC, the global leader in IT solutions for the digital enterprise. WHAT: BMC's solutions for Hadoop deliver consistent business value from Big Data initiatives by driving to production faster, integrating applications seamlessly, and continuously improving services. The company's booth at #Strata + Hadoop World New York will have experts on hand to demonstrate BMC's Control-M for Hadoop and Control-M Workload Change Manager solutions for workload automation for Hadoop jobs and workflows. Attendees can learn more about BMC's TrueSight Intelligence solution for analyzing high volumes of operational data, and go hands-on with BMC's TrueSight Capacity Optimization solution to plan, manage, and optimize the use of Hadoop infrastructure. For further information on agile application adoption, BMC's Joe Goldberg, principal solutions marketing manager, BMC will host a session titled, "Accelerate EDW modernization with the Hadoop ecosystem." The session will examine how achieving the best results in Big Data initiatives often means combining new and traditional data sources and modernizing ETL and data warehouse applications, which can experience exponential growth in complexity when moving toward enterprise-grade implementations. Mr. Goldberg will explain why freeware isn't 'free' when it comes to managing Hadoop workflows for enterprise implementations, and provide examples how companies like GoPro, Produban, Navistar, and others have taken a platform approach to managing their workflows, and how they are achieving success in their data warehouse modernization projects.

http://www.prnewswire.com/news-releases/media-alert-automate-and-accelerate-big-data-initiatives-with-bmc-at-strata--hadoop-world-new-york-300331520.html

Kyvos Insights to Demonstrate OLAP on Hadoop Solution at 2016 Strata + Hadoop World

LOS GATOS, Calif., Sept. 20, 2016 /PRNewswire/ -- #KyvosInsights, a #bigdata analytics company, today announced that it will be exhibiting and demonstrating the capabilities of its #OLAP on #Hadoop solution at booth number 663 at Strata + Hadoop World in New York City, Sept. 26-29.

Launched in June 2016, Kyvos 2.0 is a massively scalable, self-service BI on Hadoop analytics solution designed to make big data lakes ready for BI analysts. The patent-pending OLAP on Hadoop solution includes enterprise-grade functionality to improve analysts access to the data lake, advanced security support, improved performance and scalability, and integration with additional BI tools. Kyvos allows companies to build a BI Consumption Layer directly on Hadoop, which lets analysts transform their existing BI tools, giving them instant, interactive access to multi-dimensional analytics at big data scale across the enterprise, with no learning curve or programming required.

http://finance.yahoo.com/news/kyvos-insights-demonstrate-olap-hadoop-140000725.html

Sunday, September 18, 2016

Global Hadoop Services Market 2016 Is Showing Outstanding Signs in Recent Developments, Dynamic Opportunities & Emerging Projects

The intelligence report provides a framework of the global #Hadoop Services market considering the products, applications, analysis, overview and key areas of the market. The Hadoop Services market report projects the global market's valuation between Hadoop Services and Hadoop Services along with the value of the market during the forecast period 2016 to 2021. The Hadoop Services market also shades light on key market features such as the competitive rivalry within the market. This press release was orginally distributed by SBWire Brooklyn, NY -- (SBWIRE) -- 09/15/2016 -- The report presents an in-depth inquiry into the global Hadoop Services market scenario. The reports provides a complete outline of the current market state and future growth prospects, and tracks the historical growth of the global Hadoop Services market. The report also states the leading companies operating in the global Hadoop Services market. The report provides a framework of the global Hadoop Services market considering the products, applications, and key areas of the market. The research report projects the global marketa€™s valuation between Hadoop Services and Hadoop Services along with the value of the market during the forecast period. The report also highlights the chief regional segments of the global Hadoop Services market and provides their projected share in the global market.

http://m.digitaljournal.com/pr/3072709#ixzz4KemKCoXf


Thursday, August 25, 2016

SAP HANA big data strategy leans heavily on open source Hadoop tools

" #Bigdata is only going to get bigger and richer, as well as originate and flow from an increasing number of sources, both internal and external," according to a report from Forrester Research titled "Ultra-Fast Data Access Is The Key To Unleashing Full Big Data Potential." Enterprises need a "modern data analytics strategy that provides a ubiquitous, real-time data access layer to all relevant data from all different sources," the report noted. To meet the needs of these enterprises, #SAP is continuing to invest in providing business users with access to advanced analytics tools that use its #HANA in-memory, column-oriented, relational database management system, said Anne Moxie, senior analyst at Boston-based Nucleus Research. Werner Hopf, CEO of Dolphin Enterprise Solutions Corp., agreed with Moxie's assessment. Dolphin is an SAP partner based in Morgan Hill, Calif. "SAP invested a ton of development over the past two or three years to extend HANA capabilities, so it can also be used as the underlying database for transaction processing systems," he said. For example, last September, SAP announced HANA Vora, a new in-memory query engine for Hadoop that addresses the challenges companies face as they manage distributed big data, Moxie said.

http://searchsap.techtarget.com/feature/SAP-HANA-big-data-strategy-leans-heavily-on-open-source-Hadoop-tools

Tuesday, August 23, 2016

Big Data: What Could Go Wrong and What We Need to Do About It

The problem is way more complex than has been presented, and the stakes couldn’t be higher. To help you make better sense of your data and prepare for the future, I’ll address a number of #BigData challenges that have to be solved before we can say we have learned everything we can about our data – and taken all the steps we need to protect ourselves and others. First, some assumptions: 1. Network performance is not going to exceed data collection sizes, so we likely cannot move all the data to one place. 2. Given the declining Kryder rate, storage costs are not decreasing at rates seen in previous decades. With those two limitations in mind, what are the challenges for understanding, learning and making decisions from all the data we are collecting?

http://mobile.enterprisestorageforum.com/storage-management/big-data-storage-challenges-and-solutions.html

Thursday, August 11, 2016

Hortonworks Extends Streaming Analytics Reach of Hadoop

The ability to inexpensively collect massive amounts of #BigData using #Hadoop is one thing. Being able to manage all that data is another matter altogether. With that issue in mind, #Hortonworks today announced the upgrade of a Hortonworks DataFlow (HDF) streaming analytics platform with an eye toward addressing a broad range of enterprise management issues. HDF 2.0 now sports a revamped user interface that tightens the integration between the streaming analytics platform that is based on the #Apache NiFi data routing engine and Apache Kafka messaging software and Apache Ambari tools for managing Hadoop deployments. Apache #NiFi itself is based on streaming analytics software created by Onyara, which Hortonworks acquired last year. In addition, Hortonworks has announced that it has integrated HDF 2.0 with Apache Ranger software for securing access to data in Hadoop. Finally, Hortonworks revealed that it has developed Apache MiNiFi, an implementation of Apache NiFi that has been optimized to process streaming analytics on an Internet of Things ( #IoT ) gateway, while at the same time formally certifying over 150 third-party HDF 2.0 connectors. Jamie Engesser, vice president and general manager for emerging products at Hortonworks, says a truly comprehensive approach to Big Data requires an ability to process and manage data both in motion and at rest. HDF now extends the ability of Hortonworks to process data all the way out to the edge. That approach, adds Engesser, also serves to limit the amount of network bandwidth that needs to be allocated to an IoT project.

http://mobile.itbusinessedge.com/blogs/it-unmasked/hortonworks-extends-streaming-analytics-reach-of-hadoop.html

Monday, July 25, 2016

Disaster Recovery Planning for Hyper-Converged Infrastructure

Most of the chatter these days about #BigData analytics envisions a sprawl of inexpensive server/storage appliances arranged in highly scalable clustered-node configurations. This hyper-converged infrastructure ( #HCI ) is considered well-suited to the challenge of delivering a repository for a large and growing "ocean" (or "lake" or "pool") of data that is overseen by a distributed network of intelligent server controllers, all operated by a cognitive intelligence application or analytics engine. It all sounds very sci-fi. But, breaking it down, what are we really dealing with? HCI has never been well-defined. From a design perspective, it's pretty straightforward: a commodity server is connected to some storage that's usually mounted inside the server chassis (for example, internal storage) or externally connected via a bus extension interface (for example, direct-attached storage over Fibre Channel, SAS, eSATA or some other serial SCSI) all glued together with a software-defined storage ( #SDS ) stack implemented to provide control over the connected storage devices.

https://virtualizationreview.com/articles/2016/07/25/disaster-recovery-planning-for-hyperconverged-infrastructure.aspx?m=1

Monday, July 18, 2016

Splice Machine 2.0 combines HBase, Spark, NoSQL, relational...and goes open source

In the worlds of #BigData, #NoSQL and relational databases, #Splice Machine's name doesn't come up that often. But a closer look at the company's product, architectural approach and CEO put them on my radar a while back. And Version 2 of the product, which is being announced today, has made that radar dot much brighter. #Hadoop
Have RDBMS cake, eat NoSQL scaling, too
Before we look at version 2, let's cover the motivation behind v1. Specifically, Splice Machine looked long and hard at some pressing database conundrums:

The relational database model (along with SQL) works well -- best, in fact -- in many circumstances, but scaling it has always been hard.NoSQL databases are much easier to scale but the schema-less model and lack of "ACID" (Atomicity/Consistency/Isolation/Durability) guarantees can be disorienting.Hadoop scales well too, and its HDFS file system has become an important storage standard, but Hadoop's batch model can also cause dissonance for relational database professionals

The solution: create an ACID-compliant, SQL relational database on top ofApache HBase -- a NoSQL database that uses HDFS as its storage layer. Now you've got SQL, the relational model, ACID/transactional consistency, horizontal scaling and HDFS, all in one product

http://www.zdnet.com/article/splice-machine-2-0-combines-hbase-spark-nosql-relational-and-goes-open-source/

Sunday, July 17, 2016

The 10 Coolest Big Data Products Of 2016 (So Far)

Sales of big data and business analytics applications, tools and services hit nearly $122 billion last year and will grow more than 50 percent to $187 billion in 2019, according to market researcher IDC. So it's no wonder that the conveyor belt of new big data products hitting the market, both from established companies and startups, continues unabated. Here are 10 big data products that caught our attention in the first half of 2016. Some – but not all – of these made their debut at the #Strata + #Hadoop World conference in March or the Hadoop Summit in June. #AWS #Apache, #BlueData
http://m.crn.com/slide-shows/applications-os/300081349/the-10-coolest-big-data-products-of-2016-so-far.htm/pgno/0/3