Dell, EMC, Dell Technologies, Cisco,

Showing posts with label DEEP LEARNING. Show all posts
Showing posts with label DEEP LEARNING. Show all posts

Monday, July 23, 2018

Ready-to-deploy deep learning solutions

Accelerate your deep learning project deployments with Radeon Instinct™ powered solutions Deep learning adoption is lagging as companies struggle with how to make it work. Now a new ecosystem is rising to deliver the integrated pieces that ultimately will be part of one turnkey system for deep learning. Automation has proved its worth in meeting IT and business objectives. Even so, efficiencies in automation and work augmentation software can be greatly enhanced with deep learning. Yet deep learning adoption rates are low. That’s in part because the tech is difficult, and the talent pool is thin. The good news is that an ecosystem is forming and already beginning to resolve some of these issues as it continues to grow towards becoming a single turnkey system. Why it takes an ecosystem A Deloitte report found that fewer than 10% of the companies surveyed across 17 countries invested in machine learning. The chief reasons for the adoption gap is a lack of understanding on how to use the technology, an insufficient amount of data to train it with, and a shortage of talent who could make it all work. Translated in the simplest of terms, deep learning is perceived by some to be too hard to deploy for practical use. The solution for that dilemma is what it has always been for any new technology requiring esoteric skill sets and faced with a talent shortage – build an easy-to-use, turnkey system. That is, of course, easier said than done. “The ongoing digital revolution, which has been reducing frictional, transactional costs for years, has accelerated recently with tremendous increases in electronic data, the ubiquity of mobile interfaces, and the growing power of artificial intelligence (AI),” according to a McKinsey & Company report. “Together, these forces are reshaping customer expectations and creating the potential for virtually every sector with a distribution component to have its borders redrawn or redefined, at a more rapid pace than we have previously experienced.” That’s why today’s sophisticated and complex systems are commonly constructed not by a single vendor but by a strong and diverse ecosystem capable of delivering the many moving parts needed to make a single turnkey system. Especially when said systems must be equally workable for companies across industries and with diverse needs. As a result, ecosystems are growing at breathtaking speeds. McKinsey & Company analysts predict that new ecosystems are likely to entirely replace many traditional industries by 2025. Such an ecosystem is forming for machine learning. It’s seeded with four recently launched, ready-to-deploy solutions. They center on AMD’s Radeon Instinct training accelerator for machine learning, and its ROCm Open eCosystem (ROCm), an open source HPC/Hyperscale-class platform for GPU computing. AMD takes open source all the way down to the graphics card level. Open source is key to successfully wrangling machine learning systems as it leverages the skills and coding work from entire communities and makes an ecosystem functional across technologies and applications. The ROCm open ecosystem This newly forming ecosystem is optimal for beginning or expanding your deep learning efforts whether you are the IT person looking to get pre-configured deep learning technologies in place, or the scientist who just needs access to HPC systems with one of the frameworks loaded. Either way, users can quickly get to work with their data. Developers also have full and open access to the hardware and software which speeds their work in developing frameworks. Everything AMD develops for its Radeon Instinct system is open source and available on GitHub. The company also has docker containers for easier installs of ROCm drivers and frameworks which can be found on the ROCm site for Docker.  Caffe and TensorFlow machine learning frameworks are offered now, with more to follow soon. A deep learning solutions page has gone live, which features the four systems that service as the bud of the blooming ecosystem rooted in AMD technologies. The frameworks docker containers will be listed there as well. This budding machine learning ecosystem is already bearing fruit for organizations looking to launch machine learning training and applications with a minimum of technical effort and expertise by combining: Fast and easy server deployments ROCm Open eCosystem and infrastructure Deep learning framework docker containers Optimized MIOpen framework libraries The four systems forming the ecosystem center “Data science is a mix of art and science—and digital grunt work. The reality is that as much as 80 percent of the work on which data scientists spend their time can be fully or partially automated,” according to a Deloitte report. This newly forming ecosystem is focused on automating much of the machine learning processes. While complicated to achieve, the end results are far easier for organizations to use. Deloitte identified five key vectors of progress that should help foster significantly greater adoption of machine learning by making it more accessible. “Three of these advancements—automation, data reduction, and training acceleration—make machine learning easier, cheaper, and/or faster. The others—model interpretability and local machine learning—open up applications in new areas,” according to the Deloitte report. There are four prebuilt systems shaping this ecosystem early on. Each is provided by an independent partner and built on or for AMD’s Radeon Instinct and ROCm platforms, but their initial presentations are at varying levels of integration. While more partners will join the ecosystem over time, these four provide a solid bedrock for organizations looking to get started in machine learning now. 1) AMAX is providing systems with preloaded ROCm drivers and a choice of framework, either TensorFlow or Café, for machine learning, advanced rendering and HPC applications. 2) Exxact is similarly providing multi-GPU Radeon Instinct-based systems with preloaded ROCm drivers and frameworks for deep learning and HPC-class deployments, where performance per watt is important. 3) Inventec provides optimized high performance systems designed with AMD EPYC™ processors and Radeon Instinct compute technologies capable of delivering up to 100 teraflops of FP16 compute performance for deep learning and HPC workloads. 4) Supermicro is providing SuperServers supporting Radeon Instinct machine learning accelerators for AI, big data analytics, HPC, and business intelligence applications. The payoff from leveraging the technologies in a machine learning ecosystem potentially comes in many forms. “A growing number of tools and techniques for data science automation, some offered by established companies and others by venture-backed start-ups, can help reduce the time required to execute a machine learning proof of concept from months to days. And, automating data science means augmenting data scientists’ productivity in the face of severe talent shortages,” say the Deloitte researchers.

https://www.hpcwire.com/2018/07/23/ready-to-deploy-deep-learning-solutions/

Wednesday, March 28, 2018

Supermicro's New Scale-Up Artificial Intelligence and Machine Learning Systems with 8 NVIDIA Tesla V100 with NVLink GPUs Deliver Superior Performance and System Density

SAN JOSE, Calif., March 27, 2018 /PRNewswire/ -- Super Micro Computer, Inc. (NASDAQ: SMCI), a global leader in enterprise computing, storage, networking solutions and green computing technology, today is showcasing the industry's broadest selection of GPU server platforms that support NVIDIA® Tesla® V100 PCI-E and V100 SXM2 GPU accelerators at the GPU Technology Conference in the San Jose McEnery Convention Center, booth 215, through March 29. @Supermicro offers best performing #GPU servers with #Tesla V100 PCI-E and V100 SXM2 32GB GPUs For maximum acceleration of highly parallel applications like #artificialintelligence ( #AI ), #deeplearning, #selfdriving cars, #smartcities, #healthcare, #bigdata, #HPC, #virtualreality and more, Supermicro's new 4U system with next-generation @NVIDIA #NVLink™ interconnect technology is optimized for maximum performance. SuperServer 4029GP-TVRT supports eight NVIDIA Tesla V100 32GB SXM2 GPU accelerators with maximum GPU-to-GPU bandwidth for cluster and hyper-scale applications. Incorporating the latest NVIDIA NVLink technology with over five times the bandwidth of PCI-E 3.0, this system features independent GPU and CPU thermal zones to ensure uncompromised performance and stability under the most demanding workloads. "On initial internal benchmark tests, our 4029GP-TVRT system was able to achieve 5,188 images per second on ResNet-50 and 3,709 images per second on InceptionV3 workloads," said Charles Liang, President and CEO of Supermicro.  "We also see very impressive, almost linear performance increases when scaling to multiple systems using GPU Direct RDMA. With our latest innovations incorporating the new NVIDIA V100 32GB PCI-E and V100 32GB SXM2 GPUs with 2X memory in performance-optimized 1U and 4U systems with next-generation NVLink, our customers can accelerate their applications and innovations to help solve the world's most complex and challenging problems." "Enterprise customers will benefit from a new level of computing efficiency with Supermicro's high-density servers optimized for NVIDIA Tesla V100 32GB data center GPUs," said Ian Buck, vice president and general manager of accelerated computing at NVIDIA. "Twice the memory with V100 drives up to 50 percent faster results on complex deep learning and scientific applications and improves developer productivity by reducing the need to optimize for memory." "At Preferred Networks, we continue to leverage Supermicro's high-performance 4U GPU servers to successfully power our private supercomputers," said Ryosuke Okuta, CTO of Preferred Networks. "These state-of-the-art systems are already powering our current supercomputer applications, and we have already begun the process of deploying Supermicro's optimized new 4U GPU systems loaded with NVIDIA Tesla V100 32GB GPUs to drive our upcoming new private supercomputers." Supermicro is also demonstrating the performance-optimized 4U SuperServer 4029GR-TRT2 system that can support up to 10 PCI-E NVIDIA Tesla V100 accelerators with Supermicro's innovative and GPU-optimized single root complex PCI-E design, which dramatically improves GPU peer-to-peer communication performance. For even greater density, the SuperServer 1029GQ-TRT supports up to four NVIDIA Tesla V100 PCI-E GPU accelerators in only 1U of rack space and the new SuperServer 1029GQ-TVRT supports four NVIDIA Tesla V100 SXM2 32GB GPU accelerators in 1U.   With the convergence of big data analytics and machine learning, the latest NVIDIA GPU architectures, and improved machine learning algorithms, deep learning applications require the processing power of multiple GPUs that must communicate efficiently and effectively to expand the GPU network. Supermicro's single-root GPU system allows multiple NVIDIA GPUs to communicate efficiently to minimize latency and maximize throughput as measured by the NCCL P2PBandwidthTest

NVIDIA to Unleash Deep Learning in Hyperscale Datacenters

@NVIDIA CEO unveils Volta-based GV 100 for workstations, new inferencing software, technologies providing a 10x boost for deep learning, a self-driving car simulator, and more. Millions of servers powering the world’s hyperscale data centers are about to get a lot smarter. NVIDIA CEO @JensenHuang Tuesday announced new technologies and partnerships that promise to slash the cost of delivering deep learning-powered services. Speaking at the kickoff of the company’s ninth annual GPU Technology Conference, Huang described a “Cambrian Explosion” of technologies driven by GPU-powered deep learning that are bringing support for new capabilities that go far beyond accelerating images and video. “In the future, starting with this generation, starting with today, we can now accelerate voice, speech, natural language understanding and recommender systems as well as images and video,” Huang, clad in his trademark leather jacket, told an audience of 8,500 technologists, business leaders, scientists, analysts and press gathered at the San Jose Convention Center. Over the course of a two-and-a-half hour keynote, Huang also unveiled a series of advances to NVIDIA’s deep learning computing platform that deliver a 10x performance boost on deep learning workloads from just six months ago; launched GV 100, transforming workstations with 118.5 TFLOPS of deep learning performance; introduced DRIVE Constellation to run self-driving car systems for billions of simulated miles. Power to the Pros Huang’s keynote got off to a brisk start, with the launch of the new Quadro GV 100. Based on Volta, the world’s most advanced GPU architecture, Quadro GV100 packs 7.4 TFLOPS double-precision, 14.8 TFLOPS single-precision and 118.5 TFLOPS deep learning performance, and is equipped with 32GB of high-bandwidth memory capacity. NVIDIA CEO Jensen Huang launches the Quadro GV100. GV100 sports a new interconnect called NVLink 2 that extends the programming and memory model out of our GPU to a second one. They essentially function as one GPU. These two combined have 10,000 CUDA cores, 236 teraflops of Tensor Cores, all used to revolutionize modern computer graphics, with 64GB of memory. Deep Learning’s Swift Rise The announcements come as deep learning gathers momentum. In less than a decade, the computing power of GPUs has grown 20x — representing growth of 1.7x per year, far outstripping Moore’s law, Huang said. “We are all in on deep learning, and this is the result,” Huang said. Drawn to that growing power, in just five years the number of GPU developers has risen 10x to 820,000. Downloads of CUDA, our parallel computing platform, have risen 5x to 8 million. “More data, more computing are compounding together into a double exponential for AI, that’s one of the reasons why it’s moving so fast” Huang said. Bringing Deep Learning Inferencing to Millions of Servers The next step: putting deep learning to work on a massive scale. To meet this challenge, technology will have to address seven challenges: programability, latency, accuracy, size, throughput, energy efficiency and rate of learning. Together, they form the acronym PLASTER.

Monday, January 22, 2018

Google’s Vision for Mainstreaming Machine Learning

Here at The Next Platform, we’ve touched on the convergence of machine learning, HPC, and enterprise requirements looking at ways that vendors are trying to reduce the barriers to enable enterprises to leverage AI and machine learning to better address the rapid changes brought about by such emerging trends as the cloud, edge computing and mobility. At the SC17 show in November 2017, Dell EMC unveiled efforts underway to bring AI, machine learning and deep learning into the mainstream, similar to how the company and other vendors in recent years have been working to make it easier for enterprises to adopt HPC techniques for their environments. For Dell EMC, that means in part doing so through bundled, engineered systems. IBM has strategies underway, including through the integration of its PowerAI deep learning enterprise software with its Data Science Experience. Both offerings are aimed at making it easier for enterprises to embrace advance AI technologies and for developers and data scientists to develop and train machine learning models. Software vendors like SAP and Microsoft similarly are easing the way for enterprise adoption of AI and machine learning techniques. @Google has been a vocal proponent of the idea of democratizing #AI by making it easier for mainstream businesses to use. As we outlined last year, Google sees its capabilities in AI, machine learning and deep learning as a competitive advantage against larger hyperscale players like @Amazon Web Services and @Microsoft #Azure in the highly competitive #cloudcomputing arena and a key way of enticing enterprises to the #GoogleCloudPlatform to help them manage, analyze and leverage the rapidly mounting data to more quickly and effectively build products and services that their customers can use. Google in 2016 created a machine learning unit within its Cloud Platform business and rolled out a series of APIs for such jobs as natural language, language translation and vision recognition. The hyperscaler also brought  in Fei-Fei Li, at the time the director of Stanford University’s Artificial Intelligence Lab and the Stanford Vision Lab, to head up the new machine learning group

https://www.nextplatform.com/2018/01/22/googles-vision-mainstreaming-machine-learning/

Thursday, January 18, 2018

Machine Learning in Finance: Challenges, Successes & Opportunities

AI or machine learning is changing the way industries across the spectrum interact with their customers, as well as develop their processes. And nowhere is this more evident than in the financial services business. #Machinelearning in finance is pushing the industry to the edge of technological advancement.

A new @insideHPC special report, sponsored by @Dell EMC and @NVIDIA, explores the benefits, challenges and considerations involved with adopting machine learning in finance.

The guide delves into recent key machine learning innovations in financial services to give readers examples of how the technology is being used today, what’s needed to leverage machine learning, what applications and technologies to use and more. It also offers ways to learn from recent successes in the field, as well as ways to connect with the broader machine learning community.

As financial institutions look to machine learning, they should first acknowledge the benefits, and challenges.CLICK TO TWEET

First up, when considering machine learning in finance solutions, one should start with reviewing the challenges involved to make sure to consider all options and scenarios. According to the guide, challenges include regulation and compliance, cybersecurity and fraud detection, risk management — and, of course, competition.

Next up, the report explores some of the technology behind machine learning in finance. Of all of the technological innovations that have made machine learning possible, GPUs have perhaps had the most impact.

Of course, after learning about the technology behind machine learning, a company needs to decide which solution is right for them. The special report goes into detail on some of these solutions, such as the Dell EMC PowerEdge C4140 server and NVIDIA Volta. Frameworks, such as Caffe, Tensorflow, Torch, Microsoft Cognitive Toolkit (CNTK), and Apache Mahout — used to support machine learning applications — are also explored. 

After absorbing all this information, the guide offers case studies, information on professional organizations and further reading on machine learning. Research centers described in this guide include the Dell EMC Customer Solution Center, Dell EMC Machine Learning Knowledge Center and the NVIDIA Deep Learning Institute. 

The guide is ideal for those in the financial services industry who are beginning to explore the potential of machine learning, as well as those looking to expand and maximize its use.

https://insidebigdata.com/2018/01/17/machine-learning-in-finance/

Monday, January 1, 2018

NVIDIA Prohibits Use Of GeForce GPUs In Datacenters, Only Blockchain Processing Allowed

@NVIDIA made a pretty big change in its #GeForce EULA recently and this is something that could go on to cost a lot of entities an aggregate of millions if not billions of dollars in the long run. The company recently updated their EULA which now prohibits datacenter deployment of their GeForce GPUs for everything but #blockchain processing. Needless to say, this would force a shift to #Quadro and #Tesla’s in any #datacenters that were actually using GeForce cards or had planned to. Change in EULA targets use of GeForce cards in #AI and #DeepLearning applications in non-personal ‘datacenter’ use While there was no explicit restriction against using GeForce graphics cards in datacenters before, this was before deep learning took off in recent past. Previously, the company has always offered consumer variants in crippled form as far as double precision goes but now, with the advent of DNN processing, that is no longer as relevant as it once was. In fact, half precision is the name of the game now and NVIDIA GPUs are extremely good at deep neural net crunching and the consumer variants offer a much higher return on investment in CuDNN then the professional variants. “No Datacenter Deployment. The SOFTWARE is not licensed for datacenter deployment, except that blockchain processing in a datacenter is permitted.” -Extract from the EULA. I wouldn’t even be surprised to learn that there are (were?) data centers being built right now which initially planned to use GeForce cards (prompting this subtle backlash from NVIDIA) for the majority of processing. It is clear that NVIDIA feels that it has spent a lot of R&D on the DNN tech involved and if the industrial customers don’t actually pay the premium, it will cut into its profits. It is also clear that increasing prices of the GeForce cards is not an option if it wants to remain competitive and still retain its hold over the PC market. The solution? restrict datacenters from using their GeForce cards for any kind of processing, forcing them to use Teslas and Quadros for DNN work.

https://www.google.com/amp/s/wccftech.com/nvidia-geforce-eula-prohibits-datacenter-blockchain-allowed/amp/#ampshare=https://wccftech.com/nvidia-geforce-eula-prohibits-datacenter-blockchain-allowed/

Sunday, November 19, 2017

Amazon Web Services backs deep-learning format introduced by Microsoft and Facebook

The #OpenNeuralNetworkExchange ( #ONNX) #deep-earning format, introduced in September by @Microsoft and @Facebook, has a new backer following @Amazon Web Services’ decision to embrace the framework with a new open-source project. #AWS released ONNX-MXNet Thursday afternoon, which sounds like a telecom standard from the 1990s but is actually a method for allowing deep learning models built around the ONNX format to run on the @Apache MXNet framework. In supporting ONNX, envisioned as a standard way to build deep learning models, the cloud computing leader gives a significant stamp of approval to the concept of open data models. Advanced research into deep learning is a relatively new field, even though we’ve been talking about it for several years. There aren’t a lot of go-to methods for building data models, and it can be difficult to switch frameworks upon discovering that your model might benefit more from a different tactic. That was the idea behind ONNX, a bridge between open-source deep-learning frameworks like Caffe2, developed at Facebook, and Cognitive Toolkit, developed by Microsoft. It also supports the open-source PyTorch framework, and now Apache MXNet thanks to the work of AWS. These all seem like positive developments for those at the forefront of artificial intelligence research, when major players in the field agree to forge a common path that makes it easier for people to use deep-learning techniques in their own applications or businesses. All eyes now turn to Google, which has made AI and deep learning a huge part of its cloud strategy but has yet to join forces with the ONNX group
https://www.geekwire.com/2017/amazon-web-services-backs-deep-learning-format-introduced-microsoft-facebook/

Tuesday, November 14, 2017

Dell EMC Releases New HPC Solutions & PowerEdge Server

Today at #SuperComputing2017 in Denver, Colorado, @Dell EMC made a handful of announcements surrounding #HPC and #dataanalytics and how they intend to bring this technology, along with #machinelearning and #deeplearning, to the mainstream. The above-mentioned technology can bring several benefits around fraud detection, image processing, financial investment analysis and personalized medicine. To this end, the company is introducing Dell EMC Ready Bundles for Machine and Deep Learning as well as the new Dell EMC PowerEdge C4140 server aimed at #cognitiveworkloads.   #ArtificialIntelligence ( #AI ), such as deep and machine learning, has been a big buzz recently. While there are many benefits to deploying this type of technology, Dell EMC argues that many vendors don’t possess the expertise to help organizations deploy and manage AI. Dell EMC will leverage its ecosystem to help customers gain the most insights using HPC. On top of this, the company will help customers choose HPC solutions tailored to their use cases. The first part of bringing the benefits of HPC to the masses is Dell EMC's new Machine and Deep Learning Ready Bundles. Combining the knowledge of Dell EMC experts as well as the expertise of its partners, these bundles are the combination of pre-tested and validated servers, storage, networking and services optimized for machine and deep learning applications. The Ready Bundles are a result of a close partnership with Dell EMC and Intel as they two companies are working to collaborate on advancing artificial intelligence, machine learning and deep learning. Dell EMC Ready Bundles for Machine and Deep Learning benefits include: Enable faster, better, deeper data insights: Identifying, analyzing and automating data patterns empowers customers to do more with their data in a wide range of applications such as facial recognition for security, tumor diagnosis in healthcare, and better understanding of human behaviors in the retail industry. Include trusted experts: Using the knowledge and experience of Dell EMC and its strong ecosystem of technology partners helps customers get the most out of machine and deep learning solutions quickly. Maximize efficiency, security and control: Enabling customers to reduce the costs associated with moving significant amounts of data in hybrid cloud environments while minimizing risk and maximizing data control.  Dell EMC will now be supporting its HPC with its powerful 14G PowerEdge Servers and they are introducing a new PowerEdge server today designed specifically for HPC workloads, the Dell EMC PowerEdge C4140 server. As part of a joint development agreement with NVIDIA, this new sever supports up to four NVIDIA Tesla V100 GPU accelerators with PCIe and NVLink high-speed interconnect technology. The servers also leverages two Intel Xeon Scalable Processors and is ideal for intensive machine learning and deep learning applications to drive advances in scientific imaging, oil and gas exploration, financial services and other HPC industry verticals. Availability The Dell EMC Ready Bundles for Machine Learning and Deep Learning are expected to be available in the first half of 2018. TheDell EMC PowerEdge C4140 is expected to be available worldwide in December 2017.
http://www.storagereview.com/dell_emc_releases_new_hpc_solutions_poweredge_server

Monday, November 13, 2017

Dell EMC Wants to Take AI Mainstream

One of the challenges vendors are trying address when it comes to artificial intelligence is expanding the technology and its elements of machine learning and deep learning beyond the realm of hyoerscalers and some HPC centers and into the enterprise, where businesses can leverage them for such workloads as simulations, modeling, and analytics. For the past several years, system makers have been trying to crack the code that will make it easier for mainstream enterprises to adopt and deploy traditional HPC technologies, and now they want to dovetail those efforts with the expanding AI opportunity. The difference with enterprises is that they are deploying HPC and now AI to adapt to the rapid changes roiling their industries, including cloud computing, mobility and computing at the network edge. As we at The #NextPlatform have noted, system vendors like @IBM, @Lenovo, and @Hewlett Packard Enterprise have all been pushing to build out their #HPC capabilities and bring those capabilities to the masses. #HPE has been a player in the HPC and #supercomputer realm for years, and bulked up those efforts last year when it bought @SGI for $275 million. For its part, Lenovo also expanded its presence in the space when it bought IBM’s System x server business in 2014 for $2.1 billion and licensed such software as the Platform Computing middleware stack. While IBM’s HPC focus with the X86 server unit was on the high-end systems, Lenovo said last year that it has seen increasing success with smaller systems that can help bring HPC capabilities to mainstream businesses. More recently, IBM last month expanded its own efforts in this area by integrating its PowerAI deep learning enterprise software with its Data Science Experience, both of which are designed to reduce the challenges for enterprises interested in using advanced AI technologies and enable developers and data scientists to develop and train machine learning models. Software vendors like Microsoft and SAP also are pushing to make it easier for enterprises to adopt AI and deep learning technologies. Similarly, for years, Dell EMC has been growing its HPC product portfolio with an eye toward getting the technologies into the mainstream. That has included innovations around its PowerEdge C-Series modular servers for HPC clusters and efforts with such supercomputers as the Texas Advanced Computer Center (TACC) at the University of Texas at Austin and the Julich Supercomputing Centre in Germany. In 2015, the vendor launched its portfolio of HPC systems, which includes its Ready Bundle for HPC, pre-configured and pre-validated engineered systems that include the compute, storage, network, services and software support needed for enterprises to quickly deploy and use HPC technologies in their datacenters. The lineup includes Ready Bundles for HPC optimized for NFS and Lustre storage. Dell EMC also offers a range of HPC services that touch on everything from deployment and management to the cloud and finances and support for various HPC software stacks. The goal is to deliver the same kinds of capabilities that at one time had only been afforded to research and educational institutions and large businesses, according to Ravi Pendekanti, senior vice president for product management and marketing for Dell EMC’s Server Solutions business. “HPC started off as more of a scientific and research kind of oriented effort,” Pendekanti said in a recent press conference. “Over the last couple of years, it has transformed itself into being may one of the pivotal things in the financial industry. If you look at any of the emerging workloads – machine learning, deep learning and edge computing – they’ve all got something to do with HPC. The workloads that we talk about today, when we talk about things like machine learning and deep learning, we weren’t talking about 10 years ago. We’re now talking about how HPC has moved into the mainstream from the back alley of research and the scientific community.” At the SC17 supercomputing conference in Denver this week, Dell EMC is unveiling efforts to help bring AI, deep learning and machine learning to enterprises in much the same way that it’s been pushing HPC into the mainstream, through bundles engineered systems. At the same time, the company is upgrading its current HPC Ready Bundle offerings with the latest-generation PowerEdge C-Series servers. The AI bundles are designed to enable enterprises to more quickly analyze the massive amounts of data being collected and generate insights that can be applied to their businesses. The Ready Bundles for Deep Learning with Nvidia, Deep Learning with Intel and Machine Learning with Hadoop, which will become available in the first half of 2018, will be based on Dell EMC’s new PowerEdge C4140, an ultradense system that is part of Dell’s 14th Generation PowerEdge lineup. The bundles also include libraries and frameworks around such open-source HPC technologies TensorFlow, Caffe, BigDL and cuBLAS and support for such software platforms as Cloudera, Hortonworks and Spark. The bundled systems follow Dell’s announcement at the ISC 17 show in June of a strategic partnership with Nvidia in which the companies will jointly develop new products around HPC, data analytics and AI. Dell has a similar collaboration with Intel around AI, machine learning and deep learning solutions. The C4140 includes two of Intel’s 14 nanometer “Skylake” Xeon SP processors – each chip will have up to 20 cores – and four Nvidia Tesla V100 GPU accelerators with PCI-Express and NVLink interconnect technology to link the compute elements together. The 1U C4140, which will be available in December, delivers up to 500 teraflops through the Tesla V100 GPUs, which are pooled using the NVLink interconnect, and provides 2,400 watts in its power supplies to support next-generation GPUs. The system offers 24 DIMMs for up to 1.5 TB of DDR4 memory on the Xeon SPs for faster throughput and supports Linux distributions from Red Hat, Canonical and SUSE. The GPU accelerators also speed up the processing of HPC workloads while driving down costs, according to Dell EMC, which estimated that for molecular dynamic applications, a single C4140 can do the work of 19 CPU-only servers and save 12 times the cost. For financial services workloads, the server can equal the work of eight CPU-only servers and save five times the costs.

https://www.nextplatform.com/2017/11/13/dell-emc-wants-take-ai-mainstream/

Sunday, November 12, 2017

HPE Developing its Own Low Power “Neural Network” Chips

With so many chip startups targeting the future of #deeplearning training and inference, one might expect it would be far easier for tech giant @Hewlett Packard Enterprise to buy versus build. However, when it comes to select applications at the extreme edge (for space missions in particular), nothing in the ecosystem fits the bill. In the context of a broader discussion about the company’s #ExtremeEdge program focused on space-bound systems, #HPE ’s Dr. Tom Bradicich, VP and GM of #Servers, #ConvergedEdge, and #IoT systems, described a future chip that would be ideally suited for #highperformancecomputing under intense power and physical space limitations characteristic of space missions. To be more clear, he told us as much as he could—very little is known about the architecture, but there was some key elements he described. First, the architecture, called the Dot Product Engine (dig out math notes on vectors and dot products since it’s relevant here) is less of a full processor and more like an accelerator, which takes offload of certain multiplication elements common in neural network inference and broader HPC applications. Bradicich says that the prototype as it stands can be standalone for certain problems but could also be PCIe or direct sensor attached eventually to cut down on bandwidth and latency constraints. From his first description, it sounds very much like a tiny version of Nvidia’s tensor core, which (at its simplest) can handle a 4X4 matrix operation-based workload well. With the power consumption constraints of extreme edge environments, something like a Volta GPU with the TensorCore, for instance, would be far too energy hungry—not to mention demand a larger form factor. Other devices from the chip startup world have other limitations as well. HPE decided to roll its own—and has been heads down cobbling together an early software stack to support the DPEs for eventual productization (assuming that happens). As Bradicich tells us, ““These days you have to make a lot of make versus buy decisions and today, we buy a lot of Intel processors and a lot of Nvidia GPUs. But to get what we want, we are seeing we can’t get the performance and energy we need for this and we need to take a ‘make’ approach.” The tough part about the discussion was that HPE insisted on calling this a neural network processor, but how that moniker matches against what we think of with other neural network architectures is still unclear. In some ways, it almost seems like HPE could tie this to the “neuromorphic” term easier than neural network since it is a non Von Neumann architecture that sounds like it might be based on memristor memory concepts (see some of our work on The Machine if this is unfamiliar). It also sounds like there could be an FPGA angle here as well, but HPE is unable to comment on specifics of the hardware architecture. What they did say after a second round of questioning about the neural network angle is the term is being used in a broader way. “DPE is not a neural network per se, in the sense that it’s not a fixed configuration, but rather is reconfigurable, and can be used for inference of several types of neural networks (DNN, CNN, RNN). Hence it can do neural network jobs and workloads,” Bradicich clarifies. “DPE is executes linear algebra in the analog domain, which is more efficient than digital implementations, such as dedicated ASICs. And, further, it has the advantage of reconfigurability on the fly. It’s fast because it accelerates vector * matrix math, dot product multiplication, by exploiting Ohms Law on a memristor array. It can also be used for other types of operations, such as FFT, DCT, and convolution.” We put a similar question to HPE Labs Rebecca Lewington, who provided a bigger picture view of what a systems level take on these process looks like. “You need to look at the fleet of IoT devices holistically. We believe that if architected and connected, you would have a central learning engine that can gather the experiences of the entire fleet of devices. To train the neural network, you would retrain a centralized model and then push that model to the edge devices again, which is where the inference is done. That way, every device in your fleet has the ability to learn from every other device in your fleet. In the next generation, we would incorporate the Memory-Driven Computing architecture, which would enable us to attach task specific accelerators—like the Dot-Product Engine—to a centralized pool of memory. That means that we can complete the model training much more rapidly and more energy efficiently, as well as more flexibly because we can accommodate new learning frameworks as they become available.” When asked to put this chip architecture in the context of similar devices (TensorCore in the Volta GPU, Nervana, Graphcore, Wave Computing, etc) Bradicich said it is similar to several existing chip designs in that it can handle vectors well, but he says how it is designed from a hardware and software co-design standpoint makes it far faster than anything on the market—and a much better fit for the performance, power, and space requirements of extreme edge environments. “We are not afraid to buy or partner for technology like this but there is nothing like this on the marketplace horizon.” We asked what, if any, neural network frameworks this might run inference for; whether or not it can do it training, and of course, if this is really conventional neural networks we’re talking about here. No answers on any of those that point to anything we think of here at The Next Platform as true neural nets. Neuromorphic perhaps–but neural networks seems a stretch. We dislike being shy on technical details but wanted to set the stage for more information set to emerge on this architecture following the prototype demo in Spain at the end of the month during one of HPE’s events. The interesting part for now is that perhaps some of the tech giants that we guessed would have started snapping up deep learning chip startups really haven’t. The reason is becoming clear: although the architectures are promising for now, perhaps what is needed is less general purpose for many machine learning algorithms and is instead tailored for applications or environments. This takes us back to the premise we started with a few years ago; IT overall is moving from homogeneous mega-clouds and datacenters and chip concepts to specialization, novel architectures, and fine-tuned approaches to

https://www.nextplatform.com/2017/11/09/hpe-developing-low-power-neural-network-chips/

Wednesday, October 25, 2017

HPE Introduces New Set of Artificial Intelligence Platforms and Services

PALO ALTO, Calif., Oct. 25, 2017 (GLOBE NEWSWIRE) -- @Hewlett Packard Enterprise (NYSE:HPE) today announced new purpose-built platforms and services capabilities to help companies simplify the adoption of #ArtificialIntelligence, with an initial focus on a key subset of #AI known as #deeplearning. Inspired by the human brain, deep learning is typically implemented for challenging tasks such as image and facial recognition, image classification and voice recognition. To take advantage of deep learning, enterprises need a high performance compute infrastructure to build and train learning models that can manage large volumes of data to recognize patterns in audio, images, videos, text and sensor data. Many organizations lack several integral requirements to implement deep learning, including expertise and resources; sophisticated and tailored hardware and software infrastructure; and the integration capabilities required to assimilate different pieces of hardware and software to scale AI systems. To help customers overcome these challenges and realize the potential of AI, HPE is announcing the following offerings: HPE Rapid Software Installation for AI: HPE introduced an integrated hardware and software solution, purpose-built for high performance computing and deep learning applications. Based on the HPE Apollo 6500 system in collaboration with Bright Computing to enable rapid deep learning application development, this solution includes pre-configured deep learning software frameworks, libraries, automated software updates and cluster management optimized for deep learning and supports NVIDIA® Tesla V100 GPUs. HPE Deep Learning Cookbook: Built by the AI Research team at Hewlett Packard Labs, the deep learning cookbook is a set of tools to guide customers in selecting the best hardware and software environment for different deep learning tasks. These tools help enterprises estimate performance of various hardware platforms, characterize the most popular deep learning frameworks, and select the ideal hardware and software stacks to fit their individual needs. The Deep Learning Cookbook can also be used to validate the performance and tune the configuration of already purchased hardware and software stacks. One use case included in the cookbook is related to the HPE Image Classification Reference Designs. These reference designs provide customers with infrastructure configurations optimized to train image classification models for various use cases such as license plate verification and biological tissue classification. These designs are tested for performance and eliminate any guesswork, helping data scientists and IT to be more cost-effective and efficient. HPE AI Innovation Center: Designed for longer term research projects, the innovation center will serve as a platform for research collaboration between universities, enterprises on the cutting edge of AI research and HPE researchers. The centers, located in Houston, Palo Alto, and Grenoble, will give researchers for academia and enterprises access to infrastructure and tools to continue research initiatives. Enhanced HPE Centers of Excellence (CoE): Designed to assist IT departments and data scientists who are looking to accelerate their deep learning applications and realize better ROI from their deep learning deployments in the near term, the HPE CoE offer select customers access to the latest technology and expertise including the latest NVIDIA GPUs on HPE systems. The current CoE are spread across five locations including Houston; Palo Alto; Tokyo; Bangalore, India; and Grenoble, France. “We live in a world today where we’re generating copious amounts of data, and deep learning can help unleash intelligence from this data,” said Pankaj Goyal, vice president, Artificial Intelligence Business, Hewlett Packard Enterprise. “However, a ‘one size fits all’ solution doesn’t work. Each enterprise has unique needs that require a distinct approach to get started, scale and optimize its infrastructure for deep learning. At HPE, we aim to make AI real for our customers no matter where they are in their journeys with our industry-leading infrastructure portfolio, AI expertise, world-class research and ecosystem of partners.” In its mission to help make AI real for its customers, HPE offers customers a flexible consumption services for HPE infrastructure, which avoids over-provisioning, increases cost savings and scales up and down as needed to accommodate the needs of deep learning deployments. “Artificial intelligence has the ability to transform scientific data analysis, making predictions and surprising connections,” said Paul Padley, professor of physics and astronomy, Rice University. “We are at a precipice where the AI revolution can now have a profound impact on reshaping innovation, science, education and society, at large. Access to the HPE AI innovation centers will help us continue to advance our research efforts in our journey to making academic progress by using the tools and solutions available to us through HPE.” AI is becoming mainstream in the consumer world with applications such as voice interfaces, personal assistants and image tagging. However, the implications of AI go beyond mainstream consumer use cases to fields including genomic sequencing analytics, climate research, medical science, autonomous driving and robotics. These technology advancements and breakthroughs have been – and continue to be – made possible by deep learning. “As deep learning-based AI advances, it will transform science, commerce and the quality of our lives by automating tasks that don't require the most complex human thinking," said Steve Conway, senior vice president, Hyperion Research. “HPE’s infrastructure and software solutions are designed for ease-of-use and promise to play an important role in driving AI adoption into enterprises and other organizations in the next few years.” Learn more about how HPE is bringing deep learning techniques to customers in VP of Artificial Intelligence Pankaj Goyal’s latest blog

https://globenewswire.com/news-release/2017/10/25/1153239/0/en/HPE-Introduces-New-Set-of-Artificial-Intelligence-Platforms-and-Services.html

Tuesday, October 24, 2017

Designing the Future of Deep Learning 

#ArtificialIntelligence and #DeepLearning are being used to solve some of the world's biggest problems and is finding application in #autonomousdriving, #marketing and #advertising, #health and #medicine, #manufacturing, #multimedia and #entertainment, #financialservices, and so much more.  This is made possible by incredible advances in a wide range of technologies, from computation to interconnect to storage, and innovations in software libraries, frameworks, and resource management tools.  While there are many critical challenges, an open technology approach provides significant advantages. The Scaling Challenge The full deep learning story, though, must be an end-to-end technology discussion and encompass production at scale.  As we scale out deep learning workloads to the massive compute clusters required to tackle these big issues, we begin to run into the same challenges that hamper scaling of traditional high-performance computing (HPC) workloads. Ensuring optimal use of compute resources can be challenging, particularly in heterogeneous architectures that may include multiple central processing unit (CPU) architectures, such as x86, ARM64, and Power, as well as accelerators, such as graphical processing units (GPUs), field programmable gate arrays (FPGAs), tensor processing units (TPUs), etc. Architecting an optimal deep learning solution for training or inferencing, with potentially varied data types, can result in the application of one or more of these architectures and technologies. The flexibility of open technologies allows one to deploy the optimal platform at server, rack, and data center scales. One of the most important uses of deep learning is in gaining value from large data sets. The need to effectively manage large amounts of data, which may have varying ingest, processing, and persistent storage and data warehouse needs, is at the center of a modern deep learning solution. The performance requirements throughout the data workflow and processing stages can vary greatly, and, at production-scale, it can simultaneously involve data collection, training, and inference.  The balance of cost effectiveness and high performance is key to providing a properly-scaled deployment. The flexibility of open technologies, allows one to take a software-designed data center approach to the deep learning environment. Workload orchestration is another familiar challenge in the HPC realm.  A variety of tools and libraries have been developed over the years, including resource managers and job schedulers, parallel programming libraries, and other software frameworks.  As software applications have grown in complexity, with rapidly evolving dependencies, a new approach has been needed.  One such approach is containerization.  Containers allow applications to be bundled with their dependencies and deployed on a variety of compute hosts.  However, challenges have remained for providing access to compute, storage, and other resources.  Moreover, managing the deployment, monitoring, and clean-up of containerized applications presents its own set of challenges. The Open Technology Approach Penguin Computing applies its decades of expertise in high-performance and scale-out computing to deliver deep learning solutions that support customer workload requirements, whether at development or production scales.  Penguin Computing solutions feature open technologies, enabling design choices that focus on meeting the customer's needs. In the Penguin Computing AI/DL whitepaper, you will learn more about our approach to: Open Architectures for Artificial Intelligence and Deep Learning, combining flexible compute architectures, rack scale platforms, and software-defined networking and storage, to provide a scalable software-defined AI/DL environment. Discuss AI/DL strategies, providing insight into everything from specialty compute for training vs. inference to Data Lakes and high performance storage for data workflows to orchestration and workflow management tools. Deploying the AI/DL environments from development to production scale and from on-premise to hybrid to public cloud.
https://www.enterprisetech.com/2017/10/23/designing-future-deep-learning/

Thursday, October 12, 2017

Amazon and Microsoft unveil ‘Gluon’ neural network technology, teaming up on machine learning

@Microsoft and @Amazon, which surprised the tech world with a partnership between their #Cortana and #Alexa virtual assistants, are back at it again. @AmazonWebServices and Microsoft’s #AI and Research Group this morning announced a new #opensource deep learning interface called Gluon, jointly developed by the companies to let developers “prototype, build, train and deploy sophisticated machine learning models for the cloud, devices at the edge and mobile apps,” according to an announcement just released by the companies. Deep learning involves training a computer to recognize patterns or unlock insights based on a set of rules for parsing a massive pool of data. As you might expect, this is an extremely complicated and time-consuming process that requires a fair amount of skill. Cloud companies offer ways to speed up the process, but a fair amount of skill is required to get meaningful results. And most software developers interested in incorporating deep learning technology into their applications don’t have nearly the amount of expertise on hand at either AWS or Microsoft; let alone their combined expertise. Gluon will give those developers a way to tap into that expertise without having to invest nearly as much time and effort in understanding how to use machine learning techniques. Gluon allows developers to write deep-learning systems in the popular Python language and take advantage of deep learning templates developed by Microsoft and AWS. This makes it much easier to get up and running and much easier to tweak those templates for an application’s specific needs. Developers interested in learning more about the technology can check out Gluon here. @Microsoft CEO @SatyaNadella and @Amazon CEO @JeffBezos talked about collaborating at Microsoft’s CEO Summit last year, and executives acknowledged after the #Alexa -#Cortana announcement that it might not be the last partnership between them. Microsoft #Azure and @AWS compete aggressively in the cloud, but under Nadella, Microsoft has made a point of partnering strategically with its rivals. “Amazon is a very impressive company,” said Nadella at the GeekWire Summit this week. “What Jeff and his team have done is something that I’ve long admired, and I think there’s a lot that we can learn. In fact, the good news is that between Microsoft and Amazon, we have a lot of cross-pollination of talent, and I think it’s helpful for this region, by the way, which is something that Silicon Valley always had.”

https://www.geekwire.com/2017/amazon-microsoft-announce-gluon-neural-network-technology-teaming-machine-learning/

Tuesday, April 18, 2017

Lenovo HPC Strategy Update

Engineer at #Lenovo presents: Lenovo #HPC Strategy Update. “High performance computing is converging more and more with the big data topic and related infrastructure requirements in the field. Lenovo is investing in developing systems designed to resolve todays and future problems in a more efficient way and respond to the demands of Industrial and research application landscape

Improving risk analytics, shortening product development cycles, and simulating the behavior of materials at nanoscale have something in common—the need to model, forecast, and analyze complex relationships. All place enormous demands on your data and compute infrastructure.

Meet these demands and realize faster time to value, improve critical decision making, fuel innovation, and accelerate results with powerful Lenovo high-performance computing systems, storage, software, networking, and pre-integrated solutions for compute-intensive and big data applications.

Luigi Brochard holds a PhD in mathematics and is Executive Director for High Performance Computing at Lenovo. During his career at IBM he was WW Architect for IBM Deep Computing and HPC Technical Director for x86 server systems. He was nominated to Distinguished Engineer for his technical leadership and achievements in x86 systems architecture.

http://insidehpc.com/2017/04/lenovo-hpc-strategy-update/

Thursday, January 19, 2017

Intel Unveils Deep Learning Framework for Spark

Chip giant #Intel last week rolled out a new deep learning framework that runs as a #Spark job atop #Hadoop. Called #BigDL, the #opensource software is designed to take advantage of hardware acceleration capabilities that Intel has built into its Xeon CPUs. BigDL, which Intel released on #Github, is modeled after #Torch, an open source deep learning framework used in #scientificcomputing. Intel says the framework supports numeric computing via Tensor, as well as high level neural networks, and can be used to run prebuilt Caffe and Torch models on Spark. While many deep learning frameworks today leverage GPUs, Intel is taking a different route with BigDL, for obvious reasons. The new framework uses Intel’s Math Kernel Library (MKL), which enables the workload to execute as a multi-threaded Spark job and take full advantage of the multi-threading extensions Intel’s Xeon processors. All this adds up to fast execution of deep learning workloads, the company says. In fact, Intel says BigDL can run “orders of magnitude faster than out-of-box open source Caffe, Torch, or TensorFlow on a single-node Xeon” processor. That makes it “comparable with mainstream GPU,” the chip giant says. Intel says the new framework will be useful for analyzing large amounts of data on Hadoop or Spark clusters, or for adding deep learning functionality to existing Spark applications or workflows.It will also be useful, the company says, when a user wants to use share the results of deep learning workloads with other applications running on Hadoop or Spark clusters, such as ETL, data warehouses, feature engineering, “classical machine learning,” or graph analytics. Intel is waging a war for big data dominance against GPUs, and the delivery of BigDL to the open source community figures to play a part. In November, the company outlined its hardware strategy for giving developers more powerful tools for building artificial intelligence applications. The company says it will test the first AI-specific hardware, code-named “Lake Crest,” in the first half of 2017, with limited availability later in the year. The company also announced another new product, code-named “Knights Crest,” that will integrate its Xeon processors with the technology it obtained with its August acquisition of Nervana Systems, which developed deep learning products

https://www.datanami.com/2017/01/18/intel-unveils-deep-learning-framework-spark/

Sunday, January 15, 2017

Microsoft just bought an AI startup that can outperform Facebook and Google

#Microsoft announced this morning that it has acquired #Maluuba, a Toronto startup focused on using deep learning for natural language processing. #Deeplearning is an approach to #artificialintelligence currently in vogue that has driven incredible gains in the field over the last five years. As Microsoft wrote in the blog post announcing the purchase, “We’ve recently set new milestones for speech and image recognition using deep learning techniques, and with this acquisition we are, as Wayne Gretzky would say, skating to where the puck will be next — machine reading and writing.” The Verge covered Maluuba in the summer of 2016, when the startup shared the results of an AI system that could read and comprehend text with near human capability, outperforming similar systems shown off by Google and Facebook. Along with acquiring the company, Microsoft has also established closer ties with Yoshua Bengio, a pioneer in the field of deep learning who served as an advisor to Maluuba, and will now become and advisor to Microsoft’s AI division.

http://www.theverge.com/2017/1/13/14266398/microsoft-acquires-maluuba-ai-deep-learning-yoshua-bengio

Wednesday, November 11, 2015

Google researcher: Quantum computers aren’t perfect for deep learning

In the past couple of years, #Google has been trying to improve more and more of its services with artificial intelligence. Google also happens to own a #quantumcomputer — a system capable of performing certain computations faster than classical computers.

It would be reasonable to think that perhaps Google would try running AI workloads on its quantum computer from startup #DWave, which is kept at #NASA ’s Ames Research Center in Mountain View, California, right near Google headquarters.

Google is keen on advancing its capabilities in a type of AI called deep learning, which involves training artificial neural networks on a large supply of data and then getting them to make inferences about new data.

But at an event at Google headquarters last week, a Google researcher explained that the quantum computing infrastructure just isn’t the best fit for systems such as convolutional neural networks or recurrent neural networks.

Several other tech companies — including #Facebook, #Microsoft, and #Baidu — have been experimenting with deep learning in the context of image recognition, natural language processing, and speech recognition. Those other companies are large, with plenty of money to spend on infrastructure. But they don’t have quantum computers. Google does. Still, that doesn’t mean it’s always useful.

If anything, Google may be more interested in using the D-Wave machine to work on improving core Google processes like search ranking, the placement of advertisements, and spam filtering, if one report from last is correct. (And Google may well be planning to talk more about its quantum work; the company is planning to hold an event on the subject on December 8, according to a report today from 9to5Google.)

http://venturebeat.com/2015/11/11/google-researcher-quantum-computers-arent-perfect-for-deep-learning/