Wednesday, April 26, 2023

What Might Generative AI Change for Data Centers? What Will Not Change?

Generative AI and Large Languate Models, as applied to many aspects of computing, will change soem things more than others.


Nvidia CEO Jensen Huang is quite positive about what routine use of large language models will mean for sales of advanced processors, citing an acceleration in demand for Nvidia processors. Some analysts believe generative AI could add as much as $6 billion in revenue for Nvidia within three years.  


Demand driver

Nvidia products that will benefit

Degree of impact

Increased demand for compute power

Datacenter GPUs, AI accelerators

High

Increased demand for data storage

DGX systems, storage solutions

Medium

Increased demand for bandwidth

Networking solutions

Medium

New security risks

Security solutions

Low


Others in the infra business should benefit as well. Some argue infrastructure to support generative AI could reach $50 billion by 2028, for example. Some note that a single LLM training operation can cost millions of dollars. Those costs are going to be borne by app providers of all sorts. 


Data center CxOs might also tend to agree about changes generative AI processing will create. We already know that generative AI requires lots of computing cycles, which we might quantify as floating point operations per second. The actual impact from a single request depends on the size of the dataset that is interrogated, the complexity of the question or task or the type of query. 


Text responses, perhaps ironically, might require an order of magnitude more processing than generating an image. And generation of an image might require an order of magnitude more processing operations than generating a text-to-speech use case. 


Model Type

FLOPS Requirements

Large Language Model (LLM)

100+ petaFLOPS

Image Generator

10+ petaFLOPS

Speech Generator

1+ petaFLOPS


So capital investment budgets will likely be altered to add more-more processing power than we presently tend to see. As a background issue, there will be some additional demand for higher connectivity between data centers, data centers and peering points and domains to domains, though it remains unclear how much incremental demand might be generated. It is a non-zero number, but it is hard to quantify at the moment. . 


Change

Investment response

Needs driving investment

Increased demand for compute power

Investment in more powerful hardware

Generative AI workloads are computationally intensive

Increased demand for data storage

Investment in more storage capacity

Large datasets that are required for AI training and inference operations

Increased demand for bandwidth

Investment in more network capacity

To support the high-speed data transfer that is required for AI applications

New security risks

Investment in security measures

To protect data centers from the growing threat of AI-powered cyberattacks


It also seems likely there could be architectural impact as well. Edge computing makes sense to support lower-latency use cases and applications. It is unclear so far whether the training of large data sets or end user operations will actually drive edge computing very much. 


Inference operations might still need to be conducted at large hyperscale centers, not at the edge. For that reason, use of large language models might not actually cause data center architecture to shift to greater reliance on edge computing, for example. 


Higher energy requirements will likely lead to newer approaches to cooling, though. 


Architecture or design change

Possible actions by data center operators

Magnitude of capex impact

Increased demand for compute and storage resources

Deploy more powerful servers and storage systems

High

Increased need for cooling and power

Upgrade cooling and power infrastructure

High

New security and compliance requirements

Implement new security and compliance measures

Medium

Changes in data center layout and design

Reconfigure data center layout and design

Low


Saturday, April 22, 2023

IBM Eats its Own AI Dog Food

Granted, you’d expect a technology firm to tout the benefits of using the tools it sells. Consider IBM, which highlights its positioning as a hybrid cloud supplier.  “Across IBM's IT environment, we're realizing the value of hybrid cloud,” said James Kavanaugh, IBM CFO. “We reduced the average cost of running an application by 90 percent by moving from a legacy data center environment to a hybrid cloud environment running on Red Hat OpenShift.” IBM, of course, owns Red Hat. 


“By standardizing global processes and applying AIOps, we are reducing our application portfolio by more than 35 percent,” he adds. “We've automated over 24 million transactions with RPA (robotic process automation), avoiding hundreds of thousands of manual tasks and eliminating the risk of human error.”


IBM also is applying artificial intelligence. “in HR, we now handle 94 percent of our company-wide HR inquiries with our AskHR digital system, speeding up the completion of many HR tasks by up to 75 percent.”


Tuesday, April 18, 2023

Large Language Models will Save Chatbots


Chatbots almost always suck. But large language models and generative AI should vastly improve that experience. Yeah, we know: almost anything would fix something that is so unsatisfying. But generative AI will allow answers to a far-greater range of questions and an improved way of fixing matters automatically. 

Not only does generative AI allow indexing a wider range of possible information, it also aids code writing routines. And better code generation should also assist automated responses (and fixes) for stated customer issues. 

Sunday, April 16, 2023

Large Language Model Inflection Point?

For most people, it seems as though artificial intelligence has suddenly emerged as an idea and set of possibilities. In truth, AI has been gestating for many many decades. But forms of AI already are used in consumer appliances such as smart speakers, recommendation engines and search functions.


What seems to be happening now is some inflection point in adoption. But consider the explosion of interest in large language models or generative AI. Development has been under way for more than 70 years. 


Search engines, smart phones and smart speakers have been using AI to support speech interfaces. In that sense, consumers have routinely been using AI-assisted devices and apps for some time. 


Large Language Models, or Generative AI, have lots of potential applications in virtually any setting where language, questions and answers or content creation--including development of computer code--is involved. 

source: Java T Point


Compared to earlier supervised learning models, large language models are self-supervised, able to scour huge amounts of internet data to predict the next word in a sentence. 


A large language model “is a type of artificial intelligence (AI) algorithm that uses deep learning techniques and massively large data sets to understand, summarize, generate and predict new content,” consultant Sean Kerner says. 


Right now, the obvious use cases are text summarization, chatbots, search, and code generation. Other use cases undoubtedly will develop. That suggests early use for customer service, text generation, writing of code and information retrieval tasks. 


It seems clear that large language models, as a subset of AI, have reached an inflection point. What remains unclear is the degree of progress and adoption. At most inflection points there is a quantitative shift in usage that often leads to qualitative impact. 


We will see quantitative change a lot faster than potential qualitative effects, near term. The qualitative changes will take longer, but should be far deeper than we now envision. That is just the way technology change tends to happen.


Generative AI Progress: Less than You Expect, Near Term; More than You Imagine Long Term

For most people, it seems as though artificial intelligence has suddenly emerged as an idea and set of possibilities. In truth, AI has been gestating for many many decades. But forms of AI already are used in consumer appliances such as smart speakers, recommendation engines and search functions.


What seems to be happening now is some inflection point in adoption. But consider the explosion of interest in large language models or generative AI. Development has been under way for more than 70 years. 


source: Black Hawk College


“Most people overestimate what they can achieve in a year and underestimate what they can achieve in ten years” is a quote whose provenance is unknown, though some attribute it to Standord computer scientist Roy Amara. Some people call it the “Gate’s Law.”


The principle is useful for technology market forecasters, as it seems to illustrate other theorems including the S curve of product adoption. The expectation for virtually all technology forecasts is that actual adoption tends to resemble an S curve, with slow adoption at first, then eventually rapid adoption by users and finally market saturation.   


That sigmoid curve describes product life cycles, suggests how business strategy changes depending on where on any single S curve a product happens to be, and has implications for innovation and start-up strategy as well. 


source: Semantic Scholar 


Some say S curves explain overall market development, customer adoption, product usage by individual customers, sales productivity, developer productivity and sometimes investor interest. It often is used to describe adoption rates of new services and technologies, including the notion of non-linear change rates and inflection points in the adoption of consumer products and technologies.


In mathematics, the S curve is a sigmoid function. It is the basis for the Gompertz function which can be used to predict new technology adoption and is related to the Bass Model.


Another key observation is that some products or technologies can take decades to reach mass adoption.


It also can take decades before a successful innovation actually reaches commercialization. The next big thing will have first been talked about roughly 30 years ago, says technologist Greg Satell. IBM coined the term machine learning in 1959, for example, and machine learning is only now in use. 


Many times, reaping the full benefits of a major new technology can take 20 to 30 years. Alexander Fleming discovered penicillin in 1928, it didn’t arrive on the market until 1945, nearly 20 years later.


Electricity did not have a measurable impact on the economy until the early 1920s, 40 years after Edison’s plant, it can be argued.


It wasn’t until the late 1990’s, or about 30 years after 1968, that computers had a measurable effect on the US economy, many would note.



source: Wikipedia


The S curve is related to the product life cycle, as well. 


Another key principle is that successive product S curves are the pattern. A firm or an industry has to begin work on the next generation of products while existing products are still near peak levels. 


source: Strategic Thinker


There are other useful predictions one can make when using S curves. Suppliers in new markets often want to know “when” an innovation will “cross the chasm” and be adopted by the mass market. The S curve helps there as well. 


Innovations reach an adoption inflection point at around 10 percent. For those of you familiar with the notion of “crossing the chasm,” the inflection point happens when “early adopters” drive the market. The chasm is crossed at perhaps 15 percent of persons, according to technology theorist Geoffrey Moore.

source 


For most consumer technology products, the chasm gets crossed at about 10 percent household adoption. Professor Geoffrey Moore does not use a household definition, but focuses on individuals. 

source: Medium


And that is why the saying “most people overestimate what they can achieve in a year and underestimate what they can achieve in ten years” is so relevant for technology products. Linear demand is not the pattern. 


One has to assume some form of exponential or non-linear growth. And we tend to underestimate the gestation time required for some innovations, such as machine learning or artificial intelligence. 


Other processes, such as computing power, bandwidth prices or end user bandwidth consumption, are more linear. But the impact of those linear functions also tends to be non-linear. 


Each deployed use case, capability or function creates a greater surface for additional innovations. Futurist Ray Kurzweil called this the law of accelerating returns. Rates of change are not linear because positive feedback loops exist.


source: Ray Kurzweil  


Each innovation leads to further innovations and the cumulative effect is exponential. 


Think about ecosystems and network effects. Each new applied innovation becomes a new participant in an ecosystem. And as the number of participants grows, so do the possible interconnections between the discrete nodes.  

source: Linked Stars Blog 


Think of that as analogous to the way people can use one particular innovation to create another adjacent innovation. When A exists, then B can be created. When A and B exist, then C and D and E and F are possible, as existing things become the basis for creating yet other new things. 


So we often find that progress is slower than we expect, at first. But later, change seems much faster. And that is because non-linear change is the norm for technology products. 


Friday, April 7, 2023

What Era of Computing Comes Next?

By now, all of us are aware that rapid reductions in computing and storage cost, with rapid increases in capability, can enable applications, use cases and revenue models that were not feasible in the past because computing or storage costs precluded them. 


So ridesharing is possible because people have capable smartphones and mobile internet access fast enough to support that use case. Netflix and other video streaming services are possible because digital infrastructure capabilities have been improved at Moore’s Law rates. 


Applied artificial intelligence is among the capabilities that benefit directly from rapid processing improvements. A study shows that, “before 2010 training compute grew in line with Moore’s law, doubling roughly every 20 months.”


But “since the advent of deep learning in the early 2010s, the scaling of training compute has accelerated, doubling approximately every six months,” say professors Jaime Sevilla, Lennart Heim, Anson Ho, Tamay Besiroglu, Marius Hobbhahn and Pablo Villalobos in a study


source: Jaime Sevilla, Lennart Heim, Anson Ho, Tamay Besiroglu, Marius Hobbhahn and Pablo Villalobos


source: Jaime Sevilla, Lennart Heim, Anson Ho, Tamay Besiroglu, Marius Hobbhahn and Pablo Villalobos


“Our findings seem consistent with previous work, though they indicate a more moderate scaling of training compute,” the researchers say. “In particular, we identify an 18-month doubling time between 1952 and 2010, a six-month doubling time between 2010 and 2022, and a new trend of large-scale models between late 2015 and 2022, which started two to three orders of magnitude over the previous trend and displays a 10-month doubling time.”


Moore's Law and rapid increases in computing power, with corresponding reductions in price, matter hugely. It allows entrepreneurs to innovate by asking the question “ what would my business look like if computing or bandwidth no longer were barriers?” 


Does anybody doubt that near-zero pricing remains among the biggest business threats in the connectivity business? And does anybody really doubt that Moore’s Law has led to substitute products for telco voice and messaging while diminishing the cost of transporting bits? 


Has bandwidth not increased, in lead markets, at the headline level, at about the rate Moore’s Law or Nielsen’s Law predicts? 


Edholm’s Law states that internet access bandwidth at the top end increases at about the same rate as Moore’s Law likewise suggests computing power will increase.


Nielsen's Law essentially is the same as Edholm’s Law, predicting an increase in the headline speed of about 50 percent per year. 


The point is that an inflection point has been reached for applied use of artificial intelligence. As we once might have asked “what does my business look like if computing or bandwidth were essentially free,” we now must start asking questions such as “what does my business look like if artificial intelligence is available to use?” 


As when those earlier questions were asked, the cost of training is nowhere near “free.” But neither was computing or bandwidth when the founders of Microsoft and Netflix laid out their plans. 


The most-startling strategic assumption ever made by Bill Gates was his belief that horrendously-expensive computing hardware would eventually be so low cost that he could build his own business on software for ubiquitous devices. .


How startling was the assumption? Consider that, In constant dollar terms, the computing power of an Apple iPad 2, when Microsoft was founded in 1975, would have cost between US$100 million and $10 billion.


Reed Hastings, Netflix founder, apparently made a similar decision. For Bill Gates, the insight that free computing would be a reality meant he should build his business on software used by computers.


Reed Hastings came to the same conclusion as he looked at bandwidth trends in terms both of capacity and prices. At a time when dial-up modems were running at 56 kbps, Hastings extrapolated from Moore's Law to understand where bandwidth would be in the future, not where it was “right now.”


“We took out our spreadsheets and we figured we’d get 14 megabits per second to the home by 2012, which turns out is about what we will get,” says Reed Hastings, Netflix CEO. “If you drag it out to 2021, we will all have a gigabit to the home." So far, internet access speeds have increased at just about those rates.


Everyone has struggled to define what the next era of computing would look like. We might have found our answer, at least relating to nomenclature. Some say we are in the era of cloud computing. Others might prefer mobile computing or web-based computing.


The point is that we left the mainframe, mini-computer, personal computer, client-server eras. Where we are now might be considered the internet, web, cloud-based or mobile era. We have not yet agreed on a specific term. 


What comes next might well be the AI era.


MWC and AI Smartphones

Mobile World Congress was largely about artificial intelligence, hence largely about “AI” smartphones. Such devices are likely to pose issue...