Success in computing often centers on having the flexibility to know when to pivot an idea, and yet simultaneously, the steadiness to not wander off on hopeless yet seductive tangents. Pivoting is a surprisingly subtle skill.
A bit of context. We say that a project has done a pivot if it sets out in one direction but later shifts to focus on some other way of using some of the same ideas. A pivot that tosses out the technology isn't really what I mean... I'm more interested in the kind of mid-course correction that doesn't require a ground-up rethinking, yet might have a profound impact on the marketing of a product or technology.
Pivots are a universal puzzle for entrepreneurs and researchers alike. I'm pretty good at finding the right research questions, but not nearly as capable of listening to the market. I remember an episode at Reliable Network Solutions (RNS). This was a company founded by Werner Vogels, Robbert van Renesse and me, to do management solutions for what we would call cloud computing data centers today. Back then the cloud term wasn't yet common, so those were really very early days.
We launched RNS to seize an initial opportunity that was irresistible: call it the ultimate business opportunity in this nascent cloud space. In our minds, we would solve the problem for one of the earliest major players and instantly be established as the go-to company for a solution. Now, as it happened, our initial customer had special requirements, so the solution we created for them was something of a one-off (they ultimately used it, and I think they still do), but our agreement left us the freedom to create a more general product that we could sell to everyone else. So, the RNS business plan centered on this follow-on product. Our vision was that we would create a new and more scalable, general-purpose, robust, self-repairing, fast, trustworthy management infrastructure solution. We were sure it would find a quick market uptake: after all, customer zero had been almost overwhelmingly hungry for such a solution, albeit in a slightly more specialized setting unique to their datacenter setups.
RNS ended up deciding to base this general product on Astrolabe, a scalable gossip-based management information database that we came up with as a research prototype, inspired by that first product for customer zero. We had the prototype running fairly quickly, and even got a really nice ACM TOCS paper out of the work. Astrolabe had every one of our desired properties, and it was even self-organizing -- a radical departure from the usual ways of monitoring and managing datacenter systems.
Astrolabe was a lovely, elegant idea. Nonetheless, when we made this decision to bet the bank on it (as opposed to doing something much narrower, along the lines of what we did for that first customer), we let our fondness for the concept get way out ahead of what the market really wanted.
I remember one particular road trip to San Francisco. We met with the technology leaders of a large company based there, and they set up a kind of brainstorming session with us. The intent was to find great fits for Astrolabe in their largest datacenter applications. But guess what? They turned out to really need a lot less than Astrolabe... and they were nervous that Astrolabe was quite exotic and sophisticated... it was just too big a step for them.
In fact, they wanted something like what Derecho offers today: an ultra-fast tool for replicating files and other data. They would have wanted Derecho to be useable as a stand-alone command-line solution (we lack this for Derecho: someone should build one!). But they might have gone for it as a C++ library. In effect, 15 years later, I finally have what those folks were basically looking for at the time.
At any rate, we had a potential next big customer and a fantastic opportunity for which our team was almost ideally suited -- my whole career has focused on data replication. Even so, RNS just couldn't take this work on. We had early customers for our Astrolabe concept, plus customer zero continued to give us some work, and our external investors had insisted on a set of Astrolabe milestones. We were simply spread too thin, and so even though our friends in San Francisco spoke the truth, it was just impossible for us to take their advice.
But guess what? Reliable Network Solutions never did manage to monetize Astrolabe in a big way, although we did make a number of sales, and the technology ultimately found a happy home. We just sort of bumped along for a few years, making barely enough money to pay all our bills, and never hitting that home run that can transform a small company into a big one. And then, in 2001, the 9-11 terrorist attack triggered a deep tech downturn. By late 2002 our sales pipeline had frozen up (many companies pause acquisitions during tech downturns), and we had no choice but to close the doors. 35 people lost their jobs that day.
The deep lesson I learned was that a tech company always needs to be awake to the possibility that it has the strategy wrong, that its customers may be able to see the obvious even though company leaders are blind because of their enthusiasm for the technology they've been betting on, and that a pivot might actually save the day. The path forwards sometimes leads to a dead end.
This is a hard lesson to teach. I've sometimes helped the Runway "post-doc" incubator program, which is part of the Jacobs Institute at NYC Tech, the technology campus Cornell runs in New York jointly with Technion. My roles have been pretty technical/engineering ones: I've helped entrepreneurs figure out how to leverage cloud computing to reduce costs.
But teaching entrepreneurs to be introspective about their vision and to pivot wisely? This has been an elusive skill, for me. The best person I've ever seen at this is our Runway director, Fernando Gomez-Baquero. Fernando favors an approach that actually starts by accepting the ideas of the entrepreneurial team at face value. But then he asks them to validate their concept. He isn't unreasonable about this, and actually helps them come up with a plan. Just the same, he never just accepts assumptions as given. Everything needs to be tested.
This is often hard: one can stand in the hallway at a big tech show-and-tell conference and interview passing CTOs (who are often quite happy to hold forth on their visions for the future), but even if a CTO shows interest, one has to remember my experience with RNS: product one, even a successfully delivered product that gets deployed, might not be a scalable story that can be resold widely and take a company down the path to wild profits.
Fernando is also very good at the non-technology side of companies: understanding what it takes to make a product look sexy and feel right to its actual end-users. A lot of engineering teams neglect the whole look and feel aspect because they become so driven by the technical ideas and research advances they started with that they just lose track of the bigger picture. Yet you can take a fantastic technology and turn it into a horrible, unusable product. Fernando's students never make that kind of mistake: for him, technology may be the "secret sauce", but you have to make the dish itself irresistible, and convince yourself that you've gotten that part right.
This end, what Fernando does is to ask the entrepreneurs to list 20 or so assumptions about their product and their market. The first few are always easy. Take RNS: we would have said we were betting that datacenters would become really immense (they did), that managing a million machines at a time is very hard (it is), that the management framework needs to be resilient, self-repairing, stable under stress (and indeed, all of these are true). Fernando, by the way, wouldn't necessarily agree that these are three assumptions and would never split that last item into subitems. Instead, he might actually call this just one or two assumptions, and if you have additional technology details to add, he would probably be inclined to just lump them in. This gets back to my point above: For a startup company, technology is only a tiny part of the story. All these "superior properties" of the RNS solution? Fernando would just accept that. Ok, "you have technical advantages." But he would insist on seeing the next 20 items on the list.
When you are constrained to carry out a seemingly strange task, it can feel natural to list human factors: that CTOs have a significant cost exposure on datacenter management, that they know this to be true, and that they would spend money to reduce that exposure. That they might have an appetite for a novel approach because they view existing ones as poorly scalable. You might start to toss in cost estimate items, or pricing guesses. If 20 turns out to be too easy a target, Fernando would read your list, then ask for a few more.
Fernando once told me an anecdote that I remember roughly this way: Suppose that a set of entrepreneurs wanted to bring Korean ice cream to the K-12 school cafeterias in America. They might outline this vision, extoll the virtues of ice cream in tiny little dots (Korean ice cream is made by dripping the ice cream mix into liquid nitrogen "drop by drop"), and explain that for the launch, they would be going with chocolate. They would have charts showing the explosive sales growth of Korean ice cream in K-12 schools in Seoul. Profits will be immense! Done deal, right?
Well, you can test some of these hypotheses. It isn't so hard to visit K-12 cafeterias, and the head chef probably would be very happy to be interviewed in many of them. But guess what you might learn? Perhaps, FDA rules preclude schools from serving ice-cream in a format that could possibly be inhaled. Perhaps chocolate is just not a popular flavor these days strawberry is the new chocolate, or mint, or Oreo crunch. Korea happens to be a fairly warm country: the demand for ice cream is quite high during the school year there. Maybe less-so here. Anyhow, ice cream isn't considered healthy here, and not every cafeteria serves sweets.
Korean dot-style ice cream? Awesome idea! But this kind of 20-questions approach can rule out ill-fated marketing plans long before the company tries to launch its product. Our ice-cream venture could easily have been headed straight for a costly (but yet, delicious) face-plant. Perhaps the company would have failed. Launch the exact same concept on the nation's 1000 or so small downtown pedestrian malls, and it could have a chance. But here we have a small pivot: how would that shift in plan change the expense models? Would we rent space, or kisoks, or sell from machines? Would there be new supply challenges, simply to get the ice cream to the points of sale? Schools are easy to at least find: you can download a list. How would you find those 1000 small malls, and who would be selling the product at each location?
Had we done a this 20-questions challenge with Astrolabe, we might have encountered the friendly folks in California far sooner, gotten the good advice we received early enough to act on it, and realized that Astrolabe was just too exotic for our target market to accept as a product. Technically superior? Absolutely, in some dimensions -- the ones academic researchers evaluate. But not every product needs technical superiority to win in the market, and conversely, not every technical attribute is decisive in the eyes of CTO buyers.
Fernando is a wizard at guiding companies to pivot before they make that costly error and find themselves so committed to the dead end that there is simply no path except to execute the plan, wise or foolish.
Today, I look at some of my friends who lead companies, and I'm seeing examples everywhere of how smartly timed pivots can be game-changing. I was blind to that sort of insight back in 2000 when I was actually running RNS.
The art of the pivot isn't necessarily natural for everyone. By now, I understand the concept and how it can be applied, but even so, I find that I'm prone to fall in love with technology in ways that can cloud my ability to assess a market opportunity in an unbiased way. It really isn't easy.
So here's my advice: Anyone reading this blog is a technology-centric datacenter person. Perhaps you teach and do research, as I do (most of the time, anyway). Perhaps you work for a big company... perhaps you are even toying with trying to launch one. But I would argue that every one of us has something to learn from this question of technology pivots. Our students, or our colleagues, or our direct reports: they may be holding fast to some sort of incorrect belief, betting more and more heavily on it, and yet that belief may be either outright wrong, or perhaps just in need of a tweak.
You yourself may be more locked-in on some aspect of your concept than you realize. That rigidity could be your undoing!
Fernando's 20-questions challenge is awesome because it doesn't approach the issue confrontationally. After all, the entrepreneurs themselves make the list of assumptions -- it isn't as if Fernando imposes them. They turn each assumption into a question, too, Jeopardy style. Then he just urges them to validate the resulting questions, easiest first. Some pull this off... others come back eager to pivot after a few days, and open to brainstorming based on what they learned.
What an amazing idea. I wish that back then, we could have launched RNS in the Runway program! Today, I might be running all the world's cloud computing systems.
Showing posts with label cloud computing. Show all posts
Showing posts with label cloud computing. Show all posts
Saturday, 15 February 2020
Thursday, 10 October 2019
If the world was a (temporal) database...
I thought I might share a problem we've been discussing in my Cornell PhD-level class on programming at the IoT edge.
Imagine that at some point in the future, a company buys into the idea (my idea!) that we'll really need smart highway systems to ever safely use self-driving cars. A bit like air traffic control, but scaled up.
In fact a smart highway might even offer guidance and special services to "good drivers", such as permission to drive at 85mph in a special lane for self-guided cars and superior drivers... for a fee. Between the smart cars and the good drivers, there would be a lot of ways to earn revenue here. And we could even make some kind of dent in the endless gridlock that one sees in places like Silicon Valley from around 6:45am until around 7pm.
So this company, call it Smart Highways Inc, sets out to create the Linux of smart highways: a new form of operating system that the operator of the highway could license to control their infrastructure.
What would it take to make a highway intelligent? It seems to me that we would basically need to deploy a great many sensors, presumably in a pattern intendent to give us some degree of redundancy for fault-tolerance, covering such things as roadway conditions, weather, vehicles on the road, and so forth.
For each vehicle we would want to know various things about it: when it entered the system (for eventual tolls, which will be the way all of this pays for itself), its current trajectory (a path through space-time annotated with speeds and changes in speed or direction), information about the vehicle itself (is it smart? is the driver subscribed to highway guidance or driving autonomously?), and so forth.
Now we could describe a representative "app": perhaps, for the stretch of CA 101 from San Francisco to San Jose, a decision has been made to document the "worst drivers" over a one month period. (Another revenue opportunity: this data could definitely be sold to insurance companies!) How might we do this? And in particular, how might we really implement our solution?
What I like about this question is that it casts light on exactly the form of Edge IoT I've been excited about. On the one hand, there is an AI/ML aspect: automated guidance to the vehicles by the highway, and in this example, an automated judgement about the quality of driving. One would imagine that we train an expert system to take trajectories as input and output a quality metric: a driver swerving between other cars at high speed, accelerating and turning abruptly, braking abruptly, etc: all the hallmarks of poor driving!
But if you think more about this you'll quickly realize that to judge quality of driving you need a bit more information. A driver who swerves in front of another car with inches to spare, passes when there are oncoming vehicles, causes others to break or swerve to avoid collisions -- that driver is far more of a hazard than a driver who swerves suddenly to avoid a pothole or some other form of debris, or one who accelerates only while passing, and passes only when nobody is anywhere nearby in the passing lane. A driver who stays the right "when possible" is generally considered to be a better driver than one who lingers in the left, if the highway isn't overly crowded.
But if you think more about this you'll quickly realize that to judge quality of driving you need a bit more information. A driver who swerves in front of another car with inches to spare, passes when there are oncoming vehicles, causes others to break or swerve to avoid collisions -- that driver is far more of a hazard than a driver who swerves suddenly to avoid a pothole or some other form of debris, or one who accelerates only while passing, and passes only when nobody is anywhere nearby in the passing lane. A driver who stays the right "when possible" is generally considered to be a better driver than one who lingers in the left, if the highway isn't overly crowded.
A judgment is needed: was this abrupt action valid, or inappropriate? Was it good driving that evaded a problem, or reckless driving that nearly caused an accident?
So in this you can see that our expert system will need expert context information. We would want to compute the set of cars near each vehicle's trajectory, and would want to be able to query the trajectories of those cars to see if they were forced to take any kind of evasive action. We need to synthesize metrics of "roadway state" such as crowded or light traffic, perhaps identify bunches of cars (even on a lightly driven road we might see a grouping of cars that effectively blocks all the lanes), etc. Road surface and visibility clearly are relevant, and roadway debris. We would need some form of composite model covering all of these considerations.
I could elaborate but I'm hoping you can already see that we are looking at a very complicated real-time database question. What makes it interesting to me is that on the one hand, it clearly does have a great deal of relatively standard structure (like any database): a schema listing information we can collect for each vehicle (of course some "fields" may be populated for only some cars...), one for each segment of the highway, perhaps one for each driver. When we collect a set of documentation on bad drivers of the month, we end up with a database with bad-driver records in it: one linked to vehicle and driver (after all, a few people might share one vehicle and perhaps only some of the drivers are reckless), and then a series of videos or trajectory records demonstrating some of the "all time worst" behavior by that particular driver over the past month.
But on the other hand, notice that our queries also have an interesting form of locality: they are most naturally expressed as predicates over a series of temporal events that make up a particular trajectory, or a particular set of trajectories: on this date at this time, vehicle such-and-such swerved to pass an 18-wheeler truck on the right, then dove three lanes across to the left (narrowly missing the bumper of a car as it did so), accelerated abruptly, braked just as suddenly and dove to the right... Here, I'm describing some really bad behavior, but the behavior is best seen as a time-linked (and driver/vehicle-linked) series of events that are easily judged as "bad" when viewed as a group, and yet that actually would be fairly difficult to extract from a completely traditional database in which our data is separated into tables by category.
Standard databases (and even temporal one) don't offer particularly good ways to abstract these kinds of time-related event sequences if the events themselves are from a very diverse set. The tools are quite a bit better for very regular structures, and for time series data with identical events -- and that problem arises here, too. For example, when computing a vehicle trajectory from identical GPS records, we are looking at a rather clean temporal database question, and some very good work has been done on this sort of thing (check out Timescale DB, created by students of my friend and colleague, Mike Freedman!). But the full-blown problem clearly is the very diverse version, and it is much harder. I'm sure you are thinking about one level of indirection and so forth, and yes, this is how I might approach such a question -- but it would be hard even so.
In fact, is it a good idea to model a temporal trajectory a database relation? I suspect that it could be, and that representing the trajectory that way would be useful, but this particular kind of relation just lists events and their sequencing. Think also about this issue of event type mentioned above: here we have events linked by the fact that they involve some single driver (directly or perhaps indirectly -- maybe quite indirectly). Each individual event might well be of a different type: the data documenting "caused some other vehicle to take evasive action" might depend on the vehicle, and the action, and would be totally different from the data documenting "swerved across three lanes" or "passed a truck on its blind side."
Even explaining the relationships and causality can be tricky: Well, car A swerved in front of car B, which braked, causing C to brake, causing D to swerve and impact E. A is at fault, but D might be blamed by E!
In fact, as we move through the world -- you or me, as drivers of our respective vehicles, or for that matter even as pedestrians trying to cross the road, this aspect of building self-centered domain-specific temporal databases seems to be something we do very commonly, and yet don't model particularly well in today's computing infrastructures. Moreover, you and I are quite comfortable with highways that might have cars and trucks, motorcycles, double-length trucks, police enforcement vehicles, ambulances, construction vehicles... extremely distinct "entities" that are all capable of turning up on a highway, our standard ways of building databases seem a bit overly structured for dealing with this kind of information.
Think next about IoT scaling. If we had just one camera, aimed at one spot on our highway, we still could do some useful tasks with it: we could for example equip it with a radar-speed detector that would trigger photos and use that to automatically issue speeding tickets, as they do throughout Europe. But the task I described above fuses information from what may be tens of thousands of devices deployed over a highway more than 100 miles long at the location I specified, and that highway could have a quarter-million vehicles on it at peak commute hours.
As a product opportunity, Smart Highways Inc is looking at a heck of a good market -- but only if they can pull off this incredible scaling challenge. They won't simply be applying their AI/ML "driving quality" evaluation to individual drivers, using data from within the car (that easier task is the one Hari Balakrisnan's Cambridge Mobile Telematics has tackled, and even this problem has his company valued in the billions as of round A). Smart Highways Inc is looking at the cross-product version of that problem: combining data across a huge number of sensor inputs, fusing the knowledge gained, and eventually making statements that involve observations taken at multiple locations by distinct devices. Moreover, we would be doing this at highway scale, concurrently, for all of the highway all of the time.
In my lecture today, we'll be talking about MapReduce, or more properly, the Spark/Databricks version of Hadoop, which combines an open source version of MapReduce with extensions to maximize the quality of in-memory caching and introduces a big-data analytic ecosystem. The aspect of interest to me is the caching mechanism: Spark centers on a kind of cacheable query object they call a Resilient Distributed Data object, or RDD. An RDD describes a scalable computation designed to be applicable across a sharded dataset, which enables a form of SIMD computing at the granularity of files or tensors being processed on huge numbers of compute nodes in a datacenter.
The puzzle for my students, which we'll explore this afternoon, is whether the RDD model could be transferred from the batched, non-real-time settings in which it normally runs (and even more than that, functional, in the sense that Spark treats every computation as a series of read-only data transformation steps, from a static batched input set through a series of MapReduce stages to a sharded, distributed result). So our challenge is: could a graph of RDDs and an interative compute model express tasks like the Smart Highway ones?
The puzzle for my students, which we'll explore this afternoon, is whether the RDD model could be transferred from the batched, non-real-time settings in which it normally runs (and even more than that, functional, in the sense that Spark treats every computation as a series of read-only data transformation steps, from a static batched input set through a series of MapReduce stages to a sharded, distributed result). So our challenge is: could a graph of RDDs and an interative compute model express tasks like the Smart Highway ones?
RDDs are really a linkage between database models and a purely functional Lisp-style Map and Reduce functional computing model. I've always liked them, although my friends who do database research tend to view them dimnly, at best. They often feel that all of Spark is doing something a pure database could have done far better (and perhaps more easily). Still, people vote with their feet and for whatever reason, this RDD + computing style of coding is popular.
So... could we move RDDs to the edge? Spark itself, clearly, wouldn't be the proper runtime: it works in a batched way, and our setting is event-driven, with intense real-time needs. It might also entail taking actions in real-time (even pointing a camera or telling it to take a photo, or to upload one, is an action). So Spark per-se isn't quite right here. Yet Spark's RDD model feels appropriate. Tensor Flow uses a similar model, by the way, so I'm being unfair when I treat this as somehow Spark-specific. I just have more direct experience with Spark, and additionally, see Spark RDDs as a pretty clear match to the basic question of how one might start to express database queries over huge IoT sensor systems with streaming data flows. Tensor Flow has many uses, but I've seen far more work on using it within a single machine, to integrate a local computation with some form of GPU or TPU accelerator attached to that same host. And again, I realize that this may be unfair to Tensor Flow. (And beyond that I don't know anything at all about Julia, yet I hear that system name quite often lately...)
Anyhow, back to RDDs. If I'm correct, maybe someone could design an IoT Edge version of Spark, one that would actually be suitable for connecting to hundreds of thousands of sensors, and that could really perform tasks like the one outlined earlier in real-time. Could this solve our problem? It does need to happen in real-time: a smart highway generates far too much data per second to keep much of it, so a quick decision is needed that we should document the lousy driving of vehicle A when driver so-and-so is behind the wheel, because this person has caused a whole series of near accidents and actual ones -- sometimes, quite indirectly, yet always through his or her recklessness. We might need to make that determination within seconds -- otherwise the documentation (the raw video and images and radar speed data) may have been discarded.
If I was new to the field, this is the problem I personally might tackle. I've always loved problems in systems, and in my early career, systems meant databases and operating systems. Here we have a problem of that flavor.
Today, however, scale and data rates and sheer size of data objects are transforming the game. The kind of system needed would span entire datacenters, and we will need to use accelerators on the data path to have any chance at all of keeping up. So we have a mix of old and new... just the kind of problem I would love to study, if I was hungry for a hugely ambitious undertaking. And who know... if the right student knocks on my door, I might even tackle it.
Wednesday, 26 June 2019
Whiteboard analysis: IoT Edge reactive path
One of my favorite papers is the one Jim Gray wrote with Pat Helland, Patrick O'Neil and Dennis Shasha, on the costs of replicating a database over a large set of servers, which they showed to be prohibitive if you don't fragment (shard) the database into smaller and independentally accessed portions: mini-databases. In some sense, this paper gave us the modern cloud, because you can view Brewer's CAP conjecture and the eBay/Amazon BASE methodologies as both flowing from Gray's original insight.
Fundamentally, what Jim and his colleagues did was to undertake a whiteboard analysis of the scalability of concurrency control in an uncontrolled situation, where transactions are simply submitted to some big pool of servers, and then compete for locks in accordance with a two-phase locking model (one in which a transaction acquires all its locks before releasing any), and then terminates using a two-phase or three-phase commit. They show that without some mechanism to prevent lock conflicts, there is a predictable and steadily increasing rate of lock conflicts leading to delay and even deadlock/rollback/retry. The phenomenon causes overheads to rise as a polynomial in the number of servers over which you replicate the data, and quite sharply: I believe it was N^3 in the number of servers, and T^5 in the rate of transactions. So your single replicated database will have a perform collapse. With shards, using state machine replication (implemented using Derecho!) this isn't an issue, but of course we don't get the full SQL model at that point -- we end up with a form of NoSQL on the sharded database, similar to what MongoDB or Amazon's Dynamo DB offers.
Of course the "dangers" paper is iconic, but the techniques it uses are of broad value. And this was central to the way Jim approached problems: he was a huge fan in working out the critical paths and measuring costs along them. In his cloud database setup, a bit of fancy mathematics let the group he was working with turn that sort of thinking into a scalability analysis that led to a foundational insight. But even if you don't have an identical chance to change the world, it makes sense to try and follow a similar path.
This has had me thinking about paper-and-pencil analysis of the critical paths and potential consistency conflict points for large edge IoT deployments of the kind I described last week. Right now, those paths are pretty messy, if you approach it this way. Without an edge service, we would see something like this:
IoT IoT Function Micro
Sensor ---------------> Hub ---> Server ------> Service
In this example I am acting as if the function server "is" the function itself, and hiding the step in which the function server looks up the class of function that should handle this event, launches it (or perhaps had one waiting, warm-started), and then hands off the event data to the function for handling on one of its servers. Had I included this handoff the image would be more like this:
Even more concerning, many sensors can't connect directly to the cloud, and we end up cloning the architecture and running it twice: within an IoT Edge system (think of that as an operating system for a small NUMA machine or a cluster, running close to the sensors, and then relaying data to the main cloud if it can't handle the events out near the sensor device).
The Micro Service may actually be a whole graph of mutually supporting Micro Services, each running on a pool of nodes, and each interacting with some of the others. The cloud's "App Server" probably hosts these and provides elasticity if a backlog forms for one of them.
Fundamentally, what Jim and his colleagues did was to undertake a whiteboard analysis of the scalability of concurrency control in an uncontrolled situation, where transactions are simply submitted to some big pool of servers, and then compete for locks in accordance with a two-phase locking model (one in which a transaction acquires all its locks before releasing any), and then terminates using a two-phase or three-phase commit. They show that without some mechanism to prevent lock conflicts, there is a predictable and steadily increasing rate of lock conflicts leading to delay and even deadlock/rollback/retry. The phenomenon causes overheads to rise as a polynomial in the number of servers over which you replicate the data, and quite sharply: I believe it was N^3 in the number of servers, and T^5 in the rate of transactions. So your single replicated database will have a perform collapse. With shards, using state machine replication (implemented using Derecho!) this isn't an issue, but of course we don't get the full SQL model at that point -- we end up with a form of NoSQL on the sharded database, similar to what MongoDB or Amazon's Dynamo DB offers.
Of course the "dangers" paper is iconic, but the techniques it uses are of broad value. And this was central to the way Jim approached problems: he was a huge fan in working out the critical paths and measuring costs along them. In his cloud database setup, a bit of fancy mathematics let the group he was working with turn that sort of thinking into a scalability analysis that led to a foundational insight. But even if you don't have an identical chance to change the world, it makes sense to try and follow a similar path.
This has had me thinking about paper-and-pencil analysis of the critical paths and potential consistency conflict points for large edge IoT deployments of the kind I described last week. Right now, those paths are pretty messy, if you approach it this way. Without an edge service, we would see something like this:
IoT IoT Function Micro
Sensor ---------------> Hub ---> Server ------> Service
In this example I am acting as if the function server "is" the function itself, and hiding the step in which the function server looks up the class of function that should handle this event, launches it (or perhaps had one waiting, warm-started), and then hands off the event data to the function for handling on one of its servers. Had I included this handoff the image would be more like this:
IoT IoT Function Function Micro
Sensor ---------------> Hub ---> Server ------> F -----> Service
F is "your function", coded in a language like C#, F# or C++ or Python, and then encapsulated into a container of some form. You'll want to keep these programs very small and lightweight for speed. In particular, a function is not the place to do any serious computing, or to try and store anything. Real work occurs in the micro service, the one you built using Derecho. Even so, this particular step looks costly to me: without warm-starting it, launching F could take a substantial fraction of a section. And if F was warm-started, the context switch still involves some form of message passing, plus waking F up, and could still be many tens or even hundreds of milliseconds: an eternity at cloud speeds!
Even more concerning, many sensors can't connect directly to the cloud, and we end up cloning the architecture and running it twice: within an IoT Edge system (think of that as an operating system for a small NUMA machine or a cluster, running close to the sensors, and then relaying data to the main cloud if it can't handle the events out near the sensor device).
IoT Edge Edge Fcn IoT Function Micro
Sensor ---------------> Hub ---> Server -> F======> Hub ---> Server -> CF -> Service
Notice that now we have two user-supplied functions on the path. The first one will have decided that the event can't be handled out at the edge, and forwarded the request to the cloud, probably via a message queuing layer that I haven't actually shown, but represented using a double-arrow: ===>. This could have chosen to store the request and send it later, but with luck the link was up and it was passed to the cloud instantly, didn't need to sit in an arrival queue, and was instantly given to the cloud's IoT Hub, which in turn finally passed it to the cloud function server, the cloud function (CF) and the Micro Service.
The Micro Service may actually be a whole graph of mutually supporting Micro Services, each running on a pool of nodes, and each interacting with some of the others. The cloud's "App Server" probably hosts these and provides elasticity if a backlog forms for one of them.
We also have the difficulty that many sensors capture images and videos. These are initially stored on the device itself, which has substantial capacity but limited compute power. The big issue is that the first link, from sensor to the edge hub, would often be bandwidth limited. So we can't upload everything. Very likely what travels from sensor to hub is just a thumbnail and other meta-data. Then the edge function concludes that a download is needed (hopefully without too much delay), sends back a download request to the imaging device, and then the device moves the image to the cloud.
Moreover, there are industry standards for uploading photos and videos to a cloud, and those put the uploaded objects into the edge version of the blob store (short for "binary large objects"), which in turn is edge aware ands will mirror them to the main cloud blob store. Thus we have a whole pathway from IoT sensor to the edge blob server, which will eventually generate another event later to tell us that the data is ready. And as noted, for data that needs to reach the actual cloud and can't be processed at the edge, we replicate this path too, moving that image via the queuing service to the cloud.
So how long will all of this take? Latencies are high and bandwidth low for the first hop, because sensors rarely have great connectivity, and almost never have the higher levels of power required for really fast data transfers (even with 5G). So perhaps we will see a 10ms delay at that stop, plus more if the data is large. Inside the edge we should have a NUMA machine or perhaps a small cluster, and can safely assume 10G connections with latencies of 10us or less, although of course software like TCP will often impose its own delays. The big delay will probably be the handoff to the user-defined function, F.
My guess is that for an event that requires downloading a small photo, the very best performance will be something like 50ms before F sees the event (maybe even 100ms), then another 50-100 for F to request a download, then perhaps 200ms for the camera to upload the image to the blob server, and then a small delay (25ms?) for the blob server to trigger another event, F', saying "your image is ready!". We're up near 350ms and haven't done any work at all yet!
Because the function server is limited to lightweight computing, it hands off to our micro-service (a quick handoff because the service is already running; the main delay will be the binding action by which the function connects to it, and perhaps this can be done off the critical path). Call this 10ms? And then the micro service can decide what to do with this image.
Add another 75ms or so if we have to forward the request to the cloud. So the cloud might not be able to react to a photo in less than about 500ms, today.
None of this involved a Jim Gray kind of analysis of contention and backoff and retry. If you took my advice and used Derecho for any data replication, the 500ms might be the end of the story. But if you were to use a database solution like MongoDB (CosmosDB on Azure), it seems to me that you might easily see a further 250ms right there.
What should one do about these snowballing costs? One answer is that many of the early IoT applications just won't care: if the goal is to just journal that "Ken entered Gates Hall at 10am on Tuesday", a 1s delay isn't a big deal. But if the goal is to be reactive, we need to do a lot better.
I'm thinking that this is a great setting for various forms of shortcut datapaths, that could be set up after the first interaction and offer direct bypass options to move IoT events or data from the source directly to the real target. Then with RDMA in the cloud, and Derecho used to build your micro service, the 500ms could drop to perhaps 25 or 30ms, depending on the image size, and even less if the photo can be fully handled on the IoT Edge server itself.
On the other hand, if you don't use Derecho but you do need consistency, you'll get into trouble quickly: with scale (lots of these pipelines all running concurrently), and contention, it is easy to see how you could trigger Jim's "naive replication" concerns. So designers of smart highways had better beware: if they don't heed Jim's advice (and mine), by the time that smart highway warns that a car should "watch out for that reckless motorcycle approaching on your left!" it will already have zoomed past...
These are exciting times to work in computer systems. Of course a bit more funding wouldn't hurt, but we certainly will have our work cut out for us!
Saturday, 22 June 2019
Data everywhere but only a drop to drink...
One peculiarity of the IoT revolution is that it may explode the concept of big data.
The physical world is a domain of literally infinite data -- no matter how much we might hope to capture, at the very most we see only a tiny set of samples from an ocean of inaccessible information because we had no sensor in the proper place, or we didn't sample at the proper instant, or didn't have it pointing in the right direction or focused or ready to snap the photo, or we lacked bandwidth for the upload, or had no place to store the data and had to discard it, or misclassified it as "uninteresting" because the filters used to make those decisions weren't parameterized to sense the event the photo was showing.
Meanwhile, our data-hungry machine learning algorithms currently don't deal with the real world: they operate on snapshots, often ones collected ages ago. The puzzle will be to find a way to somehow compute on this incredible ocean of currently-inaccessible data while the data is still valuable: a real-time constraint. Time matters because in so many settings, conditions change extremely quickly (think of a smart highway, offering services to cars that are whizzing along at 85mph).
By computing at the back-end, AI/ML researchers have baked in very unrealistic assumptions, so that today's machine learning systems have become heavily skewed: they are very good at dealing with data acquired months ago and painstakingly tagged by an army of workers, and fairly good at using the resulting models to make decisions within a few tens of milliseconds, but in a sense consider the action of acquiring data and processing it in real-time to be part of the (offline) learning side of the game. In fact many existing systems wouldn't even work if they couldn't iterate for minutes (or longer) on data sets, and many need that data to be preprocessed in various ways, perhaps cleaned up, perhaps preloaded and cached in memory, so that a hardware accelerator can rip through the needed operations. If a smart highway were capturing data now that we would want to use to relearn vehicle trajectories so that we can react to changing conditions within fractions of a second, many aspects of this standard style of computing would have to change.
To me this points to a real problem for those intent on using machine learning everywhere and as soon as possible, but also a great research opportunity. Database and machine learning researchers need to begin to explore a new kind of system in which the data available to us is understood to be a "skim" (I learned this term when I used to work with high performance computing teams in scientific computing settings where data was getting big decades ago. For example the CERN particle accelerators capture far too much data to move data from the sensor, so even uploading "raw" data involves deciding which portions to keep, which to sample randomly, and which to completely ignore).
Beyond this issue of deciding what to include in the skim, there is the whole puzzle of supporting a dialog between the machine-learning infrastructure and the devices. I mentioned examples in which one need to predict that a photo of such and such a thing would be valuable, anticipate the timing, point the camera in the proper direction, pre-focus it (perhaps, on an expected object that isn't yet in the field of view, so that the auto-focus wouldn't be useful because the thing we want to image hasn't yet arrived), plan the timing, capture the image, and then process it -- all under real-time pressure.
I've always been fascinated by the emergence of new computing areas. To me this looks like one ripe for exploration. It wouldn't surprise me at all to see an ACM Symposium on this topic, or an ACM Transactions journal. Even at a glance one can see all the elements: a really interesting open problem that would lend itself to a theoretical formalization, but also one that will require substantial evolution of our platforms and computing systems. The area is clearly of high real-world importance and offers a real opportunity for impact, and a chance to build products. And it emerges at a juncture between systems and machine learning: a trending topic even now, so that this direction would play into gradually building momentum at the main funding agencies, which rarely can pivot on a dime, but are often good at following opportunities in a more incremental, thoughtful way.
The theoretical question would run roughly as follows. Suppose that I have a machine-learning system that lacks knowledge required to perform some task (this could be a decision or classification, or might involve some other goal, such as finding a path from A to B). The system has access to sensors, but there is a cost associated with using them (energy, repositioning, etc). Finally, we have some metric for data value: a hypothesis concerning the data we are missing that tells us how useful a particular sensor input would be. Then we can talk about the data to capture next that minimizes cost while maximizing value. Given a solution to the one-shot problem, we would then want to explore the continuous version, where the new data changes these model elements, fixed-points for problems that are static, and quality of tracking for cases where the underlying data is evolving.
The practical systems-infrastructure and O/S questions center on the capabilities of the hardware and the limitations of today's Linux-based operating system infrastructure, particularly in combination with existing offloaded compute accelerators (FPGA, TPU, GPU, even RDMA). Today's sensors run a gamut from really dumb fixed devices that don't even have storage to relatively smart sensors that can do various tasks on the device itself, have storage and some degree of intelligence about how to report data, etc. Future sensors might go further, with the ability to download logic and machine-learned models for making such decisions: I think it is very likely that we could program a device to point the camera at such and such a lane on the freeway, wait for a white vehicle moving at high speed that should arrive in the period [T0,T1], obtain a well-focused photo showing the license plate and current driver, and then report the image capture accompanied by a thumbnail. It might even be reasonable to talk about prefocusing, adjust the spectral parameters of the imaging system, selecting from a set of available lenses, etc.
Exploiting all of this will demand a new ecosystem that combines elements of machine learning on the cloud with elements of controlled logic on the sensing devices. If one thinks about the way that we refactor software, here we seem to be looking at a larger-scale refactoring in which the machine learning platform on the cloud, with "infinite storage and compute" resources, has the role of running the compute-heavy portions of the task, but where the sensors and the other elements of the solution (things like camera motion control, dynamic focus, etc) would need to participate in a cooperative way. Moreover, since we are dealing with entire IoT ecosystems, one has to visualize doing this at huge scale, with lots of sensors, lots of machine-learned models, and a shared infrastructure that imposes limits on communication bandwidth and latency, computing at the sensors, battery power, storage and so forth.
It would probably be wise to keep as much of the existing infrastructure as feasible. So perhaps that smart highway will need to compute "typical patterns" of traffic flow over a long time period with today's methodologies (no time pressure there), current vehicle trajectories over mid-term time periods using methods that work within a few seconds, and then can deal with instantaneous context (a car suddenly swerves to avoid a rock that just fell from a dumptruck onto the lane) as an ultra-urgent real-time learning task that splits into the instantaneous part ("watch out!") and the longer-term parts ("warning: obstacle in the road 0.5miles ahead, left lane") or even longer ("at mile 22, northbound, left lane, anticipate roadway debris"). This kind of hierarchy of temporality is missing in today's machine learning systems, as far as I can tell, and the more urgent forms of learning and reaction will require new tools. Yet we can preserve a lot of existing technology as we tackle these new tasks.
Data is everywhere... and that isn't going to change. It is about time that we tackle the challenge of building systems that can learn to discover context, and use current context to decide what to "look more closely" at, and with adequate time to carry out that task. This is a broad puzzle with room for everyone -- in fact you can't even consider tackling it without teams that include systems people like me as well as machine learning and vision researchers. What a great puzzle for the next generation of researchers!
The physical world is a domain of literally infinite data -- no matter how much we might hope to capture, at the very most we see only a tiny set of samples from an ocean of inaccessible information because we had no sensor in the proper place, or we didn't sample at the proper instant, or didn't have it pointing in the right direction or focused or ready to snap the photo, or we lacked bandwidth for the upload, or had no place to store the data and had to discard it, or misclassified it as "uninteresting" because the filters used to make those decisions weren't parameterized to sense the event the photo was showing.
Meanwhile, our data-hungry machine learning algorithms currently don't deal with the real world: they operate on snapshots, often ones collected ages ago. The puzzle will be to find a way to somehow compute on this incredible ocean of currently-inaccessible data while the data is still valuable: a real-time constraint. Time matters because in so many settings, conditions change extremely quickly (think of a smart highway, offering services to cars that are whizzing along at 85mph).
By computing at the back-end, AI/ML researchers have baked in very unrealistic assumptions, so that today's machine learning systems have become heavily skewed: they are very good at dealing with data acquired months ago and painstakingly tagged by an army of workers, and fairly good at using the resulting models to make decisions within a few tens of milliseconds, but in a sense consider the action of acquiring data and processing it in real-time to be part of the (offline) learning side of the game. In fact many existing systems wouldn't even work if they couldn't iterate for minutes (or longer) on data sets, and many need that data to be preprocessed in various ways, perhaps cleaned up, perhaps preloaded and cached in memory, so that a hardware accelerator can rip through the needed operations. If a smart highway were capturing data now that we would want to use to relearn vehicle trajectories so that we can react to changing conditions within fractions of a second, many aspects of this standard style of computing would have to change.
To me this points to a real problem for those intent on using machine learning everywhere and as soon as possible, but also a great research opportunity. Database and machine learning researchers need to begin to explore a new kind of system in which the data available to us is understood to be a "skim" (I learned this term when I used to work with high performance computing teams in scientific computing settings where data was getting big decades ago. For example the CERN particle accelerators capture far too much data to move data from the sensor, so even uploading "raw" data involves deciding which portions to keep, which to sample randomly, and which to completely ignore).
Beyond this issue of deciding what to include in the skim, there is the whole puzzle of supporting a dialog between the machine-learning infrastructure and the devices. I mentioned examples in which one need to predict that a photo of such and such a thing would be valuable, anticipate the timing, point the camera in the proper direction, pre-focus it (perhaps, on an expected object that isn't yet in the field of view, so that the auto-focus wouldn't be useful because the thing we want to image hasn't yet arrived), plan the timing, capture the image, and then process it -- all under real-time pressure.
I've always been fascinated by the emergence of new computing areas. To me this looks like one ripe for exploration. It wouldn't surprise me at all to see an ACM Symposium on this topic, or an ACM Transactions journal. Even at a glance one can see all the elements: a really interesting open problem that would lend itself to a theoretical formalization, but also one that will require substantial evolution of our platforms and computing systems. The area is clearly of high real-world importance and offers a real opportunity for impact, and a chance to build products. And it emerges at a juncture between systems and machine learning: a trending topic even now, so that this direction would play into gradually building momentum at the main funding agencies, which rarely can pivot on a dime, but are often good at following opportunities in a more incremental, thoughtful way.
The theoretical question would run roughly as follows. Suppose that I have a machine-learning system that lacks knowledge required to perform some task (this could be a decision or classification, or might involve some other goal, such as finding a path from A to B). The system has access to sensors, but there is a cost associated with using them (energy, repositioning, etc). Finally, we have some metric for data value: a hypothesis concerning the data we are missing that tells us how useful a particular sensor input would be. Then we can talk about the data to capture next that minimizes cost while maximizing value. Given a solution to the one-shot problem, we would then want to explore the continuous version, where the new data changes these model elements, fixed-points for problems that are static, and quality of tracking for cases where the underlying data is evolving.
The practical systems-infrastructure and O/S questions center on the capabilities of the hardware and the limitations of today's Linux-based operating system infrastructure, particularly in combination with existing offloaded compute accelerators (FPGA, TPU, GPU, even RDMA). Today's sensors run a gamut from really dumb fixed devices that don't even have storage to relatively smart sensors that can do various tasks on the device itself, have storage and some degree of intelligence about how to report data, etc. Future sensors might go further, with the ability to download logic and machine-learned models for making such decisions: I think it is very likely that we could program a device to point the camera at such and such a lane on the freeway, wait for a white vehicle moving at high speed that should arrive in the period [T0,T1], obtain a well-focused photo showing the license plate and current driver, and then report the image capture accompanied by a thumbnail. It might even be reasonable to talk about prefocusing, adjust the spectral parameters of the imaging system, selecting from a set of available lenses, etc.
Exploiting all of this will demand a new ecosystem that combines elements of machine learning on the cloud with elements of controlled logic on the sensing devices. If one thinks about the way that we refactor software, here we seem to be looking at a larger-scale refactoring in which the machine learning platform on the cloud, with "infinite storage and compute" resources, has the role of running the compute-heavy portions of the task, but where the sensors and the other elements of the solution (things like camera motion control, dynamic focus, etc) would need to participate in a cooperative way. Moreover, since we are dealing with entire IoT ecosystems, one has to visualize doing this at huge scale, with lots of sensors, lots of machine-learned models, and a shared infrastructure that imposes limits on communication bandwidth and latency, computing at the sensors, battery power, storage and so forth.
It would probably be wise to keep as much of the existing infrastructure as feasible. So perhaps that smart highway will need to compute "typical patterns" of traffic flow over a long time period with today's methodologies (no time pressure there), current vehicle trajectories over mid-term time periods using methods that work within a few seconds, and then can deal with instantaneous context (a car suddenly swerves to avoid a rock that just fell from a dumptruck onto the lane) as an ultra-urgent real-time learning task that splits into the instantaneous part ("watch out!") and the longer-term parts ("warning: obstacle in the road 0.5miles ahead, left lane") or even longer ("at mile 22, northbound, left lane, anticipate roadway debris"). This kind of hierarchy of temporality is missing in today's machine learning systems, as far as I can tell, and the more urgent forms of learning and reaction will require new tools. Yet we can preserve a lot of existing technology as we tackle these new tasks.
Data is everywhere... and that isn't going to change. It is about time that we tackle the challenge of building systems that can learn to discover context, and use current context to decide what to "look more closely" at, and with adequate time to carry out that task. This is a broad puzzle with room for everyone -- in fact you can't even consider tackling it without teams that include systems people like me as well as machine learning and vision researchers. What a great puzzle for the next generation of researchers!
Sunday, 12 May 2019
Redefining the IoT Edge
Edge computing has a dismal reputation. Although continuing miniaturization of computing elements has made it possible to put small ARM processors pretty much anywhere, general purpose tasks don’t make much sense in the edge. The most obvious reason is that no matter how powerful the processor could be, a mix of power, bandwidth and cost constraints argue against that model.
Beyond this, the interesting forms of machine learning and decision making can't possibly occur in an autonomous way. An edge sensor will have the data it captures directly and any configuration we might have pushed to it last night, but very little real-time context: if every sensor were trying to share its data with every other sensor that might be interested in that data, the resulting n^2 pattern would overwhelm even the beefiest ARM configuration. Yet exchanging smaller data summaries implies that each device will run with different mixes of detail.
This creates a computing model constrained by hard theoretical bounds. In papers written in the 1980's, Stoneybrook economics professor Pradeep Dubey studied the efficiency of game-theoretic multiparty optimization. His early results inspired follow-on research by Berkeley's Elias Koutsoupias and Christos Papadimitriou, and by my colleagues here at Cornell, Tim Roughgarten and Eva Tardos. The bottom line is unequivocal: there is a huge "price of anarchy.” In an optimization system where parties independently work towards an optimal state using non-identical data, even when they can find a Nash optimal configuration, that state can be far from the global optimal.
As a distributed protocols person who builds systems, one obvious idea would be to explore more efficient data exchange protocols for the edge: systems in which the sensors iteratively exchange subsets of data in a smarter way, using consensus to agree on the data so that they are all computing against the same inputs. There as been plenty of work on this, including some of mine. But little of it has been adopted or even deployed experimentally.
The core problem is that communication constraints make direct sensor to sensor data exchange difficult and slow. If a backlink to the cloud is available, it is almost always best to just use it. But if you do, you end up with an IoT cloud model, where data first is uploaded to the cloud, then some computed result is pushed back to the devices. The devices are no longer autonomously intelligent: they are basically peripherals of the cloud.
Optimization is at the heart of machine learning and artificial intelligence, and so all of these observations lead us towards a cloud-hosted model of IoT intelligence. Other options, for example ones in which brilliant sensors are deployed to implement a decentralized intelligent system, might enable yield collective behavior but that behavior will be suboptimal, and perhaps even unstable (or chaotic). I was once quite interested in swarm computing (it seemed like a natural outgrowth of gossip protocols, on which I was working at the time). Today, I've come to doubt that robot swarms or self-organizing convoys of smart cars can work, and if they can, that the quality of their decision-making could compete against cloud-hosted solutions.
In fact the cloud has all sorts of magical superpowers that enable it to perform operations inaccessible to the IoT sensors. Consider data fusion: with multiple overlapping cameras operated from different perspectives, we can reconstruct 3D scenes -- in effect, using the images to generate a 3D model and then painting the model with the captured data. But to do this we need lots of parallel computing and heavy processing on GPU devices. Even a swarm of brilliant sensors could never create such a fused scene given today’s communication and hardware options.
And yet, even though I believe in the remarkable power of the cloud, I'm also skeptical about an IoT model that presumes the sensors are dumb devices. Devices like cameras actually possess remarkable powers too, ones that no central system can mimic. For example, if preconfigured with some form of interest model, a smart sensor can classify images: data to upload, data to retain but report only as a thumbnail with associated metadata, and data to discard outright. A camera may be able to pivot so as to point the lens at an interesting location, or to focus in anticipation of some expected event, or to configure a multispectral image sensor. It can decide when to snap the photo, and which of several candidate images to retain (many of today's cameras take multiple images and some even do so with different depths of field or different focal points). Cameras can also do a wide range of on-device image preprocessing and compression. If we overlook these specialized capabilities, we end up with a very dumb IoT edge and a cloud unable to compensate for its limitations.
The future, then, actually will demand a form of edge computing -- but one that will center on a partnership between the cloud (or perhaps a cloud edge running on a platform near the sensor, as with Azure IoT Edge), working in close concert with the attached sensors to dynamically configure them, perhaps reconfigure them as conditions change, and even to pass them knowledge models computed on the cloud that they can use on-camera (or radar, lidar, microphone) to improve the quality of information captured. Each element has its unique capabilities and roles.
Even the IoT network is heading towards a more and more dynamic and reconfigurable model. If one sensor captures a huge and extremely interesting object, while others have nothing notable to report, it may make sense to reconfigure the WiFi network to dedicate a maximum of resources to that one WiFi link. Moments later, having pulled the video to the cloud edge, we might shift those same resources to a set of motion sensors that are watching an interesting pattern of activity, or to some other camera.
Perhaps we need a new term for this kind of edge computing, but my own instinct is to just coopt the existing term -- the bottom line is that the classic idea of edge computing hasn't really gone very far, and reviled or not, is best "known" to people who aren't even active in the field today. The next generation of edge computing will be done by a new generation of researchers and product developers, and they might as well benefit from the name recognition -- I think they can brush off the negative associations fairly easily, given that edge computing never actually took off and then collapsed, or had any kind of extensive coverage in the commercial press.
The resulting research agenda is an exciting one. We will need to develop models for computing that single globally optimal knowledge state, yet for also "compiling" elements of it to be executed remotely. We'll need to understand how to treat physical-world actions like pivoting and focusing as elements of an otherwise Van Neuman computational framework, and to include the possibility of capturing new data side by side with the possibility of iterating a stochastic gradient descent one more time. There are questions of long term knowledge (which we can compute on the back-end cloud using today's existing batched solutions), but also contextual knowledge that must be acquired on the fly, and then physical world "knowledge" such as a motion detection that might be used to trigger a camera to acquire an image. The problem poses open questions at every level: the machine learning infrastructure, the systems infrastructure on which it runs, and the devices themselves -- not brilliant and autonomous, but not dumb either. As the area matures and we gain some degree of standardization around platforms and approaches, the potential seems enormous!
So next time you teach a class on IoT and mention exciting ideas like smart highways that might sell access to high speed lanes or other services to drivers or semi-autonomous cars, pause to point out that this kind of setting is a perfect example of a future computing capability that will soon supplant past ideas of edge computing. Teach your students to think of robotic actions like pivoting a camera, or focusing it, or even configuring it to select interesting images, as one facet of a rich and complex notion of edge computing that can take us into settings inaccessible to the classical cloud, and yet equally inaccessible even to the most brilliant of autonomous sensors. Tell them about those theoretical insights: it is very hard to engineer around an impossibility proof, and if this implies that swarm computing simply won't be the winner, let them think about the implications. You'll be helping them prepare to be leaders in tomorrow's big new thing!
Beyond this, the interesting forms of machine learning and decision making can't possibly occur in an autonomous way. An edge sensor will have the data it captures directly and any configuration we might have pushed to it last night, but very little real-time context: if every sensor were trying to share its data with every other sensor that might be interested in that data, the resulting n^2 pattern would overwhelm even the beefiest ARM configuration. Yet exchanging smaller data summaries implies that each device will run with different mixes of detail.
This creates a computing model constrained by hard theoretical bounds. In papers written in the 1980's, Stoneybrook economics professor Pradeep Dubey studied the efficiency of game-theoretic multiparty optimization. His early results inspired follow-on research by Berkeley's Elias Koutsoupias and Christos Papadimitriou, and by my colleagues here at Cornell, Tim Roughgarten and Eva Tardos. The bottom line is unequivocal: there is a huge "price of anarchy.” In an optimization system where parties independently work towards an optimal state using non-identical data, even when they can find a Nash optimal configuration, that state can be far from the global optimal.
As a distributed protocols person who builds systems, one obvious idea would be to explore more efficient data exchange protocols for the edge: systems in which the sensors iteratively exchange subsets of data in a smarter way, using consensus to agree on the data so that they are all computing against the same inputs. There as been plenty of work on this, including some of mine. But little of it has been adopted or even deployed experimentally.
The core problem is that communication constraints make direct sensor to sensor data exchange difficult and slow. If a backlink to the cloud is available, it is almost always best to just use it. But if you do, you end up with an IoT cloud model, where data first is uploaded to the cloud, then some computed result is pushed back to the devices. The devices are no longer autonomously intelligent: they are basically peripherals of the cloud.
Optimization is at the heart of machine learning and artificial intelligence, and so all of these observations lead us towards a cloud-hosted model of IoT intelligence. Other options, for example ones in which brilliant sensors are deployed to implement a decentralized intelligent system, might enable yield collective behavior but that behavior will be suboptimal, and perhaps even unstable (or chaotic). I was once quite interested in swarm computing (it seemed like a natural outgrowth of gossip protocols, on which I was working at the time). Today, I've come to doubt that robot swarms or self-organizing convoys of smart cars can work, and if they can, that the quality of their decision-making could compete against cloud-hosted solutions.
In fact the cloud has all sorts of magical superpowers that enable it to perform operations inaccessible to the IoT sensors. Consider data fusion: with multiple overlapping cameras operated from different perspectives, we can reconstruct 3D scenes -- in effect, using the images to generate a 3D model and then painting the model with the captured data. But to do this we need lots of parallel computing and heavy processing on GPU devices. Even a swarm of brilliant sensors could never create such a fused scene given today’s communication and hardware options.
And yet, even though I believe in the remarkable power of the cloud, I'm also skeptical about an IoT model that presumes the sensors are dumb devices. Devices like cameras actually possess remarkable powers too, ones that no central system can mimic. For example, if preconfigured with some form of interest model, a smart sensor can classify images: data to upload, data to retain but report only as a thumbnail with associated metadata, and data to discard outright. A camera may be able to pivot so as to point the lens at an interesting location, or to focus in anticipation of some expected event, or to configure a multispectral image sensor. It can decide when to snap the photo, and which of several candidate images to retain (many of today's cameras take multiple images and some even do so with different depths of field or different focal points). Cameras can also do a wide range of on-device image preprocessing and compression. If we overlook these specialized capabilities, we end up with a very dumb IoT edge and a cloud unable to compensate for its limitations.
The future, then, actually will demand a form of edge computing -- but one that will center on a partnership between the cloud (or perhaps a cloud edge running on a platform near the sensor, as with Azure IoT Edge), working in close concert with the attached sensors to dynamically configure them, perhaps reconfigure them as conditions change, and even to pass them knowledge models computed on the cloud that they can use on-camera (or radar, lidar, microphone) to improve the quality of information captured. Each element has its unique capabilities and roles.
Even the IoT network is heading towards a more and more dynamic and reconfigurable model. If one sensor captures a huge and extremely interesting object, while others have nothing notable to report, it may make sense to reconfigure the WiFi network to dedicate a maximum of resources to that one WiFi link. Moments later, having pulled the video to the cloud edge, we might shift those same resources to a set of motion sensors that are watching an interesting pattern of activity, or to some other camera.
Perhaps we need a new term for this kind of edge computing, but my own instinct is to just coopt the existing term -- the bottom line is that the classic idea of edge computing hasn't really gone very far, and reviled or not, is best "known" to people who aren't even active in the field today. The next generation of edge computing will be done by a new generation of researchers and product developers, and they might as well benefit from the name recognition -- I think they can brush off the negative associations fairly easily, given that edge computing never actually took off and then collapsed, or had any kind of extensive coverage in the commercial press.
The resulting research agenda is an exciting one. We will need to develop models for computing that single globally optimal knowledge state, yet for also "compiling" elements of it to be executed remotely. We'll need to understand how to treat physical-world actions like pivoting and focusing as elements of an otherwise Van Neuman computational framework, and to include the possibility of capturing new data side by side with the possibility of iterating a stochastic gradient descent one more time. There are questions of long term knowledge (which we can compute on the back-end cloud using today's existing batched solutions), but also contextual knowledge that must be acquired on the fly, and then physical world "knowledge" such as a motion detection that might be used to trigger a camera to acquire an image. The problem poses open questions at every level: the machine learning infrastructure, the systems infrastructure on which it runs, and the devices themselves -- not brilliant and autonomous, but not dumb either. As the area matures and we gain some degree of standardization around platforms and approaches, the potential seems enormous!
So next time you teach a class on IoT and mention exciting ideas like smart highways that might sell access to high speed lanes or other services to drivers or semi-autonomous cars, pause to point out that this kind of setting is a perfect example of a future computing capability that will soon supplant past ideas of edge computing. Teach your students to think of robotic actions like pivoting a camera, or focusing it, or even configuring it to select interesting images, as one facet of a rich and complex notion of edge computing that can take us into settings inaccessible to the classical cloud, and yet equally inaccessible even to the most brilliant of autonomous sensors. Tell them about those theoretical insights: it is very hard to engineer around an impossibility proof, and if this implies that swarm computing simply won't be the winner, let them think about the implications. You'll be helping them prepare to be leaders in tomorrow's big new thing!
Wednesday, 13 March 2019
Intelligent IoT Services: Generic, or Bespoke?
I've been fascinated by
a puzzle that will probably play out over several years. It involves a
deep transformation of the cloud computing marketplace, centered on a choice.
In one case, IoT infrastructures will be built the way we currently build
web services that do things like intelligent recommendations or ad placements.
In the other, edge IoT will require a "new" way of developing
solutions that centers on creating new and specialized services... ones that
embody real-time logic for making decisions or even learning in real-time.
I'm going to make a case
for bespoke, handbuilt, services: the second scenario. But if I’m right,
there is hard work to be done and whoever starts first will gain a major
advantage.
So to set the stage, let
me outline the way IoT applications work today in the cloud. We have
devices deployed in some enterprise setting, perhaps a factory, or an apartment
complex, or an office building. These might be quite dumb, but they are
still network enabled: they could be things like temperature and humidity
sensors, motion detectors, microphones or cameras, etc. Because many are
dumb, even the smart ones (like cameras and videos with built-in autofocus,
deblurring, depth perception) are treated in a sort of rigid manner: the basic
model is of a device with a limited API that can be configured, and perhaps can
be patched if the firmware has issues, but then generates simple events with
meta-data that describes what happens.
In a posting a few weeks
ago, I noted that unmanaged IoT deployments are terrifying for system
administrators, so the world is rapidly shifting towards migrating IoT device
management into systems like Azure's infrastructure for Office 365.
Basically, if my company already uses Office for other workplace tasks, it
makes sense to also manage these useful (but potentially dangerous) devices
through the same system.
Azure's IoT Hub handles
that managerial role: secure connectivity to the sensors, patches guaranteed to
be pushed as soon as feasible... and in the limit, maybe nothing else. But why
stop there? My point a few weeks back was simply that even just managing
enterprise IoT will leave Azure in a position of managing immense numbers of
devices -- and hence, in a position to leverage the devices by bringing new
value to the table.
Next observation: this
will be an "app" market, not a "platform" market. In
this blog I don't often draw on marketing studies and the like, but for the
particular case, it makes sense to point to market studies that explain my
thinking (look at Lecture 28 in my CS5412 cloud computing class to
see charts from the studies I drew on).
Cloud computing, perhaps
far more than most areas of systems, is shaped by the way cloud customers
actually want to use the infrastructure. In contrast, an area like
databases or big data is about how people want to use the data, which shapes
access patterns. But they aren't trying to explicitly route their data
through FPGA devices that will transform it in some way, or doing computations
that can't keep up unless they run in GPU clusters. So, because my kind
of cloud customers migrate to the clouds that make it easier to build their
applications, they will favor the cloud that has the best support for IoT apps.
A platform story
basically offers minimal functionality, like bare metal running Linux, and
leaves the developers to do the rest. They are welcome to connect to
services but not required to do so. Sometimes this is called the hybrid
cloud.
Now, what's an
app? As I'm using the term, you would want to visualize the iPhone or
Android app store: small programs that share many common infrastructure components
(the GUI framework, the storage framework, the motion sensor and touch sensors,
etc), and then that connect to their bigger cloud-hosted servers over a Web
Services layer that tends to match nicely with the old Apache-dominated cloud
for doing highly concurrent construction of web pages. So this is the
intuition.
For IoT, though, an app
model wouldn't work in the same way -- in fact, it can't work in the
same way. First, IoT devices that want help from intelligent
machine-learning will often need support from something that learns in
real-time. In contrast, today's web architecture is all about learning
yesterday and then serving up read-only data at ultra-fast rates from scalable
caching layers that could easily be stale if the data was actually changing
rapidly. So suddenly we will need to do machine learning, decision making
and classification, and a host of other performance-intensive tasks at the
edge, under time pressure, and with data changing quite rapidly. Just
think of a service that guides a drone surveying a farming area that wants to
optimize its search strategy to "sail on the wind" and you'll be
thinking about the right issues.
Will the market want
platforms, or apps? I think the market data strongly suggests that apps
are winning. Their relatively turnkey development advantages outweigh the
limitations of programming in a somewhat constrained way. If you do look
at the slides from my course, you can see how this trend is playing out.
The big money is in apps.
And now we get to my
real puzzle. If I'm going to be creating intelligent infrastructure for
these rather limited IoT devices (limited by power, and by compute cycles, and
by bandwidth), where should the intelligence live? Not on the
devices: we just bolted them down to a point where they probably wouldn't have
the capacity. Anyhow, they lack the big picture: if 10 drones are flying
around, the cloud can build a wind map for the whole farm. But any single
drone wouldn't have enough context to create that situational picture, or to
optimize the flight plan properly. There is even a famous theoretical
result on the "cost of anarchy", showing that you don't get the
global optimum if you have a lot of autonomous agents making individually
optimal choices. No, you want the intelligence to reside in the cloud.
But where?
Today, machine
intelligence lives at the back, but the delays are too large. We can’t
control today’s drones with yesterday’s wind patterns. We need
intelligence right at the edge!
Azure and AWS both
access their IoT devices through a function layer ("lambdas" in the
case of AWS). This is an elastic service that hosts containers, launching
as many instances of your program as needed on the basis of events.
Functions of this kind are genuine programs and can do anything they need to
do, but they run what is called a "stateless" mode, meaning that they
flash into existence (or are even warm-started ahead of time, so that when the
event arrives, the delay is minimal). Then they handle the event, but they
can't save any permanent data locally, even though the container does have a
small file system that works perfectly well: as soon as the event handling
ends, the container will garbage collect itself and that local file system will
evaporate.
So, the intelligence and
knowledge and learning has to occur in a bank of servers. One scenario,
call it the PaaS mode, would be that Amazon and Microsoft pre-build a set of
very general purpose AI/ML services, and we code all our solutions by
parameterizing those and mapping everything into them. So here you have
AI-as-a-service. Seems like a guaranteed $B startup concept! But
very honestly, I'm not seeing how it can work. The machine learning you
would do to learn wind patterns and direct drones to sail on the wind is just
too different from what you need to recognize wheat blight, or to figure out
what insect is eating the corn.
The other scenario is
the "bespoke" one. My Derecho library could be useful
here. With a bespoke service, you take some tools like Derecho and build
a little cluster-hosted service of your very own, which you then tell the cloud
to host on your behalf. Then your functions or lambdas can talk to your
services, so that if an IoT event requires a decision, the path from device to
intelligence is just milliseconds. With consistent data replication, we
can even eliminate stale data issues: these services would learn as they go (or
at least, they could), and then use their most recent models to handle each new
stage of decision-making.
But without far better
tools, it will be quite annoying to create these bespoke services, and this, I
think, is the big risk to the current IoT edge opportunity: do Microsoft and
Amazon actually understand this need, and will they enlarge the coverage of
VSCode or Visual Studio or in Amazon's case, Cloud9, to "automate" as
many aspects of service creation as possible, while still leaving flexibility
for the machine learning developer to introduce the wide range of
customizations that her service might require?
What are these
automation opportunities? Some are pretty basic (but that doesn't mean
they are easy to do by hand)! To actually launch a service on a cloud,
there needs to be a control file created, typically in a JSON format, with
various fields taking on the requisite values. Often, these include
magically generated 60-hexidecimal-digit keys or other kinds of unintuitive
content. When you use these tools to create other kinds of cloud
solutions, they automate those steps. By hand, I promise that you’ll
spend an afternoon and feel pretty annoyed by the waste of your time. A
good hour will be lost on those stupid registry keys alone.
Interface definitions
are a need too. If we want functions and lambdas talking to our new
bespoke micro-services ("micro" to underscore that these aren't the
big vendor-supplied ones, like CosmosDB), the new micro-service needs to export
an interface that the lambda or function can call at runtime. Again, help
needed!
In fact the list is
surprisingly long, even though the items on it are (objectively) trivial.
The real point isn’t that these are hard to do, but rather that they are arcane
and require looking for the proper documentation, following some sort of magic
incantation, figuring out where to install the script or file, testing your
edited version of the example they give, etc. Here are a few examples:
- Launch service
- Authenticate if needed
- Register micro/service to accept RPCs
- There should be an easy way to create functions able to call the service, using those RPC APIs
- We need an efficient upload path for image objects
- There will need to be tools for garbage collection (and tools to track space use)
- … and tools for managing the collection of configuration parameter files and settings for an entire application
- .… and lifecycle tools, for pushing patches and configuration changes in a clean way.
Then there are some more
substantial needs:
- Code debugging support for issues missed in development and then arising at runtime
- Performance monitoring, hotspot visualization and performance optimization (or even, performance debugging) tools
- Ways to enable a trusted micro-service to make use of hardware accelerators like RDMA or FGPA even if the end user might not be trusted to safely to so (many accelerators save money and improve performance but are just not suitable for direct access by hordes of developers with limited skill sets. Some could destabilize the data center or crash nodes, and some might have security vulnerabilities.
This makes for a long
list, but in my view, a strong development team at Amazon or Microsoft, perhaps
allied with a strong research group to tackle the open ended tasks, could
certainly succeed. Success would open the door to mature intelligent edge
IoT. Lacking such tools, though, it is hard not to see edge IoT as being
pretty immature today: huge promise, but more substance is needed.
My bet? Well,
companies like Microsoft need periodic challenges to set in front of their
research teams. I remember that when I visited MSR Cambridge back in
2016, everyone was asking what they should be doing as researchers to enable
the next steps for the product teams... the capacity is there. And those
market slides I mentioned make it clear: The edge is a huge potential
market. So I think the pieces are in place, and that we should jump on
the IoT edge bandwagon (in some cases, “yet again”). This time, it may
really happen!
Subscribe to:
Posts (Atom)