This week’s podcast is about AgiBot, a fast rising full stack robot business in Shanghai.
You can listen to this podcast here, which has the slides and graphics mentioned. Also available at iTunes and Google Podcasts.
Here is the link to the TechMoat Consulting.
Here is the link to our Tech Tours.
Why I like AgiBot:
- Has rapidly expanded into a full suite of robots
- Going after all the humanoid industries and use cases
- Building an almost full AI tech stack. Everything except chips.
- Founded and run by Huawei executives. Moving very fast.
- A good strategy. It appears similar to Huawei.
- AgiBot’s strategy is to flood the market with affordable humanoids while growing the entire ecosystem with open datasets and open models. They are scaling fast in manufacturing, usage and training.
- They mass-produce the hardware at scale so thousands of identical robots can collect real-world data in homes and factories.
- At the same time they open-source the entire stack: the million-trajectory AgiBot World dataset, the GO-1 foundation model, the training tools, and the GenieSim simulation twin—so researchers and developers can build skills and applications on top without paying anything.
- AgiBot’s affordable humanoid platform as infrastructure. Hardware sales and ecosystem services become the revenue, not software licensing.
- Now in the top 1-2 in humanoid sales.
- AgiBot shipped approximately 5168 humanoid units in 2025.
———-
Related articles:
- More High-Tech Flex by Huawei R&D (4 of 4) (Tech Strategy)
- 6 Big Events in AI Agentic Ecommerce (Tech Strategy)
From the Concept Library, concepts for this article are:
- VLA, ViLLA Models
- World Models
From the Company Library, companies for this article are:
- AgiBot
—–transcription below
00:05
Welcome welcome everybody. My name is Jeff Towson and this is the tech strategy podcast from Techmoat Consulting and the topic for today why I really like Agibot Robotics, it’s not technically called Agibot. It’s just Agibot. There’s no robotics. added that uh Although I think it’s I think it’s called Agibot innovation. Maybe it was if you didn’t know it’s a robotics company It’s a robotics company if you’re in China.
00:35
But anyways, very interesting company. I’m not saying it’s an investment, it’s private, so you can’t do that anyways. in terms of understanding where robotics is, where it’s going, these are still early days, so who knows who’s going to win. I mean, if there’s one company I would study right now, this would be the one I’d study, which is really what I’m doing. I think you can kind of see what the future’s going to look like. You can see the business strategy. You can see the tech at literally every level of the tech.
01:03
You can see all the industry scenarios. You can see all the robot types. It’s a really good company to study for getting sort of deep into robotics and what it’s probably going to look like. Anyways, I’m going to talk about them, why I think they’re really interesting, what I’ve sort of learned. I’m spending a lot of time looking at this company. What I’ve kind of learned. Anyways, that’ll be the topic for today.
01:27
Standard disclaimer, nothing in this podcast or my writing website is investment advice. The numbers and information from me and any guests may be incorrect. The views and opinions expressed may no longer be relevant or accurate. Overall, investing is risky. This is not investment, legal or tax advice. Do your own research. And with that, let’s get into the topic. Now let’s see, concepts for today. There’s quite a few actually, but the big two, world models.
01:57
There’s a whole different type of foundation model. mean, many people argue this is kind of the future that building foundation models based on text, you know, that’s incredibly powerful for things like mathematics and coding and speech and, you know, writing. But that’s not really how human beings engage with the world. We do it by walking around and seeing things, vision, language as well, which is, you know,
02:24
what’s in our heads, we’re always thinking in terms of language and then action, we touch things. yeah, those are, now that’s a VLA model, but we’ll talk about world models and VLA models, which kind of go hand in hand. And there’s a couple of variants on VLA that are worth thinking about, VLMs, visual language models. VLA is vision language action. VLM is vision language model. And then there’s what they call Vila models, which is Vila language.
02:52
Latent action models, which is what Agibot uses and we’ll talk about both those world models and let’s just say VLA models for short those go hand in hand and you that’s kind of what these things run on now at the core They’re still transformers at the core They’re just putting them together in different structures and instead of just being text its text plus image and to some degree video as inputs so As far as I understand
03:21
which I’m still learning, the fundamental insight of Transformers still the core of these things, but we’re putting them together in different ways. That’s VLA models, world models, and that sort of thing. And that’s the big two topics for today that I’ll get into those in some degree, but let’s start it. Let’s stay at talking about the company. Then we’ll talk about the business strategy. And then to some extent, I’ll get into the technical stuff, which I think is super important, but it’s kind of a lot. I’ve been
03:50
studying this for like weeks now. So we’ll go a certain depth but not beyond that. Okay, so let’s kind of start at the beginning here, which is, uh okay, why do I like this company? Now, what do they do? This was when I was at the World AI Conference last week in Shanghai.
04:13
I mean, I probably read more about this company than any of them. And I went through a ton of companies. I mean, I’ve just got notebooks full of notes and I’ve spent in the week doing this. I was a little sick too, so that helped. I just sat in bed and read company reports on these companies, most of which are not public, but a couple of them are. Now on that list, why do I think they’re interesting? Because I think they’re playing across the entire board. I think they’re doing all the robot types.
04:44
humanoids, which can be bipedal walking, or they can be on wheels. That’s a pretty great thing. You see that in industrial scenarios a lot. They can be these robot dogs. People call those quadrupeds. I just call them robot dogs. They can be in these cleaning robots that you see everywhere. They can be in these hotel delivery robots that deliver your food to your rooms. I mean, they’re playing the whole thing.
05:11
Well, they’re not doing drones, so they’re not in the air, so I guess not everything. They’re doing all the industries and all the scenarios, industrial, logistics, warehouses, retail solutions. These are retail cubes. They’re popping up everywhere in the home, care environments. I’ll go into them in a bit, but they’re doing the whole list. And they’re building pretty much the whole AI tech stack. They’re not doing semiconductors.
05:41
Really system on a chip. mean, most of these companies are putting Nvidia chips in a little system with, you know, specialized CPUs and some memory, but they’re buying those mostly from Nvidia and some other places. OK, they’re not doing that, although I’m not sure some of might be, but Agibot’s not. And when we look at the three players that are sort of biggest in terms of number of robots being manufactured and shipped,
06:11
Well, that’s Unitree, they’re number one. It’s Agibot number two. It’s UBtech number three. And then you could also say, you know, there’s other main players like Tesla, obviously is a big one, but they’re not shipping robots yet. If you look at who’s shipping, it’s those three. And if you look at all of them, they’re all kind of playing a version of that, but their architecture, the AI tech architecture is a bit different company by company.
06:41
Agibot is doing a different architecture than VLA models, world models, things like that, then Unitree, and then UBtech. So you see an interesting sort of playbook in there. Now I like the Agibot approach to the architecture. I’ll explain why. But if you look at who’s founding these companies, a lot of them are professors, startups, these standalone companies. I’m not talking about Tencent.
07:09
They’re doing robot brains. I’m not talking about Alibaba. They’re doing robot. I’m talking about stand-aligned robot companies. ah Agibot founded by Huawei executives, senior vice presidents. I’ll give you their background and then some young guy. So their team strikes me as the most business savvy. And when I look at what they’re doing, it looks a lot to me like Huawei strategy. I’ll go into that.
07:38
Anyone who’s a senior vice president of Huawei, these people can execute at lightning speed. These are seasoned executives at deploying tech into the world. They’re very good at so I’ll go through it. And as mentioned, I think they have a good strategy. It looks a lot to me like a Huawei strategy to tell you the truth. And what else? Yeah. And then if you look at actual sales, I’ll give you some numbers. Unitree is number one in terms of
08:06
just robots shipped. In 2025, they got about 25,000 robots shipped. Now, when you break that down, okay, it looks like most of those are Robodogs. They ship a lot of these robots doing flips and things, and I don’t actually pay much attention to the Robodogs, because I don’t think they have use cases that are super interesting for businesses. They can do inspections.
08:33
They can zip around your factory or their utility plant and with their cameras or sniffers and they can sense things and do gather data. Those are useful, but I think that’s a very narrow sort of band. When you look at humanoids being shipped by unitree, it’s only four or five thousand. It’s in that range.
08:56
And it looks like Agibot’s number one for 2025. They’re probably 5000 plus, 6000 range. Hard to know exactly. But for humanoids shipped globally, know, Elon Musk is not building robo dogs. He’s building a humanoid. Well, humanoids shipped, looks like Agibot’s number one or right up their unitary is them to. Unitree probably a thousand. Tesla. I don’t know.
09:22
They haven’t sold any. The other players are in the hundreds and there’s several to watch. These are early days, so we’ll see. But that’s kind of why that’s on my. That’s my short list of why I like this company. I understand their business strategy and that’s something I like, because the other companies I don’t quite get what they’re doing. And a lot of what companies are doing is kind of based on what they can do. People are building what they have the technology and the capabilities to do.
09:51
That’s early stage. But at least with this company, I can tell what they’re building long term. Now, maybe the others will converge to that strategy as well. Who knows? OK, so let’s go through the history a little bit. This company is lightning fast. Founded in 2023, that’s what, two years and four months ago? That’s really fast. And
10:16
The two people, here’s the person everybody talks about. There’s a young man named Peng Zhihui. He’s the CTO, he’s the co-founder, he’s well known, he’s a young guy. People are sort of month by month, well maybe six months by six months. The rest of the world is sort of getting familiar with young Chinese founders they never heard about before. This month it’s been Kimmy founder.
10:45
Yang Zhilin are now sort of waking up to oh this young guy is really smart Okay Peng Zhihui. He’s going to be one of the next ones He’s CTO. He’s well known in China. He’s got a nickname called Ji Hui Jun uh Which is like wild Iron Man. I don’t quite understand why that’s the translation, but that’s he’s basically a young hardware innovator with the big
11:14
media presence, millions of followers online. He graduates, let me get the actual year. He graduates in 2018. uh From actually this is an interesting university, you should be aware of University of Electronic Science and Technology in China. I believe that’s in Chengdu. uh It’s well known. It’s for
11:39
consumer electronics, AI, and other things. He graduates his graduate studies 2018. He joins OPPO. He basically does their AI. I don’t think he was there very long. He basically gets picked up by Huawei fairly quickly under what they call their Young Genius program, or the Genius Youth Recruitment program. I actually wrote about this back in 2019 when Huawei was pivoting post Entity List Band.
12:10
They were moving from a fast follower policy, especially within their HR, to going after geniuses. And they started hiring Nobel laureates. They started putting them in their Dongguan facility, the Europe town, which is beautiful. Their young genius program, well, they picked up this young guy in July 2020.
12:37
So they kind of recognized him through this rigorous interview process and you know that this guy was something special. He’s doing video walkthroughs online of crazy stuff. Do it yourself engineering builds where he’s got a self balancing self driving bicycle. He basically put LIDAR on a bicycle and it could drive itself around with no rider. uh Minion robotic arms.
13:08
that can stitch, you know, with like sutures, the skins of grapes, which will be useful in surgery, obviously, mini computers, handheld gaming devices. Anyways, he gets millions of followers and lots of media coverage as like, dude, this guy’s creating everything. Huawei picks him up. He goes into there. Where does he go? He works in the core AI algorithm department underneath
13:37
They’re Ascend computing departments. Now they’re Ascend chips, the 910C. This is what Huawei is using to build these data centers that can match Nvidia performance with no Nvidia chips. That’s all about the Ascend platform. They have two sort of chip platforms, Kunpeng and Ascend. Kunpeng is CPUs, Ascend is GPUs. That’s their AI mothership. He gets hired as an algorithm engineer within Ascend.
14:07
Anyways, he quits 2023. Actually, he quit end of 2022, like November two or three months later, he founds Agibot 2023 February. He’s the guy everyone talks about. However, there’s another guy. There’s a guy named Deng Taihua. He’s the other co-founder. He’s the CEO. He’s the business head. This guy.
14:37
was a well, one, he’s a senior vice president of Huawei. He’s been there 20 years. To rise up and be senior vice president of Huawei, you have to be wicked smart and you have to be able to execute like crazy. What does he end up doing there after 20 years? He was in their wireless network product line. He was doing global commercial deployment of base stations, 4G.
15:06
5G, a lot of stuff. But he ends up being president of the computing product line, including the Ascend chips. So the other guy’s working for him. He’s head. He leaves and quits Huawei as well, founds his company, those two dudes. That’s really impressive. He’s basically, when they talk about, look, Ascend the computing platform.
15:32
is basically becoming the go-to GPU for China. They got 50 % of the market, which they took from Nvidia after they got kind of locked out. They’re building all this. That’s his group. Okay, that’s who’s running this. That’s what, you know, on paper, that’s really impressive. And then you think, okay, well, on paper is one thing. Look how fast they’ve been moving in two years and four months. They are really moving.
16:00
Unitree has been around for a decade. You know, they’ve been doing robot dogs from 2012, 2013, something like that. You know, their dogs are very impressive, but okay, who else is moving at this speed? Maybe only Elon Musk. I don’t know. I don’t know the Western. I don’t know this space that well. I know a couple of players. I’m digging into it as much as possible. If you look at companies that are not doing the full tech stack plus the hardware, the robots,
16:29
Okay, you can find companies that are pretty good just doing the AI brain. Tencent is impressive. Alibaba is impressive. But everything, the robots and the brain? I don’t know. On my short list now, which I’m still learning, it’s Tesla and it’s Agibot. But Unitree is doing well. don’t want to, and UBtech is doing, there’s other players and they’re selling in the hundreds. So there’s other players, it’s early days, we’ll see. But right now, this is who I got my eye on. Okay. uh
16:58
Let’s talk about what they did in the last two years. They launched February, 2023 in Shanghai. They’re Shanghai based. They’re out in Pudong for those of you who are familiar with the region. They do their manufacturing, which is all in-house at this point of the robots. Those are down near the deep water port, which is sort of south. That’s in the Lingan area, the deep water port south of Pudong airport.
17:26
But if you ever go from Pudong Airport down to Lujiazui, about halfway between those two is the high tech park. They’re in the high tech park. That’s where they are. OK, so they launch in 2023, February, August. They basically show their prototype working humanoid. They call it the A1, the Raise A1. They’ve got. Robots everywhere now, but they started with just the humanoid.
17:55
And they could, you know, it could walk and it could lift things up with its hands a little bit. That’s really what they’re doing. Phase two, which starts, we’ll call that 2024, really. They really, they really moved in 2024. The first thing they did, March, they released something very important called the Genie Operator. Now this is the AI tech stack architecture that’s in the robot.
18:24
And these robots can work without any wireless. It doesn’t have to access the cloud to operate. These things are sort of self-autonomous once you download the AI in there. They’re trained elsewhere, they’re downloaded. This is their sort of core genie operator that runs the robot, which I’ll go into this. I probably won’t get to this in this part. I’ll have to break it into two. The way you can think about the AI and the robot, there’s uh a cerebrum and there’s a cerebellum.
18:52
The cerebellum is sort of the low level uh rapid fire, high frequency signals being sent out to move the hands, to stabilize the hands, to release the pressure, to sort of low level compute simpler chips, not using a lot of energy. Then if you move up to the cerebrum, okay, that’s the AI in the robot. That’s doing a lot of visual processing.
19:18
It’s making up the plan. sending the orders to the cerebrum for what to be doing. So much more compute intensive and definitely energy much more intensive, but the frequency is actually quite low or not low, but it’s lower. So that’s go the I’ll say go, but it’s Genie operator. You know, that’s a big deal. And as I’ll talk about later, they made it open source. Everyone can use it. That’s a that’s a business decision. That’s a pretty important one.
19:48
Okay, a couple months later, August, they go from their single prototype of a humanoid
19:56
They have like the full-size humanoids. They have more compact, smaller humanoids, which are kind of research units. They have industrial humanoids, which are much more rugged, stronger, durable. They have the quadrupeds, the robo dogs. They have sweeping systems, the little things according to that. All of that 2024 comes out.
20:19
And then they open their data factory, which is in Shanghai. I believe this is in the high tech park. I know their headquarters is there. I think the data factory is there. Basically, you got to get, you got to do your training. They’re doing their training ahead of time downloading into the robot. You can do training a lot of ways. You can do SIM to real, you can have robots out there sending data back, all this. They basically create a big.
20:47
factory where they have robots walking around doing things and people are controlling whatever and they’re just gathering data by real-world scenarios. So they’re training their models with as much direct real-world stuff as possible. And then also in 2024, they released their Agibot World data set. All that data that they’re gathering, they make that open for everybody as well. So you can see there’s a very clear business strategy in here.
21:17
We are going to release a whole range of hardware, the robots. We’re going to do a lot of direct training as quickly as humanly possible, highly efficient data training. We’re going to put that into a massive data set. We’re going to let everyone in the world use it and we’re going to let everyone in the world use our AI for the Genie operator. This looks like Huawei’s standard strategy to me of we build infrastructure that everyone else builds upon.
21:45
So anyone can start using this and building on the robots, taking the data, creating skills, creating tasks, creating various services, all that, release them. Every business and developer can start to use this to build. And they give you the software for free, they give you the data for free, and they’re planning to make their money on the hardware, the robot, because you’ll probably end up building on their robots. That will get you manufacturing scale.
22:12
that will get you advantages, things like this looks like, Huawei has always said, we are in the infrastructure business. We build the stuff that everyone else uses to build with. That’s what this looks like to me right now. That’s a pretty cool strategy actually. Okay, that’s 2024. 2025, 2026, we’ll sort of put this together. Basically, they’re just raising money and scaling up.
22:38
You know, they’ve produced their thousandth unit by January 2025. Keep in mind, Elon Musk hasn’t produced a single unit. They’re at a thousand a year and six months ago. Now, I shouldn’t. Criticize Tesla, they clearly have a strategy, they’re clearly making robots, but they’re not selling them out into the wild. And then they just start showing that, like, you know, our robots can go this far by the end of 2025, they have five thousand.
23:07
robots shipped. So they’re pulling the lever on mass production as early as humanly possible. They are not doing robots that do backflips. And look at these cool videos, you know, that dance in performances on TV. They are going for mass production and rapid mass deployment as much as possible. This to me looks like really good strategy. Okay, that’s a bit about where we are.
23:35
Couple other factoids here I’ll give you. The Giga Data Factory doesn’t quite say where it is. I’ve been trying to find out where it is. I think it’s in the high tech park. They say it’s 3,000 square meters. It’s basically for physical training and validation where robots which are mostly tele-operated run real world tasks and then you have sensors in all the robots and they gather all the data.
24:05
So you have the inputs, the operator saying to do something, you have the data coming in from all the sensors. That’s how you train. The scenarios in there, apparently, I haven’t seen pictures of this, real world scenarios, residential homes, they have mockups, that’s 40 % of their training scenarios, living rooms, bathrooms, bedrooms, industrial warehouses, 20 % of the scenarios recreated, active conveyor belts, sorting systems, packaging.
24:34
stations, uh commercial catering, food handling, table clearing, retail stores, shelf stocking, item selection, cashier guidance, things like that, offices. Pretty cool. The data being collected, you get the visual data, obviously, from the cameras. You get the motion trajectory, the joint positions, head orientations. You get force feedback from all the sensors, sort of things like that, and you give it to the world. OK.
25:03
Here’s a little bit about the robot lineup. You can look these up. I’m behind on sending out emails. I owe about four emails now. I’ll catch up pretty quick. The series that are worth paying attention to, the A2 series, that’s the full-size humanoid. They have the A2 Ultra, they have the A2 Lite, they have the A2W. The W means it’s got wheels instead of biped. The A2, they have basically bipedal.
25:33
and wheeled versions, which are pretty cool. You can have basic gripper hands or you can have sort of advanced dexterous hands with, you know, really detailed sensors and all the fingers. They have the X series, which is again humanoid, but they’re smaller. That’s a lot nimble task. For those, they only do the biped. They don’t do the wheeled. They have the G series, which I mentioned. This is your industrial grade humanoids.
26:01
wheeled versions only. These are the things that move pallets around the warehouses, industrial, lot stronger. The force that the actuators can exert and the force that they can withstand when you’re moving 20-30 kilograms, very different.
26:20
The D series, which I mentioned, that’s the quadrupeds. These things are crazy. They zip all around. They can go upstairs. There’s all these interesting metrics for how steep of stairs can they go, one, how fast can they go? Most of these robots go one, two, three meters per second. Some of these robot dogs can go five meters per second. ah What kind of terrain they can go up. You know, the robot dogs like at Unitree,
26:48
They can kind of tell you, we can go up at a 30 degree angle, a 40 degree angle continuously as a robot dog just going up the stairs. And let’s say it’s 20 to 25 centimeters. I don’t see too many that can go beyond that steepness in terms of robot dogs. Unitree has a lot more dogs, by the way. And then you have the C series, which are these cleaning robots.
27:12
which are pretty interesting that the cleaning robots are actually kind of cool. You’ll see them in malls going around. They’re actually sort of a three in one pass. The first pass, I mean, as it passes over, does three things. The first thing is, it, it’s a sweeper. It’s got brooms and stuff and it sort of sweep stuff up and puts it in its, its trash bin. And then second is it’ll have a wet sort of a scrub. So it has these, you know, it uses fresh water.
27:42
well, it’s clean water, it’ll put it down there, the scrubbers will go down and it has to exert a pretty decent amount of downward force on the scrubber. So you have sort of an actuators pushing down. And then in the third, it will have a dry mop that soaks up now dry water and it will capture the dry water. So it kind of does three in one as it goes along the floor.
28:07
and then it goes to its auto station where it will recharge but it will also refill the clean water and empty the dirty water. They’re actually pretty cool.
28:18
Anyways, that’s kind of their suite of robots. And then they have supporting tech like they have dexterous hands that you can put on some of these, which are much more advanced. They have sort of teleoperation kits where you can put on VR headsets and control these things. I’ll put pictures of their robots. They’re real cool. Anyways, that’s kind of what, actually I’m almost done. Last point, and this will kind of be a summary. Okay.
28:46
They’ve got the full tech stack minus the chips. They’ve got all the robot types. They’re doing the whole AI thing. And they’re doing pretty much all the scenarios. Now there are industrial cases that are different. So I guess not all the scenarios, because the industrial cases get pretty detailed. But most of the cases people talk about for humanoids, they’re doing them. They’re not retail specialists, which some of these companies are.
29:13
We have humanoids, we have a barista, we have a bartender, and we have an ice cream server. No. In the humanoid range, in the range we’re talking about, not the full-on industrial robots, they’re doing all the scenarios, which are basically manufacturing. Industrial cases like loading and unloading. Now those can be pretty basic. Lift the pallet, well not the pallet, lift the item, put the item down.
29:41
but they can also be fairly advanced. If you have a robot with really advanced hands, they can do assembly level work of consumer electronics. They can put the components onto the chipboard. So industrial can be sort of electronics, semiconductor assembly, automotive is a big area. And depending on what level hands you could, they can be very detailed or they can just be moving things around, right? uh
30:07
line feeding, keeping things on the storage racks, keeping things moving versus adding components to a fairly advanced design. That’s kind of one bucket. Second bucket will be warehousing and logistics, sorting parcels, item fulfillment, material handling. Now, Amazon’s doing a tremendous amount in that area, obviously. Commercial services, uh retail, hospitality.
30:37
You know, welcome to our office building or welcome to our shopping mall, managing queue systems. can be a lot of lot of retail stuff here is actually pretty cool, but they can do events and marketing. They can stand in the hallway and promote things at an event or affair, stuff like that. And obviously within this, you’d have the floor maintenance, the sweeping, the cleaning, stuff like that.
31:03
Security patrol, infrastructure inspection, well that’s the, you know, that’s the Robodogs. Unitree does a lot in this. They can go up and down, not just factories, but they can go out into the wilderness and do inspections. They can go, you know, into rugged terrain. They can do outdoor stuff. They can go up and down staircases, stuff like that. You know, it can be security. It can be inspection. Inspection can be cameras. It can be sensors, sniffing. A lot of different sensors are actually kind of interesting.
31:34
And then it does AI research, which is basically the small robots, especially the small humanoids. You know, that’s where you can do a lot of training and education and things like that. Those are most of the use cases you hear about right now. Now within there, there’s a lot of security and military stuff I haven’t talked about. I didn’t see Agibot doing any of this, but I’m assuming every robot company is doing some of this. I don’t know. If you look at like some, you can see a lot of police, a lot of security stuff, a lot of
32:04
uh Fire brigades, they put hoses on the back of these robot docks are pretty interesting. Chemical spills, there’s a lot of interesting use cases out there. But anyways, that’s sort of a summary of Agibot uh with a little sort of touching on why I like them. Now let in the next part, I’m going to get into sort of the AI tech stack. Okay, then it’s going to get a lot more detailed, but yeah, call that the first pass. Okay, let’s talk about the software and the AI stuff and
32:33
This is kind of the concept, the lessons, I guess, for today. Now, OK, you’re looking at all this stuff. Most people who use these things don’t have to understand any of the AI. They just have to basically know two tools. There is basically what they call LinkCraft and LinkSoul. There’s Chinese names, Lingquang and let’s see the other one, Lingxin. But these are basically no code tools for any business or really any developer. And LinkCraft is
33:01
Yeah, this is where you’re doing motion creation. You’re doing choreography capture and you’re translating it into the training and putting it in there and then you can have it work in your shop the way you want it to work or you can create a skill or a task that you can sell to other people. You can make apps, you can make specialized things, right? They’re letting developers basically and businesses in particular, develop whatever they want for their business.
33:30
and all the software and the data sets are open source. Link soul is interesting in that this is again the other no code interaction tool. But in this case, you’re not giving it motion and tasks and skills. You’re giving it a personality and you’re telling it how you want it to interact with your employees or with customers or whatever. So you can make the voice and tone whatever you
33:58
You can give it personality. You can give it a certain character. You can train it on corporate information about your business, about the school. I you can give it an enterprise knowledge, things like that. And you can sort of do how you want it to behave within various workflows, either engaging with customers or working within your business. So interesting. So those are kind of the two tools that everyone who uses these things are going to use.
34:26
Then we switch over and we get into the architecture. Now, usually people talk about what I sort of said before, that look, there’s sort of two levels to this. There’s the cerebrum and there’s the cerebellum.
34:41
Okay, one is.
34:44
sort of high frequency, the cerebellum, that’s actions. Not, know, stand here, put your left foot forward. Oh, your left foot was, is on a pebble and it’s unsteady. So it’s this high frequency taking an action, but then also feeding back very quickly and responding. Oh, you’re, you’re holding this too tight or you’re holding it, but it’s starting to slip. So this is sort of continuously adjusting the motor and the action commands.
35:14
against real world changes and friction and all of that. They call that the cerebellum, which is kind of what your brain does all the time. You don’t think about standing and walking down the street and not falling over. Your cerebellum kind of does that for you. Now on the level above that, you have the higher level spatial reasoning. What are you seeing in your eyes? You can see the room. What is your plan to go forward? I’m going to walk across the room.
35:42
As I’m walking through the room, I’m adjusting as I go, okay, that’s your cerebrum, that’s spatial reasoning, plan generation, creating sort of a semantic map of your median environment, all of that. that’s generally, and that top level generally is AI. The lower level’s a lot of CPUs and some other things, more of that. And the top level, it can be there beyond robot or it can be in the cloud. And this is actually a difference.
36:11
because Agibot has done something interesting. Their Go model, their Genie operator, it is all sort of uh VLA models and all that, but their world models are not on their robots. Their world models are being done for training purposes, and then it is downloaded into the robot. So when the robot is actively moving, let’s say doing inference, not training,
36:40
It’s only using VLA and VLM. It is not using the world model. uh Unitree is not like that. Unitree and I think UBtech as well, their world model is running on the robot itself. Agibot has made the decision, at least at this point, to keep it off robot and do it in the training phase but not the inference. That’s a really interesting decision. I’ll talk about why I think they did that. I think that was a business decision more than anything else.
37:10
Okay, that’s the simple explanation. Now when you actually get into it more, you could actually argue that there’s four layers to all these things. I guess cerebrum is not right. You really should say cortex. Cortex would be closer. know, high level intelligence, high accuracy. You know, there’s a lot going on. High level perception, multimodal vision language, semantic mapping, spatial reasoning. There’s a ton going on. The cerebellum.
37:38
mid-level motor control, real level motor coordination, stability, balance, know, executing on your trajectory, things like that. You can actually break that into a couple more levels. ah You could say like, you know, how do you sort of translate what the cortex says to do into actual actions? You know, your hand is going to do this, that sort of thing. And you can actually say there’s a spinal cord level where
38:06
You know, your physical joints, your actuators, they have lots of tight feedback loops all the time that are happening. You could break it into four levels, I suppose. I don’t know. People do that. I’m not going to go into it that much. It’s not that interesting. The key difference, I suppose, is like, look, the energy consumption is pretty different when you go from sort of the vision language processing and all that generative AI stuff. When you go down to the lower level sensory feedback and all that. Well, the top layer
38:38
For chips, takes a lot more energy. The lower level, the chips are working much less. They’re doing way less processing, but you got energy going into the robot. But okay, that’s part of it. The chip level requirements, top level, the Cortex. The chip that I keep seeing people mention is the NVIDIA and Jetson. The Jetson Thor or the Jetson Auron, that’s their model for sort of system on a chip.
39:05
where they’re putting all this stuff on chips. It’s got memory, CPU, and all that processing together very efficient. These are basically edge computing, right? When you go down to level two, you’re talking about sort of real-time micro-press. You’re talking about a lot of ARM. I keep hearing ARM chips being mentioned. So anyways, I’ve been mapping out against various robot groups what chips they’re using at sort of the top level of the cortex and what level of the cerebral.
39:32
For the high level humanoids, the really high performing ones, every company I’ve looked at seems to be using NVIDIA Jetson series. And then below that, okay, a lot of arm and other stuff. Fine. Okay, let’s get into the models. Now, for Agibot, it’s the GE and the Go models. The Go is the Genie operator model. I’m just going to say Go.
40:02
The other one is the Genie Envision, which is GE, and their current version is two, I think they have different versions of the, I think they’re on two. So I’m just going to say the GO and the GE model, but it stands for Genie Operator, Genie Envision. Now, the GO model is basically what’s in the robot and what’s acting. And that’s when we’re talking about VLA models, VLM models, and Villa models. The Genie Environment Model, actually now,
40:31
I’ve heard it referred to as the Genie environment and the Genie in vision model. That’s the world model or the world action model, depending how you’re doing now for Agibot, that whole models in training on the inference phase, it’s just running on go for the others. Both of those are engaged. Both types of models are engaged as it’s operating in the inference phase. Now, the reason I think they’re doing that is for
40:57
Now, why would you do that? Why would you keep the world model out of the robot and just keep it at the training phase, which is what Agibot, I think, is doing? It means you can advance much more quickly. uh If you… It gets you a lot sort of better data training efficiency, and you can scale up much faster. Now, the weakness is if you don’t have a world model within your robot, your model is not going to be as robust. What a world model does, which I’ll talk about in second,
41:27
It’s a very good feedback model to see how your robot is performing within its given environment. Well, if you don’t have that in the robot, it’s only been at the training phase, your robot ultimately is going to be much less robust, but you’ll move much faster in the training phase. I think that’s what they’re doing. I think they’re going to move the world model into the robot in steps. That would be what I’d be watching for them to do. Okay, so let’s talk about these models. The Genie operator, I’m just going to talk about…
41:56
This is what’s in their G2, their A2, their X2. What’s the difference? Actually, let’s talk about what’s the difference between these VLA models. uh Vision language action, super interesting. I’ve been studying like crazy on this stuff. I’m still a very basic understanding, so I don’t want to talk like I’m not an expert in any way.
42:22
Over last couple years, I feel like I’ve gotten a really good understanding of how LLM models work, which is super important, I think, for building businesses and deploying them. I didn’t spend a lot of time looking into how veneer generators work or image generators work because that’s just sort content creation and I didn’t… It wasn’t as useful for enterprise. VLA models, yeah, I think LLM models are really what are going to end up powering agents.
42:51
and humans using AI tools. And VLA models are going to end what’s powering robots. And my sort of prediction is, look, I think businesses in the near future are going to be ah basically what I call HARO models, human agent robots. That’s going to be the three things that make up a business. Your operating model is going to be made of humans, agents, and robots.
43:18
in various combinations. That’s what a future operating model is going to be. And probably whatever service you’re providing is going to be a combination of those. okay, AI, LLM models go into AI, they go into agents, VLAs go into the robot piece. So that’s why I’ve been studying it. All right, what does it mean? uh My pretty basic understanding at this point is, oh what is the input? Well, the input is going to be vision and the input is going to be language.
43:47
So you have sort of two inputs going in, and that’s the V and the L. And then the output, when those things are combined and digested and processed and all of that, it then moves over into the action module, the A, which translates its plan into basically robot language. Move this solenoid 12 degrees to the left, apply a tiny bit more pressure with your left hand, it translates that into instructions.
44:17
that then go down into the robot and all its joints and things and tell it to act. Now the reason this looks a lot like an agent is because you need a very tight feedback loop from when you send all the little robot measurements down into, you know, move your foot this way, tighten this actuator a little bit, there’s going to be a very quick feedback loop back into the VLA model to adjust. Oh, you’re holding this item too tightly.
44:45
you’re starting to fall over to the left, rebalance. You have to have this sort of tight feedback loop. So these things, these VLA models, they run continuously, very, very quickly, the same way an agent will run very continuously once you give it a task to perform. So same type of thing. All right, so vision goes in, language goes in, and the language can be, here’s the action we want to take. We want to do this.
45:13
the Action module, it will send out the various things. Okay, fine. uh
45:22
Now that’s kind of the basic version. And that’s what I believe Unitree is doing now. When I look at Agibot, they’re doing something much more sophisticated than that. One, on top of the VLA, they’ve added a VLM. A VLM is sort of, let’s add a little bit more firepower in just the processing of vision. Basically a vision language model. We want to add more firepower to the vision and language.
45:51
sort of AI activities, why? Because we need to think about longer form activities. We’re not just telling you to move your hand 20 degrees to the left and put the cup on the table. A VLA could do that. I’m telling you not just to move the cup 20 degrees to the left, I’m telling you to make a cup of coffee for this customer. That’s sort of a longer form multi-step task.
46:22
that the robot needs to manage. So you put that in the VLM, which then will start to orchestrate the VLA. Okay, step one, grind the coffee. Here’s how you’re going to grind the coffee. The VLA will do that in various steps. Now that the coffee’s ground, the VLM will tell it, now put the coffee in the cup and add water. So the VLM is sort of a higher order orchestrator and a lot more firepower. So Agibot has VLA, but then they got VLM that they’re building on top of that. That’s one.
46:51
uh The other thing they’re doing is what they call a Villa model and That’s basically vi LLA instead of VLA. It’s VL LA and you stick an eye in there so you can say Villa That’s a vision language latent action model Now what’s the difference there? uh
47:17
Think about like if this was a car and we would look at the road, that’s vision, and we would have the language decide what to do, that’s the L, VL, what action would the car take? Okay, we’re telling the car to drive 20 feet ahead and stop. Okay, the action that comes out of that VLA model is going to be a lot of very detailed uh technical information.
47:46
uh Inject this much fuel into the carburetors. uh Turn the wheel 12 degrees to the right. The brake pedal needs to depress 20%. It’s got to give it very specific, actionable robot language that tells every part of the robot what to do. Okay, so basically that’s pretty complicated. And one of the problems with that
48:14
is if you want to train a robot to do something different now, okay, now I want you to shift lanes to the right in the car. The only way to train that is you have to put a bunch of cameras on and you have to put sensors on every little part of the car. And then you have to give it the language and the instruction, tell it what to do, and then see what it did by measuring it. Oh, the brake pads were 12 % too strong on the left wheel, right? The feedback data,
48:43
You have to basically pair all the visual data, the cameras, with very deep uh sensory data. So training takes a long time and it’s very detailed. Okay, what a latent action does, if you go from a VLA to a VLLA, uh that’s like building an intermediate representation where we have a steering wheel.
49:14
and a gear knob and a accelerator pedal and a brake pedal. Those items are sort of a latent space. They’re a sort of an abstraction layer that we’re building between all the machine language that tells every little bit of the car what to do and the instruction set. So you sort of build this intermediate thing.
49:38
where then you tell the car looks around, it sees what it can do, vision, it decides what to do, language, and then it gives those instructions to the latent space. Well, that’s basically saying, turn the steering wheel 12 degrees, press the accelerator, and then press the brake pedal. That’s what the VLA will generate. That language then goes into the action model, which then translates it to all the machine language that goes to every part of the car. Now that type of model,
50:07
is very easy to train because all you do is you put the cameras on and out of those cameras the action that it checks it against is oh we turn the steering wheel we press the brake pedal and we press it. So you can train those very quickly without all the data from the sensors and then you can just port that into model after model. So it’s very efficient for training. Now that’s kind how the future is going to be on these VLA models I think as far as I understand. If you want the actual
50:38
what it’s doing. These Villa models are basically take, they’re creating a latent space, which is a representation, a simpler representation of what the world looks like right now. That’s input number one. Input number two is they’re applying an action to that. Call that a latent action space. And those two things will create a future latent space.
51:07
And we are predicting what that will be and then seeing if it was achieved.
51:13
That’s kind how you think. You can go into that if you want. It’s super interesting. My understanding, think I’m sort of a beginner to intermediate. I’ve been studying it a lot. Anyways, latent spaces are super important. for this case, you can say, a latent space is a single observation at time.
51:34
A latent action space is sort of the relationship between two consecutive spaces in time. And your whole goal is to predict a future latent space, see if it has achieved it based on the performance data, and then feedback and improve. And so it’s kind of an abstraction layout. Okay, the key takeaways here are what? Start to think about VLAs. VLAs are going to be like, you have to understand LLMs and you have to understand VLAs.
52:03
Within VLAs, take a look at villas, VLLAs, super important. And think about this idea of pairing a VLA with a VLM on top that orchestrates the VALA. That’s the difference between move the copy from here to there to make me a cup of coffee and serve it, right? Think about that. That’s kind of, we’ll call that concept library lesson number one for today. The other one, and this will be the last one, because I know I’m kind of going a lot in here.
52:34
Everything I just talked about for Agibot, they have that in the Gini operator. The Gini environment or the GD envision, I’m not sure what the E actually stands for. That’s when you get into world models. And world models are super important. These are sort of more… World models are really sophisticated video generator models. A simple video generator model is like cling.
53:03
Here’s an image, here’s a text prompt, here’s a picture of me, the text prompt is make me do a funny dance. It then generates a short six second video of me doing a funny dance. Okay, those are really pretty simplistic in terms of video generator models, because there’s no right or wrong. It makes a prediction of what I look like doing this. But if I’m right or wrong, I’m not deploying its output.
53:32
into a robot that has to actually do something in the real world. No, it’s just the stakes are low. A video generator in a world model is you give it an input, which is an action to be taken. Walk across the room. You give it uh visual information, which is not really a video as an input. It’s more of uh a series of images, probably. You put those together, and then you predict the outcomes.
54:01
which could be it walks to the left and falls down, it walks over here and does whatever, but you basically need to predict the future and then put it into some sort of action, which is either putting it in a robot or putting it in a video game, seeing the action and then a feedback sort of corrects it over time, which is kind of how LLMs work, right? You give it some text, you tell it to do something and then it predicts lots of potential outcomes from that.
54:29
and it learns over time to predict the right one that you wanted or the correct one if it has an accuracy. Well, this is the same, but we are predicting an action within this world model basically. And the outcome is basically a predicted state. Now that would be a more sophisticated world, know, video generation model than just making a funny dance because it either has to work
54:55
in a video game where it has to work in a robot walking in a room or something like that. you can take it a bit more advanced but you can start to apply physics to that. So when it starts to predict a future state it can check that against various physics rules. Here’s the rule for gravity. Here’s the rule for conservation of energy. Now they don’t usually do that in video generation models like Kling or Minimax.
55:23
You know, it’s okay. They just kind of predict what’s going to happen based on pixels. But in this case, you actually are applying it to real world physics and you can, you know, you can incorporate that. And then the most important aspect to think about is it’s not a one shot. It’s not create a funny video of me and be done. It’s a continuous loop that can accurately predict and then manage.
55:51
actions within a physical environment or a virtual environment. So you feed it the state. Again, it’s a latent state in this case. It takes an action. It makes a prediction. It does it. That is immediately fed back in to become the next latent state that is fed in. So it’s a continuous rapid fire movement, which is how you do things in the real world. So that sort of uh rapid fire never ending loop is super important.
56:22
And yeah, so video generation is kind of the engine that sits inside these world models. They kind of, you what does that mean for a robot? It means it takes the camera view that the robot sees right now. That’s the video. It combines that with some sort of action. And then it generates pretty much the next few video frames that could happen if the robot does that. And then it predicts the most accurate or
56:49
you know, the one that it wants of all those predicted video frames. And that’s the action that goes forward. And this can be deployed in the robot either in the training phase or it can be also running during the inference when it’s actually walking around the room. And, you know, as mentioned, Agibot at this point only uses it at the training phase and Unitree and UBtech use it within the robot when it’s moving around, which is interesting. Now there are really two benefits to that. One, it’s kind of your real time planner, these world models.
57:19
It’s also a data factory. It generates a lot of data. So even in the training phase, you’re not using it then, but it’s both making your model smarter and it’s also generating a lot of data. And it can be a real time planner as well. So anyways, it’s super important. But world models and VLA models, I’m trying to get a lot smarter at. You can tell I’m somewhere between a beginner and an intermediate at this point, but give me a year and I’ll get there pretty good.
57:49
I feel pretty confident about LLMs, but it took a long time. Okay, that was a lot of content. Let me just finish here by saying, let me get back to business. Why do I like Agibot? Let me just wrap this all up, because that was kind of a crazy amount of content. It looks to me like their strategy is to move super fast in adoption, training, and manufacturing. So I think they want to flood the market.
58:18
with affordable humanoids and just sort of seed an entire ecosystem full of developers and businesses everywhere by giving them open data and open models. And they’re going for all the use cases, they’re going for all the robot types, and they have gone in a very short time to the leading manufacturer in terms of just robots moved. Now, the numbers are still small.
58:45
but they are moving fast and sort of robots out the door. eh Now, what are the benefits of that? Well, if you can get your manufacturing system to scale, well, there’s a lot of benefits to being a manufacturer at scale versus your rivals. You can offer robots cheaper than others. And that’s what Huawei has always done. They have always used their manufacturing scale as their biggest competitive weapon.
59:13
So I think that’s what they’re trying to do. Also, if you can get your ecosystem being used by everybody, the foundation model, the training tools, the simulation, the world data set, that also gets you scale, which now usually if this was just typical Huawei type stuff, you would have scale in manufacturing and you would have scale in R &D, which would make you smarter. But when you’re talking about AI systems, you get this data flywheel.
59:43
where the more people that use your stuff, the more data is generated and the smarter your models get. So you’re actually winning on three levels here. Manufacturing scale, R &D scale, and this sort of intelligence flywheel. You could also argue that there’s probably some network effects based on standardization that are going to happen. But yeah, I think that’s what they’re doing and I think they’re going to keep giving away software free. And that’s very consistent with Huawei’s like.
01:00:13
You know, we don’t sell solutions ourselves, we sell the infrastructure that everyone builds their solutions with.
01:00:21
Anyways, we’ll see. least this is a business strategy I understand. A lot of these other ones that are doing pilots and isolated use cases and we just make baristas that make coffee. don’t quite, I mean, it looks to me early stage. don’t see where it’s going. This one I think I can see the path, which I like. uh Anyways, that’s it. I keep on a couple of things that most of the sales for robots are in China. Most of them.
01:00:51
I mean, I don’t know, 80, 90 percent. It’s something like that. And most of them are still consumer facing robots. So these robot dogs, which are funny, you know, these human robots that are going into your house, that are going to retail spaces. Yeah, that’s a lot of where the industrial adoption, big surprise, is slower. It’s always slower.
01:01:16
So if you’re going for scale, you want to lean into the consumer side, which is what I think they’re doing. Anyways, there’s some exceptions to that. Linkerbot, which makes hands, is actually doing quite well on the industrial side. I’ll talk about that on another day. Anyways, that is it for me. That is kind of a pretty good depth of content for a Saturday night. It’s like 8 p.m. here. oh We’re going to have dinner here shortly and watch Mortal Kombat 2.
01:01:45
This is one of the weird things. Yeah, that’s enough content for today. But em my girlfriend, she has interesting hobbies and I don’t know where they come from. She loves Formula One. I don’t know why she’s crazy about Formula One. Every time the drivers come on and, you know, if we see one of the Formula One cars go by, she can identify which team is and that’s the Ferrari team. The other thing she loves is Mortal Kombat.
01:02:14
I don’t know where this comes from, but we play Mortal Kombat on the PlayStation and she just takes me apart. She’s been playing this for years. Anyways, the first Mortal Kombat movie, I actually thought it was pretty good, except for the lead character who was not interesting. All the other characters, the characters from the actual game, you know, are really cool. Well, the new Mortal Kombat movie is, you know, number two, which looks pretty good with Johnny Cage and, you know,
01:02:44
Anyways, it just showed up on HBO Max a day ago. So that’s our plan for a Saturday night. We’re going to order food and watch Mortal Kombat, which we’re both actually kind of excited about. She’s way more excited about this than I am. But I don’t know where this Mortal Kombat and the Formula One thing comes from. She loves Chinese mini dramas now. She loves these historical dramas like Pursuit of Jade. I don’t know where these weird hobbies come from. It’s very strange.
01:03:14
Anyways, so that’s the plan. So I’m finishing up here. It’s 8 p.m. on a Saturday. We’re going to go watch Mortal Kombat. I think we’re going to order Korean food. So that’s uh that’s pretty good for a Saturday. Anyways, that is it for me. I know this was a lot of content. I’m kind of back in professor mode. I’m going to send out a bunch of emails, one, because I’m behind. But two, there is a whole like serious topic here, I think, that needs to be covered. This is all important. uh
01:03:43
So yeah, I think we’re all going back to school a little bit again. We had to do it for software. We had to do it for AI. I think we have to do it for VLA models and real world embodied AI in the real world. Yeah, I think we’re all going back to school on this one. yeah, hunker down. Now, the good news is uh the models are complicated, but robots, I’ve been going through, I haven’t talked about this today. I’ve been going through all the technical stuff, the actuators and the costs and…
01:04:13
It’s actually pretty straightforward. The hardware side of robots is actually pretty easy. It’s much easier than EVs. These robots are not that complicated. It’s like studying your washing machine. It’s not that hard. So the hardware side, I was pleasantly surprised to find out it’s not that difficult. uh So that part’s easy. The AI stuff, yeah, that’s going to have huge implications. anyways, okay, that is it for me. I hope everyone’s doing well.
01:04:43
Talk to you probably in a couple of days. Bye bye.
——–
I am a consultant & keynote speaker on how to increase digital growth and strengthen digital AI moats.
I am the founder of TechMoat Consulting, a consulting firm specialized in increasing digital growth and strengthening digital AI moats. Get in contact here.
I write (a lot) about digital growth and digital AI strategy (3 best selling books, +2.9M followers on LinkedIn). There is a free book and email newsletter below.
My Moats and Marathons book series is a framework for building and measuring competitive advantages in digital businesses.
Note: This content (articles, podcasts, website info) is not investment advice. The information and opinions from me and any guests may be incorrect. The numbers and information may be wrong. The views expressed may no longer be relevant or accurate. Investing is risky. Do your own research.