<?xml version="1.0" encoding="UTF-8"?>
<?xml-stylesheet href="/stylesheet.xsl" type="text/xsl"?>
<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:sy="http://purl.org/rss/1.0/modules/syndication/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:podcast="https://podcastindex.org/namespace/1.0">
  <channel>
    <atom:link rel="self" type="application/rss+xml" href="https://feeds.transistor.fm/embodied-ai-101" title="MP3 Audio"/>
    <atom:link rel="hub" href="https://pubsubhubbub.appspot.com/"/>
    <podcast:podping usesPodping="true"/>
    <title>Embodied AI 101</title>
    <generator>Transistor (https://transistor.fm)</generator>
    <itunes:new-feed-url>https://feeds.transistor.fm/embodied-ai-101</itunes:new-feed-url>
    <description>Stay in the loop on research in AI and physical intelligence.</description>
    <copyright>© 2026 Shaoqing Tan</copyright>
    <podcast:guid>dc2e9af3-a6bd-5392-aadd-e305a2ce0453</podcast:guid>
    <podcast:locked>yes</podcast:locked>
    <language>en</language>
    <pubDate>Tue, 25 Aug 2026 05:10:12 -0700</pubDate>
    <lastBuildDate>Tue, 25 Aug 2026 05:11:14 -0700</lastBuildDate>
    <image>
      <url>https://img.transistorcdn.com/W67U9M8-4z2B6wcpspdoLUYtbS4QOEdWN2Nkg4375JQ/rs:fill:0:0:1/w:1400/h:1400/q:60/mb:500000/aHR0cHM6Ly9pbWct/dXBsb2FkLXByb2R1/Y3Rpb24udHJhbnNp/c3Rvci5mbS8wOGM3/YThiZDUxOTM4M2Vi/N2YzMTNkZDFiNDJh/ZDI1Mi5qcGc.jpg</url>
      <title>Embodied AI 101</title>
    </image>
    <itunes:category text="Technology"/>
    <itunes:category text="Science"/>
    <itunes:type>episodic</itunes:type>
    <itunes:author>Shaoqing Tan</itunes:author>
    <itunes:image href="https://img.transistorcdn.com/W67U9M8-4z2B6wcpspdoLUYtbS4QOEdWN2Nkg4375JQ/rs:fill:0:0:1/w:1400/h:1400/q:60/mb:500000/aHR0cHM6Ly9pbWct/dXBsb2FkLXByb2R1/Y3Rpb24udHJhbnNp/c3Rvci5mbS8wOGM3/YThiZDUxOTM4M2Vi/N2YzMTNkZDFiNDJh/ZDI1Mi5qcGc.jpg"/>
    <itunes:summary>Stay in the loop on research in AI and physical intelligence.</itunes:summary>
    <itunes:subtitle>Stay in the loop on research in AI and physical intelligence..</itunes:subtitle>
    <itunes:keywords>embodied ai technology robotics</itunes:keywords>
    <itunes:owner>
      <itunes:name>Shaoqing Tan</itunes:name>
      <itunes:email>8tzxb5lel@mozmail.com</itunes:email>
    </itunes:owner>
    <itunes:complete>No</itunes:complete>
    <itunes:explicit>No</itunes:explicit>
    <item>
      <title>Taming the Flying Tip: Constrained MPC for Dynamic Deformable Linear Objects</title>
      <itunes:title>Taming the Flying Tip: Constrained MPC for Dynamic Deformable Linear Objects</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">f126d8a6-6841-4d1a-97f1-35bc965f2aba</guid>
      <link>https://share.transistor.fm/s/3ae36f98</link>
      <description>
        <![CDATA[Deformable linear objects (DLOs) exhibit highly nonlinear dynamic behavior, complicating their control during high-speed maneuvers. Furthermore, the lack of a generic spatial representation and the difficulty of real-time state estimation...]]>
      </description>
      <content:encoded>
        <![CDATA[Deformable linear objects (DLOs) exhibit highly nonlinear dynamic behavior, complicating their control during high-speed maneuvers. Furthermore, the lack of a generic spatial representation and the difficulty of real-time state estimation...]]>
      </content:encoded>
      <pubDate>Tue, 25 Aug 2026 05:10:00 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/3ae36f98/f79351c0.mp3" length="30717953" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1920</itunes:duration>
      <itunes:summary>Deformable linear objects (DLOs) exhibit highly nonlinear dynamic behavior, complicating their control during high-speed maneuvers. Furthermore, the lack of a generic spatial representation and the difficulty of real-time state estimation...</itunes:summary>
      <itunes:subtitle>Deformable linear objects (DLOs) exhibit highly nonlinear dynamic behavior, complicating their control during high-speed maneuvers. Furthermore, the lack of a generic spatial representation and the difficulty of real-time state estimation...</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/3ae36f98/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Replay Is Not Reality: Embodied RL When Both Policy and Physics Shift</title>
      <itunes:title>Replay Is Not Reality: Embodied RL When Both Policy and Physics Shift</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">aa9f7392-4076-48b3-b102-0e9bf1116f80</guid>
      <link>https://share.transistor.fm/s/dae7d482</link>
      <description>
        <![CDATA[Abstract Embodied agents must continuously adapt to the physical world using interaction data collected across varying timescales, controllers, and environmental conditions. However, standard reinforcement learning assumes stationary dynamics and on-policy data collection, which breaks down when agents must learn from heterogeneous data sources with non-stationary dynamics and off-policy behavior.]]>
      </description>
      <content:encoded>
        <![CDATA[Abstract Embodied agents must continuously adapt to the physical world using interaction data collected across varying timescales, controllers, and environmental conditions. However, standard reinforcement learning assumes stationary dynamics and on-policy data collection, which breaks down when agents must learn from heterogeneous data sources with non-stationary dynamics and off-policy behavior.]]>
      </content:encoded>
      <pubDate>Tue, 25 Aug 2026 05:08:31 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/dae7d482/3dbc7161.mp3" length="29562714" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1848</itunes:duration>
      <itunes:summary>Abstract Embodied agents must continuously adapt to the physical world using interaction data collected across varying timescales, controllers, and environmental conditions. However, standard reinforcement learning assumes stationary dynamics and on-policy data collection, which breaks down when agents must learn from heterogeneous data sources with non-stationary dynamics and off-policy behavior.</itunes:summary>
      <itunes:subtitle>Abstract Embodied agents must continuously adapt to the physical world using interaction data collected across varying timescales, controllers, and environmental conditions. However, standard reinforcement learning assumes stationary dynamics and on-polic</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/dae7d482/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>PPO Beyond the Simulator: What a TurtleBot3 Mapless Navigation Study Really Demonstrates</title>
      <itunes:title>PPO Beyond the Simulator: What a TurtleBot3 Mapless Navigation Study Really Demonstrates</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">c39b925f-482a-499f-94f2-30949eab7b53</guid>
      <link>https://share.transistor.fm/s/f307e79b</link>
      <description>
        <![CDATA[Objective: In the field of mobile robotics, autonomous navigation in dynamic environments is one of the most challenging tasks in these environments: traditional methods based on pre-mapping and geometric planning...]]>
      </description>
      <content:encoded>
        <![CDATA[Objective: In the field of mobile robotics, autonomous navigation in dynamic environments is one of the most challenging tasks in these environments: traditional methods based on pre-mapping and geometric planning...]]>
      </content:encoded>
      <pubDate>Mon, 24 Aug 2026 05:20:46 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/f307e79b/b90bfe74.mp3" length="32989143" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2062</itunes:duration>
      <itunes:summary>Objective: In the field of mobile robotics, autonomous navigation in dynamic environments is one of the most challenging tasks in these environments: traditional methods based on pre-mapping and geometric planning...</itunes:summary>
      <itunes:subtitle>Objective: In the field of mobile robotics, autonomous navigation in dynamic environments is one of the most challenging tasks in these environments: traditional methods based on pre-mapping and geometric planning...</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/f307e79b/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>VLAs in the Wild: What 142 Papers Really Say About Generalist Robot Intelligence</title>
      <itunes:title>VLAs in the Wild: What 142 Papers Really Say About Generalist Robot Intelligence</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">1e3825b4-6ae5-430b-baf5-84f2cdcf804e</guid>
      <link>https://share.transistor.fm/s/62cb23dd</link>
      <description>
        <![CDATA[Generalist robots need to perform diverse tasks while operating in dynamic, uncertain, and unstructured environments, often around human beings. Vision-language-action (VLA) models have recently emerged as a promising and flexible approach to enabling generalist robot behavior. This is a systematic literature review of VLA models for generalist robots.]]>
      </description>
      <content:encoded>
        <![CDATA[Generalist robots need to perform diverse tasks while operating in dynamic, uncertain, and unstructured environments, often around human beings. Vision-language-action (VLA) models have recently emerged as a promising and flexible approach to enabling generalist robot behavior. This is a systematic literature review of VLA models for generalist robots.]]>
      </content:encoded>
      <pubDate>Sun, 23 Aug 2026 05:08:41 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/62cb23dd/fc4b329b.mp3" length="32656865" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2042</itunes:duration>
      <itunes:summary>Generalist robots need to perform diverse tasks while operating in dynamic, uncertain, and unstructured environments, often around human beings. Vision-language-action (VLA) models have recently emerged as a promising and flexible approach to enabling generalist robot behavior. This is a systematic literature review of VLA models for generalist robots.</itunes:summary>
      <itunes:subtitle>Generalist robots need to perform diverse tasks while operating in dynamic, uncertain, and unstructured environments, often around human beings. Vision-language-action (VLA) models have recently emerged as a promising and flexible approach to enabling gen</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/62cb23dd/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>VEDA and the Moment Vision Stops Being Enough</title>
      <itunes:title>VEDA and the Moment Vision Stops Being Enough</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">eb476f5c-6ac4-4b4b-8f7c-e44e33a560d2</guid>
      <link>https://share.transistor.fm/s/86abcd18</link>
      <description>
        <![CDATA[Industrial humanoid robots are increasingly expected to execute high-precision manipulation tasks, such as assembly, welding, and material handling, under dynamic and contact-rich working conditions.]]>
      </description>
      <content:encoded>
        <![CDATA[Industrial humanoid robots are increasingly expected to execute high-precision manipulation tasks, such as assembly, welding, and material handling, under dynamic and contact-rich working conditions.]]>
      </content:encoded>
      <pubDate>Sat, 22 Aug 2026 05:11:22 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/86abcd18/0255dee0.mp3" length="28886456" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1806</itunes:duration>
      <itunes:summary>Industrial humanoid robots are increasingly expected to execute high-precision manipulation tasks, such as assembly, welding, and material handling, under dynamic and contact-rich working conditions.</itunes:summary>
      <itunes:subtitle>Industrial humanoid robots are increasingly expected to execute high-precision manipulation tasks, such as assembly, welding, and material handling, under dynamic and contact-rich working conditions.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/86abcd18/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Three Ways to Buy Back Robot Data: Augment the Demonstration, Import the Semantics, Benchmark the Physics</title>
      <itunes:title>Three Ways to Buy Back Robot Data: Augment the Demonstration, Import the Semantics, Benchmark the Physics</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">891ac818-2dd3-4f08-8602-aaff58fed6b2</guid>
      <link>https://share.transistor.fm/s/37fefddc</link>
      <description>
        <![CDATA[Abstract Robots have the potential to assist people in homes, hospitals, warehouses, and factories, but today's systems still struggle to learn robust manipulation skills from limited data, especially in contact-rich tasks. This paper explores data-efficient robot learning approaches for contact-rich manipulation.]]>
      </description>
      <content:encoded>
        <![CDATA[Abstract Robots have the potential to assist people in homes, hospitals, warehouses, and factories, but today's systems still struggle to learn robust manipulation skills from limited data, especially in contact-rich tasks. This paper explores data-efficient robot learning approaches for contact-rich manipulation.]]>
      </content:encoded>
      <pubDate>Fri, 21 Aug 2026 05:08:59 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/37fefddc/7f7af6e2.mp3" length="31362028" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1961</itunes:duration>
      <itunes:summary>Abstract Robots have the potential to assist people in homes, hospitals, warehouses, and factories, but today's systems still struggle to learn robust manipulation skills from limited data, especially in contact-rich tasks. This paper explores data-efficient robot learning approaches for contact-rich manipulation.</itunes:summary>
      <itunes:subtitle>Abstract Robots have the potential to assist people in homes, hospitals, warehouses, and factories, but today's systems still struggle to learn robust manipulation skills from limited data, especially in contact-rich tasks. This paper explores data-effici</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/37fefddc/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>UI-Mate: Show the Desktop Agent Once, Then Let It Replan</title>
      <itunes:title>UI-Mate: Show the Desktop Agent Once, Then Let It Replan</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">c52f6fe5-7c34-4442-a636-09466b0e03de</guid>
      <link>https://share.transistor.fm/s/278054f8</link>
      <description>
        <![CDATA[A 27B model that achieves one-shot doubling of strict success rates on long-horizon office automation tasks via demonstration-conditioned policy learning.]]>
      </description>
      <content:encoded>
        <![CDATA[A 27B model that achieves one-shot doubling of strict success rates on long-horizon office automation tasks via demonstration-conditioned policy learning.]]>
      </content:encoded>
      <pubDate>Thu, 20 Aug 2026 05:21:49 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/278054f8/39090108.mp3" length="29772947" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1861</itunes:duration>
      <itunes:summary>A 27B model that achieves one-shot doubling of strict success rates on long-horizon office automation tasks via demonstration-conditioned policy learning.</itunes:summary>
      <itunes:subtitle>A 27B model that achieves one-shot doubling of strict success rates on long-horizon office automation tasks via demonstration-conditioned policy learning.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/278054f8/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>The Frozen VLM That Learns Anyway: Inside Spatial Memory Agent</title>
      <itunes:title>The Frozen VLM That Learns Anyway: Inside Spatial Memory Agent</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">0202512c-771a-4c28-885c-4cf22131a6d4</guid>
      <link>https://share.transistor.fm/s/75bff6fb</link>
      <description>
        <![CDATA[A parameter-update-free framework that distills verified experience into transferable spatial memory for frozen VLMs, achieving consistent gains across 4 VLMs and 5 spatial reasoning benchmarks.]]>
      </description>
      <content:encoded>
        <![CDATA[A parameter-update-free framework that distills verified experience into transferable spatial memory for frozen VLMs, achieving consistent gains across 4 VLMs and 5 spatial reasoning benchmarks.]]>
      </content:encoded>
      <pubDate>Thu, 20 Aug 2026 05:09:24 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/75bff6fb/80da8753.mp3" length="29869078" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1867</itunes:duration>
      <itunes:summary>A parameter-update-free framework that distills verified experience into transferable spatial memory for frozen VLMs, achieving consistent gains across 4 VLMs and 5 spatial reasoning benchmarks.</itunes:summary>
      <itunes:subtitle>A parameter-update-free framework that distills verified experience into transferable spatial memory for frozen VLMs, achieving consistent gains across 4 VLMs and 5 spatial reasoning benchmarks.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/75bff6fb/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>The LLM Proposes, the Planner Disposes: VLA-SP for Verifiable Robot Manipulation</title>
      <itunes:title>The LLM Proposes, the Planner Disposes: VLA-SP for Verifiable Robot Manipulation</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">a38c3992-3dc2-47ae-8721-4ed5bbf622dc</guid>
      <link>https://share.transistor.fm/s/f7f2dbdb</link>
      <description>
        <![CDATA[Abstract Cross-modal foundation models are increasingly used for robotic task understanding and planning. However, connecting multimodal observations and natural language instructions to symbolic planners and executable robot actions remains a challenge. This paper addresses long-horizon robotic manipulation tasks using LLM-driven symbolic planning.]]>
      </description>
      <content:encoded>
        <![CDATA[Abstract Cross-modal foundation models are increasingly used for robotic task understanding and planning. However, connecting multimodal observations and natural language instructions to symbolic planners and executable robot actions remains a challenge. This paper addresses long-horizon robotic manipulation tasks using LLM-driven symbolic planning.]]>
      </content:encoded>
      <pubDate>Wed, 19 Aug 2026 05:14:39 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/f7f2dbdb/6bc0b011.mp3" length="31272167" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1955</itunes:duration>
      <itunes:summary>Abstract Cross-modal foundation models are increasingly used for robotic task understanding and planning. However, connecting multimodal observations and natural language instructions to symbolic planners and executable robot actions remains a challenge. This paper addresses long-horizon robotic manipulation tasks using LLM-driven symbolic planning.</itunes:summary>
      <itunes:subtitle>Abstract Cross-modal foundation models are increasingly used for robotic task understanding and planning. However, connecting multimodal observations and natural language instructions to symbolic planners and executable robot actions remains a challenge. </itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/f7f2dbdb/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>When Vision Cannot Know: MulPlanLM Lets Robot Planners Lift, Weigh, and Decide</title>
      <itunes:title>When Vision Cannot Know: MulPlanLM Lets Robot Planners Lift, Weigh, and Decide</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">3cc71017-c620-4cdc-ab61-ec275c93bce3</guid>
      <link>https://share.transistor.fm/s/e92d0abe</link>
      <description>
        <![CDATA[MulPlanLM: multimodal robotic task planning with vision-language models and physical feedback]]>
      </description>
      <content:encoded>
        <![CDATA[MulPlanLM: multimodal robotic task planning with vision-language models and physical feedback]]>
      </content:encoded>
      <pubDate>Wed, 19 Aug 2026 05:12:55 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/e92d0abe/d63ee020.mp3" length="29730733" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1859</itunes:duration>
      <itunes:summary>MulPlanLM: multimodal robotic task planning with vision-language models and physical feedback</itunes:summary>
      <itunes:subtitle>MulPlanLM: multimodal robotic task planning with vision-language models and physical feedback</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/e92d0abe/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Reading Trust in Motion: A Multimodal Perception Stack for Industrial HRC</title>
      <itunes:title>Reading Trust in Motion: A Multimodal Perception Stack for Industrial HRC</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">9ae7e01c-e8a3-4269-8e67-1cb9f706883f</guid>
      <link>https://share.transistor.fm/s/f2505d8e</link>
      <description>
        <![CDATA[Abstract Industry 5.0 promotes collaborative manufacturing environments where humans and robots work together, leveraging their complementary strengths to improve productivity and safety. In such settings, trust is fundamental to ensuring effective collaboration between humans and robots. This paper presents a multi-modal model for trust evaluation in industrial human-robot collaboration settings.]]>
      </description>
      <content:encoded>
        <![CDATA[Abstract Industry 5.0 promotes collaborative manufacturing environments where humans and robots work together, leveraging their complementary strengths to improve productivity and safety. In such settings, trust is fundamental to ensuring effective collaboration between humans and robots. This paper presents a multi-modal model for trust evaluation in industrial human-robot collaboration settings.]]>
      </content:encoded>
      <pubDate>Tue, 18 Aug 2026 05:16:33 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/f2505d8e/fc8420fe.mp3" length="26728114" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1671</itunes:duration>
      <itunes:summary>Abstract Industry 5.0 promotes collaborative manufacturing environments where humans and robots work together, leveraging their complementary strengths to improve productivity and safety. In such settings, trust is fundamental to ensuring effective collaboration between humans and robots. This paper presents a multi-modal model for trust evaluation in industrial human-robot collaboration settings.</itunes:summary>
      <itunes:subtitle>Abstract Industry 5.0 promotes collaborative manufacturing environments where humans and robots work together, leveraging their complementary strengths to improve productivity and safety. In such settings, trust is fundamental to ensuring effective collab</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/f2505d8e/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Two MPCs, One Robot: Bézier Curves for Fast, Compliant Mobile Manipulation</title>
      <itunes:title>Two MPCs, One Robot: Bézier Curves for Fast, Compliant Mobile Manipulation</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">e73caffc-f56f-486b-b9b7-313aac4e5654</guid>
      <link>https://share.transistor.fm/s/f0f79501</link>
      <description>
        <![CDATA[Dual-arm mobile manipulators can transport and manipulate large objects with simple end-effectors. Interacting with dynamic environments subject to strict safety and compliance requirements, achieving whole-body motion planning online while maintaining compliance is challenging. This paper presents an efficient whole-body model predictive control approach for online compliant dual-arm mobile manipulation.]]>
      </description>
      <content:encoded>
        <![CDATA[Dual-arm mobile manipulators can transport and manipulate large objects with simple end-effectors. Interacting with dynamic environments subject to strict safety and compliance requirements, achieving whole-body motion planning online while maintaining compliance is challenging. This paper presents an efficient whole-body model predictive control approach for online compliant dual-arm mobile manipulation.]]>
      </content:encoded>
      <pubDate>Tue, 18 Aug 2026 05:13:59 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/f0f79501/f451a2d2.mp3" length="30417858" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1902</itunes:duration>
      <itunes:summary>Dual-arm mobile manipulators can transport and manipulate large objects with simple end-effectors. Interacting with dynamic environments subject to strict safety and compliance requirements, achieving whole-body motion planning online while maintaining compliance is challenging. This paper presents an efficient whole-body model predictive control approach for online compliant dual-arm mobile manipulation.</itunes:summary>
      <itunes:subtitle>Dual-arm mobile manipulators can transport and manipulate large objects with simple end-effectors. Interacting with dynamic environments subject to strict safety and compliance requirements, achieving whole-body motion planning online while maintaining co</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/f0f79501/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>The Reward Behind the Gait: Learning Transferable Locomotion from a Few Stick-Insect Steps</title>
      <itunes:title>The Reward Behind the Gait: Learning Transferable Locomotion from a Few Stick-Insect Steps</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">bb5f1458-270a-4773-aff5-972b7025cc5e</guid>
      <link>https://share.transistor.fm/s/2d3b7ecb</link>
      <description>
        <![CDATA[Insect locomotion exhibits remarkable adaptability and flexibility despite the limited scale of its nervous system. This paper extracts and models underlying leg coordination principles via adversarial inverse reinforcement learning, enabling transferable robot locomotion from limited biological data.]]>
      </description>
      <content:encoded>
        <![CDATA[Insect locomotion exhibits remarkable adaptability and flexibility despite the limited scale of its nervous system. This paper extracts and models underlying leg coordination principles via adversarial inverse reinforcement learning, enabling transferable robot locomotion from limited biological data.]]>
      </content:encoded>
      <pubDate>Mon, 17 Aug 2026 05:20:35 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/2d3b7ecb/30b39d86.mp3" length="28112813" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1758</itunes:duration>
      <itunes:summary>Insect locomotion exhibits remarkable adaptability and flexibility despite the limited scale of its nervous system. This paper extracts and models underlying leg coordination principles via adversarial inverse reinforcement learning, enabling transferable robot locomotion from limited biological data.</itunes:summary>
      <itunes:subtitle>Insect locomotion exhibits remarkable adaptability and flexibility despite the limited scale of its nervous system. This paper extracts and models underlying leg coordination principles via adversarial inverse reinforcement learning, enabling transferable</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/2d3b7ecb/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Beyond Looking Human: Bio-Functional Robot Design for Cross-Scale Dexterity and Tool Use</title>
      <itunes:title>Beyond Looking Human: Bio-Functional Robot Design for Cross-Scale Dexterity and Tool Use</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">1a247a8a-e902-4242-b228-1c668e18d99c</guid>
      <link>https://share.transistor.fm/s/8e7d79cf</link>
      <description>
        <![CDATA[Dexterous hands and robotic manipulators as physical interfaces for diverse environments. Proposes bio-functional mimicry beyond structural biomimicry, enabling cross-scale manipulation and limb-tool integration for versatile robotic tasks.]]>
      </description>
      <content:encoded>
        <![CDATA[Dexterous hands and robotic manipulators as physical interfaces for diverse environments. Proposes bio-functional mimicry beyond structural biomimicry, enabling cross-scale manipulation and limb-tool integration for versatile robotic tasks.]]>
      </content:encoded>
      <pubDate>Mon, 17 Aug 2026 05:16:18 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/8e7d79cf/c190a4a0.mp3" length="29247572" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1828</itunes:duration>
      <itunes:summary>Dexterous hands and robotic manipulators as physical interfaces for diverse environments. Proposes bio-functional mimicry beyond structural biomimicry, enabling cross-scale manipulation and limb-tool integration for versatile robotic tasks.</itunes:summary>
      <itunes:subtitle>Dexterous hands and robotic manipulators as physical interfaces for diverse environments. Proposes bio-functional mimicry beyond structural biomimicry, enabling cross-scale manipulation and limb-tool integration for versatile robotic tasks.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/8e7d79cf/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>LDA-1B and the Case for Never Throwing Robot Data Away</title>
      <itunes:title>LDA-1B and the Case for Never Throwing Robot Data Away</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">0727e67d-6b20-4abd-8504-c37044410691</guid>
      <link>https://share.transistor.fm/s/a93872ee</link>
      <description>
        <![CDATA[Scales a latent dynamics action model to 1B parameters by ingesting diverse embodied datasets into a unified training pipeline for generalist robot policies.]]>
      </description>
      <content:encoded>
        <![CDATA[Scales a latent dynamics action model to 1B parameters by ingesting diverse embodied datasets into a unified training pipeline for generalist robot policies.]]>
      </content:encoded>
      <pubDate>Sun, 16 Aug 2026 05:07:35 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/a93872ee/9e392556.mp3" length="33509502" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2095</itunes:duration>
      <itunes:summary>Scales a latent dynamics action model to 1B parameters by ingesting diverse embodied datasets into a unified training pipeline for generalist robot policies.</itunes:summary>
      <itunes:subtitle>Scales a latent dynamics action model to 1B parameters by ingesting diverse embodied datasets into a unified training pipeline for generalist robot policies.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/a93872ee/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Flex-π: One Robot Policy, Three Futures, Fifty-Six Ways to Run It</title>
      <itunes:title>Flex-π: One Robot Policy, Three Futures, Fifty-Six Ways to Run It</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">544efe33-1481-4e51-a829-f1e7ba6d806b</guid>
      <link>https://share.transistor.fm/s/fc1963e3</link>
      <description>
        <![CDATA[Jointly predicts future RGB frames, 3D pointmaps, DINO semantics, and actions within a unified world-action model. Can be deployed flexibly as a VLA, full world-action model, or hybrid from a single checkpoint.]]>
      </description>
      <content:encoded>
        <![CDATA[Jointly predicts future RGB frames, 3D pointmaps, DINO semantics, and actions within a unified world-action model. Can be deployed flexibly as a VLA, full world-action model, or hybrid from a single checkpoint.]]>
      </content:encoded>
      <pubDate>Wed, 12 Aug 2026 14:08:45 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/fc1963e3/b40a4164.mp3" length="36584010" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2287</itunes:duration>
      <itunes:summary>Jointly predicts future RGB frames, 3D pointmaps, DINO semantics, and actions within a unified world-action model. Can be deployed flexibly as a VLA, full world-action model, or hybrid from a single checkpoint.</itunes:summary>
      <itunes:subtitle>Jointly predicts future RGB frames, 3D pointmaps, DINO semantics, and actions within a unified world-action model. Can be deployed flexibly as a VLA, full world-action model, or hybrid from a single checkpoint.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/fc1963e3/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Imagined Motion, Grounded Geometry: How Rest2Art Makes Closed Objects Articulate</title>
      <itunes:title>Imagined Motion, Grounded Geometry: How Rest2Art Makes Closed Objects Articulate</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">68ece111-b18c-424f-b228-8b928fd3db95</guid>
      <link>https://share.transistor.fm/s/6bd69fc9</link>
      <description>
        <![CDATA[An ECCV 2026 paper that converts a single static closed-state image into a simulation-ready articulated asset using video diffusion for joint hypothesis generation, requiring no motion data. Enables articulated object reconstruction without dynamic observations.]]>
      </description>
      <content:encoded>
        <![CDATA[An ECCV 2026 paper that converts a single static closed-state image into a simulation-ready articulated asset using video diffusion for joint hypothesis generation, requiring no motion data. Enables articulated object reconstruction without dynamic observations.]]>
      </content:encoded>
      <pubDate>Wed, 12 Aug 2026 14:07:41 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/6bd69fc9/54373549.mp3" length="33185584" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2075</itunes:duration>
      <itunes:summary>An ECCV 2026 paper that converts a single static closed-state image into a simulation-ready articulated asset using video diffusion for joint hypothesis generation, requiring no motion data. Enables articulated object reconstruction without dynamic observations.</itunes:summary>
      <itunes:subtitle>An ECCV 2026 paper that converts a single static closed-state image into a simulation-ready articulated asset using video diffusion for joint hypothesis generation, requiring no motion data. Enables articulated object reconstruction without dynamic observ</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/6bd69fc9/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>ComBodied Agents: When the Human Becomes the World Model</title>
      <itunes:title>ComBodied Agents: When the Human Becomes the World Model</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">47dffeb9-4837-4256-8dbe-81d91095615a</guid>
      <link>https://share.transistor.fm/s/591ad4db</link>
      <description>
        <![CDATA[Proposes a new agent architecture that unifies digital and physical (embodied) action spaces, supporting human-state trajectories with explicit handling of consent, uncertainty, and reversibility for safer human-robot interaction.]]>
      </description>
      <content:encoded>
        <![CDATA[Proposes a new agent architecture that unifies digital and physical (embodied) action spaces, supporting human-state trajectories with explicit handling of consent, uncertainty, and reversibility for safer human-robot interaction.]]>
      </content:encoded>
      <pubDate>Wed, 12 Aug 2026 05:12:03 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/591ad4db/9649e813.mp3" length="30934038" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1934</itunes:duration>
      <itunes:summary>Proposes a new agent architecture that unifies digital and physical (embodied) action spaces, supporting human-state trajectories with explicit handling of consent, uncertainty, and reversibility for safer human-robot interaction.</itunes:summary>
      <itunes:subtitle>Proposes a new agent architecture that unifies digital and physical (embodied) action spaces, supporting human-state trajectories with explicit handling of consent, uncertainty, and reversibility for safer human-robot interaction.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/591ad4db/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>A Sphere Every Robot Hand Can Speak</title>
      <itunes:title>A Sphere Every Robot Hand Can Speak</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">8b5f0c80-927b-4345-aa44-2e6e5701ade2</guid>
      <link>https://share.transistor.fm/s/ca89238c</link>
      <description>
        <![CDATA[Introduces a unified hand action space that enables seamless cross-embodiment transfer for dexterous manipulation policies across different robot hands.]]>
      </description>
      <content:encoded>
        <![CDATA[Introduces a unified hand action space that enables seamless cross-embodiment transfer for dexterous manipulation policies across different robot hands.]]>
      </content:encoded>
      <pubDate>Tue, 11 Aug 2026 14:08:14 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/ca89238c/3e6b6d4a.mp3" length="32611726" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2039</itunes:duration>
      <itunes:summary>Introduces a unified hand action space that enables seamless cross-embodiment transfer for dexterous manipulation policies across different robot hands.</itunes:summary>
      <itunes:subtitle>Introduces a unified hand action space that enables seamless cross-embodiment transfer for dexterous manipulation policies across different robot hands.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/ca89238c/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>The Soft Gripper Pareto Frontier: Adaptability Without Surrendering Payload</title>
      <itunes:title>The Soft Gripper Pareto Frontier: Adaptability Without Surrendering Payload</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">f9fe23cd-5566-448f-a76a-8eabb559c9df</guid>
      <link>https://share.transistor.fm/s/b14caa2f</link>
      <description>
        <![CDATA[Soft gripper systems have inherent flexibility and remarkable interactivity security. However, balancing load-bearing capacity, deformation, complexity, and adaptability in a single gripper remains a substantial challenge.]]>
      </description>
      <content:encoded>
        <![CDATA[Soft gripper systems have inherent flexibility and remarkable interactivity security. However, balancing load-bearing capacity, deformation, complexity, and adaptability in a single gripper remains a substantial challenge.]]>
      </content:encoded>
      <pubDate>Tue, 11 Aug 2026 05:10:20 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/b14caa2f/e1c6c3d6.mp3" length="29523007" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1846</itunes:duration>
      <itunes:summary>Soft gripper systems have inherent flexibility and remarkable interactivity security. However, balancing load-bearing capacity, deformation, complexity, and adaptability in a single gripper remains a substantial challenge.</itunes:summary>
      <itunes:subtitle>Soft gripper systems have inherent flexibility and remarkable interactivity security. However, balancing load-bearing capacity, deformation, complexity, and adaptability in a single gripper remains a substantial challenge.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/b14caa2f/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Dyna-2's Million-Hour Bet: When Human Video Starts Improving Robots</title>
      <itunes:title>Dyna-2's Million-Hour Bet: When Human Video Starts Improving Robots</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">fcb4b0cb-c7bb-4f6f-ac91-3d3791d54100</guid>
      <link>https://share.transistor.fm/s/f8d1eb7e</link>
      <description>
        <![CDATA[Pre-trained on 1 million hours of human video, Dyna-2 demonstrates scaling laws for world-action models across four orders of magnitude on human data, with emergent scaling on unseen robot data and joint video + action objectives enabling cross-embodiment transfer.]]>
      </description>
      <content:encoded>
        <![CDATA[Pre-trained on 1 million hours of human video, Dyna-2 demonstrates scaling laws for world-action models across four orders of magnitude on human data, with emergent scaling on unseen robot data and joint video + action objectives enabling cross-embodiment transfer.]]>
      </content:encoded>
      <pubDate>Tue, 11 Aug 2026 05:08:30 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/f8d1eb7e/08b66933.mp3" length="29700222" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1857</itunes:duration>
      <itunes:summary>Pre-trained on 1 million hours of human video, Dyna-2 demonstrates scaling laws for world-action models across four orders of magnitude on human data, with emergent scaling on unseen robot data and joint video + action objectives enabling cross-embodiment transfer.</itunes:summary>
      <itunes:subtitle>Pre-trained on 1 million hours of human video, Dyna-2 demonstrates scaling laws for world-action models across four orders of magnitude on human data, with emergent scaling on unseen robot data and joint video + action objectives enabling cross-embodiment</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/f8d1eb7e/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Think Deeply Only When It Matters: DySL-VLA and Action-Aware Dynamic Depth</title>
      <itunes:title>Think Deeply Only When It Matters: DySL-VLA and Action-Aware Dynamic Depth</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">be431a93-bd1b-46df-9099-ccfeacd73489</guid>
      <link>https://share.transistor.fm/s/ef648a0b</link>
      <description>
        <![CDATA[Vision-Language-Action (VLA) models have shown remarkable success in robotic tasks like manipulation by fusing a language model's reasoning with a vision model's 3D understanding. However, their high computational cost remains...]]>
      </description>
      <content:encoded>
        <![CDATA[Vision-Language-Action (VLA) models have shown remarkable success in robotic tasks like manipulation by fusing a language model's reasoning with a vision model's 3D understanding. However, their high computational cost remains...]]>
      </content:encoded>
      <pubDate>Mon, 10 Aug 2026 14:10:38 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/ef648a0b/71498efd.mp3" length="31910808" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1995</itunes:duration>
      <itunes:summary>Vision-Language-Action (VLA) models have shown remarkable success in robotic tasks like manipulation by fusing a language model's reasoning with a vision model's 3D understanding. However, their high computational cost remains...</itunes:summary>
      <itunes:subtitle>Vision-Language-Action (VLA) models have shown remarkable success in robotic tasks like manipulation by fusing a language model's reasoning with a vision model's 3D understanding. However, their high computational cost remains...</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/ef648a0b/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>The 100,000-Hour Bet: How Xiaomi-Robotics-1 Tries to Make Robot Scaling Real</title>
      <itunes:title>The 100,000-Hour Bet: How Xiaomi-Robotics-1 Tries to Make Robot Scaling Real</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">7f4a30f9-5c6a-42dc-a34f-f4be6da3ef70</guid>
      <link>https://share.transistor.fm/s/e04fa49b</link>
      <description>
        <![CDATA[We present Xiaomi-Robotics-1, a foundational vision-language-action (VLA) model capable of (1) following diverse language instructions to perform a wide range of mobile manipulation tasks in unseen environments out-of-the-box, and (2) rapidly adapting to new tasks with minimal demonstrations through efficient fine-tuning.]]>
      </description>
      <content:encoded>
        <![CDATA[We present Xiaomi-Robotics-1, a foundational vision-language-action (VLA) model capable of (1) following diverse language instructions to perform a wide range of mobile manipulation tasks in unseen environments out-of-the-box, and (2) rapidly adapting to new tasks with minimal demonstrations through efficient fine-tuning.]]>
      </content:encoded>
      <pubDate>Mon, 10 Aug 2026 14:09:24 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/e04fa49b/8d3dd5b4.mp3" length="31581874" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1974</itunes:duration>
      <itunes:summary>We present Xiaomi-Robotics-1, a foundational vision-language-action (VLA) model capable of (1) following diverse language instructions to perform a wide range of mobile manipulation tasks in unseen environments out-of-the-box, and (2) rapidly adapting to new tasks with minimal demonstrations through efficient fine-tuning.</itunes:summary>
      <itunes:subtitle>We present Xiaomi-Robotics-1, a foundational vision-language-action (VLA) model capable of (1) following diverse language instructions to perform a wide range of mobile manipulation tasks in unseen environments out-of-the-box, and (2) rapidly adapting to </itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/e04fa49b/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Reason First, Move Second: A Modular Alternative to End-to-End Robot Policies</title>
      <itunes:title>Reason First, Move Second: A Modular Alternative to End-to-End Robot Policies</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">748aa10e-bcfb-485d-8e89-d826676e3521</guid>
      <link>https://share.transistor.fm/s/5cbbf48a</link>
      <description>
        <![CDATA[Explicit task reasoning empowering robotic manipulation]]>
      </description>
      <content:encoded>
        <![CDATA[Explicit task reasoning empowering robotic manipulation]]>
      </content:encoded>
      <pubDate>Mon, 10 Aug 2026 05:10:58 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/5cbbf48a/52eabefb.mp3" length="29905858" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1870</itunes:duration>
      <itunes:summary>Explicit task reasoning empowering robotic manipulation</itunes:summary>
      <itunes:subtitle>Explicit task reasoning empowering robotic manipulation</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/5cbbf48a/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>From BIM to the Build: A Multimodal Agent for Conversational Construction Robotics</title>
      <itunes:title>From BIM to the Build: A Multimodal Agent for Conversational Construction Robotics</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">6552ed29-4850-4eea-88f0-61cd92e67b07</guid>
      <link>https://share.transistor.fm/s/fa71f289</link>
      <description>
        <![CDATA[A multimodal foundation model-enabled agent for human–robot collaboration in construction]]>
      </description>
      <content:encoded>
        <![CDATA[A multimodal foundation model-enabled agent for human–robot collaboration in construction]]>
      </content:encoded>
      <pubDate>Mon, 10 Aug 2026 05:10:32 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/fa71f289/8b6dcb10.mp3" length="31829306" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1990</itunes:duration>
      <itunes:summary>A multimodal foundation model-enabled agent for human–robot collaboration in construction</itunes:summary>
      <itunes:subtitle>A multimodal foundation model-enabled agent for human–robot collaboration in construction</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/fa71f289/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>The Twin Watches, the QP Dodges: Predictive Safety for a UR10e</title>
      <itunes:title>The Twin Watches, the QP Dodges: Predictive Safety for a UR10e</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">f344fae8-c2a8-49df-adc3-888a6e5dcfa2</guid>
      <link>https://share.transistor.fm/s/9d4b5103</link>
      <description>
        <![CDATA[Safe human–robot collaboration remains a critical challenge in manufacturing. Traditional safety approaches, such as cages and proximity sensors, are often insufficient for dynamic human interaction. This paper presents a digital twin approach using deep learning for collision avoidance in industrial collaborative robot manipulation.]]>
      </description>
      <content:encoded>
        <![CDATA[Safe human–robot collaboration remains a critical challenge in manufacturing. Traditional safety approaches, such as cages and proximity sensors, are often insufficient for dynamic human interaction. This paper presents a digital twin approach using deep learning for collision avoidance in industrial collaborative robot manipulation.]]>
      </content:encoded>
      <pubDate>Sun, 09 Aug 2026 14:08:23 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/9d4b5103/d9dd3c3f.mp3" length="31703918" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1982</itunes:duration>
      <itunes:summary>Safe human–robot collaboration remains a critical challenge in manufacturing. Traditional safety approaches, such as cages and proximity sensors, are often insufficient for dynamic human interaction. This paper presents a digital twin approach using deep learning for collision avoidance in industrial collaborative robot manipulation.</itunes:summary>
      <itunes:subtitle>Safe human–robot collaboration remains a critical challenge in manufacturing. Traditional safety approaches, such as cages and proximity sensors, are often insufficient for dynamic human interaction. This paper presents a digital twin approach using deep </itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/9d4b5103/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>One Shot, Twenty Screws: Robot Vision for the Dirty Reality of Appliance Recycling</title>
      <itunes:title>One Shot, Twenty Screws: Robot Vision for the Dirty Reality of Appliance Recycling</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">f1b53253-b55e-421c-95fa-be306c970512</guid>
      <link>https://share.transistor.fm/s/4f5a8dce</link>
      <description>
        <![CDATA[Industrial-grade robust robot vision system for screw detection and removal under uneven conditions.]]>
      </description>
      <content:encoded>
        <![CDATA[Industrial-grade robust robot vision system for screw detection and removal under uneven conditions.]]>
      </content:encoded>
      <pubDate>Sun, 09 Aug 2026 14:08:02 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/4f5a8dce/92a38a13.mp3" length="29964790" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1873</itunes:duration>
      <itunes:summary>Industrial-grade robust robot vision system for screw detection and removal under uneven conditions.</itunes:summary>
      <itunes:subtitle>Industrial-grade robust robot vision system for screw detection and removal under uneven conditions.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/4f5a8dce/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Better Geometry at the Neural–Symbolic Boundary: DAIoU for Robot Manipulation</title>
      <itunes:title>Better Geometry at the Neural–Symbolic Boundary: DAIoU for Robot Manipulation</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">82614264-ca5d-4e7c-bc5d-c8f8b826a091</guid>
      <link>https://share.transistor.fm/s/e3dfa591</link>
      <description>
        <![CDATA[Distance–area–IoU fusion loss for better neuro-symbolic robot manipulation]]>
      </description>
      <content:encoded>
        <![CDATA[Distance–area–IoU fusion loss for better neuro-symbolic robot manipulation]]>
      </content:encoded>
      <pubDate>Sun, 09 Aug 2026 05:10:53 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/e3dfa591/139deaca.mp3" length="30987954" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1937</itunes:duration>
      <itunes:summary>Distance–area–IoU fusion loss for better neuro-symbolic robot manipulation</itunes:summary>
      <itunes:subtitle>Distance–area–IoU fusion loss for better neuro-symbolic robot manipulation</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/e3dfa591/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>REL Hand: A Mechanical Shortcut to More Human-Like Humanoid Dexterity</title>
      <itunes:title>REL Hand: A Mechanical Shortcut to More Human-Like Humanoid Dexterity</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">1b9507c5-e2f4-4b16-ab8e-a75fd7aad593</guid>
      <link>https://share.transistor.fm/s/71f437d1</link>
      <description>
        <![CDATA[Design and evaluation of a tendon-and-linkage hybrid-driven humanoid dexterous hand]]>
      </description>
      <content:encoded>
        <![CDATA[Design and evaluation of a tendon-and-linkage hybrid-driven humanoid dexterous hand]]>
      </content:encoded>
      <pubDate>Sun, 09 Aug 2026 05:10:30 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/71f437d1/fa418674.mp3" length="31617819" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1977</itunes:duration>
      <itunes:summary>Design and evaluation of a tendon-and-linkage hybrid-driven humanoid dexterous hand</itunes:summary>
      <itunes:subtitle>Design and evaluation of a tendon-and-linkage hybrid-driven humanoid dexterous hand</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/71f437d1/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>When the Appliance Becomes the API: Part–Function–State Models for Robot Planning</title>
      <itunes:title>When the Appliance Becomes the API: Part–Function–State Models for Robot Planning</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">0d8e3162-f1ba-48a8-b016-06887a432ed0</guid>
      <link>https://share.transistor.fm/s/ea00c1f8</link>
      <description>
        <![CDATA[Appliances Describe Themselves: Part–Function–State Modeling for Appliance Manipulation Planning]]>
      </description>
      <content:encoded>
        <![CDATA[Appliances Describe Themselves: Part–Function–State Modeling for Appliance Manipulation Planning]]>
      </content:encoded>
      <pubDate>Sat, 08 Aug 2026 14:09:13 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/ea00c1f8/c125c587.mp3" length="29341195" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1834</itunes:duration>
      <itunes:summary>Appliances Describe Themselves: Part–Function–State Modeling for Appliance Manipulation Planning</itunes:summary>
      <itunes:subtitle>Appliances Describe Themselves: Part–Function–State Modeling for Appliance Manipulation Planning</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/ea00c1f8/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>CISMG-Nav: When a Navigation Agent Stops Forgetting the Building</title>
      <itunes:title>CISMG-Nav: When a Navigation Agent Stops Forgetting the Building</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">426c3a28-3353-4f39-8a28-671a5ad7fdbf</guid>
      <link>https://share.transistor.fm/s/8665cec2</link>
      <description>
        <![CDATA[CISMG-Nav: A Cross-Task Incremental Semantic Memory Graph-Driven Vision-and-Language Navigation Framework]]>
      </description>
      <content:encoded>
        <![CDATA[CISMG-Nav: A Cross-Task Incremental Semantic Memory Graph-Driven Vision-and-Language Navigation Framework]]>
      </content:encoded>
      <pubDate>Sat, 08 Aug 2026 05:12:12 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/8665cec2/d1ee3dff.mp3" length="28502351" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1782</itunes:duration>
      <itunes:summary>CISMG-Nav: A Cross-Task Incremental Semantic Memory Graph-Driven Vision-and-Language Navigation Framework</itunes:summary>
      <itunes:subtitle>CISMG-Nav: A Cross-Task Incremental Semantic Memory Graph-Driven Vision-and-Language Navigation Framework</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/8665cec2/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>ReaDy-Go: Teaching Robots to Dodge Photorealistic People in Gaussian-Splat Digital Twins</title>
      <itunes:title>ReaDy-Go: Teaching Robots to Dodge Photorealistic People in Gaussian-Splat Digital Twins</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">9a648c74-fa00-4b6d-8ce9-66112157f84d</guid>
      <link>https://share.transistor.fm/s/556de66f</link>
      <description>
        <![CDATA[ReaDy-Go: Real-to-Sim Dynamic 3D Gaussian Splatting Simulation for Environment-Specific Visual Navigation With Moving Obstacles]]>
      </description>
      <content:encoded>
        <![CDATA[ReaDy-Go: Real-to-Sim Dynamic 3D Gaussian Splatting Simulation for Environment-Specific Visual Navigation With Moving Obstacles]]>
      </content:encoded>
      <pubDate>Sat, 08 Aug 2026 05:10:12 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/556de66f/c97a3e4c.mp3" length="33769055" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2111</itunes:duration>
      <itunes:summary>ReaDy-Go: Real-to-Sim Dynamic 3D Gaussian Splatting Simulation for Environment-Specific Visual Navigation With Moving Obstacles</itunes:summary>
      <itunes:subtitle>ReaDy-Go: Real-to-Sim Dynamic 3D Gaussian Splatting Simulation for Environment-Specific Visual Navigation With Moving Obstacles</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/556de66f/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>The Action Is Missing: Turning Human Video into Robot Control</title>
      <itunes:title>The Action Is Missing: Turning Human Video into Robot Control</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">aeba0e56-346b-4557-9d7a-bbddaae2d74b</guid>
      <link>https://share.transistor.fm/s/be2f9ed3</link>
      <description>
        <![CDATA[Recent progress in generalizable embodied control has been driven by large-scale pretraining of Vision-Language-Action (VLA) models. However, most existing approaches rely on large collections of robot demonstrations, which are costly. This survey covers scalable vision-language-action learning using human-centric data, particularly human videos, as an alternative to expensive robot demonstrations.]]>
      </description>
      <content:encoded>
        <![CDATA[Recent progress in generalizable embodied control has been driven by large-scale pretraining of Vision-Language-Action (VLA) models. However, most existing approaches rely on large collections of robot demonstrations, which are costly. This survey covers scalable vision-language-action learning using human-centric data, particularly human videos, as an alternative to expensive robot demonstrations.]]>
      </content:encoded>
      <pubDate>Fri, 07 Aug 2026 14:07:43 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/be2f9ed3/2a7d6e45.mp3" length="32867517" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2055</itunes:duration>
      <itunes:summary>Recent progress in generalizable embodied control has been driven by large-scale pretraining of Vision-Language-Action (VLA) models. However, most existing approaches rely on large collections of robot demonstrations, which are costly. This survey covers scalable vision-language-action learning using human-centric data, particularly human videos, as an alternative to expensive robot demonstrations.</itunes:summary>
      <itunes:subtitle>Recent progress in generalizable embodied control has been driven by large-scale pretraining of Vision-Language-Action (VLA) models. However, most existing approaches rely on large collections of robot demonstrations, which are costly. This survey covers </itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/be2f9ed3/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Look Where You Mean: Gaze2Act Gives VLAs a Live Intent Channel</title>
      <itunes:title>Look Where You Mean: Gaze2Act Gives VLAs a Live Intent Channel</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">8ba5b9d2-13cb-403d-bcf8-fefb77045265</guid>
      <link>https://share.transistor.fm/s/27c14246</link>
      <description>
        <![CDATA[Vision-Language-Action (VLA) models have recently shown strong potential for robot learning by following language instructions. However, in practice, language alone is often insufficient to precisely convey human intent.]]>
      </description>
      <content:encoded>
        <![CDATA[Vision-Language-Action (VLA) models have recently shown strong potential for robot learning by following language instructions. However, in practice, language alone is often insufficient to precisely convey human intent.]]>
      </content:encoded>
      <pubDate>Fri, 07 Aug 2026 05:41:28 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/27c14246/8d0e7139.mp3" length="29893319" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1869</itunes:duration>
      <itunes:summary>Vision-Language-Action (VLA) models have recently shown strong potential for robot learning by following language instructions. However, in practice, language alone is often insufficient to precisely convey human intent.</itunes:summary>
      <itunes:subtitle>Vision-Language-Action (VLA) models have recently shown strong potential for robot learning by following language instructions. However, in practice, language alone is often insufficient to precisely convey human intent.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/27c14246/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>HARP-VLA: Align the Eyes Before the Actions</title>
      <itunes:title>HARP-VLA: Align the Eyes Before the Actions</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">1c21616b-5b8d-4333-b97e-76c12ecfc9f8</guid>
      <link>https://share.transistor.fm/s/d17f4160</link>
      <description>
        <![CDATA[Learning generalizable vision-language-action (VLA) models from large-scale human videos is promising but challenging due to cross-embodiment discrepancies in both visual observations and executable actions.]]>
      </description>
      <content:encoded>
        <![CDATA[Learning generalizable vision-language-action (VLA) models from large-scale human videos is promising but challenging due to cross-embodiment discrepancies in both visual observations and executable actions.]]>
      </content:encoded>
      <pubDate>Fri, 07 Aug 2026 05:17:16 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/d17f4160/06b55d9d.mp3" length="25478834" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1593</itunes:duration>
      <itunes:summary>Learning generalizable vision-language-action (VLA) models from large-scale human videos is promising but challenging due to cross-embodiment discrepancies in both visual observations and executable actions.</itunes:summary>
      <itunes:subtitle>Learning generalizable vision-language-action (VLA) models from large-scale human videos is promising but challenging due to cross-embodiment discrepancies in both visual observations and executable actions.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/d17f4160/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Make Geometry the Policy: GAM's Bet on 3D-Native Robot Control</title>
      <itunes:title>Make Geometry the Policy: GAM's Bet on 3D-Native Robot Control</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">8d92a074-d30f-4fa3-901c-1bd35d8b7df6</guid>
      <link>https://share.transistor.fm/s/c2771525</link>
      <description>
        <![CDATA[Generalist robot policies must follow user instructions while reasoning about how objects, cameras, and robot actions interact in the 3D physical world. Recent VLAs and video world-action models lack explicit 3D geometric reasoning.]]>
      </description>
      <content:encoded>
        <![CDATA[Generalist robot policies must follow user instructions while reasoning about how objects, cameras, and robot actions interact in the 3D physical world. Recent VLAs and video world-action models lack explicit 3D geometric reasoning.]]>
      </content:encoded>
      <pubDate>Thu, 06 Aug 2026 14:08:33 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/c2771525/b62320ab.mp3" length="31149287" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1947</itunes:duration>
      <itunes:summary>Generalist robot policies must follow user instructions while reasoning about how objects, cameras, and robot actions interact in the 3D physical world. Recent VLAs and video world-action models lack explicit 3D geometric reasoning.</itunes:summary>
      <itunes:subtitle>Generalist robot policies must follow user instructions while reasoning about how objects, cameras, and robot actions interact in the 3D physical world. Recent VLAs and video world-action models lack explicit 3D geometric reasoning.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/c2771525/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>π0.7: When a Robot Foundation Model Learns How to Behave</title>
      <itunes:title>π0.7: When a Robot Foundation Model Learns How to Behave</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">4b6e83f1-4ed0-49bc-84dd-882479b861eb</guid>
      <link>https://share.transistor.fm/s/a1199918</link>
      <description>
        <![CDATA[We present a new robotic foundation model, π0.7, that can enable strong out-of-the-box performance in a wide range of scenarios, following diverse language instructions in unseen environments.]]>
      </description>
      <content:encoded>
        <![CDATA[We present a new robotic foundation model, π0.7, that can enable strong out-of-the-box performance in a wide range of scenarios, following diverse language instructions in unseen environments.]]>
      </content:encoded>
      <pubDate>Thu, 06 Aug 2026 14:08:20 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/a1199918/08aaf7ef.mp3" length="30513571" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1908</itunes:duration>
      <itunes:summary>We present a new robotic foundation model, π0.7, that can enable strong out-of-the-box performance in a wide range of scenarios, following diverse language instructions in unseen environments.</itunes:summary>
      <itunes:subtitle>We present a new robotic foundation model, π0.7, that can enable strong out-of-the-box performance in a wide range of scenarios, following diverse language instructions in unseen environments.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/a1199918/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>EgoHumanoid: Turning Human First-Person Experience into Humanoid Loco-Manipulation</title>
      <itunes:title>EgoHumanoid: Turning Human First-Person Experience into Humanoid Loco-Manipulation</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">4c321cb1-ff97-40ca-9ada-180a2bf3536a</guid>
      <link>https://share.transistor.fm/s/12c655d1</link>
      <description>
        <![CDATA[Introduces a dataset and method for learning loco-manipulation policies directly from egocentric human videos without any robot data collection, enabling real-world humanoid deployment in unstructured environments.]]>
      </description>
      <content:encoded>
        <![CDATA[Introduces a dataset and method for learning loco-manipulation policies directly from egocentric human videos without any robot data collection, enabling real-world humanoid deployment in unstructured environments.]]>
      </content:encoded>
      <pubDate>Thu, 06 Aug 2026 05:06:51 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/12c655d1/5bc3433b.mp3" length="30714191" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1920</itunes:duration>
      <itunes:summary>Introduces a dataset and method for learning loco-manipulation policies directly from egocentric human videos without any robot data collection, enabling real-world humanoid deployment in unstructured environments.</itunes:summary>
      <itunes:subtitle>Introduces a dataset and method for learning loco-manipulation policies directly from egocentric human videos without any robot data collection, enabling real-world humanoid deployment in unstructured environments.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/12c655d1/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>CycleRL: Teaching a Bicycle to Balance in Simulation—and Trusting It on Asphalt</title>
      <itunes:title>CycleRL: Teaching a Bicycle to Balance in Simulation—and Trusting It on Asphalt</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">2a360164-0c3d-451c-b52f-919d74db0d78</guid>
      <link>https://share.transistor.fm/s/4fb035d8</link>
      <description>
        <![CDATA[CycleRL: Sim-to-Real Deep Reinforcement Learning for Robust Autonomous Bicycle Control]]>
      </description>
      <content:encoded>
        <![CDATA[CycleRL: Sim-to-Real Deep Reinforcement Learning for Robust Autonomous Bicycle Control]]>
      </content:encoded>
      <pubDate>Wed, 05 Aug 2026 14:09:42 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/4fb035d8/5e189f70.mp3" length="31606952" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1976</itunes:duration>
      <itunes:summary>CycleRL: Sim-to-Real Deep Reinforcement Learning for Robust Autonomous Bicycle Control</itunes:summary>
      <itunes:subtitle>CycleRL: Sim-to-Real Deep Reinforcement Learning for Robust Autonomous Bicycle Control</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/4fb035d8/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Feeling the World Without a Force Sensor: UniFP for Legged Loco-Manipulation</title>
      <itunes:title>Feeling the World Without a Force Sensor: UniFP for Legged Loco-Manipulation</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">89f1b2c1-321b-4523-a18f-a09e933ff0d2</guid>
      <link>https://share.transistor.fm/s/f048b986</link>
      <description>
        <![CDATA[Proposes a unified RL policy that jointly learns force and position control in simulation without force sensors, enabling compliant behaviors and force-aware imitation learning. Includes a force estimator trained as an auxiliary task and integration with Diffusion Policy for contact-rich tasks on quadrupeds and humanoids.]]>
      </description>
      <content:encoded>
        <![CDATA[Proposes a unified RL policy that jointly learns force and position control in simulation without force sensors, enabling compliant behaviors and force-aware imitation learning. Includes a force estimator trained as an auxiliary task and integration with Diffusion Policy for contact-rich tasks on quadrupeds and humanoids.]]>
      </content:encoded>
      <pubDate>Wed, 05 Aug 2026 14:08:36 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/f048b986/2320b7ab.mp3" length="27245966" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1703</itunes:duration>
      <itunes:summary>Proposes a unified RL policy that jointly learns force and position control in simulation without force sensors, enabling compliant behaviors and force-aware imitation learning. Includes a force estimator trained as an auxiliary task and integration with Diffusion Policy for contact-rich tasks on quadrupeds and humanoids.</itunes:summary>
      <itunes:subtitle>Proposes a unified RL policy that jointly learns force and position control in simulation without force sensors, enabling compliant behaviors and force-aware imitation learning. Includes a force estimator trained as an auxiliary task and integration with </itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/f048b986/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>The Robot Needs a Memory: Multi-Agent Reasoning Inside a Digital Twin</title>
      <itunes:title>The Robot Needs a Memory: Multi-Agent Reasoning Inside a Digital Twin</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">4f5afd18-f62d-48ec-b41a-4d759f0b2045</guid>
      <link>https://share.transistor.fm/s/a4e0f309</link>
      <description>
        <![CDATA[Digital Twin-enabled adaptive robotics: multi-agent reasoning over language, vision, and structured database]]>
      </description>
      <content:encoded>
        <![CDATA[Digital Twin-enabled adaptive robotics: multi-agent reasoning over language, vision, and structured database]]>
      </content:encoded>
      <pubDate>Wed, 05 Aug 2026 05:13:34 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/a4e0f309/d4ed7dfa.mp3" length="27806449" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1738</itunes:duration>
      <itunes:summary>Digital Twin-enabled adaptive robotics: multi-agent reasoning over language, vision, and structured database</itunes:summary>
      <itunes:subtitle>Digital Twin-enabled adaptive robotics: multi-agent reasoning over language, vision, and structured database</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/a4e0f309/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Language Is Not Just a Prompt: Four Ways Robots Turn Words into Action</title>
      <itunes:title>Language Is Not Just a Prompt: Four Ways Robots Turn Words into Action</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">5dde641e-9053-4f3a-8bfc-55b0c1fef92e</guid>
      <link>https://share.transistor.fm/s/a92710dd</link>
      <description>
        <![CDATA[Language-conditioned robot manipulation is an emerging field aimed at enabling seamless communication and cooperation between humans and robotic agents by teaching robots to comprehend and execute instructions conveyed in natural language.]]>
      </description>
      <content:encoded>
        <![CDATA[Language-conditioned robot manipulation is an emerging field aimed at enabling seamless communication and cooperation between humans and robotic agents by teaching robots to comprehend and execute instructions conveyed in natural language.]]>
      </content:encoded>
      <pubDate>Wed, 05 Aug 2026 05:12:40 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/a92710dd/26620aa5.mp3" length="33212751" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2076</itunes:duration>
      <itunes:summary>Language-conditioned robot manipulation is an emerging field aimed at enabling seamless communication and cooperation between humans and robotic agents by teaching robots to comprehend and execute instructions conveyed in natural language.</itunes:summary>
      <itunes:subtitle>Language-conditioned robot manipulation is an emerging field aimed at enabling seamless communication and cooperation between humans and robotic agents by teaching robots to comprehend and execute instructions conveyed in natural language.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/a92710dd/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Becoming the Path: Planning Motion, Shape, and Stiffness Together</title>
      <itunes:title>Becoming the Path: Planning Motion, Shape, and Stiffness Together</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">0da97bcd-d782-413c-8e90-44c68a03db5e</guid>
      <link>https://share.transistor.fm/s/244655cf</link>
      <description>
        <![CDATA[Autonomous operation in confined spaces demands robots that can simultaneously navigate tight passages and manipulate objects. However, conventional mobile manipulators are often too large, while compliant soft manipulators are typically too weak.]]>
      </description>
      <content:encoded>
        <![CDATA[Autonomous operation in confined spaces demands robots that can simultaneously navigate tight passages and manipulate objects. However, conventional mobile manipulators are often too large, while compliant soft manipulators are typically too weak.]]>
      </content:encoded>
      <pubDate>Wed, 05 Aug 2026 03:12:44 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/244655cf/9fe300ff.mp3" length="30484732" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1906</itunes:duration>
      <itunes:summary>Autonomous operation in confined spaces demands robots that can simultaneously navigate tight passages and manipulate objects. However, conventional mobile manipulators are often too large, while compliant soft manipulators are typically too weak.</itunes:summary>
      <itunes:subtitle>Autonomous operation in confined spaces demands robots that can simultaneously navigate tight passages and manipulate objects. However, conventional mobile manipulators are often too large, while compliant soft manipulators are typically too weak.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/244655cf/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Meaning, Motion, and Two Hands: A Task Graph That Actually Plans</title>
      <itunes:title>Meaning, Motion, and Two Hands: A Task Graph That Actually Plans</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">8de46349-eb88-4da1-80f7-698b6636c96f</guid>
      <link>https://share.transistor.fm/s/2218d12c</link>
      <description>
        <![CDATA[Semantic–Geometric Task Representations for Bimanual Manipulation From Human Demonstrations to Robot Action Planning]]>
      </description>
      <content:encoded>
        <![CDATA[Semantic–Geometric Task Representations for Bimanual Manipulation From Human Demonstrations to Robot Action Planning]]>
      </content:encoded>
      <pubDate>Wed, 05 Aug 2026 03:09:56 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/2218d12c/66912b1c.mp3" length="31683438" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1981</itunes:duration>
      <itunes:summary>Semantic–Geometric Task Representations for Bimanual Manipulation From Human Demonstrations to Robot Action Planning</itunes:summary>
      <itunes:subtitle>Semantic–Geometric Task Representations for Bimanual Manipulation From Human Demonstrations to Robot Action Planning</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/2218d12c/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Safety Before Failure: Giving Robot Plans a One-Second Reflex</title>
      <itunes:title>Safety Before Failure: Giving Robot Plans a One-Second Reflex</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">4b7a681e-7dc6-444b-be8f-4ede912bdbf1</guid>
      <link>https://share.transistor.fm/s/c09f8e42</link>
      <description>
        <![CDATA[Robots that execute language-conditioned tasks in dynamic environments often rely on feedback only after an action has failed, which can be insufficient when failures involve collisions or workspace conflicts.]]>
      </description>
      <content:encoded>
        <![CDATA[Robots that execute language-conditioned tasks in dynamic environments often rely on feedback only after an action has failed, which can be insufficient when failures involve collisions or workspace conflicts.]]>
      </content:encoded>
      <pubDate>Tue, 04 Aug 2026 14:10:05 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/c09f8e42/cc357976.mp3" length="29813489" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1864</itunes:duration>
      <itunes:summary>Robots that execute language-conditioned tasks in dynamic environments often rely on feedback only after an action has failed, which can be insufficient when failures involve collisions or workspace conflicts.</itunes:summary>
      <itunes:subtitle>Robots that execute language-conditioned tasks in dynamic environments often rely on feedback only after an action has failed, which can be insufficient when failures involve collisions or workspace conflicts.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/c09f8e42/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>VL-GRiP3 and the Case for Keeping VLMs Away from the Motor Loop</title>
      <itunes:title>VL-GRiP3 and the Case for Keeping VLMs Away from the Motor Loop</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">8b2b4108-dd89-49c8-b2ae-4e13591dc9f1</guid>
      <link>https://share.transistor.fm/s/31feeb07</link>
      <description>
        <![CDATA[VL-GRiP3: A hierarchical pipeline leveraging vision-language models for autonomous robotic 3D grasping]]>
      </description>
      <content:encoded>
        <![CDATA[VL-GRiP3: A hierarchical pipeline leveraging vision-language models for autonomous robotic 3D grasping]]>
      </content:encoded>
      <pubDate>Tue, 04 Aug 2026 14:09:46 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/31feeb07/e8ea9fff.mp3" length="27625891" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1727</itunes:duration>
      <itunes:summary>VL-GRiP3: A hierarchical pipeline leveraging vision-language models for autonomous robotic 3D grasping</itunes:summary>
      <itunes:subtitle>VL-GRiP3: A hierarchical pipeline leveraging vision-language models for autonomous robotic 3D grasping</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/31feeb07/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Think Late, Act Now: TIC-VLA Turns Reasoning Latency into a Control Variable</title>
      <itunes:title>Think Late, Act Now: TIC-VLA Turns Reasoning Latency into a Control Variable</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">7e43d0ed-e80e-4f85-927e-b6fde7e6fab4</guid>
      <link>https://share.transistor.fm/s/41abf345</link>
      <description>
        <![CDATA[Robots in dynamic, human-centric environments must follow language instructions while maintaining real-time reactive control. Vision-language-action (VLA) models offer a promising framework, but they assume temporally aligned reasoning and control.]]>
      </description>
      <content:encoded>
        <![CDATA[Robots in dynamic, human-centric environments must follow language instructions while maintaining real-time reactive control. Vision-language-action (VLA) models offer a promising framework, but they assume temporally aligned reasoning and control.]]>
      </content:encoded>
      <pubDate>Tue, 04 Aug 2026 05:09:18 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/41abf345/adae1449.mp3" length="29350390" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1835</itunes:duration>
      <itunes:summary>Robots in dynamic, human-centric environments must follow language instructions while maintaining real-time reactive control. Vision-language-action (VLA) models offer a promising framework, but they assume temporally aligned reasoning and control.</itunes:summary>
      <itunes:subtitle>Robots in dynamic, human-centric environments must follow language instructions while maintaining real-time reactive control. Vision-language-action (VLA) models offer a promising framework, but they assume temporally aligned reasoning and control.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/41abf345/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>One Policy, Many Bodies: Inside Qwen-VLA's Generalist Robot Model</title>
      <itunes:title>One Policy, Many Bodies: Inside Qwen-VLA's Generalist Robot Model</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">022cf6c1-5694-44c2-b8d9-cf7ebde9357d</guid>
      <link>https://share.transistor.fm/s/746dd52b</link>
      <description>
        <![CDATA[Embodied intelligence is often studied through specialized models for individual tasks such as manipulation or navigation, resulting in fragmented capabilities and limited generalization across tasks, environments, and robot embodiments.]]>
      </description>
      <content:encoded>
        <![CDATA[Embodied intelligence is often studied through specialized models for individual tasks such as manipulation or navigation, resulting in fragmented capabilities and limited generalization across tasks, environments, and robot embodiments.]]>
      </content:encoded>
      <pubDate>Tue, 04 Aug 2026 05:08:53 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/746dd52b/a4f0d086.mp3" length="29363765" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1836</itunes:duration>
      <itunes:summary>Embodied intelligence is often studied through specialized models for individual tasks such as manipulation or navigation, resulting in fragmented capabilities and limited generalization across tasks, environments, and robot embodiments.</itunes:summary>
      <itunes:subtitle>Embodied intelligence is often studied through specialized models for individual tasks such as manipulation or navigation, resulting in fragmented capabilities and limited generalization across tasks, environments, and robot embodiments.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/746dd52b/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>VLAbot: A Human-in-the-Loop Control Plane for Long-Horizon Robot Assembly</title>
      <itunes:title>VLAbot: A Human-in-the-Loop Control Plane for Long-Horizon Robot Assembly</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">8d3c90d9-5d77-444e-9f6b-eda611149d7d</guid>
      <link>https://share.transistor.fm/s/fb8ebc97</link>
      <description>
        <![CDATA[VLAbot: A human Vision–Language–Action models interaction framework for robotic assembly]]>
      </description>
      <content:encoded>
        <![CDATA[VLAbot: A human Vision–Language–Action models interaction framework for robotic assembly]]>
      </content:encoded>
      <pubDate>Tue, 04 Aug 2026 03:13:39 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/fb8ebc97/4efd9d44.mp3" length="30702488" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1919</itunes:duration>
      <itunes:summary>VLAbot: A human Vision–Language–Action models interaction framework for robotic assembly</itunes:summary>
      <itunes:subtitle>VLAbot: A human Vision–Language–Action models interaction framework for robotic assembly</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/fb8ebc97/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Teaching Isaac Sim to Feel: A Real-Time Digital Twin for Capacitive Touch</title>
      <itunes:title>Teaching Isaac Sim to Feel: A Real-Time Digital Twin for Capacitive Touch</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">2e00d236-d39c-4347-a82d-4a108418fcaa</guid>
      <link>https://share.transistor.fm/s/0dd225d9</link>
      <description>
        <![CDATA[With advances in robotic manipulation in recent years, tactile sensing has become increasingly important in scenarios where visual information is unreliable or insufficient.]]>
      </description>
      <content:encoded>
        <![CDATA[With advances in robotic manipulation in recent years, tactile sensing has become increasingly important in scenarios where visual information is unreliable or insufficient.]]>
      </content:encoded>
      <pubDate>Tue, 04 Aug 2026 03:12:11 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/0dd225d9/709c90d8.mp3" length="28684581" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1793</itunes:duration>
      <itunes:summary>With advances in robotic manipulation in recent years, tactile sensing has become increasingly important in scenarios where visual information is unreliable or insufficient.</itunes:summary>
      <itunes:subtitle>With advances in robotic manipulation in recent years, tactile sensing has become increasingly important in scenarios where visual information is unreliable or insufficient.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/0dd225d9/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Stop Starting From Noise: Action-to-Action Flow Matching for Fast Robot Control</title>
      <itunes:title>Stop Starting From Noise: Action-to-Action Flow Matching for Fast Robot Control</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">2489a9f4-7d1c-4d1d-9cc5-22216f6a2d91</guid>
      <link>https://share.transistor.fm/s/66e71cc9</link>
      <description>
        <![CDATA[Proposes a new flow-matching approach for robot action generation that improves efficiency in learning and repurposing generalist policies. Targets improved scalability and adaptability of robot action models.]]>
      </description>
      <content:encoded>
        <![CDATA[Proposes a new flow-matching approach for robot action generation that improves efficiency in learning and repurposing generalist policies. Targets improved scalability and adaptability of robot action models.]]>
      </content:encoded>
      <pubDate>Mon, 03 Aug 2026 14:08:02 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/66e71cc9/75fd3c14.mp3" length="31822201" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1989</itunes:duration>
      <itunes:summary>Proposes a new flow-matching approach for robot action generation that improves efficiency in learning and repurposing generalist policies. Targets improved scalability and adaptability of robot action models.</itunes:summary>
      <itunes:subtitle>Proposes a new flow-matching approach for robot action generation that improves efficiency in learning and repurposing generalist policies. Targets improved scalability and adaptability of robot action models.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/66e71cc9/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>ACE-Data-0: Turning a Home into a Multimodal Robot-Learning Instrument</title>
      <itunes:title>ACE-Data-0: Turning a Home into a Multimodal Robot-Learning Instrument</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">1a199252-505c-4935-815e-d567b5d5c8a5</guid>
      <link>https://share.transistor.fm/s/a454fb6a</link>
      <description>
        <![CDATA[An embodied data engine releasing 150 hours of synchronized egocentric video, full-body motion, hand tracking, object trajectories, audio, and tactile signals from everyday household activities.]]>
      </description>
      <content:encoded>
        <![CDATA[An embodied data engine releasing 150 hours of synchronized egocentric video, full-body motion, hand tracking, object trajectories, audio, and tactile signals from everyday household activities.]]>
      </content:encoded>
      <pubDate>Mon, 03 Aug 2026 05:19:45 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/a454fb6a/db5a49e9.mp3" length="32101398" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2007</itunes:duration>
      <itunes:summary>An embodied data engine releasing 150 hours of synchronized egocentric video, full-body motion, hand tracking, object trajectories, audio, and tactile signals from everyday household activities.</itunes:summary>
      <itunes:subtitle>An embodied data engine releasing 150 hours of synchronized egocentric video, full-body motion, hand tracking, object trajectories, audio, and tactile signals from everyday household activities.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/a454fb6a/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Before the Robot Feels It: How N₀-VTLA Predicts Contact to Control It</title>
      <itunes:title>Before the Robot Feels It: How N₀-VTLA Predicts Contact to Control It</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">0a76745e-638a-4ac7-b86b-6ed635e1eb25</guid>
      <link>https://share.transistor.fm/s/cb645bd7</link>
      <description>
        <![CDATA[A vision-tactile-language-action model that predicts latent tactile tokens for fine-grained contact-rich manipulation, achieving wins on all 9 real-robot tasks and 63.8% on 20 simulation tasks.]]>
      </description>
      <content:encoded>
        <![CDATA[A vision-tactile-language-action model that predicts latent tactile tokens for fine-grained contact-rich manipulation, achieving wins on all 9 real-robot tasks and 63.8% on 20 simulation tasks.]]>
      </content:encoded>
      <pubDate>Mon, 03 Aug 2026 05:17:21 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/cb645bd7/74a4c933.mp3" length="32827811" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2052</itunes:duration>
      <itunes:summary>A vision-tactile-language-action model that predicts latent tactile tokens for fine-grained contact-rich manipulation, achieving wins on all 9 real-robot tasks and 63.8% on 20 simulation tasks.</itunes:summary>
      <itunes:subtitle>A vision-tactile-language-action model that predicts latent tactile tokens for fine-grained contact-rich manipulation, achieving wins on all 9 real-robot tasks and 63.8% on 20 simulation tasks.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/cb645bd7/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Metis Gives Foundation Models a Native, Mutable Memory</title>
      <itunes:title>Metis Gives Foundation Models a Native, Mutable Memory</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">8c3f30b8-2751-409f-98b2-b5971dec2d27</guid>
      <link>https://share.transistor.fm/s/58078b0f</link>
      <description>
        <![CDATA[Equips foundation models with persistent, dynamically evolving native memory state and gradient-free online updates, enabling continual adaptation without retraining.]]>
      </description>
      <content:encoded>
        <![CDATA[Equips foundation models with persistent, dynamically evolving native memory state and gradient-free online updates, enabling continual adaptation without retraining.]]>
      </content:encoded>
      <pubDate>Mon, 03 Aug 2026 03:10:26 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/58078b0f/072754bd.mp3" length="33098648" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2069</itunes:duration>
      <itunes:summary>Equips foundation models with persistent, dynamically evolving native memory state and gradient-free online updates, enabling continual adaptation without retraining.</itunes:summary>
      <itunes:subtitle>Equips foundation models with persistent, dynamically evolving native memory state and gradient-free online updates, enabling continual adaptation without retraining.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/58078b0f/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>TurboVLA and the Case for Taking the LLM Out of the Servo Loop</title>
      <itunes:title>TurboVLA and the Case for Taking the LLM Out of the Servo Loop</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">51e407d3-ddd6-4732-a60f-bfb053410b3e</guid>
      <link>https://share.transistor.fm/s/57443c98</link>
      <description>
        <![CDATA[Achieves high-speed inference for vision-language-action policies at 32 Hz on an RTX 4090 with less than 1 GB VRAM using a 0.2B parameter model, enabling real-time VLA deployment on consumer hardware.]]>
      </description>
      <content:encoded>
        <![CDATA[Achieves high-speed inference for vision-language-action policies at 32 Hz on an RTX 4090 with less than 1 GB VRAM using a 0.2B parameter model, enabling real-time VLA deployment on consumer hardware.]]>
      </content:encoded>
      <pubDate>Mon, 03 Aug 2026 03:08:47 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/57443c98/77507b90.mp3" length="30218074" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1889</itunes:duration>
      <itunes:summary>Achieves high-speed inference for vision-language-action policies at 32 Hz on an RTX 4090 with less than 1 GB VRAM using a 0.2B parameter model, enabling real-time VLA deployment on consumer hardware.</itunes:summary>
      <itunes:subtitle>Achieves high-speed inference for vision-language-action policies at 32 Hz on an RTX 4090 with less than 1 GB VRAM using a 0.2B parameter model, enabling real-time VLA deployment on consumer hardware.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/57443c98/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Catching on Optical Expansion: Inside Pixel2Catch's Monocular Sim-to-Real Pipeline</title>
      <itunes:title>Catching on Optical Expansion: Inside Pixel2Catch's Monocular Sim-to-Real Pipeline</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">ab8f609b-2cf8-4510-b7ef-99fcd6a88da3</guid>
      <link>https://share.transistor.fm/s/83783b0c</link>
      <description>
        <![CDATA[Pixel2Catch: Multi-Agent Sim-to-Real Transfer for Agile Manipulation With a Single RGB Camera]]>
      </description>
      <content:encoded>
        <![CDATA[Pixel2Catch: Multi-Agent Sim-to-Real Transfer for Agile Manipulation With a Single RGB Camera]]>
      </content:encoded>
      <pubDate>Sun, 02 Aug 2026 14:09:45 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/83783b0c/627b0aea.mp3" length="32434093" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2028</itunes:duration>
      <itunes:summary>Pixel2Catch: Multi-Agent Sim-to-Real Transfer for Agile Manipulation With a Single RGB Camera</itunes:summary>
      <itunes:subtitle>Pixel2Catch: Multi-Agent Sim-to-Real Transfer for Agile Manipulation With a Single RGB Camera</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/83783b0c/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Calibrate the Robot, Randomize the World: PACE for Legged Sim-to-Real</title>
      <itunes:title>Calibrate the Robot, Randomize the World: PACE for Legged Sim-to-Real</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">517f7ba3-62e2-465c-b1d9-5f6ba6839b90</guid>
      <link>https://share.transistor.fm/s/a8ca814e</link>
      <description>
        <![CDATA[Legged robots must achieve both robust locomotion and energy efficiency to be practical in real-world environments. Yet controllers trained in simulation often fail to transfer reliably, and most existing approaches are limited to specific robot morphologies. This paper presents a systematic framework for sim-to-real transfer across diverse legged robot platforms.]]>
      </description>
      <content:encoded>
        <![CDATA[Legged robots must achieve both robust locomotion and energy efficiency to be practical in real-world environments. Yet controllers trained in simulation often fail to transfer reliably, and most existing approaches are limited to specific robot morphologies. This paper presents a systematic framework for sim-to-real transfer across diverse legged robot platforms.]]>
      </content:encoded>
      <pubDate>Sun, 02 Aug 2026 14:07:58 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/a8ca814e/2dabfed8.mp3" length="29816415" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1864</itunes:duration>
      <itunes:summary>Legged robots must achieve both robust locomotion and energy efficiency to be practical in real-world environments. Yet controllers trained in simulation often fail to transfer reliably, and most existing approaches are limited to specific robot morphologies. This paper presents a systematic framework for sim-to-real transfer across diverse legged robot platforms.</itunes:summary>
      <itunes:subtitle>Legged robots must achieve both robust locomotion and energy efficiency to be practical in real-world environments. Yet controllers trained in simulation often fail to transfer reliably, and most existing approaches are limited to specific robot morpholog</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/a8ca814e/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>HumanCLAW: The Model Sees the Chair but Loses the Body</title>
      <itunes:title>HumanCLAW: The Model Sees the Chair but Loses the Body</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">6a7c0b4a-f6ef-4918-98b8-10c66ed0e3f4</guid>
      <link>https://share.transistor.fm/s/ff1340e6</link>
      <description>
        <![CDATA[Introduces a benchmark with 1,218 embodied episodes that decouples high-level VLM action decisions from low-level motor control; no current VLM solves the tasks, with the best achieving only 16.8% success, highlighting lack of embodied self-awareness.]]>
      </description>
      <content:encoded>
        <![CDATA[Introduces a benchmark with 1,218 embodied episodes that decouples high-level VLM action decisions from low-level motor control; no current VLM solves the tasks, with the best achieving only 16.8% success, highlighting lack of embodied self-awareness.]]>
      </content:encoded>
      <pubDate>Sun, 02 Aug 2026 03:09:15 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/ff1340e6/7002f914.mp3" length="33228634" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2077</itunes:duration>
      <itunes:summary>Introduces a benchmark with 1,218 embodied episodes that decouples high-level VLM action decisions from low-level motor control; no current VLM solves the tasks, with the best achieving only 16.8% success, highlighting lack of embodied self-awareness.</itunes:summary>
      <itunes:subtitle>Introduces a benchmark with 1,218 embodied episodes that decouples high-level VLM action decisions from low-level motor control; no current VLM solves the tasks, with the best achieving only 16.8% success, highlighting lack of embodied self-awareness.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/ff1340e6/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Get Up, Adapt, Survive: Recovery Skills for Small Quadrupeds</title>
      <itunes:title>Get Up, Adapt, Survive: Recovery Skills for Small Quadrupeds</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">db7d51ff-727e-4e48-a28f-6e873b212bc7</guid>
      <link>https://share.transistor.fm/s/f81e8595</link>
      <description>
        <![CDATA[Real-to-Sim-to-Real: Learning Agile and Robust Recovery Skills With Terrain Imagination via Adversarial Imitation]]>
      </description>
      <content:encoded>
        <![CDATA[Real-to-Sim-to-Real: Learning Agile and Robust Recovery Skills With Terrain Imagination via Adversarial Imitation]]>
      </content:encoded>
      <pubDate>Sat, 01 Aug 2026 14:08:06 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/f81e8595/d73cf83c.mp3" length="26919958" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1683</itunes:duration>
      <itunes:summary>Real-to-Sim-to-Real: Learning Agile and Robust Recovery Skills With Terrain Imagination via Adversarial Imitation</itunes:summary>
      <itunes:subtitle>Real-to-Sim-to-Real: Learning Agile and Robust Recovery Skills With Terrain Imagination via Adversarial Imitation</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/f81e8595/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>A Patch Only Needs a Glimpse: Breaking VLA Robots Under Partial Observability</title>
      <itunes:title>A Patch Only Needs a Glimpse: Breaking VLA Robots Under Partial Observability</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">1e9ffe13-32fa-407d-92a1-9209b0c0cb6e</guid>
      <link>https://share.transistor.fm/s/6a987141</link>
      <description>
        <![CDATA[Partially Observable Adversarial Patch Attacks on Vision-Language-Action Models in Robotics]]>
      </description>
      <content:encoded>
        <![CDATA[Partially Observable Adversarial Patch Attacks on Vision-Language-Action Models in Robotics]]>
      </content:encoded>
      <pubDate>Sat, 01 Aug 2026 14:07:12 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/6a987141/d1f626f3.mp3" length="30798619" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1925</itunes:duration>
      <itunes:summary>Partially Observable Adversarial Patch Attacks on Vision-Language-Action Models in Robotics</itunes:summary>
      <itunes:subtitle>Partially Observable Adversarial Patch Attacks on Vision-Language-Action Models in Robotics</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/6a987141/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>A Thousand Tasks, One Day, and the Case Against Learning Everything End to End</title>
      <itunes:title>A Thousand Tasks, One Day, and the Case Against Learning Everything End to End</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">fba4fbad-3413-4e9b-8733-9648e9435616</guid>
      <link>https://share.transistor.fm/s/9ca4f4ee</link>
      <description>
        <![CDATA[Presents a scalable approach to robot task learning enabling acquisition of a large number of tasks within a single day. Details were discussed in a podcast episode released July 30, 2026.]]>
      </description>
      <content:encoded>
        <![CDATA[Presents a scalable approach to robot task learning enabling acquisition of a large number of tasks within a single day. Details were discussed in a podcast episode released July 30, 2026.]]>
      </content:encoded>
      <pubDate>Sat, 01 Aug 2026 05:21:06 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/9ca4f4ee/4442a03d.mp3" length="30934038" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1934</itunes:duration>
      <itunes:summary>Presents a scalable approach to robot task learning enabling acquisition of a large number of tasks within a single day. Details were discussed in a podcast episode released July 30, 2026.</itunes:summary>
      <itunes:subtitle>Presents a scalable approach to robot task learning enabling acquisition of a large number of tasks within a single day. Details were discussed in a podcast episode released July 30, 2026.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/9ca4f4ee/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>HUG: Turning Everyday Human Grasps Into a Robot-Ready Prior</title>
      <itunes:title>HUG: Turning Everyday Human Grasps Into a Robot-Ready Prior</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">cd8a5550-5c2a-4a2d-bf3a-5ab69fdf52b8</guid>
      <link>https://share.transistor.fm/s/df643c52</link>
      <description>
        <![CDATA[Collects 1M frames (27.8 hours) of egocentric human grasping video and trains a flow-matching model to predict and retarget hand poses to robots. Achieves large zero-shot gains on diverse everyday grasping tasks without requiring robot-specific training data.]]>
      </description>
      <content:encoded>
        <![CDATA[Collects 1M frames (27.8 hours) of egocentric human grasping video and trains a flow-matching model to predict and retarget hand poses to robots. Achieves large zero-shot gains on diverse everyday grasping tasks without requiring robot-specific training data.]]>
      </content:encoded>
      <pubDate>Sat, 01 Aug 2026 05:14:41 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/df643c52/d346af6f.mp3" length="31933378" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1996</itunes:duration>
      <itunes:summary>Collects 1M frames (27.8 hours) of egocentric human grasping video and trains a flow-matching model to predict and retarget hand poses to robots. Achieves large zero-shot gains on diverse everyday grasping tasks without requiring robot-specific training data.</itunes:summary>
      <itunes:subtitle>Collects 1M frames (27.8 hours) of egocentric human grasping video and trains a flow-matching model to predict and retarget hand poses to robots. Achieves large zero-shot gains on diverse everyday grasping tasks without requiring robot-specific training d</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/df643c52/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>SkillPlug: Teaching Robot Policies to Reuse the Moves Between Tasks</title>
      <itunes:title>SkillPlug: Teaching Robot Policies to Reuse the Moves Between Tasks</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">74be8e11-c1b0-416c-b5ab-1f0672a15457</guid>
      <link>https://share.transistor.fm/s/2abd5202</link>
      <description>
        <![CDATA[SkillPlug: Unsupervised Skill Mining for Few-Shot Adaptation in Robotic Manipulation]]>
      </description>
      <content:encoded>
        <![CDATA[SkillPlug: Unsupervised Skill Mining for Few-Shot Adaptation in Robotic Manipulation]]>
      </content:encoded>
      <pubDate>Sat, 01 Aug 2026 03:12:26 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/2abd5202/d4b14b3e.mp3" length="35628555" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2227</itunes:duration>
      <itunes:summary>SkillPlug: Unsupervised Skill Mining for Few-Shot Adaptation in Robotic Manipulation</itunes:summary>
      <itunes:subtitle>SkillPlug: Unsupervised Skill Mining for Few-Shot Adaptation in Robotic Manipulation</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/2abd5202/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Expand Before You Compress: TEAM-VLA's Training-Free Shortcut to Faster Robot Policies</title>
      <itunes:title>Expand Before You Compress: TEAM-VLA's Training-Free Shortcut to Faster Robot Policies</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">bdbb0ffe-d425-4b27-8147-e72469d0cc3f</guid>
      <link>https://share.transistor.fm/s/2a174406</link>
      <description>
        <![CDATA[Token Expand-Merge: Training-Free Token Compression for Vision-Language-Action Models]]>
      </description>
      <content:encoded>
        <![CDATA[Token Expand-Merge: Training-Free Token Compression for Vision-Language-Action Models]]>
      </content:encoded>
      <pubDate>Sat, 01 Aug 2026 03:09:27 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/2a174406/dc5c3672.mp3" length="27528506" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1721</itunes:duration>
      <itunes:summary>Token Expand-Merge: Training-Free Token Compression for Vision-Language-Action Models</itunes:summary>
      <itunes:subtitle>Token Expand-Merge: Training-Free Token Compression for Vision-Language-Action Models</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/2a174406/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>One Policy, Three Bodies, Zero Messages: Inside CHORUS</title>
      <itunes:title>One Policy, Three Bodies, Zero Messages: Inside CHORUS</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">41d1f51d-43d7-4390-929d-c8a6b49fd4a7</guid>
      <link>https://share.transistor.fm/s/39672f0b</link>
      <description>
        <![CDATA[Trains a single VLA policy that enables multiple heterogeneous robots to collaborate on complex manipulation tasks using only local observations and robot IDs, avoiding centralized controllers.]]>
      </description>
      <content:encoded>
        <![CDATA[Trains a single VLA policy that enables multiple heterogeneous robots to collaborate on complex manipulation tasks using only local observations and robot IDs, avoiding centralized controllers.]]>
      </content:encoded>
      <pubDate>Wed, 29 Jul 2026 14:12:07 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/39672f0b/a989c567.mp3" length="27331230" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1709</itunes:duration>
      <itunes:summary>Trains a single VLA policy that enables multiple heterogeneous robots to collaborate on complex manipulation tasks using only local observations and robot IDs, avoiding centralized controllers.</itunes:summary>
      <itunes:subtitle>Trains a single VLA policy that enables multiple heterogeneous robots to collaborate on complex manipulation tasks using only local observations and robot IDs, avoiding centralized controllers.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/39672f0b/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Anchor the Semantics, Align the Actions: A Better Way to Fine-Tune VLAs</title>
      <itunes:title>Anchor the Semantics, Align the Actions: A Better Way to Fine-Tune VLAs</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">f10810d1-043d-4dd9-b3da-fccc426c6c17</guid>
      <link>https://share.transistor.fm/s/0f796771</link>
      <description>
        <![CDATA[Proposes a VLA fine-tuning technique that preserves pretrained VLM representations while aligning language to actions, raising real-robot generalization success from 28% to 54%.]]>
      </description>
      <content:encoded>
        <![CDATA[Proposes a VLA fine-tuning technique that preserves pretrained VLM representations while aligning language to actions, raising real-robot generalization success from 28% to 54%.]]>
      </content:encoded>
      <pubDate>Wed, 29 Jul 2026 14:10:43 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/0f796771/44ffe3ac.mp3" length="33634890" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2103</itunes:duration>
      <itunes:summary>Proposes a VLA fine-tuning technique that preserves pretrained VLM representations while aligning language to actions, raising real-robot generalization success from 28% to 54%.</itunes:summary>
      <itunes:subtitle>Proposes a VLA fine-tuning technique that preserves pretrained VLM representations while aligning language to actions, raising real-robot generalization success from 28% to 54%.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/0f796771/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Learning the Rule, Not the Gait: Stick-Insect Behavior as a Portable Robot Objective</title>
      <itunes:title>Learning the Rule, Not the Gait: Stick-Insect Behavior as a Portable Robot Objective</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">1bd946bd-bd12-46ca-a04a-7e91027a02bb</guid>
      <link>https://share.transistor.fm/s/05c6e288</link>
      <description>
        <![CDATA[Insect locomotion exhibits remarkable adaptability and flexibility despite the limited scale of its nervous system. However, the underlying principles that govern leg coordination remain difficult to extract and model. This paper presents an adversarial inverse reinforcement learning approach to infer transferable locomotor principles from limited insect behavioral data, enabling robust robot locomotion.]]>
      </description>
      <content:encoded>
        <![CDATA[Insect locomotion exhibits remarkable adaptability and flexibility despite the limited scale of its nervous system. However, the underlying principles that govern leg coordination remain difficult to extract and model. This paper presents an adversarial inverse reinforcement learning approach to infer transferable locomotor principles from limited insect behavioral data, enabling robust robot locomotion.]]>
      </content:encoded>
      <pubDate>Wed, 29 Jul 2026 05:14:23 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/05c6e288/0bcb40ba.mp3" length="28829195" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1802</itunes:duration>
      <itunes:summary>Insect locomotion exhibits remarkable adaptability and flexibility despite the limited scale of its nervous system. However, the underlying principles that govern leg coordination remain difficult to extract and model. This paper presents an adversarial inverse reinforcement learning approach to infer transferable locomotor principles from limited insect behavioral data, enabling robust robot locomotion.</itunes:summary>
      <itunes:subtitle>Insect locomotion exhibits remarkable adaptability and flexibility despite the limited scale of its nervous system. However, the underlying principles that govern leg coordination remain difficult to extract and model. This paper presents an adversarial i</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/05c6e288/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Two Brains Under Pressure: How UnderwaterVLA Connects Language Models to AUV Control</title>
      <itunes:title>Two Brains Under Pressure: How UnderwaterVLA Connects Language Models to AUV Control</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">f7f82c10-93a5-4f2d-bd8d-140be217fa2b</guid>
      <link>https://share.transistor.fm/s/2154edbe</link>
      <description>
        <![CDATA[The UnderwaterVLA dual-brain vision-language-action architecture enables robust autonomous underwater navigation]]>
      </description>
      <content:encoded>
        <![CDATA[The UnderwaterVLA dual-brain vision-language-action architecture enables robust autonomous underwater navigation]]>
      </content:encoded>
      <pubDate>Wed, 29 Jul 2026 05:13:33 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/2154edbe/1605a9bf.mp3" length="28789071" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1800</itunes:duration>
      <itunes:summary>The UnderwaterVLA dual-brain vision-language-action architecture enables robust autonomous underwater navigation</itunes:summary>
      <itunes:subtitle>The UnderwaterVLA dual-brain vision-language-action architecture enables robust autonomous underwater navigation</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/2154edbe/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>From Demo to Duty: RL-100's Real-World Reinforcement Learning Flywheel</title>
      <itunes:title>From Demo to Duty: RL-100's Real-World Reinforcement Learning Flywheel</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">6f1e9c47-6dea-4383-9a73-e6d813c24745</guid>
      <link>https://share.transistor.fm/s/baf75415</link>
      <description>
        <![CDATA[Real-world robotic manipulation in homes and factories demands reliability, efficiency, and robustness that approach or surpass skilled human operators. We present a real-world reinforcement learning (RL) framework, RL-100, for achieving performant robotic manipulation through real-world RL training.]]>
      </description>
      <content:encoded>
        <![CDATA[Real-world robotic manipulation in homes and factories demands reliability, efficiency, and robustness that approach or surpass skilled human operators. We present a real-world reinforcement learning (RL) framework, RL-100, for achieving performant robotic manipulation through real-world RL training.]]>
      </content:encoded>
      <pubDate>Tue, 28 Jul 2026 18:31:25 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/baf75415/2d43fa7a.mp3" length="31793362" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1988</itunes:duration>
      <itunes:summary>Real-world robotic manipulation in homes and factories demands reliability, efficiency, and robustness that approach or surpass skilled human operators. We present a real-world reinforcement learning (RL) framework, RL-100, for achieving performant robotic manipulation through real-world RL training.</itunes:summary>
      <itunes:subtitle>Real-world robotic manipulation in homes and factories demands reliability, efficiency, and robustness that approach or surpass skilled human operators. We present a real-world reinforcement learning (RL) framework, RL-100, for achieving performant roboti</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/baf75415/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Distilling a Swarm into 158 Million Parameters: Inside MiniUAV-VLA</title>
      <itunes:title>Distilling a Swarm into 158 Million Parameters: Inside MiniUAV-VLA</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">7f4653b0-96ff-47c7-b8a4-3692ad4fee0e</guid>
      <link>https://share.transistor.fm/s/aa73104b</link>
      <description>
        <![CDATA[Coordinating multiple unmanned aerial vehicles (UAVs) for cooperative missions requires agents that perceive their environment, reason about objectives, and generate joint actions. Vision–language–action (VLA) models unify these capabilities but lack efficiency for real-world multi-UAV deployment. This episode explores MiniUAV-VLA, a compact 158M-parameter model that uses MARL expert distillation to enable cooperative multi-UAV search and elimination missions.]]>
      </description>
      <content:encoded>
        <![CDATA[Coordinating multiple unmanned aerial vehicles (UAVs) for cooperative missions requires agents that perceive their environment, reason about objectives, and generate joint actions. Vision–language–action (VLA) models unify these capabilities but lack efficiency for real-world multi-UAV deployment. This episode explores MiniUAV-VLA, a compact 158M-parameter model that uses MARL expert distillation to enable cooperative multi-UAV search and elimination missions.]]>
      </content:encoded>
      <pubDate>Tue, 28 Jul 2026 18:30:26 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/aa73104b/bda50cf4.mp3" length="32267327" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2017</itunes:duration>
      <itunes:summary>Coordinating multiple unmanned aerial vehicles (UAVs) for cooperative missions requires agents that perceive their environment, reason about objectives, and generate joint actions. Vision–language–action (VLA) models unify these capabilities but lack efficiency for real-world multi-UAV deployment. This episode explores MiniUAV-VLA, a compact 158M-parameter model that uses MARL expert distillation to enable cooperative multi-UAV search and elimination missions.</itunes:summary>
      <itunes:subtitle>Coordinating multiple unmanned aerial vehicles (UAVs) for cooperative missions requires agents that perceive their environment, reason about objectives, and generate joint actions. Vision–language–action (VLA) models unify these capabilities but lack effi</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/aa73104b/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Imagining Locomotion: Learning a Neural World Model for Legged Robots</title>
      <itunes:title>Imagining Locomotion: Learning a Neural World Model for Legged Robots</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">c154980f-cbe5-470a-8543-29d28f14999e</guid>
      <link>https://share.transistor.fm/s/42b3adb1</link>
      <description>
        <![CDATA[A neural dynamics world model paired with model-free RL policies for quadruped and humanoid locomotion in IsaacLab, enabling long-horizon autoregressive prediction and imagined rollouts that outperform pure model-based RL in prediction accuracy, policy learning, and sim-to-real transfer.]]>
      </description>
      <content:encoded>
        <![CDATA[A neural dynamics world model paired with model-free RL policies for quadruped and humanoid locomotion in IsaacLab, enabling long-horizon autoregressive prediction and imagined rollouts that outperform pure model-based RL in prediction accuracy, policy learning, and sim-to-real transfer.]]>
      </content:encoded>
      <pubDate>Thu, 23 Jul 2026 05:17:17 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/42b3adb1/7182834d.mp3" length="31695559" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1981</itunes:duration>
      <itunes:summary>A neural dynamics world model paired with model-free RL policies for quadruped and humanoid locomotion in IsaacLab, enabling long-horizon autoregressive prediction and imagined rollouts that outperform pure model-based RL in prediction accuracy, policy learning, and sim-to-real transfer.</itunes:summary>
      <itunes:subtitle>A neural dynamics world model paired with model-free RL policies for quadruped and humanoid locomotion in IsaacLab, enabling long-horizon autoregressive prediction and imagined rollouts that outperform pure model-based RL in prediction accuracy, policy le</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/42b3adb1/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>One Model to See, Plan, and Act: Introducing EO-1 for Embodied AI</title>
      <itunes:title>One Model to See, Plan, and Act: Introducing EO-1 for Embodied AI</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">e23bbb79-499f-4d1d-a2fd-a862aee6b2f5</guid>
      <link>https://share.transistor.fm/s/5ab7b624</link>
      <description>
        <![CDATA[A 3B parameter unified decoder-only transformer that interleaves vision, text, and action tokens for perception, planning, reasoning, and control in a single model. Trained on the 1.5M-sample EO-1.5M dataset with strong results across manipulation tasks and benchmarks.]]>
      </description>
      <content:encoded>
        <![CDATA[A 3B parameter unified decoder-only transformer that interleaves vision, text, and action tokens for perception, planning, reasoning, and control in a single model. Trained on the 1.5M-sample EO-1.5M dataset with strong results across manipulation tasks and benchmarks.]]>
      </content:encoded>
      <pubDate>Thu, 23 Jul 2026 03:11:34 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/5ab7b624/9024df84.mp3" length="26221966" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1639</itunes:duration>
      <itunes:summary>A 3B parameter unified decoder-only transformer that interleaves vision, text, and action tokens for perception, planning, reasoning, and control in a single model. Trained on the 1.5M-sample EO-1.5M dataset with strong results across manipulation tasks and benchmarks.</itunes:summary>
      <itunes:subtitle>A 3B parameter unified decoder-only transformer that interleaves vision, text, and action tokens for perception, planning, reasoning, and control in a single model. Trained on the 1.5M-sample EO-1.5M dataset with strong results across manipulation tasks a</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/5ab7b624/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Video-Action Models for Robot Learning</title>
      <itunes:title>Video-Action Models for Robot Learning</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">13711d1e-9cbb-46ef-b882-fd85945b1a01</guid>
      <link>https://share.transistor.fm/s/f6ef801e</link>
      <description>
        <![CDATA[Introduces Video-Action Models (VAMs) that leverage pretrained internet-scale video models such as Cosmos-Predict2 as backbones instead of VLMs, paired with a flow-matching action decoder. Claims approximately 10x sample efficiency gains over standard vision-language-action models on real-world pick-and-place tasks.]]>
      </description>
      <content:encoded>
        <![CDATA[Introduces Video-Action Models (VAMs) that leverage pretrained internet-scale video models such as Cosmos-Predict2 as backbones instead of VLMs, paired with a flow-matching action decoder. Claims approximately 10x sample efficiency gains over standard vision-language-action models on real-world pick-and-place tasks.]]>
      </content:encoded>
      <pubDate>Thu, 23 Jul 2026 03:09:32 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/f6ef801e/22ae1b8b.mp3" length="16148732" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1010</itunes:duration>
      <itunes:summary>Introduces Video-Action Models (VAMs) that leverage pretrained internet-scale video models such as Cosmos-Predict2 as backbones instead of VLMs, paired with a flow-matching action decoder. Claims approximately 10x sample efficiency gains over standard vision-language-action models on real-world pick-and-place tasks.</itunes:summary>
      <itunes:subtitle>Introduces Video-Action Models (VAMs) that leverage pretrained internet-scale video models such as Cosmos-Predict2 as backbones instead of VLMs, paired with a flow-matching action decoder. Claims approximately 10x sample efficiency gains over standard vis</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/f6ef801e/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>FastDSAC: Unlocking the Potential of Maximum Entropy RL in High-Dimensional Humanoid Control</title>
      <itunes:title>FastDSAC: Unlocking the Potential of Maximum Entropy RL in High-Dimensional Humanoid Control</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">fd6265e9-e4e0-4720-b524-86b03cd48fe6</guid>
      <link>https://share.transistor.fm/s/d822c7cb</link>
      <description>
        <![CDATA[Introduces FastDSAC, a method that scales maximum entropy reinforcement learning to high-dimensional humanoid control tasks, improving sample efficiency and policy performance.]]>
      </description>
      <content:encoded>
        <![CDATA[Introduces FastDSAC, a method that scales maximum entropy reinforcement learning to high-dimensional humanoid control tasks, improving sample efficiency and policy performance.]]>
      </content:encoded>
      <pubDate>Wed, 22 Jul 2026 14:18:44 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/d822c7cb/e485ac1d.mp3" length="24303533" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1519</itunes:duration>
      <itunes:summary>Introduces FastDSAC, a method that scales maximum entropy reinforcement learning to high-dimensional humanoid control tasks, improving sample efficiency and policy performance.</itunes:summary>
      <itunes:subtitle>Introduces FastDSAC, a method that scales maximum entropy reinforcement learning to high-dimensional humanoid control tasks, improving sample efficiency and policy performance.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/d822c7cb/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>A Minimalist Retargeting-Guided Reinforcement Learning Recipe for Dexterous Manipulation</title>
      <itunes:title>A Minimalist Retargeting-Guided Reinforcement Learning Recipe for Dexterous Manipulation</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">18deed96-e8d5-477b-86b0-ed594aee329a</guid>
      <link>https://share.transistor.fm/s/a1828c4e</link>
      <description>
        <![CDATA[Proposes a minimalist pipeline using retargeting-guided RL for dexterous manipulation in humanoid robots. Focuses on simplifying the training recipe while maintaining strong manipulation performance.]]>
      </description>
      <content:encoded>
        <![CDATA[Proposes a minimalist pipeline using retargeting-guided RL for dexterous manipulation in humanoid robots. Focuses on simplifying the training recipe while maintaining strong manipulation performance.]]>
      </content:encoded>
      <pubDate>Wed, 22 Jul 2026 14:16:14 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/a1828c4e/69ea6d61.mp3" length="39144846" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2447</itunes:duration>
      <itunes:summary>Proposes a minimalist pipeline using retargeting-guided RL for dexterous manipulation in humanoid robots. Focuses on simplifying the training recipe while maintaining strong manipulation performance.</itunes:summary>
      <itunes:subtitle>Proposes a minimalist pipeline using retargeting-guided RL for dexterous manipulation in humanoid robots. Focuses on simplifying the training recipe while maintaining strong manipulation performance.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/a1828c4e/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>DexNDM: Closing the Reality Gap for Dexterous In-Hand Rotation</title>
      <itunes:title>DexNDM: Closing the Reality Gap for Dexterous In-Hand Rotation</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">80ed49ed-0ab7-4718-9040-a7e63068a5dd</guid>
      <link>https://share.transistor.fm/s/78ebce48</link>
      <description>
        <![CDATA[DexNDM bridges the sim-to-real gap for stable in-hand rotation of complex objects, enabling learning from biased real-world data without requiring any successful demonstrations.]]>
      </description>
      <content:encoded>
        <![CDATA[DexNDM bridges the sim-to-real gap for stable in-hand rotation of complex objects, enabling learning from biased real-world data without requiring any successful demonstrations.]]>
      </content:encoded>
      <pubDate>Wed, 22 Jul 2026 03:22:38 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/78ebce48/04d6e370.mp3" length="29456970" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1842</itunes:duration>
      <itunes:summary>DexNDM bridges the sim-to-real gap for stable in-hand rotation of complex objects, enabling learning from biased real-world data without requiring any successful demonstrations.</itunes:summary>
      <itunes:subtitle>DexNDM bridges the sim-to-real gap for stable in-hand rotation of complex objects, enabling learning from biased real-world data without requiring any successful demonstrations.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/78ebce48/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Learning Multi-Modal Trajectory Policies for Data-Efficient Robotic Manipulation</title>
      <itunes:title>Learning Multi-Modal Trajectory Policies for Data-Efficient Robotic Manipulation</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">3dbc9377-8906-40af-af5f-f9a2105e0a9d</guid>
      <link>https://share.transistor.fm/s/bae5dcc7</link>
      <description>
        <![CDATA[Proposes learning multi-modal trajectory policies to improve data efficiency in robotic manipulation tasks.]]>
      </description>
      <content:encoded>
        <![CDATA[Proposes learning multi-modal trajectory policies to improve data efficiency in robotic manipulation tasks.]]>
      </content:encoded>
      <pubDate>Wed, 22 Jul 2026 03:15:37 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/bae5dcc7/445da607.mp3" length="26945035" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1685</itunes:duration>
      <itunes:summary>Proposes learning multi-modal trajectory policies to improve data efficiency in robotic manipulation tasks.</itunes:summary>
      <itunes:subtitle>Proposes learning multi-modal trajectory policies to improve data efficiency in robotic manipulation tasks.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/bae5dcc7/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>ACE: Agentic Control for Embodied Manipulation via Zero-shot Workflow Reasoning</title>
      <itunes:title>ACE: Agentic Control for Embodied Manipulation via Zero-shot Workflow Reasoning</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">ba6b02ca-0e5a-49da-87ac-39b21d4a1fc0</guid>
      <link>https://share.transistor.fm/s/ccb448f5</link>
      <description>
        <![CDATA[A zero-shot workflow reasoning approach for agentic control of embodied manipulation, enabling robots to perform complex tasks without task-specific training.]]>
      </description>
      <content:encoded>
        <![CDATA[A zero-shot workflow reasoning approach for agentic control of embodied manipulation, enabling robots to perform complex tasks without task-specific training.]]>
      </content:encoded>
      <pubDate>Tue, 21 Jul 2026 05:17:04 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/ccb448f5/0d978aac.mp3" length="44551984" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2785</itunes:duration>
      <itunes:summary>A zero-shot workflow reasoning approach for agentic control of embodied manipulation, enabling robots to perform complex tasks without task-specific training.</itunes:summary>
      <itunes:subtitle>A zero-shot workflow reasoning approach for agentic control of embodied manipulation, enabling robots to perform complex tasks without task-specific training.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/ccb448f5/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Let the Body Follow: Coupled Egocentric Control for Whole-Body Robot Teleoperation</title>
      <itunes:title>Let the Body Follow: Coupled Egocentric Control for Whole-Body Robot Teleoperation</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">23c3bfa6-e9e0-46a8-9927-701021a602f6</guid>
      <link>https://share.transistor.fm/s/8ae43473</link>
      <description>
        <![CDATA[A coupled egocentric control approach for whole-body robot teleoperation that coordinates body and limb movements from an egocentric perspective, submitted to Humanoids 2026.]]>
      </description>
      <content:encoded>
        <![CDATA[A coupled egocentric control approach for whole-body robot teleoperation that coordinates body and limb movements from an egocentric perspective, submitted to Humanoids 2026.]]>
      </content:encoded>
      <pubDate>Tue, 21 Jul 2026 05:12:31 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/8ae43473/62a3a328.mp3" length="17648369" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1103</itunes:duration>
      <itunes:summary>A coupled egocentric control approach for whole-body robot teleoperation that coordinates body and limb movements from an egocentric perspective, submitted to Humanoids 2026.</itunes:summary>
      <itunes:subtitle>A coupled egocentric control approach for whole-body robot teleoperation that coordinates body and limb movements from an egocentric perspective, submitted to Humanoids 2026.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/8ae43473/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>GigaWorld-Policy-0.5: An Efficient Mixture-of-Transformers Policy for Real-Time Robot Control</title>
      <itunes:title>GigaWorld-Policy-0.5: An Efficient Mixture-of-Transformers Policy for Real-Time Robot Control</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">d54acd97-b5f8-4085-a58a-c308e8fd2178</guid>
      <link>https://share.transistor.fm/s/09e15d49</link>
      <description>
        <![CDATA[A Mixture-of-Transformers robot policy that achieves 85 ms inference on RTX 4090 by separating visual dynamics from action generation. Demonstrates efficient real-time robot policy execution through architectural decomposition.]]>
      </description>
      <content:encoded>
        <![CDATA[A Mixture-of-Transformers robot policy that achieves 85 ms inference on RTX 4090 by separating visual dynamics from action generation. Demonstrates efficient real-time robot policy execution through architectural decomposition.]]>
      </content:encoded>
      <pubDate>Mon, 20 Jul 2026 14:09:20 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/09e15d49/309da0a5.mp3" length="34244692" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2141</itunes:duration>
      <itunes:summary>A Mixture-of-Transformers robot policy that achieves 85 ms inference on RTX 4090 by separating visual dynamics from action generation. Demonstrates efficient real-time robot policy execution through architectural decomposition.</itunes:summary>
      <itunes:subtitle>A Mixture-of-Transformers robot policy that achieves 85 ms inference on RTX 4090 by separating visual dynamics from action generation. Demonstrates efficient real-time robot policy execution through architectural decomposition.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/09e15d49/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Xiaomi-Robotics-1: A Scalable Vision-Language-Action Foundation Model for Mobile Manipulation</title>
      <itunes:title>Xiaomi-Robotics-1: A Scalable Vision-Language-Action Foundation Model for Mobile Manipulation</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">77b207fb-8642-49dd-9244-b2adbc73b81c</guid>
      <link>https://share.transistor.fm/s/2a87c217</link>
      <description>
        <![CDATA[A scalable vision-language-action (VLA) foundation model pretrained on over 100k hours of real-world manipulation trajectories, enabling out-of-the-box mobile manipulation capabilities.]]>
      </description>
      <content:encoded>
        <![CDATA[A scalable vision-language-action (VLA) foundation model pretrained on over 100k hours of real-world manipulation trajectories, enabling out-of-the-box mobile manipulation capabilities.]]>
      </content:encoded>
      <pubDate>Mon, 20 Jul 2026 05:06:36 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/2a87c217/f5673c44.mp3" length="19288441" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1206</itunes:duration>
      <itunes:summary>A scalable vision-language-action (VLA) foundation model pretrained on over 100k hours of real-world manipulation trajectories, enabling out-of-the-box mobile manipulation capabilities.</itunes:summary>
      <itunes:subtitle>A scalable vision-language-action (VLA) foundation model pretrained on over 100k hours of real-world manipulation trajectories, enabling out-of-the-box mobile manipulation capabilities.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/2a87c217/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Local Policies Enable Zero-shot Long-Horizon Manipulation</title>
      <itunes:title>Local Policies Enable Zero-shot Long-Horizon Manipulation</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">6a670763-5e31-41ac-a8a7-f08adec2f7d8</guid>
      <link>https://share.transistor.fm/s/6eb3eece</link>
      <description>
        <![CDATA[Proposes training local policies in simulation that transfer zero-shot to real-world long-horizon robotic manipulation tasks. Addresses the challenge of generalizing learned behaviors across extended task sequences.]]>
      </description>
      <content:encoded>
        <![CDATA[Proposes training local policies in simulation that transfer zero-shot to real-world long-horizon robotic manipulation tasks. Addresses the challenge of generalizing learned behaviors across extended task sequences.]]>
      </content:encoded>
      <pubDate>Sun, 19 Jul 2026 14:16:30 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/6eb3eece/80497339.mp3" length="31845189" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1991</itunes:duration>
      <itunes:summary>Proposes training local policies in simulation that transfer zero-shot to real-world long-horizon robotic manipulation tasks. Addresses the challenge of generalizing learned behaviors across extended task sequences.</itunes:summary>
      <itunes:subtitle>Proposes training local policies in simulation that transfer zero-shot to real-world long-horizon robotic manipulation tasks. Addresses the challenge of generalizing learned behaviors across extended task sequences.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/6eb3eece/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Sim-to-Real Reinforcement Learning for Vision-Based Dexterous Manipulation on Humanoids</title>
      <itunes:title>Sim-to-Real Reinforcement Learning for Vision-Based Dexterous Manipulation on Humanoids</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">04fea893-2351-46f5-8000-fdef8cc03385</guid>
      <link>https://share.transistor.fm/s/692ed6fb</link>
      <description>
        <![CDATA[Trains perception-driven dexterous manipulation policies in simulation for zero-shot real-world transfer on humanoid robots. Focuses on bridging the sim-to-real gap for vision-based dexterous tasks.]]>
      </description>
      <content:encoded>
        <![CDATA[Trains perception-driven dexterous manipulation policies in simulation for zero-shot real-world transfer on humanoid robots. Focuses on bridging the sim-to-real gap for vision-based dexterous tasks.]]>
      </content:encoded>
      <pubDate>Sun, 19 Jul 2026 14:10:54 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/692ed6fb/9e5df87e.mp3" length="34579060" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2162</itunes:duration>
      <itunes:summary>Trains perception-driven dexterous manipulation policies in simulation for zero-shot real-world transfer on humanoid robots. Focuses on bridging the sim-to-real gap for vision-based dexterous tasks.</itunes:summary>
      <itunes:subtitle>Trains perception-driven dexterous manipulation policies in simulation for zero-shot real-world transfer on humanoid robots. Focuses on bridging the sim-to-real gap for vision-based dexterous tasks.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/692ed6fb/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>UniTracker: A Universal Motion Tracking Framework for Humanoid Robots</title>
      <itunes:title>UniTracker: A Universal Motion Tracking Framework for Humanoid Robots</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">32dd49a7-eb22-4622-b884-40d1e736ffd7</guid>
      <link>https://share.transistor.fm/s/efa535cb</link>
      <description>
        <![CDATA[A framework enabling humanoid robots to execute diverse motions within physical limits, with video demonstrations and open-source code. Targets generalizable motion tracking across varied humanoid morphologies.]]>
      </description>
      <content:encoded>
        <![CDATA[A framework enabling humanoid robots to execute diverse motions within physical limits, with video demonstrations and open-source code. Targets generalizable motion tracking across varied humanoid morphologies.]]>
      </content:encoded>
      <pubDate>Sun, 19 Jul 2026 05:18:06 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/efa535cb/8aed0823.mp3" length="25294514" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1581</itunes:duration>
      <itunes:summary>A framework enabling humanoid robots to execute diverse motions within physical limits, with video demonstrations and open-source code. Targets generalizable motion tracking across varied humanoid morphologies.</itunes:summary>
      <itunes:subtitle>A framework enabling humanoid robots to execute diverse motions within physical limits, with video demonstrations and open-source code. Targets generalizable motion tracking across varied humanoid morphologies.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/efa535cb/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>ViTacFormer: Dexterous Manipulation via Active Vision and High-Resolution Touch Sensing</title>
      <itunes:title>ViTacFormer: Dexterous Manipulation via Active Vision and High-Resolution Touch Sensing</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">91082bfa-97a5-4490-b791-1b1698496f0e</guid>
      <link>https://share.transistor.fm/s/d834c1bd</link>
      <description>
        <![CDATA[A pipeline combining active vision and high-resolution tactile sensing for dexterous manipulation on high-DoF robot hands, achieving approximately 2.5 minutes of continuous autonomous control on real-world tasks such as food preparation.]]>
      </description>
      <content:encoded>
        <![CDATA[A pipeline combining active vision and high-resolution tactile sensing for dexterous manipulation on high-DoF robot hands, achieving approximately 2.5 minutes of continuous autonomous control on real-world tasks such as food preparation.]]>
      </content:encoded>
      <pubDate>Sun, 19 Jul 2026 05:12:38 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/d834c1bd/824289ae.mp3" length="22924686" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1433</itunes:duration>
      <itunes:summary>A pipeline combining active vision and high-resolution tactile sensing for dexterous manipulation on high-DoF robot hands, achieving approximately 2.5 minutes of continuous autonomous control on real-world tasks such as food preparation.</itunes:summary>
      <itunes:subtitle>A pipeline combining active vision and high-resolution tactile sensing for dexterous manipulation on high-DoF robot hands, achieving approximately 2.5 minutes of continuous autonomous control on real-world tasks such as food preparation.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/d834c1bd/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Masked IRL: LLM-Guided Reward Disambiguation from Demonstrations and Language</title>
      <itunes:title>Masked IRL: LLM-Guided Reward Disambiguation from Demonstrations and Language</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">7a0a5434-56a9-4031-b7fb-ee4531964008</guid>
      <link>https://share.transistor.fm/s/5da2cda4</link>
      <description>
        <![CDATA[Uses language and demonstrations to learn which aspects of behavior matter for reward in imitation learning, enabling 5× faster learning by avoiding copying irrelevant motion details. Accepted at ICRA 2026.]]>
      </description>
      <content:encoded>
        <![CDATA[Uses language and demonstrations to learn which aspects of behavior matter for reward in imitation learning, enabling 5× faster learning by avoiding copying irrelevant motion details. Accepted at ICRA 2026.]]>
      </content:encoded>
      <pubDate>Sun, 19 Jul 2026 03:10:22 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/5da2cda4/77a9192d.mp3" length="26446410" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1653</itunes:duration>
      <itunes:summary>Uses language and demonstrations to learn which aspects of behavior matter for reward in imitation learning, enabling 5× faster learning by avoiding copying irrelevant motion details. Accepted at ICRA 2026.</itunes:summary>
      <itunes:subtitle>Uses language and demonstrations to learn which aspects of behavior matter for reward in imitation learning, enabling 5× faster learning by avoiding copying irrelevant motion details. Accepted at ICRA 2026.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/5da2cda4/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>FlowDAgger: Human-in-the-Loop Adaptation of Generative Robot Policies in Latent Space</title>
      <itunes:title>FlowDAgger: Human-in-the-Loop Adaptation of Generative Robot Policies in Latent Space</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">a92bc053-75f7-4248-af6b-3f06dcc2cf18</guid>
      <link>https://share.transistor.fm/s/753b1299</link>
      <description>
        <![CDATA[Steers frozen robot foundation models including VLAs, diffusion policies, and world models using human corrections via action inversion in latent space without fine-tuning the base policy. Outperforms supervised fine-tuning and latent RL with only 5–20 human interventions.]]>
      </description>
      <content:encoded>
        <![CDATA[Steers frozen robot foundation models including VLAs, diffusion policies, and world models using human corrections via action inversion in latent space without fine-tuning the base policy. Outperforms supervised fine-tuning and latent RL with only 5–20 human interventions.]]>
      </content:encoded>
      <pubDate>Sun, 19 Jul 2026 03:09:36 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/753b1299/354e74aa.mp3" length="16147896" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1010</itunes:duration>
      <itunes:summary>Steers frozen robot foundation models including VLAs, diffusion policies, and world models using human corrections via action inversion in latent space without fine-tuning the base policy. Outperforms supervised fine-tuning and latent RL with only 5–20 human interventions.</itunes:summary>
      <itunes:subtitle>Steers frozen robot foundation models including VLAs, diffusion policies, and world models using human corrections via action inversion in latent space without fine-tuning the base policy. Outperforms supervised fine-tuning and latent RL with only 5–20 hu</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/753b1299/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>RoboTTT: End-to-End Robotic Assembly with Test-Time Training</title>
      <itunes:title>RoboTTT: End-to-End Robotic Assembly with Test-Time Training</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">926e8909-0b1c-42d5-93f3-2b9ea637d8dd</guid>
      <link>https://share.transistor.fm/s/da6d7ac3</link>
      <description>
        <![CDATA[Presents a single end-to-end policy that performs precise, unscripted assembly of complex objects, measuring every grasp and alignment without speed-ups or human intervention.]]>
      </description>
      <content:encoded>
        <![CDATA[Presents a single end-to-end policy that performs precise, unscripted assembly of complex objects, measuring every grasp and alignment without speed-ups or human intervention.]]>
      </content:encoded>
      <pubDate>Sat, 18 Jul 2026 05:13:30 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/da6d7ac3/5d6f8986.mp3" length="22701078" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1419</itunes:duration>
      <itunes:summary>Presents a single end-to-end policy that performs precise, unscripted assembly of complex objects, measuring every grasp and alignment without speed-ups or human intervention.</itunes:summary>
      <itunes:subtitle>Presents a single end-to-end policy that performs precise, unscripted assembly of complex objects, measuring every grasp and alignment without speed-ups or human intervention.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/da6d7ac3/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Xiaomi-Robotics-U0: A 38B World-Foundation Model for Unified Embodied Perception and Synthesis</title>
      <itunes:title>Xiaomi-Robotics-U0: A 38B World-Foundation Model for Unified Embodied Perception and Synthesis</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">fb2f8cf1-9ba7-4179-810a-6681f1942623</guid>
      <link>https://share.transistor.fm/s/602357e9</link>
      <description>
        <![CDATA[38B autoregressive world foundation model unifying text-to-image, multi-view scene generation, embodied transfer, and video generation; ranks #1 on World Arena.]]>
      </description>
      <content:encoded>
        <![CDATA[38B autoregressive world foundation model unifying text-to-image, multi-view scene generation, embodied transfer, and video generation; ranks #1 on World Arena.]]>
      </content:encoded>
      <pubDate>Sat, 18 Jul 2026 05:07:10 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/602357e9/8a346a36.mp3" length="26778687" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1674</itunes:duration>
      <itunes:summary>38B autoregressive world foundation model unifying text-to-image, multi-view scene generation, embodied transfer, and video generation; ranks #1 on World Arena.</itunes:summary>
      <itunes:subtitle>38B autoregressive world foundation model unifying text-to-image, multi-view scene generation, embodied transfer, and video generation; ranks #1 on World Arena.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/602357e9/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>SpatialPoint: Spatial-Aware Point Prediction for Embodied Localization</title>
      <itunes:title>SpatialPoint: Spatial-Aware Point Prediction for Embodied Localization</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">14d9c9f3-7c3e-4e36-bbc9-d1748c92c0b3</guid>
      <link>https://share.transistor.fm/s/601ed68a</link>
      <description>
        <![CDATA[A spatial-aware point prediction method for embodied localization tasks. Addresses grounding and spatial reasoning in embodied AI settings.]]>
      </description>
      <content:encoded>
        <![CDATA[A spatial-aware point prediction method for embodied localization tasks. Addresses grounding and spatial reasoning in embodied AI settings.]]>
      </content:encoded>
      <pubDate>Sat, 18 Jul 2026 03:19:51 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/601ed68a/6eba7d7a.mp3" length="21882714" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1368</itunes:duration>
      <itunes:summary>A spatial-aware point prediction method for embodied localization tasks. Addresses grounding and spatial reasoning in embodied AI settings.</itunes:summary>
      <itunes:subtitle>A spatial-aware point prediction method for embodied localization tasks. Addresses grounding and spatial reasoning in embodied AI settings.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/601ed68a/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Multi-AUV Scene-Adaptive Embodied Intelligence for Multi-Target Tracking</title>
      <itunes:title>Multi-AUV Scene-Adaptive Embodied Intelligence for Multi-Target Tracking</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">5ab63e6d-93d0-4572-9055-518acccafdfb</guid>
      <link>https://share.transistor.fm/s/297f71cd</link>
      <description>
        <![CDATA[An embodied AI framework for multi-AUV (Autonomous Underwater Vehicle) multi-target tracking in ad-hoc underwater networks. Incorporates scene-adaptive policies for dynamic underwater environments.]]>
      </description>
      <content:encoded>
        <![CDATA[An embodied AI framework for multi-AUV (Autonomous Underwater Vehicle) multi-target tracking in ad-hoc underwater networks. Incorporates scene-adaptive policies for dynamic underwater environments.]]>
      </content:encoded>
      <pubDate>Sat, 18 Jul 2026 03:19:44 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/297f71cd/a71b169c.mp3" length="43611576" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2726</itunes:duration>
      <itunes:summary>An embodied AI framework for multi-AUV (Autonomous Underwater Vehicle) multi-target tracking in ad-hoc underwater networks. Incorporates scene-adaptive policies for dynamic underwater environments.</itunes:summary>
      <itunes:subtitle>An embodied AI framework for multi-AUV (Autonomous Underwater Vehicle) multi-target tracking in ad-hoc underwater networks. Incorporates scene-adaptive policies for dynamic underwater environments.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/297f71cd/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Speedup Patch: Learning a Plug-and-Play Policy to Accelerate Embodied Manipulation</title>
      <itunes:title>Speedup Patch: Learning a Plug-and-Play Policy to Accelerate Embodied Manipulation</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">d4aae317-2691-4b8f-b4ab-eae143aecbc4</guid>
      <link>https://share.transistor.fm/s/2502a647</link>
      <description>
        <![CDATA[Proposes a plug-and-play policy module designed to accelerate existing embodied manipulation policies without retraining them from scratch.]]>
      </description>
      <content:encoded>
        <![CDATA[Proposes a plug-and-play policy module designed to accelerate existing embodied manipulation policies without retraining them from scratch.]]>
      </content:encoded>
      <pubDate>Fri, 17 Jul 2026 14:09:14 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/2502a647/9a2e9d71.mp3" length="40818772" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2552</itunes:duration>
      <itunes:summary>Proposes a plug-and-play policy module designed to accelerate existing embodied manipulation policies without retraining them from scratch.</itunes:summary>
      <itunes:subtitle>Proposes a plug-and-play policy module designed to accelerate existing embodied manipulation policies without retraining them from scratch.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/2502a647/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>R3DP: Real-Time 3D-Aware Policy for Embodied Manipulation</title>
      <itunes:title>R3DP: Real-Time 3D-Aware Policy for Embodied Manipulation</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">6c2ad7de-5b5e-451c-8b16-910edc6a2331</guid>
      <link>https://share.transistor.fm/s/ccbe4f18</link>
      <description>
        <![CDATA[Introduces a real-time 3D-aware policy framework for embodied manipulation tasks, enabling spatially grounded action generation.]]>
      </description>
      <content:encoded>
        <![CDATA[Introduces a real-time 3D-aware policy framework for embodied manipulation tasks, enabling spatially grounded action generation.]]>
      </content:encoded>
      <pubDate>Fri, 17 Jul 2026 14:08:01 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/ccbe4f18/6a50b11d.mp3" length="25808186" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1613</itunes:duration>
      <itunes:summary>Introduces a real-time 3D-aware policy framework for embodied manipulation tasks, enabling spatially grounded action generation.</itunes:summary>
      <itunes:subtitle>Introduces a real-time 3D-aware policy framework for embodied manipulation tasks, enabling spatially grounded action generation.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/ccbe4f18/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>VistaVLA: Geometry- and Semantic-Aware 3D Gaussian-Grounded VLA for Robotic Manipulation</title>
      <itunes:title>VistaVLA: Geometry- and Semantic-Aware 3D Gaussian-Grounded VLA for Robotic Manipulation</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">61137c95-519f-474f-b6e7-dddbd1175c69</guid>
      <link>https://share.transistor.fm/s/de9a1c21</link>
      <description>
        <![CDATA[Presents a vision-language-action model grounded in 3D Gaussian representations that integrates geometric and semantic awareness for improved robotic manipulation.]]>
      </description>
      <content:encoded>
        <![CDATA[Presents a vision-language-action model grounded in 3D Gaussian representations that integrates geometric and semantic awareness for improved robotic manipulation.]]>
      </content:encoded>
      <pubDate>Fri, 17 Jul 2026 05:11:30 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/de9a1c21/a3d92916.mp3" length="41279363" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2580</itunes:duration>
      <itunes:summary>Presents a vision-language-action model grounded in 3D Gaussian representations that integrates geometric and semantic awareness for improved robotic manipulation.</itunes:summary>
      <itunes:subtitle>Presents a vision-language-action model grounded in 3D Gaussian representations that integrates geometric and semantic awareness for improved robotic manipulation.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/de9a1c21/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>SKooP: Symmetric Koopman Predictions for Legged Robot Reinforcement Learning</title>
      <itunes:title>SKooP: Symmetric Koopman Predictions for Legged Robot Reinforcement Learning</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">da2af69d-874c-42de-91a2-03ba55a5f7b9</guid>
      <link>https://share.transistor.fm/s/c3e534a3</link>
      <description>
        <![CDATA[Leverages symmetric Koopman operator predictions within a reinforcement learning framework to achieve faster training and better generalization for legged robot locomotion.]]>
      </description>
      <content:encoded>
        <![CDATA[Leverages symmetric Koopman operator predictions within a reinforcement learning framework to achieve faster training and better generalization for legged robot locomotion.]]>
      </content:encoded>
      <pubDate>Fri, 17 Jul 2026 05:10:38 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/c3e534a3/03b56997.mp3" length="38011758" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2376</itunes:duration>
      <itunes:summary>Leverages symmetric Koopman operator predictions within a reinforcement learning framework to achieve faster training and better generalization for legged robot locomotion.</itunes:summary>
      <itunes:subtitle>Leverages symmetric Koopman operator predictions within a reinforcement learning framework to achieve faster training and better generalization for legged robot locomotion.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/c3e534a3/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>RobotTT: A Tactile Transformer for Dexterous Manipulation</title>
      <itunes:title>RobotTT: A Tactile Transformer for Dexterous Manipulation</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">1335ee5c-25cb-4db9-bd98-2b8c98c64e02</guid>
      <link>https://share.transistor.fm/s/2ff77385</link>
      <description>
        <![CDATA[A tactile foundation model and sim-to-real pipeline designed for dexterous manipulation tasks. Combines transformer-based tactile sensing with a sim-to-real transfer approach.]]>
      </description>
      <content:encoded>
        <![CDATA[A tactile foundation model and sim-to-real pipeline designed for dexterous manipulation tasks. Combines transformer-based tactile sensing with a sim-to-real transfer approach.]]>
      </content:encoded>
      <pubDate>Fri, 17 Jul 2026 03:22:51 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/2ff77385/36fdaacb.mp3" length="20435321" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1278</itunes:duration>
      <itunes:summary>A tactile foundation model and sim-to-real pipeline designed for dexterous manipulation tasks. Combines transformer-based tactile sensing with a sim-to-real transfer approach.</itunes:summary>
      <itunes:subtitle>A tactile foundation model and sim-to-real pipeline designed for dexterous manipulation tasks. Combines transformer-based tactile sensing with a sim-to-real transfer approach.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/2ff77385/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>TAC-LOCO: Unified Whole-Body Control for Contact-Aware Locomotion and Manipulation</title>
      <itunes:title>TAC-LOCO: Unified Whole-Body Control for Contact-Aware Locomotion and Manipulation</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">ca76db17-f6fe-4bbe-ae7b-c09792a7ee09</guid>
      <link>https://share.transistor.fm/s/c7f3d4a9</link>
      <description>
        <![CDATA[A tactile-informed whole-body loco-manipulation controller for quadrupedal robots that unifies locomotion and manipulation using tactile sensing. Enables compliant and contact-aware whole-body control.]]>
      </description>
      <content:encoded>
        <![CDATA[A tactile-informed whole-body loco-manipulation controller for quadrupedal robots that unifies locomotion and manipulation using tactile sensing. Enables compliant and contact-aware whole-body control.]]>
      </content:encoded>
      <pubDate>Thu, 16 Jul 2026 14:11:13 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/c7f3d4a9/2b97a99d.mp3" length="21020046" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1314</itunes:duration>
      <itunes:summary>A tactile-informed whole-body loco-manipulation controller for quadrupedal robots that unifies locomotion and manipulation using tactile sensing. Enables compliant and contact-aware whole-body control.</itunes:summary>
      <itunes:subtitle>A tactile-informed whole-body loco-manipulation controller for quadrupedal robots that unifies locomotion and manipulation using tactile sensing. Enables compliant and contact-aware whole-body control.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/c7f3d4a9/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>TACTIC: Contact-Centric Control for Whole-Arm Manipulation</title>
      <itunes:title>TACTIC: Contact-Centric Control for Whole-Arm Manipulation</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">b97967db-6498-4ca6-b63c-db1d6460d39a</guid>
      <link>https://share.transistor.fm/s/17d72ad9</link>
      <description>
        <![CDATA[A contact-centric whole-arm manipulation framework that conditions control on both tactile and visual inputs. Targets dexterous manipulation tasks requiring rich contact feedback.]]>
      </description>
      <content:encoded>
        <![CDATA[A contact-centric whole-arm manipulation framework that conditions control on both tactile and visual inputs. Targets dexterous manipulation tasks requiring rich contact feedback.]]>
      </content:encoded>
      <pubDate>Thu, 16 Jul 2026 14:07:52 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/17d72ad9/5e26c246.mp3" length="43132594" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2696</itunes:duration>
      <itunes:summary>A contact-centric whole-arm manipulation framework that conditions control on both tactile and visual inputs. Targets dexterous manipulation tasks requiring rich contact feedback.</itunes:summary>
      <itunes:subtitle>A contact-centric whole-arm manipulation framework that conditions control on both tactile and visual inputs. Targets dexterous manipulation tasks requiring rich contact feedback.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/17d72ad9/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>SIEVE: Structure-Aware Data Selection for Imitation Learning with VLAs</title>
      <itunes:title>SIEVE: Structure-Aware Data Selection for Imitation Learning with VLAs</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">62ada1ea-6e5b-4af3-8ad8-925bd68717c3</guid>
      <link>https://share.transistor.fm/s/0663a492</link>
      <description>
        <![CDATA[SIEVE introduces structure-aware data selection specifically for imitation learning with Vision-Language-Action models, aiming to improve training efficiency and policy quality.]]>
      </description>
      <content:encoded>
        <![CDATA[SIEVE introduces structure-aware data selection specifically for imitation learning with Vision-Language-Action models, aiming to improve training efficiency and policy quality.]]>
      </content:encoded>
      <pubDate>Thu, 16 Jul 2026 05:17:43 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/0663a492/020b23a2.mp3" length="12452301" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>779</itunes:duration>
      <itunes:summary>SIEVE introduces structure-aware data selection specifically for imitation learning with Vision-Language-Action models, aiming to improve training efficiency and policy quality.</itunes:summary>
      <itunes:subtitle>SIEVE introduces structure-aware data selection specifically for imitation learning with Vision-Language-Action models, aiming to improve training efficiency and policy quality.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/0663a492/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>DexMachina: RL for Long-Horizon Bimanual Dexterous Policies</title>
      <itunes:title>DexMachina: RL for Long-Horizon Bimanual Dexterous Policies</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">61622da1-a7c5-4c66-abb1-0d7f70ffc7af</guid>
      <link>https://share.transistor.fm/s/69787b14</link>
      <description>
        <![CDATA[An RL algorithm that learns long-horizon bimanual dexterous policies for any robot hand from a single human demonstration, emphasizing generalization across hands, objects, and complex motions.]]>
      </description>
      <content:encoded>
        <![CDATA[An RL algorithm that learns long-horizon bimanual dexterous policies for any robot hand from a single human demonstration, emphasizing generalization across hands, objects, and complex motions.]]>
      </content:encoded>
      <pubDate>Thu, 16 Jul 2026 05:10:47 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/69787b14/14ddbf56.mp3" length="25512271" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1595</itunes:duration>
      <itunes:summary>An RL algorithm that learns long-horizon bimanual dexterous policies for any robot hand from a single human demonstration, emphasizing generalization across hands, objects, and complex motions.</itunes:summary>
      <itunes:subtitle>An RL algorithm that learns long-horizon bimanual dexterous policies for any robot hand from a single human demonstration, emphasizing generalization across hands, objects, and complex motions.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/69787b14/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation</title>
      <itunes:title>Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">e658c1e5-ad1a-4aa7-845d-667392475150</guid>
      <link>https://share.transistor.fm/s/471317b7</link>
      <description>
        <![CDATA[A unified platform comprising GE-Base (video diffusion model trained on 1M+ manipulation episodes), GE-Act (flow-matching action model), and GE-Sim (neural world simulator for closed-loop control). All code, models, and benchmarks are open-sourced.]]>
      </description>
      <content:encoded>
        <![CDATA[A unified platform comprising GE-Base (video diffusion model trained on 1M+ manipulation episodes), GE-Act (flow-matching action model), and GE-Sim (neural world simulator for closed-loop control). All code, models, and benchmarks are open-sourced.]]>
      </content:encoded>
      <pubDate>Thu, 16 Jul 2026 03:17:13 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/471317b7/9d60d4ab.mp3" length="25301202" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1582</itunes:duration>
      <itunes:summary>A unified platform comprising GE-Base (video diffusion model trained on 1M+ manipulation episodes), GE-Act (flow-matching action model), and GE-Sim (neural world simulator for closed-loop control). All code, models, and benchmarks are open-sourced.</itunes:summary>
      <itunes:subtitle>A unified platform comprising GE-Base (video diffusion model trained on 1M+ manipulation episodes), GE-Act (flow-matching action model), and GE-Sim (neural world simulator for closed-loop control). All code, models, and benchmarks are open-sourced.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/471317b7/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>RoboTTT: Test-Time Training for Visuomotor Policies</title>
      <itunes:title>RoboTTT: Test-Time Training for Visuomotor Policies</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">30a64c5b-3ef5-4fd8-8647-ba5a57e763ae</guid>
      <link>https://share.transistor.fm/s/71c77a56</link>
      <description>
        <![CDATA[Introduces test-time training (TTT) inside the policy to natively scale visuomotor context to 8K timesteps at constant inference cost, enabling one-shot imitation from human video demos, online self-recovery from errors, and long-horizon assembly tasks.]]>
      </description>
      <content:encoded>
        <![CDATA[Introduces test-time training (TTT) inside the policy to natively scale visuomotor context to 8K timesteps at constant inference cost, enabling one-shot imitation from human video demos, online self-recovery from errors, and long-horizon assembly tasks.]]>
      </content:encoded>
      <pubDate>Wed, 15 Jul 2026 14:09:54 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/71c77a56/592edc8e.mp3" length="36994864" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2313</itunes:duration>
      <itunes:summary>Introduces test-time training (TTT) inside the policy to natively scale visuomotor context to 8K timesteps at constant inference cost, enabling one-shot imitation from human video demos, online self-recovery from errors, and long-horizon assembly tasks.</itunes:summary>
      <itunes:subtitle>Introduces test-time training (TTT) inside the policy to natively scale visuomotor context to 8K timesteps at constant inference cost, enabling one-shot imitation from human video demos, online self-recovery from errors, and long-horizon assembly tasks.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/71c77a56/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>HY-Embodied-VLM-1.0: Efficient Physical-World Agents</title>
      <itunes:title>HY-Embodied-VLM-1.0: Efficient Physical-World Agents</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">c308ec70-b281-428c-8fa4-7bc472a1ad42</guid>
      <link>https://share.transistor.fm/s/e955f5eb</link>
      <description>
        <![CDATA[An embodied vision-language-action model with released weights, code, and paper targeting robotic manipulation and embodied AI tasks.]]>
      </description>
      <content:encoded>
        <![CDATA[An embodied vision-language-action model with released weights, code, and paper targeting robotic manipulation and embodied AI tasks.]]>
      </content:encoded>
      <pubDate>Wed, 15 Jul 2026 14:08:55 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/e955f5eb/ded10b06.mp3" length="29413920" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1839</itunes:duration>
      <itunes:summary>An embodied vision-language-action model with released weights, code, and paper targeting robotic manipulation and embodied AI tasks.</itunes:summary>
      <itunes:subtitle>An embodied vision-language-action model with released weights, code, and paper targeting robotic manipulation and embodied AI tasks.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/e955f5eb/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Learning Unified Force and Position Control for Legged Loco-Manipulation</title>
      <itunes:title>Learning Unified Force and Position Control for Legged Loco-Manipulation</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">9efa4a62-6d36-4bc0-84bb-beb5f9d48aa7</guid>
      <link>https://share.transistor.fm/s/0ab69b45</link>
      <description>
        <![CDATA[Introduces a unified RL policy that jointly handles force and position control on quadrupedal and humanoid robots without force sensors, enabling position+force tracking, compliant behaviors, and force-aware imitation learning for contact-rich tasks.]]>
      </description>
      <content:encoded>
        <![CDATA[Introduces a unified RL policy that jointly handles force and position control on quadrupedal and humanoid robots without force sensors, enabling position+force tracking, compliant behaviors, and force-aware imitation learning for contact-rich tasks.]]>
      </content:encoded>
      <pubDate>Wed, 15 Jul 2026 05:11:01 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/0ab69b45/9110667d.mp3" length="39344212" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2459</itunes:duration>
      <itunes:summary>Introduces a unified RL policy that jointly handles force and position control on quadrupedal and humanoid robots without force sensors, enabling position+force tracking, compliant behaviors, and force-aware imitation learning for contact-rich tasks.</itunes:summary>
      <itunes:subtitle>Introduces a unified RL policy that jointly handles force and position control on quadrupedal and humanoid robots without force sensors, enabling position+force tracking, compliant behaviors, and force-aware imitation learning for contact-rich tasks.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/0ab69b45/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>HapticVLA: Extending Vision-Language-Action Models to Contact-Rich Tasks Without Touch Sensors</title>
      <itunes:title>HapticVLA: Extending Vision-Language-Action Models to Contact-Rich Tasks Without Touch Sensors</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">9c7d6c13-f38a-4856-8d3a-7ac6c7663550</guid>
      <link>https://share.transistor.fm/s/6703900b</link>
      <description>
        <![CDATA[Enables contact-rich robotic manipulation using a VLA model trained with tactile sensing data but requiring no tactile input at inference time. Distills haptic knowledge into the vision-language-action policy.]]>
      </description>
      <content:encoded>
        <![CDATA[Enables contact-rich robotic manipulation using a VLA model trained with tactile sensing data but requiring no tactile input at inference time. Distills haptic knowledge into the vision-language-action policy.]]>
      </content:encoded>
      <pubDate>Tue, 14 Jul 2026 14:15:54 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/6703900b/4d8a9959.mp3" length="28073525" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1755</itunes:duration>
      <itunes:summary>Enables contact-rich robotic manipulation using a VLA model trained with tactile sensing data but requiring no tactile input at inference time. Distills haptic knowledge into the vision-language-action policy.</itunes:summary>
      <itunes:subtitle>Enables contact-rich robotic manipulation using a VLA model trained with tactile sensing data but requiring no tactile input at inference time. Distills haptic knowledge into the vision-language-action policy.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/6703900b/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>AnoleVLA: Lightweight VLA with Deep State Space Models for Mobile Manipulation</title>
      <itunes:title>AnoleVLA: Lightweight VLA with Deep State Space Models for Mobile Manipulation</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">34feaa37-0643-4e22-9ca2-4fabf78e7bd4</guid>
      <link>https://share.transistor.fm/s/44238ef5</link>
      <description>
        <![CDATA[Proposes a lightweight VLA model leveraging deep state space models (SSMs) instead of transformers for efficient mobile manipulation. Targets resource-constrained deployment scenarios with competitive performance.]]>
      </description>
      <content:encoded>
        <![CDATA[Proposes a lightweight VLA model leveraging deep state space models (SSMs) instead of transformers for efficient mobile manipulation. Targets resource-constrained deployment scenarios with competitive performance.]]>
      </content:encoded>
      <pubDate>Tue, 14 Jul 2026 14:09:29 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/44238ef5/7c8ff59b.mp3" length="29813489" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1864</itunes:duration>
      <itunes:summary>Proposes a lightweight VLA model leveraging deep state space models (SSMs) instead of transformers for efficient mobile manipulation. Targets resource-constrained deployment scenarios with competitive performance.</itunes:summary>
      <itunes:subtitle>Proposes a lightweight VLA model leveraging deep state space models (SSMs) instead of transformers for efficient mobile manipulation. Targets resource-constrained deployment scenarios with competitive performance.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/44238ef5/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>ABot-N1: Visual Language Navigation Foundation Model</title>
      <itunes:title>ABot-N1: Visual Language Navigation Foundation Model</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">ee4f1e88-fb2c-4694-bedd-3717771396a5</guid>
      <link>https://share.transistor.fm/s/6730f2e8</link>
      <description>
        <![CDATA[Decouples cognition from control in a VLM-based navigation policy, delivering 35% POI arrival gains and over 92% success rates in complex indoor and outdoor scenes.]]>
      </description>
      <content:encoded>
        <![CDATA[Decouples cognition from control in a VLM-based navigation policy, delivering 35% POI arrival gains and over 92% success rates in complex indoor and outdoor scenes.]]>
      </content:encoded>
      <pubDate>Tue, 14 Jul 2026 03:19:00 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/6730f2e8/f3a4b03a.mp3" length="37659419" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2354</itunes:duration>
      <itunes:summary>Decouples cognition from control in a VLM-based navigation policy, delivering 35% POI arrival gains and over 92% success rates in complex indoor and outdoor scenes.</itunes:summary>
      <itunes:subtitle>Decouples cognition from control in a VLM-based navigation policy, delivering 35% POI arrival gains and over 92% success rates in complex indoor and outdoor scenes.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/6730f2e8/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>ABot-AgentOS: A General-Purpose Robotic Agent Operating System</title>
      <itunes:title>ABot-AgentOS: A General-Purpose Robotic Agent Operating System</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">6dbf9e3a-3979-4896-8875-a5fe11c672b4</guid>
      <link>https://share.transistor.fm/s/3fb5d4e9</link>
      <description>
        <![CDATA[Provides scene-conditioned planning, context-isolated skill execution, multi-modal memory, and self-evolution capabilities for long-horizon embodied tasks.]]>
      </description>
      <content:encoded>
        <![CDATA[Provides scene-conditioned planning, context-isolated skill execution, multi-modal memory, and self-evolution capabilities for long-horizon embodied tasks.]]>
      </content:encoded>
      <pubDate>Tue, 14 Jul 2026 03:14:45 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/3fb5d4e9/c8f6d6aa.mp3" length="26368670" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1649</itunes:duration>
      <itunes:summary>Provides scene-conditioned planning, context-isolated skill execution, multi-modal memory, and self-evolution capabilities for long-horizon embodied tasks.</itunes:summary>
      <itunes:subtitle>Provides scene-conditioned planning, context-isolated skill execution, multi-modal memory, and self-evolution capabilities for long-horizon embodied tasks.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/3fb5d4e9/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Trust Region Policy Distillation (TOP-D)</title>
      <itunes:title>Trust Region Policy Distillation (TOP-D)</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">a5095cb9-1b53-46b1-bf2d-9d3c2a7528b5</guid>
      <link>https://share.transistor.fm/s/d737e5dd</link>
      <description>
        <![CDATA[Transforms unstable on-policy distillation into a stable training paradigm via dynamic proximal teacher construction, improving sample efficiency without additional computational overhead.]]>
      </description>
      <content:encoded>
        <![CDATA[Transforms unstable on-policy distillation into a stable training paradigm via dynamic proximal teacher construction, improving sample efficiency without additional computational overhead.]]>
      </content:encoded>
      <pubDate>Mon, 13 Jul 2026 14:13:47 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/d737e5dd/3ab49987.mp3" length="32426988" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2027</itunes:duration>
      <itunes:summary>Transforms unstable on-policy distillation into a stable training paradigm via dynamic proximal teacher construction, improving sample efficiency without additional computational overhead.</itunes:summary>
      <itunes:subtitle>Transforms unstable on-policy distillation into a stable training paradigm via dynamic proximal teacher construction, improving sample efficiency without additional computational overhead.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/d737e5dd/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>HANDOFF: Humanoid Agentic Task-Space Whole-Body Control via Distilled Complementary Teachers</title>
      <itunes:title>HANDOFF: Humanoid Agentic Task-Space Whole-Body Control via Distilled Complementary Teachers</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">c565c81c-95d0-4698-bd0a-4c6f204f64f7</guid>
      <link>https://share.transistor.fm/s/982dc788</link>
      <description>
        <![CDATA[A new humanoid whole-body control method that distills complementary teachers for agentic task-space control on humanoids. The approach enables robust whole-body coordination for complex manipulation tasks.]]>
      </description>
      <content:encoded>
        <![CDATA[A new humanoid whole-body control method that distills complementary teachers for agentic task-space control on humanoids. The approach enables robust whole-body coordination for complex manipulation tasks.]]>
      </content:encoded>
      <pubDate>Mon, 13 Jul 2026 14:13:05 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/982dc788/ecad6809.mp3" length="25408617" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1589</itunes:duration>
      <itunes:summary>A new humanoid whole-body control method that distills complementary teachers for agentic task-space control on humanoids. The approach enables robust whole-body coordination for complex manipulation tasks.</itunes:summary>
      <itunes:subtitle>A new humanoid whole-body control method that distills complementary teachers for agentic task-space control on humanoids. The approach enables robust whole-body coordination for complex manipulation tasks.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/982dc788/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>EVA-Client: A Unified Framework for Real-Robot Policy Iteration</title>
      <itunes:title>EVA-Client: A Unified Framework for Real-Robot Policy Iteration</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">c0861c5a-5ac3-4a00-80c6-9029ef00a2cb</guid>
      <link>https://share.transistor.fm/s/81f7b22e</link>
      <description>
        <![CDATA[A unified open framework for real-robot policy iteration that integrates teleoperation data collection, model training, deployment, and evaluation into a single closed-loop pipeline.]]>
      </description>
      <content:encoded>
        <![CDATA[A unified open framework for real-robot policy iteration that integrates teleoperation data collection, model training, deployment, and evaluation into a single closed-loop pipeline.]]>
      </content:encoded>
      <pubDate>Mon, 13 Jul 2026 05:07:06 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/81f7b22e/75821dde.mp3" length="24910410" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1557</itunes:duration>
      <itunes:summary>A unified open framework for real-robot policy iteration that integrates teleoperation data collection, model training, deployment, and evaluation into a single closed-loop pipeline.</itunes:summary>
      <itunes:subtitle>A unified open framework for real-robot policy iteration that integrates teleoperation data collection, model training, deployment, and evaluation into a single closed-loop pipeline.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/81f7b22e/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>UniVR-34B: A Vision-Only Foundation Model for Physical Tasks</title>
      <itunes:title>UniVR-34B: A Vision-Only Foundation Model for Physical Tasks</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">306cbecd-7f5d-46c2-a391-d911b2dd4ba3</guid>
      <link>https://share.transistor.fm/s/789ab4e8</link>
      <description>
        <![CDATA[First large-scale model to learn complex physical dynamics, visual reasoning, and long-horizon planning directly from visual demonstrations without text chains; released with 310k SFT and 3k RL samples across 16 sources alongside the VR-X benchmark.]]>
      </description>
      <content:encoded>
        <![CDATA[First large-scale model to learn complex physical dynamics, visual reasoning, and long-horizon planning directly from visual demonstrations without text chains; released with 310k SFT and 3k RL samples across 16 sources alongside the VR-X benchmark.]]>
      </content:encoded>
      <pubDate>Mon, 13 Jul 2026 03:22:53 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/789ab4e8/643f4e3a.mp3" length="42949946" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2685</itunes:duration>
      <itunes:summary>First large-scale model to learn complex physical dynamics, visual reasoning, and long-horizon planning directly from visual demonstrations without text chains; released with 310k SFT and 3k RL samples across 16 sources alongside the VR-X benchmark.</itunes:summary>
      <itunes:subtitle>First large-scale model to learn complex physical dynamics, visual reasoning, and long-horizon planning directly from visual demonstrations without text chains; released with 310k SFT and 3k RL samples across 16 sources alongside the VR-X benchmark.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/789ab4e8/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>CLAP: Converting Vision-Language Models into Vision-Language-Action Models via Language-Prompted Actions</title>
      <itunes:title>CLAP: Converting Vision-Language Models into Vision-Language-Action Models via Language-Prompted Actions</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">6d9cfa14-7732-4c7d-bf61-65cc75be81b1</guid>
      <link>https://share.transistor.fm/s/d52dbf62</link>
      <description>
        <![CDATA[Converts any pretrained vision-language model into a vision-language-action model with zero architectural changes by prepending natural-language action descriptions; reaches 90.8% success on LIBERO with a 2B model after less than 6 hours of post-training.]]>
      </description>
      <content:encoded>
        <![CDATA[Converts any pretrained vision-language model into a vision-language-action model with zero architectural changes by prepending natural-language action descriptions; reaches 90.8% success on LIBERO with a 2B model after less than 6 hours of post-training.]]>
      </content:encoded>
      <pubDate>Mon, 13 Jul 2026 03:19:11 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/d52dbf62/342d1821.mp3" length="10325306" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>646</itunes:duration>
      <itunes:summary>Converts any pretrained vision-language model into a vision-language-action model with zero architectural changes by prepending natural-language action descriptions; reaches 90.8% success on LIBERO with a 2B model after less than 6 hours of post-training.</itunes:summary>
      <itunes:subtitle>Converts any pretrained vision-language model into a vision-language-action model with zero architectural changes by prepending natural-language action descriptions; reaches 90.8% success on LIBERO with a 2B model after less than 6 hours of post-training.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/d52dbf62/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>InternVLA-A1.5: Unifying Understanding, Latent Foresight, and Action</title>
      <itunes:title>InternVLA-A1.5: Unifying Understanding, Latent Foresight, and Action</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">51ccca13-b681-4e7b-b226-a5100498d382</guid>
      <link>https://share.transistor.fm/s/d3cfb58f</link>
      <description>
        <![CDATA[InternVLA-A15 is a new Vision-Language-Action (VLA) model released by InternRobotics, targeting robotic manipulation and control tasks. The model is accompanied by a Hugging Face collection and an associated technical paper.]]>
      </description>
      <content:encoded>
        <![CDATA[InternVLA-A15 is a new Vision-Language-Action (VLA) model released by InternRobotics, targeting robotic manipulation and control tasks. The model is accompanied by a Hugging Face collection and an associated technical paper.]]>
      </content:encoded>
      <pubDate>Sun, 12 Jul 2026 14:07:51 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/d3cfb58f/9c6c5d9e.mp3" length="30926514" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1933</itunes:duration>
      <itunes:summary>InternVLA-A15 is a new Vision-Language-Action (VLA) model released by InternRobotics, targeting robotic manipulation and control tasks. The model is accompanied by a Hugging Face collection and an associated technical paper.</itunes:summary>
      <itunes:subtitle>InternVLA-A15 is a new Vision-Language-Action (VLA) model released by InternRobotics, targeting robotic manipulation and control tasks. The model is accompanied by a Hugging Face collection and an associated technical paper.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/d3cfb58f/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>MotionWAM: Towards Foundation World Action Models for Real-Time Humanoid Loco-Manipulation</title>
      <itunes:title>MotionWAM: Towards Foundation World Action Models for Real-Time Humanoid Loco-Manipulation</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">7635fe25-3e3f-4f48-b600-f4f7b985a38a</guid>
      <link>https://share.transistor.fm/s/1960dca8</link>
      <description>
        <![CDATA[Proposes foundation world-action models targeting real-time humanoid loco-manipulation, integrating world modeling with action generation for unified locomotion and manipulation control.]]>
      </description>
      <content:encoded>
        <![CDATA[Proposes foundation world-action models targeting real-time humanoid loco-manipulation, integrating world modeling with action generation for unified locomotion and manipulation control.]]>
      </content:encoded>
      <pubDate>Sun, 12 Jul 2026 09:36:13 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/1960dca8/080cd8ab.mp3" length="36138047" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2259</itunes:duration>
      <itunes:summary>Proposes foundation world-action models targeting real-time humanoid loco-manipulation, integrating world modeling with action generation for unified locomotion and manipulation control.</itunes:summary>
      <itunes:subtitle>Proposes foundation world-action models targeting real-time humanoid loco-manipulation, integrating world modeling with action generation for unified locomotion and manipulation control.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/1960dca8/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>On-Device Diffusion Transformer Policy for Efficient Robot Manipulation</title>
      <itunes:title>On-Device Diffusion Transformer Policy for Efficient Robot Manipulation</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">ecaac86f-e326-4784-9dc7-16b57fe587c7</guid>
      <link>https://share.transistor.fm/s/0f780499</link>
      <description>
        <![CDATA[LightDP introduces a framework to accelerate Diffusion Transformer-based policies for real-time, on-device deployment in robot manipulation tasks, achieving efficiency for visuomotor control without sacrificing performance. Accepted to ICCV 2025.]]>
      </description>
      <content:encoded>
        <![CDATA[LightDP introduces a framework to accelerate Diffusion Transformer-based policies for real-time, on-device deployment in robot manipulation tasks, achieving efficiency for visuomotor control without sacrificing performance. Accepted to ICCV 2025.]]>
      </content:encoded>
      <pubDate>Sun, 12 Jul 2026 09:31:23 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/0f780499/badd4f16.mp3" length="14121630" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>883</itunes:duration>
      <itunes:summary>LightDP introduces a framework to accelerate Diffusion Transformer-based policies for real-time, on-device deployment in robot manipulation tasks, achieving efficiency for visuomotor control without sacrificing performance. Accepted to ICCV 2025.</itunes:summary>
      <itunes:subtitle>LightDP introduces a framework to accelerate Diffusion Transformer-based policies for real-time, on-device deployment in robot manipulation tasks, achieving efficiency for visuomotor control without sacrificing performance. Accepted to ICCV 2025.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/0f780499/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Robotic World Model: Neural Dynamics for Locomotion</title>
      <itunes:title>Robotic World Model: Neural Dynamics for Locomotion</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">cea12aeb-e6d3-475a-b93f-1e2393bcb259</guid>
      <link>https://share.transistor.fm/s/425acfb1</link>
      <description>
        <![CDATA[A neural dynamics world model paired with model-free RL for quadruped and humanoid locomotion in IsaacLab, enabling long-horizon autoregressive prediction with policies trained in imagined rollouts that outperform model-based baselines on sim-to-real transfer.]]>
      </description>
      <content:encoded>
        <![CDATA[A neural dynamics world model paired with model-free RL for quadruped and humanoid locomotion in IsaacLab, enabling long-horizon autoregressive prediction with policies trained in imagined rollouts that outperform model-based baselines on sim-to-real transfer.]]>
      </content:encoded>
      <pubDate>Sun, 12 Jul 2026 05:09:14 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/425acfb1/8d6226d4.mp3" length="10430214" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>652</itunes:duration>
      <itunes:summary>A neural dynamics world model paired with model-free RL for quadruped and humanoid locomotion in IsaacLab, enabling long-horizon autoregressive prediction with policies trained in imagined rollouts that outperform model-based baselines on sim-to-real transfer.</itunes:summary>
      <itunes:subtitle>A neural dynamics world model paired with model-free RL for quadruped and humanoid locomotion in IsaacLab, enabling long-horizon autoregressive prediction with policies trained in imagined rollouts that outperform model-based baselines on sim-to-real tran</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/425acfb1/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>LingBot-Vision: Self-Supervised ViT with Masked Boundary Modeling</title>
      <itunes:title>LingBot-Vision: Self-Supervised ViT with Masked Boundary Modeling</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">22050a09-c2ee-4e33-8829-8fbdaaa36ce1</guid>
      <link>https://share.transistor.fm/s/28076e78</link>
      <description>
        <![CDATA[A self-supervised ViT backbone pretrained with masked boundary modeling for dense spatial perception, achieving strong zero-shot results on depth estimation, segmentation, and embodied tasks.]]>
      </description>
      <content:encoded>
        <![CDATA[A self-supervised ViT backbone pretrained with masked boundary modeling for dense spatial perception, achieving strong zero-shot results on depth estimation, segmentation, and embodied tasks.]]>
      </content:encoded>
      <pubDate>Sat, 11 Jul 2026 14:07:31 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/28076e78/5551ae7b.mp3" length="32447050" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2028</itunes:duration>
      <itunes:summary>A self-supervised ViT backbone pretrained with masked boundary modeling for dense spatial perception, achieving strong zero-shot results on depth estimation, segmentation, and embodied tasks.</itunes:summary>
      <itunes:subtitle>A self-supervised ViT backbone pretrained with masked boundary modeling for dense spatial perception, achieving strong zero-shot results on depth estimation, segmentation, and embodied tasks.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/28076e78/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>LaMem-VLA: Latent-Memory-Native VLA Framework</title>
      <itunes:title>LaMem-VLA: Latent-Memory-Native VLA Framework</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">a3f889cb-2e3a-46f0-8401-9dd63b69034f</guid>
      <link>https://share.transistor.fm/s/32d3f09a</link>
      <description>
        <![CDATA[A VLA framework that curates experience into short- and long-term memory vaults, condenses them to latent tokens, and integrates directly into VLA reasoning, evaluated on SimplerEnv and LIBERO benchmarks.]]>
      </description>
      <content:encoded>
        <![CDATA[A VLA framework that curates experience into short- and long-term memory vaults, condenses them to latent tokens, and integrates directly into VLA reasoning, evaluated on SimplerEnv and LIBERO benchmarks.]]>
      </content:encoded>
      <pubDate>Sat, 11 Jul 2026 14:07:00 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/32d3f09a/b45929c4.mp3" length="32141940" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2009</itunes:duration>
      <itunes:summary>A VLA framework that curates experience into short- and long-term memory vaults, condenses them to latent tokens, and integrates directly into VLA reasoning, evaluated on SimplerEnv and LIBERO benchmarks.</itunes:summary>
      <itunes:subtitle>A VLA framework that curates experience into short- and long-term memory vaults, condenses them to latent tokens, and integrates directly into VLA reasoning, evaluated on SimplerEnv and LIBERO benchmarks.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/32d3f09a/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>LingBot-VA 2.0: Native Video-Action Foundation Model for Robot Control</title>
      <itunes:title>LingBot-VA 2.0: Native Video-Action Foundation Model for Robot Control</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">3337a7f1-84c0-4e79-9812-b819ba090f7b</guid>
      <link>https://share.transistor.fm/s/1ad4f13b</link>
      <description>
        <![CDATA[A native video-action foundation model pretrained from scratch with a semantic visual-action tokenizer and foresight reasoning, enabling real-time robot control at ≤150 Hz on consumer GPUs without relying on retrofitted VLMs or video generators.]]>
      </description>
      <content:encoded>
        <![CDATA[A native video-action foundation model pretrained from scratch with a semantic visual-action tokenizer and foresight reasoning, enabling real-time robot control at ≤150 Hz on consumer GPUs without relying on retrofitted VLMs or video generators.]]>
      </content:encoded>
      <pubDate>Sat, 11 Jul 2026 05:10:02 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/1ad4f13b/2078fc82.mp3" length="23261143" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1454</itunes:duration>
      <itunes:summary>A native video-action foundation model pretrained from scratch with a semantic visual-action tokenizer and foresight reasoning, enabling real-time robot control at ≤150 Hz on consumer GPUs without relying on retrofitted VLMs or video generators.</itunes:summary>
      <itunes:subtitle>A native video-action foundation model pretrained from scratch with a semantic visual-action tokenizer and foresight reasoning, enabling real-time robot control at ≤150 Hz on consumer GPUs without relying on retrofitted VLMs or video generators.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/1ad4f13b/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>GigaWorld-1: Large-Scale World Model for Robot Policy Evaluation</title>
      <itunes:title>GigaWorld-1: Large-Scale World Model for Robot Policy Evaluation</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">73b15b6b-3c07-4ff1-8060-112d94c71cc6</guid>
      <link>https://share.transistor.fm/s/09351385</link>
      <description>
        <![CDATA[A large-scale world model trained on 12,980 hours of data and benchmarked across 324,000+ simulated rollouts, designed for comprehensive robot policy evaluation.]]>
      </description>
      <content:encoded>
        <![CDATA[A large-scale world model trained on 12,980 hours of data and benchmarked across 324,000+ simulated rollouts, designed for comprehensive robot policy evaluation.]]>
      </content:encoded>
      <pubDate>Sat, 11 Jul 2026 05:07:46 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/09351385/add44c63.mp3" length="23247768" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1453</itunes:duration>
      <itunes:summary>A large-scale world model trained on 12,980 hours of data and benchmarked across 324,000+ simulated rollouts, designed for comprehensive robot policy evaluation.</itunes:summary>
      <itunes:subtitle>A large-scale world model trained on 12,980 hours of data and benchmarked across 324,000+ simulated rollouts, designed for comprehensive robot policy evaluation.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/09351385/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Learning Unified Force and Position Control for Legged Loco-Manipulation</title>
      <itunes:title>Learning Unified Force and Position Control for Legged Loco-Manipulation</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">ea5722ac-360a-4ec6-8acb-b91146d8f377</guid>
      <link>https://share.transistor.fm/s/703aa5b5</link>
      <description>
        <![CDATA[Proposes a unified RL policy trained in simulation that jointly handles force and position control without force sensors, enabling compliant behaviors and improved force-aware imitation learning on quadrupedal and humanoid robots.]]>
      </description>
      <content:encoded>
        <![CDATA[Proposes a unified RL policy trained in simulation that jointly handles force and position control without force sensors, enabling compliant behaviors and improved force-aware imitation learning on quadrupedal and humanoid robots.]]>
      </content:encoded>
      <pubDate>Fri, 10 Jul 2026 14:14:20 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/703aa5b5/cdb3da5d.mp3" length="27925567" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1746</itunes:duration>
      <itunes:summary>Proposes a unified RL policy trained in simulation that jointly handles force and position control without force sensors, enabling compliant behaviors and improved force-aware imitation learning on quadrupedal and humanoid robots.</itunes:summary>
      <itunes:subtitle>Proposes a unified RL policy trained in simulation that jointly handles force and position control without force sensors, enabling compliant behaviors and improved force-aware imitation learning on quadrupedal and humanoid robots.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/703aa5b5/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Robix: Unified Vision-Language Model for Robotic Reasoning and Planning</title>
      <itunes:title>Robix: Unified Vision-Language Model for Robotic Reasoning and Planning</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">9166a5f4-8d02-46f4-a006-c41fa2d41e54</guid>
      <link>https://share.transistor.fm/s/d912d67a</link>
      <description>
        <![CDATA[A single vision-language model that unifies high-level reasoning, long-horizon planning, and human-robot interaction, serving as a cognitive layer above low-level vision-language-action models.]]>
      </description>
      <content:encoded>
        <![CDATA[A single vision-language model that unifies high-level reasoning, long-horizon planning, and human-robot interaction, serving as a cognitive layer above low-level vision-language-action models.]]>
      </content:encoded>
      <pubDate>Fri, 10 Jul 2026 05:18:52 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/d912d67a/85c855bd.mp3" length="24470717" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1530</itunes:duration>
      <itunes:summary>A single vision-language model that unifies high-level reasoning, long-horizon planning, and human-robot interaction, serving as a cognitive layer above low-level vision-language-action models.</itunes:summary>
      <itunes:subtitle>A single vision-language model that unifies high-level reasoning, long-horizon planning, and human-robot interaction, serving as a cognitive layer above low-level vision-language-action models.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/d912d67a/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>EmbodiedOneVision (EO-1): A Unified Decoder-Only Transformer for General Robot Control</title>
      <itunes:title>EmbodiedOneVision (EO-1): A Unified Decoder-Only Transformer for General Robot Control</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">547b4a0b-038e-4a26-ba1b-f7d3eead719f</guid>
      <link>https://share.transistor.fm/s/c67cdf34</link>
      <description>
        <![CDATA[A 3B-parameter open-source unified decoder-only transformer for general robot control that jointly handles perception, planning, reasoning, and action via interleaved vision-text-action pretraining on a 1.5M-sample dataset. Achieves strong results across manipulation tasks and benchmarks with fully open-sourced model, code, and dataset.]]>
      </description>
      <content:encoded>
        <![CDATA[A 3B-parameter open-source unified decoder-only transformer for general robot control that jointly handles perception, planning, reasoning, and action via interleaved vision-text-action pretraining on a 1.5M-sample dataset. Achieves strong results across manipulation tasks and benchmarks with fully open-sourced model, code, and dataset.]]>
      </content:encoded>
      <pubDate>Thu, 09 Jul 2026 14:20:24 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/c67cdf34/7491efed.mp3" length="39571582" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2474</itunes:duration>
      <itunes:summary>A 3B-parameter open-source unified decoder-only transformer for general robot control that jointly handles perception, planning, reasoning, and action via interleaved vision-text-action pretraining on a 1.5M-sample dataset. Achieves strong results across manipulation tasks and benchmarks with fully open-sourced model, code, and dataset.</itunes:summary>
      <itunes:subtitle>A 3B-parameter open-source unified decoder-only transformer for general robot control that jointly handles perception, planning, reasoning, and action via interleaved vision-text-action pretraining on a 1.5M-sample dataset. Achieves strong results across </itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/c67cdf34/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>mimic-video: Video-Action Models for Robot Learning</title>
      <itunes:title>mimic-video: Video-Action Models for Robot Learning</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">e17abb2d-3c3a-44a3-b6f8-5b734edfcd5d</guid>
      <link>https://share.transistor.fm/s/71731413</link>
      <description>
        <![CDATA[Introduces Video-Action Models (VAMs) that replace standard VLMs in VLAs with a pretrained internet-scale video backbone (Cosmos-Predict2) combined with a flow-matching action decoder. Achieves approximately 10× better sample efficiency on real-world pick-and-place tasks.]]>
      </description>
      <content:encoded>
        <![CDATA[Introduces Video-Action Models (VAMs) that replace standard VLMs in VLAs with a pretrained internet-scale video backbone (Cosmos-Predict2) combined with a flow-matching action decoder. Achieves approximately 10× better sample efficiency on real-world pick-and-place tasks.]]>
      </content:encoded>
      <pubDate>Thu, 09 Jul 2026 14:12:13 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/71731413/d2df39ac.mp3" length="34182834" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2137</itunes:duration>
      <itunes:summary>Introduces Video-Action Models (VAMs) that replace standard VLMs in VLAs with a pretrained internet-scale video backbone (Cosmos-Predict2) combined with a flow-matching action decoder. Achieves approximately 10× better sample efficiency on real-world pick-and-place tasks.</itunes:summary>
      <itunes:subtitle>Introduces Video-Action Models (VAMs) that replace standard VLMs in VLAs with a pretrained internet-scale video backbone (Cosmos-Predict2) combined with a flow-matching action decoder. Achieves approximately 10× better sample efficiency on real-world pick</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/71731413/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Lowering the Barrier: The GEM Open-Source Arm</title>
      <itunes:title>Lowering the Barrier: The GEM Open-Source Arm</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">e23bed6e-3854-4470-a53c-14b6be1a2f51</guid>
      <link>https://share.transistor.fm/s/1d83dc55</link>
      <description>
        <![CDATA[An open-source, sub-$500 7-DOF 3D-printed robotic arm with 1.2 kg payload, wrist and head cameras, and full LeRobot integration, designed to lower the barrier to entry for real manipulation research.]]>
      </description>
      <content:encoded>
        <![CDATA[An open-source, sub-$500 7-DOF 3D-printed robotic arm with 1.2 kg payload, wrist and head cameras, and full LeRobot integration, designed to lower the barrier to entry for real manipulation research.]]>
      </content:encoded>
      <pubDate>Thu, 09 Jul 2026 05:14:22 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/1d83dc55/558202df.mp3" length="29703148" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1857</itunes:duration>
      <itunes:summary>An open-source, sub-$500 7-DOF 3D-printed robotic arm with 1.2 kg payload, wrist and head cameras, and full LeRobot integration, designed to lower the barrier to entry for real manipulation research.</itunes:summary>
      <itunes:subtitle>An open-source, sub-$500 7-DOF 3D-printed robotic arm with 1.2 kg payload, wrist and head cameras, and full LeRobot integration, designed to lower the barrier to entry for real manipulation research.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/1d83dc55/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>RynnWorld-4D: A 4D Embodied World Model for Bimanual Robot Control</title>
      <itunes:title>RynnWorld-4D: A 4D Embodied World Model for Bimanual Robot Control</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">bc19084d-e990-47ed-86cf-5b0fc2a924dc</guid>
      <link>https://share.transistor.fm/s/5d248c5b</link>
      <description>
        <![CDATA[A 4D embodied world model that predicts RGB, depth, and optical flow from RGB-D input and instructions using a tri-branch diffusion architecture. Enables closed-loop bimanual robot control by bridging world prediction and policy execution.]]>
      </description>
      <content:encoded>
        <![CDATA[A 4D embodied world model that predicts RGB, depth, and optical flow from RGB-D input and instructions using a tri-branch diffusion architecture. Enables closed-loop bimanual robot control by bridging world prediction and policy execution.]]>
      </content:encoded>
      <pubDate>Wed, 08 Jul 2026 14:11:44 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/5d248c5b/a4d1d241.mp3" length="9959174" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>623</itunes:duration>
      <itunes:summary>A 4D embodied world model that predicts RGB, depth, and optical flow from RGB-D input and instructions using a tri-branch diffusion architecture. Enables closed-loop bimanual robot control by bridging world prediction and policy execution.</itunes:summary>
      <itunes:subtitle>A 4D embodied world model that predicts RGB, depth, and optical flow from RGB-D input and instructions using a tri-branch diffusion architecture. Enables closed-loop bimanual robot control by bridging world prediction and policy execution.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/5d248c5b/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>RynnWorld-Teleop: Digital Teleoperation via World Model Rendering</title>
      <itunes:title>RynnWorld-Teleop: Digital Teleoperation via World Model Rendering</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">5a3d34fd-4306-4e4c-9203-94257e4edd5a</guid>
      <link>https://share.transistor.fm/s/433922fe</link>
      <description>
        <![CDATA[A digital teleoperation system that uses hand-pose streams to drive a world model for real-time (40+ FPS) high-fidelity robot video rendering from a single image. Removes physical robots from the training loop entirely.]]>
      </description>
      <content:encoded>
        <![CDATA[A digital teleoperation system that uses hand-pose streams to drive a world model for real-time (40+ FPS) high-fidelity robot video rendering from a single image. Removes physical robots from the training loop entirely.]]>
      </content:encoded>
      <pubDate>Wed, 08 Jul 2026 14:08:55 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/433922fe/e0d4bade.mp3" length="27752114" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1735</itunes:duration>
      <itunes:summary>A digital teleoperation system that uses hand-pose streams to drive a world model for real-time (40+ FPS) high-fidelity robot video rendering from a single image. Removes physical robots from the training loop entirely.</itunes:summary>
      <itunes:subtitle>A digital teleoperation system that uses hand-pose streams to drive a world model for real-time (40+ FPS) high-fidelity robot video rendering from a single image. Removes physical robots from the training loop entirely.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/433922fe/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>VLA-Corrector: Adaptive Action Horizons through Latent Visual Monitoring</title>
      <itunes:title>VLA-Corrector: Adaptive Action Horizons through Latent Visual Monitoring</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">dd209aa1-b7da-497f-8f6e-04ab483f3cd6</guid>
      <link>https://share.transistor.fm/s/9d30b31f</link>
      <description>
        <![CDATA[A lightweight plug-in that monitors latent visual dynamics in action-chunked policies, drops stale actions on drift, and replans on-the-fly achieving +17.7 success points while keeping the VLA backbone frozen.]]>
      </description>
      <content:encoded>
        <![CDATA[A lightweight plug-in that monitors latent visual dynamics in action-chunked policies, drops stale actions on drift, and replans on-the-fly achieving +17.7 success points while keeping the VLA backbone frozen.]]>
      </content:encoded>
      <pubDate>Wed, 08 Jul 2026 05:14:06 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/9d30b31f/80b209ac.mp3" length="35208924" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2201</itunes:duration>
      <itunes:summary>A lightweight plug-in that monitors latent visual dynamics in action-chunked policies, drops stale actions on drift, and replans on-the-fly achieving +17.7 success points while keeping the VLA backbone frozen.</itunes:summary>
      <itunes:subtitle>A lightweight plug-in that monitors latent visual dynamics in action-chunked policies, drops stale actions on drift, and replans on-the-fly achieving +17.7 success points while keeping the VLA backbone frozen.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/9d30b31f/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Scenes as Objects, Not Primitives: Instance-Structured 3D Tokenization</title>
      <itunes:title>Scenes as Objects, Not Primitives: Instance-Structured 3D Tokenization</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">9d6beb91-0d9c-4a16-af23-d324a40c4962</guid>
      <link>https://share.transistor.fm/s/0d8c2bdf</link>
      <description>
        <![CDATA[A feed-forward model that decomposes unposed images into instance-structured 3D token groups without annotations, enabling unified reconstruction, segmentation, and manipulation.]]>
      </description>
      <content:encoded>
        <![CDATA[A feed-forward model that decomposes unposed images into instance-structured 3D token groups without annotations, enabling unified reconstruction, segmentation, and manipulation.]]>
      </content:encoded>
      <pubDate>Wed, 08 Jul 2026 05:09:22 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/0d8c2bdf/d09b81af.mp3" length="20356745" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1273</itunes:duration>
      <itunes:summary>A feed-forward model that decomposes unposed images into instance-structured 3D token groups without annotations, enabling unified reconstruction, segmentation, and manipulation.</itunes:summary>
      <itunes:subtitle>A feed-forward model that decomposes unposed images into instance-structured 3D token groups without annotations, enabling unified reconstruction, segmentation, and manipulation.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/0d8c2bdf/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Current as Touch: Proprioceptive Contact Feedback for Compliant Dexterous Manipulation</title>
      <itunes:title>Current as Touch: Proprioceptive Contact Feedback for Compliant Dexterous Manipulation</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">46f695f4-89f9-46a2-a6d6-2b1985290332</guid>
      <link>https://share.transistor.fm/s/a6e977cc</link>
      <description>
        <![CDATA[A proprioceptive contact feedback method leveraging motor current signals as a touch proxy to enable compliant dexterous manipulation without dedicated tactile sensors.]]>
      </description>
      <content:encoded>
        <![CDATA[A proprioceptive contact feedback method leveraging motor current signals as a touch proxy to enable compliant dexterous manipulation without dedicated tactile sensors.]]>
      </content:encoded>
      <pubDate>Tue, 07 Jul 2026 14:09:34 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/a6e977cc/df1a2d7f.mp3" length="28088154" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1756</itunes:duration>
      <itunes:summary>A proprioceptive contact feedback method leveraging motor current signals as a touch proxy to enable compliant dexterous manipulation without dedicated tactile sensors.</itunes:summary>
      <itunes:subtitle>A proprioceptive contact feedback method leveraging motor current signals as a touch proxy to enable compliant dexterous manipulation without dedicated tactile sensors.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/a6e977cc/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>DART: One-Shot VLA Policy Adaptation via Weight-Space Arithmetic</title>
      <itunes:title>DART: One-Shot VLA Policy Adaptation via Weight-Space Arithmetic</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">8ba99436-6f0a-47e5-a01b-fa892db13fb7</guid>
      <link>https://share.transistor.fm/s/39efd30c</link>
      <description>
        <![CDATA[Uses weight-space arithmetic to isolate domain shifts from task knowledge, enabling one-shot VLA policy adaptation to new cameras or embodiments.]]>
      </description>
      <content:encoded>
        <![CDATA[Uses weight-space arithmetic to isolate domain shifts from task knowledge, enabling one-shot VLA policy adaptation to new cameras or embodiments.]]>
      </content:encoded>
      <pubDate>Tue, 07 Jul 2026 05:16:14 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/39efd30c/7f89ce0e.mp3" length="22369280" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1399</itunes:duration>
      <itunes:summary>Uses weight-space arithmetic to isolate domain shifts from task knowledge, enabling one-shot VLA policy adaptation to new cameras or embodiments.</itunes:summary>
      <itunes:subtitle>Uses weight-space arithmetic to isolate domain shifts from task knowledge, enabling one-shot VLA policy adaptation to new cameras or embodiments.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/39efd30c/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Does VLA Even Know the Basics? Act2Answer Benchmark</title>
      <itunes:title>Does VLA Even Know the Basics? Act2Answer Benchmark</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">c3097403-1603-48bc-9c78-5d434dfb4dfe</guid>
      <link>https://share.transistor.fm/s/61da4fe9</link>
      <description>
        <![CDATA[Introduces a benchmark showing VLAs lose 20–40 points in commonsense/world knowledge versus source VLMs after robotics fine-tuning, evaluated via action-based answering.]]>
      </description>
      <content:encoded>
        <![CDATA[Introduces a benchmark showing VLAs lose 20–40 points in commonsense/world knowledge versus source VLMs after robotics fine-tuning, evaluated via action-based answering.]]>
      </content:encoded>
      <pubDate>Tue, 07 Jul 2026 05:12:02 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/61da4fe9/09ee348b.mp3" length="25677312" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1605</itunes:duration>
      <itunes:summary>Introduces a benchmark showing VLAs lose 20–40 points in commonsense/world knowledge versus source VLMs after robotics fine-tuning, evaluated via action-based answering.</itunes:summary>
      <itunes:subtitle>Introduces a benchmark showing VLAs lose 20–40 points in commonsense/world knowledge versus source VLMs after robotics fine-tuning, evaluated via action-based answering.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/61da4fe9/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Contact-Grounded Policy: Dexterous Visuotactile Policy with Generative Contact Grounding</title>
      <itunes:title>Contact-Grounded Policy: Dexterous Visuotactile Policy with Generative Contact Grounding</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">f02e3f94-8a68-4ee6-9fba-ec1c9f42169b</guid>
      <link>https://share.transistor.fm/s/7e8c0128</link>
      <description>
        <![CDATA[Learns dexterous manipulation policies that explicitly ground actions in generative contact predictions from visuotactile observations, improving robustness on contact-rich tasks.]]>
      </description>
      <content:encoded>
        <![CDATA[Learns dexterous manipulation policies that explicitly ground actions in generative contact predictions from visuotactile observations, improving robustness on contact-rich tasks.]]>
      </content:encoded>
      <pubDate>Mon, 06 Jul 2026 14:07:09 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/7e8c0128/29086dbb.mp3" length="25790976" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1612</itunes:duration>
      <itunes:summary>Learns dexterous manipulation policies that explicitly ground actions in generative contact predictions from visuotactile observations, improving robustness on contact-rich tasks.</itunes:summary>
      <itunes:subtitle>Learns dexterous manipulation policies that explicitly ground actions in generative contact predictions from visuotactile observations, improving robustness on contact-rich tasks.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/7e8c0128/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Freeform Preference Learning (FPL) for Robotic Manipulation</title>
      <itunes:title>Freeform Preference Learning (FPL) for Robotic Manipulation</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">cb0b6f60-b518-48c9-a59c-35774a61dfd1</guid>
      <link>https://share.transistor.fm/s/70c3ca92</link>
      <description>
        <![CDATA[Introduces multi-axis preference supervision to learn dense, language-conditioned rewards across speed/precision/subtask axes without segmentation; enables compositional generalization and better long-horizon credit assignment than single-reward baselines.]]>
      </description>
      <content:encoded>
        <![CDATA[Introduces multi-axis preference supervision to learn dense, language-conditioned rewards across speed/precision/subtask axes without segmentation; enables compositional generalization and better long-horizon credit assignment than single-reward baselines.]]>
      </content:encoded>
      <pubDate>Mon, 06 Jul 2026 05:09:07 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/70c3ca92/b500537c.mp3" length="47148032" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2947</itunes:duration>
      <itunes:summary>Introduces multi-axis preference supervision to learn dense, language-conditioned rewards across speed/precision/subtask axes without segmentation; enables compositional generalization and better long-horizon credit assignment than single-reward baselines.</itunes:summary>
      <itunes:subtitle>Introduces multi-axis preference supervision to learn dense, language-conditioned rewards across speed/precision/subtask axes without segmentation; enables compositional generalization and better long-horizon credit assignment than single-reward baselines</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/70c3ca92/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Orca: The World is in Your Mind</title>
      <itunes:title>Orca: The World is in Your Mind</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">1ef890f0-f460-4cd5-9768-cf63a2564089</guid>
      <link>https://share.transistor.fm/s/50663635</link>
      <description>
        <![CDATA[Proposes a general world foundation model leveraging Next-State-Prediction to jointly generate text, images, and embodied actions within a unified framework.]]>
      </description>
      <content:encoded>
        <![CDATA[Proposes a general world foundation model leveraging Next-State-Prediction to jointly generate text, images, and embodied actions within a unified framework.]]>
      </content:encoded>
      <pubDate>Sun, 05 Jul 2026 14:09:11 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/50663635/f8c6f3f8.mp3" length="27739648" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1734</itunes:duration>
      <itunes:summary>Proposes a general world foundation model leveraging Next-State-Prediction to jointly generate text, images, and embodied actions within a unified framework.</itunes:summary>
      <itunes:subtitle>Proposes a general world foundation model leveraging Next-State-Prediction to jointly generate text, images, and embodied actions within a unified framework.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/50663635/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Qwen-RobotNav: A Scalable Unified Navigation Model for Agentic Robotics</title>
      <itunes:title>Qwen-RobotNav: A Scalable Unified Navigation Model for Agentic Robotics</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">5d86aca2-7ff2-4227-87a7-a71f93f4cf19</guid>
      <link>https://share.transistor.fm/s/dbaffd65</link>
      <description>
        <![CDATA[A unified 2B–8B parameter model for robot navigation tasks (VLN, ObjectNav, tracking, autonomous driving) via a configurable observation protocol, with demonstrated zero-shot deployment on real quadruped robots using agentic planners.]]>
      </description>
      <content:encoded>
        <![CDATA[A unified 2B–8B parameter model for robot navigation tasks (VLN, ObjectNav, tracking, autonomous driving) via a configurable observation protocol, with demonstrated zero-shot deployment on real quadruped robots using agentic planners.]]>
      </content:encoded>
      <pubDate>Sun, 05 Jul 2026 14:08:08 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/dbaffd65/975718d3.mp3" length="35041792" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2191</itunes:duration>
      <itunes:summary>A unified 2B–8B parameter model for robot navigation tasks (VLN, ObjectNav, tracking, autonomous driving) via a configurable observation protocol, with demonstrated zero-shot deployment on real quadruped robots using agentic planners.</itunes:summary>
      <itunes:subtitle>A unified 2B–8B parameter model for robot navigation tasks (VLN, ObjectNav, tracking, autonomous driving) via a configurable observation protocol, with demonstrated zero-shot deployment on real quadruped robots using agentic planners.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/dbaffd65/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Scaling Robot Skills from Cheap Human Videos</title>
      <itunes:title>Scaling Robot Skills from Cheap Human Videos</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">d6eb6b3f-4e21-4b37-be02-3210dcadae0a</guid>
      <link>https://share.transistor.fm/s/e3476d3b</link>
      <description>
        <![CDATA[Replaces noisy 6-DoF hand poses with relative wrist translation as a shared action space between humans and bimanual robots, enabling scalable skill acquisition from inexpensive video data. This approach outperforms full-pose baselines for robot skill learning.]]>
      </description>
      <content:encoded>
        <![CDATA[Replaces noisy 6-DoF hand poses with relative wrist translation as a shared action space between humans and bimanual robots, enabling scalable skill acquisition from inexpensive video data. This approach outperforms full-pose baselines for robot skill learning.]]>
      </content:encoded>
      <pubDate>Wed, 01 Jul 2026 14:08:01 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/e3476d3b/ff71f10c.mp3" length="14300672" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>894</itunes:duration>
      <itunes:summary>Replaces noisy 6-DoF hand poses with relative wrist translation as a shared action space between humans and bimanual robots, enabling scalable skill acquisition from inexpensive video data. This approach outperforms full-pose baselines for robot skill learning.</itunes:summary>
      <itunes:subtitle>Replaces noisy 6-DoF hand poses with relative wrist translation as a shared action space between humans and bimanual robots, enabling scalable skill acquisition from inexpensive video data. This approach outperforms full-pose baselines for robot skill lea</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/e3476d3b/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>ABC: An Open Behavior Cloning Stack for Bimanual Manipulation</title>
      <itunes:title>ABC: An Open Behavior Cloning Stack for Bimanual Manipulation</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">8a31145f-2af7-4ed3-b27f-bb1280d2f4c0</guid>
      <link>https://share.transistor.fm/s/1dfc015b</link>
      <description>
        <![CDATA[Large-scale open-source framework for real-world robotic manipulation using behavior cloning, including the ABC-130K dataset with 3,500 hours and 130K+ episodes across 195 tasks. Provides hardware setups, simulators, and training recipes for Diffusion Transformers and Vision-Language-Action models.]]>
      </description>
      <content:encoded>
        <![CDATA[Large-scale open-source framework for real-world robotic manipulation using behavior cloning, including the ABC-130K dataset with 3,500 hours and 130K+ episodes across 195 tasks. Provides hardware setups, simulators, and training recipes for Diffusion Transformers and Vision-Language-Action models.]]>
      </content:encoded>
      <pubDate>Wed, 01 Jul 2026 05:16:54 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/1dfc015b/2647bc85.mp3" length="31250944" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1954</itunes:duration>
      <itunes:summary>Large-scale open-source framework for real-world robotic manipulation using behavior cloning, including the ABC-130K dataset with 3,500 hours and 130K+ episodes across 195 tasks. Provides hardware setups, simulators, and training recipes for Diffusion Transformers and Vision-Language-Action models.</itunes:summary>
      <itunes:subtitle>Large-scale open-source framework for real-world robotic manipulation using behavior cloning, including the ABC-130K dataset with 3,500 hours and 130K+ episodes across 195 tasks. Provides hardware setups, simulators, and training recipes for Diffusion Tra</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/1dfc015b/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Qwen-RobotManip: Alignment Unlocks Scale for Robotic Manipulation Foundation Models</title>
      <itunes:title>Qwen-RobotManip: Alignment Unlocks Scale for Robotic Manipulation Foundation Models</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">ee4b63bb-8b4c-4f6a-949f-703cf3d886be</guid>
      <link>https://share.transistor.fm/s/ea2d6565</link>
      <description>
        <![CDATA[Presents alignment techniques to scale robotic manipulation foundation models, building on the Qwen model family for dexterous robot control.]]>
      </description>
      <content:encoded>
        <![CDATA[Presents alignment techniques to scale robotic manipulation foundation models, building on the Qwen model family for dexterous robot control.]]>
      </content:encoded>
      <pubDate>Wed, 01 Jul 2026 05:12:07 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/ea2d6565/c580bbd9.mp3" length="19389440" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1212</itunes:duration>
      <itunes:summary>Presents alignment techniques to scale robotic manipulation foundation models, building on the Qwen model family for dexterous robot control.</itunes:summary>
      <itunes:subtitle>Presents alignment techniques to scale robotic manipulation foundation models, building on the Qwen model family for dexterous robot control.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/ea2d6565/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>ASPIRE: Automated Skill Discovery for Robotics</title>
      <itunes:title>ASPIRE: Automated Skill Discovery for Robotics</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">7335a78b-dd1f-4e79-afa4-fe667d1a24d3</guid>
      <link>https://share.transistor.fm/s/7caa42bb</link>
      <description>
        <![CDATA[Introduces the first automated system that continuously discovers, evolves, and accumulates reusable sensorimotor skills via evolutionary search over control programs, enabling compounding multi-task, sim-to-real, and cross-embodiment transfer without retraining end-to-end policies.]]>
      </description>
      <content:encoded>
        <![CDATA[Introduces the first automated system that continuously discovers, evolves, and accumulates reusable sensorimotor skills via evolutionary search over control programs, enabling compounding multi-task, sim-to-real, and cross-embodiment transfer without retraining end-to-end policies.]]>
      </content:encoded>
      <pubDate>Tue, 30 Jun 2026 14:11:15 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/7caa42bb/22a914c5.mp3" length="29518336" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1845</itunes:duration>
      <itunes:summary>Introduces the first automated system that continuously discovers, evolves, and accumulates reusable sensorimotor skills via evolutionary search over control programs, enabling compounding multi-task, sim-to-real, and cross-embodiment transfer without retraining end-to-end policies.</itunes:summary>
      <itunes:subtitle>Introduces the first automated system that continuously discovers, evolves, and accumulates reusable sensorimotor skills via evolutionary search over control programs, enabling compounding multi-task, sim-to-real, and cross-embodiment transfer without ret</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/7caa42bb/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>SERF: 4D Latent Mapping for Long-Horizon Mobile Manipulation</title>
      <itunes:title>SERF: 4D Latent Mapping for Long-Horizon Mobile Manipulation</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">316b2dcb-797a-4d92-872b-5966ff33d3fc</guid>
      <link>https://share.transistor.fm/s/2624b28e</link>
      <description>
        <![CDATA[Embeds both the robot and environment into a shared 4D latent space augmented with forward-kinematics robot points, enabling a vision-language-action model to handle dynamic scenes and long-horizon memory. Outperforms image-only VLA baselines on the BEHAVIOR-1K benchmark for mobile manipulation.]]>
      </description>
      <content:encoded>
        <![CDATA[Embeds both the robot and environment into a shared 4D latent space augmented with forward-kinematics robot points, enabling a vision-language-action model to handle dynamic scenes and long-horizon memory. Outperforms image-only VLA baselines on the BEHAVIOR-1K benchmark for mobile manipulation.]]>
      </content:encoded>
      <pubDate>Tue, 30 Jun 2026 05:48:09 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/2624b28e/635205a1.mp3" length="32793088" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2050</itunes:duration>
      <itunes:summary>Embeds both the robot and environment into a shared 4D latent space augmented with forward-kinematics robot points, enabling a vision-language-action model to handle dynamic scenes and long-horizon memory. Outperforms image-only VLA baselines on the BEHAVIOR-1K benchmark for mobile manipulation.</itunes:summary>
      <itunes:subtitle>Embeds both the robot and environment into a shared 4D latent space augmented with forward-kinematics robot points, enabling a vision-language-action model to handle dynamic scenes and long-horizon memory. Outperforms image-only VLA baselines on the BEHAV</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/2624b28e/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>ViserDex: Visual Sim-to-Real for Robust Dexterous In-Hand Reorientation</title>
      <itunes:title>ViserDex: Visual Sim-to-Real for Robust Dexterous In-Hand Reorientation</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">71c50b16-b993-4c1e-a173-35041c337010</guid>
      <link>https://share.transistor.fm/s/3a93eab2</link>
      <description>
        <![CDATA[A single-camera sim-to-real framework that uses physically consistent 3D Gaussian Splatting augmentations to achieve zero-shot transfer of dexterous in-hand reorientation policies to an Allegro hand. The approach trains entirely on consumer hardware while maintaining high fidelity to real-world dynamics.]]>
      </description>
      <content:encoded>
        <![CDATA[A single-camera sim-to-real framework that uses physically consistent 3D Gaussian Splatting augmentations to achieve zero-shot transfer of dexterous in-hand reorientation policies to an Allegro hand. The approach trains entirely on consumer hardware while maintaining high fidelity to real-world dynamics.]]>
      </content:encoded>
      <pubDate>Tue, 30 Jun 2026 05:13:30 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/3a93eab2/66e257ab.mp3" length="30902784" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1932</itunes:duration>
      <itunes:summary>A single-camera sim-to-real framework that uses physically consistent 3D Gaussian Splatting augmentations to achieve zero-shot transfer of dexterous in-hand reorientation policies to an Allegro hand. The approach trains entirely on consumer hardware while maintaining high fidelity to real-world dynamics.</itunes:summary>
      <itunes:subtitle>A single-camera sim-to-real framework that uses physically consistent 3D Gaussian Splatting augmentations to achieve zero-shot transfer of dexterous in-hand reorientation policies to an Allegro hand. The approach trains entirely on consumer hardware while</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/3a93eab2/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>DexSkin: A High-Coverage, Conformable "Electronic Skin" for Robot Fingers</title>
      <itunes:title>DexSkin: A High-Coverage, Conformable "Electronic Skin" for Robot Fingers</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">edb95e7c-57fb-4fea-9a1c-68acff653bd0</guid>
      <link>https://share.transistor.fm/s/3802d4c5</link>
      <description>
        <![CDATA[Introduces a high-coverage, conformable robotic skin hardware system designed to improve data collection and policy learning for contact-rich, dexterous manipulation tasks. The system provides rich tactile sensing coverage to enable more capable robot manipulation policies.]]>
      </description>
      <content:encoded>
        <![CDATA[Introduces a high-coverage, conformable robotic skin hardware system designed to improve data collection and policy learning for contact-rich, dexterous manipulation tasks. The system provides rich tactile sensing coverage to enable more capable robot manipulation policies.]]>
      </content:encoded>
      <pubDate>Tue, 30 Jun 2026 03:12:49 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/3802d4c5/42d1d31a.mp3" length="33273856" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2080</itunes:duration>
      <itunes:summary>Introduces a high-coverage, conformable robotic skin hardware system designed to improve data collection and policy learning for contact-rich, dexterous manipulation tasks. The system provides rich tactile sensing coverage to enable more capable robot manipulation policies.</itunes:summary>
      <itunes:subtitle>Introduces a high-coverage, conformable robotic skin hardware system designed to improve data collection and policy learning for contact-rich, dexterous manipulation tasks. The system provides rich tactile sensing coverage to enable more capable robot man</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/3802d4c5/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>EBench: A Diagnostic Benchmark for Generalist Manipulation Policies</title>
      <itunes:title>EBench: A Diagnostic Benchmark for Generalist Manipulation Policies</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">3f17dfc1-2a8c-4258-97b9-1e7161ac37a4</guid>
      <link>https://share.transistor.fm/s/28fe8c63</link>
      <description>
        <![CDATA[A CAT-scan style diagnostic benchmark for robot foundation models that evaluates policies such as π0, π0.5, and Qwen-RobotManip beyond single success rates. The benchmark is designed to distinguish genuine generalization from overfitting to demonstrations in generalist manipulation policies.]]>
      </description>
      <content:encoded>
        <![CDATA[A CAT-scan style diagnostic benchmark for robot foundation models that evaluates policies such as π0, π0.5, and Qwen-RobotManip beyond single success rates. The benchmark is designed to distinguish genuine generalization from overfitting to demonstrations in generalist manipulation policies.]]>
      </content:encoded>
      <pubDate>Tue, 30 Jun 2026 03:11:09 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/28fe8c63/bbc4fd90.mp3" length="19579904" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1224</itunes:duration>
      <itunes:summary>A CAT-scan style diagnostic benchmark for robot foundation models that evaluates policies such as π0, π0.5, and Qwen-RobotManip beyond single success rates. The benchmark is designed to distinguish genuine generalization from overfitting to demonstrations in generalist manipulation policies.</itunes:summary>
      <itunes:subtitle>A CAT-scan style diagnostic benchmark for robot foundation models that evaluates policies such as π0, π0.5, and Qwen-RobotManip beyond single success rates. The benchmark is designed to distinguish genuine generalization from overfitting to demonstrations</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/28fe8c63/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>VITRA: A Foundation for Dexterous VLA via Human Video Pretraining</title>
      <itunes:title>VITRA: A Foundation for Dexterous VLA via Human Video Pretraining</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">bb1a81d2-577f-4cf8-a13c-b7e10672e74a</guid>
      <link>https://share.transistor.fm/s/094b08db</link>
      <description>
        <![CDATA[A scalable VLA pretraining pipeline that converts unstructured egocentric human videos into robot training data, trains a dexterous hand VLA, and fine-tunes on robot data, achieving strong zero-shot generalization and real-robot dexterous manipulation.]]>
      </description>
      <content:encoded>
        <![CDATA[A scalable VLA pretraining pipeline that converts unstructured egocentric human videos into robot training data, trains a dexterous hand VLA, and fine-tunes on robot data, achieving strong zero-shot generalization and real-robot dexterous manipulation.]]>
      </content:encoded>
      <pubDate>Mon, 29 Jun 2026 14:13:41 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/094b08db/69b5523c.mp3" length="16903168" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1057</itunes:duration>
      <itunes:summary>A scalable VLA pretraining pipeline that converts unstructured egocentric human videos into robot training data, trains a dexterous hand VLA, and fine-tunes on robot data, achieving strong zero-shot generalization and real-robot dexterous manipulation.</itunes:summary>
      <itunes:subtitle>A scalable VLA pretraining pipeline that converts unstructured egocentric human videos into robot training data, trains a dexterous hand VLA, and fine-tunes on robot data, achieving strong zero-shot generalization and real-robot dexterous manipulation.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/094b08db/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>DexWM: A Dexterous Manipulation World Model from Human Videos</title>
      <itunes:title>DexWM: A Dexterous Manipulation World Model from Human Videos</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">6436a7f3-c028-4c4c-9cfa-d9f9ff6df3d9</guid>
      <link>https://share.transistor.fm/s/e96429b5</link>
      <description>
        <![CDATA[A dexterous manipulation world model pretrained on 829 hours of EgoDex human data and DROID robot data using conditioned diffusion transformers, enabling open-loop rollouts and sim-to-real transfer with minimal robot fine-tuning.]]>
      </description>
      <content:encoded>
        <![CDATA[A dexterous manipulation world model pretrained on 829 hours of EgoDex human data and DROID robot data using conditioned diffusion transformers, enabling open-loop rollouts and sim-to-real transfer with minimal robot fine-tuning.]]>
      </content:encoded>
      <pubDate>Mon, 29 Jun 2026 14:08:49 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/e96429b5/c86e2612.mp3" length="30289408" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1894</itunes:duration>
      <itunes:summary>A dexterous manipulation world model pretrained on 829 hours of EgoDex human data and DROID robot data using conditioned diffusion transformers, enabling open-loop rollouts and sim-to-real transfer with minimal robot fine-tuning.</itunes:summary>
      <itunes:subtitle>A dexterous manipulation world model pretrained on 829 hours of EgoDex human data and DROID robot data using conditioned diffusion transformers, enabling open-loop rollouts and sim-to-real transfer with minimal robot fine-tuning.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/e96429b5/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>PoLAR: Factorizing Extent and Mode in Latent Actions for Robot Policy Learning</title>
      <itunes:title>PoLAR: Factorizing Extent and Mode in Latent Actions for Robot Policy Learning</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">a0b24a84-6081-4cd7-8160-af4887a71966</guid>
      <link>https://share.transistor.fm/s/e5e4186b</link>
      <description>
        <![CDATA[Introduces PoLAR, a method that factorizes latent action representations into extent and mode components to improve robot policy learning efficiency and generalization.]]>
      </description>
      <content:encoded>
        <![CDATA[Introduces PoLAR, a method that factorizes latent action representations into extent and mode components to improve robot policy learning efficiency and generalization.]]>
      </content:encoded>
      <pubDate>Mon, 29 Jun 2026 05:14:05 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/e5e4186b/5587ac11.mp3" length="12477440" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>780</itunes:duration>
      <itunes:summary>Introduces PoLAR, a method that factorizes latent action representations into extent and mode components to improve robot policy learning efficiency and generalization.</itunes:summary>
      <itunes:subtitle>Introduces PoLAR, a method that factorizes latent action representations into extent and mode components to improve robot policy learning efficiency and generalization.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/e5e4186b/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Continual Robot Policy Learning via Variational Neural Dynamics</title>
      <itunes:title>Continual Robot Policy Learning via Variational Neural Dynamics</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">e4d91ef8-a3c9-456e-ac96-897cdce042a4</guid>
      <link>https://share.transistor.fm/s/d03ca487</link>
      <description>
        <![CDATA[Proposes a variational neural dynamics framework for continual robot policy learning, enabling robots to acquire new skills without forgetting previously learned ones.]]>
      </description>
      <content:encoded>
        <![CDATA[Proposes a variational neural dynamics framework for continual robot policy learning, enabling robots to acquire new skills without forgetting previously learned ones.]]>
      </content:encoded>
      <pubDate>Mon, 29 Jun 2026 05:13:48 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/d03ca487/8f6a712e.mp3" length="31230976" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1952</itunes:duration>
      <itunes:summary>Proposes a variational neural dynamics framework for continual robot policy learning, enabling robots to acquire new skills without forgetting previously learned ones.</itunes:summary>
      <itunes:subtitle>Proposes a variational neural dynamics framework for continual robot policy learning, enabling robots to acquire new skills without forgetting previously learned ones.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/d03ca487/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>PhysisForcing: Physics-Reinforced World Models for Robotic Manipulation</title>
      <itunes:title>PhysisForcing: Physics-Reinforced World Models for Robotic Manipulation</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">48ab3cfa-0147-4819-b921-0b805b5dcef4</guid>
      <link>https://share.transistor.fm/s/1afbeaeb</link>
      <description>
        <![CDATA[Plug-and-play training framework that enforces physical plausibility in robotic video generation models, achieving SOTA on R-Bench, PAI-Bench, and EZS-Bench. Lifts WorldArena success rate from 16% to 24% with zero extra inference cost.]]>
      </description>
      <content:encoded>
        <![CDATA[Plug-and-play training framework that enforces physical plausibility in robotic video generation models, achieving SOTA on R-Bench, PAI-Bench, and EZS-Bench. Lifts WorldArena success rate from 16% to 24% with zero extra inference cost.]]>
      </content:encoded>
      <pubDate>Mon, 29 Jun 2026 03:24:42 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/1afbeaeb/b9d476a7.mp3" length="22583808" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1412</itunes:duration>
      <itunes:summary>Plug-and-play training framework that enforces physical plausibility in robotic video generation models, achieving SOTA on R-Bench, PAI-Bench, and EZS-Bench. Lifts WorldArena success rate from 16% to 24% with zero extra inference cost.</itunes:summary>
      <itunes:subtitle>Plug-and-play training framework that enforces physical plausibility in robotic video generation models, achieving SOTA on R-Bench, PAI-Bench, and EZS-Bench. Lifts WorldArena success rate from 16% to 24% with zero extra inference cost.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/1afbeaeb/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Translation as a Bridging Action</title>
      <itunes:title>Translation as a Bridging Action</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">55e0488d-e2f7-427f-b598-721e8d5bc9f2</guid>
      <link>https://share.transistor.fm/s/898c61be</link>
      <description>
        <![CDATA[Replaces noisy 6DoF hand poses with relative wrist translation as a shared action space between cheap human videos and bimanual robots. Scales data-efficiently and outperforms full-pose baselines on manipulation tasks.]]>
      </description>
      <content:encoded>
        <![CDATA[Replaces noisy 6DoF hand poses with relative wrist translation as a shared action space between cheap human videos and bimanual robots. Scales data-efficiently and outperforms full-pose baselines on manipulation tasks.]]>
      </content:encoded>
      <pubDate>Mon, 29 Jun 2026 03:11:02 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/898c61be/c2efa127.mp3" length="35892736" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2244</itunes:duration>
      <itunes:summary>Replaces noisy 6DoF hand poses with relative wrist translation as a shared action space between cheap human videos and bimanual robots. Scales data-efficiently and outperforms full-pose baselines on manipulation tasks.</itunes:summary>
      <itunes:subtitle>Replaces noisy 6DoF hand poses with relative wrist translation as a shared action space between cheap human videos and bimanual robots. Scales data-efficiently and outperforms full-pose baselines on manipulation tasks.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/898c61be/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Play2Perfect: Dexterous Play Pretraining for Precise Assembly</title>
      <itunes:title>Play2Perfect: Dexterous Play Pretraining for Precise Assembly</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">99df9202-0721-4bc1-a53c-7750adabfdb6</guid>
      <link>https://share.transistor.fm/s/47531d33</link>
      <description>
        <![CDATA[Pre-trains a dexterous hand via unstructured 'play' interactions with objects, then fine-tunes for precise assembly tasks including 0.5 mm clearance insertions and furniture screwing, achieving 33x better sample efficiency than RL from scratch.]]>
      </description>
      <content:encoded>
        <![CDATA[Pre-trains a dexterous hand via unstructured 'play' interactions with objects, then fine-tunes for precise assembly tasks including 0.5 mm clearance insertions and furniture screwing, achieving 33x better sample efficiency than RL from scratch.]]>
      </content:encoded>
      <pubDate>Sun, 28 Jun 2026 14:13:36 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/47531d33/2b3a50fb.mp3" length="15252480" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>954</itunes:duration>
      <itunes:summary>Pre-trains a dexterous hand via unstructured 'play' interactions with objects, then fine-tunes for precise assembly tasks including 0.5 mm clearance insertions and furniture screwing, achieving 33x better sample efficiency than RL from scratch.</itunes:summary>
      <itunes:subtitle>Pre-trains a dexterous hand via unstructured 'play' interactions with objects, then fine-tunes for precise assembly tasks including 0.5 mm clearance insertions and furniture screwing, achieving 33x better sample efficiency than RL from scratch.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/47531d33/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Dexora: Open-Source VLA for High-DoF Bimanual Dexterity</title>
      <itunes:title>Dexora: Open-Source VLA for High-DoF Bimanual Dexterity</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">62c2af09-eb56-4676-8848-2d33f030ec2d</guid>
      <link>https://share.transistor.fm/s/a5947fd5</link>
      <description>
        <![CDATA[First open-source Vision-Language-Action (VLA) model for dual-arm, dual-hand 36-DoF dexterous manipulation, trained on 100K simulated and 10K real trajectories with strong cross-embodiment transfer capabilities.]]>
      </description>
      <content:encoded>
        <![CDATA[First open-source Vision-Language-Action (VLA) model for dual-arm, dual-hand 36-DoF dexterous manipulation, trained on 100K simulated and 10K real trajectories with strong cross-embodiment transfer capabilities.]]>
      </content:encoded>
      <pubDate>Sun, 28 Jun 2026 14:11:01 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/a5947fd5/4a2c475c.mp3" length="36816896" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2302</itunes:duration>
      <itunes:summary>First open-source Vision-Language-Action (VLA) model for dual-arm, dual-hand 36-DoF dexterous manipulation, trained on 100K simulated and 10K real trajectories with strong cross-embodiment transfer capabilities.</itunes:summary>
      <itunes:subtitle>First open-source Vision-Language-Action (VLA) model for dual-arm, dual-hand 36-DoF dexterous manipulation, trained on 100K simulated and 10K real trajectories with strong cross-embodiment transfer capabilities.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/a5947fd5/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>WorldVLA: Towards Autoregressive Action World Model</title>
      <itunes:title>WorldVLA: Towards Autoregressive Action World Model</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">6cec0457-d4a9-452a-870b-ff864390b4b8</guid>
      <link>https://share.transistor.fm/s/031e7377</link>
      <description>
        <![CDATA[Unifies VLA and world-modeling in a single autoregressive transformer that predicts both future images and actions. Outperforms separate VLA or world models on LIBERO simulation benchmarks.]]>
      </description>
      <content:encoded>
        <![CDATA[Unifies VLA and world-modeling in a single autoregressive transformer that predicts both future images and actions. Outperforms separate VLA or world models on LIBERO simulation benchmarks.]]>
      </content:encoded>
      <pubDate>Sun, 28 Jun 2026 05:09:08 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/031e7377/8174ac51.mp3" length="27449344" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1716</itunes:duration>
      <itunes:summary>Unifies VLA and world-modeling in a single autoregressive transformer that predicts both future images and actions. Outperforms separate VLA or world models on LIBERO simulation benchmarks.</itunes:summary>
      <itunes:subtitle>Unifies VLA and world-modeling in a single autoregressive transformer that predicts both future images and actions. Outperforms separate VLA or world models on LIBERO simulation benchmarks.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/031e7377/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>HumDex: Humanoid Dexterous Manipulation Made Easy</title>
      <itunes:title>HumDex: Humanoid Dexterous Manipulation Made Easy</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">2fdc0b74-6064-487c-8910-6e97b6f0e83f</guid>
      <link>https://share.transistor.fm/s/9f93029b</link>
      <description>
        <![CDATA[HumDex targets humanoid dexterous manipulation, aiming to simplify the development of dexterous manipulation capabilities for humanoid robots.]]>
      </description>
      <content:encoded>
        <![CDATA[HumDex targets humanoid dexterous manipulation, aiming to simplify the development of dexterous manipulation capabilities for humanoid robots.]]>
      </content:encoded>
      <pubDate>Sun, 28 Jun 2026 05:08:49 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/9f93029b/b416563a.mp3" length="31818240" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1989</itunes:duration>
      <itunes:summary>HumDex targets humanoid dexterous manipulation, aiming to simplify the development of dexterous manipulation capabilities for humanoid robots.</itunes:summary>
      <itunes:subtitle>HumDex targets humanoid dexterous manipulation, aiming to simplify the development of dexterous manipulation capabilities for humanoid robots.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/9f93029b/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>ForceBand: Learning Forceful Manipulation with sEMG</title>
      <itunes:title>ForceBand: Learning Forceful Manipulation with sEMG</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">bcb4c2f3-2fec-4f00-801b-407606e5842f</guid>
      <link>https://share.transistor.fm/s/0df2ebf5</link>
      <description>
        <![CDATA[Presents an open-source, low-cost sEMG wristband framework that extracts force signals from human muscle activity in videos, enabling zero-shot human-to-robot transfer of forceful manipulation policies across any robot, camera, or environment.]]>
      </description>
      <content:encoded>
        <![CDATA[Presents an open-source, low-cost sEMG wristband framework that extracts force signals from human muscle activity in videos, enabling zero-shot human-to-robot transfer of forceful manipulation policies across any robot, camera, or environment.]]>
      </content:encoded>
      <pubDate>Sat, 27 Jun 2026 14:11:53 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/0df2ebf5/f11ac099.mp3" length="26153472" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1635</itunes:duration>
      <itunes:summary>Presents an open-source, low-cost sEMG wristband framework that extracts force signals from human muscle activity in videos, enabling zero-shot human-to-robot transfer of forceful manipulation policies across any robot, camera, or environment.</itunes:summary>
      <itunes:subtitle>Presents an open-source, low-cost sEMG wristband framework that extracts force signals from human muscle activity in videos, enabling zero-shot human-to-robot transfer of forceful manipulation policies across any robot, camera, or environment.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/0df2ebf5/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>In-Context World Modeling for Robotic Control</title>
      <itunes:title>In-Context World Modeling for Robotic Control</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">2d0efa12-37a1-409d-bb8b-565eced42dea</guid>
      <link>https://share.transistor.fm/s/b20f27c8</link>
      <description>
        <![CDATA[Introduces ICWM, a method that learns world dynamics from just seconds of a robot's self-generated interaction data, enabling zero-shot adaptation to unseen cameras and new robot morphologies without any fine-tuning.]]>
      </description>
      <content:encoded>
        <![CDATA[Introduces ICWM, a method that learns world dynamics from just seconds of a robot's self-generated interaction data, enabling zero-shot adaptation to unseen cameras and new robot morphologies without any fine-tuning.]]>
      </content:encoded>
      <pubDate>Sat, 27 Jun 2026 14:08:16 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/b20f27c8/a0bbdb9a.mp3" length="24442368" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1528</itunes:duration>
      <itunes:summary>Introduces ICWM, a method that learns world dynamics from just seconds of a robot's self-generated interaction data, enabling zero-shot adaptation to unseen cameras and new robot morphologies without any fine-tuning.</itunes:summary>
      <itunes:subtitle>Introduces ICWM, a method that learns world dynamics from just seconds of a robot's self-generated interaction data, enabling zero-shot adaptation to unseen cameras and new robot morphologies without any fine-tuning.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/b20f27c8/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>WOLF-VLA: Vision-Language-Action for Humanoid Walking</title>
      <itunes:title>WOLF-VLA: Vision-Language-Action for Humanoid Walking</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">5e8cdb7b-e1d2-4d0e-ac81-062b6fa7afb0</guid>
      <link>https://share.transistor.fm/s/8e15e805</link>
      <description>
        <![CDATA[Introduces a framework integrating vision-language-action models for whole-body humanoid locomotion, addressing optimal control and learning for complex bipedal behaviors. Combines VLA learning with locomotion-specific control for humanoid robots.]]>
      </description>
      <content:encoded>
        <![CDATA[Introduces a framework integrating vision-language-action models for whole-body humanoid locomotion, addressing optimal control and learning for complex bipedal behaviors. Combines VLA learning with locomotion-specific control for humanoid robots.]]>
      </content:encoded>
      <pubDate>Sat, 27 Jun 2026 05:15:30 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/8e15e805/d270c33d.mp3" length="28252672" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1766</itunes:duration>
      <itunes:summary>Introduces a framework integrating vision-language-action models for whole-body humanoid locomotion, addressing optimal control and learning for complex bipedal behaviors. Combines VLA learning with locomotion-specific control for humanoid robots.</itunes:summary>
      <itunes:subtitle>Introduces a framework integrating vision-language-action models for whole-body humanoid locomotion, addressing optimal control and learning for complex bipedal behaviors. Combines VLA learning with locomotion-specific control for humanoid robots.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/8e15e805/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Motion-Focused Latent Action for Cross-Embodiment VLA from Human Videos</title>
      <itunes:title>Motion-Focused Latent Action for Cross-Embodiment VLA from Human Videos</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">02ee933f-a2f1-4c07-b1c8-664a2d508d19</guid>
      <link>https://share.transistor.fm/s/52e2b62b</link>
      <description>
        <![CDATA[Proposes a motion-focused latent action representation for cross-embodiment vision-language-action policies learned from human videos, accepted to IROS 2026.]]>
      </description>
      <content:encoded>
        <![CDATA[Proposes a motion-focused latent action representation for cross-embodiment vision-language-action policies learned from human videos, accepted to IROS 2026.]]>
      </content:encoded>
      <pubDate>Sat, 27 Jun 2026 05:14:57 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/52e2b62b/6e8a5e6e.mp3" length="29314048" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1833</itunes:duration>
      <itunes:summary>Proposes a motion-focused latent action representation for cross-embodiment vision-language-action policies learned from human videos, accepted to IROS 2026.</itunes:summary>
      <itunes:subtitle>Proposes a motion-focused latent action representation for cross-embodiment vision-language-action policies learned from human videos, accepted to IROS 2026.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/52e2b62b/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>ManiFlow: Manipulation via Rectified Flow</title>
      <itunes:title>ManiFlow: Manipulation via Rectified Flow</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">beb676be-9542-41ca-b7a8-4db07ab83b90</guid>
      <link>https://share.transistor.fm/s/7a9f409a</link>
      <description>
        <![CDATA[ManiFlow is a visuomotor imitation learning policy using consistency flow matching with a DiT-X architecture that generates high-quality actions in 1–2 steps. It works across single-arm, bimanual, and humanoid platforms using RGB or point cloud inputs.]]>
      </description>
      <content:encoded>
        <![CDATA[ManiFlow is a visuomotor imitation learning policy using consistency flow matching with a DiT-X architecture that generates high-quality actions in 1–2 steps. It works across single-arm, bimanual, and humanoid platforms using RGB or point cloud inputs.]]>
      </content:encoded>
      <pubDate>Sat, 27 Jun 2026 03:11:30 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/7a9f409a/c2179b17.mp3" length="28709888" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1795</itunes:duration>
      <itunes:summary>ManiFlow is a visuomotor imitation learning policy using consistency flow matching with a DiT-X architecture that generates high-quality actions in 1–2 steps. It works across single-arm, bimanual, and humanoid platforms using RGB or point cloud inputs.</itunes:summary>
      <itunes:subtitle>ManiFlow is a visuomotor imitation learning policy using consistency flow matching with a DiT-X architecture that generates high-quality actions in 1–2 steps. It works across single-arm, bimanual, and humanoid platforms using RGB or point cloud inputs.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/7a9f409a/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>RL-100: Toward Highly Reliable Real-World Robot Reinforcement Learning</title>
      <itunes:title>RL-100: Toward Highly Reliable Real-World Robot Reinforcement Learning</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">da92759d-8f1e-44d6-a50c-2b4a85a46d2d</guid>
      <link>https://share.transistor.fm/s/6acbca3d</link>
      <description>
        <![CDATA[RL-100 demonstrates highly reliable real-world RL manipulation achieving 900/900 success rates across 7 tasks with up to 250 consecutive trials without failure. It also shows strong robustness to disturbances and zero/few-shot adaptation capabilities.]]>
      </description>
      <content:encoded>
        <![CDATA[RL-100 demonstrates highly reliable real-world RL manipulation achieving 900/900 success rates across 7 tasks with up to 250 consecutive trials without failure. It also shows strong robustness to disturbances and zero/few-shot adaptation capabilities.]]>
      </content:encoded>
      <pubDate>Sat, 27 Jun 2026 03:10:48 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/6acbca3d/c11481fd.mp3" length="31160320" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1948</itunes:duration>
      <itunes:summary>RL-100 demonstrates highly reliable real-world RL manipulation achieving 900/900 success rates across 7 tasks with up to 250 consecutive trials without failure. It also shows strong robustness to disturbances and zero/few-shot adaptation capabilities.</itunes:summary>
      <itunes:subtitle>RL-100 demonstrates highly reliable real-world RL manipulation achieving 900/900 success rates across 7 tasks with up to 250 consecutive trials without failure. It also shows strong robustness to disturbances and zero/few-shot adaptation capabilities.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/6acbca3d/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation</title>
      <itunes:title>Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">cd261953-cfa2-4b98-9602-dd6ea24b3fb3</guid>
      <link>https://share.transistor.fm/s/bacccb3a</link>
      <description>
        <![CDATA[A video-diffusion world model trained on over 1 million manipulation episodes (3,000 hours) that includes an action model and neural simulator for closed-loop robotic manipulation control, with all code and models open-sourced.]]>
      </description>
      <content:encoded>
        <![CDATA[A video-diffusion world model trained on over 1 million manipulation episodes (3,000 hours) that includes an action model and neural simulator for closed-loop robotic manipulation control, with all code and models open-sourced.]]>
      </content:encoded>
      <pubDate>Fri, 26 Jun 2026 14:10:00 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/bacccb3a/a6c04c62.mp3" length="28729344" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1796</itunes:duration>
      <itunes:summary>A video-diffusion world model trained on over 1 million manipulation episodes (3,000 hours) that includes an action model and neural simulator for closed-loop robotic manipulation control, with all code and models open-sourced.</itunes:summary>
      <itunes:subtitle>A video-diffusion world model trained on over 1 million manipulation episodes (3,000 hours) that includes an action model and neural simulator for closed-loop robotic manipulation control, with all code and models open-sourced.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/bacccb3a/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Bi-HIL: Bilateral Control-Based Multimodal Hierarchical Imitation Learning for Long-Horizon Contact-Rich Manipulation</title>
      <itunes:title>Bi-HIL: Bilateral Control-Based Multimodal Hierarchical Imitation Learning for Long-Horizon Contact-Rich Manipulation</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">6adfec6c-b65b-48ba-8e2c-e3a93abb53ca</guid>
      <link>https://share.transistor.fm/s/6922448c</link>
      <description>
        <![CDATA[Proposes a hierarchical imitation learning framework using bilateral control, subtask-level progress tracking, and keyframe memory to handle long-horizon, contact-rich manipulation tasks.]]>
      </description>
      <content:encoded>
        <![CDATA[Proposes a hierarchical imitation learning framework using bilateral control, subtask-level progress tracking, and keyframe memory to handle long-horizon, contact-rich manipulation tasks.]]>
      </content:encoded>
      <pubDate>Fri, 26 Jun 2026 05:12:44 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/6922448c/b835f601.mp3" length="36422656" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2277</itunes:duration>
      <itunes:summary>Proposes a hierarchical imitation learning framework using bilateral control, subtask-level progress tracking, and keyframe memory to handle long-horizon, contact-rich manipulation tasks.</itunes:summary>
      <itunes:subtitle>Proposes a hierarchical imitation learning framework using bilateral control, subtask-level progress tracking, and keyframe memory to handle long-horizon, contact-rich manipulation tasks.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/6922448c/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>From Passive Observer to Active Critic: Reinforcement Learning Elicits Process Reasoning for Robotic Manipulation</title>
      <itunes:title>From Passive Observer to Active Critic: Reinforcement Learning Elicits Process Reasoning for Robotic Manipulation</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">20f67791-a8e7-441c-9907-f5f01a9d7e81</guid>
      <link>https://share.transistor.fm/s/96354d41</link>
      <description>
        <![CDATA[Uses reinforcement learning to improve process reasoning capabilities in robotic manipulation policies, shifting the model from passive observation to active critique.]]>
      </description>
      <content:encoded>
        <![CDATA[Uses reinforcement learning to improve process reasoning capabilities in robotic manipulation policies, shifting the model from passive observation to active critique.]]>
      </content:encoded>
      <pubDate>Fri, 26 Jun 2026 05:08:58 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/96354d41/34a9ee4f.mp3" length="12771840" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>799</itunes:duration>
      <itunes:summary>Uses reinforcement learning to improve process reasoning capabilities in robotic manipulation policies, shifting the model from passive observation to active critique.</itunes:summary>
      <itunes:subtitle>Uses reinforcement learning to improve process reasoning capabilities in robotic manipulation policies, shifting the model from passive observation to active critique.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/96354d41/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>ROVE: Unlocking Human Interventions for Humanoid Manipulation via Reinforcement Learning</title>
      <itunes:title>ROVE: Unlocking Human Interventions for Humanoid Manipulation via Reinforcement Learning</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">18b7ea71-6ed5-44e5-9d1f-4a4b3b089294</guid>
      <link>https://share.transistor.fm/s/bbdd0fad</link>
      <description>
        <![CDATA[ROVE leverages reinforcement learning to enable humanoid robots to benefit from human interventions during manipulation tasks.]]>
      </description>
      <content:encoded>
        <![CDATA[ROVE leverages reinforcement learning to enable humanoid robots to benefit from human interventions during manipulation tasks.]]>
      </content:encoded>
      <pubDate>Fri, 26 Jun 2026 03:12:15 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/bbdd0fad/aefefb86.mp3" length="28906496" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1807</itunes:duration>
      <itunes:summary>ROVE leverages reinforcement learning to enable humanoid robots to benefit from human interventions during manipulation tasks.</itunes:summary>
      <itunes:subtitle>ROVE leverages reinforcement learning to enable humanoid robots to benefit from human interventions during manipulation tasks.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/bbdd0fad/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>ConstrainedMimic: Safe Humanoid Robot Motion Tracking</title>
      <itunes:title>ConstrainedMimic: Safe Humanoid Robot Motion Tracking</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">dadf0604-d6c5-4d3a-9684-d44e759a75c7</guid>
      <link>https://share.transistor.fm/s/ffe6d039</link>
      <description>
        <![CDATA[A control framework for safe humanoid robot motion tracking using RL policies with real-time constraint enforcement via kinematics, dynamics, and control barrier functions.]]>
      </description>
      <content:encoded>
        <![CDATA[A control framework for safe humanoid robot motion tracking using RL policies with real-time constraint enforcement via kinematics, dynamics, and control barrier functions.]]>
      </content:encoded>
      <pubDate>Fri, 26 Jun 2026 03:11:55 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/ffe6d039/696e9573.mp3" length="42646016" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2666</itunes:duration>
      <itunes:summary>A control framework for safe humanoid robot motion tracking using RL policies with real-time constraint enforcement via kinematics, dynamics, and control barrier functions.</itunes:summary>
      <itunes:subtitle>A control framework for safe humanoid robot motion tracking using RL policies with real-time constraint enforcement via kinematics, dynamics, and control barrier functions.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/ffe6d039/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>REAL: Robust Extreme Agility via Spatio-Temporal Policy Learning and Physics-Guided Filtering</title>
      <itunes:title>REAL: Robust Extreme Agility via Spatio-Temporal Policy Learning and Physics-Guided Filtering</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">5f13589b-5db0-4abe-b686-3edf37a41921</guid>
      <link>https://share.transistor.fm/s/5299766b</link>
      <description>
        <![CDATA[Introduces spatio-temporal policy learning combined with physics-guided filtering to achieve robust and extremely agile robot control.]]>
      </description>
      <content:encoded>
        <![CDATA[Introduces spatio-temporal policy learning combined with physics-guided filtering to achieve robust and extremely agile robot control.]]>
      </content:encoded>
      <pubDate>Thu, 25 Jun 2026 14:13:34 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/5299766b/afbfda04.mp3" length="22033920" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1378</itunes:duration>
      <itunes:summary>Introduces spatio-temporal policy learning combined with physics-guided filtering to achieve robust and extremely agile robot control.</itunes:summary>
      <itunes:subtitle>Introduces spatio-temporal policy learning combined with physics-guided filtering to achieve robust and extremely agile robot control.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/5299766b/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>HiFlow: Tokenization-Free Scale-Wise Autoregressive Policy Learning via Flow Matching</title>
      <itunes:title>HiFlow: Tokenization-Free Scale-Wise Autoregressive Policy Learning via Flow Matching</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">46adc8c7-8942-4bbf-8e1e-1eb4ffd9e6d4</guid>
      <link>https://share.transistor.fm/s/67488eed</link>
      <description>
        <![CDATA[Introduces a tokenization-free autoregressive policy learning framework using flow matching across scales for robotic control.]]>
      </description>
      <content:encoded>
        <![CDATA[Introduces a tokenization-free autoregressive policy learning framework using flow matching across scales for robotic control.]]>
      </content:encoded>
      <pubDate>Thu, 25 Jun 2026 14:08:48 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/67488eed/0b4276cc.mp3" length="33197056" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2075</itunes:duration>
      <itunes:summary>Introduces a tokenization-free autoregressive policy learning framework using flow matching across scales for robotic control.</itunes:summary>
      <itunes:subtitle>Introduces a tokenization-free autoregressive policy learning framework using flow matching across scales for robotic control.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/67488eed/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Reactive Diffusion Policy: Slow-Fast Visual-Tactile Learning for Contact-Rich Manipulation</title>
      <itunes:title>Reactive Diffusion Policy: Slow-Fast Visual-Tactile Learning for Contact-Rich Manipulation</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">438dcfcd-f661-4293-9ef0-bc1e3db73388</guid>
      <link>https://share.transistor.fm/s/92aefc04</link>
      <description>
        <![CDATA[Introduces a slow-fast imitation learning framework combining diffusion-based planning with reactive tactile/force feedback for contact-rich manipulation tasks. Also includes TactAR, an AR-based teleoperation system with tactile sensing.]]>
      </description>
      <content:encoded>
        <![CDATA[Introduces a slow-fast imitation learning framework combining diffusion-based planning with reactive tactile/force feedback for contact-rich manipulation tasks. Also includes TactAR, an AR-based teleoperation system with tactile sensing.]]>
      </content:encoded>
      <pubDate>Thu, 25 Jun 2026 05:15:32 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/92aefc04/e1d769e3.mp3" length="44599296" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2788</itunes:duration>
      <itunes:summary>Introduces a slow-fast imitation learning framework combining diffusion-based planning with reactive tactile/force feedback for contact-rich manipulation tasks. Also includes TactAR, an AR-based teleoperation system with tactile sensing.</itunes:summary>
      <itunes:subtitle>Introduces a slow-fast imitation learning framework combining diffusion-based planning with reactive tactile/force feedback for contact-rich manipulation tasks. Also includes TactAR, an AR-based teleoperation system with tactile sensing.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/92aefc04/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>SARM2 + SPIRAL: Multi-Task Reward Models and RL Refinement for Long-Horizon Dexterous Manipulation</title>
      <itunes:title>SARM2 + SPIRAL: Multi-Task Reward Models and RL Refinement for Long-Horizon Dexterous Manipulation</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">5826e914-2366-4413-ad99-f5302e99c3ce</guid>
      <link>https://share.transistor.fm/s/e30dbce7</link>
      <description>
        <![CDATA[Combines scalable autonomous reward modeling with RL-based refinement to improve vision-language-action policies on long-horizon dexterous manipulation tasks via autonomous rollouts. Demonstrates significant gains over imitation learning baselines.]]>
      </description>
      <content:encoded>
        <![CDATA[Combines scalable autonomous reward modeling with RL-based refinement to improve vision-language-action policies on long-horizon dexterous manipulation tasks via autonomous rollouts. Demonstrates significant gains over imitation learning baselines.]]>
      </content:encoded>
      <pubDate>Thu, 25 Jun 2026 05:12:28 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/e30dbce7/47b75974.mp3" length="39083008" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2443</itunes:duration>
      <itunes:summary>Combines scalable autonomous reward modeling with RL-based refinement to improve vision-language-action policies on long-horizon dexterous manipulation tasks via autonomous rollouts. Demonstrates significant gains over imitation learning baselines.</itunes:summary>
      <itunes:subtitle>Combines scalable autonomous reward modeling with RL-based refinement to improve vision-language-action policies on long-horizon dexterous manipulation tasks via autonomous rollouts. Demonstrates significant gains over imitation learning baselines.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/e30dbce7/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Co-VLA: Coordination-Aware Structured Action Modeling for Dual-Arm VLA Systems</title>
      <itunes:title>Co-VLA: Coordination-Aware Structured Action Modeling for Dual-Arm VLA Systems</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">5778c298-3292-44c9-bcd5-6e2d3fad54d3</guid>
      <link>https://share.transistor.fm/s/26225352</link>
      <description>
        <![CDATA[Introduces coordination-aware structured action modeling for dual-arm robotic systems within a VLA framework. Addresses the unique challenges of bimanual manipulation through specialized action representations.]]>
      </description>
      <content:encoded>
        <![CDATA[Introduces coordination-aware structured action modeling for dual-arm robotic systems within a VLA framework. Addresses the unique challenges of bimanual manipulation through specialized action representations.]]>
      </content:encoded>
      <pubDate>Thu, 25 Jun 2026 03:34:15 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/26225352/1ef7dd7b.mp3" length="14305280" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>895</itunes:duration>
      <itunes:summary>Introduces coordination-aware structured action modeling for dual-arm robotic systems within a VLA framework. Addresses the unique challenges of bimanual manipulation through specialized action representations.</itunes:summary>
      <itunes:subtitle>Introduces coordination-aware structured action modeling for dual-arm robotic systems within a VLA framework. Addresses the unique challenges of bimanual manipulation through specialized action representations.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/26225352/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>ThinkingVLA: Interleaved Vision and Language Reasoning for Robotic Manipulation</title>
      <itunes:title>ThinkingVLA: Interleaved Vision and Language Reasoning for Robotic Manipulation</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">d94a1ec2-c3df-4759-8cfd-893932bd277e</guid>
      <link>https://share.transistor.fm/s/4cf91ddc</link>
      <description>
        <![CDATA[Proposes interleaved vision and language reasoning for robotic manipulation within a VLA framework. Aims to improve instruction following and task performance through integrated multimodal reasoning.]]>
      </description>
      <content:encoded>
        <![CDATA[Proposes interleaved vision and language reasoning for robotic manipulation within a VLA framework. Aims to improve instruction following and task performance through integrated multimodal reasoning.]]>
      </content:encoded>
      <pubDate>Thu, 25 Jun 2026 03:31:11 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/4cf91ddc/4a864d2d.mp3" length="10045952" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>626</itunes:duration>
      <itunes:summary>Proposes interleaved vision and language reasoning for robotic manipulation within a VLA framework. Aims to improve instruction following and task performance through integrated multimodal reasoning.</itunes:summary>
      <itunes:subtitle>Proposes interleaved vision and language reasoning for robotic manipulation within a VLA framework. Aims to improve instruction following and task performance through integrated multimodal reasoning.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/4cf91ddc/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Playful Agentic Robot Learning</title>
      <itunes:title>Playful Agentic Robot Learning</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">3680576f-6dee-4329-acda-62d1a82603ce</guid>
      <link>https://share.transistor.fm/s/03b62f1e</link>
      <description>
        <![CDATA[Self-directed play combined with Code-as-Policy for reusable skill acquisition and downstream manipulation tasks.]]>
      </description>
      <content:encoded>
        <![CDATA[Self-directed play combined with Code-as-Policy for reusable skill acquisition and downstream manipulation tasks.]]>
      </content:encoded>
      <pubDate>Thu, 25 Jun 2026 03:22:48 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/03b62f1e/3e3f9cda.mp3" length="28273152" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1768</itunes:duration>
      <itunes:summary>Self-directed play combined with Code-as-Policy for reusable skill acquisition and downstream manipulation tasks.</itunes:summary>
      <itunes:subtitle>Self-directed play combined with Code-as-Policy for reusable skill acquisition and downstream manipulation tasks.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/03b62f1e/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Learning Unified Force and Position Control for Legged Loco-Manipulation</title>
      <itunes:title>Learning Unified Force and Position Control for Legged Loco-Manipulation</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">b6840ed9-df56-415a-9b40-757b67fba7ab</guid>
      <link>https://share.transistor.fm/s/b422f40e</link>
      <description>
        <![CDATA[A unified RL policy for quadrupeds and humanoids that jointly handles force and position control without force sensors, enabling compliant behaviors, force-aware imitation learning, and contact-rich tasks.]]>
      </description>
      <content:encoded>
        <![CDATA[A unified RL policy for quadrupeds and humanoids that jointly handles force and position control without force sensors, enabling compliant behaviors, force-aware imitation learning, and contact-rich tasks.]]>
      </content:encoded>
      <pubDate>Wed, 24 Jun 2026 14:14:01 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/b422f40e/97a5f6b3.mp3" length="38312448" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2395</itunes:duration>
      <itunes:summary>A unified RL policy for quadrupeds and humanoids that jointly handles force and position control without force sensors, enabling compliant behaviors, force-aware imitation learning, and contact-rich tasks.</itunes:summary>
      <itunes:subtitle>A unified RL policy for quadrupeds and humanoids that jointly handles force and position control without force sensors, enabling compliant behaviors, force-aware imitation learning, and contact-rich tasks.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/b422f40e/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Robots that Collaborate: Sequential Asymmetric Imitation for Learning Coupled Robot Policies</title>
      <itunes:title>Robots that Collaborate: Sequential Asymmetric Imitation for Learning Coupled Robot Policies</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">69a4a62c-841e-4179-965d-ab0451229090</guid>
      <link>https://share.transistor.fm/s/3c61be99</link>
      <description>
        <![CDATA[Explores imitation learning approaches for multi-robot systems, focusing on policy coupling through sequential asymmetric imitation to enable collaborative robot behaviors.]]>
      </description>
      <content:encoded>
        <![CDATA[Explores imitation learning approaches for multi-robot systems, focusing on policy coupling through sequential asymmetric imitation to enable collaborative robot behaviors.]]>
      </content:encoded>
      <pubDate>Wed, 24 Jun 2026 14:11:02 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/3c61be99/a34053d3.mp3" length="25604096" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1601</itunes:duration>
      <itunes:summary>Explores imitation learning approaches for multi-robot systems, focusing on policy coupling through sequential asymmetric imitation to enable collaborative robot behaviors.</itunes:summary>
      <itunes:subtitle>Explores imitation learning approaches for multi-robot systems, focusing on policy coupling through sequential asymmetric imitation to enable collaborative robot behaviors.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/3c61be99/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>AstraBrain-WBC 0.5: A Humanoid Robot Cerebellum Foundation Model</title>
      <itunes:title>AstraBrain-WBC 0.5: A Humanoid Robot Cerebellum Foundation Model</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">b1c82e49-c030-44a8-9733-389465fc1a50</guid>
      <link>https://share.transistor.fm/s/5a6a1f10</link>
      <description>
        <![CDATA[A humanoid robot 'cerebellum' foundation model trained on 20,000 hours of human motion data that demonstrates scaling laws for robot motion control and enables zero-shot execution of unseen motions on real humanoids.]]>
      </description>
      <content:encoded>
        <![CDATA[A humanoid robot 'cerebellum' foundation model trained on 20,000 hours of human motion data that demonstrates scaling laws for robot motion control and enables zero-shot execution of unseen motions on real humanoids.]]>
      </content:encoded>
      <pubDate>Wed, 24 Jun 2026 05:17:07 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/5a6a1f10/43cc1ec5.mp3" length="13739520" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>859</itunes:duration>
      <itunes:summary>A humanoid robot 'cerebellum' foundation model trained on 20,000 hours of human motion data that demonstrates scaling laws for robot motion control and enables zero-shot execution of unseen motions on real humanoids.</itunes:summary>
      <itunes:subtitle>A humanoid robot 'cerebellum' foundation model trained on 20,000 hours of human motion data that demonstrates scaling laws for robot motion control and enables zero-shot execution of unseen motions on real humanoids.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/5a6a1f10/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>SRL: Combining SLIP Model and Reinforcement Learning for Agile Robotic Jumping</title>
      <itunes:title>SRL: Combining SLIP Model and Reinforcement Learning for Agile Robotic Jumping</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">ba23d358-ff1b-49cd-8e77-eb390f552353</guid>
      <link>https://share.transistor.fm/s/d83aad8e</link>
      <description>
        <![CDATA[Combines the Spring-Loaded Inverted Pendulum (SLIP) model with reinforcement learning to achieve agile jumping behaviors in robotic systems.]]>
      </description>
      <content:encoded>
        <![CDATA[Combines the Spring-Loaded Inverted Pendulum (SLIP) model with reinforcement learning to achieve agile jumping behaviors in robotic systems.]]>
      </content:encoded>
      <pubDate>Wed, 24 Jun 2026 05:14:53 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/d83aad8e/ca6b5b47.mp3" length="22943232" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1434</itunes:duration>
      <itunes:summary>Combines the Spring-Loaded Inverted Pendulum (SLIP) model with reinforcement learning to achieve agile jumping behaviors in robotic systems.</itunes:summary>
      <itunes:subtitle>Combines the Spring-Loaded Inverted Pendulum (SLIP) model with reinforcement learning to achieve agile jumping behaviors in robotic systems.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/d83aad8e/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>DataClaw0: Agentic Tailoring for Raw Multimodal Streams</title>
      <itunes:title>DataClaw0: Agentic Tailoring for Raw Multimodal Streams</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">7b1ea0d1-8558-48fe-92ae-143f6943c950</guid>
      <link>https://share.transistor.fm/s/9f0031a7</link>
      <description>
        <![CDATA[A 9B model that filters noise from videos, GUI, and embodied data streams, reorganizing them into dense supervision via factual anchors and semantic synthesis; trained with SFT + GRPO across five domains with benchmarks.]]>
      </description>
      <content:encoded>
        <![CDATA[A 9B model that filters noise from videos, GUI, and embodied data streams, reorganizing them into dense supervision via factual anchors and semantic synthesis; trained with SFT + GRPO across five domains with benchmarks.]]>
      </content:encoded>
      <pubDate>Wed, 24 Jun 2026 03:20:11 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/9f0031a7/1a3cdc71.mp3" length="34840064" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2178</itunes:duration>
      <itunes:summary>A 9B model that filters noise from videos, GUI, and embodied data streams, reorganizing them into dense supervision via factual anchors and semantic synthesis; trained with SFT + GRPO across five domains with benchmarks.</itunes:summary>
      <itunes:subtitle>A 9B model that filters noise from videos, GUI, and embodied data streams, reorganizing them into dense supervision via factual anchors and semantic synthesis; trained with SFT + GRPO across five domains with benchmarks.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/9f0031a7/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>ACE-Ego-0: Unifying Egocentric Human and Robotic Data for VLA Pretraining</title>
      <itunes:title>ACE-Ego-0: Unifying Egocentric Human and Robotic Data for VLA Pretraining</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">64a67b49-2ecf-4f5b-a754-efe26961375e</guid>
      <link>https://share.transistor.fm/s/04faec51</link>
      <description>
        <![CDATA[Converts 6K+ hours of mixed human/robot egocentric video into robot pseudo-actions via camera-space alignment and reliability-aware loss, achieving 72.8% on RoboCasa and 91.1% on RoboTwin.]]>
      </description>
      <content:encoded>
        <![CDATA[Converts 6K+ hours of mixed human/robot egocentric video into robot pseudo-actions via camera-space alignment and reliability-aware loss, achieving 72.8% on RoboCasa and 91.1% on RoboTwin.]]>
      </content:encoded>
      <pubDate>Wed, 24 Jun 2026 03:16:42 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/04faec51/f197da27.mp3" length="30590976" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1912</itunes:duration>
      <itunes:summary>Converts 6K+ hours of mixed human/robot egocentric video into robot pseudo-actions via camera-space alignment and reliability-aware loss, achieving 72.8% on RoboCasa and 91.1% on RoboTwin.</itunes:summary>
      <itunes:subtitle>Converts 6K+ hours of mixed human/robot egocentric video into robot pseudo-actions via camera-space alignment and reliability-aware loss, achieving 72.8% on RoboCasa and 91.1% on RoboTwin.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/04faec51/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>VERA: Video-to-Action World Model Policy</title>
      <itunes:title>VERA: Video-to-Action World Model Policy</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">638cf5a6-3c30-490f-ad84-a80b2b3b5efa</guid>
      <link>https://share.transistor.fm/s/9724bac0</link>
      <description>
        <![CDATA[A 14B-parameter video world model that converts predicted visual futures into embodiment-agnostic actions via Jacobian inverse-dynamics, enabling zero-shot cross-robot transfer across a Panda arm and 16-DoF hand with open-sourced weights and training code.]]>
      </description>
      <content:encoded>
        <![CDATA[A 14B-parameter video world model that converts predicted visual futures into embodiment-agnostic actions via Jacobian inverse-dynamics, enabling zero-shot cross-robot transfer across a Panda arm and 16-DoF hand with open-sourced weights and training code.]]>
      </content:encoded>
      <pubDate>Wed, 24 Jun 2026 00:03:04 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/9724bac0/b322790a.mp3" length="28423168" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1777</itunes:duration>
      <itunes:summary>A 14B-parameter video world model that converts predicted visual futures into embodiment-agnostic actions via Jacobian inverse-dynamics, enabling zero-shot cross-robot transfer across a Panda arm and 16-DoF hand with open-sourced weights and training code.</itunes:summary>
      <itunes:subtitle>A 14B-parameter video world model that converts predicted visual futures into embodiment-agnostic actions via Jacobian inverse-dynamics, enabling zero-shot cross-robot transfer across a Panda arm and 16-DoF hand with open-sourced weights and training code</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/9724bac0/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>GEN-1: Scaled Dexterous Manipulation Foundation Model</title>
      <itunes:title>GEN-1: Scaled Dexterous Manipulation Foundation Model</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">bebd8a68-79c4-4448-a64f-43af253e9472</guid>
      <link>https://share.transistor.fm/s/51cdd76e</link>
      <description>
        <![CDATA[A dexterous manipulation foundation model trained on 500k hours of real-world bimanual data that handles deformable objects such as cardboard folding and screw packing, featuring online retry and adaptation capabilities.]]>
      </description>
      <content:encoded>
        <![CDATA[A dexterous manipulation foundation model trained on 500k hours of real-world bimanual data that handles deformable objects such as cardboard folding and screw packing, featuring online retry and adaptation capabilities.]]>
      </content:encoded>
      <pubDate>Wed, 24 Jun 2026 00:01:21 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/51cdd76e/7569e8fd.mp3" length="44504064" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2782</itunes:duration>
      <itunes:summary>A dexterous manipulation foundation model trained on 500k hours of real-world bimanual data that handles deformable objects such as cardboard folding and screw packing, featuring online retry and adaptation capabilities.</itunes:summary>
      <itunes:subtitle>A dexterous manipulation foundation model trained on 500k hours of real-world bimanual data that handles deformable objects such as cardboard folding and screw packing, featuring online retry and adaptation capabilities.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/51cdd76e/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Efficient Hybrid SE(3)-Equivariant Visuomotor Flow Policy via Spherical Harmonics for Robot Manipulation</title>
      <itunes:title>Efficient Hybrid SE(3)-Equivariant Visuomotor Flow Policy via Spherical Harmonics for Robot Manipulation</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">8811f88f-d483-4023-bfee-acd60b5ff5a8</guid>
      <link>https://share.transistor.fm/s/992c5fe5</link>
      <description>
        <![CDATA[Develops an SE(3)-equivariant flow-based visuomotor policy leveraging spherical harmonics for efficient and geometrically consistent robot manipulation.]]>
      </description>
      <content:encoded>
        <![CDATA[Develops an SE(3)-equivariant flow-based visuomotor policy leveraging spherical harmonics for efficient and geometrically consistent robot manipulation.]]>
      </content:encoded>
      <pubDate>Mon, 22 Jun 2026 05:16:39 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/992c5fe5/ac843f65.mp3" length="22514176" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1408</itunes:duration>
      <itunes:summary>Develops an SE(3)-equivariant flow-based visuomotor policy leveraging spherical harmonics for efficient and geometrically consistent robot manipulation.</itunes:summary>
      <itunes:subtitle>Develops an SE(3)-equivariant flow-based visuomotor policy leveraging spherical harmonics for efficient and geometrically consistent robot manipulation.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/992c5fe5/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Cortical Policy: A Dual-Stream View Transformer for Robotic Manipulation</title>
      <itunes:title>Cortical Policy: A Dual-Stream View Transformer for Robotic Manipulation</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">1a0b6736-f1d3-415a-893c-85b1a611e105</guid>
      <link>https://share.transistor.fm/s/936bc44f</link>
      <description>
        <![CDATA[Introduces a dual-stream transformer architecture inspired by cortical visual processing for learning robotic manipulation policies.]]>
      </description>
      <content:encoded>
        <![CDATA[Introduces a dual-stream transformer architecture inspired by cortical visual processing for learning robotic manipulation policies.]]>
      </content:encoded>
      <pubDate>Mon, 22 Jun 2026 05:10:43 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/936bc44f/6f706423.mp3" length="26071552" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1630</itunes:duration>
      <itunes:summary>Introduces a dual-stream transformer architecture inspired by cortical visual processing for learning robotic manipulation policies.</itunes:summary>
      <itunes:subtitle>Introduces a dual-stream transformer architecture inspired by cortical visual processing for learning robotic manipulation policies.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/936bc44f/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>VisualClaw: A Self-Evolving Wearable Vision Agent</title>
      <itunes:title>VisualClaw: A Self-Evolving Wearable Vision Agent</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">c1b40d17-041a-4be6-9269-01c5695574a6</guid>
      <link>https://share.transistor.fm/s/d5310e92</link>
      <description>
        <![CDATA[An edge-filtered video streaming agent that evolves skills from memory and runs on smart glasses, reducing API costs by 98%, accompanied by the VisualClawArena benchmark dataset.]]>
      </description>
      <content:encoded>
        <![CDATA[An edge-filtered video streaming agent that evolves skills from memory and runs on smart glasses, reducing API costs by 98%, accompanied by the VisualClawArena benchmark dataset.]]>
      </content:encoded>
      <pubDate>Mon, 22 Jun 2026 03:05:59 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/d5310e92/0b239612.mp3" length="28624384" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1789</itunes:duration>
      <itunes:summary>An edge-filtered video streaming agent that evolves skills from memory and runs on smart glasses, reducing API costs by 98%, accompanied by the VisualClawArena benchmark dataset.</itunes:summary>
      <itunes:subtitle>An edge-filtered video streaming agent that evolves skills from memory and runs on smart glasses, reducing API costs by 98%, accompanied by the VisualClawArena benchmark dataset.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/d5310e92/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Kairos: A Native World Model Stack for Physical AI</title>
      <itunes:title>Kairos: A Native World Model Stack for Physical AI</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">8df9f119-773f-4f59-a270-61d82e55f50a</guid>
      <link>https://share.transistor.fm/s/5f44a660</link>
      <description>
        <![CDATA[A 4B unified architecture for world understanding, generation, and action with hybrid linear attention enabling real-time edge inference across embodiments, outperforming 14B models on embodied benchmarks.]]>
      </description>
      <content:encoded>
        <![CDATA[A 4B unified architecture for world understanding, generation, and action with hybrid linear attention enabling real-time edge inference across embodiments, outperforming 14B models on embodied benchmarks.]]>
      </content:encoded>
      <pubDate>Sun, 21 Jun 2026 14:23:15 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/5f44a660/b8529bc3.mp3" length="32311296" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2020</itunes:duration>
      <itunes:summary>A 4B unified architecture for world understanding, generation, and action with hybrid linear attention enabling real-time edge inference across embodiments, outperforming 14B models on embodied benchmarks.</itunes:summary>
      <itunes:subtitle>A 4B unified architecture for world understanding, generation, and action with hybrid linear attention enabling real-time edge inference across embodiments, outperforming 14B models on embodied benchmarks.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/5f44a660/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>DragMesh-2: A Contact-Driven Framework for Dexterous Hand–Object Interaction</title>
      <itunes:title>DragMesh-2: A Contact-Driven Framework for Dexterous Hand–Object Interaction</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">89823942-8812-4e00-ae12-deb606a39a7b</guid>
      <link>https://share.transistor.fm/s/acf8a411</link>
      <description>
        <![CDATA[A framework that trains a 51-DoF dexterous hand to open drawers and doors using only physical contact without requiring tactile sensors.]]>
      </description>
      <content:encoded>
        <![CDATA[A framework that trains a 51-DoF dexterous hand to open drawers and doors using only physical contact without requiring tactile sensors.]]>
      </content:encoded>
      <pubDate>Sun, 21 Jun 2026 14:12:42 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/acf8a411/bc132dbf.mp3" length="12587008" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>787</itunes:duration>
      <itunes:summary>A framework that trains a 51-DoF dexterous hand to open drawers and doors using only physical contact without requiring tactile sensors.</itunes:summary>
      <itunes:subtitle>A framework that trains a 51-DoF dexterous hand to open drawers and doors using only physical contact without requiring tactile sensors.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/acf8a411/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Guava: A Universal Harness for Robot Manipulation</title>
      <itunes:title>Guava: A Universal Harness for Robot Manipulation</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">f74fca31-0cd2-4509-96ea-746fcd18abf8</guid>
      <link>https://share.transistor.fm/s/0c83a1a6</link>
      <description>
        <![CDATA[A 4B open-source VLA-style model trained on fewer than 2K simulation trajectories that matches closed frontier systems on real-world manipulation tasks with zero-shot generalization to novel objects and long-horizon behaviors, including failure-recovery demonstrations.]]>
      </description>
      <content:encoded>
        <![CDATA[A 4B open-source VLA-style model trained on fewer than 2K simulation trajectories that matches closed frontier systems on real-world manipulation tasks with zero-shot generalization to novel objects and long-horizon behaviors, including failure-recovery demonstrations.]]>
      </content:encoded>
      <pubDate>Sun, 21 Jun 2026 05:10:01 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/0c83a1a6/c5e03836.mp3" length="27730432" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1734</itunes:duration>
      <itunes:summary>A 4B open-source VLA-style model trained on fewer than 2K simulation trajectories that matches closed frontier systems on real-world manipulation tasks with zero-shot generalization to novel objects and long-horizon behaviors, including failure-recovery demonstrations.</itunes:summary>
      <itunes:subtitle>A 4B open-source VLA-style model trained on fewer than 2K simulation trajectories that matches closed frontier systems on real-world manipulation tasks with zero-shot generalization to novel objects and long-horizon behaviors, including failure-recovery d</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/0c83a1a6/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Geometric Action Model for Robot Policies</title>
      <itunes:title>Geometric Action Model for Robot Policies</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">b70be4a3-d8e0-4ad8-8845-022249ddbdcc</guid>
      <link>https://share.transistor.fm/s/5ea7bef2</link>
      <description>
        <![CDATA[A new geometric action model for robot manipulation policies that focuses on structured action representations to improve policy learning and generalization.]]>
      </description>
      <content:encoded>
        <![CDATA[A new geometric action model for robot manipulation policies that focuses on structured action representations to improve policy learning and generalization.]]>
      </content:encoded>
      <pubDate>Sun, 21 Jun 2026 05:09:28 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/5ea7bef2/fb3f5b91.mp3" length="34425856" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2152</itunes:duration>
      <itunes:summary>A new geometric action model for robot manipulation policies that focuses on structured action representations to improve policy learning and generalization.</itunes:summary>
      <itunes:subtitle>A new geometric action model for robot manipulation policies that focuses on structured action representations to improve policy learning and generalization.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/5ea7bef2/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>ENPIRE: Physical AutoResearch with a Fleet of 8 Robots</title>
      <itunes:title>ENPIRE: Physical AutoResearch with a Fleet of 8 Robots</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">3c44d3be-8cf4-4269-b73d-6210494437bb</guid>
      <link>https://share.transistor.fm/s/1a1326ba</link>
      <description>
        <![CDATA[ENPIRE demonstrates fully autonomous physical AutoResearch where Codex agents control a fleet of 8 robots overnight, self-improving through real hardware rollouts on tasks like zip-tie tying and GPU installation, while discovering physical scaling laws with built-in safety harnesses and frozen reward classifiers derived from demonstrations.]]>
      </description>
      <content:encoded>
        <![CDATA[ENPIRE demonstrates fully autonomous physical AutoResearch where Codex agents control a fleet of 8 robots overnight, self-improving through real hardware rollouts on tasks like zip-tie tying and GPU installation, while discovering physical scaling laws with built-in safety harnesses and frozen reward classifiers derived from demonstrations.]]>
      </content:encoded>
      <pubDate>Sun, 21 Jun 2026 03:14:25 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/1a1326ba/d4ffa3fc.mp3" length="27455488" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1716</itunes:duration>
      <itunes:summary>ENPIRE demonstrates fully autonomous physical AutoResearch where Codex agents control a fleet of 8 robots overnight, self-improving through real hardware rollouts on tasks like zip-tie tying and GPU installation, while discovering physical scaling laws with built-in safety harnesses and frozen reward classifiers derived from demonstrations.</itunes:summary>
      <itunes:subtitle>ENPIRE demonstrates fully autonomous physical AutoResearch where Codex agents control a fleet of 8 robots overnight, self-improving through real hardware rollouts on tasks like zip-tie tying and GPU installation, while discovering physical scaling laws wi</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/1a1326ba/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>MolmoAct2: An Open Foundation Model for Real-World Robotics</title>
      <itunes:title>MolmoAct2: An Open Foundation Model for Real-World Robotics</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">f69ed734-fa1d-4c18-bd51-65892005d35a</guid>
      <link>https://share.transistor.fm/s/f6984106</link>
      <description>
        <![CDATA[An open VLA-style robotics foundation model featuring open weights, open dataset, open action tokenizer, and a depth-reasoning variant; designed to enable community experiments on real robots for manipulation and generalist policies.]]>
      </description>
      <content:encoded>
        <![CDATA[An open VLA-style robotics foundation model featuring open weights, open dataset, open action tokenizer, and a depth-reasoning variant; designed to enable community experiments on real robots for manipulation and generalist policies.]]>
      </content:encoded>
      <pubDate>Sun, 21 Jun 2026 00:43:08 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/f6984106/8c80b9f8.mp3" length="31213056" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1951</itunes:duration>
      <itunes:summary>An open VLA-style robotics foundation model featuring open weights, open dataset, open action tokenizer, and a depth-reasoning variant; designed to enable community experiments on real robots for manipulation and generalist policies.</itunes:summary>
      <itunes:subtitle>An open VLA-style robotics foundation model featuring open weights, open dataset, open action tokenizer, and a depth-reasoning variant; designed to enable community experiments on real robots for manipulation and generalist policies.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/f6984106/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Hy-Embodied-0.5-VLA: A Massive Bimanual Teleoperation Dataset for Vision-Language-Action</title>
      <itunes:title>Hy-Embodied-0.5-VLA: A Massive Bimanual Teleoperation Dataset for Vision-Language-Action</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">ffb39ab7-043a-4200-8cd7-c13c19bb6315</guid>
      <link>https://share.transistor.fm/s/29cf4790</link>
      <description>
        <![CDATA[Released a massive bimanual robot manipulation dataset with 2,163 hours and 250K+ episodes across 70+ tasks, along with a compatible VLA model for multi-view egocentric teleop. The dataset and model are fully compatible with LeRobot v3.0.]]>
      </description>
      <content:encoded>
        <![CDATA[Released a massive bimanual robot manipulation dataset with 2,163 hours and 250K+ episodes across 70+ tasks, along with a compatible VLA model for multi-view egocentric teleop. The dataset and model are fully compatible with LeRobot v3.0.]]>
      </content:encoded>
      <pubDate>Mon, 15 Jun 2026 05:27:45 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/29cf4790/1d18ed3e.mp3" length="20412928" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1272</itunes:duration>
      <itunes:summary>Released a massive bimanual robot manipulation dataset with 2,163 hours and 250K+ episodes across 70+ tasks, along with a compatible VLA model for multi-view egocentric teleop. The dataset and model are fully compatible with LeRobot v3.0.</itunes:summary>
      <itunes:subtitle>Released a massive bimanual robot manipulation dataset with 2,163 hours and 250K+ episodes across 70+ tasks, along with a compatible VLA model for multi-view egocentric teleop. The dataset and model are fully compatible with LeRobot v3.0.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/29cf4790/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Q-Guided Flow: Test-Time Gradient Guidance of Flow Policies</title>
      <itunes:title>Q-Guided Flow: Test-Time Gradient Guidance of Flow Policies</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">d0ad5cba-13f6-468d-8433-e3e5a6e0f8e9</guid>
      <link>https://share.transistor.fm/s/c3c7e426</link>
      <description>
        <![CDATA[New framework for guided flow-matching policies that improves long-horizon robotic control and sample efficiency.]]>
      </description>
      <content:encoded>
        <![CDATA[New framework for guided flow-matching policies that improves long-horizon robotic control and sample efficiency.]]>
      </content:encoded>
      <pubDate>Sun, 14 Jun 2026 14:23:14 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/c3c7e426/1567bf90.mp3" length="33957376" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2123</itunes:duration>
      <itunes:summary>New framework for guided flow-matching policies that improves long-horizon robotic control and sample efficiency.</itunes:summary>
      <itunes:subtitle>New framework for guided flow-matching policies that improves long-horizon robotic control and sample efficiency.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/c3c7e426/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Flow Reversal Steering: Guiding Diffusion-Based Robot Policies with High-Level Reasoning</title>
      <itunes:title>Flow Reversal Steering: Guiding Diffusion-Based Robot Policies with High-Level Reasoning</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">94c52237-9921-4b83-a5ae-5457e54d3ef1</guid>
      <link>https://share.transistor.fm/s/386f0d51</link>
      <description>
        <![CDATA[Introduces flow reversal steering to guide diffusion-based vision-language-action models with high-level VLM reasoning and enables RL directly in the diffusion noise space.]]>
      </description>
      <content:encoded>
        <![CDATA[Introduces flow reversal steering to guide diffusion-based vision-language-action models with high-level VLM reasoning and enables RL directly in the diffusion noise space.]]>
      </content:encoded>
      <pubDate>Sun, 14 Jun 2026 14:13:15 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/386f0d51/f1017a45.mp3" length="36137472" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2259</itunes:duration>
      <itunes:summary>Introduces flow reversal steering to guide diffusion-based vision-language-action models with high-level VLM reasoning and enables RL directly in the diffusion noise space.</itunes:summary>
      <itunes:subtitle>Introduces flow reversal steering to guide diffusion-based vision-language-action models with high-level VLM reasoning and enables RL directly in the diffusion noise space.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/386f0d51/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Test-Time Compute Scaling for Robot Policies (DIRECT)</title>
      <itunes:title>Test-Time Compute Scaling for Robot Policies (DIRECT)</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">9eba0953-a20b-43b6-90e5-ad0146558e65</guid>
      <link>https://share.transistor.fm/s/d5f31868</link>
      <description>
        <![CDATA[Larger models + more thinking + more context improve performance on some prompts but not others; a learned router enables better performance/latency trade-offs.]]>
      </description>
      <content:encoded>
        <![CDATA[Larger models + more thinking + more context improve performance on some prompts but not others; a learned router enables better performance/latency trade-offs.]]>
      </content:encoded>
      <pubDate>Sun, 14 Jun 2026 05:26:08 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/d5f31868/fd7e7bcb.mp3" length="23721472" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1483</itunes:duration>
      <itunes:summary>Larger models + more thinking + more context improve performance on some prompts but not others; a learned router enables better performance/latency trade-offs.</itunes:summary>
      <itunes:subtitle>Larger models + more thinking + more context improve performance on some prompts but not others; a learned router enables better performance/latency trade-offs.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/d5f31868/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>LabVLA: Bringing Vision-Language-Action to the Chemistry Lab</title>
      <itunes:title>LabVLA: Bringing Vision-Language-Action to the Chemistry Lab</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">f915576f-9068-498e-b4a6-664fd5c7e755</guid>
      <link>https://share.transistor.fm/s/aaaba498</link>
      <description>
        <![CDATA[RoboGenesis generates 10K+ lab scenes across 16 robot embodiments; LabVLA (Qwen3-VL + DiT flow-matching) achieves 71.1% success on LabUtopia and transfers to real Franka arms.]]>
      </description>
      <content:encoded>
        <![CDATA[RoboGenesis generates 10K+ lab scenes across 16 robot embodiments; LabVLA (Qwen3-VL + DiT flow-matching) achieves 71.1% success on LabUtopia and transfers to real Franka arms.]]>
      </content:encoded>
      <pubDate>Sun, 14 Jun 2026 05:16:11 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/aaaba498/30bf5e74.mp3" length="40073728" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2505</itunes:duration>
      <itunes:summary>RoboGenesis generates 10K+ lab scenes across 16 robot embodiments; LabVLA (Qwen3-VL + DiT flow-matching) achieves 71.1% success on LabUtopia and transfers to real Franka arms.</itunes:summary>
      <itunes:subtitle>RoboGenesis generates 10K+ lab scenes across 16 robot embodiments; LabVLA (Qwen3-VL + DiT flow-matching) achieves 71.1% success on LabUtopia and transfers to real Franka arms.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/aaaba498/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Humanoid-GPT: A Foundation Model for Zero-Shot Humanoid Control</title>
      <itunes:title>Humanoid-GPT: A Foundation Model for Zero-Shot Humanoid Control</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">4f370d16-f8f4-4602-8b49-66e9228b63a6</guid>
      <link>https://share.transistor.fm/s/932c1fab</link>
      <description>
        <![CDATA[GPT-style Transformer pretrained on 2 billion motion frames that achieves agile, generalist zero-shot control on a real Unitree G1 humanoid for tasks like soccer, dancing, and digging. Requires no fine-tuning or task-specific adaptation.]]>
      </description>
      <content:encoded>
        <![CDATA[GPT-style Transformer pretrained on 2 billion motion frames that achieves agile, generalist zero-shot control on a real Unitree G1 humanoid for tasks like soccer, dancing, and digging. Requires no fine-tuning or task-specific adaptation.]]>
      </content:encoded>
      <pubDate>Sat, 13 Jun 2026 14:17:57 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/932c1fab/a9e9ab9e.mp3" length="25193984" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1575</itunes:duration>
      <itunes:summary>GPT-style Transformer pretrained on 2 billion motion frames that achieves agile, generalist zero-shot control on a real Unitree G1 humanoid for tasks like soccer, dancing, and digging. Requires no fine-tuning or task-specific adaptation.</itunes:summary>
      <itunes:subtitle>GPT-style Transformer pretrained on 2 billion motion frames that achieves agile, generalist zero-shot control on a real Unitree G1 humanoid for tasks like soccer, dancing, and digging. Requires no fine-tuning or task-specific adaptation.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/932c1fab/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>CHORUS: Decentralized Multi-Robot Collaboration with a Single Shared VLA Model</title>
      <itunes:title>CHORUS: Decentralized Multi-Robot Collaboration with a Single Shared VLA Model</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">368917e7-8770-4f25-abc6-7b3041aea0f7</guid>
      <link>https://share.transistor.fm/s/cba23ca8</link>
      <description>
        <![CDATA[Finetunes a single Vision-Language-Action (VLA) foundation model so that any robot in a team can control any other. Outperforms both per-robot specialists and a monolithic centralized policy while scaling to large teams.]]>
      </description>
      <content:encoded>
        <![CDATA[Finetunes a single Vision-Language-Action (VLA) foundation model so that any robot in a team can control any other. Outperforms both per-robot specialists and a monolithic centralized policy while scaling to large teams.]]>
      </content:encoded>
      <pubDate>Sat, 13 Jun 2026 14:07:47 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/cba23ca8/9f44bdcc.mp3" length="35805696" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2238</itunes:duration>
      <itunes:summary>Finetunes a single Vision-Language-Action (VLA) foundation model so that any robot in a team can control any other. Outperforms both per-robot specialists and a monolithic centralized policy while scaling to large teams.</itunes:summary>
      <itunes:subtitle>Finetunes a single Vision-Language-Action (VLA) foundation model so that any robot in a team can control any other. Outperforms both per-robot specialists and a monolithic centralized policy while scaling to large teams.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/cba23ca8/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>RISE: Self-Improving Robot Policy with Compositional World Model</title>
      <itunes:title>RISE: Self-Improving Robot Policy with Compositional World Model</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">5f1e651c-d362-4c97-a047-f534a10920b8</guid>
      <link>https://share.transistor.fm/s/69c83395</link>
      <description>
        <![CDATA[Trains a compositional world model on real robot data to enable closed-loop policy improvement via future prediction and progress evaluation, bypassing both risky real-world RL and traditional sim-to-real gaps.]]>
      </description>
      <content:encoded>
        <![CDATA[Trains a compositional world model on real robot data to enable closed-loop policy improvement via future prediction and progress evaluation, bypassing both risky real-world RL and traditional sim-to-real gaps.]]>
      </content:encoded>
      <pubDate>Sat, 13 Jun 2026 05:18:47 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/69c83395/545ceb8d.mp3" length="37909504" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2370</itunes:duration>
      <itunes:summary>Trains a compositional world model on real robot data to enable closed-loop policy improvement via future prediction and progress evaluation, bypassing both risky real-world RL and traditional sim-to-real gaps.</itunes:summary>
      <itunes:subtitle>Trains a compositional world model on real robot data to enable closed-loop policy improvement via future prediction and progress evaluation, bypassing both risky real-world RL and traditional sim-to-real gaps.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/69c83395/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>EmbodiedOneVision: Interleaved Vision-Text-Action Pretraining for General Robot Control</title>
      <itunes:title>EmbodiedOneVision: Interleaved Vision-Text-Action Pretraining for General Robot Control</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">ca1ae27b-ae5a-4e5b-a379-a5ec989aff54</guid>
      <link>https://share.transistor.fm/s/637d5231</link>
      <description>
        <![CDATA[Open-source 3B unified embodied foundation model trained on 1.5M interleaved vision-text-action samples for perception, planning, and acting.]]>
      </description>
      <content:encoded>
        <![CDATA[Open-source 3B unified embodied foundation model trained on 1.5M interleaved vision-text-action samples for perception, planning, and acting.]]>
      </content:encoded>
      <pubDate>Fri, 12 Jun 2026 14:34:55 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/637d5231/6f705b99.mp3" length="34017792" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2127</itunes:duration>
      <itunes:summary>Open-source 3B unified embodied foundation model trained on 1.5M interleaved vision-text-action samples for perception, planning, and acting.</itunes:summary>
      <itunes:subtitle>Open-source 3B unified embodied foundation model trained on 1.5M interleaved vision-text-action samples for perception, planning, and acting.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/637d5231/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Robix: A Unified Model for Robot Interaction, Reasoning and Planning</title>
      <itunes:title>Robix: A Unified Model for Robot Interaction, Reasoning and Planning</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">1af824ea-a5c9-4042-b031-e623e941e77d</guid>
      <link>https://share.transistor.fm/s/3ee7e33b</link>
      <description>
        <![CDATA[A single vision-language model that unifies reasoning, task planning, and human-robot interaction for complex instructions and long-horizon tasks.]]>
      </description>
      <content:encoded>
        <![CDATA[A single vision-language model that unifies reasoning, task planning, and human-robot interaction for complex instructions and long-horizon tasks.]]>
      </content:encoded>
      <pubDate>Fri, 12 Jun 2026 14:15:36 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/3ee7e33b/6d0a75e8.mp3" length="33631232" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2102</itunes:duration>
      <itunes:summary>A single vision-language model that unifies reasoning, task planning, and human-robot interaction for complex instructions and long-horizon tasks.</itunes:summary>
      <itunes:subtitle>A single vision-language model that unifies reasoning, task planning, and human-robot interaction for complex instructions and long-horizon tasks.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/3ee7e33b/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Robotic World Model: Learning to Simulate for Robust Robot Control</title>
      <itunes:title>Robotic World Model: Learning to Simulate for Robust Robot Control</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">f0ed5e1e-fa19-46ff-9a48-dc2bc00f51e4</guid>
      <link>https://share.transistor.fm/s/d6926620</link>
      <description>
        <![CDATA[Presents a neural network-based world model for model-based reinforcement learning in robotics, focusing on sim-to-real transfer for quadrupedal and humanoid robots. Enables robust policy optimization through learned environment simulation.]]>
      </description>
      <content:encoded>
        <![CDATA[Presents a neural network-based world model for model-based reinforcement learning in robotics, focusing on sim-to-real transfer for quadrupedal and humanoid robots. Enables robust policy optimization through learned environment simulation.]]>
      </content:encoded>
      <pubDate>Fri, 12 Jun 2026 05:07:03 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/d6926620/d317ab6a.mp3" length="19109888" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1195</itunes:duration>
      <itunes:summary>Presents a neural network-based world model for model-based reinforcement learning in robotics, focusing on sim-to-real transfer for quadrupedal and humanoid robots. Enables robust policy optimization through learned environment simulation.</itunes:summary>
      <itunes:subtitle>Presents a neural network-based world model for model-based reinforcement learning in robotics, focusing on sim-to-real transfer for quadrupedal and humanoid robots. Enables robust policy optimization through learned environment simulation.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/d6926620/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>AnchorWorld: Embodied Egocentric World Simulation with View-based Evolution Customization</title>
      <itunes:title>AnchorWorld: Embodied Egocentric World Simulation with View-based Evolution Customization</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">2f8228bf-50ef-4e3e-8c8e-62cee84ec691</guid>
      <link>https://share.transistor.fm/s/aaae5b88</link>
      <description>
        <![CDATA[Embodied egocentric simulation framework that controls first-person worlds with 3D human motion and customizes evolving scenes via pose-anchored views.]]>
      </description>
      <content:encoded>
        <![CDATA[Embodied egocentric simulation framework that controls first-person worlds with 3D human motion and customizes evolving scenes via pose-anchored views.]]>
      </content:encoded>
      <pubDate>Wed, 10 Jun 2026 14:21:45 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/aaae5b88/ab1db76b.mp3" length="22286848" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1393</itunes:duration>
      <itunes:summary>Embodied egocentric simulation framework that controls first-person worlds with 3D human motion and customizes evolving scenes via pose-anchored views.</itunes:summary>
      <itunes:subtitle>Embodied egocentric simulation framework that controls first-person worlds with 3D human motion and customizes evolving scenes via pose-anchored views.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/aaae5b88/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>ArtiFixer: Few-Step Diffusion for 3D Scene Reconstruction</title>
      <itunes:title>ArtiFixer: Few-Step Diffusion for 3D Scene Reconstruction</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">e050b966-4a58-4b2e-a448-5b12f8a31a62</guid>
      <link>https://share.transistor.fm/s/593c98c5</link>
      <description>
        <![CDATA[Few-step auto-regressive diffusion model that converts broken 3D reconstructions into fully realized scenes, outperforming prior methods by 1-3 dB PSNR.]]>
      </description>
      <content:encoded>
        <![CDATA[Few-step auto-regressive diffusion model that converts broken 3D reconstructions into fully realized scenes, outperforming prior methods by 1-3 dB PSNR.]]>
      </content:encoded>
      <pubDate>Wed, 10 Jun 2026 14:10:57 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/593c98c5/63a9a359.mp3" length="25036800" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1565</itunes:duration>
      <itunes:summary>Few-step auto-regressive diffusion model that converts broken 3D reconstructions into fully realized scenes, outperforming prior methods by 1-3 dB PSNR.</itunes:summary>
      <itunes:subtitle>Few-step auto-regressive diffusion model that converts broken 3D reconstructions into fully realized scenes, outperforming prior methods by 1-3 dB PSNR.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/593c98c5/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Deployment-Time Memorization in Foundation-Model Agents</title>
      <itunes:title>Deployment-Time Memorization in Foundation-Model Agents</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">ad05e469-ffc1-46d5-9ba3-f2f22644d825</guid>
      <link>https://share.transistor.fm/s/d76febee</link>
      <description>
        <![CDATA[Examines memorization phenomena that occur during deployment of foundation model agents in practical applications.]]>
      </description>
      <content:encoded>
        <![CDATA[Examines memorization phenomena that occur during deployment of foundation model agents in practical applications.]]>
      </content:encoded>
      <pubDate>Wed, 10 Jun 2026 05:17:10 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/d76febee/fa6c5760.mp3" length="37934592" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2371</itunes:duration>
      <itunes:summary>Examines memorization phenomena that occur during deployment of foundation model agents in practical applications.</itunes:summary>
      <itunes:subtitle>Examines memorization phenomena that occur during deployment of foundation model agents in practical applications.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/d76febee/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Adversarial Machine Learning: Taxonomy, Threat Models, and Mitigation Strategies in Deep Neural Networks</title>
      <itunes:title>Adversarial Machine Learning: Taxonomy, Threat Models, and Mitigation Strategies in Deep Neural Networks</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">011e7ae9-3b29-4d1d-9273-3273bc43dd3a</guid>
      <link>https://share.transistor.fm/s/00ca06eb</link>
      <description>
        <![CDATA[Focuses on security aspects of deep neural networks, providing taxonomy and mitigation strategies for adversarial attacks.]]>
      </description>
      <content:encoded>
        <![CDATA[Focuses on security aspects of deep neural networks, providing taxonomy and mitigation strategies for adversarial attacks.]]>
      </content:encoded>
      <pubDate>Tue, 09 Jun 2026 14:08:24 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/00ca06eb/159cbf60.mp3" length="33384448" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2087</itunes:duration>
      <itunes:summary>Focuses on security aspects of deep neural networks, providing taxonomy and mitigation strategies for adversarial attacks.</itunes:summary>
      <itunes:subtitle>Focuses on security aspects of deep neural networks, providing taxonomy and mitigation strategies for adversarial attacks.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/00ca06eb/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>SoCRATES: Evaluating LLM Mediators in Conflict Scenarios</title>
      <itunes:title>SoCRATES: Evaluating LLM Mediators in Conflict Scenarios</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">0c7fb3fb-cdde-40c2-91ce-4dc2d0b0007b</guid>
      <link>https://share.transistor.fm/s/2bf01f12</link>
      <description>
        <![CDATA[First comprehensive framework for evaluating LLM mediators in real-time, emotional, socio-cognitive scenarios.]]>
      </description>
      <content:encoded>
        <![CDATA[First comprehensive framework for evaluating LLM mediators in real-time, emotional, socio-cognitive scenarios.]]>
      </content:encoded>
      <pubDate>Tue, 09 Jun 2026 05:38:49 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/2bf01f12/6f58f170.mp3" length="20182016" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1258</itunes:duration>
      <itunes:summary>First comprehensive framework for evaluating LLM mediators in real-time, emotional, socio-cognitive scenarios.</itunes:summary>
      <itunes:subtitle>First comprehensive framework for evaluating LLM mediators in real-time, emotional, socio-cognitive scenarios.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/2bf01f12/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Unembedding Matrix as a Feature Lens: Unlocking Better Text Embeddings</title>
      <itunes:title>Unembedding Matrix as a Feature Lens: Unlocking Better Text Embeddings</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">998c0492-973d-4063-81ea-b67a064d9353</guid>
      <link>https://share.transistor.fm/s/29d844d8</link>
      <description>
        <![CDATA[Improves embedding quality without extra training by using the unembedding matrix as a feature lens for text embeddings.]]>
      </description>
      <content:encoded>
        <![CDATA[Improves embedding quality without extra training by using the unembedding matrix as a feature lens for text embeddings.]]>
      </content:encoded>
      <pubDate>Tue, 09 Jun 2026 05:32:33 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/29d844d8/7dabca90.mp3" length="23655936" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1479</itunes:duration>
      <itunes:summary>Improves embedding quality without extra training by using the unembedding matrix as a feature lens for text embeddings.</itunes:summary>
      <itunes:subtitle>Improves embedding quality without extra training by using the unembedding matrix as a feature lens for text embeddings.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/29d844d8/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>LeanMarathon: Autonomous Formalization of Math Proofs on Erdős Problems</title>
      <itunes:title>LeanMarathon: Autonomous Formalization of Math Proofs on Erdős Problems</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">1864ac0e-93dc-4b13-a663-286ecf5d41a3</guid>
      <link>https://share.transistor.fm/s/3934aea6</link>
      <description>
        <![CDATA[Presents an autonomous system for formalizing mathematical proofs, specifically targeting Erdős problems. Demonstrates automated proof formalization capabilities in the Lean theorem prover.]]>
      </description>
      <content:encoded>
        <![CDATA[Presents an autonomous system for formalizing mathematical proofs, specifically targeting Erdős problems. Demonstrates automated proof formalization capabilities in the Lean theorem prover.]]>
      </content:encoded>
      <pubDate>Mon, 08 Jun 2026 14:38:40 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/3934aea6/1a5d2dd0.mp3" length="26260480" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1642</itunes:duration>
      <itunes:summary>Presents an autonomous system for formalizing mathematical proofs, specifically targeting Erdős problems. Demonstrates automated proof formalization capabilities in the Lean theorem prover.</itunes:summary>
      <itunes:subtitle>Presents an autonomous system for formalizing mathematical proofs, specifically targeting Erdős problems. Demonstrates automated proof formalization capabilities in the Lean theorem prover.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/3934aea6/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Deep Research Agents: Survey and Roadmap for Autonomous AI Research</title>
      <itunes:title>Deep Research Agents: Survey and Roadmap for Autonomous AI Research</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">d36dd184-9703-481c-abbb-2522ee424782</guid>
      <link>https://share.transistor.fm/s/0369862d</link>
      <description>
        <![CDATA[Provides a comprehensive taxonomy, benchmarks, and future directions for autonomous AI research agents. Examines systematic approaches to developing agents capable of conducting independent research.]]>
      </description>
      <content:encoded>
        <![CDATA[Provides a comprehensive taxonomy, benchmarks, and future directions for autonomous AI research agents. Examines systematic approaches to developing agents capable of conducting independent research.]]>
      </content:encoded>
      <pubDate>Mon, 08 Jun 2026 14:26:32 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/0369862d/d834e0ef.mp3" length="35885056" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2243</itunes:duration>
      <itunes:summary>Provides a comprehensive taxonomy, benchmarks, and future directions for autonomous AI research agents. Examines systematic approaches to developing agents capable of conducting independent research.</itunes:summary>
      <itunes:subtitle>Provides a comprehensive taxonomy, benchmarks, and future directions for autonomous AI research agents. Examines systematic approaches to developing agents capable of conducting independent research.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/0369862d/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Cosmos 3: Omnimodal World Models for Physical AI</title>
      <itunes:title>Cosmos 3: Omnimodal World Models for Physical AI</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">5f90d5f2-bbdc-41da-8756-99987fb05e63</guid>
      <link>https://share.transistor.fm/s/5f0f456b</link>
      <description>
        <![CDATA[Omnimodal world models explicitly designed for Physical AI and robotics applications. Enables improved simulation and control for robotic systems through multimodal understanding.]]>
      </description>
      <content:encoded>
        <![CDATA[Omnimodal world models explicitly designed for Physical AI and robotics applications. Enables improved simulation and control for robotic systems through multimodal understanding.]]>
      </content:encoded>
      <pubDate>Sun, 07 Jun 2026 14:08:40 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/5f0f456b/834a1142.mp3" length="26911232" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1682</itunes:duration>
      <itunes:summary>Omnimodal world models explicitly designed for Physical AI and robotics applications. Enables improved simulation and control for robotic systems through multimodal understanding.</itunes:summary>
      <itunes:subtitle>Omnimodal world models explicitly designed for Physical AI and robotics applications. Enables improved simulation and control for robotic systems through multimodal understanding.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/5f0f456b/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Humanoid-GPT: GPT-Style Transformer for Zero-Shot Dynamic Humanoid Control</title>
      <itunes:title>Humanoid-GPT: GPT-Style Transformer for Zero-Shot Dynamic Humanoid Control</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">3850245b-44bd-4eb0-97ca-72513c550983</guid>
      <link>https://share.transistor.fm/s/2718ecff</link>
      <description>
        <![CDATA[GPT-style Transformer trained on 2 billion motion frames enabling zero-shot dynamic humanoid control for tasks like soccer, dancing, and digging on real Unitree G1 robots without fine-tuning. Breaks the agility-generalization trade-off in humanoid robotics.]]>
      </description>
      <content:encoded>
        <![CDATA[GPT-style Transformer trained on 2 billion motion frames enabling zero-shot dynamic humanoid control for tasks like soccer, dancing, and digging on real Unitree G1 robots without fine-tuning. Breaks the agility-generalization trade-off in humanoid robotics.]]>
      </content:encoded>
      <pubDate>Sun, 07 Jun 2026 05:11:47 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/2718ecff/23314639.mp3" length="22033920" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1378</itunes:duration>
      <itunes:summary>GPT-style Transformer trained on 2 billion motion frames enabling zero-shot dynamic humanoid control for tasks like soccer, dancing, and digging on real Unitree G1 robots without fine-tuning. Breaks the agility-generalization trade-off in humanoid robotics.</itunes:summary>
      <itunes:subtitle>GPT-style Transformer trained on 2 billion motion frames enabling zero-shot dynamic humanoid control for tasks like soccer, dancing, and digging on real Unitree G1 robots without fine-tuning. Breaks the agility-generalization trade-off in humanoid robotic</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/2718ecff/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Bending Paper, Shaping Dexterity: The Robotic Origami Challenge</title>
      <itunes:title>Bending Paper, Shaping Dexterity: The Robotic Origami Challenge</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">499a8947-a593-497e-84d2-e58d9addaf22</guid>
      <link>https://share.transistor.fm/s/4861aa90</link>
      <description>
        <![CDATA[New IROS benchmark providing 500+ teleoperation episodes and physically accurate simulation assets for training policies that outperform human origami experts.]]>
      </description>
      <content:encoded>
        <![CDATA[New IROS benchmark providing 500+ teleoperation episodes and physically accurate simulation assets for training policies that outperform human origami experts.]]>
      </content:encoded>
      <pubDate>Fri, 05 Jun 2026 14:20:26 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/4861aa90/0445f726.mp3" length="27856896" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1742</itunes:duration>
      <itunes:summary>New IROS benchmark providing 500+ teleoperation episodes and physically accurate simulation assets for training policies that outperform human origami experts.</itunes:summary>
      <itunes:subtitle>New IROS benchmark providing 500+ teleoperation episodes and physically accurate simulation assets for training policies that outperform human origami experts.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/4861aa90/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>GraspGen-X: A Foundation Model for Zero-Shot 6-DoF Grasping</title>
      <itunes:title>GraspGen-X: A Foundation Model for Zero-Shot 6-DoF Grasping</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">e08178fa-c488-4d28-9b7d-928c600ff0be</guid>
      <link>https://share.transistor.fm/s/b006b1ea</link>
      <description>
        <![CDATA[First foundation model for zero-shot grasping trained on billions of simulated grasps, enabling generalized manipulation without task-specific training.]]>
      </description>
      <content:encoded>
        <![CDATA[First foundation model for zero-shot grasping trained on billions of simulated grasps, enabling generalized manipulation without task-specific training.]]>
      </content:encoded>
      <pubDate>Fri, 05 Jun 2026 14:10:23 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/b006b1ea/df603c80.mp3" length="33796608" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2113</itunes:duration>
      <itunes:summary>First foundation model for zero-shot grasping trained on billions of simulated grasps, enabling generalized manipulation without task-specific training.</itunes:summary>
      <itunes:subtitle>First foundation model for zero-shot grasping trained on billions of simulated grasps, enabling generalized manipulation without task-specific training.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/b006b1ea/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>When Does Deep RL Beat Calibrated Baselines?</title>
      <itunes:title>When Does Deep RL Beat Calibrated Baselines?</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">af059b83-1e33-46c1-9a43-d653986958b1</guid>
      <link>https://share.transistor.fm/s/dd4aadd0</link>
      <description>
        <![CDATA[Benchmark study examining when deep reinforcement learning outperforms calibrated baseline methods in adaptive resource control tasks.]]>
      </description>
      <content:encoded>
        <![CDATA[Benchmark study examining when deep reinforcement learning outperforms calibrated baseline methods in adaptive resource control tasks.]]>
      </content:encoded>
      <pubDate>Thu, 04 Jun 2026 05:24:54 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/dd4aadd0/50cdaec6.mp3" length="15071232" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>942</itunes:duration>
      <itunes:summary>Benchmark study examining when deep reinforcement learning outperforms calibrated baseline methods in adaptive resource control tasks.</itunes:summary>
      <itunes:subtitle>Benchmark study examining when deep reinforcement learning outperforms calibrated baseline methods in adaptive resource control tasks.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/dd4aadd0/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Training Deep Networks as Random Effects: An Optimization–Inference Duality</title>
      <itunes:title>Training Deep Networks as Random Effects: An Optimization–Inference Duality</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">3a744ed9-fb9f-459b-bfbe-aa79fedee7ed</guid>
      <link>https://share.transistor.fm/s/ecee17c1</link>
      <description>
        <![CDATA[Explores training dynamics of deep neural networks through a statistical lens, examining the duality between optimization and inference perspectives.]]>
      </description>
      <content:encoded>
        <![CDATA[Explores training dynamics of deep neural networks through a statistical lens, examining the duality between optimization and inference perspectives.]]>
      </content:encoded>
      <pubDate>Thu, 04 Jun 2026 05:10:17 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/ecee17c1/911cb85f.mp3" length="20697088" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1294</itunes:duration>
      <itunes:summary>Explores training dynamics of deep neural networks through a statistical lens, examining the duality between optimization and inference perspectives.</itunes:summary>
      <itunes:subtitle>Explores training dynamics of deep neural networks through a statistical lens, examining the duality between optimization and inference perspectives.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/ecee17c1/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Generative Depth Supervision for Embodied Vision-Language Models</title>
      <itunes:title>Generative Depth Supervision for Embodied Vision-Language Models</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">f839ea0a-1819-4d3f-9f3c-f8d0a90e6e6d</guid>
      <link>https://share.transistor.fm/s/717e089a</link>
      <description>
        <![CDATA[Vision-language model that adds generative depth prediction during pre-training for physical grounding; achieves SOTA on embodied benchiments and transfers directly to real-robot tasks.]]>
      </description>
      <content:encoded>
        <![CDATA[Vision-language model that adds generative depth prediction during pre-training for physical grounding; achieves SOTA on embodied benchiments and transfers directly to real-robot tasks.]]>
      </content:encoded>
      <pubDate>Tue, 02 Jun 2026 05:07:28 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/717e089a/c80dce1e.mp3" length="27442688" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1716</itunes:duration>
      <itunes:summary>Vision-language model that adds generative depth prediction during pre-training for physical grounding; achieves SOTA on embodied benchiments and transfers directly to real-robot tasks.</itunes:summary>
      <itunes:subtitle>Vision-language model that adds generative depth prediction during pre-training for physical grounding; achieves SOTA on embodied benchiments and transfers directly to real-robot tasks.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/717e089a/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>PointWorld: Scaling 3D World Models for In-The-Wild Robotic Manipulation</title>
      <itunes:title>PointWorld: Scaling 3D World Models for In-The-Wild Robotic Manipulation</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">d3d33289-eafe-43f1-ad92-0caae0286445</guid>
      <link>https://share.transistor.fm/s/15c08772</link>
      <description>
        <![CDATA[Presents a 3D point-cloud-based world model trained on mixed real/sim data that enables zero-shot grasping and articulated object handling on real robots by explicitly modeling spatial structure.]]>
      </description>
      <content:encoded>
        <![CDATA[Presents a 3D point-cloud-based world model trained on mixed real/sim data that enables zero-shot grasping and articulated object handling on real robots by explicitly modeling spatial structure.]]>
      </content:encoded>
      <pubDate>Mon, 01 Jun 2026 05:12:12 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/15c08772/949d5f9f.mp3" length="29595136" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1850</itunes:duration>
      <itunes:summary>Presents a 3D point-cloud-based world model trained on mixed real/sim data that enables zero-shot grasping and articulated object handling on real robots by explicitly modeling spatial structure.</itunes:summary>
      <itunes:subtitle>Presents a 3D point-cloud-based world model trained on mixed real/sim data that enables zero-shot grasping and articulated object handling on real robots by explicitly modeling spatial structure.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/15c08772/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding</title>
      <itunes:title>LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">fae928e1-49ef-4e6b-9e5e-3d4b1874ce52</guid>
      <link>https://share.transistor.fm/s/3cecfbf7</link>
      <description>
        <![CDATA[NVIDIA Research presents LocateAnything, a unified generative grounding and detection framework based on Parallel Box Decoding (PBD). Unlike prior VLMs that serialize bounding boxes into sequential coordinate tokens, PBD treats each box as an atomic unit and predicts all coordinates in a single forward pass. This preserves intra-box geometric coherence while achieving 2.5x faster decoding throughput. The model supports diverse localization tasks including document understanding, GUI grounding, dense object detection, and OCR localization. Built on Moon-ViT vision encoder and Qwen2.5 language decoder. Trained on LocateAnything-Data with 138M language queries and 785M bounding boxes. Achieves state-of-the-art on LVIS, M6Doc, and ScreenSpot-Pro benchmarks. Models and demo available on HuggingFace.]]>
      </description>
      <content:encoded>
        <![CDATA[NVIDIA Research presents LocateAnything, a unified generative grounding and detection framework based on Parallel Box Decoding (PBD). Unlike prior VLMs that serialize bounding boxes into sequential coordinate tokens, PBD treats each box as an atomic unit and predicts all coordinates in a single forward pass. This preserves intra-box geometric coherence while achieving 2.5x faster decoding throughput. The model supports diverse localization tasks including document understanding, GUI grounding, dense object detection, and OCR localization. Built on Moon-ViT vision encoder and Qwen2.5 language decoder. Trained on LocateAnything-Data with 138M language queries and 785M bounding boxes. Achieves state-of-the-art on LVIS, M6Doc, and ScreenSpot-Pro benchmarks. Models and demo available on HuggingFace.]]>
      </content:encoded>
      <pubDate>Sun, 31 May 2026 16:41:21 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/3cecfbf7/9f8f177d.mp3" length="31186432" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1950</itunes:duration>
      <itunes:summary>NVIDIA Research presents LocateAnything, a unified generative grounding and detection framework based on Parallel Box Decoding (PBD). Unlike prior VLMs that serialize bounding boxes into sequential coordinate tokens, PBD treats each box as an atomic unit and predicts all coordinates in a single forward pass. This preserves intra-box geometric coherence while achieving 2.5x faster decoding throughput. The model supports diverse localization tasks including document understanding, GUI grounding, dense object detection, and OCR localization. Built on Moon-ViT vision encoder and Qwen2.5 language decoder. Trained on LocateAnything-Data with 138M language queries and 785M bounding boxes. Achieves state-of-the-art on LVIS, M6Doc, and ScreenSpot-Pro benchmarks. Models and demo available on HuggingFace.</itunes:summary>
      <itunes:subtitle>NVIDIA Research presents LocateAnything, a unified generative grounding and detection framework based on Parallel Box Decoding (PBD). Unlike prior VLMs that serialize bounding boxes into sequential coordinate tokens, PBD treats each box as an atomic unit </itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/3cecfbf7/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>LT2: Linear-Time Looped Transformers</title>
      <itunes:title>LT2: Linear-Time Looped Transformers</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">5118bc6f-80ed-4d6a-a530-6894860837ac</guid>
      <link>https://share.transistor.fm/s/33bc5fd6</link>
      <description>
        <![CDATA[Replaces quadratic softmax attention in looped architectures with linear/sparse mechanisms for iterative memory refinement, achieving parity with standard looped transformers at much lower cost.]]>
      </description>
      <content:encoded>
        <![CDATA[Replaces quadratic softmax attention in looped architectures with linear/sparse mechanisms for iterative memory refinement, achieving parity with standard looped transformers at much lower cost.]]>
      </content:encoded>
      <pubDate>Sun, 31 May 2026 05:21:02 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/33bc5fd6/cf84ded5.mp3" length="40908288" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2557</itunes:duration>
      <itunes:summary>Replaces quadratic softmax attention in looped architectures with linear/sparse mechanisms for iterative memory refinement, achieving parity with standard looped transformers at much lower cost.</itunes:summary>
      <itunes:subtitle>Replaces quadratic softmax attention in looped architectures with linear/sparse mechanisms for iterative memory refinement, achieving parity with standard looped transformers at much lower cost.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/33bc5fd6/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>One Learning Rate Doesn't Fit All: Layerwise Spectral Scheduling for Transformers</title>
      <itunes:title>One Learning Rate Doesn't Fit All: Layerwise Spectral Scheduling for Transformers</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">51e52fc3-6f6d-4cc5-8696-6391d81160ff</guid>
      <link>https://share.transistor.fm/s/69013e05</link>
      <description>
        <![CDATA[Shows that modern transformers are highly heterogeneous across layers and proposes layerwise learning rates based on weight spectrum shape, yielding up to 1.5× training speedup on LLaMA/GPT-style models.]]>
      </description>
      <content:encoded>
        <![CDATA[Shows that modern transformers are highly heterogeneous across layers and proposes layerwise learning rates based on weight spectrum shape, yielding up to 1.5× training speedup on LLaMA/GPT-style models.]]>
      </content:encoded>
      <pubDate>Sun, 31 May 2026 05:07:55 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/69013e05/a236cedd.mp3" length="25507328" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1595</itunes:duration>
      <itunes:summary>Shows that modern transformers are highly heterogeneous across layers and proposes layerwise learning rates based on weight spectrum shape, yielding up to 1.5× training speedup on LLaMA/GPT-style models.</itunes:summary>
      <itunes:subtitle>Shows that modern transformers are highly heterogeneous across layers and proposes layerwise learning rates based on weight spectrum shape, yielding up to 1.5× training speedup on LLaMA/GPT-style models.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/69013e05/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>SimToolReal: Procedural Tool Generation and a Universal Objective for Zero-Shot Tool Manipulation</title>
      <itunes:title>SimToolReal: Procedural Tool Generation and a Universal Objective for Zero-Shot Tool Manipulation</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">f65f8454-8153-476d-a4b1-fa67005dece9</guid>
      <link>https://share.transistor.fm/s/dff94c99</link>
      <description>
        <![CDATA[Trains generalist policies in simulation on procedurally generated tools to move objects, enabling real-world tool use across varied shapes/sizes.]]>
      </description>
      <content:encoded>
        <![CDATA[Trains generalist policies in simulation on procedurally generated tools to move objects, enabling real-world tool use across varied shapes/sizes.]]>
      </content:encoded>
      <pubDate>Sat, 30 May 2026 14:19:08 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/dff94c99/ca1b299b.mp3" length="24625152" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1540</itunes:duration>
      <itunes:summary>Trains generalist policies in simulation on procedurally generated tools to move objects, enabling real-world tool use across varied shapes/sizes.</itunes:summary>
      <itunes:subtitle>Trains generalist policies in simulation on procedurally generated tools to move objects, enabling real-world tool use across varied shapes/sizes.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/dff94c99/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Robometer and the Future of Robotic Reward Modeling</title>
      <itunes:title>Robometer and the Future of Robotic Reward Modeling</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">7d488a14-d7c5-42ea-97a3-0c9ae37069c9</guid>
      <link>https://share.transistor.fm/s/f5c30aa2</link>
      <description>
        <![CDATA[New framework for scalable robotic reward modeling using trajectory comparisons to train general-purpose reward models.]]>
      </description>
      <content:encoded>
        <![CDATA[New framework for scalable robotic reward modeling using trajectory comparisons to train general-purpose reward models.]]>
      </content:encoded>
      <pubDate>Sat, 30 May 2026 14:07:00 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/f5c30aa2/f7a10aed.mp3" length="41579520" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2599</itunes:duration>
      <itunes:summary>New framework for scalable robotic reward modeling using trajectory comparisons to train general-purpose reward models.</itunes:summary>
      <itunes:subtitle>New framework for scalable robotic reward modeling using trajectory comparisons to train general-purpose reward models.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/f5c30aa2/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Qwen-VLA: A Generalist Vision–Language–Action Robot Model</title>
      <itunes:title>Qwen-VLA: A Generalist Vision–Language–Action Robot Model</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">a93e43b4-0c41-403c-add4-4ecd77185f6c</guid>
      <link>https://share.transistor.fm/s/0a6130ce</link>
      <description>
        <![CDATA[A single generalist VLA built on Qwen3.5-4B + 1.15B DiT flow-matching action decoder that unifies manipulation, navigation, and trajectory prediction across 11 embodiments via text-described embodiment prompts. Trained in four stages and outperforms task-specific specialists on real ALOHA and sim benchmarks without per-task fine-tuning.]]>
      </description>
      <content:encoded>
        <![CDATA[A single generalist VLA built on Qwen3.5-4B + 1.15B DiT flow-matching action decoder that unifies manipulation, navigation, and trajectory prediction across 11 embodiments via text-described embodiment prompts. Trained in four stages and outperforms task-specific specialists on real ALOHA and sim benchmarks without per-task fine-tuning.]]>
      </content:encoded>
      <pubDate>Fri, 29 May 2026 14:15:49 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/0a6130ce/29bfeb74.mp3" length="34091008" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2131</itunes:duration>
      <itunes:summary>A single generalist VLA built on Qwen3.5-4B + 1.15B DiT flow-matching action decoder that unifies manipulation, navigation, and trajectory prediction across 11 embodiments via text-described embodiment prompts. Trained in four stages and outperforms task-specific specialists on real ALOHA and sim benchmarks without per-task fine-tuning.</itunes:summary>
      <itunes:subtitle>A single generalist VLA built on Qwen3.5-4B + 1.15B DiT flow-matching action decoder that unifies manipulation, navigation, and trajectory prediction across 11 embodiments via text-described embodiment prompts. Trained in four stages and outperforms task-</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/0a6130ce/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>EXPO-FT: Sample-Efficient Reinforcement Learning Fine-Tuning for Vision-Language-Action Models</title>
      <itunes:title>EXPO-FT: Sample-Efficient Reinforcement Learning Fine-Tuning for Vision-Language-Action Models</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">09987100-f12c-4770-8363-39544a79a1f7</guid>
      <link>https://share.transistor.fm/s/b9a6cfcf</link>
      <description>
        <![CDATA[Extends the EXPO method with real-world RL post-training for VLAs using image observations, action chunking, DAgger, and on-the-fly Q-value maximization. Achieves 30/30 success on 8 challenging manipulation tasks with only ~19 min of RL data on average.]]>
      </description>
      <content:encoded>
        <![CDATA[Extends the EXPO method with real-world RL post-training for VLAs using image observations, action chunking, DAgger, and on-the-fly Q-value maximization. Achieves 30/30 success on 8 challenging manipulation tasks with only ~19 min of RL data on average.]]>
      </content:encoded>
      <pubDate>Fri, 29 May 2026 05:13:54 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/b9a6cfcf/00609338.mp3" length="32616960" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2039</itunes:duration>
      <itunes:summary>Extends the EXPO method with real-world RL post-training for VLAs using image observations, action chunking, DAgger, and on-the-fly Q-value maximization. Achieves 30/30 success on 8 challenging manipulation tasks with only ~19 min of RL data on average.</itunes:summary>
      <itunes:subtitle>Extends the EXPO method with real-world RL post-training for VLAs using image observations, action chunking, DAgger, and on-the-fly Q-value maximization. Achieves 30/30 success on 8 challenging manipulation tasks with only ~19 min of RL data on average.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/b9a6cfcf/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>RoboMeter: Learning Dense Rewards from Successes and Failures</title>
      <itunes:title>RoboMeter: Learning Dense Rewards from Successes and Failures</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">1e27be3c-b085-4ee3-b6b4-ad6637442ec7</guid>
      <link>https://share.transistor.fm/s/ef26fc2e</link>
      <description>
        <![CDATA[RoboMeter trains dense reward models from both successful and failed robot trajectories, solving a key gap in prior methods that only learn from expert demos.]]>
      </description>
      <content:encoded>
        <![CDATA[RoboMeter trains dense reward models from both successful and failed robot trajectories, solving a key gap in prior methods that only learn from expert demos.]]>
      </content:encoded>
      <pubDate>Thu, 28 May 2026 22:00:45 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/ef26fc2e/1263e210.mp3" length="35892224" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2244</itunes:duration>
      <itunes:summary>RoboMeter trains dense reward models from both successful and failed robot trajectories, solving a key gap in prior methods that only learn from expert demos.</itunes:summary>
      <itunes:subtitle>RoboMeter trains dense reward models from both successful and failed robot trajectories, solving a key gap in prior methods that only learn from expert demos.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/ef26fc2e/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>MobileGym: A Controllable, Parallel Sandbox for Mobile GUI Agents</title>
      <itunes:title>MobileGym: A Controllable, Parallel Sandbox for Mobile GUI Agents</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">41c8d82e-5e3f-4476-b3f4-4a0700692e75</guid>
      <link>https://share.transistor.fm/s/14d94c5b</link>
      <description>
        <![CDATA[Browser-hosted mobile environment with JSON state, deterministic judges, and 256 parallel rollouts. Reports +40.7 real-device points after GRPO training on 416 tasks for GUI agent development.]]>
      </description>
      <content:encoded>
        <![CDATA[Browser-hosted mobile environment with JSON state, deterministic judges, and 256 parallel rollouts. Reports +40.7 real-device points after GRPO training on 416 tasks for GUI agent development.]]>
      </content:encoded>
      <pubDate>Wed, 27 May 2026 05:34:07 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/14d94c5b/887b91ec.mp3" length="51115520" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>3195</itunes:duration>
      <itunes:summary>Browser-hosted mobile environment with JSON state, deterministic judges, and 256 parallel rollouts. Reports +40.7 real-device points after GRPO training on 416 tasks for GUI agent development.</itunes:summary>
      <itunes:subtitle>Browser-hosted mobile environment with JSON state, deterministic judges, and 256 parallel rollouts. Reports +40.7 real-device points after GRPO training on 416 tasks for GUI agent development.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/14d94c5b/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>ANY2ANY: Efficient Cross-Embodiment Transfer for Humanoid Whole-Body Tracking</title>
      <itunes:title>ANY2ANY: Efficient Cross-Embodiment Transfer for Humanoid Whole-Body Tracking</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">bace626d-6fe2-4532-9edf-73030c234d61</guid>
      <link>https://share.transistor.fm/s/48e1bd9b</link>
      <description>
        <![CDATA[Introduces a method to transfer a Unitree G1 foundation policy (Gear-Sonic) to LimX Oli/Luna humanoids using only 1% of the original compute/data. Achieves fast convergence and strong tracking performance for humanoid whole-body control.]]>
      </description>
      <content:encoded>
        <![CDATA[Introduces a method to transfer a Unitree G1 foundation policy (Gear-Sonic) to LimX Oli/Luna humanoids using only 1% of the original compute/data. Achieves fast convergence and strong tracking performance for humanoid whole-body control.]]>
      </content:encoded>
      <pubDate>Wed, 27 May 2026 05:19:30 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/48e1bd9b/cfd93236.mp3" length="19084288" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1193</itunes:duration>
      <itunes:summary>Introduces a method to transfer a Unitree G1 foundation policy (Gear-Sonic) to LimX Oli/Luna humanoids using only 1% of the original compute/data. Achieves fast convergence and strong tracking performance for humanoid whole-body control.</itunes:summary>
      <itunes:subtitle>Introduces a method to transfer a Unitree G1 foundation policy (Gear-Sonic) to LimX Oli/Luna humanoids using only 1% of the original compute/data. Achieves fast convergence and strong tracking performance for humanoid whole-body control.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/48e1bd9b/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>TriSplat: Feed-Forward 3D Reconstruction with Triangulated Meshes</title>
      <itunes:title>TriSplat: Feed-Forward 3D Reconstruction with Triangulated Meshes</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">f630ddb7-aeca-4624-a052-84a319baca94</guid>
      <link>https://share.transistor.fm/s/ada33c9a</link>
      <description>
        <![CDATA[Outputs physics-engine-compatible triangle meshes directly from sparse, unposed images without Gaussian splatting or post-processing.]]>
      </description>
      <content:encoded>
        <![CDATA[Outputs physics-engine-compatible triangle meshes directly from sparse, unposed images without Gaussian splatting or post-processing.]]>
      </content:encoded>
      <pubDate>Tue, 26 May 2026 14:31:44 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/ada33c9a/1261522b.mp3" length="41567232" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2598</itunes:duration>
      <itunes:summary>Outputs physics-engine-compatible triangle meshes directly from sparse, unposed images without Gaussian splatting or post-processing.</itunes:summary>
      <itunes:subtitle>Outputs physics-engine-compatible triangle meshes directly from sparse, unposed images without Gaussian splatting or post-processing.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/ada33c9a/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>MIKASA-Robo-VLA: A Memory-Intensive Benchmark for Vision-Language-Action Robotics</title>
      <itunes:title>MIKASA-Robo-VLA: A Memory-Intensive Benchmark for Vision-Language-Action Robotics</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">fe0dbab2-2152-47ba-9d41-844b327514e1</guid>
      <link>https://share.transistor.fm/s/5c955d26</link>
      <description>
        <![CDATA[Releases a benchmark suite for systematically evaluating memory in Vision-Language-Action policies on tabletop manipulation tasks.]]>
      </description>
      <content:encoded>
        <![CDATA[Releases a benchmark suite for systematically evaluating memory in Vision-Language-Action policies on tabletop manipulation tasks.]]>
      </content:encoded>
      <pubDate>Tue, 26 May 2026 14:11:34 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/5c955d26/4c49c468.mp3" length="27588608" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1725</itunes:duration>
      <itunes:summary>Releases a benchmark suite for systematically evaluating memory in Vision-Language-Action policies on tabletop manipulation tasks.</itunes:summary>
      <itunes:subtitle>Releases a benchmark suite for systematically evaluating memory in Vision-Language-Action policies on tabletop manipulation tasks.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/5c955d26/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>PointWorld: Scaling 3D World Models for In-The-Wild Robotic Manipulation</title>
      <itunes:title>PointWorld: Scaling 3D World Models for In-The-Wild Robotic Manipulation</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">167548be-0a25-4c6d-b404-a2a587cf7b8c</guid>
      <link>https://share.transistor.fm/s/0bd8cbe5</link>
      <description>
        <![CDATA[Introduces large-scale 3D world models pretrained on diverse real-world video to enable robust robotic manipulation policies that generalize beyond simulation.]]>
      </description>
      <content:encoded>
        <![CDATA[Introduces large-scale 3D world models pretrained on diverse real-world video to enable robust robotic manipulation policies that generalize beyond simulation.]]>
      </content:encoded>
      <pubDate>Mon, 25 May 2026 14:12:29 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/0bd8cbe5/8ca2fee7.mp3" length="28581376" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1787</itunes:duration>
      <itunes:summary>Introduces large-scale 3D world models pretrained on diverse real-world video to enable robust robotic manipulation policies that generalize beyond simulation.</itunes:summary>
      <itunes:subtitle>Introduces large-scale 3D world models pretrained on diverse real-world video to enable robust robotic manipulation policies that generalize beyond simulation.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/0bd8cbe5/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Bimanual Pegboard Manipulation: A Benchmark for Vision-Language-Action Models</title>
      <itunes:title>Bimanual Pegboard Manipulation: A Benchmark for Vision-Language-Action Models</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">97608ac4-dd59-4d49-a7b5-0814985d467a</guid>
      <link>https://share.transistor.fm/s/8b3ede31</link>
      <description>
        <![CDATA[New LeRobot-based bimanual pegboard manipulation dataset with 52 episodes, 30k frames, 3 camera views, and 14-DOF arms for VLA evaluation. Provides standardized benchmark for vision-language-action model assessment.]]>
      </description>
      <content:encoded>
        <![CDATA[New LeRobot-based bimanual pegboard manipulation dataset with 52 episodes, 30k frames, 3 camera views, and 14-DOF arms for VLA evaluation. Provides standardized benchmark for vision-language-action model assessment.]]>
      </content:encoded>
      <pubDate>Sun, 24 May 2026 14:14:31 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/8b3ede31/7aaf46a5.mp3" length="26024960" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1627</itunes:duration>
      <itunes:summary>New LeRobot-based bimanual pegboard manipulation dataset with 52 episodes, 30k frames, 3 camera views, and 14-DOF arms for VLA evaluation. Provides standardized benchmark for vision-language-action model assessment.</itunes:summary>
      <itunes:subtitle>New LeRobot-based bimanual pegboard manipulation dataset with 52 episodes, 30k frames, 3 camera views, and 14-DOF arms for VLA evaluation. Provides standardized benchmark for vision-language-action model assessment.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/8b3ede31/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>FutureSim: Replaying Real-World Events to Evaluate AI Forecasting Agents</title>
      <itunes:title>FutureSim: Replaying Real-World Events to Evaluate AI Forecasting Agents</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">4dff901a-e0b1-43fa-92ac-1b9012bd7a75</guid>
      <link>https://share.transistor.fm/s/736da7bf</link>
      <description>
        <![CDATA[A benchmark designed to test AI models' capabilities in making accurate 3-month future predictions.]]>
      </description>
      <content:encoded>
        <![CDATA[A benchmark designed to test AI models' capabilities in making accurate 3-month future predictions.]]>
      </content:encoded>
      <pubDate>Sun, 24 May 2026 05:31:47 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/736da7bf/28f0fd02.mp3" length="26113536" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1633</itunes:duration>
      <itunes:summary>A benchmark designed to test AI models' capabilities in making accurate 3-month future predictions.</itunes:summary>
      <itunes:subtitle>A benchmark designed to test AI models' capabilities in making accurate 3-month future predictions.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/736da7bf/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>AgentFloor: A Benchmark for Long-Horizon Agent Planning</title>
      <itunes:title>AgentFloor: A Benchmark for Long-Horizon Agent Planning</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">419937ba-6efb-4318-a046-dd63d206ef3f</guid>
      <link>https://share.transistor.fm/s/2ac97e12</link>
      <description>
        <![CDATA[A 30-task benchmark for evaluating long-horizon planning capabilities across 16 different AI models.]]>
      </description>
      <content:encoded>
        <![CDATA[A 30-task benchmark for evaluating long-horizon planning capabilities across 16 different AI models.]]>
      </content:encoded>
      <pubDate>Sun, 24 May 2026 05:18:22 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/2ac97e12/4d54f31d.mp3" length="33726464" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2108</itunes:duration>
      <itunes:summary>A 30-task benchmark for evaluating long-horizon planning capabilities across 16 different AI models.</itunes:summary>
      <itunes:subtitle>A 30-task benchmark for evaluating long-horizon planning capabilities across 16 different AI models.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/2ac97e12/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>AlexNet: The Deep Convolutional Network That Transformed Vision</title>
      <itunes:title>AlexNet: The Deep Convolutional Network That Transformed Vision</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">039c64ba-4c44-4d6f-b256-b3cdedb2773a</guid>
      <link>https://share.transistor.fm/s/6c9794f7</link>
      <description>
        <![CDATA[AlexNet paper that sparked the modern deep learning revolution through convolutional neural networks.]]>
      </description>
      <content:encoded>
        <![CDATA[AlexNet paper that sparked the modern deep learning revolution through convolutional neural networks.]]>
      </content:encoded>
      <pubDate>Sat, 23 May 2026 14:21:12 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/6c9794f7/462a1eb8.mp3" length="40056320" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2504</itunes:duration>
      <itunes:summary>AlexNet paper that sparked the modern deep learning revolution through convolutional neural networks.</itunes:summary>
      <itunes:subtitle>AlexNet paper that sparked the modern deep learning revolution through convolutional neural networks.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/6c9794f7/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>A Few Useful Things to Know About Machine Learning</title>
      <itunes:title>A Few Useful Things to Know About Machine Learning</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">6a6f6bc8-57f4-4b2f-8dd8-d29a13b498f4</guid>
      <link>https://share.transistor.fm/s/56578679</link>
      <description>
        <![CDATA[Practical insights into ML pitfalls and best practices for machine learning practitioners.]]>
      </description>
      <content:encoded>
        <![CDATA[Practical insights into ML pitfalls and best practices for machine learning practitioners.]]>
      </content:encoded>
      <pubDate>Sat, 23 May 2026 14:13:24 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/56578679/63de6b0b.mp3" length="44003328" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2751</itunes:duration>
      <itunes:summary>Practical insights into ML pitfalls and best practices for machine learning practitioners.</itunes:summary>
      <itunes:subtitle>Practical insights into ML pitfalls and best practices for machine learning practitioners.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/56578679/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>SimToolReal: A Universal Dexterous Tool-Use Policy</title>
      <itunes:title>SimToolReal: A Universal Dexterous Tool-Use Policy</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">d0ee6b19-ef74-49fa-923d-8f3ca5d4c2a4</guid>
      <link>https://share.transistor.fm/s/942be2ec</link>
      <description>
        <![CDATA[Introduces an object-centric sim-to-real policy that enables zero-shot dexterous tool use on physical robots without task-specific fine-tuning. Leverages simulation data for robust real-world transfer.]]>
      </description>
      <content:encoded>
        <![CDATA[Introduces an object-centric sim-to-real policy that enables zero-shot dexterous tool use on physical robots without task-specific fine-tuning. Leverages simulation data for robust real-world transfer.]]>
      </content:encoded>
      <pubDate>Sat, 23 May 2026 05:26:37 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/942be2ec/787e2f72.mp3" length="27375104" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1711</itunes:duration>
      <itunes:summary>Introduces an object-centric sim-to-real policy that enables zero-shot dexterous tool use on physical robots without task-specific fine-tuning. Leverages simulation data for robust real-world transfer.</itunes:summary>
      <itunes:subtitle>Introduces an object-centric sim-to-real policy that enables zero-shot dexterous tool use on physical robots without task-specific fine-tuning. Leverages simulation data for robust real-world transfer.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/942be2ec/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Mimic-Video: Learning Physics Priors from Web-Scale Video for Robot Dexterity</title>
      <itunes:title>Mimic-Video: Learning Physics Priors from Web-Scale Video for Robot Dexterity</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">ac0bb1d8-4e8f-48a8-8323-f9ffde7905d2</guid>
      <link>https://share.transistor.fm/s/3fd221af</link>
      <description>
        <![CDATA[Pretrains robot policies on large-scale web video to acquire dynamics and physics understanding instead of static images or VLMs. Yields faster training, better generalization, and superior dexterous manipulation results in real-world tasks.]]>
      </description>
      <content:encoded>
        <![CDATA[Pretrains robot policies on large-scale web video to acquire dynamics and physics understanding instead of static images or VLMs. Yields faster training, better generalization, and superior dexterous manipulation results in real-world tasks.]]>
      </content:encoded>
      <pubDate>Sat, 23 May 2026 05:13:14 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/3fd221af/41699fe0.mp3" length="28073472" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1755</itunes:duration>
      <itunes:summary>Pretrains robot policies on large-scale web video to acquire dynamics and physics understanding instead of static images or VLMs. Yields faster training, better generalization, and superior dexterous manipulation results in real-world tasks.</itunes:summary>
      <itunes:subtitle>Pretrains robot policies on large-scale web video to acquire dynamics and physics understanding instead of static images or VLMs. Yields faster training, better generalization, and superior dexterous manipulation results in real-world tasks.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/3fd221af/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Deep Residual Learning for Image Recognition (ResNet)</title>
      <itunes:title>Deep Residual Learning for Image Recognition (ResNet)</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">24be6699-b3fd-46e8-9cbf-da97a12f70fb</guid>
      <link>https://share.transistor.fm/s/79238f09</link>
      <description>
        <![CDATA[Introduced residual connections (ResNet) enabling training of very deep networks, still widely used in modern architectures.]]>
      </description>
      <content:encoded>
        <![CDATA[Introduced residual connections (ResNet) enabling training of very deep networks, still widely used in modern architectures.]]>
      </content:encoded>
      <pubDate>Sat, 23 May 2026 02:03:20 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/79238f09/eb01621f.mp3" length="23104512" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1444</itunes:duration>
      <itunes:summary>Introduced residual connections (ResNet) enabling training of very deep networks, still widely used in modern architectures.</itunes:summary>
      <itunes:subtitle>Introduced residual connections (ResNet) enabling training of very deep networks, still widely used in modern architectures.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/79238f09/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Attention Is All You Need – The Transformer Revolution</title>
      <itunes:title>Attention Is All You Need – The Transformer Revolution</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">446a4d0e-f1b9-406d-a371-9892e1e46fd8</guid>
      <link>https://share.transistor.fm/s/f4710e40</link>
      <description>
        <![CDATA[Introduced the Transformer architecture based purely on attention mechanisms, becoming the foundation of nearly all modern large language models.]]>
      </description>
      <content:encoded>
        <![CDATA[Introduced the Transformer architecture based purely on attention mechanisms, becoming the foundation of nearly all modern large language models.]]>
      </content:encoded>
      <pubDate>Sat, 23 May 2026 01:49:20 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/f4710e40/368001f7.mp3" length="25709568" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1607</itunes:duration>
      <itunes:summary>Introduced the Transformer architecture based purely on attention mechanisms, becoming the foundation of nearly all modern large language models.</itunes:summary>
      <itunes:subtitle>Introduced the Transformer architecture based purely on attention mechanisms, becoming the foundation of nearly all modern large language models.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/f4710e40/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>NVIDIA Cosmos: World Foundation Models for Physical AI</title>
      <itunes:title>NVIDIA Cosmos: World Foundation Models for Physical AI</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">f4663b4b-9254-4b8a-90b3-7823b945eeba</guid>
      <link>https://share.transistor.fm/s/269bf69e</link>
      <description>
        <![CDATA[World foundation models for video and physics prediction with SynthID watermarking for responsible AI practices. Developed in collaboration with Google DeepMind.]]>
      </description>
      <content:encoded>
        <![CDATA[World foundation models for video and physics prediction with SynthID watermarking for responsible AI practices. Developed in collaboration with Google DeepMind.]]>
      </content:encoded>
      <pubDate>Wed, 20 May 2026 05:19:16 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/269bf69e/fc3b37f4.mp3" length="28205568" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1763</itunes:duration>
      <itunes:summary>World foundation models for video and physics prediction with SynthID watermarking for responsible AI practices. Developed in collaboration with Google DeepMind.</itunes:summary>
      <itunes:subtitle>World foundation models for video and physics prediction with SynthID watermarking for responsible AI practices. Developed in collaboration with Google DeepMind.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/269bf69e/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>LATENT: Teaching a Humanoid to Play Tennis from Imperfect Data</title>
      <itunes:title>LATENT: Teaching a Humanoid to Play Tennis from Imperfect Data</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">4ce42358-777b-4814-880c-e228b3f9eba7</guid>
      <link>https://share.transistor.fm/s/3e75beb4</link>
      <description>
        <![CDATA[Introduces a three-stage pipeline that extracts a latent action space from noisy, low-quality human motion capture, then trains a high-level RL policy in simulation to compose and execute dynamic whole-body tennis skills. Achieves volleys at human-level performance on a humanoid robot.]]>
      </description>
      <content:encoded>
        <![CDATA[Introduces a three-stage pipeline that extracts a latent action space from noisy, low-quality human motion capture, then trains a high-level RL policy in simulation to compose and execute dynamic whole-body tennis skills. Achieves volleys at human-level performance on a humanoid robot.]]>
      </content:encoded>
      <pubDate>Tue, 19 May 2026 14:11:09 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/3e75beb4/a51deea1.mp3" length="19640832" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1224</itunes:duration>
      <itunes:summary>Introduces a three-stage pipeline that extracts a latent action space from noisy, low-quality human motion capture, then trains a high-level RL policy in simulation to compose and execute dynamic whole-body tennis skills. Achieves volleys at human-level performance on a humanoid robot.</itunes:summary>
      <itunes:subtitle>Introduces a three-stage pipeline that extracts a latent action space from noisy, low-quality human motion capture, then trains a high-level RL policy in simulation to compose and execute dynamic whole-body tennis skills. Achieves volleys at human-level p</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/3e75beb4/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>CollabVR: Collaborative Video Reasoning with Vision-Language and Video Generation Models</title>
      <itunes:title>CollabVR: Collaborative Video Reasoning with Vision-Language and Video Generation Models</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">79280c69-4381-4416-be56-53ef1ead997b</guid>
      <link>https://share.transistor.fm/s/607ce3d3</link>
      <description>
        <![CDATA[Closed-loop framework coupling Vision-Language Models with Video Generation Models at step-level granularity. Mitigates long-horizon drift and mid-clip errors in goal-directed video reasoning for robotic planning.]]>
      </description>
      <content:encoded>
        <![CDATA[Closed-loop framework coupling Vision-Language Models with Video Generation Models at step-level granularity. Mitigates long-horizon drift and mid-clip errors in goal-directed video reasoning for robotic planning.]]>
      </content:encoded>
      <pubDate>Tue, 19 May 2026 05:26:22 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/607ce3d3/e8308065.mp3" length="40468992" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2530</itunes:duration>
      <itunes:summary>Closed-loop framework coupling Vision-Language Models with Video Generation Models at step-level granularity. Mitigates long-horizon drift and mid-clip errors in goal-directed video reasoning for robotic planning.</itunes:summary>
      <itunes:subtitle>Closed-loop framework coupling Vision-Language Models with Video Generation Models at step-level granularity. Mitigates long-horizon drift and mid-clip errors in goal-directed video reasoning for robotic planning.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/607ce3d3/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>World Action Models: The Next Frontier in Embodied AI</title>
      <itunes:title>World Action Models: The Next Frontier in Embodied AI</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">b98bbf87-900e-4716-88b1-4df90a7be135</guid>
      <link>https://share.transistor.fm/s/8a029825</link>
      <description>
        <![CDATA[First systematic survey defining World Action Models (WAMs) as embodied foundation models that jointly predict future states and generate actions. Covers architectures, data ecosystems, and evaluation protocols.]]>
      </description>
      <content:encoded>
        <![CDATA[First systematic survey defining World Action Models (WAMs) as embodied foundation models that jointly predict future states and generate actions. Covers architectures, data ecosystems, and evaluation protocols.]]>
      </content:encoded>
      <pubDate>Tue, 19 May 2026 05:10:48 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/8a029825/b198c1a9.mp3" length="34847744" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2178</itunes:duration>
      <itunes:summary>First systematic survey defining World Action Models (WAMs) as embodied foundation models that jointly predict future states and generate actions. Covers architectures, data ecosystems, and evaluation protocols.</itunes:summary>
      <itunes:subtitle>First systematic survey defining World Action Models (WAMs) as embodied foundation models that jointly predict future states and generate actions. Covers architectures, data ecosystems, and evaluation protocols.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/8a029825/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Training a Whole-Body Control Foundation Model</title>
      <itunes:title>Training a Whole-Body Control Foundation Model</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">20411ce3-9d18-4fa6-9c6d-2bb44bcf5f4f</guid>
      <link>https://share.transistor.fm/s/2fdd0e1a</link>
      <description>
        <![CDATA[Describes end-to-end learning of a foundation model for adaptive whole-body humanoid control via massive simulation variation. Combines proprioceptive perception and policy adaptation across embodiments.]]>
      </description>
      <content:encoded>
        <![CDATA[Describes end-to-end learning of a foundation model for adaptive whole-body humanoid control via massive simulation variation. Combines proprioceptive perception and policy adaptation across embodiments.]]>
      </content:encoded>
      <pubDate>Mon, 18 May 2026 14:26:59 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/2fdd0e1a/2447023c.mp3" length="38027264" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2377</itunes:duration>
      <itunes:summary>Describes end-to-end learning of a foundation model for adaptive whole-body humanoid control via massive simulation variation. Combines proprioceptive perception and policy adaptation across embodiments.</itunes:summary>
      <itunes:subtitle>Describes end-to-end learning of a foundation model for adaptive whole-body humanoid control via massive simulation variation. Combines proprioceptive perception and policy adaptation across embodiments.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/2fdd0e1a/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>DexJoCo: A Unified Benchmark for Task-Oriented Dexterous Manipulation</title>
      <itunes:title>DexJoCo: A Unified Benchmark for Task-Oriented Dexterous Manipulation</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">b12daefa-dad0-4434-be73-6fdc5a577a45</guid>
      <link>https://share.transistor.fm/s/ad2d5606</link>
      <description>
        <![CDATA[Releases an open-source MuJoCo-based benchmark with 11 dexterous tasks, low-cost teleoperation hardware, and 1.1K human demonstrations. Designed to evaluate and train modern VLA/robotic policies.]]>
      </description>
      <content:encoded>
        <![CDATA[Releases an open-source MuJoCo-based benchmark with 11 dexterous tasks, low-cost teleoperation hardware, and 1.1K human demonstrations. Designed to evaluate and train modern VLA/robotic policies.]]>
      </content:encoded>
      <pubDate>Mon, 18 May 2026 14:11:24 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/ad2d5606/ce71c45e.mp3" length="41876992" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2618</itunes:duration>
      <itunes:summary>Releases an open-source MuJoCo-based benchmark with 11 dexterous tasks, low-cost teleoperation hardware, and 1.1K human demonstrations. Designed to evaluate and train modern VLA/robotic policies.</itunes:summary>
      <itunes:subtitle>Releases an open-source MuJoCo-based benchmark with 11 dexterous tasks, low-cost teleoperation hardware, and 1.1K human demonstrations. Designed to evaluate and train modern VLA/robotic policies.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/ad2d5606/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>MMSkills: Building Multimodal Skill Libraries for Visual Agents</title>
      <itunes:title>MMSkills: Building Multimodal Skill Libraries for Visual Agents</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">88dfd7cf-59b0-439d-8636-a991031bfda3</guid>
      <link>https://share.transistor.fm/s/db06d808</link>
      <description>
        <![CDATA[Skill library, demonstrations, and dataset for multi-modal robotic skill learning and manipulation tasks.]]>
      </description>
      <content:encoded>
        <![CDATA[Skill library, demonstrations, and dataset for multi-modal robotic skill learning and manipulation tasks.]]>
      </content:encoded>
      <pubDate>Mon, 18 May 2026 05:29:29 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/db06d808/206fcdae.mp3" length="18767360" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1173</itunes:duration>
      <itunes:summary>Skill library, demonstrations, and dataset for multi-modal robotic skill learning and manipulation tasks.</itunes:summary>
      <itunes:subtitle>Skill library, demonstrations, and dataset for multi-modal robotic skill learning and manipulation tasks.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/db06d808/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>PhysBrain 1.0 VLA (TwinBrainVLA): Dual-Brain Vision-Language-Action with Physics-Grounded Learning</title>
      <itunes:title>PhysBrain 1.0 VLA (TwinBrainVLA): Dual-Brain Vision-Language-Action with Physics-Grounded Learning</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">c08d54ff-9afc-47cc-95f2-fb270881917d</guid>
      <link>https://share.transistor.fm/s/6aa88d4d</link>
      <description>
        <![CDATA[Introduces dual-brain fusion Vision-Language-Action model with LangForce physics-grounded training methodology.]]>
      </description>
      <content:encoded>
        <![CDATA[Introduces dual-brain fusion Vision-Language-Action model with LangForce physics-grounded training methodology.]]>
      </content:encoded>
      <pubDate>Mon, 18 May 2026 05:16:17 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/6aa88d4d/7dabe48e.mp3" length="24930304" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1559</itunes:duration>
      <itunes:summary>Introduces dual-brain fusion Vision-Language-Action model with LangForce physics-grounded training methodology.</itunes:summary>
      <itunes:subtitle>Introduces dual-brain fusion Vision-Language-Action model with LangForce physics-grounded training methodology.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/6aa88d4d/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>MolmoAct2-LIBERO: An Open Vision-Language-Action Model for Robotics</title>
      <itunes:title>MolmoAct2-LIBERO: An Open Vision-Language-Action Model for Robotics</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">75c3179f-f9bc-4b8b-82f8-bbaf530cd2f9</guid>
      <link>https://share.transistor.fm/s/6ad08ac4</link>
      <description>
        <![CDATA[Vision-Language-Action (VLA) model fine-tuned on the merged LIBERO robotics dataset (1,693 episodes, 273k+ frames) achieving 98.25% success rate on manipulation tasks. Released with both checkpoint and dataset for VLA finetuning.]]>
      </description>
      <content:encoded>
        <![CDATA[Vision-Language-Action (VLA) model fine-tuned on the merged LIBERO robotics dataset (1,693 episodes, 273k+ frames) achieving 98.25% success rate on manipulation tasks. Released with both checkpoint and dataset for VLA finetuning.]]>
      </content:encoded>
      <pubDate>Sun, 17 May 2026 14:24:27 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/6ad08ac4/8b47b219.mp3" length="37281792" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2331</itunes:duration>
      <itunes:summary>Vision-Language-Action (VLA) model fine-tuned on the merged LIBERO robotics dataset (1,693 episodes, 273k+ frames) achieving 98.25% success rate on manipulation tasks. Released with both checkpoint and dataset for VLA finetuning.</itunes:summary>
      <itunes:subtitle>Vision-Language-Action (VLA) model fine-tuned on the merged LIBERO robotics dataset (1,693 episodes, 273k+ frames) achieving 98.25% success rate on manipulation tasks. Released with both checkpoint and dataset for VLA finetuning.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/6ad08ac4/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>SANA-WM: Efficient Minute-Scale World Modeling with Hybrid Diffusion Transformers</title>
      <itunes:title>SANA-WM: Efficient Minute-Scale World Modeling with Hybrid Diffusion Transformers</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">8cec6624-d584-4acc-a3c8-e582d2d3cd4a</guid>
      <link>https://share.transistor.fm/s/e790dc4d</link>
      <description>
        <![CDATA[A 2.6B-parameter open-source world model that generates coherent 720p, minute-long videos with precise 6-DoF camera control on a single GPU using a Hybrid Linear Diffusion Transformer + Gated DeltaNet for long-context efficiency. Targets controllable physics simulation.]]>
      </description>
      <content:encoded>
        <![CDATA[A 2.6B-parameter open-source world model that generates coherent 720p, minute-long videos with precise 6-DoF camera control on a single GPU using a Hybrid Linear Diffusion Transformer + Gated DeltaNet for long-context efficiency. Targets controllable physics simulation.]]>
      </content:encoded>
      <pubDate>Sun, 17 May 2026 14:12:10 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/e790dc4d/0fd84238.mp3" length="19662848" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1229</itunes:duration>
      <itunes:summary>A 2.6B-parameter open-source world model that generates coherent 720p, minute-long videos with precise 6-DoF camera control on a single GPU using a Hybrid Linear Diffusion Transformer + Gated DeltaNet for long-context efficiency. Targets controllable physics simulation.</itunes:summary>
      <itunes:subtitle>A 2.6B-parameter open-source world model that generates coherent 720p, minute-long videos with precise 6-DoF camera control on a single GPU using a Hybrid Linear Diffusion Transformer + Gated DeltaNet for long-context efficiency. Targets controllable phys</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/e790dc4d/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>WildClawBench: A Real-World, Long-Horizon Benchmark for AI Agents</title>
      <itunes:title>WildClawBench: A Real-World, Long-Horizon Benchmark for AI Agents</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">4ad3a4d0-def6-4f83-a00f-3f045414d5d8</guid>
      <link>https://share.transistor.fm/s/c79ac63b</link>
      <description>
        <![CDATA[New benchmark and dataset for robotic manipulation in unconstrained 'wild' environments. Includes standardized containers, leaderboards, and evaluation protocols for cross-embodiment policies.]]>
      </description>
      <content:encoded>
        <![CDATA[New benchmark and dataset for robotic manipulation in unconstrained 'wild' environments. Includes standardized containers, leaderboards, and evaluation protocols for cross-embodiment policies.]]>
      </content:encoded>
      <pubDate>Sun, 17 May 2026 05:24:48 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/c79ac63b/dc55c217.mp3" length="30854656" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1929</itunes:duration>
      <itunes:summary>New benchmark and dataset for robotic manipulation in unconstrained 'wild' environments. Includes standardized containers, leaderboards, and evaluation protocols for cross-embodiment policies.</itunes:summary>
      <itunes:subtitle>New benchmark and dataset for robotic manipulation in unconstrained 'wild' environments. Includes standardized containers, leaderboards, and evaluation protocols for cross-embodiment policies.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/c79ac63b/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>MCP-Cosmos: Bring Your Own World Model</title>
      <itunes:title>MCP-Cosmos: Bring Your Own World Model</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">603bc2e7-d27c-4f27-973a-6550c783d135</guid>
      <link>https://share.transistor.fm/s/868c73f1</link>
      <description>
        <![CDATA[Introduces a latent-space world model framework that lets agents simulate state transitions and iteratively refine plans before real-world execution. Evaluated on 20+ MCP-Bench tasks with measurable gains in tool-use success.]]>
      </description>
      <content:encoded>
        <![CDATA[Introduces a latent-space world model framework that lets agents simulate state transitions and iteratively refine plans before real-world execution. Evaluated on 20+ MCP-Bench tasks with measurable gains in tool-use success.]]>
      </content:encoded>
      <pubDate>Sun, 17 May 2026 05:15:11 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/868c73f1/81612bf2.mp3" length="23375872" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1461</itunes:duration>
      <itunes:summary>Introduces a latent-space world model framework that lets agents simulate state transitions and iteratively refine plans before real-world execution. Evaluated on 20+ MCP-Bench tasks with measurable gains in tool-use success.</itunes:summary>
      <itunes:subtitle>Introduces a latent-space world model framework that lets agents simulate state transitions and iteratively refine plans before real-world execution. Evaluated on 20+ MCP-Bench tasks with measurable gains in tool-use success.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/868c73f1/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>OpenAI o1: Teaching LLMs to Think Slow and Deep</title>
      <itunes:title>OpenAI o1: Teaching LLMs to Think Slow and Deep</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">0611a4c8-f201-4107-a541-017ceb847c80</guid>
      <link>https://share.transistor.fm/s/b683fcc7</link>
      <description>
        <![CDATA[Details OpenAI's reasoning-focused o1 model and its 'long thought' approach using test-time compute scaling. Explores how extended reasoning during inference can improve model performance on complex tasks.]]>
      </description>
      <content:encoded>
        <![CDATA[Details OpenAI's reasoning-focused o1 model and its 'long thought' approach using test-time compute scaling. Explores how extended reasoning during inference can improve model performance on complex tasks.]]>
      </content:encoded>
      <pubDate>Sat, 16 May 2026 18:41:07 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/b683fcc7/c4238528.mp3" length="13542400" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>847</itunes:duration>
      <itunes:summary>Details OpenAI's reasoning-focused o1 model and its 'long thought' approach using test-time compute scaling. Explores how extended reasoning during inference can improve model performance on complex tasks.</itunes:summary>
      <itunes:subtitle>Details OpenAI's reasoning-focused o1 model and its 'long thought' approach using test-time compute scaling. Explores how extended reasoning during inference can improve model performance on complex tasks.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/b683fcc7/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>The Llama 3 Herd of Models</title>
      <itunes:title>The Llama 3 Herd of Models</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">45bda338-8e25-4dea-98fe-662d26140075</guid>
      <link>https://share.transistor.fm/s/a3fd9a57</link>
      <description>
        <![CDATA[Comprehensive technical report on the Llama 3 family, covering architecture, training at scale, multimodal extensions, and real-world impact. Details the development of Meta's flagship open-source language model series.]]>
      </description>
      <content:encoded>
        <![CDATA[Comprehensive technical report on the Llama 3 family, covering architecture, training at scale, multimodal extensions, and real-world impact. Details the development of Meta's flagship open-source language model series.]]>
      </content:encoded>
      <pubDate>Sat, 16 May 2026 18:32:22 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/a3fd9a57/c49d32d6.mp3" length="31372288" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1961</itunes:duration>
      <itunes:summary>Comprehensive technical report on the Llama 3 family, covering architecture, training at scale, multimodal extensions, and real-world impact. Details the development of Meta's flagship open-source language model series.</itunes:summary>
      <itunes:subtitle>Comprehensive technical report on the Llama 3 family, covering architecture, training at scale, multimodal extensions, and real-world impact. Details the development of Meta's flagship open-source language model series.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/a3fd9a57/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>LATENT: Learning Athletic Humanoid Tennis Skills from Imperfect Human Motion Data</title>
      <itunes:title>LATENT: Learning Athletic Humanoid Tennis Skills from Imperfect Human Motion Data</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">453b6350-9ef6-44e8-a98d-49778e45e0a0</guid>
      <link>https://share.transistor.fm/s/d3dcf780</link>
      <description>
        <![CDATA[Introduces a three-stage pipeline that extracts a latent action space from low-quality human tennis demonstrations, then trains a high-level policy in simulation via reinforcement learning. Enables dynamic whole-body humanoid tennis play with back-and-forth volleys at human level.]]>
      </description>
      <content:encoded>
        <![CDATA[Introduces a three-stage pipeline that extracts a latent action space from low-quality human tennis demonstrations, then trains a high-level policy in simulation via reinforcement learning. Enables dynamic whole-body humanoid tennis play with back-and-forth volleys at human level.]]>
      </content:encoded>
      <pubDate>Sat, 16 May 2026 18:03:21 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/d3dcf780/d13e0605.mp3" length="30209024" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1889</itunes:duration>
      <itunes:summary>Introduces a three-stage pipeline that extracts a latent action space from low-quality human tennis demonstrations, then trains a high-level policy in simulation via reinforcement learning. Enables dynamic whole-body humanoid tennis play with back-and-forth volleys at human level.</itunes:summary>
      <itunes:subtitle>Introduces a three-stage pipeline that extracts a latent action space from low-quality human tennis demonstrations, then trains a high-level policy in simulation via reinforcement learning. Enables dynamic whole-body humanoid tennis play with back-and-for</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/d3dcf780/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>AnyFlow: Any-Step Video Diffusion for Predictive World Modeling</title>
      <itunes:title>AnyFlow: Any-Step Video Diffusion for Predictive World Modeling</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">60f34baa-12ea-49bc-ad50-2ae1cb55d630</guid>
      <link>https://share.transistor.fm/s/57e580f5</link>
      <description>
        <![CDATA[First any-step video diffusion framework using flow maps, allowing a single model to adapt to arbitrary inference budgets for scalable high-quality video generation relevant to predictive world modeling.]]>
      </description>
      <content:encoded>
        <![CDATA[First any-step video diffusion framework using flow maps, allowing a single model to adapt to arbitrary inference budgets for scalable high-quality video generation relevant to predictive world modeling.]]>
      </content:encoded>
      <pubDate>Thu, 14 May 2026 16:13:25 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/57e580f5/0e0e2059.mp3" length="13033472" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>813</itunes:duration>
      <itunes:summary>First any-step video diffusion framework using flow maps, allowing a single model to adapt to arbitrary inference budgets for scalable high-quality video generation relevant to predictive world modeling.</itunes:summary>
      <itunes:subtitle>First any-step video diffusion framework using flow maps, allowing a single model to adapt to arbitrary inference budgets for scalable high-quality video generation relevant to predictive world modeling.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/57e580f5/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title># Robotics: The Endgame</title>
      <itunes:title># Robotics: The Endgame</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">6d57df70-7f6a-4b2e-8753-641f7c0119a8</guid>
      <link>https://share.transistor.fm/s/c4319104</link>
      <description>
        <![CDATA[Technical roadmap mirroring LLM scaling: critiques VLAs, advocates video world models as second pretraining phase, introduces World Action Models (WAM), manipulation data flywheels, EgoScale with new Dexterity Scaling Law, and DreamDojo end-to-end neural physics engine for sim RL.]]>
      </description>
      <content:encoded>
        <![CDATA[Technical roadmap mirroring LLM scaling: critiques VLAs, advocates video world models as second pretraining phase, introduces World Action Models (WAM), manipulation data flywheels, EgoScale with new Dexterity Scaling Law, and DreamDojo end-to-end neural physics engine for sim RL.]]>
      </content:encoded>
      <pubDate>Thu, 14 May 2026 16:02:02 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/c4319104/ee64c916.mp3" length="32887296" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2056</itunes:duration>
      <itunes:summary>Technical roadmap mirroring LLM scaling: critiques VLAs, advocates video world models as second pretraining phase, introduces World Action Models (WAM), manipulation data flywheels, EgoScale with new Dexterity Scaling Law, and DreamDojo end-to-end neural physics engine for sim RL.</itunes:summary>
      <itunes:subtitle>Technical roadmap mirroring LLM scaling: critiques VLAs, advocates video world models as second pretraining phase, introduces World Action Models (WAM), manipulation data flywheels, EgoScale with new Dexterity Scaling Law, and DreamDojo end-to-end neural </itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/c4319104/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Claw-Eval: Toward Trustworthy and Transparent Evaluation of Autonomous Agents</title>
      <itunes:title>Claw-Eval: Toward Trustworthy and Transparent Evaluation of Autonomous Agents</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">656b6506-7b5f-45ec-9635-60c1ed1c9222</guid>
      <link>https://share.transistor.fm/s/b07d8712</link>
      <description>
        <![CDATA[Benchmark with 2,159 rubric items across 300 tasks using trajectory-aware grading and 3-trial Pass^3 scoring to mitigate luck. Evaluates agent reliability in real-world robotics settings.]]>
      </description>
      <content:encoded>
        <![CDATA[Benchmark with 2,159 rubric items across 300 tasks using trajectory-aware grading and 3-trial Pass^3 scoring to mitigate luck. Evaluates agent reliability in real-world robotics settings.]]>
      </content:encoded>
      <pubDate>Wed, 08 Apr 2026 07:19:18 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/b07d8712/7abd153f.mp3" length="27413504" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1714</itunes:duration>
      <itunes:summary>Benchmark with 2,159 rubric items across 300 tasks using trajectory-aware grading and 3-trial Pass^3 scoring to mitigate luck. Evaluates agent reliability in real-world robotics settings.</itunes:summary>
      <itunes:subtitle>Benchmark with 2,159 rubric items across 300 tasks using trajectory-aware grading and 3-trial Pass^3 scoring to mitigate luck. Evaluates agent reliability in real-world robotics settings.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/b07d8712/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>LIBERO-Para: Paraphrase Robustness in Robotic Manipulation</title>
      <itunes:title>LIBERO-Para: Paraphrase Robustness in Robotic Manipulation</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">554fb299-b0a3-484b-8806-31d6d0215697</guid>
      <link>https://share.transistor.fm/s/3b3ef07a</link>
      <description>
        <![CDATA[Reveals paraphrase fragility in VLAs causing 22-52% success drops due to task misidentification. Introduces PRIDE metric weighting success by paraphrase difficulty on LIBERO benchmark manipulation tasks.]]>
      </description>
      <content:encoded>
        <![CDATA[Reveals paraphrase fragility in VLAs causing 22-52% success drops due to task misidentification. Introduces PRIDE metric weighting success by paraphrase difficulty on LIBERO benchmark manipulation tasks.]]>
      </content:encoded>
      <pubDate>Wed, 08 Apr 2026 07:18:01 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/3b3ef07a/41910b0d.mp3" length="31015424" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1939</itunes:duration>
      <itunes:summary>Reveals paraphrase fragility in VLAs causing 22-52% success drops due to task misidentification. Introduces PRIDE metric weighting success by paraphrase difficulty on LIBERO benchmark manipulation tasks.</itunes:summary>
      <itunes:subtitle>Reveals paraphrase fragility in VLAs causing 22-52% success drops due to task misidentification. Introduces PRIDE metric weighting success by paraphrase difficulty on LIBERO benchmark manipulation tasks.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/3b3ef07a/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>YOR: Your Own Mobile Manipulator for Generalizable Robotics</title>
      <itunes:title>YOR: Your Own Mobile Manipulator for Generalizable Robotics</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">ce79c883-fb39-4231-bbf2-0db7f4a18274</guid>
      <link>https://share.transistor.fm/s/280d4189</link>
      <description>
        <![CDATA[Low-cost mobile manipulator design and training strategies for broad generalization in real-world tasks.]]>
      </description>
      <content:encoded>
        <![CDATA[Low-cost mobile manipulator design and training strategies for broad generalization in real-world tasks.]]>
      </content:encoded>
      <pubDate>Tue, 07 Apr 2026 07:41:37 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/280d4189/6f785221.mp3" length="26046976" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1628</itunes:duration>
      <itunes:summary>Low-cost mobile manipulator design and training strategies for broad generalization in real-world tasks.</itunes:summary>
      <itunes:subtitle>Low-cost mobile manipulator design and training strategies for broad generalization in real-world tasks.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/280d4189/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>EgoSim: Egocentric World Simulator for Embodied Interaction Generation</title>
      <itunes:title>EgoSim: Egocentric World Simulator for Embodied Interaction Generation</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">166c79b9-dcf9-494b-8f3d-8697ff0edbbe</guid>
      <link>https://share.transistor.fm/s/83741c31</link>
      <description>
        <![CDATA[Closed-loop egocentric video simulator maintaining persistent 3D scene state for consistent interactions, enabling cross-embodiment transfer from human videos to robotic manipulation.]]>
      </description>
      <content:encoded>
        <![CDATA[Closed-loop egocentric video simulator maintaining persistent 3D scene state for consistent interactions, enabling cross-embodiment transfer from human videos to robotic manipulation.]]>
      </content:encoded>
      <pubDate>Tue, 07 Apr 2026 07:29:11 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/83741c31/22d76f6d.mp3" length="48817664" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>3052</itunes:duration>
      <itunes:summary>Closed-loop egocentric video simulator maintaining persistent 3D scene state for consistent interactions, enabling cross-embodiment transfer from human videos to robotic manipulation.</itunes:summary>
      <itunes:subtitle>Closed-loop egocentric video simulator maintaining persistent 3D scene state for consistent interactions, enabling cross-embodiment transfer from human videos to robotic manipulation.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/83741c31/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Accelerating Video World Models: From Generative Videos to Real-Time Simulators</title>
      <itunes:title>Accelerating Video World Models: From Generative Videos to Real-Time Simulators</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">102fac41-7e57-4116-9f70-58e86599b147</guid>
      <link>https://share.transistor.fm/s/c9795299</link>
      <description>
        <![CDATA[Comprehensive survey taxonomizing efficient architectures/algorithms for video world models as simulators, targeting compute bottlenecks in embodied AI, autonomous driving, and games with techniques like short-window attention for real-time long-horizon prediction.]]>
      </description>
      <content:encoded>
        <![CDATA[Comprehensive survey taxonomizing efficient architectures/algorithms for video world models as simulators, targeting compute bottlenecks in embodied AI, autonomous driving, and games with techniques like short-window attention for real-time long-horizon prediction.]]>
      </content:encoded>
      <pubDate>Mon, 06 Apr 2026 22:17:58 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/c9795299/65189304.mp3" length="37860352" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2367</itunes:duration>
      <itunes:summary>Comprehensive survey taxonomizing efficient architectures/algorithms for video world models as simulators, targeting compute bottlenecks in embodied AI, autonomous driving, and games with techniques like short-window attention for real-time long-horizon prediction.</itunes:summary>
      <itunes:subtitle>Comprehensive survey taxonomizing efficient architectures/algorithms for video world models as simulators, targeting compute bottlenecks in embodied AI, autonomous driving, and games with techniques like short-window attention for real-time long-horizon p</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/c9795299/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>From Tokens to Thoughts: Continuous Latent Reasoning in Large Models and Robot Control</title>
      <itunes:title>From Tokens to Thoughts: Continuous Latent Reasoning in Large Models and Robot Control</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">124c43a2-1421-47b3-8ddf-8f3f5f60a3f2</guid>
      <link>https://share.transistor.fm/s/b414fdbe</link>
      <description>
        <![CDATA[Curated collection of 100+ works surveying shift to continuous latent spaces in LLMs/VLMs/VLAs for improved reasoning over discrete tokens, with relevance to robotics action modeling.]]>
      </description>
      <content:encoded>
        <![CDATA[Curated collection of 100+ works surveying shift to continuous latent spaces in LLMs/VLMs/VLAs for improved reasoning over discrete tokens, with relevance to robotics action modeling.]]>
      </content:encoded>
      <pubDate>Mon, 06 Apr 2026 22:14:05 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/b414fdbe/7a6edda5.mp3" length="25867264" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1617</itunes:duration>
      <itunes:summary>Curated collection of 100+ works surveying shift to continuous latent spaces in LLMs/VLMs/VLAs for improved reasoning over discrete tokens, with relevance to robotics action modeling.</itunes:summary>
      <itunes:subtitle>Curated collection of 100+ works surveying shift to continuous latent spaces in LLMs/VLMs/VLAs for improved reasoning over discrete tokens, with relevance to robotics action modeling.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/b414fdbe/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>CaP-X: Coding Agents for Physical eXecution</title>
      <itunes:title>CaP-X: Coding Agents for Physical eXecution</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">59353697-d362-41d4-83d5-c625eae5136f</guid>
      <link>https://share.transistor.fm/s/8da560f3</link>
      <description>
        <![CDATA[CaP-X is an open-source agentic robotics framework where LLMs/VLMs generate code to call perception and control APIs for execution across diverse simulated and real robots in CaP-Gym's 187 manipulation tasks. The framework includes CaP-Bench for evaluating frontier models and CaP-RL, which boosts a 7B model's success from 20% to 72% with minimal sim-to-real gap.]]>
      </description>
      <content:encoded>
        <![CDATA[CaP-X is an open-source agentic robotics framework where LLMs/VLMs generate code to call perception and control APIs for execution across diverse simulated and real robots in CaP-Gym's 187 manipulation tasks. The framework includes CaP-Bench for evaluating frontier models and CaP-RL, which boosts a 7B model's success from 20% to 72% with minimal sim-to-real gap.]]>
      </content:encoded>
      <pubDate>Mon, 06 Apr 2026 07:11:45 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/8da560f3/ea73e83e.mp3" length="13241344" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>828</itunes:duration>
      <itunes:summary>CaP-X is an open-source agentic robotics framework where LLMs/VLMs generate code to call perception and control APIs for execution across diverse simulated and real robots in CaP-Gym's 187 manipulation tasks. The framework includes CaP-Bench for evaluating frontier models and CaP-RL, which boosts a 7B model's success from 20% to 72% with minimal sim-to-real gap.</itunes:summary>
      <itunes:subtitle>CaP-X is an open-source agentic robotics framework where LLMs/VLMs generate code to call perception and control APIs for execution across diverse simulated and real robots in CaP-Gym's 187 manipulation tasks. The framework includes CaP-Bench for evaluatin</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/8da560f3/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>DoRA: Weight-Decomposed Low-Rank Adaptation</title>
      <itunes:title>DoRA: Weight-Decomposed Low-Rank Adaptation</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">a3426e24-29d7-4ab2-99cc-5a15f33a0871</guid>
      <link>https://share.transistor.fm/s/187113d5</link>
      <description>
        <![CDATA[An upgrade over LoRA for parameter-efficient fine-tuning, enabling better performance in LLMs by decomposing weights into magnitude and direction components.]]>
      </description>
      <content:encoded>
        <![CDATA[An upgrade over LoRA for parameter-efficient fine-tuning, enabling better performance in LLMs by decomposing weights into magnitude and direction components.]]>
      </content:encoded>
      <pubDate>Sun, 05 Apr 2026 22:30:22 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/187113d5/eead50a6.mp3" length="37735424" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2359</itunes:duration>
      <itunes:summary>An upgrade over LoRA for parameter-efficient fine-tuning, enabling better performance in LLMs by decomposing weights into magnitude and direction components.</itunes:summary>
      <itunes:subtitle>An upgrade over LoRA for parameter-efficient fine-tuning, enabling better performance in LLMs by decomposing weights into magnitude and direction components.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/187113d5/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>AI Model Collapse: What Happens When AI Trains on Its Own Outputs</title>
      <itunes:title>AI Model Collapse: What Happens When AI Trains on Its Own Outputs</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">8db30835-a2c0-441e-b9ce-396454b80bb7</guid>
      <link>https://share.transistor.fm/s/217b3d96</link>
      <description>
        <![CDATA[Seminal work showing how training on AI-generated data leads to 'model collapse' in neural networks, with urgent implications for future scaling.]]>
      </description>
      <content:encoded>
        <![CDATA[Seminal work showing how training on AI-generated data leads to 'model collapse' in neural networks, with urgent implications for future scaling.]]>
      </content:encoded>
      <pubDate>Sun, 05 Apr 2026 22:15:47 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/217b3d96/f97d546a.mp3" length="28228096" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1765</itunes:duration>
      <itunes:summary>Seminal work showing how training on AI-generated data leads to 'model collapse' in neural networks, with urgent implications for future scaling.</itunes:summary>
      <itunes:subtitle>Seminal work showing how training on AI-generated data leads to 'model collapse' in neural networks, with urgent implications for future scaling.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/217b3d96/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>PhAIL: Benchmarking Vision-Language-Action Models on Real-World Bin-Picking</title>
      <itunes:title>PhAIL: Benchmarking Vision-Language-Action Models on Real-World Bin-Picking</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">44f5296b-26ce-40c4-9d70-a9b820ac6990</guid>
      <link>https://share.transistor.fm/s/65151748</link>
      <description>
        <![CDATA[Real-world hardware evaluation of VLAs on blind bin-to-bin picking, achieving max 64 picks/hour across hundreds of runs, with full videos/data exposing gaps in production-scale robotic manipulation reliability.]]>
      </description>
      <content:encoded>
        <![CDATA[Real-world hardware evaluation of VLAs on blind bin-to-bin picking, achieving max 64 picks/hour across hundreds of runs, with full videos/data exposing gaps in production-scale robotic manipulation reliability.]]>
      </content:encoded>
      <pubDate>Sun, 05 Apr 2026 07:19:40 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/65151748/726bc139.mp3" length="32048128" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2003</itunes:duration>
      <itunes:summary>Real-world hardware evaluation of VLAs on blind bin-to-bin picking, achieving max 64 picks/hour across hundreds of runs, with full videos/data exposing gaps in production-scale robotic manipulation reliability.</itunes:summary>
      <itunes:subtitle>Real-world hardware evaluation of VLAs on blind bin-to-bin picking, achieving max 64 picks/hour across hundreds of runs, with full videos/data exposing gaps in production-scale robotic manipulation reliability.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/65151748/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Co-training Large Behavior Models: Data Modalities and Training Strategies for Robot Manipulation</title>
      <itunes:title>Co-training Large Behavior Models: Data Modalities and Training Strategies for Robot Manipulation</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">54ae145a-0dc2-4905-b5df-9e2f3f055e67</guid>
      <link>https://share.transistor.fm/s/468f608d</link>
      <description>
        <![CDATA[Comprehensive evaluation of 89 policies showing optimal co-training practices mixing real robot data with sim/egocentric human videos to boost diversity and performance in large robotics foundation models.]]>
      </description>
      <content:encoded>
        <![CDATA[Comprehensive evaluation of 89 policies showing optimal co-training practices mixing real robot data with sim/egocentric human videos to boost diversity and performance in large robotics foundation models.]]>
      </content:encoded>
      <pubDate>Sat, 04 Apr 2026 22:42:13 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/468f608d/4df549f8.mp3" length="27222016" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1702</itunes:duration>
      <itunes:summary>Comprehensive evaluation of 89 policies showing optimal co-training practices mixing real robot data with sim/egocentric human videos to boost diversity and performance in large robotics foundation models.</itunes:summary>
      <itunes:subtitle>Comprehensive evaluation of 89 policies showing optimal co-training practices mixing real robot data with sim/egocentric human videos to boost diversity and performance in large robotics foundation models.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/468f608d/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>HyDRA: Hybrid Memory for Dynamic Video World Models</title>
      <itunes:title>HyDRA: Hybrid Memory for Dynamic Video World Models</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">830666f3-360f-4393-9cf5-19da2472e8d9</guid>
      <link>https://share.transistor.fm/s/819bd53b</link>
      <description>
        <![CDATA[Novel memory system preserving dynamic object identity and motion continuity across occlusions in video world models, addressing frozen/vanishing issues for improved predictive physics in embodied AI.]]>
      </description>
      <content:encoded>
        <![CDATA[Novel memory system preserving dynamic object identity and motion continuity across occlusions in video world models, addressing frozen/vanishing issues for improved predictive physics in embodied AI.]]>
      </content:encoded>
      <pubDate>Sat, 04 Apr 2026 22:31:30 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/819bd53b/88249311.mp3" length="21002752" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1313</itunes:duration>
      <itunes:summary>Novel memory system preserving dynamic object identity and motion continuity across occlusions in video world models, addressing frozen/vanishing issues for improved predictive physics in embodied AI.</itunes:summary>
      <itunes:subtitle>Novel memory system preserving dynamic object identity and motion continuity across occlusions in video world models, addressing frozen/vanishing issues for improved predictive physics in embodied AI.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/819bd53b/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title># WildWorld: Dynamic World Modeling with Actions and Explicit State</title>
      <itunes:title># WildWorld: Dynamic World Modeling with Actions and Explicit State</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">de83c0ca-b268-4d67-96bb-4bf024e6e4d2</guid>
      <link>https://share.transistor.fm/s/3dc2a292</link>
      <description>
        <![CDATA[Massive dataset enabling dynamic world models with explicit states and actions, supporting predictive modeling for cross-embodiment robotic control.]]>
      </description>
      <content:encoded>
        <![CDATA[Massive dataset enabling dynamic world models with explicit states and actions, supporting predictive modeling for cross-embodiment robotic control.]]>
      </content:encoded>
      <pubDate>Sat, 04 Apr 2026 07:29:53 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/3dc2a292/eaede145.mp3" length="31422464" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1964</itunes:duration>
      <itunes:summary>Massive dataset enabling dynamic world models with explicit states and actions, supporting predictive modeling for cross-embodiment robotic control.</itunes:summary>
      <itunes:subtitle>Massive dataset enabling dynamic world models with explicit states and actions, supporting predictive modeling for cross-embodiment robotic control.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/3dc2a292/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Omni-WorldBench: Evaluating Interactive 4D World Models</title>
      <itunes:title>Omni-WorldBench: Evaluating Interactive 4D World Models</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">1108acde-3bc0-4fa0-acb0-08462f58b1c7</guid>
      <link>https://share.transistor.fm/s/d149c2a0</link>
      <description>
        <![CDATA[New benchmark assessing world models on interaction tasks, pushing predictive physics and video modeling towards robotics applications with action-conditioned evaluation.]]>
      </description>
      <content:encoded>
        <![CDATA[New benchmark assessing world models on interaction tasks, pushing predictive physics and video modeling towards robotics applications with action-conditioned evaluation.]]>
      </content:encoded>
      <pubDate>Sat, 04 Apr 2026 07:18:25 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/d149c2a0/0fd76ebe.mp3" length="38329856" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2396</itunes:duration>
      <itunes:summary>New benchmark assessing world models on interaction tasks, pushing predictive physics and video modeling towards robotics applications with action-conditioned evaluation.</itunes:summary>
      <itunes:subtitle>New benchmark assessing world models on interaction tasks, pushing predictive physics and video modeling towards robotics applications with action-conditioned evaluation.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/d149c2a0/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>SIMART: From Static Meshes to Sim-Ready Articulated Models</title>
      <itunes:title>SIMART: From Static Meshes to Sim-Ready Articulated Models</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">2c604135-308e-40ec-9444-93571b3c4cb5</guid>
      <link>https://share.transistor.fm/s/ee8eef07</link>
      <description>
        <![CDATA[Unified MLLM framework with Sparse 3D VQ-VAE (70% token reduction) for part-level mesh decomposition and kinematic chain prediction, enabling physics-based robotic simulation from monolithic assets.]]>
      </description>
      <content:encoded>
        <![CDATA[Unified MLLM framework with Sparse 3D VQ-VAE (70% token reduction) for part-level mesh decomposition and kinematic chain prediction, enabling physics-based robotic simulation from monolithic assets.]]>
      </content:encoded>
      <pubDate>Fri, 03 Apr 2026 22:37:59 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/ee8eef07/c71c94f8.mp3" length="36609024" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2289</itunes:duration>
      <itunes:summary>Unified MLLM framework with Sparse 3D VQ-VAE (70% token reduction) for part-level mesh decomposition and kinematic chain prediction, enabling physics-based robotic simulation from monolithic assets.</itunes:summary>
      <itunes:subtitle>Unified MLLM framework with Sparse 3D VQ-VAE (70% token reduction) for part-level mesh decomposition and kinematic chain prediction, enabling physics-based robotic simulation from monolithic assets.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/ee8eef07/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>EgoSim: An Egocentric World Simulator for Embodied Interaction</title>
      <itunes:title>EgoSim: An Egocentric World Simulator for Embodied Interaction</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">b1d5a5c2-37cf-4dee-955a-76a9b6b20faa</guid>
      <link>https://share.transistor.fm/s/8a1e2dd1</link>
      <description>
        <![CDATA[Closed-loop egocentric simulator persistently updating 3D scene state to generate spatially consistent interaction videos for continuous simulation, enabling cross-embodiment transfer from human videos to robotic manipulation tasks.]]>
      </description>
      <content:encoded>
        <![CDATA[Closed-loop egocentric simulator persistently updating 3D scene state to generate spatially consistent interaction videos for continuous simulation, enabling cross-embodiment transfer from human videos to robotic manipulation tasks.]]>
      </content:encoded>
      <pubDate>Fri, 03 Apr 2026 22:23:00 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/8a1e2dd1/0df500a1.mp3" length="34824704" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2177</itunes:duration>
      <itunes:summary>Closed-loop egocentric simulator persistently updating 3D scene state to generate spatially consistent interaction videos for continuous simulation, enabling cross-embodiment transfer from human videos to robotic manipulation tasks.</itunes:summary>
      <itunes:subtitle>Closed-loop egocentric simulator persistently updating 3D scene state to generate spatially consistent interaction videos for continuous simulation, enabling cross-embodiment transfer from human videos to robotic manipulation tasks.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/8a1e2dd1/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Digit's New Motor Cortex: Sim-to-Real RL for Whole-Body Control</title>
      <itunes:title>Digit's New Motor Cortex: Sim-to-Real RL for Whole-Body Control</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">4efe3fe4-723e-47af-8d9b-ae887f57de48</guid>
      <link>https://share.transistor.fm/s/89e819e4</link>
      <description>
        <![CDATA[AI-trained capabilities for new whole-body motions using mocap/teleop data and sim-to-real reinforcement learning, deployable overnight on hardware.]]>
      </description>
      <content:encoded>
        <![CDATA[AI-trained capabilities for new whole-body motions using mocap/teleop data and sim-to-real reinforcement learning, deployable overnight on hardware.]]>
      </content:encoded>
      <pubDate>Fri, 03 Apr 2026 07:13:57 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/89e819e4/944965c8.mp3" length="29959680" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1873</itunes:duration>
      <itunes:summary>AI-trained capabilities for new whole-body motions using mocap/teleop data and sim-to-real reinforcement learning, deployable overnight on hardware.</itunes:summary>
      <itunes:subtitle>AI-trained capabilities for new whole-body motions using mocap/teleop data and sim-to-real reinforcement learning, deployable overnight on hardware.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/89e819e4/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>EgoNav: Diffusion-Based Humanoid Navigation from Human Egocentric Video</title>
      <itunes:title>EgoNav: Diffusion-Based Humanoid Navigation from Human Egocentric Video</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">e6e13ff3-1774-44f6-b563-a4142434dd00</guid>
      <link>https://share.transistor.fm/s/8a739d58</link>
      <description>
        <![CDATA[Diffusion-based humanoid navigation trained solely on 5 hours of human egocentric video data, enabling zero-shot deployment on Unitree G1 for complex behaviors like handling glass walls, crowds, and dynamic obstacles via 360° visual memory and hybrid trajectory sampling; upcoming release of dataset, models, and code.]]>
      </description>
      <content:encoded>
        <![CDATA[Diffusion-based humanoid navigation trained solely on 5 hours of human egocentric video data, enabling zero-shot deployment on Unitree G1 for complex behaviors like handling glass walls, crowds, and dynamic obstacles via 360° visual memory and hybrid trajectory sampling; upcoming release of dataset, models, and code.]]>
      </content:encoded>
      <pubDate>Thu, 02 Apr 2026 22:32:11 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/8a739d58/08c154d8.mp3" length="40440320" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2528</itunes:duration>
      <itunes:summary>Diffusion-based humanoid navigation trained solely on 5 hours of human egocentric video data, enabling zero-shot deployment on Unitree G1 for complex behaviors like handling glass walls, crowds, and dynamic obstacles via 360° visual memory and hybrid trajectory sampling; upcoming release of dataset, models, and code.</itunes:summary>
      <itunes:subtitle>Diffusion-based humanoid navigation trained solely on 5 hours of human egocentric video data, enabling zero-shot deployment on Unitree G1 for complex behaviors like handling glass walls, crowds, and dynamic obstacles via 360° visual memory and hybrid traj</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/8a739d58/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>CaP-X: A Code-as-Policy Framework for Robot Manipulation</title>
      <itunes:title>CaP-X: A Code-as-Policy Framework for Robot Manipulation</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">3ed38932-4992-45fd-b0e4-46a12f86765a</guid>
      <link>https://share.transistor.fm/s/cd3e51d2</link>
      <description>
        <![CDATA[Comprehensive open-source agentic robotics framework treating VLMs/LLMs as code-generating APIs for perception (SAM3, Molmo) and control (IK, grasping), with CaP-Gym benchmark of 187 diverse manipulation tasks (tabletop, bimanual, mobile; sim/real) and CaP-Bench evaluating 12 frontier models; demonstrates rapid RL gains (7B model from 20% to 72% success) with strong sim-to-real transfer.]]>
      </description>
      <content:encoded>
        <![CDATA[Comprehensive open-source agentic robotics framework treating VLMs/LLMs as code-generating APIs for perception (SAM3, Molmo) and control (IK, grasping), with CaP-Gym benchmark of 187 diverse manipulation tasks (tabletop, bimanual, mobile; sim/real) and CaP-Bench evaluating 12 frontier models; demonstrates rapid RL gains (7B model from 20% to 72% success) with strong sim-to-real transfer.]]>
      </content:encoded>
      <pubDate>Thu, 02 Apr 2026 22:19:27 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/cd3e51d2/31fb79cf.mp3" length="13376000" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>836</itunes:duration>
      <itunes:summary>Comprehensive open-source agentic robotics framework treating VLMs/LLMs as code-generating APIs for perception (SAM3, Molmo) and control (IK, grasping), with CaP-Gym benchmark of 187 diverse manipulation tasks (tabletop, bimanual, mobile; sim/real) and CaP-Bench evaluating 12 frontier models; demonstrates rapid RL gains (7B model from 20% to 72% success) with strong sim-to-real transfer.</itunes:summary>
      <itunes:subtitle>Comprehensive open-source agentic robotics framework treating VLMs/LLMs as code-generating APIs for perception (SAM3, Molmo) and control (IK, grasping), with CaP-Gym benchmark of 187 diverse manipulation tasks (tabletop, bimanual, mobile; sim/real) and Ca</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/cd3e51d2/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Embodied Intelligence Breakthrough: Generalist AI’s GEN-1 Robots</title>
      <itunes:title>Embodied Intelligence Breakthrough: Generalist AI’s GEN-1 Robots</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">faff394a-526b-450b-a576-3042e09faa1c</guid>
      <link>https://share.transistor.fm/s/566b6b60</link>
      <description>
        <![CDATA[We've created GEN-1, our latest milestone in scaling robot learning. We believe it to be the first general-purpose AI model that crosses a new performance threshold: mastery of simple physical tasks. It improves average success rates to 99% on tasks where previous models achieve 64%, completes tasks roughly 3x faster than state of the art, and requires only 1 hour of robot data for each of these results. GEN-1 unlocks commercial viability across a broad range of applications—and while it cannot solve all tasks today, it is a significant step towards our mission of creating generalist intelligence for the physical world.]]>
      </description>
      <content:encoded>
        <![CDATA[We've created GEN-1, our latest milestone in scaling robot learning. We believe it to be the first general-purpose AI model that crosses a new performance threshold: mastery of simple physical tasks. It improves average success rates to 99% on tasks where previous models achieve 64%, completes tasks roughly 3x faster than state of the art, and requires only 1 hour of robot data for each of these results. GEN-1 unlocks commercial viability across a broad range of applications—and while it cannot solve all tasks today, it is a significant step towards our mission of creating generalist intelligence for the physical world.]]>
      </content:encoded>
      <pubDate>Thu, 02 Apr 2026 12:58:30 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/566b6b60/072ac9fa.mp3" length="14822400" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>927</itunes:duration>
      <itunes:summary>We've created GEN-1, our latest milestone in scaling robot learning. We believe it to be the first general-purpose AI model that crosses a new performance threshold: mastery of simple physical tasks. It improves average success rates to 99% on tasks where previous models achieve 64%, completes tasks roughly 3x faster than state of the art, and requires only 1 hour of robot data for each of these results. GEN-1 unlocks commercial viability across a broad range of applications—and while it cannot solve all tasks today, it is a significant step towards our mission of creating generalist intelligence for the physical world.</itunes:summary>
      <itunes:subtitle>We've created GEN-1, our latest milestone in scaling robot learning. We believe it to be the first general-purpose AI model that crosses a new performance threshold: mastery of simple physical tasks. It improves average success rates to 99% on tasks where</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/566b6b60/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>CaP-X: LMs' First Physical Exam</title>
      <itunes:title>CaP-X: LMs' First Physical Exam</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">c9665584-9083-466b-8544-e147bb21083d</guid>
      <link>https://share.transistor.fm/s/6e815197</link>
      <description>
        <![CDATA[A novel benchmark that evaluates language models on physical examination tasks, testing their ability to understand and perform clinical physical exam procedures in simulated environments. This work introduces a comprehensive evaluation framework for AI systems in medical/clinical settings.]]>
      </description>
      <content:encoded>
        <![CDATA[A novel benchmark that evaluates language models on physical examination tasks, testing their ability to understand and perform clinical physical exam procedures in simulated environments. This work introduces a comprehensive evaluation framework for AI systems in medical/clinical settings.]]>
      </content:encoded>
      <pubDate>Thu, 02 Apr 2026 12:43:57 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/6e815197/5afc1830.mp3" length="21207552" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1326</itunes:duration>
      <itunes:summary>A novel benchmark that evaluates language models on physical examination tasks, testing their ability to understand and perform clinical physical exam procedures in simulated environments. This work introduces a comprehensive evaluation framework for AI systems in medical/clinical settings.</itunes:summary>
      <itunes:subtitle>A novel benchmark that evaluates language models on physical examination tasks, testing their ability to understand and perform clinical physical exam procedures in simulated environments. This work introduces a comprehensive evaluation framework for AI s</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/6e815197/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>AI Model Collapse: The Danger of Training on AI-Generated Data</title>
      <itunes:title>AI Model Collapse: The Danger of Training on AI-Generated Data</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">a8444e9f-a814-4c66-9164-90fa1c64ac64</guid>
      <link>https://share.transistor.fm/s/c7adc584</link>
      <description>
        <![CDATA[Demonstrated that LLMs trained recursively on AI-generated data suffer model collapse, a degenerative process where they lose grasp of true data distributions. Sparked critical debates on data provenance and the importance of preserving human-generated training data.]]>
      </description>
      <content:encoded>
        <![CDATA[Demonstrated that LLMs trained recursively on AI-generated data suffer model collapse, a degenerative process where they lose grasp of true data distributions. Sparked critical debates on data provenance and the importance of preserving human-generated training data.]]>
      </content:encoded>
      <pubDate>Tue, 31 Mar 2026 07:36:21 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/c7adc584/cb594858.mp3" length="30248960" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1891</itunes:duration>
      <itunes:summary>Demonstrated that LLMs trained recursively on AI-generated data suffer model collapse, a degenerative process where they lose grasp of true data distributions. Sparked critical debates on data provenance and the importance of preserving human-generated training data.</itunes:summary>
      <itunes:subtitle>Demonstrated that LLMs trained recursively on AI-generated data suffer model collapse, a degenerative process where they lose grasp of true data distributions. Sparked critical debates on data provenance and the importance of preserving human-generated tr</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/c7adc584/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>High-Level Automated Reasoning with Qwen2.5-7B</title>
      <itunes:title>High-Level Automated Reasoning with Qwen2.5-7B</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">6011d202-428b-4935-b330-d7946a08cb76</guid>
      <link>https://share.transistor.fm/s/39a8fb88</link>
      <description>
        <![CDATA[Qwen2.5-7B achieved 79.6% on MATH benchmark, surpassing GPT-4o, by employing atomic reasoning actions combined with Monte Carlo Tree Search. Demonstrated that strategic reasoning architectures can enable smaller models to outperform much larger ones.]]>
      </description>
      <content:encoded>
        <![CDATA[Qwen2.5-7B achieved 79.6% on MATH benchmark, surpassing GPT-4o, by employing atomic reasoning actions combined with Monte Carlo Tree Search. Demonstrated that strategic reasoning architectures can enable smaller models to outperform much larger ones.]]>
      </content:encoded>
      <pubDate>Tue, 31 Mar 2026 07:35:15 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/39a8fb88/b3b6d323.mp3" length="26634752" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1665</itunes:duration>
      <itunes:summary>Qwen2.5-7B achieved 79.6% on MATH benchmark, surpassing GPT-4o, by employing atomic reasoning actions combined with Monte Carlo Tree Search. Demonstrated that strategic reasoning architectures can enable smaller models to outperform much larger ones.</itunes:summary>
      <itunes:subtitle>Qwen2.5-7B achieved 79.6% on MATH benchmark, surpassing GPT-4o, by employing atomic reasoning actions combined with Monte Carlo Tree Search. Demonstrated that strategic reasoning architectures can enable smaller models to outperform much larger ones.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/39a8fb88/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Co-Training Large Behavior Models: Multimodal Data for Robot Manipulation</title>
      <itunes:title>Co-Training Large Behavior Models: Multimodal Data for Robot Manipulation</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">0595fed9-a77a-46e7-8806-8062b7cb5220</guid>
      <link>https://share.transistor.fm/s/bcf99c35</link>
      <description>
        <![CDATA[Explores data modalities and co-training strategies to enhance large behavior models (foundation models) for improved performance in robot manipulation tasks, supporting end-to-end learning and cross-embodiment generalization.]]>
      </description>
      <content:encoded>
        <![CDATA[Explores data modalities and co-training strategies to enhance large behavior models (foundation models) for improved performance in robot manipulation tasks, supporting end-to-end learning and cross-embodiment generalization.]]>
      </content:encoded>
      <pubDate>Mon, 30 Mar 2026 22:19:22 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/bcf99c35/4554d69d.mp3" length="31812096" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1989</itunes:duration>
      <itunes:summary>Explores data modalities and co-training strategies to enhance large behavior models (foundation models) for improved performance in robot manipulation tasks, supporting end-to-end learning and cross-embodiment generalization.</itunes:summary>
      <itunes:subtitle>Explores data modalities and co-training strategies to enhance large behavior models (foundation models) for improved performance in robot manipulation tasks, supporting end-to-end learning and cross-embodiment generalization.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/bcf99c35/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>HyDRA: Hybrid Memory for Dynamic Video World Models</title>
      <itunes:title>HyDRA: Hybrid Memory for Dynamic Video World Models</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">23622abe-24bf-46b4-aff3-2c982acdd258</guid>
      <link>https://share.transistor.fm/s/951f9f06</link>
      <description>
        <![CDATA[Memory architecture preserving identity and motion continuity for out-of-view dynamic subjects, addressing frozen/vanishing issues in video world models.]]>
      </description>
      <content:encoded>
        <![CDATA[Memory architecture preserving identity and motion continuity for out-of-view dynamic subjects, addressing frozen/vanishing issues in video world models.]]>
      </content:encoded>
      <pubDate>Sun, 29 Mar 2026 22:20:50 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/951f9f06/addb8d4c.mp3" length="34514432" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2158</itunes:duration>
      <itunes:summary>Memory architecture preserving identity and motion continuity for out-of-view dynamic subjects, addressing frozen/vanishing issues in video world models.</itunes:summary>
      <itunes:subtitle>Memory architecture preserving identity and motion continuity for out-of-view dynamic subjects, addressing frozen/vanishing issues in video world models.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/951f9f06/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>DexWM: Leveraging Human Videos for Dexterous Robot World Models</title>
      <itunes:title>DexWM: Leveraging Human Videos for Dexterous Robot World Models</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">156aeea0-9fe8-4a42-bb94-bd1c3e3bf10b</guid>
      <link>https://share.transistor.fm/s/64be43f8</link>
      <description>
        <![CDATA[Dataset of robot trajectories designed for training world models to learn dexterous hand-object interactions directly from human videos.]]>
      </description>
      <content:encoded>
        <![CDATA[Dataset of robot trajectories designed for training world models to learn dexterous hand-object interactions directly from human videos.]]>
      </content:encoded>
      <pubDate>Sun, 29 Mar 2026 22:18:43 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/64be43f8/945d4bed.mp3" length="30125056" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1883</itunes:duration>
      <itunes:summary>Dataset of robot trajectories designed for training world models to learn dexterous hand-object interactions directly from human videos.</itunes:summary>
      <itunes:subtitle>Dataset of robot trajectories designed for training world models to learn dexterous hand-object interactions directly from human videos.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/64be43f8/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>World Models in Robotics</title>
      <itunes:title>World Models in Robotics</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">c910ec03-1aff-4520-bce6-b4ab3cdec0d1</guid>
      <link>https://share.transistor.fm/s/2bd996b5</link>
      <description>
        <![CDATA[Technical survey categorizing world models into action-conditioned, video-inverse dynamics, and joint world-action models (WAMs), discussing their generalization, video data leverage, and trends for closing the robotics data gap.]]>
      </description>
      <content:encoded>
        <![CDATA[Technical survey categorizing world models into action-conditioned, video-inverse dynamics, and joint world-action models (WAMs), discussing their generalization, video data leverage, and trends for closing the robotics data gap.]]>
      </content:encoded>
      <pubDate>Sun, 29 Mar 2026 07:14:29 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/2bd996b5/0fb04438.mp3" length="25766912" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1611</itunes:duration>
      <itunes:summary>Technical survey categorizing world models into action-conditioned, video-inverse dynamics, and joint world-action models (WAMs), discussing their generalization, video data leverage, and trends for closing the robotics data gap.</itunes:summary>
      <itunes:subtitle>Technical survey categorizing world models into action-conditioned, video-inverse dynamics, and joint world-action models (WAMs), discussing their generalization, video data leverage, and trends for closing the robotics data gap.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/2bd996b5/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>SIMART: Decomposing Monolithic Meshes into Sim-Ready Articulated Assets</title>
      <itunes:title>SIMART: Decomposing Monolithic Meshes into Sim-Ready Articulated Assets</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">9d0dad82-3b27-481d-a93c-b005296f6620</guid>
      <link>https://share.transistor.fm/s/1e0bfc14</link>
      <description>
        <![CDATA[Unified MLLM framework with Sparse 3D VQ-VAE that reduces tokens by 70% for efficient part-level decomposition and kinematic prediction in physics-based robotic simulations.]]>
      </description>
      <content:encoded>
        <![CDATA[Unified MLLM framework with Sparse 3D VQ-VAE that reduces tokens by 70% for efficient part-level decomposition and kinematic prediction in physics-based robotic simulations.]]>
      </content:encoded>
      <pubDate>Sat, 28 Mar 2026 07:21:21 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/1e0bfc14/08ba830b.mp3" length="43369472" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2711</itunes:duration>
      <itunes:summary>Unified MLLM framework with Sparse 3D VQ-VAE that reduces tokens by 70% for efficient part-level decomposition and kinematic prediction in physics-based robotic simulations.</itunes:summary>
      <itunes:subtitle>Unified MLLM framework with Sparse 3D VQ-VAE that reduces tokens by 70% for efficient part-level decomposition and kinematic prediction in physics-based robotic simulations.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/1e0bfc14/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>LeWorldModel: A Stable JEPA World Model from Pixels</title>
      <itunes:title>LeWorldModel: A Stable JEPA World Model from Pixels</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">a87d50dc-b8f1-472a-a43c-38ffbad989e4</guid>
      <link>https://share.transistor.fm/s/a996f27a</link>
      <description>
        <![CDATA[Stable end-to-end JEPA world model trained directly from pixels using simple MSE prediction loss and SIGReg anti-collapse regularization, enabling efficient latent planning under 1 second on 15M params with emergent spatial structure outperforming prior methods.]]>
      </description>
      <content:encoded>
        <![CDATA[Stable end-to-end JEPA world model trained directly from pixels using simple MSE prediction loss and SIGReg anti-collapse regularization, enabling efficient latent planning under 1 second on 15M params with emergent spatial structure outperforming prior methods.]]>
      </content:encoded>
      <pubDate>Fri, 27 Mar 2026 22:16:24 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/a996f27a/3e660550.mp3" length="13365248" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>836</itunes:duration>
      <itunes:summary>Stable end-to-end JEPA world model trained directly from pixels using simple MSE prediction loss and SIGReg anti-collapse regularization, enabling efficient latent planning under 1 second on 15M params with emergent spatial structure outperforming prior methods.</itunes:summary>
      <itunes:subtitle>Stable end-to-end JEPA world model trained directly from pixels using simple MSE prediction loss and SIGReg anti-collapse regularization, enabling efficient latent planning under 1 second on 15M params with emergent spatial structure outperforming prior m</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/a996f27a/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>World Models for Robots: The Next Big Leap?</title>
      <itunes:title>World Models for Robots: The Next Big Leap?</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">c0e1f103-4f9b-432b-9bad-20fa114a2426</guid>
      <link>https://share.transistor.fm/s/03edfd29</link>
      <description>
        <![CDATA[Technical overview defining world models in robotics, their potential to solve diverse problems via video prediction, and key enablers like scale.]]>
      </description>
      <content:encoded>
        <![CDATA[Technical overview defining world models in robotics, their potential to solve diverse problems via video prediction, and key enablers like scale.]]>
      </content:encoded>
      <pubDate>Fri, 27 Mar 2026 07:39:49 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/03edfd29/8e9289f7.mp3" length="19534848" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1221</itunes:duration>
      <itunes:summary>Technical overview defining world models in robotics, their potential to solve diverse problems via video prediction, and key enablers like scale.</itunes:summary>
      <itunes:subtitle>Technical overview defining world models in robotics, their potential to solve diverse problems via video prediction, and key enablers like scale.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/03edfd29/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Harnessing Long-Running AI in Embodied Systems</title>
      <itunes:title>Harnessing Long-Running AI in Embodied Systems</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">ab34a505-c88b-4ad7-aa6c-48feaef4686b</guid>
      <link>https://share.transistor.fm/s/c2a00a1d</link>
      <description>
        <![CDATA[As AI moves from quick Q&amp;A to marathon tasks, designers grapple with continuity. This episode explores how Anthropics harness design principles translate to embodied AI - robots that need to maintain context across long-running missions.]]>
      </description>
      <content:encoded>
        <![CDATA[As AI moves from quick Q&amp;A to marathon tasks, designers grapple with continuity. This episode explores how Anthropics harness design principles translate to embodied AI - robots that need to maintain context across long-running missions.]]>
      </content:encoded>
      <pubDate>Fri, 27 Mar 2026 00:11:43 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/c2a00a1d/b3efe5fc.mp3" length="26078720" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1630</itunes:duration>
      <itunes:summary>As AI moves from quick Q&amp;amp;A to marathon tasks, designers grapple with continuity. This episode explores how Anthropics harness design principles translate to embodied AI - robots that need to maintain context across long-running missions.</itunes:summary>
      <itunes:subtitle>As AI moves from quick Q&amp;amp;A to marathon tasks, designers grapple with continuity. This episode explores how Anthropics harness design principles translate to embodied AI - robots that need to maintain context across long-running missions.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/c2a00a1d/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>HoMMI: Learning Whole-Body Mobile Manipulation from Human Demonstrations</title>
      <itunes:title>HoMMI: Learning Whole-Body Mobile Manipulation from Human Demonstrations</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">8aed0ce0-2291-4d3a-be47-1a6679fb06b3</guid>
      <link>https://share.transistor.fm/s/e34191c6</link>
      <description>
        <![CDATA[Whole-Body Mobile Manipulation Interface (HoMMI) that learns bimanual and whole-body manipulation, long-horizon navigation, and active perception directly from egocentric human demonstrations without teleoperation.]]>
      </description>
      <content:encoded>
        <![CDATA[Whole-Body Mobile Manipulation Interface (HoMMI) that learns bimanual and whole-body manipulation, long-horizon navigation, and active perception directly from egocentric human demonstrations without teleoperation.]]>
      </content:encoded>
      <pubDate>Wed, 25 Mar 2026 22:29:11 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/e34191c6/b7c7e89a.mp3" length="16363008" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1020</itunes:duration>
      <itunes:summary>Whole-Body Mobile Manipulation Interface (HoMMI) that learns bimanual and whole-body manipulation, long-horizon navigation, and active perception directly from egocentric human demonstrations without teleoperation.</itunes:summary>
      <itunes:subtitle>Whole-Body Mobile Manipulation Interface (HoMMI) that learns bimanual and whole-body manipulation, long-horizon navigation, and active perception directly from egocentric human demonstrations without teleoperation.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/e34191c6/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>TurboQuant: Redefining AI Efficiency with Extreme Compression</title>
      <itunes:title>TurboQuant: Redefining AI Efficiency with Extreme Compression</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">e58b50dd-daaa-4e9c-98f8-ae918b436368</guid>
      <link>https://share.transistor.fm/s/86d6b9f8</link>
      <description>
        <![CDATA[<p>This episode explores TurboQuant, a revolutionary set of quantization algorithms from Google Research that redefines AI efficiency through extreme compression.</p><p>We dive deep into how TurboQuant addresses one of AI's most pressing challenges: the memory bottleneck created by high-dimensional vectors in key-value caches. The research introduces theoretically grounded quantization methods that enable massive compression for large language models and vector search engines without sacrificing performance.</p><p>Key topics covered:</p><ul><li>The theoretical foundations of TurboQuant's quantization algorithms</li><li>How extreme compression works for LLMs and vector search engines</li><li>Impact on high-dimensional vectors and key-value cache memory bottlenecks</li><li>Performance metrics and comparisons with existing methods</li><li>Practical implications for AI deployment and efficiency</li></ul><p>Links:<br>Paper: https://arxiv.org/pdf/2504.19874<br>Blog: https://research.google/blog/turboquant-redefining-ai-efficiency-with-extreme-compression/</p>]]>
      </description>
      <content:encoded>
        <![CDATA[<p>This episode explores TurboQuant, a revolutionary set of quantization algorithms from Google Research that redefines AI efficiency through extreme compression.</p><p>We dive deep into how TurboQuant addresses one of AI's most pressing challenges: the memory bottleneck created by high-dimensional vectors in key-value caches. The research introduces theoretically grounded quantization methods that enable massive compression for large language models and vector search engines without sacrificing performance.</p><p>Key topics covered:</p><ul><li>The theoretical foundations of TurboQuant's quantization algorithms</li><li>How extreme compression works for LLMs and vector search engines</li><li>Impact on high-dimensional vectors and key-value cache memory bottlenecks</li><li>Performance metrics and comparisons with existing methods</li><li>Practical implications for AI deployment and efficiency</li></ul><p>Links:<br>Paper: https://arxiv.org/pdf/2504.19874<br>Blog: https://research.google/blog/turboquant-redefining-ai-efficiency-with-extreme-compression/</p>]]>
      </content:encoded>
      <pubDate>Wed, 25 Mar 2026 17:52:48 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/86d6b9f8/84c47914.mp3" length="19600384" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1225</itunes:duration>
      <itunes:summary>Google Research introduces TurboQuant, a breakthrough in quantization algorithms that enables massive compression for LLMs and vector search engines, solving critical memory bottlenecks in AI systems through theoretically grounded extreme compression techniques.</itunes:summary>
      <itunes:subtitle>Google Research introduces TurboQuant, a breakthrough in quantization algorithms that enables massive compression for LLMs and vector search engines, solving critical memory bottlenecks in AI systems through theoretically grounded extreme compression tech</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/86d6b9f8/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>DexWM: Learning Dexterous Object Manipulation from Human Videos</title>
      <itunes:title>DexWM: Learning Dexterous Object Manipulation from Human Videos</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">fed7e6b4-6143-478a-af46-ead7da86ec8f</guid>
      <link>https://share.transistor.fm/s/515495e8</link>
      <description>
        <![CDATA[Dataset of robot trajectories designed for training world models that learn dexterous hand-object interactions from human videos, released on Hugging Face.]]>
      </description>
      <content:encoded>
        <![CDATA[Dataset of robot trajectories designed for training world models that learn dexterous hand-object interactions from human videos, released on Hugging Face.]]>
      </content:encoded>
      <pubDate>Wed, 25 Mar 2026 07:19:46 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/515495e8/914f7bff.mp3" length="30930944" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1934</itunes:duration>
      <itunes:summary>Dataset of robot trajectories designed for training world models that learn dexterous hand-object interactions from human videos, released on Hugging Face.</itunes:summary>
      <itunes:subtitle>Dataset of robot trajectories designed for training world models that learn dexterous hand-object interactions from human videos, released on Hugging Face.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/515495e8/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>FlashAttention-3: Fast &amp; Accurate Attention with Asynchrony &amp; Low-Precision</title>
      <itunes:title>FlashAttention-3: Fast &amp; Accurate Attention with Asynchrony &amp; Low-Precision</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">dc804725-7ff1-4f06-88a1-2b2423f7f5f3</guid>
      <link>https://share.transistor.fm/s/438d3ecf</link>
      <description>
        <![CDATA[Major efficiency leap for Transformer attention mechanisms, enabling faster training/inference on long sequences with low-precision compute.]]>
      </description>
      <content:encoded>
        <![CDATA[Major efficiency leap for Transformer attention mechanisms, enabling faster training/inference on long sequences with low-precision compute.]]>
      </content:encoded>
      <pubDate>Tue, 24 Mar 2026 22:54:40 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/438d3ecf/e1b5ba82.mp3" length="16724992" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1046</itunes:duration>
      <itunes:summary>Major efficiency leap for Transformer attention mechanisms, enabling faster training/inference on long sequences with low-precision compute.</itunes:summary>
      <itunes:subtitle>Major efficiency leap for Transformer attention mechanisms, enabling faster training/inference on long sequences with low-precision compute.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/438d3ecf/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>When AI Trains on Its Own Output: The Model Collapse Problem</title>
      <itunes:title>When AI Trains on Its Own Output: The Model Collapse Problem</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">53875c2e-37b0-4701-a3e0-410b44a10e56</guid>
      <link>https://share.transistor.fm/s/6cf58275</link>
      <description>
        <![CDATA[Warns of "model collapse" in LLMs trained on synthetic data from prior models, urging preservation of human-generated data. One of 2024's most influential papers.]]>
      </description>
      <content:encoded>
        <![CDATA[Warns of "model collapse" in LLMs trained on synthetic data from prior models, urging preservation of human-generated data. One of 2024's most influential papers.]]>
      </content:encoded>
      <pubDate>Tue, 24 Mar 2026 22:39:06 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/6cf58275/07ebec32.mp3" length="23684096" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1481</itunes:duration>
      <itunes:summary>Warns of "model collapse" in LLMs trained on synthetic data from prior models, urging preservation of human-generated data. One of 2024's most influential papers.</itunes:summary>
      <itunes:subtitle>Warns of "model collapse" in LLMs trained on synthetic data from prior models, urging preservation of human-generated data. One of 2024's most influential papers.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/6cf58275/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>MolmoBot: A Vision-Language Model for Zero-Shot Robot Manipulation</title>
      <itunes:title>MolmoBot: A Vision-Language Model for Zero-Shot Robot Manipulation</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">13c2834c-0651-4361-87da-2afba9742684</guid>
      <link>https://share.transistor.fm/s/fc1abf47</link>
      <description>
        <![CDATA[Vision-language model (VLM) for zero-shot robot manipulation, trained entirely in simulation without real-world data; achieves 79.2% success rate on real-world tabletop tasks, outperforming π₀.₅ baseline at 39.2%.]]>
      </description>
      <content:encoded>
        <![CDATA[Vision-language model (VLM) for zero-shot robot manipulation, trained entirely in simulation without real-world data; achieves 79.2% success rate on real-world tabletop tasks, outperforming π₀.₅ baseline at 39.2%.]]>
      </content:encoded>
      <pubDate>Tue, 24 Mar 2026 07:22:35 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/fc1abf47/1505cf54.mp3" length="36572160" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2286</itunes:duration>
      <itunes:summary>Vision-language model (VLM) for zero-shot robot manipulation, trained entirely in simulation without real-world data; achieves 79.2% success rate on real-world tabletop tasks, outperforming π₀.₅ baseline at 39.2%.</itunes:summary>
      <itunes:subtitle>Vision-language model (VLM) for zero-shot robot manipulation, trained entirely in simulation without real-world data; achieves 79.2% success rate on real-world tabletop tasks, outperforming π₀.₅ baseline at 39.2%.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/fc1abf47/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>LeWorldModel: Stable End-to-End JEPA from Pixels</title>
      <itunes:title>LeWorldModel: Stable End-to-End JEPA from Pixels</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">e43e7bca-b815-403b-9e20-67abdaf2a10d</guid>
      <link>https://share.transistor.fm/s/6913c086</link>
      <description>
        <![CDATA[A stable end-to-end Joint Embedding Predictive Architecture (JEPA) trained directly from pixels that enables robust world modeling for embodied AI systems.]]>
      </description>
      <content:encoded>
        <![CDATA[A stable end-to-end Joint Embedding Predictive Architecture (JEPA) trained directly from pixels that enables robust world modeling for embodied AI systems.]]>
      </content:encoded>
      <pubDate>Tue, 24 Mar 2026 01:12:52 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/6913c086/88848ce1.mp3" length="12611072" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>789</itunes:duration>
      <itunes:summary>A stable end-to-end Joint Embedding Predictive Architecture (JEPA) trained directly from pixels that enables robust world modeling for embodied AI systems.</itunes:summary>
      <itunes:subtitle>A stable end-to-end Joint Embedding Predictive Architecture (JEPA) trained directly from pixels that enables robust world modeling for embodied AI systems.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/6913c086/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>EgoVerse: An Egocentric Data Ecosystem for Scaling Robot Learning</title>
      <itunes:title>EgoVerse: An Egocentric Data Ecosystem for Scaling Robot Learning</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">1f098822-fdc2-4f97-8f27-7c485e73ba6d</guid>
      <link>https://share.transistor.fm/s/a163eab0</link>
      <description>
        <![CDATA[Ecosystem with over 1300 hours of egocentric human video data spanning 240 scenes and 2000+ tasks, designed for scalable robot policy training via behavior cloning; includes cloud infrastructure, data viewer, and human-to-robot transfer algorithms to enable cross-embodiment learning without teleoperation.]]>
      </description>
      <content:encoded>
        <![CDATA[Ecosystem with over 1300 hours of egocentric human video data spanning 240 scenes and 2000+ tasks, designed for scalable robot policy training via behavior cloning; includes cloud infrastructure, data viewer, and human-to-robot transfer algorithms to enable cross-embodiment learning without teleoperation.]]>
      </content:encoded>
      <pubDate>Mon, 23 Mar 2026 22:18:26 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/a163eab0/b7fb449b.mp3" length="40259072" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>2517</itunes:duration>
      <itunes:summary>Ecosystem with over 1300 hours of egocentric human video data spanning 240 scenes and 2000+ tasks, designed for scalable robot policy training via behavior cloning; includes cloud infrastructure, data viewer, and human-to-robot transfer algorithms to enable cross-embodiment learning without teleoperation.</itunes:summary>
      <itunes:subtitle>Ecosystem with over 1300 hours of egocentric human video data spanning 240 scenes and 2000+ tasks, designed for scalable robot policy training via behavior cloning; includes cloud infrastructure, data viewer, and human-to-robot transfer algorithms to enab</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/a163eab0/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>HSImul3R: Physics-Driven Reconstruction of Human–Scene Interactions</title>
      <itunes:title>HSImul3R: Physics-Driven Reconstruction of Human–Scene Interactions</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">9c5a18d4-07a8-46eb-9e68-0b86c49eb745</guid>
      <link>https://share.transistor.fm/s/8e30b95b</link>
      <description>
        <![CDATA[Physics-in-the-loop bi-directional optimization pipeline reconstructing stable, simulation-ready 3D human-scene interactions from casual videos, deployable directly to humanoid robots for world modeling and manipulation.]]>
      </description>
      <content:encoded>
        <![CDATA[Physics-in-the-loop bi-directional optimization pipeline reconstructing stable, simulation-ready 3D human-scene interactions from casual videos, deployable directly to humanoid robots for world modeling and manipulation.]]>
      </content:encoded>
      <pubDate>Mon, 23 Mar 2026 22:15:11 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/8e30b95b/1131c480.mp3" length="26994688" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1688</itunes:duration>
      <itunes:summary>Physics-in-the-loop bi-directional optimization pipeline reconstructing stable, simulation-ready 3D human-scene interactions from casual videos, deployable directly to humanoid robots for world modeling and manipulation.</itunes:summary>
      <itunes:subtitle>Physics-in-the-loop bi-directional optimization pipeline reconstructing stable, simulation-ready 3D human-scene interactions from casual videos, deployable directly to humanoid robots for world modeling and manipulation.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/8e30b95b/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>MolmoSpaces: A Large-Scale Open Ecosystem for Robot Navigation and Manipulation</title>
      <itunes:title>MolmoSpaces: A Large-Scale Open Ecosystem for Robot Navigation and Manipulation</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">34fb3f2f-2a5a-443a-a0f8-d93e6031b83c</guid>
      <link>https://share.transistor.fm/s/3f56c98e</link>
      <description>
        <![CDATA[Open-source suite of large-scale simulation environments and benchmarks designed for advancing end-to-end learning in robot navigation and manipulation across multiple embodiments.]]>
      </description>
      <content:encoded>
        <![CDATA[Open-source suite of large-scale simulation environments and benchmarks designed for advancing end-to-end learning in robot navigation and manipulation across multiple embodiments.]]>
      </content:encoded>
      <pubDate>Mon, 23 Mar 2026 11:48:08 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/3f56c98e/95172d6c.mp3" length="28189696" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1762</itunes:duration>
      <itunes:summary>Open-source suite of large-scale simulation environments and benchmarks designed for advancing end-to-end learning in robot navigation and manipulation across multiple embodiments.</itunes:summary>
      <itunes:subtitle>Open-source suite of large-scale simulation environments and benchmarks designed for advancing end-to-end learning in robot navigation and manipulation across multiple embodiments.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/3f56c98e/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>DreamZero: World Action Models Are Zero-Shot Policies</title>
      <itunes:title>DreamZero: World Action Models Are Zero-Shot Policies</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">084662d4-8e3a-4b1c-ae1b-34f3bd140878</guid>
      <link>https://share.transistor.fm/s/938915b1</link>
      <description>
        <![CDATA[Introduces World Action Models (WAMs), a family of 14B-parameter autoregressive diffusion models that jointly predict video and robotic actions to enable zero-shot generalization across manipulation tasks, outperforming fine-tuned Vision-Language-Action models on benchmarks like MolmoSpaces and RoboArena.]]>
      </description>
      <content:encoded>
        <![CDATA[Introduces World Action Models (WAMs), a family of 14B-parameter autoregressive diffusion models that jointly predict video and robotic actions to enable zero-shot generalization across manipulation tasks, outperforming fine-tuned Vision-Language-Action models on benchmarks like MolmoSpaces and RoboArena.]]>
      </content:encoded>
      <pubDate>Mon, 23 Mar 2026 11:36:32 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/938915b1/f07d1ff5.mp3" length="25711616" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1607</itunes:duration>
      <itunes:summary>Introduces World Action Models (WAMs), a family of 14B-parameter autoregressive diffusion models that jointly predict video and robotic actions to enable zero-shot generalization across manipulation tasks, outperforming fine-tuned Vision-Language-Action models on benchmarks like MolmoSpaces and RoboArena.</itunes:summary>
      <itunes:subtitle>Introduces World Action Models (WAMs), a family of 14B-parameter autoregressive diffusion models that jointly predict video and robotic actions to enable zero-shot generalization across manipulation tasks, outperforming fine-tuned Vision-Language-Action m</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/938915b1/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Kinema4D: A 4D Generative Simulator for Embodied AI</title>
      <itunes:title>Kinema4D: A 4D Generative Simulator for Embodied AI</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">7751c723-1306-4bd6-a1ec-151801fc4783</guid>
      <link>https://share.transistor.fm/s/06ec58aa</link>
      <description>
        <![CDATA[An action-conditioned 4D generative robotic simulator that disentangles precise kinematic control from environmental dynamics, facilitating physically-plausible simulations of complex robot-world interactions for training and world modeling.]]>
      </description>
      <content:encoded>
        <![CDATA[An action-conditioned 4D generative robotic simulator that disentangles precise kinematic control from environmental dynamics, facilitating physically-plausible simulations of complex robot-world interactions for training and world modeling.]]>
      </content:encoded>
      <pubDate>Sun, 22 Mar 2026 19:16:00 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/06ec58aa/bac41a79.mp3" length="29848064" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1866</itunes:duration>
      <itunes:summary>An action-conditioned 4D generative robotic simulator that disentangles precise kinematic control from environmental dynamics, facilitating physically-plausible simulations of complex robot-world interactions for training and world modeling.</itunes:summary>
      <itunes:subtitle>An action-conditioned 4D generative robotic simulator that disentangles precise kinematic control from environmental dynamics, facilitating physically-plausible simulations of complex robot-world interactions for training and world modeling.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/06ec58aa/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>VEGA-3D: Teaching multimodal LLMs spatial reasoning through video generation</title>
      <itunes:title>VEGA-3D: Teaching multimodal LLMs spatial reasoning through video generation</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">17a61ba0-80b5-488f-9706-9db44ab2a0ac</guid>
      <link>https://share.transistor.fm/s/c5eed771</link>
      <description>
        <![CDATA[A plug-and-play framework extracts implicit 3D priors from video diffusion models to enhance multimodal LLMs with spatial reasoning capabilities, enabling improved geometric scene understanding and embodied decision-making without explicit 3D supervision.]]>
      </description>
      <content:encoded>
        <![CDATA[A plug-and-play framework extracts implicit 3D priors from video diffusion models to enhance multimodal LLMs with spatial reasoning capabilities, enabling improved geometric scene understanding and embodied decision-making without explicit 3D supervision.]]>
      </content:encoded>
      <pubDate>Sun, 22 Mar 2026 19:02:22 -0700</pubDate>
      <author>Shaoqing Tan</author>
      <enclosure url="https://media.transistor.fm/c5eed771/16b751d6.mp3" length="31138816" type="audio/mpeg"/>
      <itunes:author>Shaoqing Tan</itunes:author>
      <itunes:duration>1947</itunes:duration>
      <itunes:summary>A plug-and-play framework extracts implicit 3D priors from video diffusion models to enhance multimodal LLMs with spatial reasoning capabilities, enabling improved geometric scene understanding and embodied decision-making without explicit 3D supervision.</itunes:summary>
      <itunes:subtitle>A plug-and-play framework extracts implicit 3D priors from video diffusion models to enhance multimodal LLMs with spatial reasoning capabilities, enabling improved geometric scene understanding and embodied decision-making without explicit 3D supervision.</itunes:subtitle>
      <itunes:keywords>embodied ai technology robotics</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/c5eed771/transcript.txt" type="text/plain"/>
    </item>
  </channel>
</rss>
