<?xml version="1.0" encoding="UTF-8"?>
<?xml-stylesheet href="/stylesheet.xsl" type="text/xsl"?>
<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:sy="http://purl.org/rss/1.0/modules/syndication/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:podcast="https://podcastindex.org/namespace/1.0">
  <channel>
    <atom:link rel="self" type="application/rss+xml" href="https://feeds.transistor.fm/the-experimentation-edge" title="MP3 Audio"/>
    <atom:link rel="hub" href="https://pubsubhubbub.appspot.com/"/>
    <podcast:podping usesPodping="true"/>
    <title>The Experimentation Edge</title>
    <generator>Transistor (https://transistor.fm)</generator>
    <itunes:new-feed-url>https://feeds.transistor.fm/the-experimentation-edge</itunes:new-feed-url>
    <description>How do product teams decide what to build and what not to? The Experimentation Edge is the podcast where product, growth, and engineering leaders share how A/B testing, feature flags, and experimentation drive real business outcomes — backed by named companies and real numbers. From DoorDash's 12,000 A/B tests a year to Atlassian's experimentation-led product win to UPS's $500M experimentation team, each episode goes deep with operators running experimentation programs at scale.

Hosted by Ashley Stirrup, CMO at GrowthBook and a 25-year executive in data and experimentation. For product managers, engineers, data scientists, and growth leaders at B2B tech companies who care about experimentation culture, statistical rigor, and shipping with confidence. No marketing speak. Just operators explaining what they shipped, what moved the needle, and how experimentation reshaped their teams.

Topics: A/B testing, experimentation, growth experimentation, product experimentation, tech experimentation, feature flags, experimentation culture, statistical significance, marketplace experimentation, conversion rate optimization, experimentation at scale.</description>
    <copyright>Ashley Stirrup</copyright>
    <podcast:guid>e03ff1cc-4468-5ca2-a16f-a7a84195031d</podcast:guid>
    <podcast:locked>yes</podcast:locked>
    <language>en</language>
    <pubDate>Thu, 23 Jul 2026 11:53:11 -0600</pubDate>
    <lastBuildDate>Thu, 23 Jul 2026 11:54:11 -0600</lastBuildDate>
    <link>https://the-experimentation-edge.transistor.fm/</link>
    <image>
      <url>https://img.transistorcdn.com/swQmJWlBr0i0NiPZuyg9ULjIj7vBdCykVhI-iVV7qEc/rs:fill:0:0:1/w:1400/h:1400/q:60/mb:500000/aHR0cHM6Ly9pbWct/dXBsb2FkLXByb2R1/Y3Rpb24udHJhbnNp/c3Rvci5mbS80YTFk/MGU1MjJlODhlNjJh/MTdlZTZkN2Q1ODY5/OTdjYy5wbmc.jpg</url>
      <title>The Experimentation Edge</title>
      <link>https://the-experimentation-edge.transistor.fm/</link>
    </image>
    <itunes:category text="Business"/>
    <itunes:category text="Technology"/>
    <itunes:type>episodic</itunes:type>
    <itunes:author>Growthbook</itunes:author>
    <itunes:image href="https://img.transistorcdn.com/swQmJWlBr0i0NiPZuyg9ULjIj7vBdCykVhI-iVV7qEc/rs:fill:0:0:1/w:1400/h:1400/q:60/mb:500000/aHR0cHM6Ly9pbWct/dXBsb2FkLXByb2R1/Y3Rpb24udHJhbnNp/c3Rvci5mbS80YTFk/MGU1MjJlODhlNjJh/MTdlZTZkN2Q1ODY5/OTdjYy5wbmc.jpg"/>
    <itunes:summary>How do product teams decide what to build and what not to? The Experimentation Edge is the podcast where product, growth, and engineering leaders share how A/B testing, feature flags, and experimentation drive real business outcomes — backed by named companies and real numbers. From DoorDash's 12,000 A/B tests a year to Atlassian's experimentation-led product win to UPS's $500M experimentation team, each episode goes deep with operators running experimentation programs at scale.

Hosted by Ashley Stirrup, CMO at GrowthBook and a 25-year executive in data and experimentation. For product managers, engineers, data scientists, and growth leaders at B2B tech companies who care about experimentation culture, statistical rigor, and shipping with confidence. No marketing speak. Just operators explaining what they shipped, what moved the needle, and how experimentation reshaped their teams.

Topics: A/B testing, experimentation, growth experimentation, product experimentation, tech experimentation, feature flags, experimentation culture, statistical significance, marketplace experimentation, conversion rate optimization, experimentation at scale.</itunes:summary>
    <itunes:subtitle>How do product teams decide what to build and what not to.</itunes:subtitle>
    <itunes:keywords>A/B testing,experimentation,product management,growth strategy,feature flags,product experimentation,data-driven decisions,conversion optimization,product leadership,GrowthBook,experimentation culture,product analytics,hypothesis testing,growth marketing,product strategy,CPO,VP product,experimentation platform,product metrics,evidence-based product</itunes:keywords>
    <itunes:owner>
      <itunes:name>Ashley Stirrup</itunes:name>
    </itunes:owner>
    <itunes:complete>No</itunes:complete>
    <itunes:explicit>No</itunes:explicit>
    <item>
      <title>How Cogniteer Built an Experimentation Engine From Scratch</title>
      <itunes:episode>29</itunes:episode>
      <podcast:episode>29</podcast:episode>
      <itunes:title>How Cogniteer Built an Experimentation Engine From Scratch</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">7c584f3b-1d18-4191-a934-244256334b80</guid>
      <link>https://share.transistor.fm/s/9016b4fd</link>
      <description>
        <![CDATA[<p><strong><br>Summary</strong></p><p>In this episode of The Experimentation Edge, host Ashley Stirrup talks with Fabian Hans, founder and behavioral psychologist at Cogniteer, a consultancy that helps enterprises build in-house experimentation programs and raise both test velocity and win rate. Drawing on fifteen years in conversion rate optimization, Fabian explains why mass-producing the same A/B tests across clients quietly kills learning, why most ecommerce drop-offs are structural rather than your fault, and how matching the interface to how people actually buy, new versus returning, B2C versus B2B, can move conversion far more than another button. It is a practical, psychology-grounded conversation for product managers, engineers, data scientists, and growth leaders who want their experimentation programs to compound understanding, not just volume.</p><p><strong><br>Chapters</strong></p><p>00:00 Introduction</p><p>01:25 From agency mass production to in house deep dives</p><p>04:05 Why some products resist selling online</p><p>06:35 The drop offs every ecommerce shop shares</p><p>07:45 The 50% win rate test Cogniteer reused</p><p>09:25 Why alignment beats developer resources</p><p>12:45 Two teams, two goals, one broken checkout</p><p>17:05 Selling water dispensers without a product catalog</p><p>22:15 Designing every experiment to lose</p><p>26:15 Personalizing buyers and where AI takes experimentation</p><p><strong><br>Takeaways</strong></p><p>-Deep dives beat mass produced tests, because understanding one business's users uncovers bigger levers than reusing the same test across many clients.</p><p>-Many ecommerce drop offs are structural, since the basket and product page leak in roughly 80% of shops because it is ecommerce, not because of your product.</p><p>-Product to channel fit decides what sells online, so books and fashion judge well on a screen while perfume and washing machines need cues the interface cannot fully provide.</p><p>-The real bottleneck is alignment, not developer resources, so agree on the problem and its hierarchy before anyone builds a variation.</p><p>-Match the interface to how people actually buy, because new buyers need information, returning buyers want speed, and B2B buyers often want a solution and an offer instead of a product catalog.</p><p><br><strong><br>Connect with the Guest</strong></p><p>LinkedIn: <a href="https://www.linkedin.com/in/fabianhans-cogniteer/">https://www.linkedin.com/in/fabianhans-cogniteer/</a></p><p>Website: <a href="https://www.cogniteer.de/">https://www.cogniteer.de/</a></p><p><strong>Sponsor</strong><br>GrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide.</p><p>Go to <a href="https://www.growthbook.io/?utm_source=edge-podcast&amp;utm_medium=podcast&amp;utm_campaign=episode-all">http://growthbook.io</a></p>]]>
      </description>
      <content:encoded>
        <![CDATA[<p><strong><br>Summary</strong></p><p>In this episode of The Experimentation Edge, host Ashley Stirrup talks with Fabian Hans, founder and behavioral psychologist at Cogniteer, a consultancy that helps enterprises build in-house experimentation programs and raise both test velocity and win rate. Drawing on fifteen years in conversion rate optimization, Fabian explains why mass-producing the same A/B tests across clients quietly kills learning, why most ecommerce drop-offs are structural rather than your fault, and how matching the interface to how people actually buy, new versus returning, B2C versus B2B, can move conversion far more than another button. It is a practical, psychology-grounded conversation for product managers, engineers, data scientists, and growth leaders who want their experimentation programs to compound understanding, not just volume.</p><p><strong><br>Chapters</strong></p><p>00:00 Introduction</p><p>01:25 From agency mass production to in house deep dives</p><p>04:05 Why some products resist selling online</p><p>06:35 The drop offs every ecommerce shop shares</p><p>07:45 The 50% win rate test Cogniteer reused</p><p>09:25 Why alignment beats developer resources</p><p>12:45 Two teams, two goals, one broken checkout</p><p>17:05 Selling water dispensers without a product catalog</p><p>22:15 Designing every experiment to lose</p><p>26:15 Personalizing buyers and where AI takes experimentation</p><p><strong><br>Takeaways</strong></p><p>-Deep dives beat mass produced tests, because understanding one business's users uncovers bigger levers than reusing the same test across many clients.</p><p>-Many ecommerce drop offs are structural, since the basket and product page leak in roughly 80% of shops because it is ecommerce, not because of your product.</p><p>-Product to channel fit decides what sells online, so books and fashion judge well on a screen while perfume and washing machines need cues the interface cannot fully provide.</p><p>-The real bottleneck is alignment, not developer resources, so agree on the problem and its hierarchy before anyone builds a variation.</p><p>-Match the interface to how people actually buy, because new buyers need information, returning buyers want speed, and B2B buyers often want a solution and an offer instead of a product catalog.</p><p><br><strong><br>Connect with the Guest</strong></p><p>LinkedIn: <a href="https://www.linkedin.com/in/fabianhans-cogniteer/">https://www.linkedin.com/in/fabianhans-cogniteer/</a></p><p>Website: <a href="https://www.cogniteer.de/">https://www.cogniteer.de/</a></p><p><strong>Sponsor</strong><br>GrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide.</p><p>Go to <a href="https://www.growthbook.io/?utm_source=edge-podcast&amp;utm_medium=podcast&amp;utm_campaign=episode-all">http://growthbook.io</a></p>]]>
      </content:encoded>
      <pubDate>Thu, 23 Jul 2026 10:31:15 -0600</pubDate>
      <author>Growthbook</author>
      <enclosure url="https://media.transistor.fm/9016b4fd/61d56fba.mp3" length="64027845" type="audio/mpeg"/>
      <itunes:author>Growthbook</itunes:author>
      <itunes:image href="https://img.transistorcdn.com/Gmq8dnjDci5g5UY4PMn---Hcj-Y73woGe2SxwuV56Mw/rs:fill:0:0:1/w:1400/h:1400/q:60/mb:500000/aHR0cHM6Ly9pbWct/dXBsb2FkLXByb2R1/Y3Rpb24udHJhbnNp/c3Rvci5mbS85Njg1/Y2RkMGY3MTdhNTYx/MmQ3MjE2NDZiMGY4/MjdlNi5wbmc.jpg"/>
      <itunes:duration>2000</itunes:duration>
      <itunes:summary>
        <![CDATA[<p><strong><br>Summary</strong></p><p>In this episode of The Experimentation Edge, host Ashley Stirrup talks with Fabian Hans, founder and behavioral psychologist at Cogniteer, a consultancy that helps enterprises build in-house experimentation programs and raise both test velocity and win rate. Drawing on fifteen years in conversion rate optimization, Fabian explains why mass-producing the same A/B tests across clients quietly kills learning, why most ecommerce drop-offs are structural rather than your fault, and how matching the interface to how people actually buy, new versus returning, B2C versus B2B, can move conversion far more than another button. It is a practical, psychology-grounded conversation for product managers, engineers, data scientists, and growth leaders who want their experimentation programs to compound understanding, not just volume.</p><p><strong><br>Chapters</strong></p><p>00:00 Introduction</p><p>01:25 From agency mass production to in house deep dives</p><p>04:05 Why some products resist selling online</p><p>06:35 The drop offs every ecommerce shop shares</p><p>07:45 The 50% win rate test Cogniteer reused</p><p>09:25 Why alignment beats developer resources</p><p>12:45 Two teams, two goals, one broken checkout</p><p>17:05 Selling water dispensers without a product catalog</p><p>22:15 Designing every experiment to lose</p><p>26:15 Personalizing buyers and where AI takes experimentation</p><p><strong><br>Takeaways</strong></p><p>-Deep dives beat mass produced tests, because understanding one business's users uncovers bigger levers than reusing the same test across many clients.</p><p>-Many ecommerce drop offs are structural, since the basket and product page leak in roughly 80% of shops because it is ecommerce, not because of your product.</p><p>-Product to channel fit decides what sells online, so books and fashion judge well on a screen while perfume and washing machines need cues the interface cannot fully provide.</p><p>-The real bottleneck is alignment, not developer resources, so agree on the problem and its hierarchy before anyone builds a variation.</p><p>-Match the interface to how people actually buy, because new buyers need information, returning buyers want speed, and B2B buyers often want a solution and an offer instead of a product catalog.</p><p><br><strong><br>Connect with the Guest</strong></p><p>LinkedIn: <a href="https://www.linkedin.com/in/fabianhans-cogniteer/">https://www.linkedin.com/in/fabianhans-cogniteer/</a></p><p>Website: <a href="https://www.cogniteer.de/">https://www.cogniteer.de/</a></p><p><strong>Sponsor</strong><br>GrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide.</p><p>Go to <a href="https://www.growthbook.io/?utm_source=edge-podcast&amp;utm_medium=podcast&amp;utm_campaign=episode-all">http://growthbook.io</a></p>]]>
      </itunes:summary>
      <itunes:keywords>a/b testing,conversion rate optimization,experimentation program,cro,ecommerce experimentation,test velocity,win rate,cart abandonment,checkout optimization,product detail page optimization,upsell strategy,b2b lead generation,website personalization,user behavior,behavioral psychology,cogniteer,fabian hans,the experimentation edge,growthbook,ron kohavi</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/9016b4fd/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>How Fin A/B Tests Millions of Samples in Days</title>
      <itunes:episode>28</itunes:episode>
      <podcast:episode>28</podcast:episode>
      <itunes:title>How Fin A/B Tests Millions of Samples in Days</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">ddc43f6a-c3ab-4196-9598-14bbcd369c71</guid>
      <link>https://share.transistor.fm/s/0dd7272a</link>
      <description>
        <![CDATA[<p><strong>Summary</strong></p><p>On this episode of The Experimentation Edge, host Ashley Stirrup talks with Pedro Tabacof, Principal Machine Learning Scientist at Fin (formerly Intercom), about how one of the most advanced AI customer support agents in the world is built on relentless experimentation. Pedro explains why unit tests don't work on non-deterministic AI, how Fin runs up to two dozen concurrent A/B tests pulling millions of samples in days, and shares two counterintuitive experiments: one where slowing the agent down improved every metric, and one where adding more context made Fin more helpful and more prone to fake promises until a targeted prompt fix kept the upside without the hallucinations. It's a candid look for product managers, engineers, and data scientists at how a $100M ARR AI product actually ships improvements.<br></p><p><strong>Chapters</strong></p><p>00:00 Welcome and what Fin actually does</p><p>02:00 How Fin became Anthropic's first line of support</p><p>02:30 Why Fin sells resolutions not deflections</p><p>06:00 Owning the stack with custom models</p><p>10:40 Pedro's path from fuzzy logic to AI</p><p>12:55 Why A/B testing is the only gold standard for AI</p><p>15:50 Do no harm testing on every change</p><p>18:00 The latency experiment that shocked the team</p><p>27:30 When more context made Fin hallucinate</p><p>30:15 Win rates and the future of AI driven experimentation<br></p><p><strong>Takeaways</strong></p><p>-Faster is not always better. Fin increased latency artificially and positive feedback went up, likely because a small delay makes an AI feel like it is doing real work.</p><p>-You cannot unit test a non-deterministic AI. A/B testing at scale, millions of samples in days, is the only reliable way to know a change actually helped.</p><p>-Adding more conversation history made Fin more helpful and more prone to fake promises, until a targeted prompt fix removed the hallucinations and kept most of the gain.</p><p>-A losing experiment is often a winner with one broken part. Diagnose which element hurts the experience, fix only that, and rerun.</p><p>-Fin A/B tests everything, even one-character prompt changes and many bug fixes, and treats a 20 to 30 percent win rate as a healthy sign of a real experimentation program.<br></p><p><strong>Connect with the Guest</strong></p><p>LinkedIn: <a href="https://www.linkedin.com/in/tabacof/">https://www.linkedin.com/in/tabacof/</a><br> Website: <a href="https://fin.ai">https://fin.ai</a></p><p><strong>Sponsor</strong><br>GrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide.</p><p>Go to <a href="https://www.growthbook.io/?utm_source=edge-podcast&amp;utm_medium=podcast&amp;utm_campaign=episode-all">http://growthbook.io</a></p>]]>
      </description>
      <content:encoded>
        <![CDATA[<p><strong>Summary</strong></p><p>On this episode of The Experimentation Edge, host Ashley Stirrup talks with Pedro Tabacof, Principal Machine Learning Scientist at Fin (formerly Intercom), about how one of the most advanced AI customer support agents in the world is built on relentless experimentation. Pedro explains why unit tests don't work on non-deterministic AI, how Fin runs up to two dozen concurrent A/B tests pulling millions of samples in days, and shares two counterintuitive experiments: one where slowing the agent down improved every metric, and one where adding more context made Fin more helpful and more prone to fake promises until a targeted prompt fix kept the upside without the hallucinations. It's a candid look for product managers, engineers, and data scientists at how a $100M ARR AI product actually ships improvements.<br></p><p><strong>Chapters</strong></p><p>00:00 Welcome and what Fin actually does</p><p>02:00 How Fin became Anthropic's first line of support</p><p>02:30 Why Fin sells resolutions not deflections</p><p>06:00 Owning the stack with custom models</p><p>10:40 Pedro's path from fuzzy logic to AI</p><p>12:55 Why A/B testing is the only gold standard for AI</p><p>15:50 Do no harm testing on every change</p><p>18:00 The latency experiment that shocked the team</p><p>27:30 When more context made Fin hallucinate</p><p>30:15 Win rates and the future of AI driven experimentation<br></p><p><strong>Takeaways</strong></p><p>-Faster is not always better. Fin increased latency artificially and positive feedback went up, likely because a small delay makes an AI feel like it is doing real work.</p><p>-You cannot unit test a non-deterministic AI. A/B testing at scale, millions of samples in days, is the only reliable way to know a change actually helped.</p><p>-Adding more conversation history made Fin more helpful and more prone to fake promises, until a targeted prompt fix removed the hallucinations and kept most of the gain.</p><p>-A losing experiment is often a winner with one broken part. Diagnose which element hurts the experience, fix only that, and rerun.</p><p>-Fin A/B tests everything, even one-character prompt changes and many bug fixes, and treats a 20 to 30 percent win rate as a healthy sign of a real experimentation program.<br></p><p><strong>Connect with the Guest</strong></p><p>LinkedIn: <a href="https://www.linkedin.com/in/tabacof/">https://www.linkedin.com/in/tabacof/</a><br> Website: <a href="https://fin.ai">https://fin.ai</a></p><p><strong>Sponsor</strong><br>GrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide.</p><p>Go to <a href="https://www.growthbook.io/?utm_source=edge-podcast&amp;utm_medium=podcast&amp;utm_campaign=episode-all">http://growthbook.io</a></p>]]>
      </content:encoded>
      <pubDate>Tue, 21 Jul 2026 08:51:38 -0600</pubDate>
      <author>Growthbook</author>
      <enclosure url="https://media.transistor.fm/0dd7272a/fa5786be.mp3" length="87512778" type="audio/mpeg"/>
      <itunes:author>Growthbook</itunes:author>
      <itunes:image href="https://img.transistorcdn.com/0EDFROubI1PE3JXqp1E-VYS-IiXlzaqXFJaG57btTQM/rs:fill:0:0:1/w:1400/h:1400/q:60/mb:500000/aHR0cHM6Ly9pbWct/dXBsb2FkLXByb2R1/Y3Rpb24udHJhbnNp/c3Rvci5mbS83Nzg1/OTQ3NDdmNGFmOTBh/ZjAzMzVkNGE4YmM1/NzcxOS5wbmc.jpg"/>
      <itunes:duration>2734</itunes:duration>
      <itunes:summary>
        <![CDATA[<p><strong>Summary</strong></p><p>On this episode of The Experimentation Edge, host Ashley Stirrup talks with Pedro Tabacof, Principal Machine Learning Scientist at Fin (formerly Intercom), about how one of the most advanced AI customer support agents in the world is built on relentless experimentation. Pedro explains why unit tests don't work on non-deterministic AI, how Fin runs up to two dozen concurrent A/B tests pulling millions of samples in days, and shares two counterintuitive experiments: one where slowing the agent down improved every metric, and one where adding more context made Fin more helpful and more prone to fake promises until a targeted prompt fix kept the upside without the hallucinations. It's a candid look for product managers, engineers, and data scientists at how a $100M ARR AI product actually ships improvements.<br></p><p><strong>Chapters</strong></p><p>00:00 Welcome and what Fin actually does</p><p>02:00 How Fin became Anthropic's first line of support</p><p>02:30 Why Fin sells resolutions not deflections</p><p>06:00 Owning the stack with custom models</p><p>10:40 Pedro's path from fuzzy logic to AI</p><p>12:55 Why A/B testing is the only gold standard for AI</p><p>15:50 Do no harm testing on every change</p><p>18:00 The latency experiment that shocked the team</p><p>27:30 When more context made Fin hallucinate</p><p>30:15 Win rates and the future of AI driven experimentation<br></p><p><strong>Takeaways</strong></p><p>-Faster is not always better. Fin increased latency artificially and positive feedback went up, likely because a small delay makes an AI feel like it is doing real work.</p><p>-You cannot unit test a non-deterministic AI. A/B testing at scale, millions of samples in days, is the only reliable way to know a change actually helped.</p><p>-Adding more conversation history made Fin more helpful and more prone to fake promises, until a targeted prompt fix removed the hallucinations and kept most of the gain.</p><p>-A losing experiment is often a winner with one broken part. Diagnose which element hurts the experience, fix only that, and rerun.</p><p>-Fin A/B tests everything, even one-character prompt changes and many bug fixes, and treats a 20 to 30 percent win rate as a healthy sign of a real experimentation program.<br></p><p><strong>Connect with the Guest</strong></p><p>LinkedIn: <a href="https://www.linkedin.com/in/tabacof/">https://www.linkedin.com/in/tabacof/</a><br> Website: <a href="https://fin.ai">https://fin.ai</a></p><p><strong>Sponsor</strong><br>GrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide.</p><p>Go to <a href="https://www.growthbook.io/?utm_source=edge-podcast&amp;utm_medium=podcast&amp;utm_campaign=episode-all">http://growthbook.io</a></p>]]>
      </itunes:summary>
      <itunes:keywords>a/b testing,experimentation,fin,intercom,ai customer support,ai agents,llm evaluation,hallucinations,guardrail metrics,resolution rate,statistical significance,latency experiment,feature flags,custom models,prompt engineering,experimentation culture,product experimentation,machine learning,growthbook,the experimentation edge</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/0dd7272a/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>How Kargo turns losing experiments into competitive edges</title>
      <itunes:episode>27</itunes:episode>
      <podcast:episode>27</podcast:episode>
      <itunes:title>How Kargo turns losing experiments into competitive edges</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">e626b5a9-7a29-4248-9950-43f4945bf8ce</guid>
      <link>https://share.transistor.fm/s/7263f7c1</link>
      <description>
        <![CDATA[<p><strong><br>Summary</strong></p><p>In this episode of The Experimentation Edge, host Ashley Stirrup, CMO of GrowthBook, sits down with James Falzone, Director of Product Management at Kargo, to unpack how a high scale ad tech marketplace turns failure into its biggest advantage. James explains how Kargo connects advertisers to publishers through real time auctions that resolve in milliseconds across up to 10 billion ad requests a day, why experimentation is embedded in the company's culture rather than siloed in a team, and what happened when a winning click optimization model failed completely after being copied to a new customer type. The conversation is built for product managers, data scientists, engineers, and growth leaders who want a practical, honest view of running experiments at scale, learning from losses, and keeping AI grounded in solid infrastructure.</p><p><strong><br>Chapters</strong></p><p>00:00 Welcome and introducing James Falzone</p><p>01:45 What Kargo does and how real time ad auctions work</p><p>04:45 Why experimentation is embedded in Kargo's culture</p><p>07:45 The three things every marketplace has to deliver</p><p>10:15 The experiment that failed: click optimization on third party demand</p><p>12:15 A bad result versus a bad experiment</p><p>13:45 Why different customer types need different signals</p><p>15:30 Putting "where did you fail?" on every retro</p><p>18:45 How experimentation evolves with AI</p><p>21:15 Better not bigger: the closing takeaway</p><p><strong><br>Takeaways</strong></p><p>-A bad result is not a bad experiment. If you're not failing, you're probably not trying anything new.</p><p>-The same metrics and signals don't apply to every customer type. Bad results often come from a lack of context, not bad tech.</p><p>-Metrics and signals you test against should always be business driven, not ported from the last thing that worked.</p><p>-Put failure on the agenda. A biweekly "where did you fail?" retro turns one person's dead end into the whole team's shortcut.</p><p>-AI's biggest unlock is access. More people can run experiments, but it has to be built on solid ML and infrastructure. Better, not bigger.</p><p><strong><br>Connect with the Guest</strong></p><p>LinkedIn: <a href="https://www.linkedin.com/in/jamesafalzone/">https://www.linkedin.com/in/jamesafalzone/</a></p><p>Website: <a href="https://kargo.com">https://kargo.com</a></p><p><strong>Sponsor</strong><br>GrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide.</p><p>Go to <a href="https://www.growthbook.io/?utm_source=edge-podcast&amp;utm_medium=podcast&amp;utm_campaign=episode-all">http://growthbook.io</a></p>]]>
      </description>
      <content:encoded>
        <![CDATA[<p><strong><br>Summary</strong></p><p>In this episode of The Experimentation Edge, host Ashley Stirrup, CMO of GrowthBook, sits down with James Falzone, Director of Product Management at Kargo, to unpack how a high scale ad tech marketplace turns failure into its biggest advantage. James explains how Kargo connects advertisers to publishers through real time auctions that resolve in milliseconds across up to 10 billion ad requests a day, why experimentation is embedded in the company's culture rather than siloed in a team, and what happened when a winning click optimization model failed completely after being copied to a new customer type. The conversation is built for product managers, data scientists, engineers, and growth leaders who want a practical, honest view of running experiments at scale, learning from losses, and keeping AI grounded in solid infrastructure.</p><p><strong><br>Chapters</strong></p><p>00:00 Welcome and introducing James Falzone</p><p>01:45 What Kargo does and how real time ad auctions work</p><p>04:45 Why experimentation is embedded in Kargo's culture</p><p>07:45 The three things every marketplace has to deliver</p><p>10:15 The experiment that failed: click optimization on third party demand</p><p>12:15 A bad result versus a bad experiment</p><p>13:45 Why different customer types need different signals</p><p>15:30 Putting "where did you fail?" on every retro</p><p>18:45 How experimentation evolves with AI</p><p>21:15 Better not bigger: the closing takeaway</p><p><strong><br>Takeaways</strong></p><p>-A bad result is not a bad experiment. If you're not failing, you're probably not trying anything new.</p><p>-The same metrics and signals don't apply to every customer type. Bad results often come from a lack of context, not bad tech.</p><p>-Metrics and signals you test against should always be business driven, not ported from the last thing that worked.</p><p>-Put failure on the agenda. A biweekly "where did you fail?" retro turns one person's dead end into the whole team's shortcut.</p><p>-AI's biggest unlock is access. More people can run experiments, but it has to be built on solid ML and infrastructure. Better, not bigger.</p><p><strong><br>Connect with the Guest</strong></p><p>LinkedIn: <a href="https://www.linkedin.com/in/jamesafalzone/">https://www.linkedin.com/in/jamesafalzone/</a></p><p>Website: <a href="https://kargo.com">https://kargo.com</a></p><p><strong>Sponsor</strong><br>GrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide.</p><p>Go to <a href="https://www.growthbook.io/?utm_source=edge-podcast&amp;utm_medium=podcast&amp;utm_campaign=episode-all">http://growthbook.io</a></p>]]>
      </content:encoded>
      <pubDate>Tue, 14 Jul 2026 08:32:16 -0600</pubDate>
      <author>Growthbook</author>
      <enclosure url="https://media.transistor.fm/7263f7c1/f0fb5f51.mp3" length="42354584" type="audio/mpeg"/>
      <itunes:author>Growthbook</itunes:author>
      <itunes:image href="https://img.transistorcdn.com/B8ynlvMRh3nrGRqxAN2mwfTuJ8tuD7FJDdZSllVgBiA/rs:fill:0:0:1/w:1400/h:1400/q:60/mb:500000/aHR0cHM6Ly9pbWct/dXBsb2FkLXByb2R1/Y3Rpb24udHJhbnNp/c3Rvci5mbS84MjQ5/M2Q3YTA5NmJkYjU1/NTBmNzFiZDFjNjI5/NzBjMi5wbmc.jpg"/>
      <itunes:duration>1323</itunes:duration>
      <itunes:summary>
        <![CDATA[<p><strong><br>Summary</strong></p><p>In this episode of The Experimentation Edge, host Ashley Stirrup, CMO of GrowthBook, sits down with James Falzone, Director of Product Management at Kargo, to unpack how a high scale ad tech marketplace turns failure into its biggest advantage. James explains how Kargo connects advertisers to publishers through real time auctions that resolve in milliseconds across up to 10 billion ad requests a day, why experimentation is embedded in the company's culture rather than siloed in a team, and what happened when a winning click optimization model failed completely after being copied to a new customer type. The conversation is built for product managers, data scientists, engineers, and growth leaders who want a practical, honest view of running experiments at scale, learning from losses, and keeping AI grounded in solid infrastructure.</p><p><strong><br>Chapters</strong></p><p>00:00 Welcome and introducing James Falzone</p><p>01:45 What Kargo does and how real time ad auctions work</p><p>04:45 Why experimentation is embedded in Kargo's culture</p><p>07:45 The three things every marketplace has to deliver</p><p>10:15 The experiment that failed: click optimization on third party demand</p><p>12:15 A bad result versus a bad experiment</p><p>13:45 Why different customer types need different signals</p><p>15:30 Putting "where did you fail?" on every retro</p><p>18:45 How experimentation evolves with AI</p><p>21:15 Better not bigger: the closing takeaway</p><p><strong><br>Takeaways</strong></p><p>-A bad result is not a bad experiment. If you're not failing, you're probably not trying anything new.</p><p>-The same metrics and signals don't apply to every customer type. Bad results often come from a lack of context, not bad tech.</p><p>-Metrics and signals you test against should always be business driven, not ported from the last thing that worked.</p><p>-Put failure on the agenda. A biweekly "where did you fail?" retro turns one person's dead end into the whole team's shortcut.</p><p>-AI's biggest unlock is access. More people can run experiments, but it has to be built on solid ML and infrastructure. Better, not bigger.</p><p><strong><br>Connect with the Guest</strong></p><p>LinkedIn: <a href="https://www.linkedin.com/in/jamesafalzone/">https://www.linkedin.com/in/jamesafalzone/</a></p><p>Website: <a href="https://kargo.com">https://kargo.com</a></p><p><strong>Sponsor</strong><br>GrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide.</p><p>Go to <a href="https://www.growthbook.io/?utm_source=edge-podcast&amp;utm_medium=podcast&amp;utm_campaign=episode-all">http://growthbook.io</a></p>]]>
      </itunes:summary>
      <itunes:keywords>a/b testing,experimentation,product management,reverse trial,conversion rate optimization,b2b saas,north star metrics,okrs,guardrail metrics,feature flags,statistical significance,trial conversion,paywall optimization,freemium,product analytics,AI in experimentation,uber eats,codecademy,diligent,the experimentation edge</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/7263f7c1/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>The 'wine effect' and other surprises that reshaped how Box runs e-commerce experiments</title>
      <itunes:episode>28</itunes:episode>
      <podcast:episode>28</podcast:episode>
      <itunes:title>The 'wine effect' and other surprises that reshaped how Box runs e-commerce experiments</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">4aab842c-d786-46ea-a376-ed1b24657f72</guid>
      <link>https://share.transistor.fm/s/b226edb1</link>
      <description>
        <![CDATA[<p><strong>Summary</strong></p><p>In this episode of The Experimentation Edge, host Ashley Stirrup talks with Danielle Oleen, Director of E-commerce at Box, about what it really takes to build a culture of experimentation inside a B2B company. Drawing on 15 years across B2C and B2B at Wayfair, Drizly, Zoom, and now Box, Danielle explains why experimentation belongs to every product team and not just e-commerce, walks through a pricing page saga of one win and two losses that exposed the limits of simplification, and shares the "wine effect" test that won for a reason no one predicted. It's a practical, story rich conversation for product managers, growth leaders, and anyone trying to make better decisions with data.</p><p><strong>Chapters</strong></p><p>00:45 Meet Danielle Oleen and Box's reinvention</p><p>02:45 Owning the entire customer life cycle</p><p>04:45 Why experimentation matters even without a checkout</p><p>07:45 The feature that's used but hidden</p><p>11:45 Proving ROI with a scrappy manual test</p><p>12:45 Building a culture that shares wins and losses</p><p>16:45 The pyramid strategy for prioritizing tests</p><p>18:45 The simplification tightrope on the pricing page</p><p>24:45 When a test wins for the wrong reason</p><p>27:45 Where experimentation at Box goes next</p><p><strong>Takeaways</strong></p><p>- Experimentation isn't only for e-commerce. Any product with a funnel, even an AI chatbot, can be measured and improved through testing.</p><p>- Simplification has a limit. Removing too much can strip away the cues and context buyers actually need to decide.</p><p>- Share losses as openly as wins. Wins build credibility, and losses build the psychological safety a testing culture runs on.</p><p>- Prioritize like a pyramid. Fix the widest-impact experiences first, then optimize down into smaller cohorts.</p><p>- Surprising results are the point. A test can win for a reason you never hypothesized, like the "wine effect," and that's where the real learning lives.</p><p><strong>Connect with the Guest</strong></p><p>LinkedIn: <a href="https://www.linkedin.com/in/dolean1/">https://www.linkedin.com/in/dolean1/</a></p><p>Website: <a href="https://www.box.com">https://www.box.com</a></p><p><strong>Sponsor</strong><br>GrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide.</p><p>Go to <a href="https://www.growthbook.io/?utm_source=edge-podcast&amp;utm_medium=podcast&amp;utm_campaign=episode-all">http://growthbook.io</a></p>]]>
      </description>
      <content:encoded>
        <![CDATA[<p><strong>Summary</strong></p><p>In this episode of The Experimentation Edge, host Ashley Stirrup talks with Danielle Oleen, Director of E-commerce at Box, about what it really takes to build a culture of experimentation inside a B2B company. Drawing on 15 years across B2C and B2B at Wayfair, Drizly, Zoom, and now Box, Danielle explains why experimentation belongs to every product team and not just e-commerce, walks through a pricing page saga of one win and two losses that exposed the limits of simplification, and shares the "wine effect" test that won for a reason no one predicted. It's a practical, story rich conversation for product managers, growth leaders, and anyone trying to make better decisions with data.</p><p><strong>Chapters</strong></p><p>00:45 Meet Danielle Oleen and Box's reinvention</p><p>02:45 Owning the entire customer life cycle</p><p>04:45 Why experimentation matters even without a checkout</p><p>07:45 The feature that's used but hidden</p><p>11:45 Proving ROI with a scrappy manual test</p><p>12:45 Building a culture that shares wins and losses</p><p>16:45 The pyramid strategy for prioritizing tests</p><p>18:45 The simplification tightrope on the pricing page</p><p>24:45 When a test wins for the wrong reason</p><p>27:45 Where experimentation at Box goes next</p><p><strong>Takeaways</strong></p><p>- Experimentation isn't only for e-commerce. Any product with a funnel, even an AI chatbot, can be measured and improved through testing.</p><p>- Simplification has a limit. Removing too much can strip away the cues and context buyers actually need to decide.</p><p>- Share losses as openly as wins. Wins build credibility, and losses build the psychological safety a testing culture runs on.</p><p>- Prioritize like a pyramid. Fix the widest-impact experiences first, then optimize down into smaller cohorts.</p><p>- Surprising results are the point. A test can win for a reason you never hypothesized, like the "wine effect," and that's where the real learning lives.</p><p><strong>Connect with the Guest</strong></p><p>LinkedIn: <a href="https://www.linkedin.com/in/dolean1/">https://www.linkedin.com/in/dolean1/</a></p><p>Website: <a href="https://www.box.com">https://www.box.com</a></p><p><strong>Sponsor</strong><br>GrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide.</p><p>Go to <a href="https://www.growthbook.io/?utm_source=edge-podcast&amp;utm_medium=podcast&amp;utm_campaign=episode-all">http://growthbook.io</a></p>]]>
      </content:encoded>
      <pubDate>Thu, 09 Jul 2026 06:18:32 -0600</pubDate>
      <author>Growthbook</author>
      <enclosure url="https://media.transistor.fm/b226edb1/3b6bc40b.mp3" length="60033844" type="audio/mpeg"/>
      <itunes:author>Growthbook</itunes:author>
      <itunes:image href="https://img.transistorcdn.com/-CYmUHEfDrKCDSCHAssonKNDDlaXstEj1JAOpiI8fwM/rs:fill:0:0:1/w:1400/h:1400/q:60/mb:500000/aHR0cHM6Ly9pbWct/dXBsb2FkLXByb2R1/Y3Rpb24udHJhbnNp/c3Rvci5mbS8yZDUz/MGQ0ZGU1NzMwYjIy/ZDA0MjE2Njg3MmMx/NTI3Yi5wbmc.jpg"/>
      <itunes:duration>1875</itunes:duration>
      <itunes:summary>
        <![CDATA[<p><strong>Summary</strong></p><p>In this episode of The Experimentation Edge, host Ashley Stirrup talks with Danielle Oleen, Director of E-commerce at Box, about what it really takes to build a culture of experimentation inside a B2B company. Drawing on 15 years across B2C and B2B at Wayfair, Drizly, Zoom, and now Box, Danielle explains why experimentation belongs to every product team and not just e-commerce, walks through a pricing page saga of one win and two losses that exposed the limits of simplification, and shares the "wine effect" test that won for a reason no one predicted. It's a practical, story rich conversation for product managers, growth leaders, and anyone trying to make better decisions with data.</p><p><strong>Chapters</strong></p><p>00:45 Meet Danielle Oleen and Box's reinvention</p><p>02:45 Owning the entire customer life cycle</p><p>04:45 Why experimentation matters even without a checkout</p><p>07:45 The feature that's used but hidden</p><p>11:45 Proving ROI with a scrappy manual test</p><p>12:45 Building a culture that shares wins and losses</p><p>16:45 The pyramid strategy for prioritizing tests</p><p>18:45 The simplification tightrope on the pricing page</p><p>24:45 When a test wins for the wrong reason</p><p>27:45 Where experimentation at Box goes next</p><p><strong>Takeaways</strong></p><p>- Experimentation isn't only for e-commerce. Any product with a funnel, even an AI chatbot, can be measured and improved through testing.</p><p>- Simplification has a limit. Removing too much can strip away the cues and context buyers actually need to decide.</p><p>- Share losses as openly as wins. Wins build credibility, and losses build the psychological safety a testing culture runs on.</p><p>- Prioritize like a pyramid. Fix the widest-impact experiences first, then optimize down into smaller cohorts.</p><p>- Surprising results are the point. A test can win for a reason you never hypothesized, like the "wine effect," and that's where the real learning lives.</p><p><strong>Connect with the Guest</strong></p><p>LinkedIn: <a href="https://www.linkedin.com/in/dolean1/">https://www.linkedin.com/in/dolean1/</a></p><p>Website: <a href="https://www.box.com">https://www.box.com</a></p><p><strong>Sponsor</strong><br>GrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide.</p><p>Go to <a href="https://www.growthbook.io/?utm_source=edge-podcast&amp;utm_medium=podcast&amp;utm_campaign=episode-all">http://growthbook.io</a></p>]]>
      </itunes:summary>
      <itunes:keywords>a/b testing,experimentation,b2b ecommerce,conversion rate optimization,pricing page optimization,product experimentation,growthbook,the experimentation edge,box,checkout optimization,cognitive overload,experimentation culture,revenue growth,churn reduction,product management,self-service saas,average order value,customer lifecycle,monetization testing,ai in experimentation</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/b226edb1/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Dilligent explains why moving on from an experiment might cost you</title>
      <itunes:episode>26</itunes:episode>
      <podcast:episode>26</podcast:episode>
      <itunes:title>Dilligent explains why moving on from an experiment might cost you</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">92d2eac6-45c4-4dfb-835e-a64a137e9e5b</guid>
      <link>https://share.transistor.fm/s/af75d403</link>
      <description>
        <![CDATA[<p><strong>Summary</strong></p><p>Dan Layfield, Director of Product Management at Diligent, joins host Ashley Stirrup on The Experimentation Edge to trace what fifteen years of A/B testing across Codecademy, Uber Eats, and the Fortune 1000 boardroom actually taught him. He breaks down the Codecademy trial-model rebuild that took four months and several rounds to deliver a 35% conversion lift, why moving on from a losing experiment too early is one of a PM's costliest mistakes, how to escape the B2B feature factory with metrics that genuinely ladder up, why retention should ride a product's natural use case instead of fighting it, and where AI is already replacing weeks of research and analysis. It's a practitioner's guide for product managers, growth leaders, data scientists, and engineers bringing experimentation rigor to both B2C and B2B.</p><p><br><strong><br>Chapters</strong></p><p>00:45 Meet Dan Layfield and Diligent</p><p>01:45 Two worlds of experimentation, Codecademy and Uber</p><p>03:45 The trial model that lifted conversion 35%</p><p>06:20 What to do with a losing experiment</p><p>08:50 Two flavors of experimentation</p><p>09:45 Reading forty metrics at Uber Eats</p><p>13:10 Escaping the B2B feature factory</p><p>16:45 Anchoring the North Star to real usage</p><p>19:15 Where AI fits in research and analysis</p><p><br><strong><br>Takeaways</strong></p><ul><li>A losing experiment is often inconclusive, not negative; treat it as a map of the funnel rather than a verdict, and know when a big problem is worth another round.</li><li>Persistence paid off at Codecademy: four months and three to four rounds of trial-model testing produced a 35% conversion increase.</li><li>Separate your two experimentation modes; high-volume CRO chases many small wins, while big, uncertain bets are worth taking multiple shots to de-risk.</li><li>Most B2B product teams are feature factories; the fix is a top-down OKR system, and planning usually breaks in the connections between layers, not inside them.</li><li>Anchor retention and engagement to the product's natural use case, and use AI to synthesize research and simple A/B analysis in hours instead of weeks.</li></ul><p><br><strong><br>Connect with the Guest</strong></p><p>LinkedIn: <a href="https://www.linkedin.com/in/layfield/">https://www.linkedin.com/in/layfield/</a></p><p>Website: <a href="https://www.diligent.com">https://www.diligent.com</a></p><p><strong>Sponsor</strong><br>GrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide.</p><p>Go to <a href="http://growthbook.io/?utm_source=edge-podcast&amp;utm_medium=podcast&amp;utm_campaign=episode-#">http://growthbook.io</a></p>]]>
      </description>
      <content:encoded>
        <![CDATA[<p><strong>Summary</strong></p><p>Dan Layfield, Director of Product Management at Diligent, joins host Ashley Stirrup on The Experimentation Edge to trace what fifteen years of A/B testing across Codecademy, Uber Eats, and the Fortune 1000 boardroom actually taught him. He breaks down the Codecademy trial-model rebuild that took four months and several rounds to deliver a 35% conversion lift, why moving on from a losing experiment too early is one of a PM's costliest mistakes, how to escape the B2B feature factory with metrics that genuinely ladder up, why retention should ride a product's natural use case instead of fighting it, and where AI is already replacing weeks of research and analysis. It's a practitioner's guide for product managers, growth leaders, data scientists, and engineers bringing experimentation rigor to both B2C and B2B.</p><p><br><strong><br>Chapters</strong></p><p>00:45 Meet Dan Layfield and Diligent</p><p>01:45 Two worlds of experimentation, Codecademy and Uber</p><p>03:45 The trial model that lifted conversion 35%</p><p>06:20 What to do with a losing experiment</p><p>08:50 Two flavors of experimentation</p><p>09:45 Reading forty metrics at Uber Eats</p><p>13:10 Escaping the B2B feature factory</p><p>16:45 Anchoring the North Star to real usage</p><p>19:15 Where AI fits in research and analysis</p><p><br><strong><br>Takeaways</strong></p><ul><li>A losing experiment is often inconclusive, not negative; treat it as a map of the funnel rather than a verdict, and know when a big problem is worth another round.</li><li>Persistence paid off at Codecademy: four months and three to four rounds of trial-model testing produced a 35% conversion increase.</li><li>Separate your two experimentation modes; high-volume CRO chases many small wins, while big, uncertain bets are worth taking multiple shots to de-risk.</li><li>Most B2B product teams are feature factories; the fix is a top-down OKR system, and planning usually breaks in the connections between layers, not inside them.</li><li>Anchor retention and engagement to the product's natural use case, and use AI to synthesize research and simple A/B analysis in hours instead of weeks.</li></ul><p><br><strong><br>Connect with the Guest</strong></p><p>LinkedIn: <a href="https://www.linkedin.com/in/layfield/">https://www.linkedin.com/in/layfield/</a></p><p>Website: <a href="https://www.diligent.com">https://www.diligent.com</a></p><p><strong>Sponsor</strong><br>GrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide.</p><p>Go to <a href="http://growthbook.io/?utm_source=edge-podcast&amp;utm_medium=podcast&amp;utm_campaign=episode-#">http://growthbook.io</a></p>]]>
      </content:encoded>
      <pubDate>Tue, 07 Jul 2026 09:09:48 -0600</pubDate>
      <author>Growthbook</author>
      <enclosure url="https://media.transistor.fm/af75d403/19cabc08.mp3" length="41990283" type="audio/mpeg"/>
      <itunes:author>Growthbook</itunes:author>
      <itunes:image href="https://img.transistorcdn.com/NYpPHhN-59994gBmnC7BuGeRlnlRDhjrxRpGp54A0SQ/rs:fill:0:0:1/w:1400/h:1400/q:60/mb:500000/aHR0cHM6Ly9pbWct/dXBsb2FkLXByb2R1/Y3Rpb24udHJhbnNp/c3Rvci5mbS8zNGJk/MWEzOGUyMzRjMzgz/ZjBiZGVmODNjZWM5/NWViNi5wbmc.jpg"/>
      <itunes:duration>1311</itunes:duration>
      <itunes:summary>
        <![CDATA[<p><strong>Summary</strong></p><p>Dan Layfield, Director of Product Management at Diligent, joins host Ashley Stirrup on The Experimentation Edge to trace what fifteen years of A/B testing across Codecademy, Uber Eats, and the Fortune 1000 boardroom actually taught him. He breaks down the Codecademy trial-model rebuild that took four months and several rounds to deliver a 35% conversion lift, why moving on from a losing experiment too early is one of a PM's costliest mistakes, how to escape the B2B feature factory with metrics that genuinely ladder up, why retention should ride a product's natural use case instead of fighting it, and where AI is already replacing weeks of research and analysis. It's a practitioner's guide for product managers, growth leaders, data scientists, and engineers bringing experimentation rigor to both B2C and B2B.</p><p><br><strong><br>Chapters</strong></p><p>00:45 Meet Dan Layfield and Diligent</p><p>01:45 Two worlds of experimentation, Codecademy and Uber</p><p>03:45 The trial model that lifted conversion 35%</p><p>06:20 What to do with a losing experiment</p><p>08:50 Two flavors of experimentation</p><p>09:45 Reading forty metrics at Uber Eats</p><p>13:10 Escaping the B2B feature factory</p><p>16:45 Anchoring the North Star to real usage</p><p>19:15 Where AI fits in research and analysis</p><p><br><strong><br>Takeaways</strong></p><ul><li>A losing experiment is often inconclusive, not negative; treat it as a map of the funnel rather than a verdict, and know when a big problem is worth another round.</li><li>Persistence paid off at Codecademy: four months and three to four rounds of trial-model testing produced a 35% conversion increase.</li><li>Separate your two experimentation modes; high-volume CRO chases many small wins, while big, uncertain bets are worth taking multiple shots to de-risk.</li><li>Most B2B product teams are feature factories; the fix is a top-down OKR system, and planning usually breaks in the connections between layers, not inside them.</li><li>Anchor retention and engagement to the product's natural use case, and use AI to synthesize research and simple A/B analysis in hours instead of weeks.</li></ul><p><br><strong><br>Connect with the Guest</strong></p><p>LinkedIn: <a href="https://www.linkedin.com/in/layfield/">https://www.linkedin.com/in/layfield/</a></p><p>Website: <a href="https://www.diligent.com">https://www.diligent.com</a></p><p><strong>Sponsor</strong><br>GrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide.</p><p>Go to <a href="http://growthbook.io/?utm_source=edge-podcast&amp;utm_medium=podcast&amp;utm_campaign=episode-#">http://growthbook.io</a></p>]]>
      </itunes:summary>
      <itunes:keywords>a/b testing,experimentation,product management,reverse trial,conversion rate optimization,b2b saas,north star metrics,okrs,guardrail metrics,feature flags,statistical significance,trial conversion,paywall optimization,freemium,product analytics,AI in experimentation,uber eats,codecademy,diligent,the experimentation edge</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/af75d403/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>The metric Stitch Fix says every experimenter should chase</title>
      <itunes:episode>24</itunes:episode>
      <podcast:episode>24</podcast:episode>
      <itunes:title>The metric Stitch Fix says every experimenter should chase</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">91a99f9a-7c35-4176-93d8-1ce8999b5115</guid>
      <link>https://share.transistor.fm/s/67dbaad3</link>
      <description>
        <![CDATA[<p><strong>Summary</strong></p><p>In this episode of The Experimentation Edge, GrowthBook CMO Ashley Stirrup sits down with Nick Beyler, data science manager at Stitch Fix, where he leads the decision and insights team and owns the company's internal experimentation platform. Nick shares why the metric he most wants is the one he can't measure yet, a North Star that predicts a client's long-term value from their earliest behaviors, and why the most impactful experiment learnings tend to come from adoption friction rather than product bugs. He makes the case that if you're only testing winners you're not taking enough risks, explains how guardrails make that risk safe, and looks ahead to a new in-house platform and the promise of agentic AI. It's a practical, statistician's-eye view of experimentation for product managers, data scientists, and engineers building serious testing programs.</p><p><br><strong><br>Chapters</strong></p><p>00:00 Cold open and welcome to the show</p><p>01:45 What Stitch Fix actually does</p><p>04:15 Balancing AI with the human stylist</p><p>05:15 From public policy to the A/B testing adrenaline rush</p><p>07:15 Inside the weekly experimentation review group</p><p>08:45 The AI style assistant and listening to qualitative feedback</p><p>10:45 Why adoption friction beats product bugs</p><p>13:45 Testing for losers and building guardrails</p><p>15:45 Keep rate, successful fixes, and the holy grail metric</p><p>18:15 The new platform and the promise of agentic AI</p><p><br><strong><br>Takeaways</strong></p><ul><li>The most impactful experiment learnings usually come from adoption friction, not product bugs. By the time a big feature reaches A/B testing, it's often already a winner, so the open question is how and where to introduce it.</li><li>A losing test is a finding, not a failure. If every experiment wins, you're not taking enough risk to learn anything new.</li><li>Guardrails and stopping criteria are what make risk-taking safe, especially when the experience is as personal as shopping.</li><li>The most valuable North Star metric is the one you can't measure yet, long-term client value, and causal-inference modeling helps predict it from short-term behavior.</li><li>Quantitative results are only half the story. Direct, qualitative client feedback inside an experiment often reshapes the rollout more than the numbers do.</li></ul><p><strong><br>Connect with the Guest</strong></p><p>LinkedIn: <a href="https://www.linkedin.com/in/nick-beyler-381864119/">https://www.linkedin.com/in/nick-beyler-381864119/</a> <br>Website: <a href="https://www.stitchfix.com">https://www.stitchfix.com</a></p><p><strong>Sponsor</strong><br>GrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide.</p><p>Go to <a href="http://growthbook.io?utm_source=edge-podcast&amp;utm_medium=podcast&amp;utm_campaign=episode-25">http://growthbook.io</a></p>]]>
      </description>
      <content:encoded>
        <![CDATA[<p><strong>Summary</strong></p><p>In this episode of The Experimentation Edge, GrowthBook CMO Ashley Stirrup sits down with Nick Beyler, data science manager at Stitch Fix, where he leads the decision and insights team and owns the company's internal experimentation platform. Nick shares why the metric he most wants is the one he can't measure yet, a North Star that predicts a client's long-term value from their earliest behaviors, and why the most impactful experiment learnings tend to come from adoption friction rather than product bugs. He makes the case that if you're only testing winners you're not taking enough risks, explains how guardrails make that risk safe, and looks ahead to a new in-house platform and the promise of agentic AI. It's a practical, statistician's-eye view of experimentation for product managers, data scientists, and engineers building serious testing programs.</p><p><br><strong><br>Chapters</strong></p><p>00:00 Cold open and welcome to the show</p><p>01:45 What Stitch Fix actually does</p><p>04:15 Balancing AI with the human stylist</p><p>05:15 From public policy to the A/B testing adrenaline rush</p><p>07:15 Inside the weekly experimentation review group</p><p>08:45 The AI style assistant and listening to qualitative feedback</p><p>10:45 Why adoption friction beats product bugs</p><p>13:45 Testing for losers and building guardrails</p><p>15:45 Keep rate, successful fixes, and the holy grail metric</p><p>18:15 The new platform and the promise of agentic AI</p><p><br><strong><br>Takeaways</strong></p><ul><li>The most impactful experiment learnings usually come from adoption friction, not product bugs. By the time a big feature reaches A/B testing, it's often already a winner, so the open question is how and where to introduce it.</li><li>A losing test is a finding, not a failure. If every experiment wins, you're not taking enough risk to learn anything new.</li><li>Guardrails and stopping criteria are what make risk-taking safe, especially when the experience is as personal as shopping.</li><li>The most valuable North Star metric is the one you can't measure yet, long-term client value, and causal-inference modeling helps predict it from short-term behavior.</li><li>Quantitative results are only half the story. Direct, qualitative client feedback inside an experiment often reshapes the rollout more than the numbers do.</li></ul><p><strong><br>Connect with the Guest</strong></p><p>LinkedIn: <a href="https://www.linkedin.com/in/nick-beyler-381864119/">https://www.linkedin.com/in/nick-beyler-381864119/</a> <br>Website: <a href="https://www.stitchfix.com">https://www.stitchfix.com</a></p><p><strong>Sponsor</strong><br>GrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide.</p><p>Go to <a href="http://growthbook.io?utm_source=edge-podcast&amp;utm_medium=podcast&amp;utm_campaign=episode-25">http://growthbook.io</a></p>]]>
      </content:encoded>
      <pubDate>Thu, 02 Jul 2026 10:08:51 -0600</pubDate>
      <author>Growthbook</author>
      <enclosure url="https://media.transistor.fm/67dbaad3/04543059.mp3" length="39765210" type="audio/mpeg"/>
      <itunes:author>Growthbook</itunes:author>
      <itunes:image href="https://img.transistorcdn.com/rY-mYsBAYbFby9635tcOatHyOvLMlr5lfKaDHGqHdQM/rs:fill:0:0:1/w:1400/h:1400/q:60/mb:500000/aHR0cHM6Ly9pbWct/dXBsb2FkLXByb2R1/Y3Rpb24udHJhbnNp/c3Rvci5mbS81NzJl/YzI3ZDIxYTQ4OWMy/ODI5Y2YxNWE4MTIx/NjA0Zi5wbmc.jpg"/>
      <itunes:duration>1242</itunes:duration>
      <itunes:summary>
        <![CDATA[<p><strong>Summary</strong></p><p>In this episode of The Experimentation Edge, GrowthBook CMO Ashley Stirrup sits down with Nick Beyler, data science manager at Stitch Fix, where he leads the decision and insights team and owns the company's internal experimentation platform. Nick shares why the metric he most wants is the one he can't measure yet, a North Star that predicts a client's long-term value from their earliest behaviors, and why the most impactful experiment learnings tend to come from adoption friction rather than product bugs. He makes the case that if you're only testing winners you're not taking enough risks, explains how guardrails make that risk safe, and looks ahead to a new in-house platform and the promise of agentic AI. It's a practical, statistician's-eye view of experimentation for product managers, data scientists, and engineers building serious testing programs.</p><p><br><strong><br>Chapters</strong></p><p>00:00 Cold open and welcome to the show</p><p>01:45 What Stitch Fix actually does</p><p>04:15 Balancing AI with the human stylist</p><p>05:15 From public policy to the A/B testing adrenaline rush</p><p>07:15 Inside the weekly experimentation review group</p><p>08:45 The AI style assistant and listening to qualitative feedback</p><p>10:45 Why adoption friction beats product bugs</p><p>13:45 Testing for losers and building guardrails</p><p>15:45 Keep rate, successful fixes, and the holy grail metric</p><p>18:15 The new platform and the promise of agentic AI</p><p><br><strong><br>Takeaways</strong></p><ul><li>The most impactful experiment learnings usually come from adoption friction, not product bugs. By the time a big feature reaches A/B testing, it's often already a winner, so the open question is how and where to introduce it.</li><li>A losing test is a finding, not a failure. If every experiment wins, you're not taking enough risk to learn anything new.</li><li>Guardrails and stopping criteria are what make risk-taking safe, especially when the experience is as personal as shopping.</li><li>The most valuable North Star metric is the one you can't measure yet, long-term client value, and causal-inference modeling helps predict it from short-term behavior.</li><li>Quantitative results are only half the story. Direct, qualitative client feedback inside an experiment often reshapes the rollout more than the numbers do.</li></ul><p><strong><br>Connect with the Guest</strong></p><p>LinkedIn: <a href="https://www.linkedin.com/in/nick-beyler-381864119/">https://www.linkedin.com/in/nick-beyler-381864119/</a> <br>Website: <a href="https://www.stitchfix.com">https://www.stitchfix.com</a></p><p><strong>Sponsor</strong><br>GrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide.</p><p>Go to <a href="http://growthbook.io?utm_source=edge-podcast&amp;utm_medium=podcast&amp;utm_campaign=episode-25">http://growthbook.io</a></p>]]>
      </itunes:summary>
      <itunes:keywords>experimentation,a/b testing,stitch fix,nick beyler,growthbook,the experimentation edge,causal inference,north star metric,long-term value,adoption friction,guardrail metrics,product experimentation,AI experimentation,data science,personalization,keep rate,sequential testing,agentic AI,feature rollout,experimentation program</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/67dbaad3/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>What the Expedia Group cannot measure, it cannot ship</title>
      <itunes:episode>23</itunes:episode>
      <podcast:episode>23</podcast:episode>
      <itunes:title>What the Expedia Group cannot measure, it cannot ship</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">f1c9aa62-371f-4bcf-8df9-2676e6a64278</guid>
      <link>https://share.transistor.fm/s/0dc99dd8</link>
      <description>
        <![CDATA[<p><strong><br>Summary</strong></p><p>Amir Moghaddam, Director of Software Engineering at Expedia Group, joins host Ashley Stirrup on The Experimentation Edge to make the case that measurement is not a reporting step but a gate: what you cannot measure, you cannot ship. Drawing on nearly four years at DoorDash and his current work leading Expedia's air booking platform, Amir explains why he refuses to label experiments winners or losers, how a "failed" pricing test pushed his team toward full personalization, and why a three sided marketplace forces hard trade-offs between competing metrics. The conversation closes on how the same experimentation discipline now applies to shipping and measuring AI. Built for product managers, engineers, data scientists, and growth leaders who care about rigor over opinion.</p><p><strong><br>Chapters</strong></p><p>00:00 Cold open<br>00:50 Meet Amir and the air booking platform at Expedia<br>03:10 DoorDash, growth, and a 70 experiment year<br>04:20 Three kinds of experimentation at Expedia<br>06:30 AI velocity and the new frontier model pace<br>08:30 What you cannot measure, you cannot ship<br>10:45 The DoorDash carousel and the price experiment<br>12:45 The three sided marketplace and competing metrics<br>16:55 There are no losing experiments<br>20:45 Predictability, LLMs, and Expedia's road ahead</p><p><strong><br>Takeaways</strong></p><ul><li>"What you cannot measure, you cannot ship" — if you can't measure an outcome, you can't decide whether it's better, so you're just debating opinions.</li><li>Measurement spans three live dimensions: spend (more with less), speed (sprints instead of quarters), and quality, with guardrail "do no harm" metrics on top.</li><li>There are no losing experiments. A flat result is a signal to either refine the hypothesis or step back and look from a completely different angle.</li><li>DoorDash's price experiment proved price by itself doesn't predict orders. Different customers want different things at different times, which pushed the team toward personalization.</li><li>A three sided marketplace (buyers, merchants, Dashers) makes metrics compete. Running the test is easy; deciding what to optimize when goals conflict is the real work.</li></ul><p><br><strong><br>Connect with the Guest</strong></p><p>LinkedIn: <a href="https://www.linkedin.com/in/amirmoghaddam">https://www.linkedin.com/in/amirmoghaddam</a><br>Website: <a href="https://www.expediagroup.com">https://www.expediagroup.com</a></p><p><strong>Sponsor</strong><br>GrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide. </p><p>Go to <a href="http://growthbook.io">growthbook.io</a></p>]]>
      </description>
      <content:encoded>
        <![CDATA[<p><strong><br>Summary</strong></p><p>Amir Moghaddam, Director of Software Engineering at Expedia Group, joins host Ashley Stirrup on The Experimentation Edge to make the case that measurement is not a reporting step but a gate: what you cannot measure, you cannot ship. Drawing on nearly four years at DoorDash and his current work leading Expedia's air booking platform, Amir explains why he refuses to label experiments winners or losers, how a "failed" pricing test pushed his team toward full personalization, and why a three sided marketplace forces hard trade-offs between competing metrics. The conversation closes on how the same experimentation discipline now applies to shipping and measuring AI. Built for product managers, engineers, data scientists, and growth leaders who care about rigor over opinion.</p><p><strong><br>Chapters</strong></p><p>00:00 Cold open<br>00:50 Meet Amir and the air booking platform at Expedia<br>03:10 DoorDash, growth, and a 70 experiment year<br>04:20 Three kinds of experimentation at Expedia<br>06:30 AI velocity and the new frontier model pace<br>08:30 What you cannot measure, you cannot ship<br>10:45 The DoorDash carousel and the price experiment<br>12:45 The three sided marketplace and competing metrics<br>16:55 There are no losing experiments<br>20:45 Predictability, LLMs, and Expedia's road ahead</p><p><strong><br>Takeaways</strong></p><ul><li>"What you cannot measure, you cannot ship" — if you can't measure an outcome, you can't decide whether it's better, so you're just debating opinions.</li><li>Measurement spans three live dimensions: spend (more with less), speed (sprints instead of quarters), and quality, with guardrail "do no harm" metrics on top.</li><li>There are no losing experiments. A flat result is a signal to either refine the hypothesis or step back and look from a completely different angle.</li><li>DoorDash's price experiment proved price by itself doesn't predict orders. Different customers want different things at different times, which pushed the team toward personalization.</li><li>A three sided marketplace (buyers, merchants, Dashers) makes metrics compete. Running the test is easy; deciding what to optimize when goals conflict is the real work.</li></ul><p><br><strong><br>Connect with the Guest</strong></p><p>LinkedIn: <a href="https://www.linkedin.com/in/amirmoghaddam">https://www.linkedin.com/in/amirmoghaddam</a><br>Website: <a href="https://www.expediagroup.com">https://www.expediagroup.com</a></p><p><strong>Sponsor</strong><br>GrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide. </p><p>Go to <a href="http://growthbook.io">growthbook.io</a></p>]]>
      </content:encoded>
      <pubDate>Wed, 01 Jul 2026 09:04:41 -0600</pubDate>
      <author>Growthbook</author>
      <enclosure url="https://media.transistor.fm/0dc99dd8/5965e9dc.mp3" length="41784393" type="audio/mpeg"/>
      <itunes:author>Growthbook</itunes:author>
      <itunes:image href="https://img.transistorcdn.com/MgJCSQCEWNpn-t-JulaVs__Y_Rr53fbCjs-XYeOkq3E/rs:fill:0:0:1/w:1400/h:1400/q:60/mb:500000/aHR0cHM6Ly9pbWct/dXBsb2FkLXByb2R1/Y3Rpb24udHJhbnNp/c3Rvci5mbS9kN2Rk/ZWViZTQyYjlkMDhk/MzMxYjVlODA2ZDgw/NDU4ZC5wbmc.jpg"/>
      <itunes:duration>1739</itunes:duration>
      <itunes:summary>
        <![CDATA[<p><strong><br>Summary</strong></p><p>Amir Moghaddam, Director of Software Engineering at Expedia Group, joins host Ashley Stirrup on The Experimentation Edge to make the case that measurement is not a reporting step but a gate: what you cannot measure, you cannot ship. Drawing on nearly four years at DoorDash and his current work leading Expedia's air booking platform, Amir explains why he refuses to label experiments winners or losers, how a "failed" pricing test pushed his team toward full personalization, and why a three sided marketplace forces hard trade-offs between competing metrics. The conversation closes on how the same experimentation discipline now applies to shipping and measuring AI. Built for product managers, engineers, data scientists, and growth leaders who care about rigor over opinion.</p><p><strong><br>Chapters</strong></p><p>00:00 Cold open<br>00:50 Meet Amir and the air booking platform at Expedia<br>03:10 DoorDash, growth, and a 70 experiment year<br>04:20 Three kinds of experimentation at Expedia<br>06:30 AI velocity and the new frontier model pace<br>08:30 What you cannot measure, you cannot ship<br>10:45 The DoorDash carousel and the price experiment<br>12:45 The three sided marketplace and competing metrics<br>16:55 There are no losing experiments<br>20:45 Predictability, LLMs, and Expedia's road ahead</p><p><strong><br>Takeaways</strong></p><ul><li>"What you cannot measure, you cannot ship" — if you can't measure an outcome, you can't decide whether it's better, so you're just debating opinions.</li><li>Measurement spans three live dimensions: spend (more with less), speed (sprints instead of quarters), and quality, with guardrail "do no harm" metrics on top.</li><li>There are no losing experiments. A flat result is a signal to either refine the hypothesis or step back and look from a completely different angle.</li><li>DoorDash's price experiment proved price by itself doesn't predict orders. Different customers want different things at different times, which pushed the team toward personalization.</li><li>A three sided marketplace (buyers, merchants, Dashers) makes metrics compete. Running the test is easy; deciding what to optimize when goals conflict is the real work.</li></ul><p><br><strong><br>Connect with the Guest</strong></p><p>LinkedIn: <a href="https://www.linkedin.com/in/amirmoghaddam">https://www.linkedin.com/in/amirmoghaddam</a><br>Website: <a href="https://www.expediagroup.com">https://www.expediagroup.com</a></p><p><strong>Sponsor</strong><br>GrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide. </p><p>Go to <a href="http://growthbook.io">growthbook.io</a></p>]]>
      </itunes:summary>
      <itunes:keywords>experimentation,a/b testing,feature flags,hypothesis testing,product experimentation,measurement,guardrail metrics,doordash,expedia,personalization,recommendation systems,multi-armed bandit,three sided marketplace,competing metrics,AI experimentation,llm non-determinism,cross-model validation,testing velocity,product growth,the experimentation edge</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/0dc99dd8/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>How Fin went from weeks to hours of analysis using AI</title>
      <itunes:episode>22</itunes:episode>
      <podcast:episode>22</podcast:episode>
      <itunes:title>How Fin went from weeks to hours of analysis using AI</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">4087a783-bd08-48ea-b60d-c8eaabcdbec4</guid>
      <link>https://share.transistor.fm/s/d47fee72</link>
      <description>
        <![CDATA[<p><strong>Summary</strong></p><p>In this episode of The Experimentation Edge, host Ashley Stirrup sits down with Raunak Kumar, senior manager of GTM analytics at Fin (formerly Intercom), to unpack how experimentation actually works when the data is messy and the traffic is thin. Drawing on nearly 12 years in marketing analytics across Atlassian, Stripe, and Fin, Raunak explains how AI tools like Claude Code have collapsed analysis from weeks to hours and freed his team to clear its experiment backlog, why declining organic search traffic and a 5x jump in untagged ChatGPT referrals are forcing teams to rethink attribution, and how the most valuable experiments are often the ones that "lose." From a Jira Service Desk bundling test that won on trials but had to be rolled back, to a Stripe contact form that was quietly blocking real buyers, this conversation is a practical guide for product managers, engineers, data scientists, and growth marketers who want to learn more from every test they run.</p><p><br><strong><br>Chapters</strong></p><p>0:45 Welcome and what the show is about<br>1:45 Raunak's role and 12 years in marketing analytics<br>2:45 How AI and Claude Code changed the analyst's day<br>4:15 LLMs, declining organic traffic, and the 5x ChatGPT jump<br>5:15 Two kinds of experiments at Fin: on page and off page<br>7:15 The Jira Service Desk bundling experiment<br>10:45 Why the trial winner became a rollback<br>11:45 Contextual onboarding turns the loser into a winner<br>14:45 Reading an experiment that loses<br>18:45 What's next: incrementality, connected TV, and testing creative</p><p><br><strong><br>Takeaways</strong></p><ul><li>AI has collapsed marketing analysis from weeks to hours, and the real payoff is a cleared experiment backlog plus analysts who compete on the questions they ask, not the speed they query.</li><li>Organic search traffic is declining as ChatGPT, Gemini's AI mode, and Claude answer buyers in place; Fin saw a 5x rise in ChatGPT referrals, but LLMs don't tag that traffic, so attribution has to be proven through experiments.</li><li>A guardrail metric saved Atlassian from a costly mistake: bundling Jira Service Desk lifted trials more than 50 percent but tanked activation and paid conversion, forcing a rollback.</li><li>A failed test can hold the real winner; contextual onboarding matched to user intent roughly doubled activation and became the default variant after the bundling experiment was rolled back.</li><li>In low-volume B2B, read losing experiments for sub-segment signal; a "failed" Stripe form simplification revealed the form was blocking legitimate small-business buyers using Gmail.</li></ul><p><br><strong><br>Connect with the Guest</strong></p><p>LinkedIn: <a href="http://linkedin.com/in/raunakkumar1991">http://linkedin.com/in/raunakkumar1991</a><br>Website: <a href="https://fin.ai">https://fin.ai</a></p><p><strong>Sponsor</strong><br>Growthbook helps you ship features with confidence by bringing experimentation and feature flagging into one open-source platform. No more guessing whether that new checkout flow actually moved the needle, waiting weeks for data team bandwidth, or flying blind on rollouts.</p><p>Growthbook gives you a single place to run A/B tests, manage feature flags, and analyze results against your existing data warehouse.</p><p>With powerful stats built in, it takes the complexity out of experimentation, helps you catch regressions before they hit every user, and makes it easy to test ideas that keep your product improving and your metrics moving in the right direction.</p><p>See a demo at <a href="https://www.growthbook.io/">https://www.growthbook.io/</a></p>]]>
      </description>
      <content:encoded>
        <![CDATA[<p><strong>Summary</strong></p><p>In this episode of The Experimentation Edge, host Ashley Stirrup sits down with Raunak Kumar, senior manager of GTM analytics at Fin (formerly Intercom), to unpack how experimentation actually works when the data is messy and the traffic is thin. Drawing on nearly 12 years in marketing analytics across Atlassian, Stripe, and Fin, Raunak explains how AI tools like Claude Code have collapsed analysis from weeks to hours and freed his team to clear its experiment backlog, why declining organic search traffic and a 5x jump in untagged ChatGPT referrals are forcing teams to rethink attribution, and how the most valuable experiments are often the ones that "lose." From a Jira Service Desk bundling test that won on trials but had to be rolled back, to a Stripe contact form that was quietly blocking real buyers, this conversation is a practical guide for product managers, engineers, data scientists, and growth marketers who want to learn more from every test they run.</p><p><br><strong><br>Chapters</strong></p><p>0:45 Welcome and what the show is about<br>1:45 Raunak's role and 12 years in marketing analytics<br>2:45 How AI and Claude Code changed the analyst's day<br>4:15 LLMs, declining organic traffic, and the 5x ChatGPT jump<br>5:15 Two kinds of experiments at Fin: on page and off page<br>7:15 The Jira Service Desk bundling experiment<br>10:45 Why the trial winner became a rollback<br>11:45 Contextual onboarding turns the loser into a winner<br>14:45 Reading an experiment that loses<br>18:45 What's next: incrementality, connected TV, and testing creative</p><p><br><strong><br>Takeaways</strong></p><ul><li>AI has collapsed marketing analysis from weeks to hours, and the real payoff is a cleared experiment backlog plus analysts who compete on the questions they ask, not the speed they query.</li><li>Organic search traffic is declining as ChatGPT, Gemini's AI mode, and Claude answer buyers in place; Fin saw a 5x rise in ChatGPT referrals, but LLMs don't tag that traffic, so attribution has to be proven through experiments.</li><li>A guardrail metric saved Atlassian from a costly mistake: bundling Jira Service Desk lifted trials more than 50 percent but tanked activation and paid conversion, forcing a rollback.</li><li>A failed test can hold the real winner; contextual onboarding matched to user intent roughly doubled activation and became the default variant after the bundling experiment was rolled back.</li><li>In low-volume B2B, read losing experiments for sub-segment signal; a "failed" Stripe form simplification revealed the form was blocking legitimate small-business buyers using Gmail.</li></ul><p><br><strong><br>Connect with the Guest</strong></p><p>LinkedIn: <a href="http://linkedin.com/in/raunakkumar1991">http://linkedin.com/in/raunakkumar1991</a><br>Website: <a href="https://fin.ai">https://fin.ai</a></p><p><strong>Sponsor</strong><br>Growthbook helps you ship features with confidence by bringing experimentation and feature flagging into one open-source platform. No more guessing whether that new checkout flow actually moved the needle, waiting weeks for data team bandwidth, or flying blind on rollouts.</p><p>Growthbook gives you a single place to run A/B tests, manage feature flags, and analyze results against your existing data warehouse.</p><p>With powerful stats built in, it takes the complexity out of experimentation, helps you catch regressions before they hit every user, and makes it easy to test ideas that keep your product improving and your metrics moving in the right direction.</p><p>See a demo at <a href="https://www.growthbook.io/">https://www.growthbook.io/</a></p>]]>
      </content:encoded>
      <pubDate>Tue, 30 Jun 2026 10:17:53 -0600</pubDate>
      <author>Growthbook</author>
      <enclosure url="https://media.transistor.fm/d47fee72/fa28664e.mp3" length="33327568" type="audio/mpeg"/>
      <itunes:author>Growthbook</itunes:author>
      <itunes:image href="https://img.transistorcdn.com/aCKWWi6t15JRVVj8HN5unsTLf263DIbVt8keb7_nCDA/rs:fill:0:0:1/w:1400/h:1400/q:60/mb:500000/aHR0cHM6Ly9pbWct/dXBsb2FkLXByb2R1/Y3Rpb24udHJhbnNp/c3Rvci5mbS8xY2Q3/M2IyNWU0ZjQ5OGNh/YjA1ZmIzNjcyODZi/YzQ2OS5wbmc.jpg"/>
      <itunes:duration>1387</itunes:duration>
      <itunes:summary>
        <![CDATA[<p><strong>Summary</strong></p><p>In this episode of The Experimentation Edge, host Ashley Stirrup sits down with Raunak Kumar, senior manager of GTM analytics at Fin (formerly Intercom), to unpack how experimentation actually works when the data is messy and the traffic is thin. Drawing on nearly 12 years in marketing analytics across Atlassian, Stripe, and Fin, Raunak explains how AI tools like Claude Code have collapsed analysis from weeks to hours and freed his team to clear its experiment backlog, why declining organic search traffic and a 5x jump in untagged ChatGPT referrals are forcing teams to rethink attribution, and how the most valuable experiments are often the ones that "lose." From a Jira Service Desk bundling test that won on trials but had to be rolled back, to a Stripe contact form that was quietly blocking real buyers, this conversation is a practical guide for product managers, engineers, data scientists, and growth marketers who want to learn more from every test they run.</p><p><br><strong><br>Chapters</strong></p><p>0:45 Welcome and what the show is about<br>1:45 Raunak's role and 12 years in marketing analytics<br>2:45 How AI and Claude Code changed the analyst's day<br>4:15 LLMs, declining organic traffic, and the 5x ChatGPT jump<br>5:15 Two kinds of experiments at Fin: on page and off page<br>7:15 The Jira Service Desk bundling experiment<br>10:45 Why the trial winner became a rollback<br>11:45 Contextual onboarding turns the loser into a winner<br>14:45 Reading an experiment that loses<br>18:45 What's next: incrementality, connected TV, and testing creative</p><p><br><strong><br>Takeaways</strong></p><ul><li>AI has collapsed marketing analysis from weeks to hours, and the real payoff is a cleared experiment backlog plus analysts who compete on the questions they ask, not the speed they query.</li><li>Organic search traffic is declining as ChatGPT, Gemini's AI mode, and Claude answer buyers in place; Fin saw a 5x rise in ChatGPT referrals, but LLMs don't tag that traffic, so attribution has to be proven through experiments.</li><li>A guardrail metric saved Atlassian from a costly mistake: bundling Jira Service Desk lifted trials more than 50 percent but tanked activation and paid conversion, forcing a rollback.</li><li>A failed test can hold the real winner; contextual onboarding matched to user intent roughly doubled activation and became the default variant after the bundling experiment was rolled back.</li><li>In low-volume B2B, read losing experiments for sub-segment signal; a "failed" Stripe form simplification revealed the form was blocking legitimate small-business buyers using Gmail.</li></ul><p><br><strong><br>Connect with the Guest</strong></p><p>LinkedIn: <a href="http://linkedin.com/in/raunakkumar1991">http://linkedin.com/in/raunakkumar1991</a><br>Website: <a href="https://fin.ai">https://fin.ai</a></p><p><strong>Sponsor</strong><br>Growthbook helps you ship features with confidence by bringing experimentation and feature flagging into one open-source platform. No more guessing whether that new checkout flow actually moved the needle, waiting weeks for data team bandwidth, or flying blind on rollouts.</p><p>Growthbook gives you a single place to run A/B tests, manage feature flags, and analyze results against your existing data warehouse.</p><p>With powerful stats built in, it takes the complexity out of experimentation, helps you catch regressions before they hit every user, and makes it easy to test ideas that keep your product improving and your metrics moving in the right direction.</p><p>See a demo at <a href="https://www.growthbook.io/">https://www.growthbook.io/</a></p>]]>
      </itunes:summary>
      <itunes:keywords>a/b testing,experimentation,growth marketing,gtm analytics,marketing attribution,b2b marketing analytics,incrementality testing,geo lift study,contextual onboarding,activation rate,counter-metrics,product bundling,connected tv advertising,landing page optimization,conversion rate optimization,llm referral traffic,organic traffic decline,ai in marketing,claude code,customer acquisition</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/d47fee72/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Inside The Home Depot's experimentation at a $25B scale</title>
      <itunes:episode>21</itunes:episode>
      <podcast:episode>21</podcast:episode>
      <itunes:title>Inside The Home Depot's experimentation at a $25B scale</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">fe2cadd0-39d5-413a-b3a8-182232d39537</guid>
      <link>https://share.transistor.fm/s/810b4c1d</link>
      <description>
        <![CDATA[<p><strong>Summary</strong><br>What does experimentation look like inside a $150 billion retailer? In this episode of The Experimentation Edge, host Ashley Stirrup talks with Kim Ting Li, Senior Manager of Experimentation at The Home Depot, where one centralized team tests every major change to a $25 billion online business. Kim explains how 40 people serve 40–50 business teams, why executives join test readouts and ping analysts directly, how every result since 2020 lives in a searchable library, and why scaling beyond hundreds of experiments per year depends on server-side testing capabilities more than AI. For product, data, and engineering leaders building or scaling experimentation programs.<br></p><p><strong>Chapters</strong><br> 00:00 Intro<br> 00:45 From neuroscience research to Home Depot<br> 01:45 A $150B enterprise, a $25B online business<br> 02:45 The centralized experimentation model<br> 03:45 Inside the 40-person team<br> 04:30 Readouts, blast emails, and the experiment library<br> 05:40 Executive visibility and the golden rule<br> 06:15 "If you won't act on a bad result, don't run the test"<br> 11:15 Learning from losing tests<br> 12:30 Scaling up: AI, server-side testing, and what's next<br></p><p><strong>Takeaways</strong></p><ul><li>One centralized team of about 40 people tests every major change to Home Depot's $25B online business, serving 40–50 business teams with consistent hypothesis and analysis standards.</li><li>Executive engagement is real at Home Depot: leaders join 30-minute readouts, search the experiment library, and ping analysts directly because they treat A/B testing as the golden rule for measuring incrementality.</li><li>Institutional memory is infrastructure — every test result since 2020 lives in a centralized, searchable archive so no one re-runs a question the company already answered.</li><li>Kim's stakeholder filter: if you wouldn't do anything differently after a bad result, don't run the test.</li><li>Scaling past low hundreds of experiments per year is a capabilities problem before it's an AI problem — Home Depot is moving from client-side to server-side testing so winners release quickly, end to end.<p><br></p></li></ul><p><strong>Connect with the Guest</strong><br> LinkedIn: <a href="https://www.linkedin.com/in/kimtingli">https://www.linkedin.com/in/kimtingli</a><br> Website: <a href="https://www.homedepot.com">https://www.homedepot.com</a></p><p><strong>Sponsor</strong><br>Growthbook helps you ship features with confidence by bringing experimentation and feature flagging into one open-source platform. No more guessing whether that new checkout flow actually moved the needle, waiting weeks for data team bandwidth, or flying blind on rollouts.</p><p>Growthbook gives you a single place to run A/B tests, manage feature flags, and analyze results against your existing data warehouse.</p><p>With powerful stats built in, it takes the complexity out of experimentation, helps you catch regressions before they hit every user, and makes it easy to test ideas that keep your product improving and your metrics moving in the right direction.</p><p>See a demo at <a href="https://www.growthbook.io/">https://www.growthbook.io/</a></p>]]>
      </description>
      <content:encoded>
        <![CDATA[<p><strong>Summary</strong><br>What does experimentation look like inside a $150 billion retailer? In this episode of The Experimentation Edge, host Ashley Stirrup talks with Kim Ting Li, Senior Manager of Experimentation at The Home Depot, where one centralized team tests every major change to a $25 billion online business. Kim explains how 40 people serve 40–50 business teams, why executives join test readouts and ping analysts directly, how every result since 2020 lives in a searchable library, and why scaling beyond hundreds of experiments per year depends on server-side testing capabilities more than AI. For product, data, and engineering leaders building or scaling experimentation programs.<br></p><p><strong>Chapters</strong><br> 00:00 Intro<br> 00:45 From neuroscience research to Home Depot<br> 01:45 A $150B enterprise, a $25B online business<br> 02:45 The centralized experimentation model<br> 03:45 Inside the 40-person team<br> 04:30 Readouts, blast emails, and the experiment library<br> 05:40 Executive visibility and the golden rule<br> 06:15 "If you won't act on a bad result, don't run the test"<br> 11:15 Learning from losing tests<br> 12:30 Scaling up: AI, server-side testing, and what's next<br></p><p><strong>Takeaways</strong></p><ul><li>One centralized team of about 40 people tests every major change to Home Depot's $25B online business, serving 40–50 business teams with consistent hypothesis and analysis standards.</li><li>Executive engagement is real at Home Depot: leaders join 30-minute readouts, search the experiment library, and ping analysts directly because they treat A/B testing as the golden rule for measuring incrementality.</li><li>Institutional memory is infrastructure — every test result since 2020 lives in a centralized, searchable archive so no one re-runs a question the company already answered.</li><li>Kim's stakeholder filter: if you wouldn't do anything differently after a bad result, don't run the test.</li><li>Scaling past low hundreds of experiments per year is a capabilities problem before it's an AI problem — Home Depot is moving from client-side to server-side testing so winners release quickly, end to end.<p><br></p></li></ul><p><strong>Connect with the Guest</strong><br> LinkedIn: <a href="https://www.linkedin.com/in/kimtingli">https://www.linkedin.com/in/kimtingli</a><br> Website: <a href="https://www.homedepot.com">https://www.homedepot.com</a></p><p><strong>Sponsor</strong><br>Growthbook helps you ship features with confidence by bringing experimentation and feature flagging into one open-source platform. No more guessing whether that new checkout flow actually moved the needle, waiting weeks for data team bandwidth, or flying blind on rollouts.</p><p>Growthbook gives you a single place to run A/B tests, manage feature flags, and analyze results against your existing data warehouse.</p><p>With powerful stats built in, it takes the complexity out of experimentation, helps you catch regressions before they hit every user, and makes it easy to test ideas that keep your product improving and your metrics moving in the right direction.</p><p>See a demo at <a href="https://www.growthbook.io/">https://www.growthbook.io/</a></p>]]>
      </content:encoded>
      <pubDate>Mon, 29 Jun 2026 10:19:04 -0600</pubDate>
      <author>Growthbook</author>
      <enclosure url="https://media.transistor.fm/810b4c1d/f2110166.mp3" length="11291631" type="audio/mpeg"/>
      <itunes:author>Growthbook</itunes:author>
      <itunes:image href="https://img.transistorcdn.com/3R5I0-tOjNriZ5xWk4IBgcJTdRX6CWBzkIDzQvzvH90/rs:fill:0:0:1/w:1400/h:1400/q:60/mb:500000/aHR0cHM6Ly9pbWct/dXBsb2FkLXByb2R1/Y3Rpb24udHJhbnNp/c3Rvci5mbS8wZTY3/ZjhmNmNiNzhmOTFj/NjBlN2RjZGY2MzUy/YjJjZC5wbmc.jpg"/>
      <itunes:duration>703</itunes:duration>
      <itunes:summary>
        <![CDATA[<p><strong>Summary</strong><br>What does experimentation look like inside a $150 billion retailer? In this episode of The Experimentation Edge, host Ashley Stirrup talks with Kim Ting Li, Senior Manager of Experimentation at The Home Depot, where one centralized team tests every major change to a $25 billion online business. Kim explains how 40 people serve 40–50 business teams, why executives join test readouts and ping analysts directly, how every result since 2020 lives in a searchable library, and why scaling beyond hundreds of experiments per year depends on server-side testing capabilities more than AI. For product, data, and engineering leaders building or scaling experimentation programs.<br></p><p><strong>Chapters</strong><br> 00:00 Intro<br> 00:45 From neuroscience research to Home Depot<br> 01:45 A $150B enterprise, a $25B online business<br> 02:45 The centralized experimentation model<br> 03:45 Inside the 40-person team<br> 04:30 Readouts, blast emails, and the experiment library<br> 05:40 Executive visibility and the golden rule<br> 06:15 "If you won't act on a bad result, don't run the test"<br> 11:15 Learning from losing tests<br> 12:30 Scaling up: AI, server-side testing, and what's next<br></p><p><strong>Takeaways</strong></p><ul><li>One centralized team of about 40 people tests every major change to Home Depot's $25B online business, serving 40–50 business teams with consistent hypothesis and analysis standards.</li><li>Executive engagement is real at Home Depot: leaders join 30-minute readouts, search the experiment library, and ping analysts directly because they treat A/B testing as the golden rule for measuring incrementality.</li><li>Institutional memory is infrastructure — every test result since 2020 lives in a centralized, searchable archive so no one re-runs a question the company already answered.</li><li>Kim's stakeholder filter: if you wouldn't do anything differently after a bad result, don't run the test.</li><li>Scaling past low hundreds of experiments per year is a capabilities problem before it's an AI problem — Home Depot is moving from client-side to server-side testing so winners release quickly, end to end.<p><br></p></li></ul><p><strong>Connect with the Guest</strong><br> LinkedIn: <a href="https://www.linkedin.com/in/kimtingli">https://www.linkedin.com/in/kimtingli</a><br> Website: <a href="https://www.homedepot.com">https://www.homedepot.com</a></p><p><strong>Sponsor</strong><br>Growthbook helps you ship features with confidence by bringing experimentation and feature flagging into one open-source platform. No more guessing whether that new checkout flow actually moved the needle, waiting weeks for data team bandwidth, or flying blind on rollouts.</p><p>Growthbook gives you a single place to run A/B tests, manage feature flags, and analyze results against your existing data warehouse.</p><p>With powerful stats built in, it takes the complexity out of experimentation, helps you catch regressions before they hit every user, and makes it easy to test ideas that keep your product improving and your metrics moving in the right direction.</p><p>See a demo at <a href="https://www.growthbook.io/">https://www.growthbook.io/</a></p>]]>
      </itunes:summary>
      <itunes:keywords>ai in hr,hr automation,hr operations,ai adoption in hr,hr as a product,employee experience,people operations,hr ai strategy,hr process improvement,hr tech stack,ai service desk for hr,hr service delivery,build vs buy hr ai,hcm ai agents,employee self-service,hr digital transformation,hr workflows,ai agents for hr,people analytics,hr automation strategy</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/810b4c1d/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>How Disney picks which experiments to run</title>
      <itunes:episode>20</itunes:episode>
      <podcast:episode>20</podcast:episode>
      <itunes:title>How Disney picks which experiments to run</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">92f77812-c8a6-46e0-bf3b-1acde637806c</guid>
      <link>https://share.transistor.fm/s/ab473451</link>
      <description>
        <![CDATA[<p><strong>Summary</strong><br> What does it look like to kill a multimillion dollar feature before anyone builds it? In this episode of The Experimentation Edge, host Ashley Stirrup talks with Crystal Ammari, a digital product optimization and experimentation strategy leader whose career spans Nike and The Walt Disney Company. Crystal shares the "dry test" that used a single fake button to measure demand for video chat (4 million users, 106 clicks), why she reframes experimentation as savings and gains rather than wins and losses, how a misconfigured tool, not bad methodology, made tests take six months, and how a stuck Disney team went from "we don't know where to start" to 110 scored and prioritized test ideas. For product, data, and engineering leaders building or scaling experimentation programs.<br></p><p><strong>Chapters</strong><br> 00:00 Intro<br> 00:45 The mindset shift from shipping to results<br> 02:00 Why testing took six months, a tooling problem<br> 03:15 The dev team that laughed, and the vendor who agreed<br> 04:50 An executive demand for video chat<br> 05:35 Dry testing with a fake button<br> 06:30 106 clicks and a multimillion dollar save<br> 07:30 Savings and gains, not wins and losses<br> 08:45 The Disney team that didn't know where to start<br> 10:30 From low engagement to 110 prioritized ideas<br> 12:45 Just get something live, and where AI fits next<br></p><p><strong>Takeaways</strong></p><ul><li>A "dry test", a fake "Click here to video chat" button that grayed out on click — measured real demand without building the feature. Of roughly 4 million users, only 106 clicked, killing a multimillion dollar build.</li><li>Reframe experiment outcomes as savings and gains, not wins and losses. A "losing" test saves you from a costly mistake, which keeps teams focused on learning instead of fearing failure.</li><li>Slow experimentation is often a tooling problem, not a methodology problem. One program's six month test cycle came from rebuilding every page instead of overlaying changes the way the tool intended.</li><li>Getting a stuck team unstuck starts with data and a workshop. A Disney team went from "we don't know where to start" to 110 scored, prioritized test ideas, using Contentsquare heatmaps to diagnose low engagement first.</li><li>The biggest thing that gets a team testing is to just do it. Stop designing the perfect experiment and get something simple live to take away the mystery.<p><br></p></li></ul><p><strong>Connect with the Guest</strong><br> LinkedIn: <a href="https://www.linkedin.com/in/crystal-ammari/">https://www.linkedin.com/in/crystal-ammari/</a><br> Website: <a href="https://thewaltdisneycompany.com">https://thewaltdisneycompany.com</a></p><p><strong>Sponsor</strong><br>Growthbook helps you ship features with confidence by bringing experimentation and feature flagging into one open-source platform. No more guessing whether that new checkout flow actually moved the needle, waiting weeks for data team bandwidth, or flying blind on rollouts.</p><p>Growthbook gives you a single place to run A/B tests, manage feature flags, and analyze results against your existing data warehouse.</p><p>With powerful stats built in, it takes the complexity out of experimentation, helps you catch regressions before they hit every user, and makes it easy to test ideas that keep your product improving and your metrics moving in the right direction.</p><p>See a demo at <a href="https://www.growthbook.io/">https://www.growthbook.io/</a></p>]]>
      </description>
      <content:encoded>
        <![CDATA[<p><strong>Summary</strong><br> What does it look like to kill a multimillion dollar feature before anyone builds it? In this episode of The Experimentation Edge, host Ashley Stirrup talks with Crystal Ammari, a digital product optimization and experimentation strategy leader whose career spans Nike and The Walt Disney Company. Crystal shares the "dry test" that used a single fake button to measure demand for video chat (4 million users, 106 clicks), why she reframes experimentation as savings and gains rather than wins and losses, how a misconfigured tool, not bad methodology, made tests take six months, and how a stuck Disney team went from "we don't know where to start" to 110 scored and prioritized test ideas. For product, data, and engineering leaders building or scaling experimentation programs.<br></p><p><strong>Chapters</strong><br> 00:00 Intro<br> 00:45 The mindset shift from shipping to results<br> 02:00 Why testing took six months, a tooling problem<br> 03:15 The dev team that laughed, and the vendor who agreed<br> 04:50 An executive demand for video chat<br> 05:35 Dry testing with a fake button<br> 06:30 106 clicks and a multimillion dollar save<br> 07:30 Savings and gains, not wins and losses<br> 08:45 The Disney team that didn't know where to start<br> 10:30 From low engagement to 110 prioritized ideas<br> 12:45 Just get something live, and where AI fits next<br></p><p><strong>Takeaways</strong></p><ul><li>A "dry test", a fake "Click here to video chat" button that grayed out on click — measured real demand without building the feature. Of roughly 4 million users, only 106 clicked, killing a multimillion dollar build.</li><li>Reframe experiment outcomes as savings and gains, not wins and losses. A "losing" test saves you from a costly mistake, which keeps teams focused on learning instead of fearing failure.</li><li>Slow experimentation is often a tooling problem, not a methodology problem. One program's six month test cycle came from rebuilding every page instead of overlaying changes the way the tool intended.</li><li>Getting a stuck team unstuck starts with data and a workshop. A Disney team went from "we don't know where to start" to 110 scored, prioritized test ideas, using Contentsquare heatmaps to diagnose low engagement first.</li><li>The biggest thing that gets a team testing is to just do it. Stop designing the perfect experiment and get something simple live to take away the mystery.<p><br></p></li></ul><p><strong>Connect with the Guest</strong><br> LinkedIn: <a href="https://www.linkedin.com/in/crystal-ammari/">https://www.linkedin.com/in/crystal-ammari/</a><br> Website: <a href="https://thewaltdisneycompany.com">https://thewaltdisneycompany.com</a></p><p><strong>Sponsor</strong><br>Growthbook helps you ship features with confidence by bringing experimentation and feature flagging into one open-source platform. No more guessing whether that new checkout flow actually moved the needle, waiting weeks for data team bandwidth, or flying blind on rollouts.</p><p>Growthbook gives you a single place to run A/B tests, manage feature flags, and analyze results against your existing data warehouse.</p><p>With powerful stats built in, it takes the complexity out of experimentation, helps you catch regressions before they hit every user, and makes it easy to test ideas that keep your product improving and your metrics moving in the right direction.</p><p>See a demo at <a href="https://www.growthbook.io/">https://www.growthbook.io/</a></p>]]>
      </content:encoded>
      <pubDate>Mon, 29 Jun 2026 08:44:37 -0600</pubDate>
      <author>Growthbook</author>
      <enclosure url="https://media.transistor.fm/ab473451/d81a6bab.mp3" length="31296453" type="audio/mpeg"/>
      <itunes:author>Growthbook</itunes:author>
      <itunes:image href="https://img.transistorcdn.com/QQ-D5XjTH5xRAjWUtMRKQ6N1fBk4ptNqueFORErwlfA/rs:fill:0:0:1/w:1400/h:1400/q:60/mb:500000/aHR0cHM6Ly9pbWct/dXBsb2FkLXByb2R1/Y3Rpb24udHJhbnNp/c3Rvci5mbS84OTI1/YzhkOGQ3ZGYzNzhm/MWU1MjhhZTVjYzk3/MWNhNi5wbmc.jpg"/>
      <itunes:duration>1953</itunes:duration>
      <itunes:summary>
        <![CDATA[<p><strong>Summary</strong><br> What does it look like to kill a multimillion dollar feature before anyone builds it? In this episode of The Experimentation Edge, host Ashley Stirrup talks with Crystal Ammari, a digital product optimization and experimentation strategy leader whose career spans Nike and The Walt Disney Company. Crystal shares the "dry test" that used a single fake button to measure demand for video chat (4 million users, 106 clicks), why she reframes experimentation as savings and gains rather than wins and losses, how a misconfigured tool, not bad methodology, made tests take six months, and how a stuck Disney team went from "we don't know where to start" to 110 scored and prioritized test ideas. For product, data, and engineering leaders building or scaling experimentation programs.<br></p><p><strong>Chapters</strong><br> 00:00 Intro<br> 00:45 The mindset shift from shipping to results<br> 02:00 Why testing took six months, a tooling problem<br> 03:15 The dev team that laughed, and the vendor who agreed<br> 04:50 An executive demand for video chat<br> 05:35 Dry testing with a fake button<br> 06:30 106 clicks and a multimillion dollar save<br> 07:30 Savings and gains, not wins and losses<br> 08:45 The Disney team that didn't know where to start<br> 10:30 From low engagement to 110 prioritized ideas<br> 12:45 Just get something live, and where AI fits next<br></p><p><strong>Takeaways</strong></p><ul><li>A "dry test", a fake "Click here to video chat" button that grayed out on click — measured real demand without building the feature. Of roughly 4 million users, only 106 clicked, killing a multimillion dollar build.</li><li>Reframe experiment outcomes as savings and gains, not wins and losses. A "losing" test saves you from a costly mistake, which keeps teams focused on learning instead of fearing failure.</li><li>Slow experimentation is often a tooling problem, not a methodology problem. One program's six month test cycle came from rebuilding every page instead of overlaying changes the way the tool intended.</li><li>Getting a stuck team unstuck starts with data and a workshop. A Disney team went from "we don't know where to start" to 110 scored, prioritized test ideas, using Contentsquare heatmaps to diagnose low engagement first.</li><li>The biggest thing that gets a team testing is to just do it. Stop designing the perfect experiment and get something simple live to take away the mystery.<p><br></p></li></ul><p><strong>Connect with the Guest</strong><br> LinkedIn: <a href="https://www.linkedin.com/in/crystal-ammari/">https://www.linkedin.com/in/crystal-ammari/</a><br> Website: <a href="https://thewaltdisneycompany.com">https://thewaltdisneycompany.com</a></p><p><strong>Sponsor</strong><br>Growthbook helps you ship features with confidence by bringing experimentation and feature flagging into one open-source platform. No more guessing whether that new checkout flow actually moved the needle, waiting weeks for data team bandwidth, or flying blind on rollouts.</p><p>Growthbook gives you a single place to run A/B tests, manage feature flags, and analyze results against your existing data warehouse.</p><p>With powerful stats built in, it takes the complexity out of experimentation, helps you catch regressions before they hit every user, and makes it easy to test ideas that keep your product improving and your metrics moving in the right direction.</p><p>See a demo at <a href="https://www.growthbook.io/">https://www.growthbook.io/</a></p>]]>
      </itunes:summary>
      <itunes:keywords>dry testing,experimentation,ab testing,a/b testing,fake door testing,test before you build,experimentation strategy,digital product optimization,losses avoided,testing velocity,feature flags,experimentation culture,conversion optimization,contentsquare,heatmaps,test prioritization,experiment workshop,product experimentation,disney experimentation,growthbook</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/ab473451/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Ship faster, measure better: experimentation tips from JPMorgan Chase</title>
      <itunes:episode>19</itunes:episode>
      <podcast:episode>19</podcast:episode>
      <itunes:title>Ship faster, measure better: experimentation tips from JPMorgan Chase</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">d919cb8c-af1e-483d-8278-7f179aa32f44</guid>
      <link>https://share.transistor.fm/s/6cf031b3</link>
      <description>
        <![CDATA[<p><strong>Summary</strong></p><p>How do you know if the thing you just shipped actually worked? On this episode of The Experimentation Edge, host Ashley Stirrup, CMO of GrowthBook, sits down with Kevin Yang, Executive Director and Head of Experimentation at JPMorgan Chase, who has spent six years building experimentation across Chase's digital platforms. Kevin shares how his team turned experimentation into more than a billion dollars of estimated value, why the losing experiments matter more than the winners, and the simple chart exercise he uses to prove that a million-dollar change is invisible without a control group. He and Ashley also dig into measuring engagement without chasing vanity metrics, planning for failure to defeat confirmation bias, and why AI is pushing experimentation into a golden era. It's a practical look for product managers, data scientists, and engineers at how a bank operating at massive scale makes better decisions.</p><p><strong>Chapters</strong></p><p>00:00 Welcome to the experimentation edge</p><p>01:45 Kevin's role leading experimentation at chase</p><p>04:15 Why chase invested in experimentation</p><p>06:45 A billion dollars and the value of losers</p><p>12:45 Plan for failure to beat confirmation bias</p><p>14:30 The million dollar change you can't see</p><p>18:45 Sharing learnings and experimentation wrapped</p><p>20:45 Engagement without vanity metrics</p><p>22:00 Experimentation's golden era with AI</p><p>23:30 Why AI needs more experimentation, not less</p><p><strong>Takeaways</strong></p><ul><li>Chase estimates over a billion dollars of value from experimentation, and most of the lasting learning comes from the losing tests, not the winners.</li><li>A control group is non-negotiable: at scale, a change worth millions is invisible under noise and seasonality, and no one can spot it by eye.</li><li>Treat engagement carefully. For a bank, more time in the app isn't a win; trust, fast task completion, and healthy repeat engagement are.</li><li>Plan for failure before you run a test. A pre-built playbook for a loss prevents confirmation bias and keeps teams from gaming the metrics.</li><li>AI is ushering in a golden era for experimentation, because shipping faster only compounds mistakes unless you measure what you ship.</li></ul><p><strong>Connect with the Guest</strong></p><p><strong>LinkedIn:</strong> <a href="https://www.linkedin.com/in/kevintyang">https://www.linkedin.com/in/kevintyang</a></p><p><strong>Website:</strong> <a href="https://www.jpmorganchase.com">https://www.jpmorganchase.com</a></p><p><strong>Sponsor</strong><br>Growthbook helps you ship features with confidence by bringing experimentation and feature flagging into one open-source platform. No more guessing whether that new checkout flow actually moved the needle, waiting weeks for data team bandwidth, or flying blind on rollouts.</p><p>Growthbook gives you a single place to run A/B tests, manage feature flags, and analyze results against your existing data warehouse.</p><p>With powerful stats built in, it takes the complexity out of experimentation, helps you catch regressions before they hit every user, and makes it easy to test ideas that keep your product improving and your metrics moving in the right direction.</p><p>See a demo at <a href="https://www.growthbook.io/">https://www.growthbook.io/</a></p>]]>
      </description>
      <content:encoded>
        <![CDATA[<p><strong>Summary</strong></p><p>How do you know if the thing you just shipped actually worked? On this episode of The Experimentation Edge, host Ashley Stirrup, CMO of GrowthBook, sits down with Kevin Yang, Executive Director and Head of Experimentation at JPMorgan Chase, who has spent six years building experimentation across Chase's digital platforms. Kevin shares how his team turned experimentation into more than a billion dollars of estimated value, why the losing experiments matter more than the winners, and the simple chart exercise he uses to prove that a million-dollar change is invisible without a control group. He and Ashley also dig into measuring engagement without chasing vanity metrics, planning for failure to defeat confirmation bias, and why AI is pushing experimentation into a golden era. It's a practical look for product managers, data scientists, and engineers at how a bank operating at massive scale makes better decisions.</p><p><strong>Chapters</strong></p><p>00:00 Welcome to the experimentation edge</p><p>01:45 Kevin's role leading experimentation at chase</p><p>04:15 Why chase invested in experimentation</p><p>06:45 A billion dollars and the value of losers</p><p>12:45 Plan for failure to beat confirmation bias</p><p>14:30 The million dollar change you can't see</p><p>18:45 Sharing learnings and experimentation wrapped</p><p>20:45 Engagement without vanity metrics</p><p>22:00 Experimentation's golden era with AI</p><p>23:30 Why AI needs more experimentation, not less</p><p><strong>Takeaways</strong></p><ul><li>Chase estimates over a billion dollars of value from experimentation, and most of the lasting learning comes from the losing tests, not the winners.</li><li>A control group is non-negotiable: at scale, a change worth millions is invisible under noise and seasonality, and no one can spot it by eye.</li><li>Treat engagement carefully. For a bank, more time in the app isn't a win; trust, fast task completion, and healthy repeat engagement are.</li><li>Plan for failure before you run a test. A pre-built playbook for a loss prevents confirmation bias and keeps teams from gaming the metrics.</li><li>AI is ushering in a golden era for experimentation, because shipping faster only compounds mistakes unless you measure what you ship.</li></ul><p><strong>Connect with the Guest</strong></p><p><strong>LinkedIn:</strong> <a href="https://www.linkedin.com/in/kevintyang">https://www.linkedin.com/in/kevintyang</a></p><p><strong>Website:</strong> <a href="https://www.jpmorganchase.com">https://www.jpmorganchase.com</a></p><p><strong>Sponsor</strong><br>Growthbook helps you ship features with confidence by bringing experimentation and feature flagging into one open-source platform. No more guessing whether that new checkout flow actually moved the needle, waiting weeks for data team bandwidth, or flying blind on rollouts.</p><p>Growthbook gives you a single place to run A/B tests, manage feature flags, and analyze results against your existing data warehouse.</p><p>With powerful stats built in, it takes the complexity out of experimentation, helps you catch regressions before they hit every user, and makes it easy to test ideas that keep your product improving and your metrics moving in the right direction.</p><p>See a demo at <a href="https://www.growthbook.io/">https://www.growthbook.io/</a></p>]]>
      </content:encoded>
      <pubDate>Thu, 25 Jun 2026 07:05:10 -0600</pubDate>
      <author>Growthbook</author>
      <enclosure url="https://media.transistor.fm/6cf031b3/f31ed5a9.mp3" length="37666919" type="audio/mpeg"/>
      <itunes:author>Growthbook</itunes:author>
      <itunes:image href="https://img.transistorcdn.com/XUwnjrDKzBqMhv6OSX6QQ5rFK1H4w85cuHvTEfPJkik/rs:fill:0:0:1/w:1400/h:1400/q:60/mb:500000/aHR0cHM6Ly9pbWct/dXBsb2FkLXByb2R1/Y3Rpb24udHJhbnNp/c3Rvci5mbS81ZTM1/OTY0MjkzNjkwYjY0/MDUxMTgzNGM1N2Zl/ZWM3OS5wbmc.jpg"/>
      <itunes:duration>1568</itunes:duration>
      <itunes:summary>
        <![CDATA[<p><strong>Summary</strong></p><p>How do you know if the thing you just shipped actually worked? On this episode of The Experimentation Edge, host Ashley Stirrup, CMO of GrowthBook, sits down with Kevin Yang, Executive Director and Head of Experimentation at JPMorgan Chase, who has spent six years building experimentation across Chase's digital platforms. Kevin shares how his team turned experimentation into more than a billion dollars of estimated value, why the losing experiments matter more than the winners, and the simple chart exercise he uses to prove that a million-dollar change is invisible without a control group. He and Ashley also dig into measuring engagement without chasing vanity metrics, planning for failure to defeat confirmation bias, and why AI is pushing experimentation into a golden era. It's a practical look for product managers, data scientists, and engineers at how a bank operating at massive scale makes better decisions.</p><p><strong>Chapters</strong></p><p>00:00 Welcome to the experimentation edge</p><p>01:45 Kevin's role leading experimentation at chase</p><p>04:15 Why chase invested in experimentation</p><p>06:45 A billion dollars and the value of losers</p><p>12:45 Plan for failure to beat confirmation bias</p><p>14:30 The million dollar change you can't see</p><p>18:45 Sharing learnings and experimentation wrapped</p><p>20:45 Engagement without vanity metrics</p><p>22:00 Experimentation's golden era with AI</p><p>23:30 Why AI needs more experimentation, not less</p><p><strong>Takeaways</strong></p><ul><li>Chase estimates over a billion dollars of value from experimentation, and most of the lasting learning comes from the losing tests, not the winners.</li><li>A control group is non-negotiable: at scale, a change worth millions is invisible under noise and seasonality, and no one can spot it by eye.</li><li>Treat engagement carefully. For a bank, more time in the app isn't a win; trust, fast task completion, and healthy repeat engagement are.</li><li>Plan for failure before you run a test. A pre-built playbook for a loss prevents confirmation bias and keeps teams from gaming the metrics.</li><li>AI is ushering in a golden era for experimentation, because shipping faster only compounds mistakes unless you measure what you ship.</li></ul><p><strong>Connect with the Guest</strong></p><p><strong>LinkedIn:</strong> <a href="https://www.linkedin.com/in/kevintyang">https://www.linkedin.com/in/kevintyang</a></p><p><strong>Website:</strong> <a href="https://www.jpmorganchase.com">https://www.jpmorganchase.com</a></p><p><strong>Sponsor</strong><br>Growthbook helps you ship features with confidence by bringing experimentation and feature flagging into one open-source platform. No more guessing whether that new checkout flow actually moved the needle, waiting weeks for data team bandwidth, or flying blind on rollouts.</p><p>Growthbook gives you a single place to run A/B tests, manage feature flags, and analyze results against your existing data warehouse.</p><p>With powerful stats built in, it takes the complexity out of experimentation, helps you catch regressions before they hit every user, and makes it easy to test ideas that keep your product improving and your metrics moving in the right direction.</p><p>See a demo at <a href="https://www.growthbook.io/">https://www.growthbook.io/</a></p>]]>
      </itunes:summary>
      <itunes:keywords>experimentation,a/b testing,jpmorgan chase,kevin yang,ashley stirrup,growthbook,the experimentation edge,control group,product experimentation,digital banking,online experimentation,experimentation platform,data driven decisions,conversion optimization,AI experimentation,personalization models,experimentation culture,hypothesis testing,product analytics,statistical significance</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/6cf031b3/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Twitch on why false negatives kill product ideas</title>
      <itunes:episode>18</itunes:episode>
      <podcast:episode>18</podcast:episode>
      <itunes:title>Twitch on why false negatives kill product ideas</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">04ff2e01-f188-4645-99b5-901ed3f69fbe</guid>
      <link>https://share.transistor.fm/s/5dc7b86f</link>
      <description>
        <![CDATA[<p><strong>Summary</strong> </p><p>How do you make a high-stakes product decision when the safe choice is to never test it at all? In this episode of The Experimentation Edge, host Ashley Stirrup talks with Arun Bodapati, director of data science at Twitch, about the discipline behind trustworthy experimentation. Drawing on his experience at Schwab, Uber, and Twitch, Arun explains why false negatives are the most dangerous result a team can produce, what hygiene to nail before you push play, and how Twitch used geo-fenced experiments and causal inference to finally settle a pricing question it had avoided for years. It's a practical conversation for product managers, engineers, data scientists, and growth leaders who want experiments that hold up  and earn executive trust.</p><p> </p><p><strong>Chapters</strong></p><p>00:00 Welcome and introduction</p><p>01:15 Arun's background and marketing experimentation at Schwab</p><p>04:15 Uber's mature, experiment-driven culture</p><p>06:30 Coming to Twitch: from Python notebooks to a shared standard</p><p>08:30 The pricing problem Twitch had long avoided</p><p>10:30 Geo-fenced experiments, matched markets, and elasticity</p><p>13:15 The gifted-subs surprise and testing promotions</p><p>16:15 The discipline that matters before you push play</p><p>18:15 Why false negatives are worse than false positives</p><p>20:05 Enrollment triggers and broad explore experiments</p><p>22:45 AI, the Kiro tool, and what's next for experimentation<br></p><p><strong>Takeaways</strong> </p><ul><li>False negatives are more dangerous than false positives — they get institutionalized as "we tried that, it didn't work" and quietly kill good ideas for years.</li><li>The most valuable experiment work happens before you push play: clear enrollment logic, a plain-English hypothesis, and no optimizing ahead of the test.</li><li>If an intervention sounds weak when you write it out in plain English, don't run the experiment — you're just wasting time.</li><li>Run a broad explore experiment first; small, over-narrowed populations lack power and raise the odds of a false negative. Find the responsive segment with heterogeneous treatment effects afterward.</li><li>Twitch used geo-fenced experiments with matched markets and causal inference to measure true price elasticity, turning a feared pricing decision into a measured, accretive one.</li></ul><p><strong>Connect with the Guest</strong> </p><p>LinkedIn: <a href="https://www.linkedin.com/in/abodapati/">https://www.linkedin.com/in/abodapati/</a></p><p>Website: <a href="https://www.twitch.tv">https://www.twitch.tv</a></p><p><strong>Sponsor</strong><br>Growthbook helps you ship features with confidence by bringing experimentation and feature flagging into one open-source platform. No more guessing whether that new checkout flow actually moved the needle, waiting weeks for data team bandwidth, or flying blind on rollouts.</p><p>Growthbook gives you a single place to run A/B tests, manage feature flags, and analyze results against your existing data warehouse.</p><p>With powerful stats built in, it takes the complexity out of experimentation, helps you catch regressions before they hit every user, and makes it easy to test ideas that keep your product improving and your metrics moving in the right direction.</p><p>See a demo at <a href="https://www.growthbook.io/">https://www.growthbook.io/</a></p>]]>
      </description>
      <content:encoded>
        <![CDATA[<p><strong>Summary</strong> </p><p>How do you make a high-stakes product decision when the safe choice is to never test it at all? In this episode of The Experimentation Edge, host Ashley Stirrup talks with Arun Bodapati, director of data science at Twitch, about the discipline behind trustworthy experimentation. Drawing on his experience at Schwab, Uber, and Twitch, Arun explains why false negatives are the most dangerous result a team can produce, what hygiene to nail before you push play, and how Twitch used geo-fenced experiments and causal inference to finally settle a pricing question it had avoided for years. It's a practical conversation for product managers, engineers, data scientists, and growth leaders who want experiments that hold up  and earn executive trust.</p><p> </p><p><strong>Chapters</strong></p><p>00:00 Welcome and introduction</p><p>01:15 Arun's background and marketing experimentation at Schwab</p><p>04:15 Uber's mature, experiment-driven culture</p><p>06:30 Coming to Twitch: from Python notebooks to a shared standard</p><p>08:30 The pricing problem Twitch had long avoided</p><p>10:30 Geo-fenced experiments, matched markets, and elasticity</p><p>13:15 The gifted-subs surprise and testing promotions</p><p>16:15 The discipline that matters before you push play</p><p>18:15 Why false negatives are worse than false positives</p><p>20:05 Enrollment triggers and broad explore experiments</p><p>22:45 AI, the Kiro tool, and what's next for experimentation<br></p><p><strong>Takeaways</strong> </p><ul><li>False negatives are more dangerous than false positives — they get institutionalized as "we tried that, it didn't work" and quietly kill good ideas for years.</li><li>The most valuable experiment work happens before you push play: clear enrollment logic, a plain-English hypothesis, and no optimizing ahead of the test.</li><li>If an intervention sounds weak when you write it out in plain English, don't run the experiment — you're just wasting time.</li><li>Run a broad explore experiment first; small, over-narrowed populations lack power and raise the odds of a false negative. Find the responsive segment with heterogeneous treatment effects afterward.</li><li>Twitch used geo-fenced experiments with matched markets and causal inference to measure true price elasticity, turning a feared pricing decision into a measured, accretive one.</li></ul><p><strong>Connect with the Guest</strong> </p><p>LinkedIn: <a href="https://www.linkedin.com/in/abodapati/">https://www.linkedin.com/in/abodapati/</a></p><p>Website: <a href="https://www.twitch.tv">https://www.twitch.tv</a></p><p><strong>Sponsor</strong><br>Growthbook helps you ship features with confidence by bringing experimentation and feature flagging into one open-source platform. No more guessing whether that new checkout flow actually moved the needle, waiting weeks for data team bandwidth, or flying blind on rollouts.</p><p>Growthbook gives you a single place to run A/B tests, manage feature flags, and analyze results against your existing data warehouse.</p><p>With powerful stats built in, it takes the complexity out of experimentation, helps you catch regressions before they hit every user, and makes it easy to test ideas that keep your product improving and your metrics moving in the right direction.</p><p>See a demo at <a href="https://www.growthbook.io/">https://www.growthbook.io/</a></p>]]>
      </content:encoded>
      <pubDate>Wed, 24 Jun 2026 10:03:28 -0600</pubDate>
      <author>Growthbook</author>
      <enclosure url="https://media.transistor.fm/5dc7b86f/24c165a0.mp3" length="40692347" type="audio/mpeg"/>
      <itunes:author>Growthbook</itunes:author>
      <itunes:image href="https://img.transistorcdn.com/1epWkHtt-CQTn7Ff1InHgbNuO1gXsBRD0WlfUIgb4po/rs:fill:0:0:1/w:1400/h:1400/q:60/mb:500000/aHR0cHM6Ly9pbWct/dXBsb2FkLXByb2R1/Y3Rpb24udHJhbnNp/c3Rvci5mbS8zMmUx/MmYxNjUxMmM3ZjRl/ZDYzOGRjMjNjMWYx/N2Y2Ny5wbmc.jpg"/>
      <itunes:duration>1694</itunes:duration>
      <itunes:summary>
        <![CDATA[<p><strong>Summary</strong> </p><p>How do you make a high-stakes product decision when the safe choice is to never test it at all? In this episode of The Experimentation Edge, host Ashley Stirrup talks with Arun Bodapati, director of data science at Twitch, about the discipline behind trustworthy experimentation. Drawing on his experience at Schwab, Uber, and Twitch, Arun explains why false negatives are the most dangerous result a team can produce, what hygiene to nail before you push play, and how Twitch used geo-fenced experiments and causal inference to finally settle a pricing question it had avoided for years. It's a practical conversation for product managers, engineers, data scientists, and growth leaders who want experiments that hold up  and earn executive trust.</p><p> </p><p><strong>Chapters</strong></p><p>00:00 Welcome and introduction</p><p>01:15 Arun's background and marketing experimentation at Schwab</p><p>04:15 Uber's mature, experiment-driven culture</p><p>06:30 Coming to Twitch: from Python notebooks to a shared standard</p><p>08:30 The pricing problem Twitch had long avoided</p><p>10:30 Geo-fenced experiments, matched markets, and elasticity</p><p>13:15 The gifted-subs surprise and testing promotions</p><p>16:15 The discipline that matters before you push play</p><p>18:15 Why false negatives are worse than false positives</p><p>20:05 Enrollment triggers and broad explore experiments</p><p>22:45 AI, the Kiro tool, and what's next for experimentation<br></p><p><strong>Takeaways</strong> </p><ul><li>False negatives are more dangerous than false positives — they get institutionalized as "we tried that, it didn't work" and quietly kill good ideas for years.</li><li>The most valuable experiment work happens before you push play: clear enrollment logic, a plain-English hypothesis, and no optimizing ahead of the test.</li><li>If an intervention sounds weak when you write it out in plain English, don't run the experiment — you're just wasting time.</li><li>Run a broad explore experiment first; small, over-narrowed populations lack power and raise the odds of a false negative. Find the responsive segment with heterogeneous treatment effects afterward.</li><li>Twitch used geo-fenced experiments with matched markets and causal inference to measure true price elasticity, turning a feared pricing decision into a measured, accretive one.</li></ul><p><strong>Connect with the Guest</strong> </p><p>LinkedIn: <a href="https://www.linkedin.com/in/abodapati/">https://www.linkedin.com/in/abodapati/</a></p><p>Website: <a href="https://www.twitch.tv">https://www.twitch.tv</a></p><p><strong>Sponsor</strong><br>Growthbook helps you ship features with confidence by bringing experimentation and feature flagging into one open-source platform. No more guessing whether that new checkout flow actually moved the needle, waiting weeks for data team bandwidth, or flying blind on rollouts.</p><p>Growthbook gives you a single place to run A/B tests, manage feature flags, and analyze results against your existing data warehouse.</p><p>With powerful stats built in, it takes the complexity out of experimentation, helps you catch regressions before they hit every user, and makes it easy to test ideas that keep your product improving and your metrics moving in the right direction.</p><p>See a demo at <a href="https://www.growthbook.io/">https://www.growthbook.io/</a></p>]]>
      </itunes:summary>
      <itunes:keywords>experimentation,a/b testing,false negatives,product experimentation,causal inference,price elasticity,geo-fenced experiments,subscription pricing,twitch data science,experimentation culture,enrollment logic,experiment hygiene,statistical power,heterogeneous treatment effects,feature flags,bandits,thompson sampling,growthbook,experimentation program,data-driven decisions</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/5dc7b86f/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Squarespace killed its blank template and built something better</title>
      <itunes:episode>17</itunes:episode>
      <podcast:episode>17</podcast:episode>
      <itunes:title>Squarespace killed its blank template and built something better</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">42ca59b1-d634-4924-824f-2f6c7832021e</guid>
      <link>https://share.transistor.fm/s/3d779878</link>
      <description>
        <![CDATA[<p><strong>Summary<br></strong>What do you do when your big launch increases engagement and tanks conversion? On this episode of The Experimentation Edge, host Ashley Stirrup talks with Lina Blackman, Director of Product Analytics at Squarespace, about the blank template launch that flopped — and how its learnings became Blueprint, Squarespace's AI-guided website builder. Lina explains how her embedded analyst team runs 150–200 experiments a year for 3 million customers, the two questions she asks every time a test loses, why teams only need one or two big wins a quarter, how Squarespace calibrates statistical certainty to business stakes, and where AI belongs (and doesn't) in the A/B testing workflow. For product managers, data scientists, and experimentation leaders who want to extract more learning from every test.</p><p>Chapters 00:00 Introduction: Lina Blackman, Director of Product Analytics at Squarespace 01:45 Squarespace's business and 3 million website customers 02:30 Decentralized analysts, centralized experimentation program 04:15 150–200 experiments a year: onboarding, mobile, checkout, pricing 04:55 The blank template disaster that became Blueprint AI 07:45 Two questions for every losing test 09:30 Moving ship-first teams up the experimentation maturity curve 12:30 A/B test logs and insights rituals 13:30 North Star metrics and the KPI tree 16:35 AI in the A/B testing workflow — and what stays manual.<br></p><p><strong>Takeaways</strong></p><ul><li>Stated preference lies: users asked for a blank canvas, but behavior demanded guided design — and only the experiment could referee.</li><li>Close every losing test with two questions: did it work for a granular segment, and is the idea worth further investment?</li><li>One or two big wins a quarter is a healthy hit rate when you run 150–200 experiments a year.</li><li>Calibrate certainty to stakes — tight bounds on revenue and pricing tests, wider bounds on engagement tests so teams don't spin on noise.</li><li>Hand AI the mundane parts of the workflow (tracking, assignment setup), but if AI runs the brief and the analysis, ask why you're running the test at all.<p><br></p></li></ul><p><strong>Connect with the Guest</strong><br><strong>LinkedIn:</strong> <a href="https://www.linkedin.com/in/linanguyen">https://www.linkedin.com/in/linanguyen</a><br>Website: <a href="https://www.squarespace.com">https://www.squarespace.com</a></p><p><strong>Sponsor</strong><br>Growthbook helps you ship features with confidence by bringing experimentation and feature flagging into one open-source platform. No more guessing whether that new checkout flow actually moved the needle, waiting weeks for data team bandwidth, or flying blind on rollouts.</p><p>Growthbook gives you a single place to run A/B tests, manage feature flags, and analyze results against your existing data warehouse.</p><p>With powerful stats built in, it takes the complexity out of experimentation, helps you catch regressions before they hit every user, and makes it easy to test ideas that keep your product improving and your metrics moving in the right direction.</p><p>See a demo at <a href="https://www.growthbook.io/">https://www.growthbook.io/</a></p>]]>
      </description>
      <content:encoded>
        <![CDATA[<p><strong>Summary<br></strong>What do you do when your big launch increases engagement and tanks conversion? On this episode of The Experimentation Edge, host Ashley Stirrup talks with Lina Blackman, Director of Product Analytics at Squarespace, about the blank template launch that flopped — and how its learnings became Blueprint, Squarespace's AI-guided website builder. Lina explains how her embedded analyst team runs 150–200 experiments a year for 3 million customers, the two questions she asks every time a test loses, why teams only need one or two big wins a quarter, how Squarespace calibrates statistical certainty to business stakes, and where AI belongs (and doesn't) in the A/B testing workflow. For product managers, data scientists, and experimentation leaders who want to extract more learning from every test.</p><p>Chapters 00:00 Introduction: Lina Blackman, Director of Product Analytics at Squarespace 01:45 Squarespace's business and 3 million website customers 02:30 Decentralized analysts, centralized experimentation program 04:15 150–200 experiments a year: onboarding, mobile, checkout, pricing 04:55 The blank template disaster that became Blueprint AI 07:45 Two questions for every losing test 09:30 Moving ship-first teams up the experimentation maturity curve 12:30 A/B test logs and insights rituals 13:30 North Star metrics and the KPI tree 16:35 AI in the A/B testing workflow — and what stays manual.<br></p><p><strong>Takeaways</strong></p><ul><li>Stated preference lies: users asked for a blank canvas, but behavior demanded guided design — and only the experiment could referee.</li><li>Close every losing test with two questions: did it work for a granular segment, and is the idea worth further investment?</li><li>One or two big wins a quarter is a healthy hit rate when you run 150–200 experiments a year.</li><li>Calibrate certainty to stakes — tight bounds on revenue and pricing tests, wider bounds on engagement tests so teams don't spin on noise.</li><li>Hand AI the mundane parts of the workflow (tracking, assignment setup), but if AI runs the brief and the analysis, ask why you're running the test at all.<p><br></p></li></ul><p><strong>Connect with the Guest</strong><br><strong>LinkedIn:</strong> <a href="https://www.linkedin.com/in/linanguyen">https://www.linkedin.com/in/linanguyen</a><br>Website: <a href="https://www.squarespace.com">https://www.squarespace.com</a></p><p><strong>Sponsor</strong><br>Growthbook helps you ship features with confidence by bringing experimentation and feature flagging into one open-source platform. No more guessing whether that new checkout flow actually moved the needle, waiting weeks for data team bandwidth, or flying blind on rollouts.</p><p>Growthbook gives you a single place to run A/B tests, manage feature flags, and analyze results against your existing data warehouse.</p><p>With powerful stats built in, it takes the complexity out of experimentation, helps you catch regressions before they hit every user, and makes it easy to test ideas that keep your product improving and your metrics moving in the right direction.</p><p>See a demo at <a href="https://www.growthbook.io/">https://www.growthbook.io/</a></p>]]>
      </content:encoded>
      <pubDate>Tue, 23 Jun 2026 07:19:03 -0600</pubDate>
      <author>Growthbook</author>
      <enclosure url="https://media.transistor.fm/3d779878/8fdba122.mp3" length="32663230" type="audio/mpeg"/>
      <itunes:author>Growthbook</itunes:author>
      <itunes:image href="https://img.transistorcdn.com/7MLyqPTG_8DnuYgVehUiEqoMKHOEJ9iV8BwtJ1NShu8/rs:fill:0:0:1/w:1400/h:1400/q:60/mb:500000/aHR0cHM6Ly9pbWct/dXBsb2FkLXByb2R1/Y3Rpb24udHJhbnNp/c3Rvci5mbS85ZTYz/OWUxN2Q2Yjg5NjNk/NzIyYTY0OWQyOGE4/NGU5My5wbmc.jpg"/>
      <itunes:duration>1359</itunes:duration>
      <itunes:summary>
        <![CDATA[<p><strong>Summary<br></strong>What do you do when your big launch increases engagement and tanks conversion? On this episode of The Experimentation Edge, host Ashley Stirrup talks with Lina Blackman, Director of Product Analytics at Squarespace, about the blank template launch that flopped — and how its learnings became Blueprint, Squarespace's AI-guided website builder. Lina explains how her embedded analyst team runs 150–200 experiments a year for 3 million customers, the two questions she asks every time a test loses, why teams only need one or two big wins a quarter, how Squarespace calibrates statistical certainty to business stakes, and where AI belongs (and doesn't) in the A/B testing workflow. For product managers, data scientists, and experimentation leaders who want to extract more learning from every test.</p><p>Chapters 00:00 Introduction: Lina Blackman, Director of Product Analytics at Squarespace 01:45 Squarespace's business and 3 million website customers 02:30 Decentralized analysts, centralized experimentation program 04:15 150–200 experiments a year: onboarding, mobile, checkout, pricing 04:55 The blank template disaster that became Blueprint AI 07:45 Two questions for every losing test 09:30 Moving ship-first teams up the experimentation maturity curve 12:30 A/B test logs and insights rituals 13:30 North Star metrics and the KPI tree 16:35 AI in the A/B testing workflow — and what stays manual.<br></p><p><strong>Takeaways</strong></p><ul><li>Stated preference lies: users asked for a blank canvas, but behavior demanded guided design — and only the experiment could referee.</li><li>Close every losing test with two questions: did it work for a granular segment, and is the idea worth further investment?</li><li>One or two big wins a quarter is a healthy hit rate when you run 150–200 experiments a year.</li><li>Calibrate certainty to stakes — tight bounds on revenue and pricing tests, wider bounds on engagement tests so teams don't spin on noise.</li><li>Hand AI the mundane parts of the workflow (tracking, assignment setup), but if AI runs the brief and the analysis, ask why you're running the test at all.<p><br></p></li></ul><p><strong>Connect with the Guest</strong><br><strong>LinkedIn:</strong> <a href="https://www.linkedin.com/in/linanguyen">https://www.linkedin.com/in/linanguyen</a><br>Website: <a href="https://www.squarespace.com">https://www.squarespace.com</a></p><p><strong>Sponsor</strong><br>Growthbook helps you ship features with confidence by bringing experimentation and feature flagging into one open-source platform. No more guessing whether that new checkout flow actually moved the needle, waiting weeks for data team bandwidth, or flying blind on rollouts.</p><p>Growthbook gives you a single place to run A/B tests, manage feature flags, and analyze results against your existing data warehouse.</p><p>With powerful stats built in, it takes the complexity out of experimentation, helps you catch regressions before they hit every user, and makes it easy to test ideas that keep your product improving and your metrics moving in the right direction.</p><p>See a demo at <a href="https://www.growthbook.io/">https://www.growthbook.io/</a></p>]]>
      </itunes:summary>
      <itunes:keywords>ab testing,product analytics,experimentation program,squarespace experimentation,failed ab tests,losing experiments,blueprint ai,user segmentation,product onboarding,north star metric,kpi tree,experimentation maturity,testing velocity,guardrail metrics,statistical significance,ai in experimentation,llm product features,experimentation culture,decision matrix,growthbook</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
    </item>
    <item>
      <title>Signet Jeweler's "View All" page made more money by showing less</title>
      <itunes:episode>16</itunes:episode>
      <podcast:episode>16</podcast:episode>
      <itunes:title>Signet Jeweler's "View All" page made more money by showing less</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">c31b010d-8fb9-4e7f-ac2d-71672419dc66</guid>
      <link>https://share.transistor.fm/s/9eec445a</link>
      <description>
        <![CDATA[<p><strong>Summary<br></strong>Craig Kistler, VP of Experience Design, Personalization, and Experimentation at Signet Jewelers (the parent company of Kay, Jared, Zales, Peoples, and Banter), joins host Ashley Stirrup on The Experimentation Edge to unpack how a hybrid online-and-in-store jewelry retailer runs experimentation at scale. Craig shares the counterintuitive "view all" experiment where his team blocked the product grid, added friction on purpose, and grew revenue; why he optimizes for revenue per visitor instead of conversion rate; and how Signet deliberately traded a high-volume testing program for fewer, higher-value experiments. A practical listen for product managers, designers, and experimentation leaders building programs that compound.<br></p><p><strong>Chapters<br></strong>00:45 What Signet Jewelers actually is (Kay, Jared, Zales, and more)<br>01:45 From art school to UX to experimentation: Craig's background<br>04:45 How experimentation is organized: a centralized model across brands<br>06:45 From 40–50 tests a quarter to 15–25 value-driven experiments<br>08:45 The "view all" experiment: adding friction to grow revenue<br>12:45 One product page, many stakeholders: financing, warranties, chat<br>15:45 Extracting learnings when an experiment loses<br>18:45 Why revenue per visitor beats conversion as the north star<br>20:15 Intent-based personalization and "engagement season is every day"<br>24:45 Bringing the whole org along by tying insights to dollars.<br></p><p><strong>Takeaways</strong></p><ul><li>Friction can increase revenue. Blocking the "view all" grid and forcing a style choice sent shoppers deeper and lifted conversion and revenue, because the extra click added value.</li><li>The three-click rule is conditional. Clicks only hurt when they're empty; a click that narrows thousands of options to dozens is a feature, not a cost.</li><li>Revenue per visitor is the honest north star. Conversion rate can be gamed to 100% by making everything free or cutting bounce-heavy traffic; revenue per visitor can't.</li><li>Fewer, bigger experiments beat high volume. Signet went from 40–50 tests a quarter to 15–25 because complex, value-driven tests produce reusable insights that small tweaks don't.</li><li>Tie every result to dollars. Translating experiment outcomes into revenue is how Craig keeps financing, warranty, and chat stakeholders aligned and gets executives to act.<p></p></li></ul><p><strong>Connect with the Guest<br></strong>LinkedIn: <a href="https://www.linkedin.com/in/craigkistler/">https://www.linkedin.com/in/craigkistler/</a><br>Website: <a href="https://www.signetjewelers.com">https://www.signetjewelers.com</a></p><p><strong>Sponsor</strong><br>Growthbook helps you ship features with confidence by bringing experimentation and feature flagging into one open-source platform. No more guessing whether that new checkout flow actually moved the needle, waiting weeks for data team bandwidth, or flying blind on rollouts.</p><p>Growthbook gives you a single place to run A/B tests, manage feature flags, and analyze results against your existing data warehouse.</p><p>With powerful stats built in, it takes the complexity out of experimentation, helps you catch regressions before they hit every user, and makes it easy to test ideas that keep your product improving and your metrics moving in the right direction.</p><p>See a demo at <a href="https://www.growthbook.io/">https://www.growthbook.io/</a></p>]]>
      </description>
      <content:encoded>
        <![CDATA[<p><strong>Summary<br></strong>Craig Kistler, VP of Experience Design, Personalization, and Experimentation at Signet Jewelers (the parent company of Kay, Jared, Zales, Peoples, and Banter), joins host Ashley Stirrup on The Experimentation Edge to unpack how a hybrid online-and-in-store jewelry retailer runs experimentation at scale. Craig shares the counterintuitive "view all" experiment where his team blocked the product grid, added friction on purpose, and grew revenue; why he optimizes for revenue per visitor instead of conversion rate; and how Signet deliberately traded a high-volume testing program for fewer, higher-value experiments. A practical listen for product managers, designers, and experimentation leaders building programs that compound.<br></p><p><strong>Chapters<br></strong>00:45 What Signet Jewelers actually is (Kay, Jared, Zales, and more)<br>01:45 From art school to UX to experimentation: Craig's background<br>04:45 How experimentation is organized: a centralized model across brands<br>06:45 From 40–50 tests a quarter to 15–25 value-driven experiments<br>08:45 The "view all" experiment: adding friction to grow revenue<br>12:45 One product page, many stakeholders: financing, warranties, chat<br>15:45 Extracting learnings when an experiment loses<br>18:45 Why revenue per visitor beats conversion as the north star<br>20:15 Intent-based personalization and "engagement season is every day"<br>24:45 Bringing the whole org along by tying insights to dollars.<br></p><p><strong>Takeaways</strong></p><ul><li>Friction can increase revenue. Blocking the "view all" grid and forcing a style choice sent shoppers deeper and lifted conversion and revenue, because the extra click added value.</li><li>The three-click rule is conditional. Clicks only hurt when they're empty; a click that narrows thousands of options to dozens is a feature, not a cost.</li><li>Revenue per visitor is the honest north star. Conversion rate can be gamed to 100% by making everything free or cutting bounce-heavy traffic; revenue per visitor can't.</li><li>Fewer, bigger experiments beat high volume. Signet went from 40–50 tests a quarter to 15–25 because complex, value-driven tests produce reusable insights that small tweaks don't.</li><li>Tie every result to dollars. Translating experiment outcomes into revenue is how Craig keeps financing, warranty, and chat stakeholders aligned and gets executives to act.<p></p></li></ul><p><strong>Connect with the Guest<br></strong>LinkedIn: <a href="https://www.linkedin.com/in/craigkistler/">https://www.linkedin.com/in/craigkistler/</a><br>Website: <a href="https://www.signetjewelers.com">https://www.signetjewelers.com</a></p><p><strong>Sponsor</strong><br>Growthbook helps you ship features with confidence by bringing experimentation and feature flagging into one open-source platform. No more guessing whether that new checkout flow actually moved the needle, waiting weeks for data team bandwidth, or flying blind on rollouts.</p><p>Growthbook gives you a single place to run A/B tests, manage feature flags, and analyze results against your existing data warehouse.</p><p>With powerful stats built in, it takes the complexity out of experimentation, helps you catch regressions before they hit every user, and makes it easy to test ideas that keep your product improving and your metrics moving in the right direction.</p><p>See a demo at <a href="https://www.growthbook.io/">https://www.growthbook.io/</a></p>]]>
      </content:encoded>
      <pubDate>Wed, 17 Jun 2026 08:00:00 -0600</pubDate>
      <author>Growthbook</author>
      <enclosure url="https://media.transistor.fm/9eec445a/03d3734a.mp3" length="39620756" type="audio/mpeg"/>
      <itunes:author>Growthbook</itunes:author>
      <itunes:image href="https://img.transistorcdn.com/O0zsewo5CBmFVWtqDy6eYWSqILlmRADyqT9UmMTSVpA/rs:fill:0:0:1/w:1400/h:1400/q:60/mb:500000/aHR0cHM6Ly9pbWct/dXBsb2FkLXByb2R1/Y3Rpb24udHJhbnNp/c3Rvci5mbS9jNzY5/NDZiZWY1MjU1ZGQ5/OGY4ZDExNGVlNDFl/M2YxOC5wbmc.jpg"/>
      <itunes:duration>1649</itunes:duration>
      <itunes:summary>
        <![CDATA[<p><strong>Summary<br></strong>Craig Kistler, VP of Experience Design, Personalization, and Experimentation at Signet Jewelers (the parent company of Kay, Jared, Zales, Peoples, and Banter), joins host Ashley Stirrup on The Experimentation Edge to unpack how a hybrid online-and-in-store jewelry retailer runs experimentation at scale. Craig shares the counterintuitive "view all" experiment where his team blocked the product grid, added friction on purpose, and grew revenue; why he optimizes for revenue per visitor instead of conversion rate; and how Signet deliberately traded a high-volume testing program for fewer, higher-value experiments. A practical listen for product managers, designers, and experimentation leaders building programs that compound.<br></p><p><strong>Chapters<br></strong>00:45 What Signet Jewelers actually is (Kay, Jared, Zales, and more)<br>01:45 From art school to UX to experimentation: Craig's background<br>04:45 How experimentation is organized: a centralized model across brands<br>06:45 From 40–50 tests a quarter to 15–25 value-driven experiments<br>08:45 The "view all" experiment: adding friction to grow revenue<br>12:45 One product page, many stakeholders: financing, warranties, chat<br>15:45 Extracting learnings when an experiment loses<br>18:45 Why revenue per visitor beats conversion as the north star<br>20:15 Intent-based personalization and "engagement season is every day"<br>24:45 Bringing the whole org along by tying insights to dollars.<br></p><p><strong>Takeaways</strong></p><ul><li>Friction can increase revenue. Blocking the "view all" grid and forcing a style choice sent shoppers deeper and lifted conversion and revenue, because the extra click added value.</li><li>The three-click rule is conditional. Clicks only hurt when they're empty; a click that narrows thousands of options to dozens is a feature, not a cost.</li><li>Revenue per visitor is the honest north star. Conversion rate can be gamed to 100% by making everything free or cutting bounce-heavy traffic; revenue per visitor can't.</li><li>Fewer, bigger experiments beat high volume. Signet went from 40–50 tests a quarter to 15–25 because complex, value-driven tests produce reusable insights that small tweaks don't.</li><li>Tie every result to dollars. Translating experiment outcomes into revenue is how Craig keeps financing, warranty, and chat stakeholders aligned and gets executives to act.<p></p></li></ul><p><strong>Connect with the Guest<br></strong>LinkedIn: <a href="https://www.linkedin.com/in/craigkistler/">https://www.linkedin.com/in/craigkistler/</a><br>Website: <a href="https://www.signetjewelers.com">https://www.signetjewelers.com</a></p><p><strong>Sponsor</strong><br>Growthbook helps you ship features with confidence by bringing experimentation and feature flagging into one open-source platform. No more guessing whether that new checkout flow actually moved the needle, waiting weeks for data team bandwidth, or flying blind on rollouts.</p><p>Growthbook gives you a single place to run A/B tests, manage feature flags, and analyze results against your existing data warehouse.</p><p>With powerful stats built in, it takes the complexity out of experimentation, helps you catch regressions before they hit every user, and makes it easy to test ideas that keep your product improving and your metrics moving in the right direction.</p><p>See a demo at <a href="https://www.growthbook.io/">https://www.growthbook.io/</a></p>]]>
      </itunes:summary>
      <itunes:keywords>experimentation,a/b testing,revenue per visitor,conversion rate optimization,ecommerce experimentation,product detail page,product listing page,signet jewelers,craig kistler,the experimentation edge,growthbook,intent-based personalization,ux friction,view all page,testing velocity,north star metric,guardrail metrics,experimentation program,ecommerce personalization,ashley stirrup</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/9eec445a/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>RingCentral's DART framework: The four metrics that actually measure AI agents</title>
      <itunes:episode>15</itunes:episode>
      <podcast:episode>15</podcast:episode>
      <itunes:title>RingCentral's DART framework: The four metrics that actually measure AI agents</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">0f763bde-30c2-46a3-af10-b1799d83ac68</guid>
      <link>https://share.transistor.fm/s/ecb85e39</link>
      <description>
        <![CDATA[<p><strong>Summary<br></strong>RingCentral's Director of Product Management for AI Products, Mayank Agarwal, joins host Ashley Stirrup to dismantle the metrics most teams use to judge AI agents. Drawing on his background founding an AI-first quantitative trading firm and scaling Groupon's bookable marketplace, Mayank explains why accuracy and thumbs-up/down feedback both mislead, and introduces DART — a four-metric behavioral framework (decay, acceptance, relevance, task completion) ported from how he measured trading strategies. He also breaks down a Groupon flash-discount experiment that backfired and the scarcity pivot that fixed it. Essential listening for product managers, engineers, and data scientists building or measuring AI features.</p><p><strong><br>Chapters</strong></p><p>00:00 Welcome and Mayank's path from quant trading to RingCentral AI<br>02:45 Why experimentation has to be owned cross-functionally<br>04:55 Small experiments that compounded to a 12% lift at Groupon<br>06:45 Why accuracy and thumbs-up/down fail for AI agents<br>08:15 The DART framework, metric by metric<br>12:45 Applying DART to AI-generated smart notes<br>14:55 The Groupon flash-sale that dropped conversion<br>16:45 Swapping price urgency for scarcity and social proof<br>19:45 North Star metrics, guardrails, and Goodhart's law<br>26:45 The future: experimenting on — and for — AI agents</p><p><br></p><p><strong><br>Takeaways</strong></p><ul><li>Accuracy is a comfortable lie. It grades a narrow test set and can stay high while the agent fails real users.</li><li>Thumbs-up/down feedback is sparse and skewed. Unhappy users rarely rate — they just quietly stop using the product.</li><li>DART measures behavior, not opinions. Four signals read off logs and transcripts: decay, acceptance, relevance, and task completion.</li><li>Acceptance rate is the trust metric. The share of output users keep without editing is the strongest available proxy for trust.</li><li>A losing experiment is paid-for information. Groupon's flash-sale flop revealed the lever was wrong, not the goal — scarcity beat price-based urgency.</li></ul><p><br><strong><br>Connect with the Guest<br></strong>LinkedIn:<a href="https://www.linkedin.com/in/mayank-agarwal-6223b04a/"> https://www.linkedin.com/in/mayank-agarwal-6223b04a/</a><br>Website:<a href="https://www.ringcentral.com"> https://www.ringcentral.com<br></a><br></p><p><br></p><p><strong>Sponsor</strong><br>Growthbook helps you ship features with confidence by bringing experimentation and feature flagging into one open-source platform. No more guessing whether that new checkout flow actually moved the needle, waiting weeks for data team bandwidth, or flying blind on rollouts.</p><p>Growthbook gives you a single place to run A/B tests, manage feature flags, and analyze results against your existing data warehouse.</p><p>With powerful stats built in, it takes the complexity out of experimentation, helps you catch regressions before they hit every user, and makes it easy to test ideas that keep your product improving and your metrics moving in the right direction.</p><p>See a demo at <a href="https://www.growthbook.io/">https://www.growthbook.io/</a></p>]]>
      </description>
      <content:encoded>
        <![CDATA[<p><strong>Summary<br></strong>RingCentral's Director of Product Management for AI Products, Mayank Agarwal, joins host Ashley Stirrup to dismantle the metrics most teams use to judge AI agents. Drawing on his background founding an AI-first quantitative trading firm and scaling Groupon's bookable marketplace, Mayank explains why accuracy and thumbs-up/down feedback both mislead, and introduces DART — a four-metric behavioral framework (decay, acceptance, relevance, task completion) ported from how he measured trading strategies. He also breaks down a Groupon flash-discount experiment that backfired and the scarcity pivot that fixed it. Essential listening for product managers, engineers, and data scientists building or measuring AI features.</p><p><strong><br>Chapters</strong></p><p>00:00 Welcome and Mayank's path from quant trading to RingCentral AI<br>02:45 Why experimentation has to be owned cross-functionally<br>04:55 Small experiments that compounded to a 12% lift at Groupon<br>06:45 Why accuracy and thumbs-up/down fail for AI agents<br>08:15 The DART framework, metric by metric<br>12:45 Applying DART to AI-generated smart notes<br>14:55 The Groupon flash-sale that dropped conversion<br>16:45 Swapping price urgency for scarcity and social proof<br>19:45 North Star metrics, guardrails, and Goodhart's law<br>26:45 The future: experimenting on — and for — AI agents</p><p><br></p><p><strong><br>Takeaways</strong></p><ul><li>Accuracy is a comfortable lie. It grades a narrow test set and can stay high while the agent fails real users.</li><li>Thumbs-up/down feedback is sparse and skewed. Unhappy users rarely rate — they just quietly stop using the product.</li><li>DART measures behavior, not opinions. Four signals read off logs and transcripts: decay, acceptance, relevance, and task completion.</li><li>Acceptance rate is the trust metric. The share of output users keep without editing is the strongest available proxy for trust.</li><li>A losing experiment is paid-for information. Groupon's flash-sale flop revealed the lever was wrong, not the goal — scarcity beat price-based urgency.</li></ul><p><br><strong><br>Connect with the Guest<br></strong>LinkedIn:<a href="https://www.linkedin.com/in/mayank-agarwal-6223b04a/"> https://www.linkedin.com/in/mayank-agarwal-6223b04a/</a><br>Website:<a href="https://www.ringcentral.com"> https://www.ringcentral.com<br></a><br></p><p><br></p><p><strong>Sponsor</strong><br>Growthbook helps you ship features with confidence by bringing experimentation and feature flagging into one open-source platform. No more guessing whether that new checkout flow actually moved the needle, waiting weeks for data team bandwidth, or flying blind on rollouts.</p><p>Growthbook gives you a single place to run A/B tests, manage feature flags, and analyze results against your existing data warehouse.</p><p>With powerful stats built in, it takes the complexity out of experimentation, helps you catch regressions before they hit every user, and makes it easy to test ideas that keep your product improving and your metrics moving in the right direction.</p><p>See a demo at <a href="https://www.growthbook.io/">https://www.growthbook.io/</a></p>]]>
      </content:encoded>
      <pubDate>Mon, 15 Jun 2026 08:00:00 -0600</pubDate>
      <author>Growthbook</author>
      <enclosure url="https://media.transistor.fm/ecb85e39/b932d6bd.mp3" length="46635806" type="audio/mpeg"/>
      <itunes:author>Growthbook</itunes:author>
      <itunes:image href="https://img.transistorcdn.com/vgPrybtOhGtcG22xo-AuPARiD2Uyal_mMKLiRXJ6Gu4/rs:fill:0:0:1/w:1400/h:1400/q:60/mb:500000/aHR0cHM6Ly9pbWct/dXBsb2FkLXByb2R1/Y3Rpb24udHJhbnNp/c3Rvci5mbS9jNzk2/MGY5OWY3NTgyNGYx/N2FkMDdjODBkMjE2/YTE0Zi5wbmc.jpg"/>
      <itunes:duration>1942</itunes:duration>
      <itunes:summary>
        <![CDATA[<p><strong>Summary<br></strong>RingCentral's Director of Product Management for AI Products, Mayank Agarwal, joins host Ashley Stirrup to dismantle the metrics most teams use to judge AI agents. Drawing on his background founding an AI-first quantitative trading firm and scaling Groupon's bookable marketplace, Mayank explains why accuracy and thumbs-up/down feedback both mislead, and introduces DART — a four-metric behavioral framework (decay, acceptance, relevance, task completion) ported from how he measured trading strategies. He also breaks down a Groupon flash-discount experiment that backfired and the scarcity pivot that fixed it. Essential listening for product managers, engineers, and data scientists building or measuring AI features.</p><p><strong><br>Chapters</strong></p><p>00:00 Welcome and Mayank's path from quant trading to RingCentral AI<br>02:45 Why experimentation has to be owned cross-functionally<br>04:55 Small experiments that compounded to a 12% lift at Groupon<br>06:45 Why accuracy and thumbs-up/down fail for AI agents<br>08:15 The DART framework, metric by metric<br>12:45 Applying DART to AI-generated smart notes<br>14:55 The Groupon flash-sale that dropped conversion<br>16:45 Swapping price urgency for scarcity and social proof<br>19:45 North Star metrics, guardrails, and Goodhart's law<br>26:45 The future: experimenting on — and for — AI agents</p><p><br></p><p><strong><br>Takeaways</strong></p><ul><li>Accuracy is a comfortable lie. It grades a narrow test set and can stay high while the agent fails real users.</li><li>Thumbs-up/down feedback is sparse and skewed. Unhappy users rarely rate — they just quietly stop using the product.</li><li>DART measures behavior, not opinions. Four signals read off logs and transcripts: decay, acceptance, relevance, and task completion.</li><li>Acceptance rate is the trust metric. The share of output users keep without editing is the strongest available proxy for trust.</li><li>A losing experiment is paid-for information. Groupon's flash-sale flop revealed the lever was wrong, not the goal — scarcity beat price-based urgency.</li></ul><p><br><strong><br>Connect with the Guest<br></strong>LinkedIn:<a href="https://www.linkedin.com/in/mayank-agarwal-6223b04a/"> https://www.linkedin.com/in/mayank-agarwal-6223b04a/</a><br>Website:<a href="https://www.ringcentral.com"> https://www.ringcentral.com<br></a><br></p><p><br></p><p><strong>Sponsor</strong><br>Growthbook helps you ship features with confidence by bringing experimentation and feature flagging into one open-source platform. No more guessing whether that new checkout flow actually moved the needle, waiting weeks for data team bandwidth, or flying blind on rollouts.</p><p>Growthbook gives you a single place to run A/B tests, manage feature flags, and analyze results against your existing data warehouse.</p><p>With powerful stats built in, it takes the complexity out of experimentation, helps you catch regressions before they hit every user, and makes it easy to test ideas that keep your product improving and your metrics moving in the right direction.</p><p>See a demo at <a href="https://www.growthbook.io/">https://www.growthbook.io/</a></p>]]>
      </itunes:summary>
      <itunes:keywords>experimentation,a/b testing,ai agents,dart framework,ai metrics,acceptance rate,task completion rate,signal to noise,decay rate,behavioral metrics,guardrail metrics,goodhart's law,product management,ringcentral,groupon,conversion rate optimization,scarcity marketing,social proof,llm evaluation,north star metric</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/ecb85e39/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>The 2% close rate increase that turned Ford Credit's product teams into believers</title>
      <itunes:episode>14</itunes:episode>
      <podcast:episode>14</podcast:episode>
      <itunes:title>The 2% close rate increase that turned Ford Credit's product teams into believers</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">e7e90d7c-e7b8-4f2e-9eb8-a54609113851</guid>
      <link>https://share.transistor.fm/s/c3607cb8</link>
      <description>
        <![CDATA[<p><strong>Summary</strong></p><p>On this edition of The Experimentation Edge, Ashley Stirrup talks with Geoffrey Bell, Experimentation Product Specialist at Ford Credit, about building an experimentation practice inside a captive auto lender. Geoffrey shares the losing test that earned his program credibility, the "experimentation piggy bank" he picked up at Microsoft, and the breakthrough of connecting online experiments to offline dealership receivables. The throughline: a program proves its worth not just by the wins it ships, but by the expensive mistakes it prevents and the revenue it can finally trace. It's for product managers, data scientists, and growth leaders who want experimentation taken seriously by the business.</p><p><strong>Chapters</strong></p><p>00:00 Intro</p><p>01:15 Geoffrey's path: Lowe's, Microsoft, Ford Credit</p><p>07:15 How Ford Credit fits with Ford Motor</p><p>10:15 The teams behind every Ford Credit page</p><p>15:15 The vehicle selector test that lost on purpose</p><p>19:15 Why feature placement beats feature ideas</p><p>21:15 Personalization and the shrinking-audience problem</p><p>25:15 Telling the story when a test loses</p><p>30:45 Connecting an online test to an offline car sale</p><p>33:55 The experimentation piggy bank</p><p><strong>Takeaways</strong></p><p>1. Losing tests often create more value than winners because they stop expensive mistakes before they ship.</p><p>2. Measure experimentation two ways: the revenue you earn from wins and the revenue you save by killing bad experiences.</p><p>3. A feature that fails early in a flow can succeed later; placement and timing often matter more than the idea itself.</p><p>4. Connecting online experiments to offline outcomes like receivables turns a small lift into a number leadership cares about.</p><p>5. When you struggle to land a result, lead with the story of what the customer did, then bring the numbers.</p><p><br></p><p><br><strong>Connect with the Guest<br></strong>LinkedIn: <a href="https://www.linkedin.com/in/geoffrey-bell-62a03617/">https://www.linkedin.com/in/geoffrey-bell-62a03617/</a> <br>Website: <a href="https://www.ford.com/finance/">https://www.ford.com/finance/</a></p><p><br><strong>Sponsor</strong><br>Growthbook helps you ship features with confidence by bringing experimentation and feature flagging into one open-source platform. No more guessing whether that new checkout flow actually moved the needle, waiting weeks for data team bandwidth, or flying blind on rollouts.</p><p>Growthbook gives you a single place to run A/B tests, manage feature flags, and analyze results against your existing data warehouse.</p><p>With powerful stats built in, it takes the complexity out of experimentation, helps you catch regressions before they hit every user, and makes it easy to test ideas that keep your product improving and your metrics moving in the right direction.</p><p><br>See a demo at <a href="https://www.growthbook.io/">https://www.growthbook.io/</a> </p>
<ul><li>(00:00) - Intro</li>
<li>(01:15) - Geoffrey's path: Lowe's, Microsoft, Ford Credit</li>
<li>(07:15) - How Ford Credit fits with Ford Motor</li>
<li>(10:15) - The teams behind every Ford Credit page</li>
<li>(15:15) - The vehicle selector test that lost on purpose</li>
<li>(19:15) - Why feature placement beats feature ideas</li>
<li>(21:15) - Personalization and the shrinking-audience problem</li>
<li>(25:15) - Telling the story when a test loses</li>
<li>(30:45) - Connecting an online test to an offline car sale</li>
<li>(33:55) - The experimentation piggy bank</li>
</ul>]]>
      </description>
      <content:encoded>
        <![CDATA[<p><strong>Summary</strong></p><p>On this edition of The Experimentation Edge, Ashley Stirrup talks with Geoffrey Bell, Experimentation Product Specialist at Ford Credit, about building an experimentation practice inside a captive auto lender. Geoffrey shares the losing test that earned his program credibility, the "experimentation piggy bank" he picked up at Microsoft, and the breakthrough of connecting online experiments to offline dealership receivables. The throughline: a program proves its worth not just by the wins it ships, but by the expensive mistakes it prevents and the revenue it can finally trace. It's for product managers, data scientists, and growth leaders who want experimentation taken seriously by the business.</p><p><strong>Chapters</strong></p><p>00:00 Intro</p><p>01:15 Geoffrey's path: Lowe's, Microsoft, Ford Credit</p><p>07:15 How Ford Credit fits with Ford Motor</p><p>10:15 The teams behind every Ford Credit page</p><p>15:15 The vehicle selector test that lost on purpose</p><p>19:15 Why feature placement beats feature ideas</p><p>21:15 Personalization and the shrinking-audience problem</p><p>25:15 Telling the story when a test loses</p><p>30:45 Connecting an online test to an offline car sale</p><p>33:55 The experimentation piggy bank</p><p><strong>Takeaways</strong></p><p>1. Losing tests often create more value than winners because they stop expensive mistakes before they ship.</p><p>2. Measure experimentation two ways: the revenue you earn from wins and the revenue you save by killing bad experiences.</p><p>3. A feature that fails early in a flow can succeed later; placement and timing often matter more than the idea itself.</p><p>4. Connecting online experiments to offline outcomes like receivables turns a small lift into a number leadership cares about.</p><p>5. When you struggle to land a result, lead with the story of what the customer did, then bring the numbers.</p><p><br></p><p><br><strong>Connect with the Guest<br></strong>LinkedIn: <a href="https://www.linkedin.com/in/geoffrey-bell-62a03617/">https://www.linkedin.com/in/geoffrey-bell-62a03617/</a> <br>Website: <a href="https://www.ford.com/finance/">https://www.ford.com/finance/</a></p><p><br><strong>Sponsor</strong><br>Growthbook helps you ship features with confidence by bringing experimentation and feature flagging into one open-source platform. No more guessing whether that new checkout flow actually moved the needle, waiting weeks for data team bandwidth, or flying blind on rollouts.</p><p>Growthbook gives you a single place to run A/B tests, manage feature flags, and analyze results against your existing data warehouse.</p><p>With powerful stats built in, it takes the complexity out of experimentation, helps you catch regressions before they hit every user, and makes it easy to test ideas that keep your product improving and your metrics moving in the right direction.</p><p><br>See a demo at <a href="https://www.growthbook.io/">https://www.growthbook.io/</a> </p>
<ul><li>(00:00) - Intro</li>
<li>(01:15) - Geoffrey's path: Lowe's, Microsoft, Ford Credit</li>
<li>(07:15) - How Ford Credit fits with Ford Motor</li>
<li>(10:15) - The teams behind every Ford Credit page</li>
<li>(15:15) - The vehicle selector test that lost on purpose</li>
<li>(19:15) - Why feature placement beats feature ideas</li>
<li>(21:15) - Personalization and the shrinking-audience problem</li>
<li>(25:15) - Telling the story when a test loses</li>
<li>(30:45) - Connecting an online test to an offline car sale</li>
<li>(33:55) - The experimentation piggy bank</li>
</ul>]]>
      </content:encoded>
      <pubDate>Tue, 02 Jun 2026 08:47:53 -0600</pubDate>
      <author>Growthbook</author>
      <enclosure url="https://media.transistor.fm/c3607cb8/b92b83cc.mp3" length="35423913" type="audio/mpeg"/>
      <itunes:author>Growthbook</itunes:author>
      <itunes:image href="https://img.transistorcdn.com/IhgAJES-DtJ28fiBfcLny2FonoW-YqNhZxyOyIPChyA/rs:fill:0:0:1/w:1400/h:1400/q:60/mb:500000/aHR0cHM6Ly9pbWct/dXBsb2FkLXByb2R1/Y3Rpb24udHJhbnNp/c3Rvci5mbS9lZmI0/ZWI4OTNlZDJmOTBi/ODliNzJiZGZmNDIy/Njc0NS5wbmc.jpg"/>
      <itunes:duration>2211</itunes:duration>
      <itunes:summary>
        <![CDATA[<p><strong>Summary</strong></p><p>On this edition of The Experimentation Edge, Ashley Stirrup talks with Geoffrey Bell, Experimentation Product Specialist at Ford Credit, about building an experimentation practice inside a captive auto lender. Geoffrey shares the losing test that earned his program credibility, the "experimentation piggy bank" he picked up at Microsoft, and the breakthrough of connecting online experiments to offline dealership receivables. The throughline: a program proves its worth not just by the wins it ships, but by the expensive mistakes it prevents and the revenue it can finally trace. It's for product managers, data scientists, and growth leaders who want experimentation taken seriously by the business.</p><p><strong>Chapters</strong></p><p>00:00 Intro</p><p>01:15 Geoffrey's path: Lowe's, Microsoft, Ford Credit</p><p>07:15 How Ford Credit fits with Ford Motor</p><p>10:15 The teams behind every Ford Credit page</p><p>15:15 The vehicle selector test that lost on purpose</p><p>19:15 Why feature placement beats feature ideas</p><p>21:15 Personalization and the shrinking-audience problem</p><p>25:15 Telling the story when a test loses</p><p>30:45 Connecting an online test to an offline car sale</p><p>33:55 The experimentation piggy bank</p><p><strong>Takeaways</strong></p><p>1. Losing tests often create more value than winners because they stop expensive mistakes before they ship.</p><p>2. Measure experimentation two ways: the revenue you earn from wins and the revenue you save by killing bad experiences.</p><p>3. A feature that fails early in a flow can succeed later; placement and timing often matter more than the idea itself.</p><p>4. Connecting online experiments to offline outcomes like receivables turns a small lift into a number leadership cares about.</p><p>5. When you struggle to land a result, lead with the story of what the customer did, then bring the numbers.</p><p><br></p><p><br><strong>Connect with the Guest<br></strong>LinkedIn: <a href="https://www.linkedin.com/in/geoffrey-bell-62a03617/">https://www.linkedin.com/in/geoffrey-bell-62a03617/</a> <br>Website: <a href="https://www.ford.com/finance/">https://www.ford.com/finance/</a></p><p><br><strong>Sponsor</strong><br>Growthbook helps you ship features with confidence by bringing experimentation and feature flagging into one open-source platform. No more guessing whether that new checkout flow actually moved the needle, waiting weeks for data team bandwidth, or flying blind on rollouts.</p><p>Growthbook gives you a single place to run A/B tests, manage feature flags, and analyze results against your existing data warehouse.</p><p>With powerful stats built in, it takes the complexity out of experimentation, helps you catch regressions before they hit every user, and makes it easy to test ideas that keep your product improving and your metrics moving in the right direction.</p><p><br>See a demo at <a href="https://www.growthbook.io/">https://www.growthbook.io/</a> </p>
<ul><li>(00:00) - Intro</li>
<li>(01:15) - Geoffrey's path: Lowe's, Microsoft, Ford Credit</li>
<li>(07:15) - How Ford Credit fits with Ford Motor</li>
<li>(10:15) - The teams behind every Ford Credit page</li>
<li>(15:15) - The vehicle selector test that lost on purpose</li>
<li>(19:15) - Why feature placement beats feature ideas</li>
<li>(21:15) - Personalization and the shrinking-audience problem</li>
<li>(25:15) - Telling the story when a test loses</li>
<li>(30:45) - Connecting an online test to an offline car sale</li>
<li>(33:55) - The experimentation piggy bank</li>
</ul>]]>
      </itunes:summary>
      <itunes:keywords>a/b testing,experimentation,feature flags,ford credit,geoffrey bell,prequalification testing,vehicle selector test,experimentation piggy bank,offline conversion,receivables,personalization testing,sample size,close rate,adobe target,experimentation roi,how to prove experimentation roi,why ab tests fail,online to offline attribution,experimentation culture,contentsquare</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/c3607cb8/transcript.txt" type="text/plain"/>
      <podcast:chapters url="https://share.transistor.fm/s/c3607cb8/chapters.json" type="application/json+chapters"/>
    </item>
    <item>
      <title>Atlassian on the talent product turnaround from A/B testing</title>
      <itunes:episode>13</itunes:episode>
      <podcast:episode>13</podcast:episode>
      <itunes:title>Atlassian on the talent product turnaround from A/B testing</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">17d64360-385f-46c5-bad0-c9e1d1568b2f</guid>
      <link>https://share.transistor.fm/s/69fb860f</link>
      <description>
        <![CDATA[<p>This episode of The Experimentation Edge explores how A/B testing, feature flags, and user research transformed Atlassian's talent product after it failed with its first users. Andrew Willingham — 11 years at Amazon, now Head of Legal and People Products at Atlassian — shares how product experimentation works when you can't test at scale, why your customer and your user are not the same person, and how the metrics you choose decide which experiments you can even run.</p><p><strong>Summary</strong><br>Andrew Willingham, Head of Legal and People Products at Atlassian, spent 11 years at Amazon before joining Atlassian a year ago. His path from running A/B tests on millions of Amazon shoppers to building talent management software for a few hundred thousand employees forced a fundamental shift: when you can't run tests at scale, you have to sit with your actual users and watch them fail. He shares how building a talent review product for Amazon's HR specialists completely flopped when handed to HRBPs — and why that failure taught him more than any winning experiment. Now at Atlassian, he's applying that same rigor to reimagining hiring processes with AI, testing everything from recruiter screens to interview sequences that the industry has run the same way for decades.</p><p><strong>Timestamps</strong><br>03:09 From marketing Amazon's mobile app to building HR software for 1.5 million associates  <br>08:19 Why a talent review product loved by IO psych experts flopped with actual HRBPs  <br>11:11 How A/B testing helps product managers escape opinion-based politics  <br>15:25 Testing copy that changes behavior: "We'll generate that status report for you"  <br>17:20 The two North Star metrics Andrew optimizes: efficiency and quality  <br>19:05 Khan Academy's metric trap: measuring cognitive engagement, not just completion  <br>21:10 Why product managers resist experimentation — and what changes when you admit you don't know  </p><p><strong>Takeaways</strong><br>- Your customer and your user may not be the same person — building for HR specialists instead of the HRBPs who actually run talent reviews resulted in a feature nobody could use.  <br>- When you can't test at scale, desk rides replace A/B tests — sitting with users and watching them struggle reveals failures faster than any dashboard.  <br>- Experimentation short-circuits political debates by removing opinion from product decisions.  <br>- Test metrics before you test features — usage time could signal engagement or just mean your product takes too long to do its job.  <br>- The experiments that fail deliver the most valuable learnings, especially when you expected a slam dunk.  </p><p><br><strong>Connect with the guest</strong><br>Andrew Willingham on LinkedIn: <a href="https://www.linkedin.com/in/andrewwillingham/">https://www.linkedin.com/in/andrewwillingham/</a><br>Learn more about Atlassian: <a href="https://www.atlassian.com/">https://www.atlassian.com/</a></p><p><br><strong>Sponsor</strong><br>Growthbook helps you ship features with confidence by bringing experimentation and feature flagging into one open-source platform. No more guessing whether that new checkout flow actually moved the needle, waiting weeks for data team bandwidth, or flying blind on rollouts.</p><p>Growthbook gives you a single place to run A/B tests, manage feature flags, and analyze results against your existing data warehouse.</p><p>With powerful stats built in, it takes the complexity out of experimentation, helps you catch regressions before they hit every user, and makes it easy to test ideas that keep your product improving and your metrics moving in the right direction.</p><p>See a demo at <a href="https://www.growthbook.io/">https://www.growthbook.io/</a></p><p>Topics: A/B testing, product experimentation, feature flags, user research, talent management, qualitative research, metric design, experimentation at scale, growth experimentation.</p>
<ul><li>(03:09) - From marketing Amazon's mobile app to building HR software for 1.5 million associates </li>
<li>(08:19) - Why a talent review product loved by IO psych experts flopped with actual HRBPs </li>
<li>(11:11) - How A/B testing helps product managers escape opinion-based politics </li>
<li>(15:25) - Testing copy that changes behavior: "We'll generate that status report for you" </li>
<li>(17:20) - The two North Star metrics Andrew optimizes: efficiency and quality </li>
<li>(19:05) - Khan Academy's metric trap: measuring cognitive engagement, not just completion </li>
<li>(21:10) - Why product managers resist experimentation — and what changes when you admit you don't know </li>
</ul>]]>
      </description>
      <content:encoded>
        <![CDATA[<p>This episode of The Experimentation Edge explores how A/B testing, feature flags, and user research transformed Atlassian's talent product after it failed with its first users. Andrew Willingham — 11 years at Amazon, now Head of Legal and People Products at Atlassian — shares how product experimentation works when you can't test at scale, why your customer and your user are not the same person, and how the metrics you choose decide which experiments you can even run.</p><p><strong>Summary</strong><br>Andrew Willingham, Head of Legal and People Products at Atlassian, spent 11 years at Amazon before joining Atlassian a year ago. His path from running A/B tests on millions of Amazon shoppers to building talent management software for a few hundred thousand employees forced a fundamental shift: when you can't run tests at scale, you have to sit with your actual users and watch them fail. He shares how building a talent review product for Amazon's HR specialists completely flopped when handed to HRBPs — and why that failure taught him more than any winning experiment. Now at Atlassian, he's applying that same rigor to reimagining hiring processes with AI, testing everything from recruiter screens to interview sequences that the industry has run the same way for decades.</p><p><strong>Timestamps</strong><br>03:09 From marketing Amazon's mobile app to building HR software for 1.5 million associates  <br>08:19 Why a talent review product loved by IO psych experts flopped with actual HRBPs  <br>11:11 How A/B testing helps product managers escape opinion-based politics  <br>15:25 Testing copy that changes behavior: "We'll generate that status report for you"  <br>17:20 The two North Star metrics Andrew optimizes: efficiency and quality  <br>19:05 Khan Academy's metric trap: measuring cognitive engagement, not just completion  <br>21:10 Why product managers resist experimentation — and what changes when you admit you don't know  </p><p><strong>Takeaways</strong><br>- Your customer and your user may not be the same person — building for HR specialists instead of the HRBPs who actually run talent reviews resulted in a feature nobody could use.  <br>- When you can't test at scale, desk rides replace A/B tests — sitting with users and watching them struggle reveals failures faster than any dashboard.  <br>- Experimentation short-circuits political debates by removing opinion from product decisions.  <br>- Test metrics before you test features — usage time could signal engagement or just mean your product takes too long to do its job.  <br>- The experiments that fail deliver the most valuable learnings, especially when you expected a slam dunk.  </p><p><br><strong>Connect with the guest</strong><br>Andrew Willingham on LinkedIn: <a href="https://www.linkedin.com/in/andrewwillingham/">https://www.linkedin.com/in/andrewwillingham/</a><br>Learn more about Atlassian: <a href="https://www.atlassian.com/">https://www.atlassian.com/</a></p><p><br><strong>Sponsor</strong><br>Growthbook helps you ship features with confidence by bringing experimentation and feature flagging into one open-source platform. No more guessing whether that new checkout flow actually moved the needle, waiting weeks for data team bandwidth, or flying blind on rollouts.</p><p>Growthbook gives you a single place to run A/B tests, manage feature flags, and analyze results against your existing data warehouse.</p><p>With powerful stats built in, it takes the complexity out of experimentation, helps you catch regressions before they hit every user, and makes it easy to test ideas that keep your product improving and your metrics moving in the right direction.</p><p>See a demo at <a href="https://www.growthbook.io/">https://www.growthbook.io/</a></p><p>Topics: A/B testing, product experimentation, feature flags, user research, talent management, qualitative research, metric design, experimentation at scale, growth experimentation.</p>
<ul><li>(03:09) - From marketing Amazon's mobile app to building HR software for 1.5 million associates </li>
<li>(08:19) - Why a talent review product loved by IO psych experts flopped with actual HRBPs </li>
<li>(11:11) - How A/B testing helps product managers escape opinion-based politics </li>
<li>(15:25) - Testing copy that changes behavior: "We'll generate that status report for you" </li>
<li>(17:20) - The two North Star metrics Andrew optimizes: efficiency and quality </li>
<li>(19:05) - Khan Academy's metric trap: measuring cognitive engagement, not just completion </li>
<li>(21:10) - Why product managers resist experimentation — and what changes when you admit you don't know </li>
</ul>]]>
      </content:encoded>
      <pubDate>Wed, 13 May 2026 06:00:00 -0600</pubDate>
      <author>Growthbook</author>
      <enclosure url="https://media.transistor.fm/69fb860f/42caa581.mp3" length="21711071" type="audio/mpeg"/>
      <itunes:author>Growthbook</itunes:author>
      <itunes:image href="https://img.transistorcdn.com/ulTUu5Mi7dEBlrV4U_4bRIypJCAythnyBy3_fOR3Uts/rs:fill:0:0:1/w:1400/h:1400/q:60/mb:500000/aHR0cHM6Ly9pbWct/dXBsb2FkLXByb2R1/Y3Rpb24udHJhbnNp/c3Rvci5mbS84MjM1/MTJlYTljNzhmZGQ0/NTkyMjg1MDcyMzI1/ZTk3Yi5wbmc.jpg"/>
      <itunes:duration>1354</itunes:duration>
      <itunes:summary>
        <![CDATA[<p>This episode of The Experimentation Edge explores how A/B testing, feature flags, and user research transformed Atlassian's talent product after it failed with its first users. Andrew Willingham — 11 years at Amazon, now Head of Legal and People Products at Atlassian — shares how product experimentation works when you can't test at scale, why your customer and your user are not the same person, and how the metrics you choose decide which experiments you can even run.</p><p><strong>Summary</strong><br>Andrew Willingham, Head of Legal and People Products at Atlassian, spent 11 years at Amazon before joining Atlassian a year ago. His path from running A/B tests on millions of Amazon shoppers to building talent management software for a few hundred thousand employees forced a fundamental shift: when you can't run tests at scale, you have to sit with your actual users and watch them fail. He shares how building a talent review product for Amazon's HR specialists completely flopped when handed to HRBPs — and why that failure taught him more than any winning experiment. Now at Atlassian, he's applying that same rigor to reimagining hiring processes with AI, testing everything from recruiter screens to interview sequences that the industry has run the same way for decades.</p><p><strong>Timestamps</strong><br>03:09 From marketing Amazon's mobile app to building HR software for 1.5 million associates  <br>08:19 Why a talent review product loved by IO psych experts flopped with actual HRBPs  <br>11:11 How A/B testing helps product managers escape opinion-based politics  <br>15:25 Testing copy that changes behavior: "We'll generate that status report for you"  <br>17:20 The two North Star metrics Andrew optimizes: efficiency and quality  <br>19:05 Khan Academy's metric trap: measuring cognitive engagement, not just completion  <br>21:10 Why product managers resist experimentation — and what changes when you admit you don't know  </p><p><strong>Takeaways</strong><br>- Your customer and your user may not be the same person — building for HR specialists instead of the HRBPs who actually run talent reviews resulted in a feature nobody could use.  <br>- When you can't test at scale, desk rides replace A/B tests — sitting with users and watching them struggle reveals failures faster than any dashboard.  <br>- Experimentation short-circuits political debates by removing opinion from product decisions.  <br>- Test metrics before you test features — usage time could signal engagement or just mean your product takes too long to do its job.  <br>- The experiments that fail deliver the most valuable learnings, especially when you expected a slam dunk.  </p><p><br><strong>Connect with the guest</strong><br>Andrew Willingham on LinkedIn: <a href="https://www.linkedin.com/in/andrewwillingham/">https://www.linkedin.com/in/andrewwillingham/</a><br>Learn more about Atlassian: <a href="https://www.atlassian.com/">https://www.atlassian.com/</a></p><p><br><strong>Sponsor</strong><br>Growthbook helps you ship features with confidence by bringing experimentation and feature flagging into one open-source platform. No more guessing whether that new checkout flow actually moved the needle, waiting weeks for data team bandwidth, or flying blind on rollouts.</p><p>Growthbook gives you a single place to run A/B tests, manage feature flags, and analyze results against your existing data warehouse.</p><p>With powerful stats built in, it takes the complexity out of experimentation, helps you catch regressions before they hit every user, and makes it easy to test ideas that keep your product improving and your metrics moving in the right direction.</p><p>See a demo at <a href="https://www.growthbook.io/">https://www.growthbook.io/</a></p><p>Topics: A/B testing, product experimentation, feature flags, user research, talent management, qualitative research, metric design, experimentation at scale, growth experimentation.</p>
<ul><li>(03:09) - From marketing Amazon's mobile app to building HR software for 1.5 million associates </li>
<li>(08:19) - Why a talent review product loved by IO psych experts flopped with actual HRBPs </li>
<li>(11:11) - How A/B testing helps product managers escape opinion-based politics </li>
<li>(15:25) - Testing copy that changes behavior: "We'll generate that status report for you" </li>
<li>(17:20) - The two North Star metrics Andrew optimizes: efficiency and quality </li>
<li>(19:05) - Khan Academy's metric trap: measuring cognitive engagement, not just completion </li>
<li>(21:10) - Why product managers resist experimentation — and what changes when you admit you don't know </li>
</ul>]]>
      </itunes:summary>
      <itunes:keywords>product management,talent management,HR technology,HRIS,Amazon experimentation culture,A/B testing,behavioral metrics,user research,talent acquisition,AI in HR,hiring process optimization,quality metrics,employee experience,experimentation at scale,product strategy,qualitative research,quantitative research,HR product development,metric design,learning velocity,product experimentation,feature flags,tech experimentation,growth experimentation</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/69fb860f/transcript.txt" type="text/plain"/>
      <podcast:chapters url="https://share.transistor.fm/s/69fb860f/chapters.json" type="application/json+chapters"/>
    </item>
    <item>
      <title>How DoorDash saved millions with one A/B test</title>
      <itunes:episode>12</itunes:episode>
      <podcast:episode>12</podcast:episode>
      <itunes:title>How DoorDash saved millions with one A/B test</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">5e97cd6d-9977-4d41-be92-a1ece2a8c412</guid>
      <link>https://share.transistor.fm/s/59877618</link>
      <description>
        <![CDATA[<p>This episode of The Experimentation Edge unpacks how DoorDash's experimentation platform runs 12,000+ A/B tests per year across 42 million monthly active users — and now powers merchant-led testing on menu pricing and promotions. Ilya Izrailevsky, Senior Engineering Manager leading the platform, explains how feature flags, marketplace experimentation, and CEO-level experiment reviews built a multi-million-dollar experimentation culture across consumers, dashers, and merchants.</p><p><strong>Summary</strong><br>Most companies struggle to scale experimentation beyond engineering teams. DoorDash runs over 12,000 experiments per year across 42 million monthly active users — and now they're enabling restaurant owners to run their own A/B tests on menu pricing and promotions. Ilya Izrailevsky, Senior Engineering Manager leading DoorDash's experimentation platform, shares how the company built a three-sided marketplace testing program that balances consumers, dashers, and merchants across 40+ countries. From his time scaling search at Amazon (where offline model evaluation narrowed hundreds of candidates down to 10 for live testing) to preventing DashPass churn at DoorDash, Ilya reveals what happens when experimentation scales beyond product teams — and why CEO-level experiment review emails drive cultural change faster than any training program.</p><p>One standout learning: expanding delivery radius to 11+ miles increased grocery orders but tanked retail conversions. The lesson wasn't about distance — it was that one metric approach breaks in multi-dimensional marketplaces. DoorDash now segments experimentation by vertical, behavior pattern, and regional market, using AI agents to mine institutional knowledge from past tests and auto-generate experiment summaries that ship company-wide within hours of readout.</p><p><br><strong>Timestamps</strong><br>00:40 From building Wasabi (Intuit's open-source platform) to running ML at Amazon and Uber  <br>03:04 Why product velocity without experimentation creates feature bloat, not impact  <br>05:32 Scaling search at Amazon: billions of products, 10 visible results, 25% win rate  <br>08:22 Offline evaluation as a filter — golden data sets cut model candidates before live traffic  <br>10:23 DoorDash's three-sided marketplace: 300 million feature flag evaluations per second  <br>12:38 CEO Tony Xu reads every experiment email and replies with alternative hypotheses  <br>13:33 Democratization at scale: enabling merchants to A/B test menu pricing and promotions  <br>17:05 DashPass churn experiment uncovered value perception gap — became a full product area  <br>22:03 Why expanding delivery radius killed retail orders but boosted grocery conversions  <br>24:16 No single North Star metric — balancing consumer quality, dasher earnings, merchant mix  <br>27:29 Four-dimensional scale: democratization, global expansion, new verticals, AI agents  <br>31:03 Agentic experimentation: AI mines past tests to generate hypotheses and debug imbalance</p><p><br><strong>Takeaways</strong><br>- Win rate matters less than learnings per test — DoorDash ships company-wide experiment summaries (win or lose) that the CEO actively reads and responds to, creating cultural accountability around testing rigor.<br>- Offline evaluation acts as a pre-filter for model velocity — Amazon's search team used golden data sets to cut hundreds of ML candidates down to 10 for live A/B testing, preventing wasted experiment slots.<br>- One-size metrics break in multi-dimensional marketplaces — DoorDash balances consumer retention, dasher utilization, and merchant inventory mix across verticals because optimizing one side degrades the ecosystem.<br>- Democratization requires opinionated templates, not open-ended tools — enabling non-technical users to run tests means embedding success metrics and guardrails into pre-built experiment configs.<br>- AI scales institutional knowledge, not just analysis speed — mining past experiment readouts to auto-generate new hypotheses turns your testing history into a compounding advantage.</p><p><br><strong>Connect with the guest</strong><br>LinkedIn: <a href="https://www.linkedin.com/in/ilyaizrailevsky/">https://www.linkedin.com/in/ilyaizrailevsky/</a><br>Learn more about DoorDash: <a href="https://www.doordash.com/">https://www.doordash.com/</a></p><p><strong>Sponsor</strong><br>Growthbook helps you ship features with confidence by bringing experimentation and feature flagging into one open-source platform. No more guessing whether that new checkout flow actually moved the needle, waiting weeks for data team bandwidth, or flying blind on rollouts.</p><p>Growthbook gives you a single place to run A/B tests, manage feature flags, and analyze results against your existing data warehouse.</p><p>With powerful stats built in, it takes the complexity out of experimentation, helps you catch regressions before they hit every user, and makes it easy to test ideas that keep your product improving and your metrics moving in the right direction.</p><p>See a demo at <a href="https://www.growthbook.io/">https://www.growthbook.io/</a></p><p>Topics: A/B testing, experimentation platform, feature flags, marketplace experimentation, machine learning, growth experimentation, statistical significance, experimentation culture, agentic AI workflows.</p>
<ul><li>(00:40) - From building Wasabi (Intuit's open-source platform) to running ML at Amazon and Uber </li>
<li>(03:04) - Why product velocity without experimentation creates feature bloat, not impact </li>
<li>(05:32) - Scaling search at Amazon: billions of products, 10 visible results, 25% win rate </li>
<li>(08:22) - Offline evaluation as a filter — golden data sets cut model candidates before live traffic </li>
<li>(10:23) - DoorDash's three-sided marketplace: 300 million feature flag evaluations per second </li>
<li>(12:38) - CEO Tony Xu reads every experiment email and replies with alternative hypotheses </li>
<li>(13:33) - Democratization at scale: enabling merchants to A/B test menu pricing and promotions </li>
<li>(17:05) - DashPass churn experiment uncovered value perception gap — became a full product area </li>
<li>(22:03) - Why expanding delivery radius killed retail orders but boosted grocery conversions </li>
<li>(24:16) - No single North Star metric — balancing consumer quality, dasher earnings, merchant mix </li>
<li>(27:29) - Four-dimensional scale: democratization, global expansion, new verticals, AI agents </li>
<li>(31:03) - Agentic experimentation: AI mines past tests to generate hypotheses and debug imbalance</li>
</ul>]]>
      </description>
      <content:encoded>
        <![CDATA[<p>This episode of The Experimentation Edge unpacks how DoorDash's experimentation platform runs 12,000+ A/B tests per year across 42 million monthly active users — and now powers merchant-led testing on menu pricing and promotions. Ilya Izrailevsky, Senior Engineering Manager leading the platform, explains how feature flags, marketplace experimentation, and CEO-level experiment reviews built a multi-million-dollar experimentation culture across consumers, dashers, and merchants.</p><p><strong>Summary</strong><br>Most companies struggle to scale experimentation beyond engineering teams. DoorDash runs over 12,000 experiments per year across 42 million monthly active users — and now they're enabling restaurant owners to run their own A/B tests on menu pricing and promotions. Ilya Izrailevsky, Senior Engineering Manager leading DoorDash's experimentation platform, shares how the company built a three-sided marketplace testing program that balances consumers, dashers, and merchants across 40+ countries. From his time scaling search at Amazon (where offline model evaluation narrowed hundreds of candidates down to 10 for live testing) to preventing DashPass churn at DoorDash, Ilya reveals what happens when experimentation scales beyond product teams — and why CEO-level experiment review emails drive cultural change faster than any training program.</p><p>One standout learning: expanding delivery radius to 11+ miles increased grocery orders but tanked retail conversions. The lesson wasn't about distance — it was that one metric approach breaks in multi-dimensional marketplaces. DoorDash now segments experimentation by vertical, behavior pattern, and regional market, using AI agents to mine institutional knowledge from past tests and auto-generate experiment summaries that ship company-wide within hours of readout.</p><p><br><strong>Timestamps</strong><br>00:40 From building Wasabi (Intuit's open-source platform) to running ML at Amazon and Uber  <br>03:04 Why product velocity without experimentation creates feature bloat, not impact  <br>05:32 Scaling search at Amazon: billions of products, 10 visible results, 25% win rate  <br>08:22 Offline evaluation as a filter — golden data sets cut model candidates before live traffic  <br>10:23 DoorDash's three-sided marketplace: 300 million feature flag evaluations per second  <br>12:38 CEO Tony Xu reads every experiment email and replies with alternative hypotheses  <br>13:33 Democratization at scale: enabling merchants to A/B test menu pricing and promotions  <br>17:05 DashPass churn experiment uncovered value perception gap — became a full product area  <br>22:03 Why expanding delivery radius killed retail orders but boosted grocery conversions  <br>24:16 No single North Star metric — balancing consumer quality, dasher earnings, merchant mix  <br>27:29 Four-dimensional scale: democratization, global expansion, new verticals, AI agents  <br>31:03 Agentic experimentation: AI mines past tests to generate hypotheses and debug imbalance</p><p><br><strong>Takeaways</strong><br>- Win rate matters less than learnings per test — DoorDash ships company-wide experiment summaries (win or lose) that the CEO actively reads and responds to, creating cultural accountability around testing rigor.<br>- Offline evaluation acts as a pre-filter for model velocity — Amazon's search team used golden data sets to cut hundreds of ML candidates down to 10 for live A/B testing, preventing wasted experiment slots.<br>- One-size metrics break in multi-dimensional marketplaces — DoorDash balances consumer retention, dasher utilization, and merchant inventory mix across verticals because optimizing one side degrades the ecosystem.<br>- Democratization requires opinionated templates, not open-ended tools — enabling non-technical users to run tests means embedding success metrics and guardrails into pre-built experiment configs.<br>- AI scales institutional knowledge, not just analysis speed — mining past experiment readouts to auto-generate new hypotheses turns your testing history into a compounding advantage.</p><p><br><strong>Connect with the guest</strong><br>LinkedIn: <a href="https://www.linkedin.com/in/ilyaizrailevsky/">https://www.linkedin.com/in/ilyaizrailevsky/</a><br>Learn more about DoorDash: <a href="https://www.doordash.com/">https://www.doordash.com/</a></p><p><strong>Sponsor</strong><br>Growthbook helps you ship features with confidence by bringing experimentation and feature flagging into one open-source platform. No more guessing whether that new checkout flow actually moved the needle, waiting weeks for data team bandwidth, or flying blind on rollouts.</p><p>Growthbook gives you a single place to run A/B tests, manage feature flags, and analyze results against your existing data warehouse.</p><p>With powerful stats built in, it takes the complexity out of experimentation, helps you catch regressions before they hit every user, and makes it easy to test ideas that keep your product improving and your metrics moving in the right direction.</p><p>See a demo at <a href="https://www.growthbook.io/">https://www.growthbook.io/</a></p><p>Topics: A/B testing, experimentation platform, feature flags, marketplace experimentation, machine learning, growth experimentation, statistical significance, experimentation culture, agentic AI workflows.</p>
<ul><li>(00:40) - From building Wasabi (Intuit's open-source platform) to running ML at Amazon and Uber </li>
<li>(03:04) - Why product velocity without experimentation creates feature bloat, not impact </li>
<li>(05:32) - Scaling search at Amazon: billions of products, 10 visible results, 25% win rate </li>
<li>(08:22) - Offline evaluation as a filter — golden data sets cut model candidates before live traffic </li>
<li>(10:23) - DoorDash's three-sided marketplace: 300 million feature flag evaluations per second </li>
<li>(12:38) - CEO Tony Xu reads every experiment email and replies with alternative hypotheses </li>
<li>(13:33) - Democratization at scale: enabling merchants to A/B test menu pricing and promotions </li>
<li>(17:05) - DashPass churn experiment uncovered value perception gap — became a full product area </li>
<li>(22:03) - Why expanding delivery radius killed retail orders but boosted grocery conversions </li>
<li>(24:16) - No single North Star metric — balancing consumer quality, dasher earnings, merchant mix </li>
<li>(27:29) - Four-dimensional scale: democratization, global expansion, new verticals, AI agents </li>
<li>(31:03) - Agentic experimentation: AI mines past tests to generate hypotheses and debug imbalance</li>
</ul>]]>
      </content:encoded>
      <pubDate>Tue, 12 May 2026 06:00:00 -0600</pubDate>
      <author>Growthbook</author>
      <enclosure url="https://media.transistor.fm/59877618/b95fea3b.mp3" length="30055189" type="audio/mpeg"/>
      <itunes:author>Growthbook</itunes:author>
      <itunes:image href="https://img.transistorcdn.com/1UaQNgXc09Z0lmewBY91yncvftGJVPKt1B8h5WRrR2I/rs:fill:0:0:1/w:1400/h:1400/q:60/mb:500000/aHR0cHM6Ly9pbWct/dXBsb2FkLXByb2R1/Y3Rpb24udHJhbnNp/c3Rvci5mbS9hMTM4/OTRmOWRjOThjMzY5/ZjRjN2YwYzYxNzk0/MTFkZS5wbmc.jpg"/>
      <itunes:duration>1875</itunes:duration>
      <itunes:summary>
        <![CDATA[<p>This episode of The Experimentation Edge unpacks how DoorDash's experimentation platform runs 12,000+ A/B tests per year across 42 million monthly active users — and now powers merchant-led testing on menu pricing and promotions. Ilya Izrailevsky, Senior Engineering Manager leading the platform, explains how feature flags, marketplace experimentation, and CEO-level experiment reviews built a multi-million-dollar experimentation culture across consumers, dashers, and merchants.</p><p><strong>Summary</strong><br>Most companies struggle to scale experimentation beyond engineering teams. DoorDash runs over 12,000 experiments per year across 42 million monthly active users — and now they're enabling restaurant owners to run their own A/B tests on menu pricing and promotions. Ilya Izrailevsky, Senior Engineering Manager leading DoorDash's experimentation platform, shares how the company built a three-sided marketplace testing program that balances consumers, dashers, and merchants across 40+ countries. From his time scaling search at Amazon (where offline model evaluation narrowed hundreds of candidates down to 10 for live testing) to preventing DashPass churn at DoorDash, Ilya reveals what happens when experimentation scales beyond product teams — and why CEO-level experiment review emails drive cultural change faster than any training program.</p><p>One standout learning: expanding delivery radius to 11+ miles increased grocery orders but tanked retail conversions. The lesson wasn't about distance — it was that one metric approach breaks in multi-dimensional marketplaces. DoorDash now segments experimentation by vertical, behavior pattern, and regional market, using AI agents to mine institutional knowledge from past tests and auto-generate experiment summaries that ship company-wide within hours of readout.</p><p><br><strong>Timestamps</strong><br>00:40 From building Wasabi (Intuit's open-source platform) to running ML at Amazon and Uber  <br>03:04 Why product velocity without experimentation creates feature bloat, not impact  <br>05:32 Scaling search at Amazon: billions of products, 10 visible results, 25% win rate  <br>08:22 Offline evaluation as a filter — golden data sets cut model candidates before live traffic  <br>10:23 DoorDash's three-sided marketplace: 300 million feature flag evaluations per second  <br>12:38 CEO Tony Xu reads every experiment email and replies with alternative hypotheses  <br>13:33 Democratization at scale: enabling merchants to A/B test menu pricing and promotions  <br>17:05 DashPass churn experiment uncovered value perception gap — became a full product area  <br>22:03 Why expanding delivery radius killed retail orders but boosted grocery conversions  <br>24:16 No single North Star metric — balancing consumer quality, dasher earnings, merchant mix  <br>27:29 Four-dimensional scale: democratization, global expansion, new verticals, AI agents  <br>31:03 Agentic experimentation: AI mines past tests to generate hypotheses and debug imbalance</p><p><br><strong>Takeaways</strong><br>- Win rate matters less than learnings per test — DoorDash ships company-wide experiment summaries (win or lose) that the CEO actively reads and responds to, creating cultural accountability around testing rigor.<br>- Offline evaluation acts as a pre-filter for model velocity — Amazon's search team used golden data sets to cut hundreds of ML candidates down to 10 for live A/B testing, preventing wasted experiment slots.<br>- One-size metrics break in multi-dimensional marketplaces — DoorDash balances consumer retention, dasher utilization, and merchant inventory mix across verticals because optimizing one side degrades the ecosystem.<br>- Democratization requires opinionated templates, not open-ended tools — enabling non-technical users to run tests means embedding success metrics and guardrails into pre-built experiment configs.<br>- AI scales institutional knowledge, not just analysis speed — mining past experiment readouts to auto-generate new hypotheses turns your testing history into a compounding advantage.</p><p><br><strong>Connect with the guest</strong><br>LinkedIn: <a href="https://www.linkedin.com/in/ilyaizrailevsky/">https://www.linkedin.com/in/ilyaizrailevsky/</a><br>Learn more about DoorDash: <a href="https://www.doordash.com/">https://www.doordash.com/</a></p><p><strong>Sponsor</strong><br>Growthbook helps you ship features with confidence by bringing experimentation and feature flagging into one open-source platform. No more guessing whether that new checkout flow actually moved the needle, waiting weeks for data team bandwidth, or flying blind on rollouts.</p><p>Growthbook gives you a single place to run A/B tests, manage feature flags, and analyze results against your existing data warehouse.</p><p>With powerful stats built in, it takes the complexity out of experimentation, helps you catch regressions before they hit every user, and makes it easy to test ideas that keep your product improving and your metrics moving in the right direction.</p><p>See a demo at <a href="https://www.growthbook.io/">https://www.growthbook.io/</a></p><p>Topics: A/B testing, experimentation platform, feature flags, marketplace experimentation, machine learning, growth experimentation, statistical significance, experimentation culture, agentic AI workflows.</p>
<ul><li>(00:40) - From building Wasabi (Intuit's open-source platform) to running ML at Amazon and Uber </li>
<li>(03:04) - Why product velocity without experimentation creates feature bloat, not impact </li>
<li>(05:32) - Scaling search at Amazon: billions of products, 10 visible results, 25% win rate </li>
<li>(08:22) - Offline evaluation as a filter — golden data sets cut model candidates before live traffic </li>
<li>(10:23) - DoorDash's three-sided marketplace: 300 million feature flag evaluations per second </li>
<li>(12:38) - CEO Tony Xu reads every experiment email and replies with alternative hypotheses </li>
<li>(13:33) - Democratization at scale: enabling merchants to A/B test menu pricing and promotions </li>
<li>(17:05) - DashPass churn experiment uncovered value perception gap — became a full product area </li>
<li>(22:03) - Why expanding delivery radius killed retail orders but boosted grocery conversions </li>
<li>(24:16) - No single North Star metric — balancing consumer quality, dasher earnings, merchant mix </li>
<li>(27:29) - Four-dimensional scale: democratization, global expansion, new verticals, AI agents </li>
<li>(31:03) - Agentic experimentation: AI mines past tests to generate hypotheses and debug imbalance</li>
</ul>]]>
      </itunes:summary>
      <itunes:keywords>a/b testing,experimentation,machine learning optimization,search ranking,multi-dimensional optimization,win rate,guardrail metrics,offline experimentation,doordash experimentation,experimentation program,democratization of experimentation,experiment lifecycle,ai-powered experimentation,marketplace experimentation,consumer behavior testing,dashpass retention,experimentation culture,ceo experiment reviews,experiment templates,agentic ai workflows,tech experimentation,growth experimentation,product experimentation,feature flags</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/59877618/transcript.txt" type="text/plain"/>
      <podcast:chapters url="https://share.transistor.fm/s/59877618/chapters.json" type="application/json+chapters"/>
    </item>
    <item>
      <title>How UPS generated half a billion from 80+ apps with A/B testing</title>
      <itunes:episode>11</itunes:episode>
      <podcast:episode>11</podcast:episode>
      <itunes:title>How UPS generated half a billion from 80+ apps with A/B testing</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">37d8b34d-0487-430a-8acb-1bf1d854b497</guid>
      <link>https://share.transistor.fm/s/7b0fed20</link>
      <description>
        <![CDATA[<p>This episode of The Experimentation Edge shows how UPS's A/B testing program drove $500M+ in incremental revenue across 80+ customer-facing applications. Dave Massey — head of the J.E.D.I. team (Journey Experience and Design Innovation) — walks through the first test that proved UX could move revenue, how he defended counterintuitive results to skeptical execs, and how a small experimentation team can override opinion with data at enterprise scale.</p><p><strong>Summary</strong><br>Dave Massey walked into UPS in 2016 and immediately got pulled into a meeting about AB testing tools. By the end of the day, he owned the platform—and the problem: UPS hadn't run a single meaningful experiment. Three years later, senior leadership gave him a hard number to hit. Prove UX could move revenue, or the pilot dies. His first test—removing navigation from the checkout flow—delivered $35 million in incremental revenue. Senior leaders didn't believe it. They made him defend the results upside down and sideways. When the dust settled, the data held. Today, Massey's team has driven over half a billion dollars in incremental revenue by treating UPS.com like the e-commerce business it actually is.</p><p>Massey's approach is simple: test everything, especially what senior leaders think will work. His team, Journey Experience and Design Innovation (nicknamed J.E.D.I.), has built a reputation for saying no with data, not opinion. When a business unit demanded required recipient emails to capture customer data, J.E.D.I. ran the test in 24 hours and killed it. Conversion tanked. Two years later, the international team asked for the same feature—but framed it as a customs solution. That test passed. Same feature, different reason, different outcome. That's the edge Massey's team delivers: rigorous hypothesis design, a UX research team embedded in the experimentation workflow, and zero tolerance for untested ideas.</p><p><br><strong>Timestamps<br></strong>03:09 Dave's first day at UPS: inheriting an AB testing tool with no program  <br>05:59 Senior leadership's ultimatum: prove UX ROI or kill the pilot  <br>08:38 First test result: $35M from removing navigation in checkout  <br>09:48 Defending the numbers: how Massey's team survived scrutiny  <br>11:07 Why a data-driven engineering culture made experimentation inevitable  <br>16:12 Team size: 80 people supporting almost 80 customer-facing applications  <br>19:08 The 24-hour test: when required email fields killed conversion  <br>22:28 Why Massey embeds UX research inside the experimentation team  <br>24:41 AI at UPS: treating it as a tool, not a replacement  </p><p><br><strong>Takeaways<br></strong>- Massey's first test removed navigation from UPS's shipping checkout flow and delivered $35 million in incremental revenue—proving e-commerce best practices apply even when customers think "this is just a tool, not e-commerce."  <br>- J.E.D.I.'s win rate stays high because UX research and experimentation teams operate under the same leader, giving the program both behavioral metrics and voice-of-customer insight before tests ever launch.  <br>- When senior leaders push ideas, Massey's team tests them instead of arguing—then delivers results that either validate the idea or identify three better alternatives the data actually supports.  <br>- The same feature (required recipient email) failed for customer data capture but passed for international customs—proof that framing and customer benefit matter more than the feature itself.  <br>- UPS runs everything centrally now, but the real win is that demand for testing has decentralized—business units across the company now come to J.E.D.I. asking to test their ideas.</p><p><br><strong>Connect with the guest</strong><br>LinkedIn: <a href="https://www.linkedin.com/in/masseycreates/">https://www.linkedin.com/in/masseycreates/</a><a href="https://www.linkedin.com/in/davemassey"> </a><br>Learn more about UPS: <a href="https://www.ups.com">https://www.ups.com</a></p><p><br><strong>Sponsor</strong><br>Growthbook helps you ship features with confidence by bringing experimentation and feature flagging into one open-source platform. No more guessing whether that new checkout flow actually moved the needle, waiting weeks for data team bandwidth, or flying blind on rollouts.</p><p>Growthbook gives you a single place to run A/B tests, manage feature flags, and analyze results against your existing data warehouse.</p><p>With powerful stats built in, it takes the complexity out of experimentation, helps you catch regressions before they hit every user, and makes it easy to test ideas that keep your product improving and your metrics moving in the right direction.</p><p>See a demo at <a href="https://www.growthbook.io/">https://www.growthbook.io/</a></p><p>Topics: A/B testing, experimentation, conversion rate optimization, feature flags, UX research, e-commerce experimentation, statistical significance, experimentation team building, growth experimentation, sequential testing.</p>
<ul><li>(03:09) - Dave's first day at UPS: inheriting an AB testing tool with no program </li>
<li>(05:59) - Senior leadership's ultimatum: prove UX ROI or kill the pilot </li>
<li>(08:38) - First test result: $35M from removing navigation in checkout </li>
<li>(09:48) - Defending the numbers: how Massey's team survived scrutiny </li>
<li>(11:07) - Why a data-driven engineering culture made experimentation inevitable </li>
<li>(16:12) - Team size: 80 people supporting almost 80 customer-facing applications </li>
<li>(19:08) - The 24-hour test: when required email fields killed conversion </li>
<li>(22:28) - Why Massey embeds UX research inside the experimentation team </li>
<li>(24:41) - AI at UPS: treating it as a tool, not a replacement </li>
</ul>]]>
      </description>
      <content:encoded>
        <![CDATA[<p>This episode of The Experimentation Edge shows how UPS's A/B testing program drove $500M+ in incremental revenue across 80+ customer-facing applications. Dave Massey — head of the J.E.D.I. team (Journey Experience and Design Innovation) — walks through the first test that proved UX could move revenue, how he defended counterintuitive results to skeptical execs, and how a small experimentation team can override opinion with data at enterprise scale.</p><p><strong>Summary</strong><br>Dave Massey walked into UPS in 2016 and immediately got pulled into a meeting about AB testing tools. By the end of the day, he owned the platform—and the problem: UPS hadn't run a single meaningful experiment. Three years later, senior leadership gave him a hard number to hit. Prove UX could move revenue, or the pilot dies. His first test—removing navigation from the checkout flow—delivered $35 million in incremental revenue. Senior leaders didn't believe it. They made him defend the results upside down and sideways. When the dust settled, the data held. Today, Massey's team has driven over half a billion dollars in incremental revenue by treating UPS.com like the e-commerce business it actually is.</p><p>Massey's approach is simple: test everything, especially what senior leaders think will work. His team, Journey Experience and Design Innovation (nicknamed J.E.D.I.), has built a reputation for saying no with data, not opinion. When a business unit demanded required recipient emails to capture customer data, J.E.D.I. ran the test in 24 hours and killed it. Conversion tanked. Two years later, the international team asked for the same feature—but framed it as a customs solution. That test passed. Same feature, different reason, different outcome. That's the edge Massey's team delivers: rigorous hypothesis design, a UX research team embedded in the experimentation workflow, and zero tolerance for untested ideas.</p><p><br><strong>Timestamps<br></strong>03:09 Dave's first day at UPS: inheriting an AB testing tool with no program  <br>05:59 Senior leadership's ultimatum: prove UX ROI or kill the pilot  <br>08:38 First test result: $35M from removing navigation in checkout  <br>09:48 Defending the numbers: how Massey's team survived scrutiny  <br>11:07 Why a data-driven engineering culture made experimentation inevitable  <br>16:12 Team size: 80 people supporting almost 80 customer-facing applications  <br>19:08 The 24-hour test: when required email fields killed conversion  <br>22:28 Why Massey embeds UX research inside the experimentation team  <br>24:41 AI at UPS: treating it as a tool, not a replacement  </p><p><br><strong>Takeaways<br></strong>- Massey's first test removed navigation from UPS's shipping checkout flow and delivered $35 million in incremental revenue—proving e-commerce best practices apply even when customers think "this is just a tool, not e-commerce."  <br>- J.E.D.I.'s win rate stays high because UX research and experimentation teams operate under the same leader, giving the program both behavioral metrics and voice-of-customer insight before tests ever launch.  <br>- When senior leaders push ideas, Massey's team tests them instead of arguing—then delivers results that either validate the idea or identify three better alternatives the data actually supports.  <br>- The same feature (required recipient email) failed for customer data capture but passed for international customs—proof that framing and customer benefit matter more than the feature itself.  <br>- UPS runs everything centrally now, but the real win is that demand for testing has decentralized—business units across the company now come to J.E.D.I. asking to test their ideas.</p><p><br><strong>Connect with the guest</strong><br>LinkedIn: <a href="https://www.linkedin.com/in/masseycreates/">https://www.linkedin.com/in/masseycreates/</a><a href="https://www.linkedin.com/in/davemassey"> </a><br>Learn more about UPS: <a href="https://www.ups.com">https://www.ups.com</a></p><p><br><strong>Sponsor</strong><br>Growthbook helps you ship features with confidence by bringing experimentation and feature flagging into one open-source platform. No more guessing whether that new checkout flow actually moved the needle, waiting weeks for data team bandwidth, or flying blind on rollouts.</p><p>Growthbook gives you a single place to run A/B tests, manage feature flags, and analyze results against your existing data warehouse.</p><p>With powerful stats built in, it takes the complexity out of experimentation, helps you catch regressions before they hit every user, and makes it easy to test ideas that keep your product improving and your metrics moving in the right direction.</p><p>See a demo at <a href="https://www.growthbook.io/">https://www.growthbook.io/</a></p><p>Topics: A/B testing, experimentation, conversion rate optimization, feature flags, UX research, e-commerce experimentation, statistical significance, experimentation team building, growth experimentation, sequential testing.</p>
<ul><li>(03:09) - Dave's first day at UPS: inheriting an AB testing tool with no program </li>
<li>(05:59) - Senior leadership's ultimatum: prove UX ROI or kill the pilot </li>
<li>(08:38) - First test result: $35M from removing navigation in checkout </li>
<li>(09:48) - Defending the numbers: how Massey's team survived scrutiny </li>
<li>(11:07) - Why a data-driven engineering culture made experimentation inevitable </li>
<li>(16:12) - Team size: 80 people supporting almost 80 customer-facing applications </li>
<li>(19:08) - The 24-hour test: when required email fields killed conversion </li>
<li>(22:28) - Why Massey embeds UX research inside the experimentation team </li>
<li>(24:41) - AI at UPS: treating it as a tool, not a replacement </li>
</ul>]]>
      </content:encoded>
      <pubDate>Mon, 11 May 2026 06:00:00 -0600</pubDate>
      <author>Growthbook</author>
      <enclosure url="https://media.transistor.fm/7b0fed20/74587a02.mp3" length="22283046" type="audio/mpeg"/>
      <itunes:author>Growthbook</itunes:author>
      <itunes:image href="https://img.transistorcdn.com/xemB3fCyovr4nhTHoLEGRLZ11ibzgkeYl5DHp_5Gvrg/rs:fill:0:0:1/w:1400/h:1400/q:60/mb:500000/aHR0cHM6Ly9pbWct/dXBsb2FkLXByb2R1/Y3Rpb24udHJhbnNp/c3Rvci5mbS80NmUz/YmY3MDI0OGJiZGIy/NzRiNDI5MjZkOWQ5/NzY2MS5wbmc.jpg"/>
      <itunes:duration>1390</itunes:duration>
      <itunes:summary>
        <![CDATA[<p>This episode of The Experimentation Edge shows how UPS's A/B testing program drove $500M+ in incremental revenue across 80+ customer-facing applications. Dave Massey — head of the J.E.D.I. team (Journey Experience and Design Innovation) — walks through the first test that proved UX could move revenue, how he defended counterintuitive results to skeptical execs, and how a small experimentation team can override opinion with data at enterprise scale.</p><p><strong>Summary</strong><br>Dave Massey walked into UPS in 2016 and immediately got pulled into a meeting about AB testing tools. By the end of the day, he owned the platform—and the problem: UPS hadn't run a single meaningful experiment. Three years later, senior leadership gave him a hard number to hit. Prove UX could move revenue, or the pilot dies. His first test—removing navigation from the checkout flow—delivered $35 million in incremental revenue. Senior leaders didn't believe it. They made him defend the results upside down and sideways. When the dust settled, the data held. Today, Massey's team has driven over half a billion dollars in incremental revenue by treating UPS.com like the e-commerce business it actually is.</p><p>Massey's approach is simple: test everything, especially what senior leaders think will work. His team, Journey Experience and Design Innovation (nicknamed J.E.D.I.), has built a reputation for saying no with data, not opinion. When a business unit demanded required recipient emails to capture customer data, J.E.D.I. ran the test in 24 hours and killed it. Conversion tanked. Two years later, the international team asked for the same feature—but framed it as a customs solution. That test passed. Same feature, different reason, different outcome. That's the edge Massey's team delivers: rigorous hypothesis design, a UX research team embedded in the experimentation workflow, and zero tolerance for untested ideas.</p><p><br><strong>Timestamps<br></strong>03:09 Dave's first day at UPS: inheriting an AB testing tool with no program  <br>05:59 Senior leadership's ultimatum: prove UX ROI or kill the pilot  <br>08:38 First test result: $35M from removing navigation in checkout  <br>09:48 Defending the numbers: how Massey's team survived scrutiny  <br>11:07 Why a data-driven engineering culture made experimentation inevitable  <br>16:12 Team size: 80 people supporting almost 80 customer-facing applications  <br>19:08 The 24-hour test: when required email fields killed conversion  <br>22:28 Why Massey embeds UX research inside the experimentation team  <br>24:41 AI at UPS: treating it as a tool, not a replacement  </p><p><br><strong>Takeaways<br></strong>- Massey's first test removed navigation from UPS's shipping checkout flow and delivered $35 million in incremental revenue—proving e-commerce best practices apply even when customers think "this is just a tool, not e-commerce."  <br>- J.E.D.I.'s win rate stays high because UX research and experimentation teams operate under the same leader, giving the program both behavioral metrics and voice-of-customer insight before tests ever launch.  <br>- When senior leaders push ideas, Massey's team tests them instead of arguing—then delivers results that either validate the idea or identify three better alternatives the data actually supports.  <br>- The same feature (required recipient email) failed for customer data capture but passed for international customs—proof that framing and customer benefit matter more than the feature itself.  <br>- UPS runs everything centrally now, but the real win is that demand for testing has decentralized—business units across the company now come to J.E.D.I. asking to test their ideas.</p><p><br><strong>Connect with the guest</strong><br>LinkedIn: <a href="https://www.linkedin.com/in/masseycreates/">https://www.linkedin.com/in/masseycreates/</a><a href="https://www.linkedin.com/in/davemassey"> </a><br>Learn more about UPS: <a href="https://www.ups.com">https://www.ups.com</a></p><p><br><strong>Sponsor</strong><br>Growthbook helps you ship features with confidence by bringing experimentation and feature flagging into one open-source platform. No more guessing whether that new checkout flow actually moved the needle, waiting weeks for data team bandwidth, or flying blind on rollouts.</p><p>Growthbook gives you a single place to run A/B tests, manage feature flags, and analyze results against your existing data warehouse.</p><p>With powerful stats built in, it takes the complexity out of experimentation, helps you catch regressions before they hit every user, and makes it easy to test ideas that keep your product improving and your metrics moving in the right direction.</p><p>See a demo at <a href="https://www.growthbook.io/">https://www.growthbook.io/</a></p><p>Topics: A/B testing, experimentation, conversion rate optimization, feature flags, UX research, e-commerce experimentation, statistical significance, experimentation team building, growth experimentation, sequential testing.</p>
<ul><li>(03:09) - Dave's first day at UPS: inheriting an AB testing tool with no program </li>
<li>(05:59) - Senior leadership's ultimatum: prove UX ROI or kill the pilot </li>
<li>(08:38) - First test result: $35M from removing navigation in checkout </li>
<li>(09:48) - Defending the numbers: how Massey's team survived scrutiny </li>
<li>(11:07) - Why a data-driven engineering culture made experimentation inevitable </li>
<li>(16:12) - Team size: 80 people supporting almost 80 customer-facing applications </li>
<li>(19:08) - The 24-hour test: when required email fields killed conversion </li>
<li>(22:28) - Why Massey embeds UX research inside the experimentation team </li>
<li>(24:41) - AI at UPS: treating it as a tool, not a replacement </li>
</ul>]]>
      </itunes:summary>
      <itunes:keywords>experimentation,a/b testing,conversion rate optimization,statistical significance,testing velocity,user research,personalization,ecommerce optimization,checkout flow,feature flagging,experimentation roi,win rate,product experimentation,data-driven decision making,behavioral metrics,customer experience testing,experimentation team building,sequential testing,ux research,guardrail metrics,tech experimentation,growth experimentation,feature flags</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/7b0fed20/transcript.txt" type="text/plain"/>
      <podcast:chapters url="https://share.transistor.fm/s/7b0fed20/chapters.json" type="application/json+chapters"/>
    </item>
    <item>
      <title>How Experimentation Led to Annual Growth at Fanatics</title>
      <itunes:episode>10</itunes:episode>
      <podcast:episode>10</podcast:episode>
      <itunes:title>How Experimentation Led to Annual Growth at Fanatics</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">c35052ae-96e1-4626-9a22-3b184fad5709</guid>
      <link>https://share.transistor.fm/s/6a9adab0</link>
      <description>
        <![CDATA[<p><strong>Summary</strong></p><p>Most e-commerce companies test a handful of features each month. Fanatics runs nearly 100 experiments monthly and delivers a big portion of the company's total annual growth through experimentation alone. Medha Umarji, VP of Growth and Experimentation at the multi-billion dollar sports merchandising retailer, explains how she built a program that scales from 10 tests per month to 100—and maintains enough rigor to spot false positives before they become costly decisions.</p><p>The difference isn't tooling or headcount. It's culture. When your CEO reads Excel spreadsheets for fun and actively wants data to prove him wrong, you stop debating whether to test and start debating how to test smarter. Medha shares the frameworks Fanatics uses to balance speed with rigor: a "do no harm" track for brand plays that won't show up in conversion metrics, a small-sample framework for teams that can't hit statistical significance thresholds, and an experimentation Wiki that feeds a continuous iteration flywheel. One surprising test on ad removal initially showed 95% statistical significance—until they replicated it and found the result was a false positive. The lesson: even at scale, you need to double-click on causality.<br></p><p><strong>Timestamps</strong></p><p>03:09 How Fanatics scaled from 10 to 100 experiments per month over 10 years</p><p>05:25 Why some leadership teams embrace experimentation and others resist it</p><p>07:06 How experimentation consistently delivers a big portion of Fanatics' annual growth</p><p>08:20 What happens when your CEO consumes Excel spreadsheets and questions everything</p><p>10:35 How top-down humility shapes an entire company's testing culture</p><p>12:10 The ad removal test that looked like a 95% win—then failed replication</p><p>15:55 How Fanatics built an experimentation Wiki that powers their growth engine</p><p>22:45 The "do no harm" framework for features that don't measure cleanly in A/B tests</p><p>25:20 Why lowering barriers to adoption matters more than statistical perfection early on</p><p>26:27 Your odds of winning at experimentation are worse than roulette<br></p><p><strong>Takeaways</strong></p><ul><li>Replication catches false positives: A 95% confidence level still means 1 in 20 results are noise—if a critical test outcome can't be explained through micro-metrics, run it again before committing resources.</li><li>Top-down buy-in shifts the conversation from "why test?" to "how do we test?": When leadership treats data as the tiebreaker, teams stop defending opinions and start building better experiments.</li><li>Frameworks like "do no harm" and "small sample" expand who can test: Not every initiative needs 30,000 orders to ship value—lower the barrier for teams that can't hit statistical thresholds while protecting core KPIs.</li><li>Documenting experiments in a centralized Wiki creates a growth flywheel: Fanatics' Wiki feeds their roadmap with iterations on already-built features, reducing tech dependency and accelerating velocity.</li><li>Micro-metrics establish causality beyond top-line KPIs: If revenue moves but scroll depth, cart adds, and product views don't follow the same pattern, question the result before declaring a win.</li></ul><p><br><strong>Connect with the guest</strong></p><p>LinkedIn: <a href="https://www.linkedin.com/in/medhaumarji/">https://www.linkedin.com/in/medhaumarji/</a></p><p><strong>Learn more about Fanatics</strong></p><p><a href="https://www.fanatics.com/">https://www.fanatics.com/</a></p>]]>
      </description>
      <content:encoded>
        <![CDATA[<p><strong>Summary</strong></p><p>Most e-commerce companies test a handful of features each month. Fanatics runs nearly 100 experiments monthly and delivers a big portion of the company's total annual growth through experimentation alone. Medha Umarji, VP of Growth and Experimentation at the multi-billion dollar sports merchandising retailer, explains how she built a program that scales from 10 tests per month to 100—and maintains enough rigor to spot false positives before they become costly decisions.</p><p>The difference isn't tooling or headcount. It's culture. When your CEO reads Excel spreadsheets for fun and actively wants data to prove him wrong, you stop debating whether to test and start debating how to test smarter. Medha shares the frameworks Fanatics uses to balance speed with rigor: a "do no harm" track for brand plays that won't show up in conversion metrics, a small-sample framework for teams that can't hit statistical significance thresholds, and an experimentation Wiki that feeds a continuous iteration flywheel. One surprising test on ad removal initially showed 95% statistical significance—until they replicated it and found the result was a false positive. The lesson: even at scale, you need to double-click on causality.<br></p><p><strong>Timestamps</strong></p><p>03:09 How Fanatics scaled from 10 to 100 experiments per month over 10 years</p><p>05:25 Why some leadership teams embrace experimentation and others resist it</p><p>07:06 How experimentation consistently delivers a big portion of Fanatics' annual growth</p><p>08:20 What happens when your CEO consumes Excel spreadsheets and questions everything</p><p>10:35 How top-down humility shapes an entire company's testing culture</p><p>12:10 The ad removal test that looked like a 95% win—then failed replication</p><p>15:55 How Fanatics built an experimentation Wiki that powers their growth engine</p><p>22:45 The "do no harm" framework for features that don't measure cleanly in A/B tests</p><p>25:20 Why lowering barriers to adoption matters more than statistical perfection early on</p><p>26:27 Your odds of winning at experimentation are worse than roulette<br></p><p><strong>Takeaways</strong></p><ul><li>Replication catches false positives: A 95% confidence level still means 1 in 20 results are noise—if a critical test outcome can't be explained through micro-metrics, run it again before committing resources.</li><li>Top-down buy-in shifts the conversation from "why test?" to "how do we test?": When leadership treats data as the tiebreaker, teams stop defending opinions and start building better experiments.</li><li>Frameworks like "do no harm" and "small sample" expand who can test: Not every initiative needs 30,000 orders to ship value—lower the barrier for teams that can't hit statistical thresholds while protecting core KPIs.</li><li>Documenting experiments in a centralized Wiki creates a growth flywheel: Fanatics' Wiki feeds their roadmap with iterations on already-built features, reducing tech dependency and accelerating velocity.</li><li>Micro-metrics establish causality beyond top-line KPIs: If revenue moves but scroll depth, cart adds, and product views don't follow the same pattern, question the result before declaring a win.</li></ul><p><br><strong>Connect with the guest</strong></p><p>LinkedIn: <a href="https://www.linkedin.com/in/medhaumarji/">https://www.linkedin.com/in/medhaumarji/</a></p><p><strong>Learn more about Fanatics</strong></p><p><a href="https://www.fanatics.com/">https://www.fanatics.com/</a></p>]]>
      </content:encoded>
      <pubDate>Thu, 07 May 2026 06:00:00 -0600</pubDate>
      <author>Growthbook</author>
      <enclosure url="https://media.transistor.fm/6a9adab0/d3357f35.mp3" length="27477331" type="audio/mpeg"/>
      <itunes:author>Growthbook</itunes:author>
      <itunes:image href="https://img.transistorcdn.com/jfZB9oGxWX42tXoQqAYsXxxg3rGq3z4zTvsBb0C1IDk/rs:fill:0:0:1/w:1400/h:1400/q:60/mb:500000/aHR0cHM6Ly9pbWct/dXBsb2FkLXByb2R1/Y3Rpb24udHJhbnNp/c3Rvci5mbS8zYmMy/OTY4MGM0YWFhYTE0/NDBlNzQwMjZhODUw/YTJhZC5qcGc.jpg"/>
      <itunes:duration>1718</itunes:duration>
      <itunes:summary>How Experimentation Leads to Annual Growth Every Year at Fanatics</itunes:summary>
      <itunes:subtitle>How Experimentation Leads to Annual Growth Every Year at Fanatics</itunes:subtitle>
      <itunes:keywords>experimentation program, A/B testing, e-commerce optimization, testing velocity, data-driven decision making, CEO buy-in, experimentation culture, do no harm framework, non-inferiority testing, experimentation wiki, false positive rate, 95% statistical significance, replicating test results, micro-metrics analysis, guardrail metrics, small sample testing, experimentation roadmap, feature iteration, VP Growth, retail experimentation</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/6a9adab0/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Inside Chess.com's Plan to Run 1,000 Experiments in a Single Year</title>
      <itunes:episode>9</itunes:episode>
      <podcast:episode>9</podcast:episode>
      <itunes:title>Inside Chess.com's Plan to Run 1,000 Experiments in a Single Year</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">a951862b-60e8-486a-a8e2-ec0e438b7a10</guid>
      <link>https://share.transistor.fm/s/5ebd1b8f</link>
      <description>
        <![CDATA[<p><strong>Summary</strong></p><p><br>Chess.com ran its first A/B test in 2023. Two years later, the team is on track to run 1,000 experiments in a single year—and they've already shipped 195 in Q1. </p><p><br>In this episode, Ashley Stirrup sits down with Nafis Shaikh, Director of Product Management at Chess.com, to get inside the experimentation engine powering one of the world's most beloved gaming products. </p><p><br>Nafis brings experience from Zynga and Prodigy and a refreshingly honest take on what changes when a product built on passion suddenly has to serve a 10-million-DAU user base that spans absolute beginners to rated FIDE players. He and Ashley get into why one-size-fits-all doesn't actually fit anyone, how to measure an AI coach when you can't tell whether users have their volume on, and a game review experiment that completely upended the team's assumptions about how players want to learn. </p><p><br>Nafis also shares practical advice for product managers trying to introduce experimentation culture to organizations that have never done it before—starting with a simple pre/post test rather than a fancy platform. If you lead product, care about experimentation maturity, or just want to hear how a classic product is scaling its learning loop, this one's worth your time.</p><p><strong><br>Timestamps</strong></p><ul><li>[00:35] – Chess.com's experimentation origin story and the 1,000-test goal</li><li>[05:01] – Designing for a user base that spans beginners to FIDE-rated players</li><li>[07:30] – The four metrics dimensions Nafis uses to evaluate tests</li><li>[12:03] – How do you A/B test an AI coach when you can't tell who's listening?</li><li>[15:49] – Embracing humility and the shift away from "we know what works"</li><li>[20:50] – The game review test that surprised everyone: 80% of users review wins</li><li>[24:06] – Advice for PMs introducing experimentation at a new company</li><li>[29:15] – The onboarding debate and personalization from session zero<p></p></li></ul><p><strong><br>Takeaways</strong></p><ul><li>Scale test volume to learning speed, not just shipping speed</li><li>Build hypotheses around user psychology, not just KPI movement</li><li>Accept that being wrong is the point—experimentation only works when leadership embraces humility</li><li>Start simple if you're new to experimentation; a clean pre/post comparison beats a fancy platform you don't use</li><li>Reposition features around how users actually feel, not how you assume they should feel</li><li>Design onboarding around the shortest path to value, not the longest path to personalization<p></p></li></ul><p><strong><br>Guest LinkedIn:</strong> <a href="https://www.linkedin.com/in/nafis-shaikh-20161916/">https://www.linkedin.com/in/nafis-shaikh-20161916/</a></p><p><strong>Company website:</strong> <a href="https://www.chess.com">https://www.chess.com</a></p>]]>
      </description>
      <content:encoded>
        <![CDATA[<p><strong>Summary</strong></p><p><br>Chess.com ran its first A/B test in 2023. Two years later, the team is on track to run 1,000 experiments in a single year—and they've already shipped 195 in Q1. </p><p><br>In this episode, Ashley Stirrup sits down with Nafis Shaikh, Director of Product Management at Chess.com, to get inside the experimentation engine powering one of the world's most beloved gaming products. </p><p><br>Nafis brings experience from Zynga and Prodigy and a refreshingly honest take on what changes when a product built on passion suddenly has to serve a 10-million-DAU user base that spans absolute beginners to rated FIDE players. He and Ashley get into why one-size-fits-all doesn't actually fit anyone, how to measure an AI coach when you can't tell whether users have their volume on, and a game review experiment that completely upended the team's assumptions about how players want to learn. </p><p><br>Nafis also shares practical advice for product managers trying to introduce experimentation culture to organizations that have never done it before—starting with a simple pre/post test rather than a fancy platform. If you lead product, care about experimentation maturity, or just want to hear how a classic product is scaling its learning loop, this one's worth your time.</p><p><strong><br>Timestamps</strong></p><ul><li>[00:35] – Chess.com's experimentation origin story and the 1,000-test goal</li><li>[05:01] – Designing for a user base that spans beginners to FIDE-rated players</li><li>[07:30] – The four metrics dimensions Nafis uses to evaluate tests</li><li>[12:03] – How do you A/B test an AI coach when you can't tell who's listening?</li><li>[15:49] – Embracing humility and the shift away from "we know what works"</li><li>[20:50] – The game review test that surprised everyone: 80% of users review wins</li><li>[24:06] – Advice for PMs introducing experimentation at a new company</li><li>[29:15] – The onboarding debate and personalization from session zero<p></p></li></ul><p><strong><br>Takeaways</strong></p><ul><li>Scale test volume to learning speed, not just shipping speed</li><li>Build hypotheses around user psychology, not just KPI movement</li><li>Accept that being wrong is the point—experimentation only works when leadership embraces humility</li><li>Start simple if you're new to experimentation; a clean pre/post comparison beats a fancy platform you don't use</li><li>Reposition features around how users actually feel, not how you assume they should feel</li><li>Design onboarding around the shortest path to value, not the longest path to personalization<p></p></li></ul><p><strong><br>Guest LinkedIn:</strong> <a href="https://www.linkedin.com/in/nafis-shaikh-20161916/">https://www.linkedin.com/in/nafis-shaikh-20161916/</a></p><p><strong>Company website:</strong> <a href="https://www.chess.com">https://www.chess.com</a></p>]]>
      </content:encoded>
      <pubDate>Tue, 21 Apr 2026 06:00:00 -0600</pubDate>
      <author>Growthbook</author>
      <enclosure url="https://media.transistor.fm/5ebd1b8f/8114a915.mp3" length="30552351" type="audio/mpeg"/>
      <itunes:author>Growthbook</itunes:author>
      <itunes:image href="https://img.transistorcdn.com/F27wmxVbKyoD-z8yQHePT9M_yQrlpMz2HNM1aTxCIg8/rs:fill:0:0:1/w:1400/h:1400/q:60/mb:500000/aHR0cHM6Ly9pbWct/dXBsb2FkLXByb2R1/Y3Rpb24udHJhbnNp/c3Rvci5mbS83ZDlm/ODRjZWJhZThkOGZi/MjhlMmRmM2U3NGFh/NjQ0My5wbmc.jpg"/>
      <itunes:duration>1907</itunes:duration>
      <itunes:summary>
        <![CDATA[<p><strong>Summary</strong></p><p><br>Chess.com ran its first A/B test in 2023. Two years later, the team is on track to run 1,000 experiments in a single year—and they've already shipped 195 in Q1. </p><p><br>In this episode, Ashley Stirrup sits down with Nafis Shaikh, Director of Product Management at Chess.com, to get inside the experimentation engine powering one of the world's most beloved gaming products. </p><p><br>Nafis brings experience from Zynga and Prodigy and a refreshingly honest take on what changes when a product built on passion suddenly has to serve a 10-million-DAU user base that spans absolute beginners to rated FIDE players. He and Ashley get into why one-size-fits-all doesn't actually fit anyone, how to measure an AI coach when you can't tell whether users have their volume on, and a game review experiment that completely upended the team's assumptions about how players want to learn. </p><p><br>Nafis also shares practical advice for product managers trying to introduce experimentation culture to organizations that have never done it before—starting with a simple pre/post test rather than a fancy platform. If you lead product, care about experimentation maturity, or just want to hear how a classic product is scaling its learning loop, this one's worth your time.</p><p><strong><br>Timestamps</strong></p><ul><li>[00:35] – Chess.com's experimentation origin story and the 1,000-test goal</li><li>[05:01] – Designing for a user base that spans beginners to FIDE-rated players</li><li>[07:30] – The four metrics dimensions Nafis uses to evaluate tests</li><li>[12:03] – How do you A/B test an AI coach when you can't tell who's listening?</li><li>[15:49] – Embracing humility and the shift away from "we know what works"</li><li>[20:50] – The game review test that surprised everyone: 80% of users review wins</li><li>[24:06] – Advice for PMs introducing experimentation at a new company</li><li>[29:15] – The onboarding debate and personalization from session zero<p></p></li></ul><p><strong><br>Takeaways</strong></p><ul><li>Scale test volume to learning speed, not just shipping speed</li><li>Build hypotheses around user psychology, not just KPI movement</li><li>Accept that being wrong is the point—experimentation only works when leadership embraces humility</li><li>Start simple if you're new to experimentation; a clean pre/post comparison beats a fancy platform you don't use</li><li>Reposition features around how users actually feel, not how you assume they should feel</li><li>Design onboarding around the shortest path to value, not the longest path to personalization<p></p></li></ul><p><strong><br>Guest LinkedIn:</strong> <a href="https://www.linkedin.com/in/nafis-shaikh-20161916/">https://www.linkedin.com/in/nafis-shaikh-20161916/</a></p><p><strong>Company website:</strong> <a href="https://www.chess.com">https://www.chess.com</a></p>]]>
      </itunes:summary>
      <itunes:keywords>product management, product experimentation, A/B testing strategy, experimentation culture, product management metrics, AI coaching product, gaming product management, user segmentation testing, onboarding experimentation, feature flag testing, Zynga product management, Prodigy Education, experimentation maturity, hypothesis-driven product development, personalization strategy, freemium conversion testing, game review feature, product KPIs, experimentation scaling</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/5ebd1b8f/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Ancestry on AI Storytelling, Paywalls, and Decision Quality</title>
      <itunes:episode>8</itunes:episode>
      <podcast:episode>8</podcast:episode>
      <itunes:title>Ancestry on AI Storytelling, Paywalls, and Decision Quality</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">b0ad794d-dddd-4707-8db3-72272dcf6600</guid>
      <link>https://share.transistor.fm/s/87484f97</link>
      <description>
        <![CDATA[<p><strong>Summary</strong></p><p>What happens when A/B testing stops being a tool and becomes your operating system? Suresh Teckchandani, VP of Product &amp; Technology at Ancestry (formerly PayPal and eBay), shares how the team scaled experimentation from isolated tests to capability-building that drives roadmap and revenue. He details the “growth metering” and paywall experiments that unlocked a 5.3% lift in key engagement and improved conversions—then became platform features. Suresh explains Ancestry’s centralized experimentation platform with self-serve access for PMs and engineers, why “obvious” UX changes can be the riskiest, and how removing friction actually hurt engagement by 20–25% due to user mental models. He also breaks down a major growth lever: AI-powered storytelling that turns raw records into narratives, delivering 30%+ CTR lift and a 5x increase in story views. You’ll learn how Ancestry balances input vs. output metrics, when not to test, and why the best leaders optimize for decision quality over win counts—with clean baselines, right audiences, adequate sample sizes, and true statistical significance.</p><p><strong>Timestamps</strong></p><p>[00:45] – Ancestry’s experimentation maturity: metering, paywalls, and the 5.3% lift that unlocked capabilities</p><p>[03:48] – From isolated tests to a capability mindset: experimentation as an operating system</p><p>[05:34] – Balancing wins with learning: zooming out for subscription engagement and NPS</p><p>[08:49] – Operating model: centralized platform, self-serve dashboards, and baseline resets</p><p>[11:59] – Counterintuitive UX lesson: removing friction backfired (–20–25% CTR); respect mental models</p><p>[15:10] – AI storytelling as a growth lever: record comparisons into narratives, 30%+ CTR and 5x views</p><p>[18:27] – Input vs. output metrics: when to roll back and how to link short- and long-term outcomes</p><p>[30:00] – Parting advice: test what changes CX, avoid vanity testing, and optimize decision quality</p><p><strong>Takeaways</strong></p><p>- Build capabilities, not just tests—use experiments to unlock platform features (e.g., metering, paywalls).</p><p>- Democratize experimentation with a centralized platform and self-serve tooling; reset baselines regularly.</p><p>- Test “obvious” UX changes; preserve helpful friction and align with user mental models.</p><p>- Turn data into narratives with AI to deepen engagement and increase discovery.</p><p>- Define input and output metrics; ship only what improves core outcomes (retention, sign-ups), and roll back fast if not.</p><p>- Optimize for decision quality: right audience, sufficient sample sizes, clean baselines, and true statistical significance.</p><p><br><strong>Sponsor</strong><br>Growthbook helps you ship features with confidence by bringing experimentation and feature flagging into one open-source platform. No more guessing whether that new checkout flow actually moved the needle, waiting weeks for data team bandwidth, or flying blind on rollouts.</p><p>Growthbook gives you a single place to run A/B tests, manage feature flags, and analyze results against your existing data warehouse.</p><p>With powerful stats built in, it takes the complexity out of experimentation, helps you catch regressions before they hit every user, and makes it easy to test ideas that keep your product improving and your metrics moving in the right direction.</p><p>See a demo at <a href="https://www.growthbook.io/">https://www.growthbook.io/</a></p><p><br></p>]]>
      </description>
      <content:encoded>
        <![CDATA[<p><strong>Summary</strong></p><p>What happens when A/B testing stops being a tool and becomes your operating system? Suresh Teckchandani, VP of Product &amp; Technology at Ancestry (formerly PayPal and eBay), shares how the team scaled experimentation from isolated tests to capability-building that drives roadmap and revenue. He details the “growth metering” and paywall experiments that unlocked a 5.3% lift in key engagement and improved conversions—then became platform features. Suresh explains Ancestry’s centralized experimentation platform with self-serve access for PMs and engineers, why “obvious” UX changes can be the riskiest, and how removing friction actually hurt engagement by 20–25% due to user mental models. He also breaks down a major growth lever: AI-powered storytelling that turns raw records into narratives, delivering 30%+ CTR lift and a 5x increase in story views. You’ll learn how Ancestry balances input vs. output metrics, when not to test, and why the best leaders optimize for decision quality over win counts—with clean baselines, right audiences, adequate sample sizes, and true statistical significance.</p><p><strong>Timestamps</strong></p><p>[00:45] – Ancestry’s experimentation maturity: metering, paywalls, and the 5.3% lift that unlocked capabilities</p><p>[03:48] – From isolated tests to a capability mindset: experimentation as an operating system</p><p>[05:34] – Balancing wins with learning: zooming out for subscription engagement and NPS</p><p>[08:49] – Operating model: centralized platform, self-serve dashboards, and baseline resets</p><p>[11:59] – Counterintuitive UX lesson: removing friction backfired (–20–25% CTR); respect mental models</p><p>[15:10] – AI storytelling as a growth lever: record comparisons into narratives, 30%+ CTR and 5x views</p><p>[18:27] – Input vs. output metrics: when to roll back and how to link short- and long-term outcomes</p><p>[30:00] – Parting advice: test what changes CX, avoid vanity testing, and optimize decision quality</p><p><strong>Takeaways</strong></p><p>- Build capabilities, not just tests—use experiments to unlock platform features (e.g., metering, paywalls).</p><p>- Democratize experimentation with a centralized platform and self-serve tooling; reset baselines regularly.</p><p>- Test “obvious” UX changes; preserve helpful friction and align with user mental models.</p><p>- Turn data into narratives with AI to deepen engagement and increase discovery.</p><p>- Define input and output metrics; ship only what improves core outcomes (retention, sign-ups), and roll back fast if not.</p><p>- Optimize for decision quality: right audience, sufficient sample sizes, clean baselines, and true statistical significance.</p><p><br><strong>Sponsor</strong><br>Growthbook helps you ship features with confidence by bringing experimentation and feature flagging into one open-source platform. No more guessing whether that new checkout flow actually moved the needle, waiting weeks for data team bandwidth, or flying blind on rollouts.</p><p>Growthbook gives you a single place to run A/B tests, manage feature flags, and analyze results against your existing data warehouse.</p><p>With powerful stats built in, it takes the complexity out of experimentation, helps you catch regressions before they hit every user, and makes it easy to test ideas that keep your product improving and your metrics moving in the right direction.</p><p>See a demo at <a href="https://www.growthbook.io/">https://www.growthbook.io/</a></p><p><br></p>]]>
      </content:encoded>
      <pubDate>Tue, 14 Apr 2026 06:00:00 -0600</pubDate>
      <author>The Experimentation Edge</author>
      <enclosure url="https://media.transistor.fm/87484f97/bfb4b339.mp3" length="24231135" type="audio/mpeg"/>
      <itunes:author>The Experimentation Edge</itunes:author>
      <itunes:image href="https://img.transistorcdn.com/kL0uERD2NwT_lilK3KeGtET2Tv2Ft2rdYxHDXICof9I/rs:fill:0:0:1/w:1400/h:1400/q:60/mb:500000/aHR0cHM6Ly9pbWct/dXBsb2FkLXByb2R1/Y3Rpb24udHJhbnNp/c3Rvci5mbS9jMWNk/ODI3M2M1ZTlhMzc1/Njc5Y2UyOGUzMzlj/YTdlMy5qcGc.jpg"/>
      <itunes:duration>1511</itunes:duration>
      <itunes:summary>
        <![CDATA[<p><strong>Summary</strong></p><p>What happens when A/B testing stops being a tool and becomes your operating system? Suresh Teckchandani, VP of Product &amp; Technology at Ancestry (formerly PayPal and eBay), shares how the team scaled experimentation from isolated tests to capability-building that drives roadmap and revenue. He details the “growth metering” and paywall experiments that unlocked a 5.3% lift in key engagement and improved conversions—then became platform features. Suresh explains Ancestry’s centralized experimentation platform with self-serve access for PMs and engineers, why “obvious” UX changes can be the riskiest, and how removing friction actually hurt engagement by 20–25% due to user mental models. He also breaks down a major growth lever: AI-powered storytelling that turns raw records into narratives, delivering 30%+ CTR lift and a 5x increase in story views. You’ll learn how Ancestry balances input vs. output metrics, when not to test, and why the best leaders optimize for decision quality over win counts—with clean baselines, right audiences, adequate sample sizes, and true statistical significance.</p><p><strong>Timestamps</strong></p><p>[00:45] – Ancestry’s experimentation maturity: metering, paywalls, and the 5.3% lift that unlocked capabilities</p><p>[03:48] – From isolated tests to a capability mindset: experimentation as an operating system</p><p>[05:34] – Balancing wins with learning: zooming out for subscription engagement and NPS</p><p>[08:49] – Operating model: centralized platform, self-serve dashboards, and baseline resets</p><p>[11:59] – Counterintuitive UX lesson: removing friction backfired (–20–25% CTR); respect mental models</p><p>[15:10] – AI storytelling as a growth lever: record comparisons into narratives, 30%+ CTR and 5x views</p><p>[18:27] – Input vs. output metrics: when to roll back and how to link short- and long-term outcomes</p><p>[30:00] – Parting advice: test what changes CX, avoid vanity testing, and optimize decision quality</p><p><strong>Takeaways</strong></p><p>- Build capabilities, not just tests—use experiments to unlock platform features (e.g., metering, paywalls).</p><p>- Democratize experimentation with a centralized platform and self-serve tooling; reset baselines regularly.</p><p>- Test “obvious” UX changes; preserve helpful friction and align with user mental models.</p><p>- Turn data into narratives with AI to deepen engagement and increase discovery.</p><p>- Define input and output metrics; ship only what improves core outcomes (retention, sign-ups), and roll back fast if not.</p><p>- Optimize for decision quality: right audience, sufficient sample sizes, clean baselines, and true statistical significance.</p><p><br><strong>Sponsor</strong><br>Growthbook helps you ship features with confidence by bringing experimentation and feature flagging into one open-source platform. No more guessing whether that new checkout flow actually moved the needle, waiting weeks for data team bandwidth, or flying blind on rollouts.</p><p>Growthbook gives you a single place to run A/B tests, manage feature flags, and analyze results against your existing data warehouse.</p><p>With powerful stats built in, it takes the complexity out of experimentation, helps you catch regressions before they hit every user, and makes it easy to test ideas that keep your product improving and your metrics moving in the right direction.</p><p>See a demo at <a href="https://www.growthbook.io/">https://www.growthbook.io/</a></p><p><br></p>]]>
      </itunes:summary>
      <itunes:keywords>A/B testing,experimentation,product management,Ancestry,experimentation at scale,product experimentation,AI storytelling,user engagement,product metrics,feature testing,experimentation culture,VP product,subscription growth,conversion optimization,UX testing,data-driven product,test and learn,paywall optimization,product innovation,experimentation platform</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/87484f97/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Fyxer's engineering playbook to go from $1M to $35M ARR</title>
      <itunes:episode>7</itunes:episode>
      <podcast:episode>7</podcast:episode>
      <itunes:title>Fyxer's engineering playbook to go from $1M to $35M ARR</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">02648e43-ae26-4719-8179-d2a7f0c463dc</guid>
      <link>https://share.transistor.fm/s/cf8beed6</link>
      <description>
        <![CDATA[<p>Summary</p><p>How do you drive hypergrowth without guessing? Kameron Tanseli, Head of Growth Engineering at Fyxer—an AI assistant for your email—breaks down the experimentation playbook that helped the company scale from $1M to $35M ARR, with sights set on $100–$150M. Kameron explains how startups should think about A/B testing differently: de-risk big bets, not just button colors. He shares a risk-based approach to when to run rigorous tests vs. ship-and-measure, why a 25% win rate is a sign you’re testing ambitiously, and how PLG features should be shipped first, then rapidly iterated to drive usage. You’ll hear how Fyxer uses AI to speed the entire lifecycle—Claude, Cursor desktop cloud agents, GrowthBook, and BigQuery—plus how a Slack-first changelog and an internal “AI data scientist” democratize insights. Kameron also details turning everyday product usage into growth loops, personalizing signup paths, and measuring success by movement in global ARR, not just local metrics. He closes with candid advice for new growth engineers: expect to struggle early, be T-shaped, and adopt your customer’s language.</p><p><br></p><p>Timestamps</p><p>[00:34] – Startup A/B testing mindset: de-risking big bets with only a 25% win rate</p><p>[02:45] – When to A/B test vs. ship: risk appetite, funnel stage, and non-inferiority tests</p><p>[04:43] – 360 experiments with 4 people: scaling to 1,000 using AI and Cursor cloud agents</p><p>[08:22] – Separating feature impact from momentum: PLG and trial model moves ARR</p><p>[10:29] – Ship PLG features, then iterate to drive usage; measuring DAU and revenue impact</p><p>[11:40] – Habit loops to growth loops: turning product features into PLG (scheduling case study)</p><p>[16:47] – Building an experimentation culture: founder buy-in, Slack changelog, shared data</p><p>[26:50] – The modern growth stack: Claude, Cursor, GrowthBook, BigQuery, and DOT in Slack</p><p><br></p><p>Takeaways</p><p>- Prioritize by risk: run rigorous A/B tests where you have volume; use before/after or non-inferiority for low-risk in-product changes.</p><p>- Test big levers—not just UI: pricing models, usage limits, onboarding pathways—and judge success by ARR movement, not micro-metrics.</p><p>- Ship first, then optimize: launch PLG features and immediately run experiments to increase adoption; track daily active usage per feature.</p><p>- Build growth loops from habits: design shareable artifacts and personalized signup paths; drive users back to your domain to capture value.</p><p>- Scale experimentation with AI: use Cursor desktop/cloud agents for parallel builds and visual QA; orchestrate docs/analysis via Claude; automate cleanups and reporting.</p><p>- Make experimentation company-wide: centralize data (BigQuery), broadcast wins/losses in Slack via GrowthBook, and auto-correlate metric dips to releases.</p><p><br></p>]]>
      </description>
      <content:encoded>
        <![CDATA[<p>Summary</p><p>How do you drive hypergrowth without guessing? Kameron Tanseli, Head of Growth Engineering at Fyxer—an AI assistant for your email—breaks down the experimentation playbook that helped the company scale from $1M to $35M ARR, with sights set on $100–$150M. Kameron explains how startups should think about A/B testing differently: de-risk big bets, not just button colors. He shares a risk-based approach to when to run rigorous tests vs. ship-and-measure, why a 25% win rate is a sign you’re testing ambitiously, and how PLG features should be shipped first, then rapidly iterated to drive usage. You’ll hear how Fyxer uses AI to speed the entire lifecycle—Claude, Cursor desktop cloud agents, GrowthBook, and BigQuery—plus how a Slack-first changelog and an internal “AI data scientist” democratize insights. Kameron also details turning everyday product usage into growth loops, personalizing signup paths, and measuring success by movement in global ARR, not just local metrics. He closes with candid advice for new growth engineers: expect to struggle early, be T-shaped, and adopt your customer’s language.</p><p><br></p><p>Timestamps</p><p>[00:34] – Startup A/B testing mindset: de-risking big bets with only a 25% win rate</p><p>[02:45] – When to A/B test vs. ship: risk appetite, funnel stage, and non-inferiority tests</p><p>[04:43] – 360 experiments with 4 people: scaling to 1,000 using AI and Cursor cloud agents</p><p>[08:22] – Separating feature impact from momentum: PLG and trial model moves ARR</p><p>[10:29] – Ship PLG features, then iterate to drive usage; measuring DAU and revenue impact</p><p>[11:40] – Habit loops to growth loops: turning product features into PLG (scheduling case study)</p><p>[16:47] – Building an experimentation culture: founder buy-in, Slack changelog, shared data</p><p>[26:50] – The modern growth stack: Claude, Cursor, GrowthBook, BigQuery, and DOT in Slack</p><p><br></p><p>Takeaways</p><p>- Prioritize by risk: run rigorous A/B tests where you have volume; use before/after or non-inferiority for low-risk in-product changes.</p><p>- Test big levers—not just UI: pricing models, usage limits, onboarding pathways—and judge success by ARR movement, not micro-metrics.</p><p>- Ship first, then optimize: launch PLG features and immediately run experiments to increase adoption; track daily active usage per feature.</p><p>- Build growth loops from habits: design shareable artifacts and personalized signup paths; drive users back to your domain to capture value.</p><p>- Scale experimentation with AI: use Cursor desktop/cloud agents for parallel builds and visual QA; orchestrate docs/analysis via Claude; automate cleanups and reporting.</p><p>- Make experimentation company-wide: centralize data (BigQuery), broadcast wins/losses in Slack via GrowthBook, and auto-correlate metric dips to releases.</p><p><br></p>]]>
      </content:encoded>
      <pubDate>Thu, 02 Apr 2026 13:34:09 -0600</pubDate>
      <author>Growthbook</author>
      <enclosure url="https://media.transistor.fm/cf8beed6/44b21d50.mp3" length="29579244" type="audio/mpeg"/>
      <itunes:author>Growthbook</itunes:author>
      <itunes:image href="https://img.transistorcdn.com/Tsql0to9rV_1FkNzCAmK6JCuGEAMVv8TDXZh3__7_is/rs:fill:0:0:1/w:1400/h:1400/q:60/mb:500000/aHR0cHM6Ly9pbWct/dXBsb2FkLXByb2R1/Y3Rpb24udHJhbnNp/c3Rvci5mbS9lNGM3/YjQyOGM4MDg1Mjc1/YmE2Y2I3NThjNjkz/M2UzYS5qcGc.jpg"/>
      <itunes:duration>1847</itunes:duration>
      <itunes:summary>
        <![CDATA[<p>Summary</p><p>How do you drive hypergrowth without guessing? Kameron Tanseli, Head of Growth Engineering at Fyxer—an AI assistant for your email—breaks down the experimentation playbook that helped the company scale from $1M to $35M ARR, with sights set on $100–$150M. Kameron explains how startups should think about A/B testing differently: de-risk big bets, not just button colors. He shares a risk-based approach to when to run rigorous tests vs. ship-and-measure, why a 25% win rate is a sign you’re testing ambitiously, and how PLG features should be shipped first, then rapidly iterated to drive usage. You’ll hear how Fyxer uses AI to speed the entire lifecycle—Claude, Cursor desktop cloud agents, GrowthBook, and BigQuery—plus how a Slack-first changelog and an internal “AI data scientist” democratize insights. Kameron also details turning everyday product usage into growth loops, personalizing signup paths, and measuring success by movement in global ARR, not just local metrics. He closes with candid advice for new growth engineers: expect to struggle early, be T-shaped, and adopt your customer’s language.</p><p><br></p><p>Timestamps</p><p>[00:34] – Startup A/B testing mindset: de-risking big bets with only a 25% win rate</p><p>[02:45] – When to A/B test vs. ship: risk appetite, funnel stage, and non-inferiority tests</p><p>[04:43] – 360 experiments with 4 people: scaling to 1,000 using AI and Cursor cloud agents</p><p>[08:22] – Separating feature impact from momentum: PLG and trial model moves ARR</p><p>[10:29] – Ship PLG features, then iterate to drive usage; measuring DAU and revenue impact</p><p>[11:40] – Habit loops to growth loops: turning product features into PLG (scheduling case study)</p><p>[16:47] – Building an experimentation culture: founder buy-in, Slack changelog, shared data</p><p>[26:50] – The modern growth stack: Claude, Cursor, GrowthBook, BigQuery, and DOT in Slack</p><p><br></p><p>Takeaways</p><p>- Prioritize by risk: run rigorous A/B tests where you have volume; use before/after or non-inferiority for low-risk in-product changes.</p><p>- Test big levers—not just UI: pricing models, usage limits, onboarding pathways—and judge success by ARR movement, not micro-metrics.</p><p>- Ship first, then optimize: launch PLG features and immediately run experiments to increase adoption; track daily active usage per feature.</p><p>- Build growth loops from habits: design shareable artifacts and personalized signup paths; drive users back to your domain to capture value.</p><p>- Scale experimentation with AI: use Cursor desktop/cloud agents for parallel builds and visual QA; orchestrate docs/analysis via Claude; automate cleanups and reporting.</p><p>- Make experimentation company-wide: centralize data (BigQuery), broadcast wins/losses in Slack via GrowthBook, and auto-correlate metric dips to releases.</p><p><br></p>]]>
      </itunes:summary>
      <itunes:keywords>experimentation, A/B-testing, feature-flagging, product-experimentation, experimentation-culture, data-driven-decisions, product-development, growth-engineering, startup-experimentation, PLG, product-led-growth, growth-loops, habitual-loops, viral-growth, conversion-rate-optimization, non-inferiority-testing, win-rate, experiment-velocity, AI-assisted-development, Cursor, Claude, GrowthBook, Fixer, hyper-growth, startup-growth, pricing-experiments, B2B-SaaS, onboarding-optimization, personalization, developer-productivity, experiment-automation, data-analysis, BigQuery, Slack-automation, experiment-cleanup, guardrail-metrics, before-and-after-testing, risk-appetite, exploitation-vs-exploration, T-shaped-skills, growth-culture, founder-led-experimentation, ARR-growth, series-B, CRO</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/cf8beed6/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>The 5 pixels that cost LinkedIn a million dollars a month</title>
      <itunes:episode>6</itunes:episode>
      <podcast:episode>6</podcast:episode>
      <itunes:title>The 5 pixels that cost LinkedIn a million dollars a month</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">d917baa4-8cf1-4079-a812-427d04cc0afa</guid>
      <link>https://share.transistor.fm/s/316ada1b</link>
      <description>
        <![CDATA[<p>Summary</p><p>How do you build a culture where nothing ships without evidence—and leaders actually act on the data? Makram Mansour, Head of Marketplace at ID.me and former experimentation leader at LinkedIn and Intuit, shares the systems, mindsets, and guardrails behind “experimenting everywhere.” At LinkedIn, he helped support 10,000+ annual experiments with 2,000 weekly platform users, and he explains the hard-earned lessons (like a 5px UI tweak causing a million-dollar ad loss) that led to a “test before release” mandate. At Intuit, he operationalized “fail forward,” partnering with HR to rewrite OKRs so teams are rewarded for learning, not just launching. Makram breaks down why to shift from MVP to MVT (minimum viable test), how to surface leap-of-faith assumptions with PRFAQs and “unit of one” prototypes, and where AI now unlocks faster, safer front-end testing. He also details critical guardrails—cost visibility for AI infrastructure, ethical and inclusion metrics, and the people-process-technology triad—plus practical ways to remove bottlenecks via a center of excellence. If you’re starting from scratch or scaling your program, you’ll learn how to personalize responsibly at the top of the funnel, define your North Star and signposts, and stack early wins while building influence across the org.</p><p><br></p><p>Timestamps</p><p>[00:45] – Makram’s path: running experimentation at LinkedIn and Intuit, and why nothing ships without an A/B test</p><p>[02:15] – Costly lessons: 5px banner change, algorithm tweaks, and the case for rigorous guardrails</p><p>[06:40] – Leadership discipline: killing features (voice meetups, LinkedIn Stories) and changing OKRs to reward learning</p><p>[11:05] – People, process, technology: top-down and bottom-up tracks, and embedding “fail forward”</p><p>[13:40] – From MVP to MVT: validating leap-of-faith assumptions, PRFAQ, and rapid “unit of one” prototypes</p><p>[15:55] – Bottlenecks and unlocks: engineering/data science capacity, centers of excellence, and AI for fast front-end tests</p><p>[22:45] – Personalization at the top of funnel: avoid waste, design reviews, and right-size testing before building</p><p>[25:45] – Guardrail metrics that matter: AI infra costs, ethics/compliance, and fairness-by-design</p><p>[29:45] – ID.me now: zero-to-one builds, vision-to-values, North Star and leading indicators</p><p>[33:30] – How to start at a new org: crawl-walk-run, small wins, relationships, and over-communication</p><p><br></p><p>Takeaways</p><p>- Shift from MVP to MVT: list leap-of-faith assumptions and design minimum viable tests before you build.</p><p>- Institutionalize learning: align OKRs with “fail forward,” and be willing to kill low-performing features quickly.</p><p>- Build the triad: pair an easy-to-use platform with training, top-down sponsorship, and clear launch processes.</p><p>- Add real guardrails: track AI infrastructure costs, ethics/compliance, and inclusion metrics alongside growth KPIs.</p><p>- Unblock teams: create a center of excellence for data science and enable rapid variants with AI-powered tooling.</p><p>- Start small and visible: rack up quick wins, over-communicate progress, and grow influence through relationships.</p><p><br></p>]]>
      </description>
      <content:encoded>
        <![CDATA[<p>Summary</p><p>How do you build a culture where nothing ships without evidence—and leaders actually act on the data? Makram Mansour, Head of Marketplace at ID.me and former experimentation leader at LinkedIn and Intuit, shares the systems, mindsets, and guardrails behind “experimenting everywhere.” At LinkedIn, he helped support 10,000+ annual experiments with 2,000 weekly platform users, and he explains the hard-earned lessons (like a 5px UI tweak causing a million-dollar ad loss) that led to a “test before release” mandate. At Intuit, he operationalized “fail forward,” partnering with HR to rewrite OKRs so teams are rewarded for learning, not just launching. Makram breaks down why to shift from MVP to MVT (minimum viable test), how to surface leap-of-faith assumptions with PRFAQs and “unit of one” prototypes, and where AI now unlocks faster, safer front-end testing. He also details critical guardrails—cost visibility for AI infrastructure, ethical and inclusion metrics, and the people-process-technology triad—plus practical ways to remove bottlenecks via a center of excellence. If you’re starting from scratch or scaling your program, you’ll learn how to personalize responsibly at the top of the funnel, define your North Star and signposts, and stack early wins while building influence across the org.</p><p><br></p><p>Timestamps</p><p>[00:45] – Makram’s path: running experimentation at LinkedIn and Intuit, and why nothing ships without an A/B test</p><p>[02:15] – Costly lessons: 5px banner change, algorithm tweaks, and the case for rigorous guardrails</p><p>[06:40] – Leadership discipline: killing features (voice meetups, LinkedIn Stories) and changing OKRs to reward learning</p><p>[11:05] – People, process, technology: top-down and bottom-up tracks, and embedding “fail forward”</p><p>[13:40] – From MVP to MVT: validating leap-of-faith assumptions, PRFAQ, and rapid “unit of one” prototypes</p><p>[15:55] – Bottlenecks and unlocks: engineering/data science capacity, centers of excellence, and AI for fast front-end tests</p><p>[22:45] – Personalization at the top of funnel: avoid waste, design reviews, and right-size testing before building</p><p>[25:45] – Guardrail metrics that matter: AI infra costs, ethics/compliance, and fairness-by-design</p><p>[29:45] – ID.me now: zero-to-one builds, vision-to-values, North Star and leading indicators</p><p>[33:30] – How to start at a new org: crawl-walk-run, small wins, relationships, and over-communication</p><p><br></p><p>Takeaways</p><p>- Shift from MVP to MVT: list leap-of-faith assumptions and design minimum viable tests before you build.</p><p>- Institutionalize learning: align OKRs with “fail forward,” and be willing to kill low-performing features quickly.</p><p>- Build the triad: pair an easy-to-use platform with training, top-down sponsorship, and clear launch processes.</p><p>- Add real guardrails: track AI infrastructure costs, ethics/compliance, and inclusion metrics alongside growth KPIs.</p><p>- Unblock teams: create a center of excellence for data science and enable rapid variants with AI-powered tooling.</p><p>- Start small and visible: rack up quick wins, over-communicate progress, and grow influence through relationships.</p><p><br></p>]]>
      </content:encoded>
      <pubDate>Thu, 02 Apr 2026 13:30:43 -0600</pubDate>
      <author>Growthbook</author>
      <enclosure url="https://media.transistor.fm/316ada1b/23090b42.mp3" length="36383883" type="audio/mpeg"/>
      <itunes:author>Growthbook</itunes:author>
      <itunes:image href="https://img.transistorcdn.com/-uQbemIFStvqUb3nBCzSJmqKvgKiJLtEuY1ZXUwJ0co/rs:fill:0:0:1/w:1400/h:1400/q:60/mb:500000/aHR0cHM6Ly9pbWct/dXBsb2FkLXByb2R1/Y3Rpb24udHJhbnNp/c3Rvci5mbS9jOTE3/OTBjYzNlYjM2ZDkw/YTIzMTFmMzM2NTQ2/YjA4NC5wbmc.jpg"/>
      <itunes:duration>2272</itunes:duration>
      <itunes:summary>
        <![CDATA[<p>Summary</p><p>How do you build a culture where nothing ships without evidence—and leaders actually act on the data? Makram Mansour, Head of Marketplace at ID.me and former experimentation leader at LinkedIn and Intuit, shares the systems, mindsets, and guardrails behind “experimenting everywhere.” At LinkedIn, he helped support 10,000+ annual experiments with 2,000 weekly platform users, and he explains the hard-earned lessons (like a 5px UI tweak causing a million-dollar ad loss) that led to a “test before release” mandate. At Intuit, he operationalized “fail forward,” partnering with HR to rewrite OKRs so teams are rewarded for learning, not just launching. Makram breaks down why to shift from MVP to MVT (minimum viable test), how to surface leap-of-faith assumptions with PRFAQs and “unit of one” prototypes, and where AI now unlocks faster, safer front-end testing. He also details critical guardrails—cost visibility for AI infrastructure, ethical and inclusion metrics, and the people-process-technology triad—plus practical ways to remove bottlenecks via a center of excellence. If you’re starting from scratch or scaling your program, you’ll learn how to personalize responsibly at the top of the funnel, define your North Star and signposts, and stack early wins while building influence across the org.</p><p><br></p><p>Timestamps</p><p>[00:45] – Makram’s path: running experimentation at LinkedIn and Intuit, and why nothing ships without an A/B test</p><p>[02:15] – Costly lessons: 5px banner change, algorithm tweaks, and the case for rigorous guardrails</p><p>[06:40] – Leadership discipline: killing features (voice meetups, LinkedIn Stories) and changing OKRs to reward learning</p><p>[11:05] – People, process, technology: top-down and bottom-up tracks, and embedding “fail forward”</p><p>[13:40] – From MVP to MVT: validating leap-of-faith assumptions, PRFAQ, and rapid “unit of one” prototypes</p><p>[15:55] – Bottlenecks and unlocks: engineering/data science capacity, centers of excellence, and AI for fast front-end tests</p><p>[22:45] – Personalization at the top of funnel: avoid waste, design reviews, and right-size testing before building</p><p>[25:45] – Guardrail metrics that matter: AI infra costs, ethics/compliance, and fairness-by-design</p><p>[29:45] – ID.me now: zero-to-one builds, vision-to-values, North Star and leading indicators</p><p>[33:30] – How to start at a new org: crawl-walk-run, small wins, relationships, and over-communication</p><p><br></p><p>Takeaways</p><p>- Shift from MVP to MVT: list leap-of-faith assumptions and design minimum viable tests before you build.</p><p>- Institutionalize learning: align OKRs with “fail forward,” and be willing to kill low-performing features quickly.</p><p>- Build the triad: pair an easy-to-use platform with training, top-down sponsorship, and clear launch processes.</p><p>- Add real guardrails: track AI infrastructure costs, ethics/compliance, and inclusion metrics alongside growth KPIs.</p><p>- Unblock teams: create a center of excellence for data science and enable rapid variants with AI-powered tooling.</p><p>- Start small and visible: rack up quick wins, over-communicate progress, and grow influence through relationships.</p><p><br></p>]]>
      </itunes:summary>
      <itunes:keywords>experimentation, A/B-testing, feature-flagging, product-experimentation, experimentation-culture, data-driven-decisions, product-development, experimentation-platform, North-Star-metrics, guardrail-metrics, hypothesis-testing, fail-forward, minimum-viable-test, personalization, customer-driven-innovation, experimentation-at-scale, LinkedIn, Intuit, id.me, leadership-buy-in, center-of-excellence, experiment-velocity, AI-experimentation, infrastructure-cost, ROI-measurement, product-management, growth-teams, MVT, leap-of-faith-assumptions, experimentation-maturity, crawl-walk-run, experiment-design, KPIs, HIPPO-effect, network-effects, marketplace-experimentation, sphere-of-influence, change-management, people-process-technology</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/316ada1b/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Truist on shipping faster and more safely with AI and human-in-the-loop banking</title>
      <itunes:episode>5</itunes:episode>
      <podcast:episode>5</podcast:episode>
      <itunes:title>Truist on shipping faster and more safely with AI and human-in-the-loop banking</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">91320eba-3686-4b4f-b8ae-af386d6ae881</guid>
      <link>https://share.transistor.fm/s/e0e5d644</link>
      <description>
        <![CDATA[<p>How do you boost developer velocity in a highly regulated industry—without sacrificing safety or customer trust? Charles Williams, Senior Vice President and Software Engineering Director at Truist (formed from the BB&amp;T and SunTrust merger), shares how his team elevates developer experience to ship faster and more reliably. Charles breaks down shifting quality “left” with automation, measuring success with both DORA metrics and developer sentiment, and why human-in-the-loop is non-negotiable for AI in finance. He details Truist’s governance model—steering committees, enterprise architecture, and clear guardrails—to avoid tool sprawl while building a purpose-built AI ecosystem: Microsoft Copilot for productivity, GitLab’s AI-enabled DevSecOps platform for engineering, and separate consumer-facing capabilities. Expect practical insights on starting with low-risk, high-yield use cases (unit tests, docs, security triage), tracking AI utilization, and upskilling teams in prompt engineering so developers can “manage” AI agents effectively. Charles also explores the path to personalized experiences balanced with privacy, why branches should be enhanced—not reduced—by AI, and the cultural skills leaders need now: empathy, neurodiversity awareness, and change management. He closes with where AI is driving ROI first—developer onboarding and pipeline productivity—with code quality gains following close behind.</p><p><br></p><p>Timestamps</p><p>[00:02] – Truist overview and Charles’s mandate: improving developer experience at scale</p><p>[00:56] – AI as a strategic priority; shifting quality left with automation to remove bottlenecks</p><p>[02:19] – Measuring success: DORA metrics plus sentiment—eliminating toil to drive happiness</p><p>[04:33] – Human-in-the-loop AI for high-stakes finance; customer and internal use cases</p><p>[07:25] – How Truist evaluates tools: personas, pain points, and starting with tests, docs, security</p><p>[08:35] – The stack: Microsoft Copilot, GitLab’s AI gateway approach, and tracking utilization</p><p>[10:36] – New skills and culture: prompt engineering, “managing” AI agents, and strong governance</p><p>[20:45] – What’s next: personalization vs privacy, fintech agility + bank stability, and where AI pays off now</p><p><br></p><p>Takeaways</p><p>- Shift quality left with automated checks so developers catch issues early without human gatekeeping.</p><p>- Measure DORA metrics and developer sentiment; remove mundane toil to increase speed and satisfaction.</p><p>- Keep humans in the loop for AI-assisted coding and customer answers—trust but verify in regulated contexts.</p><p>- Build an AI ecosystem with clear purposes (productivity, engineering, consumer) and a steering committee to avoid duplication.</p><p>- Start with low-risk, high-yield AI use cases—unit tests, documentation, and security triage—to build confidence and momentum.</p><p>- Upskill teams in prompt engineering and AI oversight so developers can effectively direct and review AI “agents.”</p><p><br></p>]]>
      </description>
      <content:encoded>
        <![CDATA[<p>How do you boost developer velocity in a highly regulated industry—without sacrificing safety or customer trust? Charles Williams, Senior Vice President and Software Engineering Director at Truist (formed from the BB&amp;T and SunTrust merger), shares how his team elevates developer experience to ship faster and more reliably. Charles breaks down shifting quality “left” with automation, measuring success with both DORA metrics and developer sentiment, and why human-in-the-loop is non-negotiable for AI in finance. He details Truist’s governance model—steering committees, enterprise architecture, and clear guardrails—to avoid tool sprawl while building a purpose-built AI ecosystem: Microsoft Copilot for productivity, GitLab’s AI-enabled DevSecOps platform for engineering, and separate consumer-facing capabilities. Expect practical insights on starting with low-risk, high-yield use cases (unit tests, docs, security triage), tracking AI utilization, and upskilling teams in prompt engineering so developers can “manage” AI agents effectively. Charles also explores the path to personalized experiences balanced with privacy, why branches should be enhanced—not reduced—by AI, and the cultural skills leaders need now: empathy, neurodiversity awareness, and change management. He closes with where AI is driving ROI first—developer onboarding and pipeline productivity—with code quality gains following close behind.</p><p><br></p><p>Timestamps</p><p>[00:02] – Truist overview and Charles’s mandate: improving developer experience at scale</p><p>[00:56] – AI as a strategic priority; shifting quality left with automation to remove bottlenecks</p><p>[02:19] – Measuring success: DORA metrics plus sentiment—eliminating toil to drive happiness</p><p>[04:33] – Human-in-the-loop AI for high-stakes finance; customer and internal use cases</p><p>[07:25] – How Truist evaluates tools: personas, pain points, and starting with tests, docs, security</p><p>[08:35] – The stack: Microsoft Copilot, GitLab’s AI gateway approach, and tracking utilization</p><p>[10:36] – New skills and culture: prompt engineering, “managing” AI agents, and strong governance</p><p>[20:45] – What’s next: personalization vs privacy, fintech agility + bank stability, and where AI pays off now</p><p><br></p><p>Takeaways</p><p>- Shift quality left with automated checks so developers catch issues early without human gatekeeping.</p><p>- Measure DORA metrics and developer sentiment; remove mundane toil to increase speed and satisfaction.</p><p>- Keep humans in the loop for AI-assisted coding and customer answers—trust but verify in regulated contexts.</p><p>- Build an AI ecosystem with clear purposes (productivity, engineering, consumer) and a steering committee to avoid duplication.</p><p>- Start with low-risk, high-yield AI use cases—unit tests, documentation, and security triage—to build confidence and momentum.</p><p>- Upskill teams in prompt engineering and AI oversight so developers can effectively direct and review AI “agents.”</p><p><br></p>]]>
      </content:encoded>
      <pubDate>Thu, 02 Apr 2026 13:24:48 -0600</pubDate>
      <author>Growthbook</author>
      <enclosure url="https://media.transistor.fm/e0e5d644/084ba601.mp3" length="26898727" type="audio/mpeg"/>
      <itunes:author>Growthbook</itunes:author>
      <itunes:image href="https://img.transistorcdn.com/HEpKLq6MvVN6TCZylcbOgld5XS28KXeB3545kRs-JwE/rs:fill:0:0:1/w:1400/h:1400/q:60/mb:500000/aHR0cHM6Ly9pbWct/dXBsb2FkLXByb2R1/Y3Rpb24udHJhbnNp/c3Rvci5mbS9jNTQx/Y2Q0YWM4NTEzZjBk/YmI0N2RiMDhjYjMy/YjYxMy5qcGc.jpg"/>
      <itunes:duration>1679</itunes:duration>
      <itunes:summary>
        <![CDATA[<p>How do you boost developer velocity in a highly regulated industry—without sacrificing safety or customer trust? Charles Williams, Senior Vice President and Software Engineering Director at Truist (formed from the BB&amp;T and SunTrust merger), shares how his team elevates developer experience to ship faster and more reliably. Charles breaks down shifting quality “left” with automation, measuring success with both DORA metrics and developer sentiment, and why human-in-the-loop is non-negotiable for AI in finance. He details Truist’s governance model—steering committees, enterprise architecture, and clear guardrails—to avoid tool sprawl while building a purpose-built AI ecosystem: Microsoft Copilot for productivity, GitLab’s AI-enabled DevSecOps platform for engineering, and separate consumer-facing capabilities. Expect practical insights on starting with low-risk, high-yield use cases (unit tests, docs, security triage), tracking AI utilization, and upskilling teams in prompt engineering so developers can “manage” AI agents effectively. Charles also explores the path to personalized experiences balanced with privacy, why branches should be enhanced—not reduced—by AI, and the cultural skills leaders need now: empathy, neurodiversity awareness, and change management. He closes with where AI is driving ROI first—developer onboarding and pipeline productivity—with code quality gains following close behind.</p><p><br></p><p>Timestamps</p><p>[00:02] – Truist overview and Charles’s mandate: improving developer experience at scale</p><p>[00:56] – AI as a strategic priority; shifting quality left with automation to remove bottlenecks</p><p>[02:19] – Measuring success: DORA metrics plus sentiment—eliminating toil to drive happiness</p><p>[04:33] – Human-in-the-loop AI for high-stakes finance; customer and internal use cases</p><p>[07:25] – How Truist evaluates tools: personas, pain points, and starting with tests, docs, security</p><p>[08:35] – The stack: Microsoft Copilot, GitLab’s AI gateway approach, and tracking utilization</p><p>[10:36] – New skills and culture: prompt engineering, “managing” AI agents, and strong governance</p><p>[20:45] – What’s next: personalization vs privacy, fintech agility + bank stability, and where AI pays off now</p><p><br></p><p>Takeaways</p><p>- Shift quality left with automated checks so developers catch issues early without human gatekeeping.</p><p>- Measure DORA metrics and developer sentiment; remove mundane toil to increase speed and satisfaction.</p><p>- Keep humans in the loop for AI-assisted coding and customer answers—trust but verify in regulated contexts.</p><p>- Build an AI ecosystem with clear purposes (productivity, engineering, consumer) and a steering committee to avoid duplication.</p><p>- Start with low-risk, high-yield AI use cases—unit tests, documentation, and security triage—to build confidence and momentum.</p><p>- Upskill teams in prompt engineering and AI oversight so developers can effectively direct and review AI “agents.”</p><p><br></p>]]>
      </itunes:summary>
      <itunes:keywords>experimentation, A/B-testing, feature-flagging, product-experimentation, experimentation-culture, data-driven-decisions, product-development, developer-experience, AI-in-development, enterprise-AI, code-quality, developer-productivity, shift-left-testing, CI-CD-pipelines, fintech-innovation, banking-technology, AI-governance, human-in-the-loop, change-management, experimentation-strategy, AI-in-financial-services, developer-onboarding, prompt-engineering, AI-adoption, GitLab, Truist, DORA-metrics, customer-experience, machine-learning, neurodiversity</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/e0e5d644/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>From chatbots to open-world agents at Microsoft: evals, go-live metrics, and copilot velocity</title>
      <itunes:episode>4</itunes:episode>
      <podcast:episode>4</podcast:episode>
      <itunes:title>From chatbots to open-world agents at Microsoft: evals, go-live metrics, and copilot velocity</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">d7943904-0cd0-4d98-bdd4-57026366eb7e</guid>
      <link>https://share.transistor.fm/s/cb699bef</link>
      <description>
        <![CDATA[<p>AI is moving so fast that what “good” looked like a few months ago is already outdated. </p><p>So how do you measure value, ship safely, and scale what works? Marco Casalaina, VP of Products, Core AI and AI Futurist at Microsoft, joins to unpack how his team builds and evaluates next‑gen AI—at hyperspeed. </p><p>Marco leads the AI Futures team and previously led Azure OpenAI, Azure Cognitive Services, Responsible AI, and AI Studio; before Microsoft, he ran Salesforce Einstein. </p><p>He explains why enterprise value is best measured by go‑lives and real usage, how Microsoft’s Foundry equips developers with agent‑specific evals (tool call accuracy, task adherence), and why old metrics like “accepted completions” don’t fit modern dev loops. </p><p>We dig into model routing now productized across model families, orchestration frameworks and the Copilot SDK, shared memory experiments, and the rise of self‑verifying agents that iterate to defined thresholds. </p><p>Expect concrete examples—from rewriting docs for coding agents to Ralph loops with browser testing—and practical advice for leaders: major in evals, set acceptable error rates by use case, and get hands‑on with the tools daily.</p><p><br></p><p><strong>Timestamps</strong></p><p>[00:45] – Guest intro and Microsoft’s enterprise AI focus</p><p>[02:07] – Measuring value: go‑lives, telemetry thresholds, and token volume</p><p>[03:43] – From chatbots to agents: Foundry evals (tool calls, task adherence) and A/B testing in Microsoft 365 Copilot</p><p>[06:17] – When “good” changes monthly: model routing productized across model families</p><p>[07:23] – Orchestration and Copilot SDK: agents that create their own tools; OpenClaw and shared memory experiments</p><p>[11:45] – Engagement redefined: coding agents read your docs; writing for agents vs. humans</p><p>[14:55] – New dev loops: why accepted completions died; Ralph loops and self‑verifying builds</p><p>[17:06] – Evals in practice and guardrails: thresholds, non‑determinism, and out‑of‑domain tests; how to keep up without burning out</p><p><br></p><p><strong>Takeaways</strong></p><p>- Measure value by go‑lives and real usage (token volume), not time in portals or playgrounds.</p><p>- Evolve evals for agents: track tool call accuracy/success and task completion/adherence; A/B test models and strategies.</p><p>- Productize adaptability with model routing to match tasks to the right model family as capabilities shift.</p><p>- Build self‑verification into workflows: pair agents with automated testing (e.g., browser runners) and iterate to thresholds, not perfection.</p><p>- Write for agents as readers: tighten documentation, ship vetted code samples, and monitor bot traffic patterns.</p><p>- Guardrail open‑world agents: add out‑of‑domain evals and explicit capability limits; set acceptable error rates based on the stakes of your use case.</p>]]>
      </description>
      <content:encoded>
        <![CDATA[<p>AI is moving so fast that what “good” looked like a few months ago is already outdated. </p><p>So how do you measure value, ship safely, and scale what works? Marco Casalaina, VP of Products, Core AI and AI Futurist at Microsoft, joins to unpack how his team builds and evaluates next‑gen AI—at hyperspeed. </p><p>Marco leads the AI Futures team and previously led Azure OpenAI, Azure Cognitive Services, Responsible AI, and AI Studio; before Microsoft, he ran Salesforce Einstein. </p><p>He explains why enterprise value is best measured by go‑lives and real usage, how Microsoft’s Foundry equips developers with agent‑specific evals (tool call accuracy, task adherence), and why old metrics like “accepted completions” don’t fit modern dev loops. </p><p>We dig into model routing now productized across model families, orchestration frameworks and the Copilot SDK, shared memory experiments, and the rise of self‑verifying agents that iterate to defined thresholds. </p><p>Expect concrete examples—from rewriting docs for coding agents to Ralph loops with browser testing—and practical advice for leaders: major in evals, set acceptable error rates by use case, and get hands‑on with the tools daily.</p><p><br></p><p><strong>Timestamps</strong></p><p>[00:45] – Guest intro and Microsoft’s enterprise AI focus</p><p>[02:07] – Measuring value: go‑lives, telemetry thresholds, and token volume</p><p>[03:43] – From chatbots to agents: Foundry evals (tool calls, task adherence) and A/B testing in Microsoft 365 Copilot</p><p>[06:17] – When “good” changes monthly: model routing productized across model families</p><p>[07:23] – Orchestration and Copilot SDK: agents that create their own tools; OpenClaw and shared memory experiments</p><p>[11:45] – Engagement redefined: coding agents read your docs; writing for agents vs. humans</p><p>[14:55] – New dev loops: why accepted completions died; Ralph loops and self‑verifying builds</p><p>[17:06] – Evals in practice and guardrails: thresholds, non‑determinism, and out‑of‑domain tests; how to keep up without burning out</p><p><br></p><p><strong>Takeaways</strong></p><p>- Measure value by go‑lives and real usage (token volume), not time in portals or playgrounds.</p><p>- Evolve evals for agents: track tool call accuracy/success and task completion/adherence; A/B test models and strategies.</p><p>- Productize adaptability with model routing to match tasks to the right model family as capabilities shift.</p><p>- Build self‑verification into workflows: pair agents with automated testing (e.g., browser runners) and iterate to thresholds, not perfection.</p><p>- Write for agents as readers: tighten documentation, ship vetted code samples, and monitor bot traffic patterns.</p><p>- Guardrail open‑world agents: add out‑of‑domain evals and explicit capability limits; set acceptable error rates based on the stakes of your use case.</p>]]>
      </content:encoded>
      <pubDate>Tue, 17 Mar 2026 01:00:00 -0600</pubDate>
      <author>Ashley Stirrup</author>
      <enclosure url="https://media.transistor.fm/cb699bef/ce3b7b9b.mp3" length="27294738" type="audio/mpeg"/>
      <itunes:author>Ashley Stirrup</itunes:author>
      <itunes:image href="https://img.transistorcdn.com/9a4onDNsv69nArjMHqcRfRJdZQqkm4fNo5_aSSeEHB4/rs:fill:0:0:1/w:1400/h:1400/q:60/mb:500000/aHR0cHM6Ly9pbWct/dXBsb2FkLXByb2R1/Y3Rpb24udHJhbnNp/c3Rvci5mbS8yYWJj/ZjUxNGVjMmYyNmY1/ZTE2MmQxN2Y3MTdi/NzEwOS5qcGc.jpg"/>
      <itunes:duration>1706</itunes:duration>
      <itunes:summary>AI is moving so fast that what “good” looked like a few months ago is already outdated. So how do you measure value, ship safely, and scale what works? Marco Casalaina, VP of Products, Core AI and AI Futurist at Microsoft, joins to unpack how his team builds and evaluates next‑gen AI—at hyperspeed. Marco leads the AI Futures team and previously led Azure OpenAI, Azure Cognitive Services, Responsible AI, and AI Studio; before Microsoft, he ran Salesforce Einstein. He explains why enterprise value is best measured by go‑lives and real usage, how Microsoft’s Foundry equips developers with agent‑specific evals (tool call accuracy, task adherence), and why old metrics like “accepted completions” don’t fit modern dev loops. We dig into model routing now productized across model families, orchestration frameworks and the Copilot SDK, shared memory experiments, and the rise of self‑verifying agents that iterate to defined thresholds. Expect concrete examples—from rewriting docs for coding agents to Ralph loops with browser testing—and practical advice for leaders: major in evals, set acceptable error rates by use case, and get hands‑on with the tools daily.
Timestamps[00:45] – Guest intro and Microsoft’s enterprise AI focus[02:07] – Measuring value: go‑lives, telemetry thresholds, and token volume[03:43] – From chatbots to agents: Foundry evals (tool calls, task adherence) and A/B testing in Microsoft 365 Copilot[06:17] – When “good” changes monthly: model routing productized across model families[07:23] – Orchestration and Copilot SDK: agents that create their own tools; OpenClaw and shared memory experiments[11:45] – Engagement redefined: coding agents read your docs; writing for agents vs. humans[14:55] – New dev loops: why accepted completions died; Ralph loops and self‑verifying builds[17:06] – Evals in practice and guardrails: thresholds, non‑determinism, and out‑of‑domain tests; how to keep up without burning out
Takeaways- Measure value by go‑lives and real usage (token volume), not time in portals or playgrounds.- Evolve evals for agents: track tool call accuracy/success and task completion/adherence; A/B test models and strategies.- Productize adaptability with model routing to match tasks to the right model family as capabilities shift.- Build self‑verification into workflows: pair agents with automated testing (e.g., browser runners) and iterate to thresholds, not perfection.- Write for agents as readers: tighten documentation, ship vetted code samples, and monitor bot traffic patterns.- Guardrail open‑world agents: add out‑of‑domain evals and explicit capability limits; set acceptable error rates based on the stakes of your use case.</itunes:summary>
      <itunes:subtitle>AI is moving so fast that what “good” looked like a few months ago is already outdated. So how do you measure value, ship safely, and scale what works? Marco Casalaina, VP of Products, Core AI and AI Futurist at Microsoft, joins to unpack how his team bui</itunes:subtitle>
      <itunes:keywords>experimentation, A/B-testing, feature-flagging, product-experimentation, experimentation-culture, data-driven-decisions, product-development, Microsoft, Marco Casalaina, Copilot, AI-agents, model-evaluation, go-live-metrics, task-completion, tool-call-accuracy, groundedness-testing, coherence-evaluation, fluency-metrics, model-routing, Azure-OpenAI, AI-governance, responsible-AI, LLM-evaluation, evaluation-thresholds, token-volume, agentic-workflows, error-tolerance, self-verification, Ralph-loop, eval-metrics, task-adherence, non-deterministic-testing</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/cb699bef/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Upwork on AI-Driven Ops at scale</title>
      <itunes:episode>3</itunes:episode>
      <podcast:episode>3</podcast:episode>
      <itunes:title>Upwork on AI-Driven Ops at scale</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">1d8c0760-23e3-4439-bed9-82f7481748bb</guid>
      <link>https://share.transistor.fm/s/73141a1f</link>
      <description>
        <![CDATA[<p><strong>Summary</strong></p><p>What do you test rigorously—and what do you ship fast and fix forward—when every change could impact millions? </p><p>Vinoj Kumar, Vice President of Engineering at Upwork, leads at the intersection of infrastructure and product, where feedback loops are longer and the blast radius is wider. </p><p>He shares a pragmatic framework for experimentation—blast radius x reversibility—that sets testing rigor, plus how he measures success in product terms: faster search, resilient marketplace trust, and developer velocity. Vinoj explains why “high engagement” can mask low-quality experiences, how his team instrumented an internal NL chatbot with turns-to-success and downstream signals (like fewer JIRA tickets), and how a composite metric—cost per quality inference (CPQI)—aligns finance, engineering, and data science by uniting cloud costs, performance, and model accuracy. </p><p>He details where AI is already paying off (build pipelines, incident detection, testing), how to monitor model drift post-launch, and why some wins on paper must be killed in production to protect trust—like a high-hit-rate caching project that surfaced stale profile data. </p><p>Expect concrete practices: shadow traffic, slow canaries, synthetic staging that mirrors reality, feature flags, LLMs-as-judges, and the mindset to tie infrastructure to business outcomes.</p><p><strong>Timestamps</strong></p><p>[00:45] – Guest intro: Infrastructure meets product—and why experimentation looks different</p><p>[01:36] – Deciding what to test: blast radius x reversibility; canaries, shadow traffic, ship-and-monitor</p><p>[03:09] – Defining “good”: internal dev metrics vs. marketplace outcomes—and when engagement lies</p><p>[06:23] – Case study: “Talk to Data” chatbot—thumbs, turns-to-success, and reduced JIRA tickets</p><p>[09:45] – CPQI: a composite metric for cost, performance, and model quality that breaks silos</p><p>[16:55] – AI in engineering: build-time gains, MTTR/MTTD, agentic testing, and drift monitoring</p><p>[24:06] – The caching miss: 92% hit rate, stale data, trust risks—and what to do instead</p><p>[29:12] – Career advice: balance stability with bold experiments; always link infra to business value</p><p><br></p><p><strong>Takeaways</strong></p><p>- Decide testing rigor with blast radius x reversibility; reserve heavy testing for irreversible, high-impact systems.</p><p>- Measure quality by efficiency and success ratio—not raw clicks or query counts.</p><p>- Instrument NL tools with “turns to success” and track downstream impact (e.g., fewer ad hoc data tickets).</p><p>- Build composite metrics (e.g., CPQI) to align finance, engineering, and data science around shared outcomes.</p><p>- Use AI to accelerate builds, detect incidents sooner, and evaluate models; watch MTTR and MTTD.</p><p>- Treat ML features as living systems: feature-flag rollouts, realistic staging, drift monitoring, and LLM-as-judge evaluations—and be willing to kill “wins” that erode trust.</p>]]>
      </description>
      <content:encoded>
        <![CDATA[<p><strong>Summary</strong></p><p>What do you test rigorously—and what do you ship fast and fix forward—when every change could impact millions? </p><p>Vinoj Kumar, Vice President of Engineering at Upwork, leads at the intersection of infrastructure and product, where feedback loops are longer and the blast radius is wider. </p><p>He shares a pragmatic framework for experimentation—blast radius x reversibility—that sets testing rigor, plus how he measures success in product terms: faster search, resilient marketplace trust, and developer velocity. Vinoj explains why “high engagement” can mask low-quality experiences, how his team instrumented an internal NL chatbot with turns-to-success and downstream signals (like fewer JIRA tickets), and how a composite metric—cost per quality inference (CPQI)—aligns finance, engineering, and data science by uniting cloud costs, performance, and model accuracy. </p><p>He details where AI is already paying off (build pipelines, incident detection, testing), how to monitor model drift post-launch, and why some wins on paper must be killed in production to protect trust—like a high-hit-rate caching project that surfaced stale profile data. </p><p>Expect concrete practices: shadow traffic, slow canaries, synthetic staging that mirrors reality, feature flags, LLMs-as-judges, and the mindset to tie infrastructure to business outcomes.</p><p><strong>Timestamps</strong></p><p>[00:45] – Guest intro: Infrastructure meets product—and why experimentation looks different</p><p>[01:36] – Deciding what to test: blast radius x reversibility; canaries, shadow traffic, ship-and-monitor</p><p>[03:09] – Defining “good”: internal dev metrics vs. marketplace outcomes—and when engagement lies</p><p>[06:23] – Case study: “Talk to Data” chatbot—thumbs, turns-to-success, and reduced JIRA tickets</p><p>[09:45] – CPQI: a composite metric for cost, performance, and model quality that breaks silos</p><p>[16:55] – AI in engineering: build-time gains, MTTR/MTTD, agentic testing, and drift monitoring</p><p>[24:06] – The caching miss: 92% hit rate, stale data, trust risks—and what to do instead</p><p>[29:12] – Career advice: balance stability with bold experiments; always link infra to business value</p><p><br></p><p><strong>Takeaways</strong></p><p>- Decide testing rigor with blast radius x reversibility; reserve heavy testing for irreversible, high-impact systems.</p><p>- Measure quality by efficiency and success ratio—not raw clicks or query counts.</p><p>- Instrument NL tools with “turns to success” and track downstream impact (e.g., fewer ad hoc data tickets).</p><p>- Build composite metrics (e.g., CPQI) to align finance, engineering, and data science around shared outcomes.</p><p>- Use AI to accelerate builds, detect incidents sooner, and evaluate models; watch MTTR and MTTD.</p><p>- Treat ML features as living systems: feature-flag rollouts, realistic staging, drift monitoring, and LLM-as-judge evaluations—and be willing to kill “wins” that erode trust.</p>]]>
      </content:encoded>
      <pubDate>Wed, 11 Mar 2026 01:00:00 -0600</pubDate>
      <author>Ashley Stirrup</author>
      <enclosure url="https://media.transistor.fm/73141a1f/0bc3f76d.mp3" length="29436325" type="audio/mpeg"/>
      <itunes:author>Ashley Stirrup</itunes:author>
      <itunes:image href="https://img.transistorcdn.com/fXyB_SNb_FkVOwrGaNEkppqbSZ5XrZGC3oLDC2u8Rr8/rs:fill:0:0:1/w:1400/h:1400/q:60/mb:500000/aHR0cHM6Ly9pbWct/dXBsb2FkLXByb2R1/Y3Rpb24udHJhbnNp/c3Rvci5mbS8zZWY3/MzNlZjg2YzRiNDY5/MDg2OTEwMjQ1OTM2/MjIxNC5qcGc.jpg"/>
      <itunes:duration>1840</itunes:duration>
      <itunes:summary>SummaryWhat do you test rigorously—and what do you ship fast and fix forward—when every change could impact millions? Vinoj Kumar, Vice President of Engineering at Upwork, leads at the intersection of infrastructure and product, where feedback loops are longer and the blast radius is wider. He shares a pragmatic framework for experimentation—blast radius x reversibility—that sets testing rigor, plus how he measures success in product terms: faster search, resilient marketplace trust, and developer velocity. Vinoj explains why “high engagement” can mask low-quality experiences, how his team instrumented an internal NL chatbot with turns-to-success and downstream signals (like fewer JIRA tickets), and how a composite metric—cost per quality inference (CPQI)—aligns finance, engineering, and data science by uniting cloud costs, performance, and model accuracy. He details where AI is already paying off (build pipelines, incident detection, testing), how to monitor model drift post-launch, and why some wins on paper must be killed in production to protect trust—like a high-hit-rate caching project that surfaced stale profile data. Expect concrete practices: shadow traffic, slow canaries, synthetic staging that mirrors reality, feature flags, LLMs-as-judges, and the mindset to tie infrastructure to business outcomes.Timestamps[00:45] – Guest intro: Infrastructure meets product—and why experimentation looks different[01:36] – Deciding what to test: blast radius x reversibility; canaries, shadow traffic, ship-and-monitor[03:09] – Defining “good”: internal dev metrics vs. marketplace outcomes—and when engagement lies[06:23] – Case study: “Talk to Data” chatbot—thumbs, turns-to-success, and reduced JIRA tickets[09:45] – CPQI: a composite metric for cost, performance, and model quality that breaks silos[16:55] – AI in engineering: build-time gains, MTTR/MTTD, agentic testing, and drift monitoring[24:06] – The caching miss: 92% hit rate, stale data, trust risks—and what to do instead[29:12] – Career advice: balance stability with bold experiments; always link infra to business value
Takeaways- Decide testing rigor with blast radius x reversibility; reserve heavy testing for irreversible, high-impact systems.- Measure quality by efficiency and success ratio—not raw clicks or query counts.- Instrument NL tools with “turns to success” and track downstream impact (e.g., fewer ad hoc data tickets).- Build composite metrics (e.g., CPQI) to align finance, engineering, and data science around shared outcomes.- Use AI to accelerate builds, detect incidents sooner, and evaluate models; watch MTTR and MTTD.- Treat ML features as living systems: feature-flag rollouts, realistic staging, drift monitoring, and LLM-as-judge evaluations—and be willing to kill “wins” that erode trust.</itunes:summary>
      <itunes:subtitle>SummaryWhat do you test rigorously—and what do you ship fast and fix forward—when every change could impact millions? Vinoj Kumar, Vice President of Engineering at Upwork, leads at the intersection of infrastructure and product, where feedback loops are l</itunes:subtitle>
      <itunes:keywords>experimentation, A/B-testing, feature-flagging, product-experimentation, experimentation-culture, data-driven-decisions, product-development, Upwork, Vinoj Kumar, blast-radius, CPQI, infrastructure-optimization, platform-engineering, search-relevance, machine-learning-models, cost-per-quality-inference, performance-metrics, developer-experience, marketplace-metrics, freelancer-matching, inference-cost, model-accuracy, latency-optimization, quality-measurement, natural-language-tools, talk-to-data, turns-to-success, friction-measurement, composite-metrics, infrastructure-scaling, AI-agents, build-efficiency</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/73141a1f/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>How Moxie Pest Control boosted conversions 5% with data and coaching</title>
      <itunes:episode>2</itunes:episode>
      <podcast:episode>2</podcast:episode>
      <itunes:title>How Moxie Pest Control boosted conversions 5% with data and coaching</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">8bfa7c21-9558-48ed-89b2-1515c3644d0d</guid>
      <link>https://share.transistor.fm/s/366bbe93</link>
      <description>
        <![CDATA[<p>Think pest control isn’t a digital business? Think again. </p><p>Raj Mehta, Vice President of Product and Technology at Moxie Pest Control, outlines how he turned a spreadsheet-run operation into a data-driven engine across 9,000+ daily calls. </p><p>Raj shares how consolidating fragmented systems into a data lake unlocked automation—from shrinking lead routing from 20–25 minutes to under 30 seconds—to deploying AI-powered call intelligence that scores every sales and retention conversation against a playbook. He breaks down a practical roadmap for traditional businesses: build MVPs, pilot in one branch with a trained feedback team, iterate fast, then scale. </p><p>You’ll hear how he positioned AI as a growth amplifier (not a job cutter), the difference between deterministic automation and LLM use cases, and the measurable impact: a 5% lift in conversion that compounds in a recurring-revenue model. Plus, Raj’s concise advice for leaders bringing AI into operations without breaking trust or momentum.</p><p><br></p><p><strong>Timestamps</strong></p><p>[00:45] – Guest intro: Raj Mehta, Moxie’s tech transformation, and 9,000+ daily calls</p><p>[01:20] – Starting point: 90% of ops in spreadsheets; why a data lake became the foundation</p><p>[02:45] – Automating time-to-lead: 25 minutes to &lt;30 seconds and a 5% conversion lift</p><p>[04:27] – Roadmap design: MVPs, single-branch pilots, and scaling what works</p><p>[06:05] – Culture building: framing AI as growth and upskilling, not headcount cuts</p><p>[07:34] – Two lanes of automation: deterministic scripts vs. LLM-driven workflows</p><p>[08:45] – Call intelligence: scoring every sales/retention call and coaching at scale</p><p>[14:05] – Impact and advice: recurring revenue compounding and Raj’s playbook for getting started</p><p><br></p><p><strong>Takeaways</strong></p><p>- Build a single source of truth (data lake) to power automation and AI reliably.</p><p>- Cut time-to-lead with workflow automation and track the downstream impact on conversion.</p><p>- Pilot in one branch with a trained “feedback team,” iterate, then roll out—don’t scale too soon.</p><p>- Position AI as a growth multiplier; retain and upskill top performers to shape the culture.</p><p>- Separate deterministic automation from LLM use cases; do deep discovery with frontline teams.</p><p>- Use AI call intelligence to score every call against your playbook, surface coaching themes, and save manager time.</p><p><br></p>]]>
      </description>
      <content:encoded>
        <![CDATA[<p>Think pest control isn’t a digital business? Think again. </p><p>Raj Mehta, Vice President of Product and Technology at Moxie Pest Control, outlines how he turned a spreadsheet-run operation into a data-driven engine across 9,000+ daily calls. </p><p>Raj shares how consolidating fragmented systems into a data lake unlocked automation—from shrinking lead routing from 20–25 minutes to under 30 seconds—to deploying AI-powered call intelligence that scores every sales and retention conversation against a playbook. He breaks down a practical roadmap for traditional businesses: build MVPs, pilot in one branch with a trained feedback team, iterate fast, then scale. </p><p>You’ll hear how he positioned AI as a growth amplifier (not a job cutter), the difference between deterministic automation and LLM use cases, and the measurable impact: a 5% lift in conversion that compounds in a recurring-revenue model. Plus, Raj’s concise advice for leaders bringing AI into operations without breaking trust or momentum.</p><p><br></p><p><strong>Timestamps</strong></p><p>[00:45] – Guest intro: Raj Mehta, Moxie’s tech transformation, and 9,000+ daily calls</p><p>[01:20] – Starting point: 90% of ops in spreadsheets; why a data lake became the foundation</p><p>[02:45] – Automating time-to-lead: 25 minutes to &lt;30 seconds and a 5% conversion lift</p><p>[04:27] – Roadmap design: MVPs, single-branch pilots, and scaling what works</p><p>[06:05] – Culture building: framing AI as growth and upskilling, not headcount cuts</p><p>[07:34] – Two lanes of automation: deterministic scripts vs. LLM-driven workflows</p><p>[08:45] – Call intelligence: scoring every sales/retention call and coaching at scale</p><p>[14:05] – Impact and advice: recurring revenue compounding and Raj’s playbook for getting started</p><p><br></p><p><strong>Takeaways</strong></p><p>- Build a single source of truth (data lake) to power automation and AI reliably.</p><p>- Cut time-to-lead with workflow automation and track the downstream impact on conversion.</p><p>- Pilot in one branch with a trained “feedback team,” iterate, then roll out—don’t scale too soon.</p><p>- Position AI as a growth multiplier; retain and upskill top performers to shape the culture.</p><p>- Separate deterministic automation from LLM use cases; do deep discovery with frontline teams.</p><p>- Use AI call intelligence to score every call against your playbook, surface coaching themes, and save manager time.</p><p><br></p>]]>
      </content:encoded>
      <pubDate>Thu, 05 Mar 2026 00:00:00 -0700</pubDate>
      <author>Ashley Stirrup</author>
      <enclosure url="https://media.transistor.fm/366bbe93/58926313.mp3" length="16394655" type="audio/mpeg"/>
      <itunes:author>Ashley Stirrup</itunes:author>
      <itunes:image href="https://img.transistorcdn.com/2xwSt78q-6Po5bNPd6tH1uLi_4fSqlrSGaQxS3_izcQ/rs:fill:0:0:1/w:1400/h:1400/q:60/mb:500000/aHR0cHM6Ly9pbWct/dXBsb2FkLXByb2R1/Y3Rpb24udHJhbnNp/c3Rvci5mbS8yYmY4/OTExZjM0ZGYzOTNl/MmZjMDY1MGZkMThk/NzExZC5qcGc.jpg"/>
      <itunes:duration>1025</itunes:duration>
      <itunes:summary>Think pest control isn’t a digital business? Think again. Raj Mehta, Vice President of Product and Technology at Moxie Pest Control, outlines how he turned a spreadsheet-run operation into a data-driven engine across 9,000+ daily calls. Raj shares how consolidating fragmented systems into a data lake unlocked automation—from shrinking lead routing from 20–25 minutes to under 30 seconds—to deploying AI-powered call intelligence that scores every sales and retention conversation against a playbook. He breaks down a practical roadmap for traditional businesses: build MVPs, pilot in one branch with a trained feedback team, iterate fast, then scale. You’ll hear how he positioned AI as a growth amplifier (not a job cutter), the difference between deterministic automation and LLM use cases, and the measurable impact: a 5% lift in conversion that compounds in a recurring-revenue model. Plus, Raj’s concise advice for leaders bringing AI into operations without breaking trust or momentum.
Timestamps[00:45] – Guest intro: Raj Mehta, Moxie’s tech transformation, and 9,000+ daily calls[01:20] – Starting point: 90% of ops in spreadsheets; why a data lake became the foundation[02:45] – Automating time-to-lead: 25 minutes to &amp;lt;30 seconds and a 5% conversion lift[04:27] – Roadmap design: MVPs, single-branch pilots, and scaling what works[06:05] – Culture building: framing AI as growth and upskilling, not headcount cuts[07:34] – Two lanes of automation: deterministic scripts vs. LLM-driven workflows[08:45] – Call intelligence: scoring every sales/retention call and coaching at scale[14:05] – Impact and advice: recurring revenue compounding and Raj’s playbook for getting started
Takeaways- Build a single source of truth (data lake) to power automation and AI reliably.- Cut time-to-lead with workflow automation and track the downstream impact on conversion.- Pilot in one branch with a trained “feedback team,” iterate, then roll out—don’t scale too soon.- Position AI as a growth multiplier; retain and upskill top performers to shape the culture.- Separate deterministic automation from LLM use cases; do deep discovery with frontline teams.- Use AI call intelligence to score every call against your playbook, surface coaching themes, and save manager time.</itunes:summary>
      <itunes:subtitle>Think pest control isn’t a digital business? Think again. Raj Mehta, Vice President of Product and Technology at Moxie Pest Control, outlines how he turned a spreadsheet-run operation into a data-driven engine across 9,000+ daily calls. Raj shares how con</itunes:subtitle>
      <itunes:keywords>experimentation, A/B-testing, feature-flagging, product-experimentation, experimentation-culture, data-driven-decisions, product-development, Moxie Pest Control, Raj Mehta, data-consolidation, call-center-optimization, lead-conversion, automation, call-coaching, AI-powered-insights, NPS-scores, customer-retention, sales-playbook, workflow-automation, MVP-testing, phased-rollout, cultural-buy-in, AI-adoption, operational-efficiency, phone-system-automation, lead-scoring, conversation-analysis, team-enablement, scaling-automation</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/366bbe93/transcript.txt" type="text/plain"/>
    </item>
    <item>
      <title>Typeform on how to stop running experiments and start earning them</title>
      <itunes:episode>1</itunes:episode>
      <podcast:episode>1</podcast:episode>
      <itunes:title>Typeform on how to stop running experiments and start earning them</itunes:title>
      <itunes:episodeType>full</itunes:episodeType>
      <guid isPermaLink="false">b6e9183d-a2bb-416f-aa6c-30ba920c78c0</guid>
      <link>https://share.transistor.fm/s/bd2615c2</link>
      <description>
        <![CDATA[<p>If 80% of A/B tests fail, how do you de-risk decisions that touch pricing, product, and brand? Aleksandra (Aleks) Bass, Chief Product &amp; Technology Officer at Typeform, shares how her team “earns the right to A/B test” with medical-grade rigor—moving from literature reviews and user tests to simulated trials before exposing changes to customers. She details Typeform’s repositioning from “forms” to an AI engagement platform—and the pricing and packaging bet behind it: a 15% drop in new business count offset by a 32% increase in ASP and a 25% lift in annual attach. </p><p>Aleks unpacks how Typeform AI acts as a co-pilot that doubled activation and boosted one-day conversion, plus the design shifts (CTA altitude and onboarding) that increased adoption. She also reveals why they moved video features down-tier and what a head-to-head test showed: video interviewers generated 14x more words and 10x fewer skipped questions with comparable completion time. </p><p>Finally, Aleks breaks down the cultural side—eliminating “anti-knowledge,” standardizing experiment design, and creating a cross-functional review that prevents false learnings—along with how her data engineering team evaluates LLMs for quality, latency, and trust.</p><p><br></p><p>Timestamps</p><p>[00:45] – Rethinking experimentation: “earn the right to A/B test” with staged rigor  </p><p>[03:31] – Pricing and packaging shift: from forms to flows, ASP up 32%, annual attach up 25%  </p><p>[05:34] – Typeform AI as a co-pilot: doubling activation and lifting one-day conversion  </p><p>[07:27] – Adoption lessons: elevating AI CTAs and reducing friction to use  </p><p>[10:02] – Behind the scenes: model selection, quality bars, and why MVP can backfire in AI  </p><p>[12:40] – Moving video down-tier: demand signals, cannibalization checks, and net gains  </p><p>[14:52] – Video vs. standard forms: 14x more words, 10x fewer skips, similar completion time  </p><p>[20:22] – Building an experimentation culture: process resistance, “anti-knowledge,” and cross-functional review  </p><p>[30:16] – Leader playbook: visibility, empathy, and incentives for rigorous testing</p><p><br></p><p>Takeaways</p><p>- Implement a staged experimentation funnel—discovery, simulation, then customer A/B—to reduce risk.  </p><p>- Use pricing experiments to trade volume for revenue quality; pair higher monthly prices with stronger annual discounts to grow annual attach.  </p><p>- Treat AI as an activation lever: elevate AI-first CTAs and streamline onboarding to boost adoption.  </p><p>- Add video interviewer options to increase response richness (14x more words) while keeping completion rates steady.  </p><p>- Enforce experiment hygiene: change one variable at a time, randomize at the right unit (account vs. user), and run long enough for effect size.  </p><p>- Purge “anti-knowledge” by standardizing design, instituting cross-functional reviews, and only codifying learnings supported by repeatable data.</p><p><br></p>]]>
      </description>
      <content:encoded>
        <![CDATA[<p>If 80% of A/B tests fail, how do you de-risk decisions that touch pricing, product, and brand? Aleksandra (Aleks) Bass, Chief Product &amp; Technology Officer at Typeform, shares how her team “earns the right to A/B test” with medical-grade rigor—moving from literature reviews and user tests to simulated trials before exposing changes to customers. She details Typeform’s repositioning from “forms” to an AI engagement platform—and the pricing and packaging bet behind it: a 15% drop in new business count offset by a 32% increase in ASP and a 25% lift in annual attach. </p><p>Aleks unpacks how Typeform AI acts as a co-pilot that doubled activation and boosted one-day conversion, plus the design shifts (CTA altitude and onboarding) that increased adoption. She also reveals why they moved video features down-tier and what a head-to-head test showed: video interviewers generated 14x more words and 10x fewer skipped questions with comparable completion time. </p><p>Finally, Aleks breaks down the cultural side—eliminating “anti-knowledge,” standardizing experiment design, and creating a cross-functional review that prevents false learnings—along with how her data engineering team evaluates LLMs for quality, latency, and trust.</p><p><br></p><p>Timestamps</p><p>[00:45] – Rethinking experimentation: “earn the right to A/B test” with staged rigor  </p><p>[03:31] – Pricing and packaging shift: from forms to flows, ASP up 32%, annual attach up 25%  </p><p>[05:34] – Typeform AI as a co-pilot: doubling activation and lifting one-day conversion  </p><p>[07:27] – Adoption lessons: elevating AI CTAs and reducing friction to use  </p><p>[10:02] – Behind the scenes: model selection, quality bars, and why MVP can backfire in AI  </p><p>[12:40] – Moving video down-tier: demand signals, cannibalization checks, and net gains  </p><p>[14:52] – Video vs. standard forms: 14x more words, 10x fewer skips, similar completion time  </p><p>[20:22] – Building an experimentation culture: process resistance, “anti-knowledge,” and cross-functional review  </p><p>[30:16] – Leader playbook: visibility, empathy, and incentives for rigorous testing</p><p><br></p><p>Takeaways</p><p>- Implement a staged experimentation funnel—discovery, simulation, then customer A/B—to reduce risk.  </p><p>- Use pricing experiments to trade volume for revenue quality; pair higher monthly prices with stronger annual discounts to grow annual attach.  </p><p>- Treat AI as an activation lever: elevate AI-first CTAs and streamline onboarding to boost adoption.  </p><p>- Add video interviewer options to increase response richness (14x more words) while keeping completion rates steady.  </p><p>- Enforce experiment hygiene: change one variable at a time, randomize at the right unit (account vs. user), and run long enough for effect size.  </p><p>- Purge “anti-knowledge” by standardizing design, instituting cross-functional reviews, and only codifying learnings supported by repeatable data.</p><p><br></p>]]>
      </content:encoded>
      <pubDate>Tue, 24 Feb 2026 11:03:37 -0700</pubDate>
      <author>Ashley Stirrup</author>
      <enclosure url="https://media.transistor.fm/bd2615c2/3ee4fe19.mp3" length="30641157" type="audio/mpeg"/>
      <itunes:author>Ashley Stirrup</itunes:author>
      <itunes:image href="https://img.transistorcdn.com/UZhp9goFTvSyc1XTO5WXWcFsjet5PU9wOU_jC80Wh9A/rs:fill:0:0:1/w:1400/h:1400/q:60/mb:500000/aHR0cHM6Ly9pbWct/dXBsb2FkLXByb2R1/Y3Rpb24udHJhbnNp/c3Rvci5mbS80ZmE4/OGM2MDAyNDU4MjJl/NDEwNTgzN2RmNTM1/MDJkOS5qcGc.jpg"/>
      <itunes:duration>1915</itunes:duration>
      <itunes:summary>If 80% of A/B tests fail, how do you de-risk decisions that touch pricing, product, and brand? Aleksandra (Aleks) Bass, Chief Product &amp;amp; Technology Officer at Typeform, shares how her team “earns the right to A/B test” with medical-grade rigor—moving from literature reviews and user tests to simulated trials before exposing changes to customers. She details Typeform’s repositioning from “forms” to an AI engagement platform—and the pricing and packaging bet behind it: a 15% drop in new business count offset by a 32% increase in ASP and a 25% lift in annual attach. Aleks unpacks how Typeform AI acts as a co-pilot that doubled activation and boosted one-day conversion, plus the design shifts (CTA altitude and onboarding) that increased adoption. She also reveals why they moved video features down-tier and what a head-to-head test showed: video interviewers generated 14x more words and 10x fewer skipped questions with comparable completion time. Finally, Aleks breaks down the cultural side—eliminating “anti-knowledge,” standardizing experiment design, and creating a cross-functional review that prevents false learnings—along with how her data engineering team evaluates LLMs for quality, latency, and trust.
Timestamps[00:45] – Rethinking experimentation: “earn the right to A/B test” with staged rigor  [03:31] – Pricing and packaging shift: from forms to flows, ASP up 32%, annual attach up 25%  [05:34] – Typeform AI as a co-pilot: doubling activation and lifting one-day conversion  [07:27] – Adoption lessons: elevating AI CTAs and reducing friction to use  [10:02] – Behind the scenes: model selection, quality bars, and why MVP can backfire in AI  [12:40] – Moving video down-tier: demand signals, cannibalization checks, and net gains  [14:52] – Video vs. standard forms: 14x more words, 10x fewer skips, similar completion time  [20:22] – Building an experimentation culture: process resistance, “anti-knowledge,” and cross-functional review  [30:16] – Leader playbook: visibility, empathy, and incentives for rigorous testing
Takeaways- Implement a staged experimentation funnel—discovery, simulation, then customer A/B—to reduce risk.  - Use pricing experiments to trade volume for revenue quality; pair higher monthly prices with stronger annual discounts to grow annual attach.  - Treat AI as an activation lever: elevate AI-first CTAs and streamline onboarding to boost adoption.  - Add video interviewer options to increase response richness (14x more words) while keeping completion rates steady.  - Enforce experiment hygiene: change one variable at a time, randomize at the right unit (account vs. user), and run long enough for effect size.  - Purge “anti-knowledge” by standardizing design, instituting cross-functional reviews, and only codifying learnings supported by repeatable data.</itunes:summary>
      <itunes:subtitle>If 80% of A/B tests fail, how do you de-risk decisions that touch pricing, product, and brand? Aleksandra (Aleks) Bass, Chief Product &amp;amp; Technology Officer at Typeform, shares how her team “earns the right to A/B test” with medical-grade rigor—moving f</itunes:subtitle>
      <itunes:keywords>experimentation, A/B-testing, feature-flagging, product-experimentation, experimentation-culture, data-driven-decisions, product-development, Typeform, Alex Bass, AI-form-generation, survey-data, pricing-strategy, activation-metrics, statistical-rigor, hypothesis-testing, research-methodology, medical-experimentation-approach, user-testing, conversion-rate-optimization, annual-attach-rate, AI-adoption, form-completion, customer-satisfaction, response-quality, interview-experience, video-forms, accessibility-testing, experimentation-standards, cross-functional-teams, anti-knowledge, belief-validation, humility-in-data</itunes:keywords>
      <itunes:explicit>No</itunes:explicit>
      <podcast:transcript url="https://share.transistor.fm/s/bd2615c2/transcript.txt" type="text/plain"/>
    </item>
  </channel>
</rss>
