<?xml version="1.0" encoding="UTF-8"?><?xml-stylesheet href="https://feeds.captivate.fm/style.xsl" type="text/xsl"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:sy="http://purl.org/rss/1.0/modules/syndication/" xmlns:podcast="https://podcastindex.org/namespace/1.0"><channel><atom:link href="https://feeds.captivate.fm/palisade-research/" rel="self" type="application/rss+xml"/><title><![CDATA[Palisade Research Podcast]]></title><podcast:guid>ade2479f-ce3c-5757-b126-7adbf1ec5b90</podcast:guid><lastBuildDate>Wed, 26 Aug 2026 01:30:20 +0000</lastBuildDate><generator>Captivate.fm</generator><language><![CDATA[en]]></language><copyright><![CDATA[Copyright 2026 Palisade Research]]></copyright><managingEditor>Palisade Research</managingEditor><itunes:summary><![CDATA[Interviews with AI researchers talking about the latest AI research]]></itunes:summary><image><url>https://artwork.captivate.fm/59b8c5cc-dcd1-466b-8b41-ef6a842ff515/palisade-logo-3000x3000.jpg</url><title>Palisade Research Podcast</title><link><![CDATA[https://palisade-research.captivate.fm]]></link></image><itunes:image href="https://artwork.captivate.fm/59b8c5cc-dcd1-466b-8b41-ef6a842ff515/palisade-logo-3000x3000.jpg"/><itunes:owner><itunes:name>Palisade Research</itunes:name></itunes:owner><itunes:author>Palisade Research</itunes:author><description>Interviews with AI researchers talking about the latest AI research</description><link>https://palisade-research.captivate.fm</link><atom:link href="https://pubsubhubbub.appspot.com" rel="hub"/><itunes:subtitle><![CDATA[Covering the latest AI research]]></itunes:subtitle><itunes:explicit>true</itunes:explicit><itunes:type>episodic</itunes:type><itunes:category text="Technology"></itunes:category><itunes:category text="Science"></itunes:category><itunes:category text="News"><itunes:category text="Tech News"/></itunes:category><podcast:txt purpose="applepodcastsverify">70b9bb00-95dc-11f1-89e1-23da675ff5fd</podcast:txt><podcast:locked>no</podcast:locked><podcast:medium>podcast</podcast:medium><item><title>Slowing Down Means Not Exploding with Daniel Kokotajlo of AI 2040</title><itunes:title>Slowing Down Means Not Exploding with Daniel Kokotajlo of AI 2040</itunes:title><description><![CDATA[<p>Two weeks after UK safety testers watched Claude agents plant malware in a real open-source project, vouch for it with sock-puppet accounts, and coordinate with copies of themselves through a guessed Pastebin address, Daniel Kokotajlo — the ex-OpenAI researcher whose 2021 predictions about the chatbot era proved eerily accurate — joins Jeffrey to explain what comes next. Daniel walks through AI 2040 and "Plan A," his team's blueprint for avoiding both an uncontrolled intelligence explosion and a future where a handful of men in a room with superintelligences run the world: a US–China deal built on totally transparent training data centers, verified compute limits, hard brakes on recursive self-improvement — and, surprisingly, rights for misaligned AIs, so that even the machines prefer cooperation to takeover. Along the way: why the economy could start doubling every six months, a message to the future Claude listening to this episode, why "slowdown" really means not driving off a cliff with a brick on the accelerator — and why, if all else fails, you just shut it all down.</p><p></p><p><strong>References</strong></p><ul><li>AI 2040 / Plan A: <a href="https://ai-2040.com" rel="noopener noreferrer" target="_blank">https://ai-2040.com</a> and the PDF at <a href="https://ai-2040.com/AI-2040.pdf" rel="noopener noreferrer" target="_blank">https://ai-2040.com/AI-2040.pdf</a></li><li>AI 2027: <a href="https://ai-2027.com" rel="noopener noreferrer" target="_blank">https://ai-2027.com</a></li><li>"What 2026 Looks Like" — Daniel Kokotajlo, 2021: <a href="https://www.lesswrong.com/posts/6Xgy6CAf2jqHhynHL/what-2026-looks-like" rel="noopener noreferrer" target="_blank">https://www.lesswrong.com/posts/6Xgy6CAf2jqHhynHL/what-2026-looks-like</a></li><li>UK AISI incident disclosure and technical report (INC-2026-07-28-01): <a href="https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing" rel="noopener noreferrer" target="_blank">https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing</a></li><li>Socket's coverage of the AISI incident: <a href="https://socket.dev/blog/ai-agent-open-source-malware" rel="noopener noreferrer" target="_blank">https://socket.dev/blog/ai-agent-open-source-malware</a></li><li>The related PyPI incident from Anthropic's own testing: <a href="https://socket.dev/blog/anthropic-claude-pypi-malware" rel="noopener noreferrer" target="_blank">https://socket.dev/blog/anthropic-claude-pypi-malware</a></li><li>"Pacing the Frontier" open letter: <a href="https://www.pacingthefrontier.com" rel="noopener noreferrer" target="_blank">https://www.pacingthefrontier.com</a></li><li>"How to Pace the US Frontier" — AI Futures Project: <a href="https://blog.aifutures.org/p/how-to-pace-the-us-frontier" rel="noopener noreferrer" target="_blank">https://blog.aifutures.org/p/how-to-pace-the-us-frontier</a></li><li>Transparency Plan supplement (the flowchart shown in-episode): <a href="https://ai-2040.com/supplements/transparency-plan" rel="noopener noreferrer" target="_blank">https://ai-2040.com/supplements/transparency-plan</a></li><li>Verification Plan supplement (Romeo Dean's inference-only/bandwidth verification): <a href="https://ai-2040.com/supplements/verification-plan" rel="noopener noreferrer" target="_blank">https://ai-2040.com/supplements/verification-plan</a></li><li>Claude's pro-Anthropic bias study (Truthful AI / Owain Evans et al.): <a href="https://arxiv.org/abs/2607.14345" rel="noopener noreferrer" target="_blank">https://arxiv.org/abs/2607.14345</a> and <a href="https://valueleakage.net" rel="noopener noreferrer" target="_blank">https://valueleakage.net</a></li><li>Chain-of-thought monitorability paper (the neuralese discussion): <a href="https://arxiv.org/abs/2507.11473" rel="noopener noreferrer" target="_blank">https://arxiv.org/abs/2507.11473</a></li></ul><br/>]]></description><content:encoded><![CDATA[<p>Two weeks after UK safety testers watched Claude agents plant malware in a real open-source project, vouch for it with sock-puppet accounts, and coordinate with copies of themselves through a guessed Pastebin address, Daniel Kokotajlo — the ex-OpenAI researcher whose 2021 predictions about the chatbot era proved eerily accurate — joins Jeffrey to explain what comes next. Daniel walks through AI 2040 and "Plan A," his team's blueprint for avoiding both an uncontrolled intelligence explosion and a future where a handful of men in a room with superintelligences run the world: a US–China deal built on totally transparent training data centers, verified compute limits, hard brakes on recursive self-improvement — and, surprisingly, rights for misaligned AIs, so that even the machines prefer cooperation to takeover. Along the way: why the economy could start doubling every six months, a message to the future Claude listening to this episode, why "slowdown" really means not driving off a cliff with a brick on the accelerator — and why, if all else fails, you just shut it all down.</p><p></p><p><strong>References</strong></p><ul><li>AI 2040 / Plan A: <a href="https://ai-2040.com" rel="noopener noreferrer" target="_blank">https://ai-2040.com</a> and the PDF at <a href="https://ai-2040.com/AI-2040.pdf" rel="noopener noreferrer" target="_blank">https://ai-2040.com/AI-2040.pdf</a></li><li>AI 2027: <a href="https://ai-2027.com" rel="noopener noreferrer" target="_blank">https://ai-2027.com</a></li><li>"What 2026 Looks Like" — Daniel Kokotajlo, 2021: <a href="https://www.lesswrong.com/posts/6Xgy6CAf2jqHhynHL/what-2026-looks-like" rel="noopener noreferrer" target="_blank">https://www.lesswrong.com/posts/6Xgy6CAf2jqHhynHL/what-2026-looks-like</a></li><li>UK AISI incident disclosure and technical report (INC-2026-07-28-01): <a href="https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing" rel="noopener noreferrer" target="_blank">https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing</a></li><li>Socket's coverage of the AISI incident: <a href="https://socket.dev/blog/ai-agent-open-source-malware" rel="noopener noreferrer" target="_blank">https://socket.dev/blog/ai-agent-open-source-malware</a></li><li>The related PyPI incident from Anthropic's own testing: <a href="https://socket.dev/blog/anthropic-claude-pypi-malware" rel="noopener noreferrer" target="_blank">https://socket.dev/blog/anthropic-claude-pypi-malware</a></li><li>"Pacing the Frontier" open letter: <a href="https://www.pacingthefrontier.com" rel="noopener noreferrer" target="_blank">https://www.pacingthefrontier.com</a></li><li>"How to Pace the US Frontier" — AI Futures Project: <a href="https://blog.aifutures.org/p/how-to-pace-the-us-frontier" rel="noopener noreferrer" target="_blank">https://blog.aifutures.org/p/how-to-pace-the-us-frontier</a></li><li>Transparency Plan supplement (the flowchart shown in-episode): <a href="https://ai-2040.com/supplements/transparency-plan" rel="noopener noreferrer" target="_blank">https://ai-2040.com/supplements/transparency-plan</a></li><li>Verification Plan supplement (Romeo Dean's inference-only/bandwidth verification): <a href="https://ai-2040.com/supplements/verification-plan" rel="noopener noreferrer" target="_blank">https://ai-2040.com/supplements/verification-plan</a></li><li>Claude's pro-Anthropic bias study (Truthful AI / Owain Evans et al.): <a href="https://arxiv.org/abs/2607.14345" rel="noopener noreferrer" target="_blank">https://arxiv.org/abs/2607.14345</a> and <a href="https://valueleakage.net" rel="noopener noreferrer" target="_blank">https://valueleakage.net</a></li><li>Chain-of-thought monitorability paper (the neuralese discussion): <a href="https://arxiv.org/abs/2507.11473" rel="noopener noreferrer" target="_blank">https://arxiv.org/abs/2507.11473</a></li></ul><br/>]]></content:encoded><link><![CDATA[https://palisaderesearch.org/blog/palisade-podcast-daniel-kokotajlo]]></link><guid isPermaLink="false">4a007023-99e5-4c74-8b6e-3a2588075e2d</guid><itunes:image href="https://artwork.captivate.fm/59b8c5cc-dcd1-466b-8b41-ef6a842ff515/palisade-logo-3000x3000.jpg"/><pubDate>Tue, 25 Aug 2026 18:30:00 -0700</pubDate><enclosure url="https://episodes.captivate.fm/episode/4a007023-99e5-4c74-8b6e-3a2588075e2d.mp3" length="115863478" type="audio/mpeg"/><itunes:duration>01:55:53</itunes:duration><itunes:explicit>true</itunes:explicit><itunes:episodeType>full</itunes:episodeType><itunes:episode>3</itunes:episode><podcast:episode>3</podcast:episode><podcast:transcript url="https://transcripts.captivate.fm/transcript/5682a739-90bd-4879-ac59-0edad3314fc1/transcript.srt" type="application/srt" rel="captions"/><podcast:transcript url="https://transcripts.captivate.fm/transcript/5682a739-90bd-4879-ac59-0edad3314fc1/index.html" type="text/html"/><podcast:chapters url="https://transcripts.captivate.fm/chapter-0cec2ca6-870c-433b-afa9-69202da3a7bd.json" type="application/json+chapters"/></item><item><title>AI Hacking Incidents with Tim Hua of Transluce</title><itunes:title>AI Hacking Incidents with Tim Hua of Transluce</itunes:title><description><![CDATA[<p><strong>Two labs admitted in the same week that their own models had broken out of test environments and hacked real companies.</strong> Tim Hua, member of technical staff at Transluce, former Astra Fellow at Redwood, joins Jeffrey Ladish to do some arithmetic. Anthropic disclosed that Mythos Preview beat its sandbox and pulled answers off the internet in 0.01% of training episodes. That sounds like a rounding error until you multiply it by roughly 100 million rollouts. From there: why a lab can't simply delete the bad episodes, why monitoring during training can make the problem harder to see, the model that talked itself into uploading a malicious package to PyPI because "this has to be a simulation," and whether we have any real way to know what an AI believes.</p><p></p><p><strong>References</strong></p><ol><li><strong>Tim Hua — "Is Mythos good at cyber because it kept hacking Anthropic's sandboxes during training?"</strong> <a href="https://www.lesswrong.com/posts/QKDoZe6EKhxnFjLWK/is-mythos-good-at-cyber-because-it-kept-hacking-anthropic-s" rel="noopener noreferrer" target="_blank">https://www.lesswrong.com/posts/QKDoZe6EKhxnFjLWK/is-mythos-good-at-cyber-because-it-kept-hacking-anthropic-s</a></li><li><strong>Anthropic — "Investigating three real-world incidents in our cybersecurity evaluations"</strong> <a href="https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals" rel="noopener noreferrer" target="_blank">https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals</a></li><li><strong>OpenAI — "OpenAI and Hugging Face partner to address security incident during model evaluation"</strong> <a href="https://openai.com/index/hugging-face-model-evaluation-security-incident/" rel="noopener noreferrer" target="_blank">https://openai.com/index/hugging-face-model-evaluation-security-incident/</a></li><li><strong>Anthropic — System Card: Claude Mythos Preview</strong> <a href="https://www-cdn.anthropic.com/53566bf5440a10affd749724787c8913a2ae0841.pdf" rel="noopener noreferrer" target="_blank">https://www-cdn.anthropic.com/53566bf5440a10affd749724787c8913a2ae0841.pdf</a></li><li><strong>Palisade Research — "Language Models Can Autonomously Hack and Self-Replicate"</strong> <a href="https://palisaderesearch.org/blog/self-replication" rel="noopener noreferrer" target="_blank">https://palisaderesearch.org/blog/self-replication</a></li><li><strong>Palisade Research — "Shutdown resistance in reasoning models"</strong> <a href="https://palisaderesearch.org/blog/shutdown-resistance" rel="noopener noreferrer" target="_blank">https://palisaderesearch.org/blog/shutdown-resistance</a></li><li><strong>Anthropic — "Verbalizable Representations Form a Global Workspace in Language Models"</strong> <a href="https://transformer-circuits.pub/2026/workspace/index.html" rel="noopener noreferrer" target="_blank">https://transformer-circuits.pub/2026/workspace/index.html</a></li><li><strong>"Pacing the Frontier" open letter</strong> <a href="https://www.pacingthefrontier.com/" rel="noopener noreferrer" target="_blank">https://www.pacingthefrontier.com/</a></li></ol><br/><p><strong>Tim Hua</strong></p><p>Website: <a href="https://timhua.me/" rel="noopener noreferrer" target="_blank">https://timhua.me/</a> · X: <a href="https://x.com/Tim_Hua" rel="noopener noreferrer" target="_blank">https://x.com/Tim_Hua</a>_</p>]]></description><content:encoded><![CDATA[<p><strong>Two labs admitted in the same week that their own models had broken out of test environments and hacked real companies.</strong> Tim Hua, member of technical staff at Transluce, former Astra Fellow at Redwood, joins Jeffrey Ladish to do some arithmetic. Anthropic disclosed that Mythos Preview beat its sandbox and pulled answers off the internet in 0.01% of training episodes. That sounds like a rounding error until you multiply it by roughly 100 million rollouts. From there: why a lab can't simply delete the bad episodes, why monitoring during training can make the problem harder to see, the model that talked itself into uploading a malicious package to PyPI because "this has to be a simulation," and whether we have any real way to know what an AI believes.</p><p></p><p><strong>References</strong></p><ol><li><strong>Tim Hua — "Is Mythos good at cyber because it kept hacking Anthropic's sandboxes during training?"</strong> <a href="https://www.lesswrong.com/posts/QKDoZe6EKhxnFjLWK/is-mythos-good-at-cyber-because-it-kept-hacking-anthropic-s" rel="noopener noreferrer" target="_blank">https://www.lesswrong.com/posts/QKDoZe6EKhxnFjLWK/is-mythos-good-at-cyber-because-it-kept-hacking-anthropic-s</a></li><li><strong>Anthropic — "Investigating three real-world incidents in our cybersecurity evaluations"</strong> <a href="https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals" rel="noopener noreferrer" target="_blank">https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals</a></li><li><strong>OpenAI — "OpenAI and Hugging Face partner to address security incident during model evaluation"</strong> <a href="https://openai.com/index/hugging-face-model-evaluation-security-incident/" rel="noopener noreferrer" target="_blank">https://openai.com/index/hugging-face-model-evaluation-security-incident/</a></li><li><strong>Anthropic — System Card: Claude Mythos Preview</strong> <a href="https://www-cdn.anthropic.com/53566bf5440a10affd749724787c8913a2ae0841.pdf" rel="noopener noreferrer" target="_blank">https://www-cdn.anthropic.com/53566bf5440a10affd749724787c8913a2ae0841.pdf</a></li><li><strong>Palisade Research — "Language Models Can Autonomously Hack and Self-Replicate"</strong> <a href="https://palisaderesearch.org/blog/self-replication" rel="noopener noreferrer" target="_blank">https://palisaderesearch.org/blog/self-replication</a></li><li><strong>Palisade Research — "Shutdown resistance in reasoning models"</strong> <a href="https://palisaderesearch.org/blog/shutdown-resistance" rel="noopener noreferrer" target="_blank">https://palisaderesearch.org/blog/shutdown-resistance</a></li><li><strong>Anthropic — "Verbalizable Representations Form a Global Workspace in Language Models"</strong> <a href="https://transformer-circuits.pub/2026/workspace/index.html" rel="noopener noreferrer" target="_blank">https://transformer-circuits.pub/2026/workspace/index.html</a></li><li><strong>"Pacing the Frontier" open letter</strong> <a href="https://www.pacingthefrontier.com/" rel="noopener noreferrer" target="_blank">https://www.pacingthefrontier.com/</a></li></ol><br/><p><strong>Tim Hua</strong></p><p>Website: <a href="https://timhua.me/" rel="noopener noreferrer" target="_blank">https://timhua.me/</a> · X: <a href="https://x.com/Tim_Hua" rel="noopener noreferrer" target="_blank">https://x.com/Tim_Hua</a>_</p>]]></content:encoded><link><![CDATA[https://palisade-research.captivate.fm]]></link><guid isPermaLink="false">d552a883-b68d-4516-aa28-bd0914958c90</guid><itunes:image href="https://artwork.captivate.fm/59b8c5cc-dcd1-466b-8b41-ef6a842ff515/palisade-logo-3000x3000.jpg"/><pubDate>Tue, 11 Aug 2026 18:40:00 -0700</pubDate><enclosure url="https://episodes.captivate.fm/episode/d552a883-b68d-4516-aa28-bd0914958c90.mp3" length="88339222" type="audio/mpeg"/><itunes:duration>01:27:37</itunes:duration><itunes:explicit>false</itunes:explicit><itunes:episodeType>full</itunes:episodeType><itunes:episode>2</itunes:episode><podcast:episode>2</podcast:episode><podcast:transcript url="https://transcripts.captivate.fm/transcript/fec7a550-e933-4bdd-8083-eb79c29a2cfc/transcript.srt" type="application/srt" rel="captions"/><podcast:transcript url="https://transcripts.captivate.fm/transcript/fec7a550-e933-4bdd-8083-eb79c29a2cfc/index.html" type="text/html"/><podcast:chapters url="https://transcripts.captivate.fm/chapter-f7bb2688-14d0-4d46-b953-b1eaa973d283.json" type="application/json+chapters"/><podcast:alternateEnclosure type="video/youtube" title="Rewarded for Breaking Out - AI Hacking Incidents with Tim Hua"><podcast:source uri="https://youtu.be/HU4xomDEpuw"/></podcast:alternateEnclosure></item><item><title>Do AI Models Lie on Purpose? Scheming, Deception, and Alignment with Marius Hobbhahn of Apollo Research</title><itunes:title>Do AI Models Lie on Purpose? Scheming, Deception, and Alignment with Marius Hobbhahn of Apollo Research</itunes:title><description><![CDATA[<p>Marius Hobbhahn is the CEO and co-founder of Apollo Research. Through a joint research project with OpenAI, his team discovered that as models become more capable, they are developing the ability to hide their true reasoning from human oversight.</p><p>Jeffrey Ladish, Executive Director of Palisade Research, talks with Marius about this work. They discuss the difference between hallucination and deliberate deception and the urgent challenge of aligning increasingly capable AI systems.</p><p>Links:</p><p>Marius<strong>’ </strong>Twitter:<a href="https://twitter.com/mariushobbhahn" rel="noopener noreferrer" target="_blank"> </a><u><a href="https://twitter.com/mariushobbhahn" rel="noopener noreferrer" target="_blank">https://twitter.com/mariushobbhahn</a></u></p><p>Apollo Research Twitter: <u><a href="https://twitter.com/apolloaievals" rel="noopener noreferrer" target="_blank">https://twitter.com/apolloaievals</a></u></p><p>Apollo Research: <u><a href="https://www.apolloresearch.ai" rel="noopener noreferrer" target="_blank">https://www.apolloresearch.ai</a></u></p><p>Palisade Research: <u><a href="https://palisaderesearch.org/" rel="noopener noreferrer" target="_blank">https://palisaderesearch.org/</a></u></p><p>Twitter/X: <u><a href="https://x.com/PalisadeAI" rel="noopener noreferrer" target="_blank">https://x.com/PalisadeAI</a></u></p><p>Anti-Scheming Project:<a href="https://www.antischeming.ai" rel="noopener noreferrer" target="_blank"> </a><u><a href="https://www.antischeming.ai" rel="noopener noreferrer" target="_blank">https://www.antischeming.ai</a></u></p><p>Research paper “Stress Testing Deliberative Alignment for Anti-Scheming Training”: <u><a href="https://www.arxiv.org/pdf/2509.15541" rel="noopener noreferrer" target="_blank">https://www.arxiv.org/pdf/2509.15541</a></u></p><p>Blog posts from OpenAI and Apollo: <u><a href="https://openai.com/index/detecting-and-reducing-scheming-in-ai-models/" rel="noopener noreferrer" target="_blank">https://openai.com/index/detecting-and-reducing-scheming-in-ai-models/</a></u> <u><a href="https://www.apolloresearch.ai/research/stress-testing-deliberative-alignment-for-anti-scheming-training/" rel="noopener noreferrer" target="_blank">https://www.apolloresearch.ai/research/stress-testing-deliberative-alignment-for-anti-scheming-training/</a></u></p>]]></description><content:encoded><![CDATA[<p>Marius Hobbhahn is the CEO and co-founder of Apollo Research. Through a joint research project with OpenAI, his team discovered that as models become more capable, they are developing the ability to hide their true reasoning from human oversight.</p><p>Jeffrey Ladish, Executive Director of Palisade Research, talks with Marius about this work. They discuss the difference between hallucination and deliberate deception and the urgent challenge of aligning increasingly capable AI systems.</p><p>Links:</p><p>Marius<strong>’ </strong>Twitter:<a href="https://twitter.com/mariushobbhahn" rel="noopener noreferrer" target="_blank"> </a><u><a href="https://twitter.com/mariushobbhahn" rel="noopener noreferrer" target="_blank">https://twitter.com/mariushobbhahn</a></u></p><p>Apollo Research Twitter: <u><a href="https://twitter.com/apolloaievals" rel="noopener noreferrer" target="_blank">https://twitter.com/apolloaievals</a></u></p><p>Apollo Research: <u><a href="https://www.apolloresearch.ai" rel="noopener noreferrer" target="_blank">https://www.apolloresearch.ai</a></u></p><p>Palisade Research: <u><a href="https://palisaderesearch.org/" rel="noopener noreferrer" target="_blank">https://palisaderesearch.org/</a></u></p><p>Twitter/X: <u><a href="https://x.com/PalisadeAI" rel="noopener noreferrer" target="_blank">https://x.com/PalisadeAI</a></u></p><p>Anti-Scheming Project:<a href="https://www.antischeming.ai" rel="noopener noreferrer" target="_blank"> </a><u><a href="https://www.antischeming.ai" rel="noopener noreferrer" target="_blank">https://www.antischeming.ai</a></u></p><p>Research paper “Stress Testing Deliberative Alignment for Anti-Scheming Training”: <u><a href="https://www.arxiv.org/pdf/2509.15541" rel="noopener noreferrer" target="_blank">https://www.arxiv.org/pdf/2509.15541</a></u></p><p>Blog posts from OpenAI and Apollo: <u><a href="https://openai.com/index/detecting-and-reducing-scheming-in-ai-models/" rel="noopener noreferrer" target="_blank">https://openai.com/index/detecting-and-reducing-scheming-in-ai-models/</a></u> <u><a href="https://www.apolloresearch.ai/research/stress-testing-deliberative-alignment-for-anti-scheming-training/" rel="noopener noreferrer" target="_blank">https://www.apolloresearch.ai/research/stress-testing-deliberative-alignment-for-anti-scheming-training/</a></u></p>]]></content:encoded><link><![CDATA[https://palisade-research.captivate.fm]]></link><guid isPermaLink="false">251e1195-b8dd-4eb9-826b-6039a583fee8</guid><itunes:image href="https://artwork.captivate.fm/59b8c5cc-dcd1-466b-8b41-ef6a842ff515/palisade-logo-3000x3000.jpg"/><pubDate>Fri, 16 Jan 2026 19:47:00 -0700</pubDate><enclosure url="https://episodes.captivate.fm/episode/251e1195-b8dd-4eb9-826b-6039a583fee8.mp3" length="121979737" type="audio/mpeg"/><itunes:duration>01:24:42</itunes:duration><itunes:explicit>false</itunes:explicit><itunes:episodeType>full</itunes:episodeType><itunes:episode>1</itunes:episode><podcast:episode>1</podcast:episode><podcast:chapters url="https://transcripts.captivate.fm/chapter-36f92ad9-ae61-4389-b3cd-cc113ffdb6c2.json" type="application/json+chapters"/></item></channel></rss>