` 深度解析:概率动作列表、随机性来源与采样实战)
人工智能强化学习深度学习【免费下载链接】open_spielOpenSpiel is a collection of environments and algorithms for research in general reinforcement learning and search/planning in games.项目地址https://gitcode.com/gh_mirrors/op/open_spiel点击查看免费下载导读本文围绕 OpenSpiel 核心 State API 中的chance_outcomes()方法展开讲解如何获取随机节点Chance Node的概率分布、理解is_chance_node()与其配合的判定逻辑并通过 Leduc Poker、外部采样 MCCFR、CFR 等真实用例展示如何读取、校验与采样这一分布。读完本文你将掌握在 OpenSpielPython 与 C中处理随机性来源的标准姿势以及kExplicitStochastic/kSampledStochastic两种随机模式对chance_outcomes()行为的影响。1.chance_outcomes()是什么chance_outcomes()是 State 类提供的核心方法之一返回一个(action, probability)二元组列表表示当前状态下的随机动作及其概率分布。其 C 签名在 spiel.h 中定义// Get the chance outcomes and their probabilities. // // Chance actions do not have a separate UID space from regular actions. // // Note: what is returned here depending on the games chance_mode (in // its GameType): // - Option 1. kExplicit. All chance node outcomes are returned along with // their respective probabilities. Then State::ApplyAction(...) is // deterministic. // - Option 2. kSampled. Return a dummy single action here with probability // 1, and then State::ApplyAction(...) does the real sampling. In this // case, the game has to maintain its own RNG. virtual ActionsAndProbs ChanceOutcomes() const { SpielFatalError(ChanceOutcomes unimplemented!); }其中ActionsAndProbs即std::vectorstd::pairAction, double。几个关键语义返回值(动作 ID, 概率)列表概率之和应为 1合法离散分布只在随机节点调用才有意义非随机节点调用会触发SpielFatalError(ChanceOutcomes unimplemented!)或返回空/未定义结果因此调用前务必用is_chance_node()判定动作 ID 与普通动作共享编号空间随机动作没有独立 UID 空间见 spiel.h这保证了整个博弈树中动作编号的统一性随机性来源由chance_mode决定若游戏声明为kExplicitStochasticchance_outcomes()返回完整分布且ApplyAction是确定性的若为kSampledStochastic则只返回一个概率为 1 的占位动作真正的采样在ApplyAction内部完成需要游戏自身维护 RNG。ChanceMode枚举定义在 spiel.henum class ChanceMode { kDeterministic, // No chance nodes kExplicitStochastic, // Has at least one chance node, all with // deterministic ApplyAction() kSampledStochastic, // At least one chance node with non-deterministic // ApplyAction() };源码注释spiel.h建议实现随机游戏时优先采用kExplicitStochastic因为学习算法能看到完整的随机结果分布而不只是单个采样结果。2. Python 中的典型用法以 Leduc Poker 为例在 Python 中通过pyspiel访问该方法其绑定位于 open_spiel/python/pybind11/pyspiel.cc.def(chance_outcomes, State::ChanceOutcomes)。Leduc Poker 是验证chance_outcomes()的经典游戏发牌由随机节点完成每轮从牌堆中抽取手牌。文档示例state_chance_outcomes.md完整复现如下import pyspiel import numpy as np game pyspiel.load_game(leduc_poker) state game.new_initial_state() # First players private card. print(state.chance_outcomes()) # Output: # [(0, 0.16666666666666666), (1, 0.16666666666666666), (2, 0.16666666666666666), (3, 0.16666666666666666), (4, 0.16666666666666666), (5, 0.16666666666666666)] state.apply_action(0) # Second players private card. outcomes state.chance_outcomes() print() # Output: # [(1, 0.2), (2, 0.2), (3, 0.2), (4, 0.2), (5, 0.2)] # Sampling an outcome and applying it. action_list, prob_list zip(*outcomes) action np.random.choice(action_list, pprob_list) state.apply_action(action)这段代码展示了两个关键观察概率随牌堆变化初始 6 张牌每张概率1/6 ≈ 0.1667第一位玩家抽走动作0后剩余 5 张牌每张概率变为1/5 0.2且动作0不再出现——chance_outcomes()返回的是当前状态下真实的剩余牌分布不是固定常量手动采样 应用把(action, probability)拆成两个列表用np.random.choice(action_list, pprob_list)按概率采样再apply_action(action)推进状态。2.1 源码印证Leduc 发牌分布的生成逻辑Leduc 的ChanceOutcomes()实现在 open_spiel/games/leduc_poker/leduc_poker.ccstd::vectorstd::pairAction, double LeducState::ChanceOutcomes() const { SPIEL_CHECK_TRUE(IsChanceNode()); std::vectorstd::pairAction, double outcomes; if (suit_isomorphism_) { const double p 1.0 / deck_size_; // Consecutive cards in deck are viewed identically. for (int card 0; card deck_.size() / 2; card) { if (deck_[card * 2] ! kInvalidCard deck_[card * 2 1] ! kInvalidCard) { outcomes.push_back({card, p * 2}); } else if (deck_[card * 2] ! kInvalidCard || deck_[card * 2 1] ! kInvalidCard) { outcomes.push_back({card, p}); } } return outcomes; } const double p 1.0 / deck_size_; for (int card 0; card deck_.size(); card) { // This card is still in the deck, prob is 1/decksize. if (deck_[card] ! kInvalidCard) outcomes.push_back({card, p}); } return outcomes; }实现细节值得注意第一行SPIEL_CHECK_TRUE(IsChanceNode())明确要求该方法只能在随机节点被调用这正是 Python 示例中初始状态发牌节点可调用的原因非同构模式下概率恒为1 / deck_size_deck_size_是剩余牌数见 leduc_poker.h 注释 Number of cards remaining与示例中 1/6、1/5 的输出完全吻合同构suit_isomorphism_模式下花色被折叠连续两张花色牌视为同一张概率相应翻倍p * 2这是可配置游戏参数引起的分布差异。3. 配合is_chance_node()使用先判定再取值由于chance_outcomes()只在随机节点有定义实践中总是先调用is_chance_node()判断当前状态是否为随机节点。该方法的 C 默认实现在 spiel.h// Is this state a chance node? Chance nodes are states whose actions // represent stochastic outcomes. Chance or Nature is thought of as a // player with a fixed (randomized) policy. virtual bool IsChanceNode() const { return CurrentPlayer() kChancePlayerId; }Python 绑定同样位于 pyspiel.cc.def(is_chance_node, State::IsChanceNode)。kChancePlayerId是 OpenSpiel 中自然/环境玩家的特殊 IDkChancePlayerId -1一类的保留值。从语义上讲随机性被抽象为一名使用固定随机策略的玩家这使得博弈树可以统一地按玩家轮流行动的框架处理随机节点只是行动者恰为自然。is_chance_node()的判定测试state_is_chance_node.mdimport pyspiel game pyspiel.load_game(tic_tac_toe) state game.new_initial_state() print(state.is_chance_node()) # Output: False game pyspiel.load_game(leduc_poker) state game.new_initial_state() print(state.is_chance_node()) # Output: True game pyspiel.load_game(matrix_sh) state game.new_initial_state() print(state.is_chance_node()) # Output: Falsetic_tac_toe完全确定性的完美信息游戏初始状态是玩家 0 的决策节点返回Falseleduc_poker初始状态就是发牌随机节点返回Truematrix_sh单回合同时行动的矩阵博弈没有随机节点返回False。这种先is_chance_node()判定、再chance_outcomes()取值、最后apply_action()推进的三步模式是遍历含随机性博弈树的标准套路。LegalActions()的默认实现甚至直接复用了这一判定逻辑spiel.hvirtual std::vectorAction LegalActions(Player player) const { if (!IsTerminal() player CurrentPlayer()) { return IsChanceNode() ? LegalChanceOutcomes() : LegalActions(); } else { return {}; } }即在随机节点上合法动作列表就是合法随机结果LegalChanceOutcomes()默认由ChanceOutcomes()取出所有动作 ID 构成见 spiel.h。因此state.legal_actions()在随机节点同样可用返回的正是可抽到的全部结果。4. 深入OpenSpiel 内部如何采样随机结果4.1 C 侧SampleAction从随机分布中抽取一个结果OpenSpiel 提供全局函数SampleAction声明于 spiel.h实现于 spiel.ccstd::pairAction, double SampleAction(const ActionsAndProbs outcomes, double z) { SPIEL_CHECK_GE(z, 0); SPIEL_CHECK_LT(z, 1); // Special case for one-item lists. if (outcomes.size() 1) { SPIEL_CHECK_FLOAT_EQ(outcomes[0].second, 1.0); return outcomes[0]; } // First do a check that this is indeed a proper discrete distribution. double sum 0; for (const std::pairAction, double outcome : outcomes) { double prob outcome.second; SPIEL_CHECK_PROB(prob); sum prob; } SPIEL_CHECK_FLOAT_EQ(sum, 1.0); // Now sample an outcome. sum 0; for (const std::pairAction, double outcome : outcomes) { double prob outcome.second; if (sum z z (sum prob)) { return outcome; } sum prob; } SpielFatalError(Failed to sample an action, this should never happen.); }要点该函数是累加区间采样把[0, 1)按概率切成连续区间z落在哪个区间就选哪个动作复杂度 O(n)采样前会做分布合法性校验概率必须在[0, 1]内且总和为 1SPIEL_CHECK_PROB、SPIEL_CHECK_FLOAT_EQ这提醒调用方chance_outcomes()返回的分布理论上必须归一化另有一个接受随机数生成器的重载spiel.ccSampleAction(outcomes, absl::Uniform(rng, 0.0, 1.0))把z的生成与采样封装在一起。4.2 Python 侧np.random.choice的等价替代Python 中除文档示例的np.random.choice(action_list, pprob_list)外pyspiel也直接暴露了 C 的采样能力。更贴近 OpenSpiel 内部风格的做法是import pyspiel game pyspiel.load_game(leduc_poker) state game.new_initial_state() outcomes state.chance_outcomes() actions [a for a, p in outcomes] probs [p for a, p in outcomes] assert abs(sum(probs) - 1.0) 1e-9 # 概率和应为 1 action pyspiel.sample_action(outcomes) # 等价于 C SampleAction state.apply_action(action)不过以当前仓库 pyspiel.cc 的绑定为准chance_outcomes直接返回(action, prob)列表配合zip(*outcomes)拆解、np.random.choice采样即可无需依赖额外 API。需要强调的一点是不要手动修改返回的概率值OpenSpiel 约定该分布已归一化算法代码如SampleAction会按合法离散分布处理。5. 实战价值算法中的标准消费模式chance_outcomes()并非孤立 API它是众多强化学习与博弈论算法处理随机性的枢纽。仓库中大量算法通过is_chance_node()ChanceOutcomes()SampleAction的组合遍历含随机的博弈树。5.1 CFR遍历完整随机分布CFRCounterfactual Regret Minimization在初始化信息状态节点时对每个随机节点穷举其全部随机结果并递归展开子节点open_spiel/algorithms/cfr.ccif (state.IsChanceNode()) { for (const auto action_prob : state.ChanceOutcomes()) { InitializeInfostateNodes(*state.Child(action_prob.first)); } return; }这里action_prob.first是动作 IDstate.Child(...)生成对应子状态。CFR 依赖完整分布计算期望后悔值因此要求游戏采用kExplicitStochastic模式——这也解释了 spiel.h 注释中显式随机更有利于学习算法的建议。5.2 外部采样 MCCFR按分布随机抽样与 CFR 的穷举不同外部采样 MCCFR 在随机节点按概率抽样一次open_spiel/algorithms/external_sampling_mccfr.cc} else if (state.IsChanceNode()) { Action action SampleAction(state.ChanceOutcomes(), dist_(*rng)).first;同样的模式还出现在 outcome_sampling_mccfr.cc、mcts.cc、evaluate_bots.cc 等处——MCTS 在随机节点通过SampleAction(working_state-ChanceOutcomes(), rng_)抽取环境结果mcts.cc。可见穷举分布 vs 按分布采样是确定性搜索类算法CFR、值迭代与随机采样类算法MCCFR、MCTS、DQN 环境模拟在随机节点上的核心差异而两者的共同输入都是chance_outcomes()。5.3 智能体评估统一的环境随机性接口open_spiel/algorithms/evaluate_bots.cc 中评估器在随机节点同样调用SampleAction(state-ChanceOutcomes(), rng)保证所有被评估的 bot 面对一致的环境随机性接口。6. 常见问题与注意事项在非随机节点调用chance_outcomes()C 侧 Leduc 实现直接SPIEL_CHECK_TRUE(IsChanceNode())断言多数游戏会崩溃/报错调用前务必用is_chance_node()判定或用state.get_type()返回StateType::kChance确认节点类型参见 state_get_type.md 中对 Chance 类型的说明。kSampledStochastic模式的假分布若游戏声明该模式chance_outcomes()只返回一个概率为 1 的占位动作真实随机性藏在ApplyAction内部。此时不要依赖该分布做概率计算从源码角度看这种模式下的状态序列化也存在限制spiel.h 注释明确说明默认序列化方案对kSampledStochastic游戏不适用。概率浮点误差返回概率是double比较时建议用容差如abs(sum(probs) - 1.0) 1e-9不要用精确比较。返回值顺序ChanceOutcomes()无强制排序保证但LegalChanceOutcomes()文档约定合法动作升序spiel.h若你的算法依赖顺序请基于legal_actions()或自行排序。7. 关联 API 速览chance_outcomes()与以下 API 协同工作相关文档见 api_reference.md方法作用对应文档is_chance_node()判断当前是否为随机节点state_is_chance_node.mdchance_outcomes()返回 (动作, 概率) 列表state_chance_outcomes.mdapply_action(action)应用随机结果推进状态state_apply_action.mdlegal_actions()随机节点上返回合法随机结果state_legal_actions.mdmax_chance_outcomes()单个随机节点的最大结果数上限game_max_chance_outcomes.mdmax_chance_nodes_in_history()任意历史中随机节点最大数量game_max_chance_nodes_in_history.md其中game.max_chance_outcomes()是游戏级元信息随机节点结果数的上界C 默认返回 0 表示无随机节点见 spiel.h可用于为随机动作预留张量/缓存空间Leduc 的实现返回剩余牌数leduc_poker.cc 附近。8. 小结chance_outcomes()是 OpenSpiel 处理环境随机性的统一入口它把自然的随机选择暴露为普通动作的带概率列表让 CFR 等算法可以穷举、让 MCCFR/MCTS 等算法可以采样。理解它需要同时掌握三个层面接口语义返回归一化的 (动作, 概率) 对仅随机节点有效、游戏声明kExplicitStochastic给出完整分布kSampledStochastic只给占位动作、消费模式is_chance_node()判定 →chance_outcomes()取值 →apply_action()/SampleAction()推进。以此为起点即可在自己的博弈算法中正确处理任何含随机性的游戏环境。赞分享人工智能强化学习深度学习【免费下载链接】open_spielOpenSpiel is a collection of environments and algorithms for research in general reinforcement learning and search/planning in games.项目地址https://gitcode.com/gh_mirrors/op/open_spiel点击查看免费下载相关推荐PyTorch3D sample_pdf 深度解析NeRF 分层采样中的概率密度采样器PyTorch3D sample_pdf 深度解析NeRF 分层采样中的概率密度采样器 本篇文章聚焦 PyTorch3D 中 pytorch3d.render人工智能深度学习计算机视觉图形学微信、QQ、TIM 防撤回完整指南RevokeMsgPatcher 补丁操作步骤与报错排查微信、QQ、TIM 防撤回完整指南RevokeMsgPatcher 补丁操作步骤与报错排查 消息刚发出来转眼就要被撤走还能不能留住RevokeMsgPa桌面应用即时通讯Loki日志采样算法深度解析随机采样与系统采样的终极对比指南Loki日志采样算法深度解析随机采样与系统采样的终极对比指南 Loki是一个开源、高扩展性和多租户的日志聚合系统由Grafana Labs开发。它主要用于收可观测性日志分析后端微服务对象存储云原生上一篇IcemacOS菜单栏管理的终极解决方案让你的桌面整洁如新下一篇Marketing Skills 项目中的 SavvyCal 调度平台集成指南API 与 CLI 实战全解创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考