Action decision method, action decision device, and computing device

CN122366595BActive Publication Date: 2026-09-18INST OF AUTOMATION CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610830367.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-09
Publication Date
2026-09-18
Estimated Expiration
2046-06-09

AI Technical Summary

Technical Problem

由于每步决策都面临多分支选择,随着序列长度的增加,联合动作空间维度呈指数级膨胀,导致穷举搜索算法的计算复杂度急剧上升,在实际应用中面临“维度灾难”的严峻挑战

Benefits of technology

[0015] In the action decision method according to an exemplary embodiment of the present disclosure, since only the non-leaf nodes corresponding to the preferred candidate action scheme are expanded without expanding the leaf nodes corresponding to the preferred candidate action scheme during the process of expanding nodes (e.g., expanding nodes step by step), the computational complexity in the action decision process can be greatly reduced (e.g., the computational complexity is reduced from exponential to linear) and it has good scalability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122366595B_ABST
    Figure CN122366595B_ABST
Patent Text Reader

Abstract

An action decision method, an action decision device and a computing device are provided. The action decision method comprises: constructing a root node of a search tree based on initial state information of an intelligent mobile platform; generating the search tree by expanding nodes layer by layer downwards from the root node, the nodes of the first layer corresponding to the root node; determining an optimal action sequence of the intelligent mobile platform based on the search tree, the step of generating the search tree comprising: generating a corresponding candidate action scheme based on current state information of a current node of a current layer, the current node being a non-leaf node; determining a preferred candidate action scheme from the candidate action schemes; determining a node corresponding to the preferred candidate action scheme of a next layer as a non-leaf node, and determining the remaining nodes of the next layer as leaf nodes; predicting state information of the non-leaf node of the next layer based on the current state information of the current node of the current layer and an action history and the preferred candidate action scheme; and stopping expanding the nodes when the next layer is the last layer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of intelligent mobile platforms, and more specifically, to an action decision-making method, an action decision-making device, and a computing device. Background Technology

[0002] Intelligent mobile platforms typically involve action decisions during operation. In these decisions, the platform needs to make a series of coupled decisions over consecutive time steps. Each decision not only directly affects the current state (e.g., the environment) but also alters the subsequent action space through state transition mechanisms, ultimately forming a complete sequence of actions with temporal dependencies. Because each decision involves multiple branch choices, the dimensionality of the joint action space expands exponentially with the sequence length, leading to a sharp increase in the computational complexity of exhaustive search algorithms and posing a severe challenge of the "curse of dimensionality" in practical applications. Summary of the Invention

[0003] The purpose of this disclosure is to provide an action decision-making method, an action decision-making device, and a computing device.

[0004] In one general aspect, an action decision-making method is provided. The action decision-making method includes: constructing the root node of a search tree based on initial state information of an intelligent mobile platform; generating the search tree by expanding nodes layer by layer downwards from the root node, wherein the search tree includes nodes in N layers, the N layers including a first layer to an Nth layer, and the nodes in the first layer correspond to the root node, where N is an integer greater than or equal to 2; and determining the optimal action sequence of the intelligent mobile platform based on the generated search tree, wherein the step of generating the search tree by expanding nodes layer by layer downwards from the root node includes: generating one or more candidate actions corresponding to the current state of the current node in the current layer based on the current state information of the current node in the current layer. The algorithm proposes a plan, wherein the current node is a non-leaf node; a preferred candidate action plan is determined from the one or more candidate action plans; the node in the next layer corresponding to the current node of the current layer and corresponding to the preferred candidate action plan is determined as the non-leaf node of the next layer, and the node in the next layer corresponding to the remaining candidate action plans (excluding the preferred candidate action plan) of the one or more candidate action plans is determined as the leaf node of the next layer; the state information of the non-leaf node of the next layer is predicted based on the current state information and action history of the current node of the current layer and the preferred candidate action plan; and when the next layer corresponds to the Nth layer, node expansion is stopped.

[0005] Optionally, the step of determining a preferred candidate action scheme from the one or more candidate action schemes includes: receiving a high-level semantic instruction related to the decision intent corresponding to the current node of the current layer; and determining the candidate action scheme that matches the decision intent among the one or more candidate action schemes as the preferred candidate action scheme.

[0006] Optionally, the step of determining the candidate action scheme that matches the decision intent among the one or more candidate action schemes as the preferred candidate action scheme includes: identifying the decision intent by inputting the high-level semantic instruction corresponding to the decision intent into the inference model and evaluating the degree of matching between the one or more candidate action schemes and the decision intent; and determining the candidate action scheme with the highest degree of matching among the one or more candidate action schemes as the preferred candidate action scheme.

[0007] Optionally, the inference model includes at least one of a multimodal large language model, an expert system, and a scoring network, wherein the multimodal large language model is configured to identify a decision intent from a high-level semantic instruction corresponding to the decision intent and evaluate a first degree of matching between the one or more candidate action schemes and the decision intent; the expert system is configured to identify a decision intent from a high-level semantic instruction corresponding to the decision intent and evaluate a second degree of matching between the one or more candidate action schemes and the decision intent; and the scoring network is configured to identify a decision intent from a high-level semantic instruction corresponding to the decision intent and evaluate a third degree of matching between the one or more candidate action schemes and the decision intent, wherein the step of identifying the decision intent and evaluating the degree of matching between the one or more candidate action schemes and the decision intent by inputting the high-level semantic instruction corresponding to the decision intent into the inference model includes: evaluating the degree of matching between the one or more candidate action schemes and the decision intent based on the degree of matching between the first degree of matching, the second degree of matching, and the third degree of matching corresponding to at least one of the multimodal large language model, the expert system, and the scoring network.

[0008] Optionally, the step of generating one or more candidate action schemes corresponding to the current state of the current node in the current layer includes: generating a set of actionable descriptions that satisfy candidate rules based on the current state information of the current node in the current layer and a set of meta-actions, wherein the set of meta-actions includes all possible atomic action types in the decision-making scenario, and the candidate rules include at least one of enumeration-based policies, template-based policies, and sampling-based policies; converting the set of actionable descriptions into an action scheme representation corresponding to the set of actionable descriptions by calling a generative model, wherein the action scheme representation includes at least one of trajectory prediction, state prediction, and visualization image; and generating the one or more candidate action schemes into a set of actionable descriptions and an action scheme representation corresponding to the set of actionable descriptions.

[0009] Optionally, the generative model includes at least one of a conditional generative model and a rule engine that incorporates domain knowledge.

[0010] Optionally, the step of determining the optimal action sequence of the intelligent mobile platform based on the generated search tree includes: determining the optimal action sequence of the intelligent mobile platform by combining the preferred candidate action schemes corresponding to the root nodes of the first layer and the preferred candidate action schemes corresponding to the non-leaf nodes of the second to N-1 layers.

[0011] Optionally, at least two layers from the first to the Nth layer have different node degrees, wherein the node degree of a single layer represents the number of candidate action schemes that a single node of the single layer can generate.

[0012] In one general aspect, an action decision-making device is provided. The action decision-making device includes: a search tree generation module configured to: construct the root node of a search tree based on the initial state information of an intelligent mobile platform, and generate the search tree by expanding nodes layer by layer downwards from the root node, wherein the search tree includes nodes in N layers, the N layers including a first layer to an Nth layer, and the nodes in the first layer correspond to the root node, wherein N is an integer greater than or equal to 2; and an optimal action sequence determination module configured to: determine the optimal action sequence of the intelligent mobile platform based on the generated search tree, wherein the search tree generation module is configured to: generate a sequence corresponding to the current state of the current node in the current layer based on the current state information of the current node in the current layer. One or more candidate action schemes, wherein the current node is a non-leaf node; a preferred candidate action scheme is determined from the one or more candidate action schemes; the node in the next layer corresponding to the current node of the current layer and corresponding to the preferred candidate action scheme is determined as the non-leaf node of the next layer, and the node in the next layer corresponding to the remaining candidate action schemes in the one or more candidate action schemes other than the preferred candidate action scheme is determined as the leaf node of the next layer; the state information of the non-leaf node of the next layer is predicted based on the current state information and action history of the current node of the current layer and the preferred candidate action scheme; and when the next layer corresponds to the Nth layer, node expansion is stopped.

[0013] In one general aspect, a computer-readable storage medium is provided storing a computer program, wherein, when the computer program is executed by a processor, any of the action decision methods described above are implemented.

[0014] In one general aspect, a computing device is provided. The computing device includes: a processor; and a memory storing a computer program that, when executed by the processor, implements any of the action decision methods described above.

[0015] In the action decision method according to an exemplary embodiment of the present disclosure, since only the non-leaf nodes corresponding to the preferred candidate action scheme are expanded without expanding the leaf nodes corresponding to the preferred candidate action scheme during the process of expanding nodes (e.g., expanding nodes step by step), the computational complexity in the action decision process can be greatly reduced (e.g., the computational complexity is reduced from exponential to linear) and it has good scalability.

[0016] In the action decision-making method according to an exemplary embodiment of the present disclosure, since different node degrees are supported at different levels, it is able to adapt to changes in the degree of criticality during the decision-making process.

[0017] In the action decision method according to an exemplary embodiment of the present disclosure, since a search tree can be generated by expanding only the non-leaf nodes corresponding to the preferred candidate action schemes without expanding the leaf nodes corresponding to the preferred candidate action schemes during the process of expanding nodes (e.g., expanding nodes step by step), and the optimal action sequence of the intelligent mobile platform can be determined based on the generated search tree, efficient search can be achieved in complex decision spaces (e.g., exponential decision spaces), and the computational complexity in the action decision process can be greatly reduced (e.g., the computational complexity can be reduced from exponential to linear), and good scalability can be achieved.

[0018] In the action decision method according to an exemplary embodiment of the present disclosure, since the optimal action sequence of the intelligent mobile platform can be determined based on the preferred candidate action schemes of each layer, the action sequence of the intelligent mobile platform is determined to be optimal.

[0019] In the action decision method according to an exemplary embodiment of the present disclosure, since candidate rules can adopt a variety of construction strategies and support the combined use of multiple strategies, both coverage and diversity can be taken into account.

[0020] In the action decision-making method according to an exemplary embodiment of the present disclosure, since a rich variety of candidate options can be generated using a generative model, the limitations of traditional enumeration methods can be overcome and the ability to discover optimal solutions can be enhanced.

[0021] In the action decision-making method according to exemplary embodiments of the present disclosure, a modular design can be adopted to support the replacement and combination of different generation models, thereby adapting to the needs of different technical conditions and application scenarios.

[0022] In the action decision-making method according to exemplary embodiments of the present disclosure, since preferred candidate action schemes can be screened based on high-level semantic instructions during the node expansion process, a sequence search capability guided by high-level semantic instructions is realized. Users can describe their strategy goals or style preferences through natural language, and the action decision-making method or action decision-making device according to the example embodiments of the present disclosure can thereby discover action sequences that match their intentions in the decision space, improving the goal orientation and controllability of the search. Furthermore, ensuring that each decision step undergoes semantic evaluation during the node expansion process avoids the limitations of traditional search methods in semantic understanding, improving the consistency between search results and user intent.

[0023] In the action decision method according to an exemplary embodiment of the present disclosure, the final matching degree can be determined by the matching degree generated based on different inference models, thus improving the accuracy of determining the final matching degree.

[0024] In the action decision-making method according to exemplary embodiments of the present disclosure, a modular design can be adopted to support the replacement and combination of different inference models, thereby adapting to the needs of different technical conditions and application scenarios. Attached Figure Description

[0025] Figure 1 A flowchart illustrating an action decision method according to an exemplary embodiment of the present disclosure is shown.

[0026] Figure 2 A flowchart illustrating a method for generating a search tree by expanding nodes downwards layer by layer from the root node, according to an exemplary embodiment of the present disclosure.

[0027] Figure 3 A schematic diagram of a search tree according to an exemplary embodiment of the present disclosure is shown.

[0028] Figure 4 A schematic diagram of an action decision method according to an exemplary embodiment of the present disclosure is shown.

[0029] Figure 5 An action decision device according to an example embodiment of the present disclosure is shown.

[0030] Figure 6 A block diagram of a computing device according to an exemplary embodiment of the present disclosure is shown. Detailed Implementation

[0031] The following detailed descriptions are provided to help the reader gain a comprehensive understanding of the methods, apparatus, and / or action decision-making methods or apparatus according to exemplary embodiments of this disclosure. However, after understanding this disclosure, various changes, modifications, and equivalents of the methods, apparatus, and / or action decision-making methods or apparatus according to exemplary embodiments of this disclosure will become apparent. For example, the order of operations described herein is merely illustrative and is not limited to those orders set forth herein, but may be changed as will become clear after understanding this disclosure, except for operations that must occur in a specific order. Furthermore, for clarity and conciseness, descriptions of features known in the art may be omitted.

[0032] The features described herein may be implemented in different forms and should not be construed as limited to the examples described herein. Rather, the examples described herein are provided only to illustrate some of the many possible ways to implement the methods, apparatus and / or action decision methods or apparatus according to exemplary embodiments of the present disclosure, many of which will become clear upon understanding the disclosure of this application.

[0033] As used herein, the term “and / or” includes any one of the associated listed items and any combination of any two or more.

[0034] Although terms such as “first,” “second,” and “third” may be used herein to describe various components, assemblies, regions, layers, or parts, these components, assemblies, regions, layers, or parts should not be limited by these terms. Rather, these terms are used only to distinguish one component, assembly, region, layer, or part from another. Thus, without departing from the teaching of the examples described herein, the first component, first assembly, first region, first layer, or first part referred to as the first component, first assembly, first region, first layer, or first part may also be referred to as the second component, second assembly, second region, second layer, or second part.

[0035] In the specification, when an element (such as a layer, region, or substrate) is described as being "on" another element, "connected to," or "bonded to" another element, the element may be directly "on" another element, directly "connected to," or "bonded to" the other element, or one or more other elements may be present in between. Conversely, when an element is described as being "directly on" another element, "directly connected to," or "directly bonded to" another element, no other elements may be present in between.

[0036] The terminology used herein is for the purpose of describing various examples only and is not intended to limit disclosure. Unless the context clearly indicates otherwise, the singular form is intended to include the plural form as well. The terms “comprising,” “including,” and “having” indicate the presence of the described features, quantities, operations, components, elements, and / or combinations thereof, but do not preclude the presence or addition of one or more other features, quantities, operations, components, elements, and / or combinations thereof.

[0037] Unless otherwise defined, all terms used herein (including technical and scientific terms) shall have the same meaning as would be normally understood by one of ordinary skill in the art to which this disclosure pertains upon learning of this disclosure. Unless expressly defined herein, terms (such as those defined in a general dictionary) shall be interpreted as having a meaning consistent with their meaning in the context of the relevant field and in this disclosure, and shall not be interpreted in an idealized or overly formalistic manner.

[0038] Furthermore, in the description of the examples, detailed descriptions of well-known related structures or functions will be omitted when it is believed that such detailed descriptions would lead to a vague interpretation of this disclosure.

[0039] In this disclosure, an action can also be referred to as a meta-action. A meta-action is an atomic operation unit in a sequential decision-making process and is a fundamental element constituting a complex decision sequence. By defining a standardized set of meta-actions and attribute structures, a unified decision vocabulary can be provided for the subsequent search process.

[0040] The meta-action set encompasses all possible atomic action types within a decision-making scenario. In intelligent mobile platform (e.g., robotic) task scenarios, meta-actions may include, but are not limited to, grasping, placing, moving, and rotating. Optionally, meta-actions can be further refined and combined in various ways. For example, grasping may include grasping forward, grasping downward, etc. Moving may include moving forward, moving backward, moving left, moving right, etc. In one example, the definition of the meta-action set can balance completeness and atomicity, ensuring that complex decision-making behaviors can be expressed through combination.

[0041] The attribute structure of a meta-action defines the semantic information needed to describe a specific action instance. Common attributes include action type, executing entity, and target location. Scene-specific attributes can be extended according to application requirements. The design of the attribute structure facilitates the generation and parsing of natural language descriptions.

[0042] An action sequence can be defined as an ordered combination of meta-actions. Given a set of meta-actions... A length of The action sequence can be represented as ,in Indicates the first The meta-actions of a step. The rules for combining action sequences define the legal connections between meta-actions, which are used to constrain the search space. For example, it can be defined that certain action types must be followed by specific subsequent actions, or that certain action combinations are not feasible under certain conditions.

[0043] In the following description, embodiments will be presented in detail with reference to the accompanying drawings. However, embodiments may be implemented in various forms and are not limited to those described herein.

[0044] Figure 1 A flowchart illustrating an action decision method according to an exemplary embodiment of the present disclosure is shown.

[0045] Reference Figure 1 In operation S110, the root node of the search tree can be constructed based on the initial state information of the intelligent mobile platform.

[0046] A smart mobile platform can be any intelligent platform with mobility capabilities. For example, by way of example only, a smart mobile platform can include, but is not limited to, robots, intelligent vehicles, etc.

[0047] The state information of an intelligent mobile platform may include at least one of any information associated with action decisions. In one example, the state information of an intelligent mobile platform may include, but is not limited to, at least one of the following: the location information of the intelligent mobile platform (e.g., the location information of some components of the intelligent mobile platform), the velocity information of the location information of the intelligent mobile platform (e.g., the velocity information of some components of the intelligent mobile platform), and contextual information related to action decisions.

[0048] A search tree can be the core data structure for organizing a sequence of decision-making search processes. The root node can represent the initial state of the search, typically corresponding to the initial configuration or historical context of the decision scenario. In one example, the root node's state representation contains complete scenario information, such as state variables like the position and velocity of each entity, as well as contextual information relevant to the decision.

[0049] In operation S120, a search tree can be generated by expanding the nodes downwards layer by layer from the root node.

[0050] In an example embodiment, the operation of generating a search tree by expanding nodes layer by layer downwards from the root node may include: generating one or more candidate action schemes corresponding to the current state of the current node in the current layer based on the current state information of the current node in the current layer, wherein the current node is a non-leaf node; determining a preferred candidate action scheme from the one or more candidate action schemes; determining the node in the next layer corresponding to the preferred candidate action scheme as a non-leaf node in the next layer, and determining the node corresponding to the remaining candidate action schemes (excluding the preferred candidate action scheme) in the one or more candidate action schemes as leaf nodes in the next layer; predicting the state information of the non-leaf nodes in the next layer based on the current state information and action history of the current node in the current layer and the preferred candidate action scheme; and stopping node expansion when the next layer corresponds to the Nth layer. In this example embodiment, since only the non-leaf nodes corresponding to the preferred candidate action scheme are expanded without expanding the leaf nodes corresponding to the preferred candidate action scheme during the node expansion process (e.g., expanding nodes layer by layer), the computational complexity in the action decision process can be greatly reduced (e.g., reducing the computational complexity from exponential to linear), and it has good scalability. This will be discussed later in conjunction with... Figure 2 This example embodiment will now be described in more detail.

[0051] The search tree can include nodes across N layers. The N layers range from layer 1 to layer N, with each node in layer 1 corresponding to the root node. N is an integer greater than or equal to 2. In one example embodiment, at least two layers from layer 1 to layer N have different node degrees, where the node degree of a single layer represents the number of candidate action schemes that a single node in that layer can generate. Because different node degrees can be set at different levels, it can adapt to changes in the criticality of decisions.

[0052] The hierarchical structure of the search tree corresponds to the number of steps in the decision sequence. Layer nodes correspond to the execution of The state after the first step, where k is an integer from 1 to N. From the root node to the nth node... The path of a layer node corresponds to a path of length... The sequence of actions.

[0053] The branching structure of a node corresponds to the candidate action options in the current state. Each non-leaf node can be expanded into multiple child nodes, with the number of child nodes equal to the number of candidate solutions for that step. The search process involves selecting a branch at each level for expansion, ultimately forming a complete path from the root to the leaf.

[0054] The depth parameter of the search tree (e.g., N-1) defines the maximum number of steps in the search, corresponding to the length of the final output action sequence. The depth parameter can be configured according to specific application requirements. Shallower depths are suitable for short-term action decisions, while deeper depths are suitable for long-term planning tasks.

[0055] The node degree parameter defines the number of candidate solutions at each step and determines the width of the search tree. A larger node degree results in a wider search space, but also increases computational overhead. In one example embodiment, different node degrees can be set at different levels to accommodate changes in the criticality of decisions.

[0056] In operation S130, the optimal sequence of actions for the intelligent mobile platform can be determined based on the generated search tree.

[0057] Since a search tree can be generated by expanding only the non-leaf nodes corresponding to the preferred candidate action schemes without expanding the leaf nodes corresponding to the preferred candidate action schemes during the expansion of nodes (e.g., expanding nodes step by step), and the optimal action sequence of the intelligent mobile platform can be determined based on the generated search tree, efficient search can be achieved in complex decision spaces (e.g., exponential decision spaces), and the computational complexity in the action decision process can be greatly reduced (e.g., reducing the computational complexity from exponential to linear), and it has good scalability.

[0058] In one example embodiment, the optimal action sequence of the intelligent mobile platform can be determined by combining the preferred candidate action schemes corresponding to the root nodes of the first layer and the preferred candidate action schemes corresponding to the non-leaf nodes of the second to N-1th layers. Since the optimal action sequence of the intelligent mobile platform can be determined based on the preferred candidate action schemes of each layer, the determined action sequence of the intelligent mobile platform is optimal.

[0059] Figure 2 A flowchart illustrating a method for generating a search tree by expanding nodes downwards layer by layer from the root node, according to an exemplary embodiment of the present disclosure.

[0060] Reference Figure 2 In operation S210, one or more candidate action schemes corresponding to the current state of the current node of the current layer can be generated based on the current state information of the current node of the current layer.

[0061] One or more candidate action schemes corresponding to the current state of the current node in the current layer can be generated using various existing methods based on the current state information of the current node in the current layer.

[0062] In one example embodiment, a set of action descriptions that satisfy candidate rules can be generated based on the current state information of the current node in the current layer and the set of meta-actions. The set of meta-actions includes all possible atomic action types in the decision-making scenario. Candidate rules include at least one of enumeration-based strategies, template-based strategies, and sampling-based strategies. The enumeration-based strategy can traverse each action type in the meta-action set and filter feasible options in combination with the current state. The template-based strategy predefines action description templates and generates specific descriptions (e.g., a set of action descriptions that satisfy candidate rules) by filling in state-related parameters. The sampling-based strategy samples from the probability distribution of the action space to generate candidate descriptions (e.g., a set of action descriptions that satisfy candidate rules). Since candidate rules can adopt multiple construction strategies and support the combined use of multiple strategies, both coverage and diversity can be considered.

[0063] Then, the set of action descriptions can be transformed into action plan representations corresponding to the set of action descriptions by invoking a generative model. The action plan representation can be tailored to the application scenario. In one example, the action plan representation includes at least one of trajectory prediction, state prediction, and visualization. Trajectory prediction can represent the motion trajectory of each entity (e.g., components of a smart mobile component) after the action is performed. State prediction can represent the state configuration of the scene after the action is performed. Visualization can render the trajectory or state as an image for easy intuitive understanding. Because a rich variety of candidate options can be generated using generative models, the limitations of traditional enumeration methods can be overcome, enhancing the ability to discover optimal solutions.

[0064] The generative model can be a pre-trained generative model. It can be implemented using various techniques. In one example embodiment, the generative model may include at least one of a conditional generative model and a rule engine incorporating domain knowledge. In one example, the conditional generative model can generate action scheme representations (e.g., corresponding trajectory or state predictions) based on action descriptions. The rule engine incorporating domain knowledge can directly compute action scheme representations based on a set of actionable action descriptions. A modular design can be adopted to support the replacement and combination of different generative models, thereby adapting to the needs of different technical conditions and application scenarios.

[0065] Next, one or more candidate action schemes can be generated into a set of action descriptions and corresponding action scheme representations. In other words, the result of candidate action scheme generation includes a set of pairs of action descriptions and scheme representations for subsequent processing.

[0066] In operation S220, a preferred candidate action scheme can be determined from one or more candidate action schemes.

[0067] A preferred candidate action can be determined from one or more candidate action options using various existing methods.

[0068] In one example embodiment, a high-level semantic instruction related to the decision intent corresponding to the current node of the current layer is received, and a candidate action scheme matching the decision intent from one or more candidate action schemes is determined as the preferred candidate action scheme. Since the preferred candidate action scheme can be filtered based on the high-level semantic instruction during the node expansion process, a sequence search capability guided by high-level semantic instructions is realized. Users can describe their strategy goals or style preferences using natural language, and the action decision-making method or action decision-making device according to the example embodiment of this disclosure can thereby discover action sequences that match the intent in the decision space, improving the goal orientation and controllability of the search. Furthermore, ensuring that each decision step undergoes semantic evaluation during the node expansion process avoids the limitations of traditional search methods in semantic understanding, improving the consistency between search results and user intent.

[0069] High-level semantic instructions can be provided by users in natural language, describing desired policy objectives, style preferences, or constraints. In a non-restrictive example, in a robotics task scenario, a high-level semantic instruction could be "plan a safe execution path," and the decision intent could be "a highly safe execution path." The expression of high-level semantic instructions can be flexible and natural, without strictly adhering to a predefined format.

[0070] In one embodiment, the decision intent is identified by inputting a high-level semantic instruction corresponding to the decision intent into the inference model, and the degree of matching between one or more candidate action schemes and the decision intent is evaluated. The candidate action scheme with the highest degree of matching among the one or more candidate action schemes is determined as the preferred candidate action scheme.

[0071] The inference model can be a pre-trained inference model. In one example, the inference model may include at least one of a multimodal large language model, an expert system, and a scoring network. In one example embodiment, the multimodal large language model is configured to identify a decision intent from a high-level semantic instruction corresponding to the decision intent and evaluate a first degree of matching between one or more candidate action schemes and the decision intent; the expert system is configured to identify the decision intent from the high-level semantic instruction corresponding to the decision intent and evaluate a second degree of matching between one or more candidate action schemes and the decision intent; and the scoring network is configured to identify the decision intent from the high-level semantic instruction corresponding to the decision intent and evaluate a third degree of matching between one or more candidate action schemes and the decision intent. The degree of matching between one or more candidate action schemes and the decision intent can be evaluated based on the degree of matching between the first, second, and third degree of matching and at least one of the multimodal large language model, expert system, and scoring network. In this example embodiment, the final degree of matching can be determined based on the degree of matching generated by different inference models, thus improving the accuracy of determining the final degree of matching.

[0072] For example, the inference model can be responsible for understanding the intent of semantic instructions and evaluating the degree of matching between each candidate solution and the instruction. The inference model can have multimodal understanding capabilities, capable of processing textual descriptions of instructions and actions, as well as image-based visualizations of solutions. Optionally, the inference model can also have reasoning and judgment capabilities, able to comprehensively consider multiple factors such as instruction intent, scene context, and solution features to make an evaluation.

[0073] The inference model can be implemented using various technologies. In one example, a multimodal large language model possesses powerful semantic understanding and reasoning capabilities, directly receiving instructions and candidate information in the form of prompt words and outputting evaluation conclusions. In another example, an expert system encodes domain knowledge into a rule base and performs evaluation through rule matching. In yet another example, a scoring network models the evaluation as a regression task, outputting matching scores for each candidate solution. The action decision-making method or device according to the example embodiments of this disclosure can adopt a modular design, supporting the replacement and combination of different inference models, thereby adapting to the needs of different technical conditions and application scenarios.

[0074] In operation S230, the node corresponding to the preferred candidate action scheme in the next layer corresponding to the current node in the current layer can be determined as the non-leaf node of the next layer, and the node corresponding to the remaining candidate action schemes other than the preferred candidate action scheme in one or more candidate action schemes can be determined as the leaf node of the next layer.

[0075] In operation S240, the state information of non-leaf nodes in the next layer can be predicted based on the current state information and action history of the current node in the current layer, as well as the preferred candidate action scheme.

[0076] The state information of non-leaf nodes in the next layer can be predicted in various ways based on the current state information and action history of the current node in the current layer, as well as the preferred candidate action schemes. For example, the state information of non-leaf nodes in the next layer can be predicted through a state transition mechanism based on the current state information and action history of the current node in the current layer, as well as the preferred candidate action schemes.

[0077] In operation S250, when the next layer corresponds to the Nth layer, node expansion is stopped.

[0078] Figure 3 A schematic diagram of a search tree according to an exemplary embodiment of the present disclosure is shown.

[0079] Reference Figure 3 The search tree according to exemplary embodiments of this disclosure may include nodes across multiple layers. By way of example only, the multiple layers may include layers one through four. A node N100 in the first layer may correspond to the root node of the search tree. In one example, the initial state and / or historical context of the intelligent mobile platform is defined as the root node of the search tree.

[0080] A search tree can be generated by expanding nodes layer by layer downwards from the root node. For example, based on the current state information of the current node in the current layer, one or more candidate action schemes corresponding to the current state of the current node in the current layer are generated, where the current node is a non-leaf node. A preferred candidate action scheme is determined from one or more candidate action schemes. The nodes in the next layer corresponding to the current node in the current layer that correspond to the preferred candidate action scheme are determined as non-leaf nodes in the next layer, and the nodes in the next layer corresponding to the remaining candidate action schemes (excluding the preferred candidate action scheme) in the one or more candidate action schemes are determined as leaf nodes in the next layer. The state information of the non-leaf nodes in the next layer is predicted based on the current state information and action history of the current node in the current layer, as well as the preferred candidate action scheme. When the next layer corresponds to the third layer, node expansion stops.

[0081] More specifically, when the current layer is the first layer and the current non-leaf node is the root node, one or more candidate action schemes corresponding to the state of the first layer's root node are generated based on the state information of the first layer's root node. A preferred candidate action scheme is determined from the one or more candidate action schemes corresponding to the state of the first layer's root node. The node N220 in the next layer (i.e., the second layer) corresponding to the preferred candidate action scheme is determined as a non-leaf node of the second layer, and the nodes N210 and N230 in the second layer corresponding to the remaining candidate action schemes (excluding the preferred candidate action scheme) in the one or more candidate action schemes are determined as leaf nodes of the second layer. The state information of the non-leaf nodes in the second layer is predicted based on the state information and action history of the first layer's root node and the preferred candidate action scheme.

[0082] When the current layer is the second layer and the current non-leaf node is node N220, one or more candidate action schemes corresponding to the state of node N220 in the second layer are generated based on the state information of node N220 in the second layer. A preferred candidate action scheme is determined from the one or more candidate action schemes corresponding to the state of node N220 in the second layer. Node N330 in the next layer (i.e., the third layer) corresponding to the preferred candidate action scheme is determined as a non-leaf node in the third layer, and nodes N310 and N320 in the third layer corresponding to the remaining candidate action schemes (excluding the preferred candidate action scheme) in the one or more candidate action schemes are determined as leaf nodes in the third layer. The state information of non-leaf nodes in the fourth layer is predicted based on the state information and action history of node N330 in the third layer, as well as the preferred candidate action scheme.

[0083] Since the fourth level is the last level of the search tree, the nodes of the fourth level will not be expanded further.

[0084] The optimal action sequence of the intelligent mobile platform can be determined by combining the preferred candidate action schemes corresponding to the root node of the first layer and the preferred candidate action schemes corresponding to the non-leaf nodes of the second to third layers. In other words, the optimal action sequence can be determined by combining the preferred candidate action schemes corresponding to node N100 of the first layer (from...). Figure 3 The solid arrow between node N100 in the first layer and node N220 in the second layer represents the preferred candidate action scheme corresponding to node N220 in the second layer (by...). Figure 3 The solid arrow between node N220 in the second layer and node N330 in the third layer (represented by the arrow) and the preferred candidate action scheme corresponding to node N330 in the third layer (by...) Figure 3 The solid arrow between node N330 in the third layer and node N430 in the fourth layer is used to determine the optimal action sequence of the intelligent mobile platform.

[0085] Figure 4A schematic diagram of an action decision method according to an exemplary embodiment of the present disclosure is shown.

[0086] The action decision method according to an exemplary embodiment of the present disclosure can model the search process as a tree traversal process that starts from the root node and expands downwards layer by layer. Figure 4 The root node 410 represents the starting state of the search, which typically corresponds to the initial configuration or historical context of the decision-making scenario. The state representation of the root node 410 includes complete scenario information (e.g., entity position 411, entity velocity 412, and other state variables) as well as decision-related context information (e.g., decision-related context 413).

[0087] The candidate solution generation module 420 is responsible for generating multiple feasible candidate action solutions 430 at each decision-making step, providing options for subsequent evaluation and selection. Candidate action solutions 430 may include solution 1 431 and solution 2 432, etc. The candidate solution generation module 420 may include a candidate rule construction unit and a solution generation unit.

[0088] The candidate rule construction unit generates a set of feasible action descriptions for the current step based on the current node state and the meta-action definition. The design of the candidate rules integrates domain knowledge and state constraints to ensure that the generated candidate actions are semantically reasonable and physically feasible.

[0089] Candidate rules can be constructed using various strategies. An enumeration-based strategy iterates through each action type in the meta-action set and filters feasible options based on the current state. A template-based strategy predefines action description templates and generates specific descriptions by filling in state-related parameters. A sampling-based strategy samples from the probability distribution of the action space to generate candidate descriptions. The action decision-making method or device according to the example embodiments of this disclosure supports the combined use of multiple strategies to balance coverage and diversity.

[0090] The solution generation unit calls the generation model to transform the action description into a specific solution representation. The form of the solution representation depends on the application scenario and may include: trajectory prediction, representing the motion trajectory of each entity after the action is executed; state prediction, representing the state configuration of the scene after the action is executed; and visualization image, rendering the trajectory or state into an image for easy intuitive understanding.

[0091] Generative models can be implemented using various techniques. Conditional generative models use action descriptions as conditions to generate corresponding trajectory or state predictions; reinforcement learning policy networks output action distributions based on the current state and sample to generate candidates; rule engines combine domain knowledge to directly compute candidate solutions. The action decision-making method or device according to the example embodiments of this disclosure adopts a modular design, supporting the replacement and combination of different generative models.

[0092] The candidate solution generation results include a set of paired action descriptions and solution representations, which are then processed by the semantic reasoning evaluation module.

[0093] The semantic reasoning evaluation module 450 is responsible for evaluating candidate solutions and selecting the optimal solution based on user-provided semantic instructions 440 (e.g., high-level semantic instructions). (e.g., selecting solution k 460 from multiple feasible candidate action solutions 430). The core capability of this module lies in understanding abstract semantic intent and transforming it into evaluation criteria for specific solutions.

[0094] For example, the evaluation input of the semantic reasoning evaluation module 450 may include three parts: high-level semantic instructions, the state information and action history of the current node, and the representation of each candidate solution. The state information and action history provide the context for decision-making, and the candidate solution representation provides the options to be evaluated.

[0095] The semantic reasoning evaluation module 450 can use a reasoning model. The reasoning model is responsible for understanding the intent of semantic instructions and evaluating the degree of matching between each candidate solution and the instruction. The reasoning model has multimodal understanding capabilities, capable of processing textual instructions and action descriptions, as well as image-based solution visualizations. The reasoning model also needs to possess reasoning and judgment capabilities, able to comprehensively consider factors such as instruction intent, scene context, and solution features to make an evaluation.

[0096] The reasoning model can be implemented using various technologies. Multimodal large language models possess powerful semantic understanding and reasoning capabilities, directly receiving instructions and candidate information in the form of prompts and outputting evaluation conclusions. Expert systems encode domain knowledge into a rule base and perform evaluation through rule matching. Scoring networks model the evaluation as a regression task, outputting matching scores for each candidate solution. The action decision-making method or device according to the example embodiments of this disclosure adopts a modular design, supporting the replacement and combination of different reasoning models.

[0097] The evaluation output of the semantic reasoning evaluation module 450 is a selection conclusion, which identifies the candidate solution with the highest score or the one that best matches the instruction intent (e.g., solution k 460). The selection conclusion is then passed to the iterative search scheduling module 470 to update the current node state and advance to the next iteration.

[0098] The iterative search scheduling module 470 can coordinate the candidate solution generation module 420 and the semantic reasoning evaluation module 450 to complete the entire search process from the initial state to the final action sequence through iterative loops.

[0099] The initialization phase sets search parameters. The root node state is determined by the initial scene configuration provided by the user; high-level semantic instructions are input by the user in natural language; the maximum search depth N is set according to application requirements, determining the length of the output action sequence. According to the example embodiments of this disclosure, the action decision method or action decision device initializes the action sequence record as an empty list, with the current node pointing to the root node.

[0100] The iterative search phase executes M rounds of loops, with each round corresponding to one decision. M can be a positive integer.

[0101] First, the candidate solution generation module is invoked. The candidate rule construction unit generates a set of action descriptions based on the current node state; the solution generation unit calls the generation model to transform each action description into a specific solution representation, outputting a set of candidate solutions. Second, the semantic reasoning evaluation module is invoked. The high-level semantic instruction, the current node state and action history, and the set of candidate solutions are taken as input; the reasoning model evaluates the matching degree between each candidate solution and the instruction, outputting the selection of the optimal candidate solution. Then, a state update is performed. The selected action is added to the action sequence record; the current node state is updated according to the selected solution representation. The state update can use trajectory prediction results from the solution or be calculated through a state transition function. Finally, the termination condition is checked. If M iterations have been completed, the search terminates; otherwise, the next iteration begins.

[0102] The output phase returns the search results. The complete action sequence record is the discovered optimal action sequence, corresponding to the search path from the root node to the final node. The action decision-making method or device according to the example embodiments of this disclosure can simultaneously output the scheme representation for each step, facilitating intuitive understanding and verification by the user.

[0103] The core advantages of the iterative mechanism are: the generation module ensures that there are a wide variety of candidate options at each step; the evaluation module ensures that the choice at each step conforms to the high-level semantic intent; and the iterative mechanism links single-step decisions into a complete sequence to achieve long-term planning goals.

[0104] After completing M iterations, the optimal action sequence for the intelligent mobile platform can be output.

[0105] The action decision-making method or device according to the example embodiments of this disclosure supports multiple parameter configurations to adapt to different application scenarios. The search depth N is configurable to adapt to the needs of decision sequences of different lengths; the node degree is configurable to balance the search breadth and computational overhead; the candidate rules are configurable to incorporate constraint knowledge from different domains; and the generation model and the inference model are replaceable to adapt to different technical conditions and accuracy requirements.

[0106] The action decision-making method or device according to the example embodiments of this disclosure supports two usage modes: single-step search and multi-step search. In single-step search mode, the depth M is set to 1, and the action decision-making method or device according to the example embodiments of this disclosure performs only one iteration, outputting a single optimal action plan, suitable for real-time decision-making scenarios. In multi-step search mode, the depth M is greater than 1, and the action decision-making method or device according to the example embodiments of this disclosure performs multiple iterations, outputting a complete action sequence, suitable for planning and decision-making scenarios.

[0107] The action decision-making method or device according to the example embodiments of this disclosure adopts a modular architecture design, with each functional module interacting through a standardized interface, supporting independent upgrades and replacements. The candidate solution generation module can be replaced with different generation model implementations; the semantic reasoning evaluation module can be replaced with different reasoning model implementations; and the search strategy can be extended to variant forms such as parallel search and bundle search.

[0108] Figure 5 An action decision device according to an example embodiment of the present disclosure is shown.

[0109] Reference Figure 5 The action decision device 500 includes a search tree generation module 510 and an optimal action sequence determination module 520.

[0110] The search tree generation module 510 is configured to construct the root node of the search tree based on the initial state information of the intelligent mobile platform, and to generate the search tree by expanding the nodes downwards layer by layer from the root node. The search tree consists of nodes in N layers. The N layers consist of layers one through N, and the nodes in the first layer correspond to the root node, where N is an integer greater than or equal to 2.

[0111] For example, the search tree generation module can be configured to: generate one or more candidate action schemes corresponding to the current state of the current node in the current layer based on the current state information of the current node in the current layer, wherein the current node is a non-leaf node; determine the preferred candidate action scheme from the one or more candidate action schemes; determine the node in the next layer corresponding to the preferred candidate action scheme as the non-leaf node of the next layer, and determine the node in the next layer corresponding to the remaining candidate action schemes in the one or more candidate action schemes other than the preferred candidate action scheme as the leaf node of the next layer; predict the state information of the non-leaf node in the next layer based on the current state information and action history of the current node in the current layer and the preferred candidate action scheme; and stop expanding nodes when the next layer corresponds to the Nth layer.

[0112] The optimal action sequence determination module 520 is configured to determine the optimal action sequence of the intelligent mobile platform based on the generated search tree.

[0113] Reference Figures 1 to 4 One or more of the above describe the operations performed by the search tree generation module 510 to generate the search tree and by the optimal action sequence determination module 520 to determine the optimal action sequence of the intelligent mobile platform. Here, to avoid redundancy, the operations performed by the search tree generation module 510 to generate the search tree and by the optimal action sequence determination module 520 to determine the optimal action sequence of the intelligent mobile platform will not be repeated.

[0114] Figure 6 A block diagram of a computing device according to an exemplary embodiment of the present disclosure is shown.

[0115] Reference Figure 6 A computing device 600 according to an example embodiment of the present disclosure may include a processor 610 and a memory 620. Here, the memory 620 stores a computer program, wherein the computer program, when executed by the processor 610, implements a reference... Figures 1 to 4 The arbitrary method described. For the sake of brevity, the reference executed by the computing device 600 will not be described again here. Figures 1 to 4 Any method described.

[0116] Optionally, the computing device 600 can generally follow the technical route of "meta-action definition - search tree construction - candidate solution generation - semantic reasoning evaluation - iterative search", and can implement various modules. This method coordinates the generation module and the evaluation module, and efficiently discovers the optimal action sequence that conforms to the high-level semantic instructions in the exponential decision space through a cyclic iterative mechanism.

[0117] The computing device 600 first establishes a formal representation of the decision-making process. The meta-action definition module determines all possible atomic action types in the decision-making scenario, defines the attribute structure of each meta-action (including semantic attributes such as action type, participating entities, target location, and action style), and constrains the legal connection relationships between meta-actions, forming a complete decision vocabulary and combination rules.

[0118] Based on the formal representation, the search tree construction module organizes the search process in a tree structure. The initial state or historical context is defined as the root node of the search tree, and the state representation of each node contains current scene information and the history of actions already executed. The parent-child relationship between nodes corresponds to the state transition after performing a certain action, and the depth parameter of the search tree corresponds to the preset number of decision steps.

[0119] During the search process, the candidate solution generation module generates multiple candidate action solutions at each step. The candidate rule construction unit generates a set of feasible action descriptions for the current step based on the current node state and the meta-action definition; the solution generation unit calls the generation model to transform each action description into a specific solution representation, including trajectory prediction, state prediction, and visualization images.

[0120] For each candidate solution, the semantic reasoning evaluation module evaluates them based on the high-level semantic instructions provided by the user. This module receives the desired policy objective, style preference, or constraints as input, combines the current node's state information and action history, understands the instruction intent through a reasoning model, evaluates the matching degree of each candidate solution, and finally outputs the solution with the highest evaluation score as the selection for the current step.

[0121] The above generation and evaluation process is coordinated by the iterative search scheduling module. After initializing the root node state, high-level semantic instructions, and maximum search depth, the system sequentially calls the generation and evaluation modules in each iteration to generate and select a solution. Based on the selection result, the current node state is updated and the system proceeds to the next step. After reaching the maximum search depth, the system outputs the complete action sequence from the root node to the current node, which is the optimal decision path that conforms to the high-level semantic instructions.

[0122] The various modules in the action decision-making apparatus of this disclosure shown can be configured as software, hardware, firmware, or any combination thereof to perform specific functions. For example, each module may correspond to a dedicated integrated circuit, pure software code, or a module combining software and hardware. Furthermore, one or more functions implemented by each module may also be uniformly executed by components in a physical entity device (e.g., a processor, client, or server).

[0123] Furthermore, the action decision-making method of this disclosure described herein can be implemented by a program (or instructions) recorded on a computer-readable storage medium. For example, according to an exemplary embodiment of this disclosure, a computer-readable storage medium storing instructions may be provided, wherein when the instructions are executed by at least one computing device, they cause at least one computing device to perform the action decision-making method according to this disclosure.

[0124] The computer program in the aforementioned computer-readable storage medium can run in an environment deployed in computer devices such as clients, hosts, agent devices, and servers. It should be noted that the computer program can also be used to perform additional steps beyond those described above, or to perform more specific processing while performing the above steps. The details of these additional steps and further processing are already described in the reference... Figures 1 to 4 One or more of the related methods are mentioned in the description of the relevant methods, so they will not be repeated here to avoid repetition.

[0125] It should be noted that each module in the action decision-making apparatus according to the exemplary embodiments of the present disclosure can rely entirely on the operation of the computer program to realize the corresponding function. That is, each module corresponds to each step in the functional architecture of the computer program, so that the entire action decision-making method or action decision-making apparatus according to the exemplary embodiments of the present disclosure is called by a special software package (e.g., a lib library) to realize the corresponding function.

[0126] On the other hand, the various modules according to the various embodiments of this disclosure can also be implemented by hardware, software, firmware, middleware, microcode, or any combination thereof. When implemented in software, firmware, middleware, or microcode, the program code or code segment for performing the corresponding operation can be stored in a computer-readable medium such as a storage medium, so that a processor can perform the corresponding operation by reading and running the corresponding program code or code segment.

[0127] For example, exemplary embodiments of the present disclosure can also be implemented as a computing device including a storage component and a processor, wherein the storage component stores a set of computer-executable instructions, and when the set of computer-executable instructions is executed by the processor, an action decision method according to exemplary embodiments of the present disclosure is performed.

[0128] Specifically, the computing device can be deployed on a server or client, or on node devices in a distributed network environment. Furthermore, the computing device can be a PC, tablet, personal digital assistant, smartphone, web application, or other device capable of executing the aforementioned set of instructions.

[0129] Here, the computing device is not necessarily a single computing device, but can be any collection of devices or circuits capable of executing the above instructions (or instruction sets) individually or in combination. The computing device can also be part of a manager that integrates control of an action decision device according to an example embodiment of the present disclosure or an action decision device according to an example embodiment of the present disclosure, or can be configured to be interconnected with a portable electronic device locally or remotely (e.g., via wireless transmission).

[0130] In a computing device, the processor may include a central processing unit (CPU), a graphics processing unit (GPU), a programmable logic device, a dedicated processor, an action decision method or device according to exemplary embodiments of the present disclosure, a microcontroller, or a microprocessor. By way of example and not limitation, the processor may also include analog processors, digital processors, microprocessors, multi-core processors, processor arrays, network processors, etc.

[0131] Some operations described in the action decision method according to exemplary embodiments of the present disclosure can be implemented in software, some operations can be implemented in hardware, and some operations can also be implemented in a combination of software and hardware.

[0132] The processor can execute instructions or code stored in one of the storage components, which can also store data. Instructions and data can also be sent and received over a network via a network interface device, which can employ any known transport protocol.

[0133] The storage component can be integrated with the processor, for example, by placing RAM or flash memory within an integrated circuit microprocessor. Alternatively, the storage component can include a separate device, such as an external disk drive, storage array, or any other storage device that can be used by the action decision method or action decision apparatus according to the example embodiments of this disclosure. The storage component and the processor can be operatively coupled, or can communicate with each other, for example, via I / O ports, network connections, etc., enabling the processor to read files stored in the storage component.

[0134] In addition, the computing device may include a video display (such as a liquid crystal display) and a user interaction interface (such as a keyboard, mouse, touch input device, etc.). All components of the computing device may be interconnected via a bus and / or network.

[0135] The action decision method according to exemplary embodiments of this disclosure can be described as various interconnected or coupled functional blocks or functional diagrams. However, these functional blocks or functional diagrams can be equally integrated into a single logic device or operate according to non-precise boundaries.

[0136] Therefore, refer to Figures 1 to 4 At least one of the action decision-making methods described herein can be implemented by an action decision-making apparatus or computing device according to an example embodiment of the present disclosure, including at least one processor and at least one memory storing instructions.

[0137] According to an exemplary embodiment of the present disclosure, at least one computing device is a computing device for executing an action decision method according to an exemplary embodiment of the present disclosure, and a storage device stores a set of computer-executable instructions. When the set of computer-executable instructions is executed by at least one computing device, the action decision method described with reference to FIGA is executed.

[0138] The foregoing has described various exemplary embodiments of this disclosure. It should be understood that the foregoing description is exemplary only and not exhaustive, and this disclosure is not limited to the disclosed exemplary embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of this disclosure. Therefore, the scope of protection of this disclosure should be determined by the scope of the claims.

Claims

1. A method for action decision making for intelligent mobile platforms, characterized in that, The action decision-making method includes: The root node of the search tree is constructed based on the initial state information of the intelligent mobile platform, wherein the intelligent mobile platform is an intelligent platform with mobility capabilities, and the initial state information of the intelligent mobile platform includes at least one of the following: the location information of the intelligent mobile platform, the speed information of the location information of the intelligent mobile platform, and the context information related to action decision-making. A search tree is generated by expanding nodes downwards level by level from the root node. The search tree comprises N levels of nodes, from the first level to the Nth level, where the nodes in the first level correspond to the root node. N is an integer greater than or equal to 2. The optimal action sequence of the intelligent mobile platform is determined based on the generated search tree. The optimal action sequence is defined as an ordered combination of the optimal meta-actions of the intelligent mobile platform, and a meta-action is an atomic operation unit in the sequence decision-making process of the intelligent mobile platform. The steps involved in generating a search tree by expanding nodes downwards from the root node include: Based on the current state information of the current node in the current layer, generate one or more candidate action schemes corresponding to the current state of the current node in the current layer, where the current node is a non-leaf node; Determine the preferred candidate action scheme from the one or more candidate action schemes; The node in the next layer corresponding to the current node of the current layer and the node corresponding to the preferred candidate action scheme is determined as the non-leaf node of the next layer, and the node in the next layer corresponding to the remaining candidate action schemes other than the preferred candidate action scheme in the one or more candidate action schemes is determined as the leaf node of the next layer. The state information of the non-leaf nodes in the next layer is predicted based on the current state information and action history of the current node in the current layer, as well as the preferred candidate action schemes; and When the next layer corresponds to the Nth layer, stop expanding the node.

2. The action decision method of claim 1, wherein, The step of determining a preferred candidate action scheme from the one or more candidate action schemes includes: Receive high-level semantic instructions related to the decision intent corresponding to the current node of the current layer; and The candidate action scheme that matches the decision intention among the one or more candidate action schemes is determined as the preferred candidate action scheme.

3. The action decision method of claim 2, wherein, The step of determining the candidate action plan that matches the decision intention from among the one or more candidate action plans as the preferred candidate action plan includes: The decision intent is identified by inputting high-level semantic instructions corresponding to the decision intent into the inference model, and the degree of matching between the one or more candidate action schemes and the decision intent is evaluated; and The candidate action scheme with the highest matching degree among the one or more candidate action schemes is determined as the preferred candidate action scheme.

4. The action decision method of claim 3, wherein, Inference models include at least one of the following: multimodal large language models, expert systems, and scoring networks. The multimodal large language model is configured to identify the decision intent from the high-level semantic instructions corresponding to the decision intent and evaluate a first degree of matching between the one or more candidate action schemes and the decision intent. The expert system is configured to identify the decision intent from the high-level semantic instructions corresponding to the decision intent and evaluate a second degree of matching between the one or more candidate action schemes and the decision intent. The scoring network is configured to identify the decision intent from the high-level semantic instructions corresponding to the decision intent and evaluate a third degree of matching between the one or more candidate action schemes and the decision intent. The step of identifying the decision intent and evaluating the degree of matching between the one or more candidate action schemes and the decision intent by inputting high-level semantic instructions corresponding to the decision intent into the inference model includes: The matching degree between the one or more candidate action schemes and the decision intention is evaluated based on the matching degree corresponding to at least one of the multimodal large language model, the expert system, and the scoring network, among the first matching degree, the second matching degree, and the third matching degree.

5. The action decision method of claim 1, wherein, The steps for generating one or more candidate action schemes corresponding to the current state of the current node in the current layer include: Based on the current state information of the current node in the current layer and the set of meta-actions, a set of action descriptions that satisfy the candidate rules is generated. The set of meta-actions includes all possible atomic action types in the decision-making scenario. The candidate rules include at least one of enumeration-based policies, template-based policies, and sampling-based policies. The set of action descriptions is transformed into action scheme representations corresponding to the set of action descriptions by calling the generative model. The action scheme representation includes at least one of trajectory prediction, state prediction and visualization image. The one or more candidate action schemes are generated into a set of action descriptions and an action scheme representation corresponding to the set of action descriptions.

6. The action decision method of claim 5, wherein, The generative model includes at least one of a conditional generative model and a rule engine that incorporates domain knowledge.

7. The action decision method of claim 1, wherein, The steps for determining the optimal action sequence for an intelligent mobile platform based on the generated search tree include: The optimal action sequence of the intelligent mobile platform is determined by combining the preferred candidate action schemes corresponding to the root node of the first layer and the preferred candidate action schemes corresponding to the non-leaf nodes of the second to N-1 layers.

8. The action decision method of claim 1, wherein, At least two layers from the first to the Nth layer have different node degrees, where the node degree of a single layer represents the number of candidate action schemes that a single node of that single layer can generate.

9. An action decision apparatus for an intelligent mobile platform, characterized by, The action decision-making device includes: The search tree generation module is configured to: construct the root node of the search tree based on the initial state information of the intelligent mobile platform, and generate the search tree by expanding the nodes layer by layer downwards from the root node. The search tree includes nodes in N layers, where the N layers range from the first layer to the Nth layer, and the nodes in the first layer correspond to the root node. N is an integer greater than or equal to 2. The intelligent mobile platform is a mobile intelligent platform, and the initial state information of the intelligent mobile platform includes at least one of the following: the location information of the intelligent mobile platform, the speed information of the location information of the intelligent mobile platform, and context information related to action decisions. The optimal action sequence determination module is configured to determine the optimal action sequence of the intelligent mobile platform based on the generated search tree. The optimal action sequence of the intelligent mobile platform is defined as an ordered combination of the optimal meta-actions of the intelligent mobile platform, and a meta-action is an atomic operation unit in the sequence decision-making process of the intelligent mobile platform. The search tree generation module is configured as follows: Based on the current state information of the current node in the current layer, generate one or more candidate action schemes corresponding to the current state of the current node in the current layer, where the current node is a non-leaf node; Determine the preferred candidate action scheme from the one or more candidate action schemes; The node in the next layer corresponding to the current node of the current layer and the node corresponding to the preferred candidate action scheme is determined as the non-leaf node of the next layer, and the node in the next layer corresponding to the remaining candidate action schemes other than the preferred candidate action scheme in the one or more candidate action schemes is determined as the leaf node of the next layer. The state information of the non-leaf nodes in the next layer is predicted based on the current state information and action history of the current node in the current layer, as well as the preferred candidate action schemes; and When the next layer corresponds to the Nth layer, stop expanding the node.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the action decision method as described in any one of claims 1 to 8.

11. A computing device for an intelligent mobile platform, characterized in that, include: processor; A memory that stores a computer program, which, when executed by a processor, implements the action decision method as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Spacecraft sequence game method and device based on Monte Carlo tree search and medium

    CN116039956A

  • Automatic driving planning method based on autoregressive traffic flow deduction and tree search

    CN121180246A