Robot strategy generation method and device, computer and storage medium
By using confidence triggers and evaluating multiple candidate future trajectories, the problems of inaccurate decision-making and low efficiency of visual language models in complex tasks are solved, thereby improving the efficiency and robustness of robot operation.
Patent Information
- Application Number
- CN202511679270.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-14
- Publication Date
- 2026-01-13
AI Technical Summary
Existing visual language models suffer from inefficient and inaccurate state value assessment, random and short-sighted decision-making, and excessive inference latency when dealing with complex and long-cycle robot operation tasks, resulting in insufficient planning and decision-making capabilities.
The reliability of the initial plan is evaluated by a confidence trigger. If the confidence level is lower than the threshold, parallel candidate future trajectories are generated, the advantage is evaluated and the action decoding is corrected to generate the optimal action.
It improves the efficiency of task execution and the robustness of decision-making, accurately identifies and corrects low-quality action planning, and achieves a balance between efficiency and performance.
Smart Images

Figure CN121315969A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent device strategy optimization technology, and in particular to a robot strategy generation method, apparatus, computer and storage medium. Background Technology
[0002] In recent years, Visual Language Models (VLMs) incorporating massive amounts of internet knowledge have shown great potential in advanced robotic task planning. These models can understand visual scenes and natural language instructions and generate corresponding action plans. However, even state-of-the-art VLMs still face many difficulties when dealing with complex physical reasoning, especially when it comes to precise physical concepts and long-term planning. To improve the performance of pre-trained robot policies, "test-time scaling" has become a promising research direction, which optimizes and refines actions by utilizing additional computational resources during the deployment phase. The most similar approach to this invention is ReflectVLM, which proposes a two-phase test-time scaling framework to enhance the long-term planning capabilities of VLMs, including a proposal phase and a reflection phase. Its core idea is to train the VLM to critically evaluate and optimize its decisions by examining the possible future consequences of actions.
[0003] Existing visual language models (VLMs) exhibit significant limitations in their planning and decision-making capabilities when handling complex, long-cycle robotic maneuvers. Specifically, existing methods (such as reflective planning) face the following technical challenges when guiding VLMs to correct their behavior: State value assessment is inefficient and inaccurate: Existing methods rely on implicitly learning state value from noisy visual predictions of the future. This learning method is inefficient and the supervision signal is vague, making it difficult to accurately assess the long-term rewards of actions.
[0004] The randomness and short-sightedness of decision-making: Existing methods usually only evaluate a single, most greedy future path, ignoring other possibilities, which makes the decision highly sensitive to prediction errors and environmental changes, and lacks robustness.
[0005] Excessive inference latency: The serial "reasoning-imagination-re-reasoning" workflow forces reflection at each decision-making step, which greatly increases computational overhead and time latency, making it unsuitable for scenarios requiring rapid response.
[0006] In view of the shortcomings of the prior art, the present invention aims to propose a new robot strategy optimization framework to solve the above-mentioned technical problems. Summary of the Invention
[0007] This application provides a robot strategy generation method, apparatus, computer, and storage medium to solve the aforementioned technical problems existing in the prior art.
[0008] In view of the above, the first aspect of this application provides a robot policy generation method, the method comprising: Step S101: Based on the acquired current state image and target state image, the VLM policy network generates a candidate action list; Step S102: Take the candidate action with the highest probability in the candidate action list as the initial plan, and obtain the top-level hidden state vector of the last token of the VLM strategy website when generating the initial plan; Step S103: Use the top-level hidden state vector as the input of the confidence trigger, and determine whether to execute the initial plan based on the comparison result between the confidence of the initial plan output by the confidence trigger and the first preset threshold. Step S104: If the comparison result shows that the confidence level of the initial plan is lower than the first preset threshold, then a trajectory set containing several different, parallel candidate future trajectories is generated. Step S105: Evaluate the advantage of each candidate future trajectory in the trajectory set and determine the evaluation value of each candidate future trajectory; Step S106: The VLM policy network generates the output probability distribution of each candidate future trajectory based on the current state image, the target state image, the trajectory set, and the corresponding evaluation value. Step S107: Correct the action decoding based on the output probability distribution of each candidate future trajectory, generate the robot's optimal action, and execute it.
[0009] Optionally, after step S103, the method further includes: S108. If the comparison result shows that the confidence level of the initial plan is not lower than the first preset threshold, then the candidate action corresponding to the initial plan is executed as the optimal action of the robot.
[0010] Optionally, the confidence trigger is specifically a pre-trained multilayer perceptron classifier.
[0011] Optionally, step S105 further includes: The evaluation values of each candidate future trajectory are converted into natural language descriptive evaluation values.
[0012] Optionally, the evaluation value is specifically the amount of distance reduction from the current state to the future state.
[0013] Optionally, step S107 specifically includes: The decoding method is determined based on the similarity between the output probability distributions of each candidate future trajectory; If the similarity is lower than the second preset threshold, the decoding method is complementary decoding; otherwise, the decoding method is comparative decoding. The robot's optimal action is generated and executed by modifying the decoding method described above.
[0014] A second aspect of this application provides a robot strategy generation apparatus, the apparatus comprising: An initial plan processing unit is used to generate a candidate action list based on the acquired current state image and target state image by the VLM policy network; select the candidate action with the highest probability in the candidate action list as the initial plan, and obtain the top-level hidden state vector of the last token of the VLM policy network when generating the initial plan; use the top-level hidden state vector as the input of the confidence trigger, and determine whether to execute the initial plan based on the comparison result between the confidence of the initial plan output by the confidence trigger and a first preset threshold; The trajectory generation unit is used to generate a trajectory set containing several different, parallel candidate future trajectories if the comparison result shows that the confidence level of the initial plan is lower than the first preset threshold. A trajectory evaluation unit is used to evaluate the advantage of each candidate future trajectory in the trajectory set and determine the evaluation value of each candidate future trajectory. The decoding and optimization unit is used by the VLM policy network to generate the output probability distribution of each candidate future trajectory based on the current state image, the target state image, the trajectory set and the corresponding evaluation value; to modify the action decoding based on the output probability distribution of each candidate future trajectory, generate the robot's optimal action and execute it.
[0015] A third aspect of this application provides a robot, the robot including a processor and a memory: The memory is used to store program code and transmit the program code to the processor; The processor is configured to execute the steps of the robot strategy generation method as described in the first aspect above, according to the instructions in the program code.
[0016] A fourth aspect of this application provides a computer-readable storage medium for storing program code for performing the method described in the first aspect above.
[0017] As can be seen from the above technical solutions, the embodiments of this application have the following advantages: This application provides a robot strategy generation method, apparatus, computer, and storage medium. A confidence trigger determines whether the initial plan generated by the VLM (Virtual Machine Learning) strategy network should be directly executed, eliminating unnecessary re-evaluation and improving task execution efficiency. If the confidence level is below a preset threshold, a re-evaluation phase is initiated. This phase generates multiple candidate future trajectories guided by multiple candidate actions and evaluates the advantages of each candidate future trajectory, providing more accurate supervision signals and enhancing decision robustness. The evaluation results are used as feedback to the VLM strategy network, which aggregates the feedback information from all trajectories during the decoding phase to correct the initial plan, generate the optimal action, and execute it. This application can accurately identify and correct low-quality action planning, achieving a balance between efficiency and performance. Attached Figure Description
[0018] Figure 1 This is a flowchart of the robot strategy generation method in the embodiments of this application; Figure 2 This is a schematic diagram of the robot strategy generation device in the embodiments of this application; Figure 3 This is a schematic diagram of the robot in the embodiments of this application. Detailed Implementation
[0019] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present application.
[0020] The closest approximation to this invention is ReflectVLM (see Y. Feng et al.'s paper "Reflectiveplanning: Vision-language models for multi-stage long-horizon roboticmanipulation"). This approach proposes a two-stage test-time extension framework to enhance the long-term planning capabilities of VLM: Proposal Stage: VLM first generates a preliminary action based solely on the current and target visual states.
[0021] Reflection Stage: Subsequently, the framework uses a dynamic model (such as a diffusion model) to "imagine" the future visual state resulting from the initial action. Finally, VLM analyzes this imagined future and modifies its initial action accordingly.
[0022] The core idea of this method is to train the VLM to critically evaluate and optimize its decisions by examining the potential future consequences of actions.
[0023] However, this plan has certain drawbacks, including: 1. Ineffective Implicit Value Learning: ReflectVLM allows the VLM to implicitly learn the quality of an action from a predicted future image. This method of supervision is highly ambiguous and subjective, as the model may focus on task-irrelevant visual noise in the image, leading to inefficient learning and poor generalization.
[0024] 2. Lack of robustness in decision-making: This method evaluates only a single greedy future trajectory, and this single-path evaluation is highly stochastic. If the predicted single future is biased, or if the path is not the optimal choice, it will directly lead to incorrect decisions.
[0025] 3. High computational cost: Its fixed process of "reasoning-imagination-re-reasoning" turns single-step reasoning into serial multi-step reasoning, which significantly increases the delay of each decision-making step and results in low overall task execution efficiency.
[0026] In response to this, this application provides a robot strategy generation method to solve the above problems.
[0027] For easier understanding, please refer to Figure 1 , Figure 1 This is a flowchart of the robot policy generation method in the embodiments of this application, such as... Figure 1 As shown, specifically: Step S101: Based on the acquired current state image and target state image, the VLM policy network generates a candidate action list; It's important to note that in this step, the VLM policy network receives two key inputs: the current state image (a visual representation of the robot's current environment) and the target state image (a visual representation of the desired environmental state for the robot). Based on these two inputs, the VLM policy network, through its internal neural network structure and pre-trained knowledge, generates a list of possible actions as candidate actions. These actions are predicted based on the probability of transitioning from the current state to the target state.
[0028] Suppose a robot is currently on a cluttered table and needs to move a red ball to the other end of the table. The current state image is the cluttered scene on the table, and the target state image is the scene where the red ball is on the other end of the table. The VLM policy network may generate a list of candidate actions including "move arm to the left", "grab the red ball", "move arm to the right", etc.
[0029] Step S102: Take the candidate action with the highest probability in the candidate action list as the initial plan, and obtain the top-level hidden state vector of the last token of the VLM strategy website when generating the initial plan; It's important to note that the action with the highest probability is selected as the initial plan from the generated candidate action list. Furthermore, to subsequently evaluate the reliability of this initial plan, the top-level hidden state vector of the last token (usually the last element in the sequence) used by the VLM policy network to generate this plan needs to be obtained. This hidden state vector contains the network's internal state information at the time of action generation and can be used for subsequent confidence assessment.
[0030] Step S103: Use the top-level hidden state vector as the input of the confidence trigger, and determine whether to execute the initial plan based on the comparison result between the confidence of the initial plan output by the confidence trigger and the first preset threshold. It should be noted that the top-level hidden state vector obtained in step S102 is input into a confidence trigger, which outputs a confidence value representing the reliability of the initial plan. This confidence value is compared with a preset first threshold. If the confidence value is higher than or equal to the threshold, the initial plan is considered reliable enough and can be executed directly; otherwise, the subsequent optimization steps are performed.
[0031] In the example above, suppose the action "grabbing the red ball" has the highest probability and is selected as the initial plan. However, the confidence trigger outputs a confidence value of 0.7, while the first preset threshold is 0.8. Because 0.7 < 0.8, the initial plan "grabbing the red ball" is considered unreliable and needs further optimization.
[0032] In this embodiment, the confidence trigger can be a pre-trained multilayer perceptron (MLP) classifier used to evaluate the reliability of the initial plan. It learns from a large amount of sample data and is able to output a confidence value based on the input hidden state vector.
[0033] Step S104: If the comparison result shows that the confidence level of the initial plan is lower than the first preset threshold, then generate a trajectory set containing several different, parallel candidate future trajectories. It's important to note that when the confidence level of the initial plan falls below a threshold, more possibilities need to be explored. At this point, the system generates multiple distinct future trajectories, each containing a series of actions and a corresponding predicted image. These trajectories are generated in parallel to evaluate the impact of different action sequences on the future state.
[0034] In the example above, the system generated three different candidate future trajectories, each containing five actions. The action sequence for the first trajectory is "move arm to the left → grab the red ball → move arm to the right → put down the ball → return to the original position". The action sequence for the second trajectory is "extend arm directly → grab the red ball → quickly move to the target position → put down the ball". The action sequence for the third trajectory is "adjust camera angle → confirm the ball's position → grab the red ball → slowly move to the target position → put down the ball".
[0035] In this embodiment of the application, either a beam search multi-path planning strategy or a Monte Carlo tree search (MCTS) multi-path parallel exploration strategy can be selected to generate the trajectory set, wherein: Beam Search: Suitable for fixed-length sequence generation tasks, it can generate multiple candidate paths in parallel and has high computational efficiency, but may get stuck in local optima; Monte Carlo Tree Search (MCTS): Suitable for dynamic environments and long-term planning tasks, it can dynamically adjust the search depth and optimize the path by backtracking and updating, but it has a high computational cost.
[0036] Depending on the specific task requirements (such as task length, environmental dynamism, computing resources, etc.), a suitable multi-path search strategy can be selected.
[0037] Step S105: Evaluate the advantages of each candidate future trajectory in the trajectory set and determine the evaluation value of each candidate future trajectory; It should be noted that each candidate future trajectory is evaluated, and its advantage value is calculated. This advantage value can be based on various indicators, such as the reduction in distance from the current state to the future state, the efficiency of completing the task, energy consumption, etc. These evaluation values allow for comparison of the merits of different trajectories. In this application, the reduction in distance from the current state to the future state is used as the evaluation value.
[0038] For the three trajectories mentioned above, let's assume the evaluation value is the distance reduction. The distance reduction for the first trajectory is 80%, for the second trajectory it's 70%, and for the third trajectory it's 85%. Therefore, the third trajectory has the highest advantage value.
[0039] Furthermore, step S105 also includes: The evaluation values of each candidate future trajectory are converted into natural language descriptive evaluation values.
[0040] It should be noted that when evaluating candidate future trajectories, in addition to calculating numerical evaluation values (such as distance reduction), these evaluation values can also be converted into natural language descriptions for better understanding and interpretation.
[0041] For the three trajectories mentioned above, the evaluation values, when converted into natural language descriptions, might be: the first trajectory "good distance reduction, but many movements"; the second trajectory "rapid movements, but average distance reduction"; and the third trajectory "excellent distance reduction, smooth movements".
[0042] Step S106: The VLM policy network generates the output probability distribution of each candidate future trajectory based on the current state image, the target state image, the trajectory set, and the corresponding evaluation value. It should be noted that the VLM policy network combines the current state image, the target state image, each candidate future trajectory, and their evaluation values to recalculate the output probability distribution for each trajectory. This probability distribution reflects the likelihood of each trajectory achieving the target state.
[0043] In the example above, the VLM policy network calculated the output probability distributions for the three trajectories as follows: 0.2 for the first trajectory, 0.1 for the second trajectory, and 0.7 for the third trajectory. This indicates that the third trajectory is most likely to successfully complete the task.
[0044] Step S107: Correct the action decoding based on the output probability distribution of each candidate future trajectory, generate the robot's optimal action, and execute it.
[0045] It should be noted that, based on the output probability distribution of each candidate future trajectory, the trajectory with the highest probability is selected, and the specific action sequence is decoded from it. Then, this action sequence is modified to generate the robot's optimal action, which is then executed.
[0046] In the example above, based on the probability distribution described, the third trajectory is selected. The corrected action sequence might be "adjust the camera angle → confirm the ball's position → grab the red ball → slowly move to the target position → put the ball down". The robot executes this optimal action sequence and successfully moves the red ball to the target position.
[0047] Furthermore, step S107 specifically includes: The decoding method is determined based on the similarity between the output probability distributions of each candidate future trajectory; If the similarity is lower than the second preset threshold, the decoding method is complementary decoding; otherwise, the decoding method is comparative decoding. The robot's optimal action is generated and executed by correcting the decoding method.
[0048] It should be noted that when correcting the action decoding, the similarity between the output probability distributions of each candidate future trajectory is first compared. If the similarity is lower than the second preset threshold, it indicates that there are large differences between the trajectories. In this case, a complementary decoding method is used to combine the advantages of each trajectory to generate the optimal action. If the similarity is higher than or equal to the threshold, it indicates that the trajectories are relatively similar. In this case, a comparative decoding method is used to select the optimal trajectory and correct the action.
[0049] Similarity comparison can be achieved using methods such as Jensen-Shannon divergence, Kullback-Leibler divergence, multi-head attention, or Transformer architecture. By dynamically combining the output distributions of different future trajectories and selecting appropriate decoding strategies based on the distribution differences, the system can generate more robust correction actions.
[0050] Furthermore, step S103 includes the following: S108. If the comparison result shows that the confidence level of the initial plan is not lower than the first preset threshold, then the candidate action corresponding to the initial plan will be executed as the optimal action of the robot.
[0051] It should be noted that if the confidence level of the initial plan is higher than or equal to the first preset threshold in step S103, it means that the plan is reliable enough and can be executed directly without proceeding to the subsequent optimization steps.
[0052] Please see Figure 2 , Figure 2 This is a schematic diagram of the robot strategy generation device in an embodiment of this application, such as... Figure 2 As shown, specifically: The initial plan processing unit 201 is used to generate a candidate action list based on the acquired current state image and target state image by the VLM policy network; take the candidate action with the highest probability in the candidate action list as the initial plan, and obtain the top-level hidden state vector of the last token of the VLM policy website when generating the initial plan; take the top-level hidden state vector as the input of the confidence trigger, and determine whether to execute the initial plan based on the comparison result of the confidence of the initial plan output by the confidence trigger and the first preset threshold. The trajectory generation unit 202 is used to generate a trajectory set containing several different, parallel candidate future trajectories if the comparison result shows that the confidence level of the initial plan is lower than a first preset threshold. The trajectory evaluation unit 203 is used to evaluate the advantages of each candidate future trajectory in the trajectory set and determine the evaluation value of each candidate future trajectory. The decoding and optimization unit 204 is used by the VLM policy network to generate the output probability distribution of each candidate future trajectory based on the current state image, the target state image, the trajectory set and the corresponding evaluation value; to modify the action decoding based on the output probability distribution of each candidate future trajectory, generate the robot's optimal action and execute it.
[0053] Another embodiment of the present invention provides a robot, such as Figure 3 As shown, robot 10 includes: One or more processors 110 and memory 120, Figure 3 The following description uses a processor 110 as an example. The processor 110 and the memory 120 can be connected via a bus or other means. Figure 3 Taking the example of a connection between China and Israel via a bus.
[0054] Processor 110 is used to implement various control logics of robot 10. It can be a general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), microcontroller, ARM (Acorn RISC Machine) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination of these components. Furthermore, processor 110 can also be any conventional processor, microprocessor, or state machine. Processor 110 can also be implemented as a combination of computing devices, such as a combination of DSP and microprocessor, multiple microprocessors, one or more microprocessors combined with DSP and / or any other such configuration.
[0055] The memory 120, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions corresponding to the method for constructing the multilingual phoneme representation model in this embodiment of the invention. The processor 110 executes various functional applications and data processing of the robot 10 by running the non-volatile software programs, instructions, and units stored in the memory 120, thereby implementing the method for constructing the multilingual phoneme representation model in the above-described method embodiment.
[0056] The memory 120 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the device 10. Furthermore, the memory 120 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory 120 may optionally include memory remotely located relative to the processor 110, and these remote memories may be connected to the robot 10 via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0057] One or more units are stored in memory 120, and when executed by one or more processors 110, perform the following steps: Based on the acquired current state image and target state image, the VLM policy network generates a list of candidate actions; The candidate action with the highest probability in the candidate action list is taken as the initial plan, and the top-level hidden state vector of the last token of the VLM strategy website when generating the initial plan is obtained. The top-level hidden state vector is used as the input to the confidence trigger. The decision on whether to execute the initial plan is based on the comparison between the confidence of the initial plan output by the confidence trigger and the first preset threshold. If the comparison result shows that the confidence level of the initial plan is lower than the first preset threshold, then a trajectory set containing several different, parallel candidate future trajectories is generated. The advantage of each candidate future trajectory in the trajectory set is evaluated, and the evaluation value of each candidate future trajectory is determined. The VLM policy network generates the output probability distribution of each candidate future trajectory based on the current state image, the target state image, the trajectory set, and the corresponding evaluation value. The robot's optimal action is generated and executed by modifying the action decoding based on the output probability distribution of each candidate future trajectory.
[0058] This invention provides a non-volatile computer-readable storage medium storing computer-executable instructions. When these computer-executable instructions are executed by one or more processors, they implement any one of the robot strategy generation methods described in the above embodiments.
[0059] As examples, non-volatile storage media can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) as external cache memory. By way of illustration and not limitation, RAM can be obtained in many forms such as synchronous RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), Synchlink DRAM (SLDRAM), and direct Rambus RAM (DRRAM). The memory components or memories disclosed in the operating environment described herein are intended to include one or more of these and / or any other suitable types of memory.
[0060] This application provides a robot strategy generation method, apparatus, computer, and storage medium. A confidence trigger determines whether the initial plan generated by the VLM (Virtual Machine Learning) strategy network should be directly executed, eliminating unnecessary re-evaluation and improving task execution efficiency. If the confidence level is below a preset threshold, a re-evaluation phase is initiated. This phase generates multiple candidate future trajectories guided by multiple candidate actions and evaluates the advantages of each candidate future trajectory, providing more accurate monitoring signals and enhancing decision robustness. The evaluation results are used as feedback to the VLM strategy network, which aggregates the feedback information from all trajectories during the decoding phase to correct the initial plan, generate the optimal action, and execute it. This application can accurately identify and correct low-quality action planning, achieving a balance between efficiency and performance.
[0061] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0062] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented, for example, in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0063] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0064] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between apparatuses or units through some interfaces, and may be electrical, mechanical, or other forms.
[0065] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0066] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0067] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes: USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, optical disks, and other media capable of storing program code.
[0068] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
[0069] It should be noted that if any software tools or components not belonging to our company appear in the embodiments of this application, they are merely for illustrative purposes and do not represent actual use.
Claims
1. A robot strategy generation method, characterized in that, include: Step S101: Based on the acquired current state image and target state image, the VLM policy network generates a candidate action list; Step S102: Take the candidate action with the highest probability in the candidate action list as the initial plan, and obtain the top-level hidden state vector of the last token of the VLM strategy website when generating the initial plan; Step S103: Use the top-level hidden state vector as the input of the confidence trigger, and determine whether to execute the initial plan based on the comparison result between the confidence of the initial plan output by the confidence trigger and the first preset threshold. Step S104: If the comparison result shows that the confidence level of the initial plan is lower than the first preset threshold, then a trajectory set containing several different, parallel candidate future trajectories is generated. Step S105: Evaluate the advantage of each candidate future trajectory in the trajectory set and determine the evaluation value of each candidate future trajectory; Step S106: The VLM policy network generates the output probability distribution of each candidate future trajectory based on the current state image, the target state image, the trajectory set, and the corresponding evaluation value. Step S107: Correct the action decoding based on the output probability distribution of each candidate future trajectory, generate the robot's optimal action, and execute it.
2. The robot strategy generation method according to claim 1, characterized in that, Following step S103, the following is also included: S108. If the comparison result shows that the confidence level of the initial plan is not lower than the first preset threshold, then the candidate action corresponding to the initial plan is executed as the optimal action of the robot.
3. The robot strategy generation method according to claim 1, characterized in that, The confidence trigger is specifically a pre-trained multilayer perceptron classifier.
4. The robot strategy generation method according to claim 1, characterized in that, Step S105 further includes: The evaluation values of each candidate future trajectory are converted into natural language descriptive evaluation values.
5. The robot strategy generation method according to claim 1, characterized in that, The evaluation value is specifically the amount of distance reduction from the current state to the future state.
6. The robot strategy generation method according to claim 1, characterized in that, Step S107 specifically involves: The decoding method is determined based on the similarity between the output probability distributions of each candidate future trajectory; If the similarity is lower than the second preset threshold, the decoding method is complementary decoding; otherwise, the decoding method is comparative decoding. The robot's optimal action is generated and executed by modifying the decoding method described above.
7. A robot strategy generation device, characterized in that, include: The initial planning processing unit is used to generate a list of candidate actions for the VLM policy network based on the acquired current state image and target state image. The candidate action with the highest probability in the candidate action list is taken as the initial plan, and the top-level hidden state vector of the last token of the VLM strategy website when generating the initial plan is obtained; the top-level hidden state vector is taken as the input of the confidence trigger, and the initial plan is executed based on the comparison result between the confidence of the initial plan output by the confidence trigger and the first preset threshold. The trajectory generation unit is used to generate a trajectory set containing several different, parallel candidate future trajectories if the comparison result shows that the confidence level of the initial plan is lower than the first preset threshold. A trajectory evaluation unit is used to evaluate the advantage of each candidate future trajectory in the trajectory set and determine the evaluation value of each candidate future trajectory. The decoding and optimization unit is used by the VLM policy network to generate the output probability distribution of each candidate future trajectory based on the current state image, the target state image, the trajectory set and the corresponding evaluation value; to modify the action decoding based on the output probability distribution of each candidate future trajectory, generate the robot's optimal action and execute it.
8. The robot strategy generation apparatus according to claim 7, characterized in that, Also includes: An execution unit is configured to execute the candidate action corresponding to the initial plan as the optimal action of the robot if the comparison result shows that the confidence level of the initial plan is not lower than the first preset threshold.
9. A robot, characterized in that, The device includes a processor and a memory: The memory is used to store program code and transmit the program code to the processor; The processor is configured to execute the robot strategy generation method according to any one of claims 1-6 according to the instructions in the program code.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store program code for executing the robot strategy generation method according to any one of claims 1-6.
Citation Information
Cited By
Method and device for determining execution strategy, equipment, storage medium and program product
CN121893298A