An Embodied Intelligence Robot Perception and Decision-making Method and System for Complex Scenarios
Through self-supervised learning and Q-learning combined with UCB algorithm, random confidence thresholds are generated for decision optimization, which solves the problem of embodied intelligent robots over-relying on known optimal actions in complex scenarios, and improves perception accuracy and task execution efficiency.
Patent Information
- Application Number
- CN202510600685.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-05-12
AI Technical Summary
In the prior art, embodied intelligent robots are prone to over-rely relying on known optimal actions in complex scenarios, resulting in falling into local optimal solutions, difficulty in dealing with changing environments and real-time decision-making needs, lacking the comprehensive processing ability of multimodal information, resulting in limited adaptability and decision-making accuracy.
Environmental information is obtained through self-supervised learning, a state action table is generated using the Q-learning algorithm, and the action confidence is calculated in combination with the UCB algorithm to generate a random confidence threshold, and filter the optimization confidence threshold for perceived decisions, and optimize the decision process.
It improves the perceived accuracy and robustness of embodied intelligent robots in dynamic environments, enhances their adaptability and task execution efficiency in complex scenarios, avoids local optimal solutions, and achieves more efficient decision-making and task completion.
Smart Images

Figure CN120103719B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent decision-making, and particularly to a perception and decision-making method and system for an embodied intelligent robot facing complex scenarios. Background Art
[0002] In the current existing technologies, the perception and decision-making ability of an embodied intelligent robot facing complex scenarios mainly relies on traditional perception and decision-making algorithms. These methods usually analyze through static data models and are difficult to effectively handle the changing environment and real-time decision-making requirements in complex scenarios. The disadvantages of the existing technologies include: insufficient response ability to dynamic changes, poor real-time performance, limited processing ability. Especially when facing a rapidly changing environment, problems such as decision-making lag or insufficient accuracy are likely to occur. In addition, these technologies usually lack the ability to comprehensively process multi-modal information and are difficult to accurately capture and understand complex perception information, resulting in limitations in the adaptability and decision-making accuracy of the robot in complex scenarios.
[0003] Traditional Q-learning mainly selects actions by maximizing the current state-action value. This method easily causes the embodied intelligent robot to overly rely on known optimal actions and thus fall into local optimal solutions. This limitation stems from the fact that Q-learning overly focuses on the current reward information and ignores the exploration of other possible actions. Therefore, the embodied intelligent robot may miss those actions with lower initial rewards but greater long-term potential, restricting the diversity of strategies and the achievement of global optimal solutions.
[0004] The current existing technologies have the problem that the embodied intelligent robot overly relies on known optimal actions and thus falls into local optimal solutions. Summary of the Invention
[0005] The present invention provides a perception and decision-making method for an embodied intelligent robot facing complex scenarios, and its main purpose is to solve the problem that the embodied intelligent robot overly relies on known optimal actions and thus falls into local optimal solutions.
[0006] In a first aspect, to achieve the above object, a perception and decision-making method for an embodied intelligent robot facing complex scenarios provided by the present invention includes:
[0007] Obtain the initial perception data of a preset embodied intelligent robot for a preset scenario, and perform self-supervised learning on the initial perception data to obtain the environmental information of the preset scenario;
[0008] Obtain the target task of the preset embodied intelligent robot, and perform an initial decision on the environmental information according to the target task to obtain a state-action table;
[0009] Obtain the selection times of each action in the state-action table, and generate the initial confidence of each action according to the selection times;
[0010] Generate a number of different random confidence thresholds to be screened by using the initial confidence and a preset variation parameter;
[0011] Perform policy selection on each action in the state-action table according to the confidence thresholds to be screened, and obtain a number of policy scores;
[0012] Use the policy scores to screen out the optimized confidence threshold, and use the optimized confidence threshold to perform perception decision on the environmental information to obtain the final decision result.
[0013] In a second aspect, the present invention also provides an embodied intelligent robot perception decision-making system for complex scenarios, and the system includes:
[0014] An environmental perception module, configured to obtain the initial perception data of a preset embodied intelligent robot for a preset scenario, and perform self-supervised learning on the initial perception data to obtain the environmental information of the preset scenario;
[0015] An initial decision-making module, configured to obtain the target task of the preset embodied intelligent robot, and perform an initial decision on the environmental information according to the target task to obtain a state-action table;
[0016] An initial confidence generation module, configured to obtain the selection times of each action in the state-action table, and generate the initial confidence of each action according to the selection times;
[0017] A confidence threshold to be screened generation module, configured to generate a number of different random confidence thresholds to be screened by using the initial confidence and a preset variation parameter;
[0018] A confidence threshold to be screened selection module, configured to perform policy selection on each action in the state-action table according to the confidence thresholds to be screened, and obtain a number of policy scores;
[0019] A final decision-making module, configured to use the policy scores to screen out the optimized confidence threshold, and use the optimized confidence threshold to perform perception decision on the environmental information to obtain the final decision result.
[0020] The present invention obtains the initial perception data of a preset embodied intelligent robot for a preset scenario, performs self-supervised learning on the initial perception data to obtain the environmental information of the preset scenario, obtains the target task of the preset embodied intelligent robot, makes an initial decision on the environmental information according to the target task to obtain a state-action table, obtains the number of times each action in the state-action table is selected, and generates an initial confidence level for each action according to the number of selections. Using the initial confidence level and a preset change parameter, a number of different random confidence threshold values to be screened are generated. Policy selection is performed on each action in the state-action table according to the confidence threshold values to be screened to obtain a number of policy scores. The optimized confidence threshold value is screened out using the policy scores, and the environmental information is perceptually decided using the optimized confidence threshold value to obtain a final decision result, effectively solving the problem that the embodied intelligent robot overly relies on known optimal actions and thus falls into a local optimal solution. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0022] Figure 1 It is a flowchart of a method for perceptual decision-making of an embodied intelligent robot for complex scenarios provided by an embodiment of the present invention;
[0023] Figure 2 It is a block diagram of a system for perceptual decision-making of an embodied intelligent robot for complex scenarios provided by an embodiment of the present invention;
[0024] The realization, functional characteristics, and advantages of the objectives of the present invention will be further described in conjunction with the embodiments with reference to the drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0025] In order to enable those skilled in the art to better understand the technical solutions of the present disclosure, and to fully understand how the present disclosure uses technical means to solve technical problems and the implementation process of achieving corresponding technical effects and be able to implement accordingly, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all of the embodiments. The embodiments of the present disclosure and each feature in the embodiments can be combined with each other on the premise of not conflicting, and the formed technical solutions are all within the protection scope of the present disclosure. Based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present disclosure.
[0026] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily limit to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0027] The embodiment of the present application provides a perception and decision-making method for an embodied intelligent robot for complex scenarios. The perception and decision-making method for an embodied intelligent robot for complex scenarios can be executed by software or hardware installed on a terminal device or a server device. The server includes but is not limited to: a single server, a server cluster, a cloud server, or a cloud server cluster, etc. The server can be an independent server or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.
[0028] Refer to Figure 1 As shown, it is a schematic flowchart of a perception and decision-making method for an embodied intelligent robot for complex scenarios provided by an embodiment of the present invention. In this embodiment, the perception and decision-making method for an embodied intelligent robot for complex scenarios includes:
[0029] S1. Obtain the initial perception data of a preset embodied intelligent robot for a preset scenario, and perform self-supervised learning on the initial perception data to obtain the environmental information of the preset scenario.
[0030] In the embodiments of the present invention, the embodied intelligent robot has a perception ability. The perception devices of the embodied intelligent robot, such as visual sensors, auditory sensors, tactile sensors, position sensors, etc., collect the initial perception data in a complex environment through these perception devices to assist the embodied intelligent robot in perceiving objects, obstacles, people, etc. in space. By learning the similarities and differences of the initial perception data, the embodied intelligent robot can automatically obtain effective features from the unlabeled initial perception data to obtain environmental information.
[0031] Specifically, the performing self-supervised learning on the initial perception data to obtain the environmental information of the preset scenario includes:
[0032] Perform data alignment and timestamp synchronization on the initial perception data to obtain synchronized perception data;
[0033] Obtain positive samples, negative samples, and a preset comparison factor in the known perception data, and generate a contrast loss function according to the positive samples, the negative samples, and the preset comparison factor;
[0034] Minimize the contrast loss function to obtain a minimum loss function;
[0035] Optimize a preset self-supervised learning model using the minimum loss function to obtain a contrast learning model;
[0036] Analyze the similarity between the synchronized perception data and the positive samples using the contrast learning model to obtain environmental similarity features;
[0037] Generate the environmental information of the preset scenario according to the environmental similarity features.
[0038] Specifically, the analyzing the similarity between the synchronized perception data and the positive samples using the contrast learning model to obtain environmental similarity features includes:
[0039] Compare the synchronized perception data and the positive samples using the contrast learning model to obtain environmental contrast features;
[0040] Calculate the similarity between each environmental contrast feature and the positive samples one by one;
[0041] Select the positive samples whose similarity is greater than a preset similarity threshold;
[0042] Extract the environmental similarity features similar to the synchronized perception data from the selected positive samples.
[0043] Specifically, the initial perception data in a complex environment is collected by a perception device, and the data alignment and timestamp synchronization are performed on the initial perception data to obtain synchronized perception data. The specific operation steps are as follows: obtain the timestamp of each initial perception data, summarize the data with the same timestamp, and sort each initial perception data in the order of the timestamps to obtain synchronized perception data.
[0044] Obtain the positive samples, negative samples, and a preset comparison factor in the known perception data, and generate a contrast loss function according to the positive samples, the negative samples, and the preset comparison factor. Among them, the positive samples represent data from different perspectives or different timestamps of the same data source and the same instance. For example, images from the same object captured at different times can be regarded as positive samples. The negative samples represent the perception data from different objects, different classes, or different environments, and are samples with a large difference from the current initial perception data. The calculation formula is as follows:
[0045]
[0046] Among them, represents the th positive sample, represents the th positive sample, represents the th negative sample, represents the preset comparison factor, represents the total number of negative samples.
[0047] Minimize the contrast loss function to obtain the minimum loss function, and use the minimum loss function to optimize a preset self-supervised learning model to obtain a contrast learning model. The calculation formula is as follows:
[0048]
[0049]
[0050] Among them, represents the th positive sample, represents the th positive sample, represents the th negative sample, represents the preset comparison factor, represents the total number of negative samples, represents the preset self-supervised learning model.
[0051] Analyze the similarity between the synchronous perception data and the positive samples using the contrastive learning model to obtain environmental similarity features, and generate environmental information for the preset scenario based on the environmental similarity features. The specific operation is as follows: Input the synchronous perception data into the contrastive learning model to obtain environmental contrast features, calculate the similarity between the environmental contrast features and each feature of the positive samples, extract the features with similarity greater than the preset similarity threshold to obtain the environmental similarity features of the synchronous perception data. The calculation formula for similarity is as follows:
[0052]
[0053] Among them, represents the positive sample, represents the environmental contrast features of the synchronous perception data.
[0054] Generate the environmental information for the preset scenario based on the environmental similarity features, including: Environmental status: such as maps, obstacles, paths, terrain, etc. Status changes: such as whether the path is blocked, the appearance and movement of obstacles, etc.
[0055] Effectively process the initial perception data through self-supervised learning, enabling the embodied intelligent robot to learn the environmental similarity features from the synchronous perception data without labeled data. By optimizing the contrastive learning model, the robot can more accurately identify and analyze the important features in the environment, thereby improving the accuracy and robustness of perception and decision-making. It can not only enable the embodied intelligent robot to self-adjust the perception model in a dynamic environment but also enhance its adaptability in different scenarios, enabling it to perform tasks and make decisions more efficiently in complex environments.
[0056] S2. Obtain the target task of the preset embodied intelligent robot, and make an initial decision on the environmental information according to the target task to obtain a state-action table.
[0057] In the embodiment of the present invention, obtain the target task of the embodied intelligent robot, and make an initial decision on the environmental information according to the target task based on the Q-learning algorithm to obtain a state-action table. Among them, the state-action table includes: The state of the embodied intelligent robot, such as: the position, direction, etc. of the embodied intelligent robot; The actions of the embodied intelligent robot, such as: the embodied intelligent robot selects to move forward, backward, turn, etc.
[0058] Specifically, making the initial decision on the environmental information according to the target task to obtain a state-action table includes:
[0059] Obtain the initial state and initial action of the preset embodied intelligent robot;
[0060] Randomly combine the initial state and the initial actions, and summarize each combination into an initial combination table;
[0061] Generate a target action path according to the target task and the environmental information;
[0062] Update the initial state according to the target action path to obtain an updated state;
[0063] Generate an initial target action for the target action path according to the updated state;
[0064] Obtain a preset learning rate, a preset discount factor, and the reward value of each initial target action;
[0065] Obtain the update probability of obtaining the updated state, and calculate the state-action value of each combination in the initial combination table according to the update probability;
[0066] Update the state-action value by using the preset learning rate, the preset discount factor, and the reward value to obtain a state-action table.
[0067] Specifically, obtain the target task of the preset embodied intelligent robot, where the target task is to enable the embodied intelligent robot to interact with a preset environment, such as completing a certain path planning, reaching a certain specified position, or completing a certain specific operation.
[0068] Generate a target action path according to the target task and the environmental information, and use the action corresponding to the target action path in the initial combination table as the initial target action. The specific operation steps are as follows: Determine the starting point and the target point of the embodied robot, analyze each action and state on each path, consider the feasibility and efficiency of the path, and select the path with the highest feasibility as the target action path, and use the actions on the target action path as the initial target actions.
[0069] The initial combination table is a two-dimensional matrix, where the elements of the two-dimensional matrix are composed of each state and action of the embodied intelligent robot. Obtain the update probability of obtaining the updated state, and calculate the state-action value of each combination in the initial combination table according to the update probability. The calculation formula is as follows:
[0070]
[0071] where, represents the probability of transferring to the updated state when taking the initial action in the initial state , represents the initial state, represents the initial action, represents the updated state, Represents the initial target action, Represents the preset discount factor, Represents the reward value, Represents the maximum state-action value of all initial target actions in the updated state below.
[0072] Update the state-action value using the preset learning rate, the preset discount factor, and the reward value. By continuously interacting with the environmental information and updating the state-action value, obtain the final state-action table. The calculation formula is as follows:
[0073]
[0074] Among them, Represents the state-action value of taking the initial action in the initial state , Represents the initial state, Represents the initial action, Represents the updated state, Represents the initial target action, Represents the preset discount factor, Represents the reward value, Represents the maximum state-action value of all initial target actions in the updated state below, Represents the preset learning rate.
[0075] By randomly combining the initial state and actions, and gradually selecting the most suitable actions according to the target task and environmental information, the embodied intelligent robot can achieve autonomous decision-making in a dynamic environment. By updating the initial combination table using the learning rate, discount factor, and reward value, the robot can continuously optimize its action strategy, gradually improve the efficiency and accuracy of task completion, enhance the adaptive ability of the embodied intelligent robot, enable it to effectively perceive and make decisions in a changing environment, and thus better execute complex tasks.
[0076] S3. Obtain the selection times of each action in the state-action table, and generate the initial confidence of each action according to the selection times.
[0077] In the embodiment of the present invention, generating the initial confidence according to the selection times of each action in the state-action table represents the degree of trust of each action, providing data support for better decision-making of the subsequent embodied intelligent robot.
[0078] Specifically, the obtaining the selection times of each action in the state-action table and generating the initial confidence of each action according to the selection times includes:
[0079] Randomly select one state from the state-action table as the target state;
[0080] Count the number of times the target state is selected;
[0081] Count the number of times the action corresponding to the target state in the state-action table is selected;
[0082] Obtain the state-action value of the target state and the action corresponding to the target state, and calculate the initial confidence of each action in the state-action table according to the state-action value, the state selection times, and the action selection times.
[0083] Specifically, calculate the confidence of each action through the UCB algorithm, providing a data basis for the subsequent embodied intelligent robot in the exploration stage. When an action is selected a small number of times, the confidence will be adjusted higher to encourage more exploration; while when an action is selected multiple times, the confidence gradually decreases to avoid local optima.
[0084] Obtain the state-action value of the target state and the action corresponding to the target state, and calculate the initial confidence of each action in the state-action table according to the state-action value, the state selection times, and the action selection times. The calculation formula is as follows:
[0085]
[0086] where, represents the state-action value of taking the action corresponding to the target state under the target state under the target state of the action corresponding to the target state, represents the target state, represents the action corresponding to the target state, represents the state selection times, represents the action selection times, represents under the target state of taking the action corresponding to the target state of the initial confidence.
[0087] By calculating the selection times of each action and its relevance to the target state, it helps the embodied intelligent robot evaluate and update the confidence of actions. Through the combination of state-action value, state selection times, and action selection times, the robot can optimize the decision-making process based on historical experience and action feedback, so as to more accurately select the optimal action when facing complex tasks, improve the learning efficiency and adaptability of the robot in a dynamic environment, be able to better make autonomous decisions in a changing scenario, and increase the success rate of task execution.
[0088] S4. Generate a number of different random confidence threshold values to be screened using the initial confidence level and a preset variation parameter.
[0089] In an embodiment of the present invention, a number of different random confidence threshold values to be screened are generated using the initial confidence level and a preset variation parameter through a normal distribution, providing a data basis for subsequently screening out the optimal confidence threshold value for decision-making.
[0090] Specifically, the generating a number of different random confidence threshold values to be screened using the initial confidence level and a preset variation parameter includes:
[0091] Calculate the average confidence level of the initial confidence level;
[0092] Take the preset variation parameter as the standard deviation of the initial confidence level;
[0093] Generate a number of different random confidence threshold values to be screened using the average confidence level and the standard deviation.
[0094] Specifically, calculate the average confidence level of the initial confidence level, take the preset variation parameter as the standard deviation of the initial confidence level, and generate a number of different random confidence threshold values to be screened through a normal distribution using the average confidence level and the standard deviation. The calculation formula is as follows:
[0095]
[0096]
[0097]
[0098] Wherein, represents the initial confidence level for taking the action corresponding to the target state in the target state , represents a normal distribution with a mean of and a standard deviation of . The confidence level range is described using the average confidence level and the standard deviation through a normal distribution, and multiple values are randomly selected from this confidence level range as the confidence threshold values to be screened. The confidence threshold values to be screened are used for further decision-making to increase the diversity of decision-making.
[0099] By utilizing the average value and standard deviation of the initial confidence, and combining with a preset variation parameter, multiple random confidence thresholds to be screened are generated, enabling the embodied intelligent robot to introduce a certain degree of randomness and diversity in the decision-making process, which helps the robot explore and optimize the decision-making strategy under different environmental conditions and avoid falling into local optimal solutions. By adjusting the confidence threshold, the robot can more flexibly respond to environmental changes, improve the robustness and adaptability of decision-making, and thus make more appropriate action selections in complex or uncertain task scenarios.
[0100] S5. Perform policy selection on each action in the state-action table according to the confidence threshold to be screened, and obtain a number of policy scores.
[0101] In the embodiment of the present invention, performing policy selection on each action in the state-action table according to the confidence threshold to be screened, and obtaining a number of policy scores provides a data analysis basis for subsequent screening of the optimized confidence.
[0102] Specifically, the performing policy selection on each action in the state-action table according to the confidence threshold to be screened, and obtaining a number of policy scores includes:
[0103] Obtain the screening task of the preset embodied intelligent robot and the screening action path of the screening task;
[0104] Select one of the confidence thresholds to be screened one by one as the target confidence threshold;
[0105] Select one action in the state-action table one by one as the target screening action;
[0106] Obtain the initial confidence of each target screening action, and determine whether the initial confidence of the target screening action is greater than or equal to the target confidence threshold;
[0107] If the initial confidence of the target screening action is less than the target confidence threshold, randomly select an action in the state-action table for exploration;
[0108] If the initial confidence of the target screening action is greater than or equal to the target confidence threshold, calculate the state-action update value of the target screening action, and screen out the action corresponding to the maximum state-action update value to obtain the selected action;
[0109] Use the selected action to obtain the completion time of the screening action path;
[0110] Obtain the length of the screening action path, and calculate the policy score of each target confidence threshold in the screening action path according to the selected action, the completion time, and the length.
[0111] Specifically, the state-action update value of each action in the state-action table is calculated using the formula in S2 to obtain the selected action. The calculation formula is as follows:
[0112]
[0113] where represents the probability of transferring to the state corresponding to the selected action when taking the target screening action in the screening state ; represents the screening state, represents the target screening action, represents the state corresponding to the selected action, represents the selected action, represents the preset discount factor, represents the selection reward value of the selected action, and represents the state-action update value of all selected actions in the state corresponding to the selected action .
[0114] Obtain the length of the screening action path, and calculate the policy score of each target confidence threshold in the screening action path according to the selected action, the completion time, and the length. The calculation formula is as follows:
[0115]
[0116] where represents the completion time, represents the length, represents the discount factor of the completion time, and
[0117]
[0118] By comprehensively considering the target task, the confidence threshold, the action selection in the state-action table, and the policy evaluation, the embodied intelligent robot can be more accurate and efficient when performing complex tasks. The embodied intelligent robot makes different action selections for different confidence thresholds, calculates the completion time and length of the target action path, further helps to evaluate the actual effects of the policies of different confidence thresholds, and screens out the optimized confidence threshold according to the policy score to achieve a better task completion path and a higher success rate.S6. Screen out the optimized confidence threshold using the policy score, and perform perceptual decision-making on the environmental information using the optimized confidence threshold to obtain the final decision result.
[0119] In the embodiments of the present invention, an optimized confidence threshold is selected from a number of confidence thresholds to be screened through a scatter plot, and the optimized confidence threshold is used to help the embodied intelligent robot obtain decision-making actions, better perceive and make decisions on environmental information, and obtain the final decision result.
[0120] Specifically, the step of selecting the optimized confidence threshold by using the policy score includes:
[0121] Generating a scatter plot of the policy score and the confidence threshold to be screened corresponding to the policy score;
[0122] Taking the confidence threshold to be screened corresponding to the point with the highest vertical axis in the scatter plot as the optimized confidence threshold.
[0123] Specifically, the step of using the optimized confidence threshold to perceive and make decisions on the environmental information to obtain the final decision result includes:
[0124] Generating a decision action path of the preset embodied intelligent robot according to the environmental information and the target task;
[0125] Obtaining the actions to be screened according to the decision action path, and calculating the action confidence of the actions to be screened;
[0126] Selecting the actions to be screened whose action confidence is greater than or equal to the optimized confidence threshold as decision-making actions;
[0127] Generating the final decision result of the preset embodied intelligent robot according to the decision-making actions, the decision action path and the environmental information.
[0128] Specifically, the horizontal axis of the scatter plot represents the value of each confidence threshold to be screened, and the vertical axis represents the policy score corresponding to each confidence threshold to be screened, so that the policy scores corresponding to each confidence threshold to be screened can be directly compared, and the confidence threshold to be screened corresponding to the point with the highest vertical axis is selected as the optimized confidence threshold.
[0129] The action confidence of the actions to be screened is obtained according to the initial confidence calculation formula in S3, the action confidence is compared with the selected optimized confidence threshold, and the actions to be screened whose action confidence is greater than or equal to the optimized confidence threshold are selected as decision-making actions, and the initial target actions with an initial confidence less than the optimized confidence threshold are explored.
[0130] Generate the final decision result of the preset embodied intelligent robot according to the decision-making action, the decision-making action path, and the environmental information. In a specific embodiment, the environmental information is as follows: The embodied intelligent robot is in a warehouse with a flat ground but multiple obstacles, such as stacked goods, shelves, etc. The perception data is that the sensor detects the position where the goods are stacked, which is about 2 meters away from the embodied intelligent robot. There is a shelf about 3 meters in front of the embodied intelligent robot and it cannot pass directly. There is enough space on the left side and it can bypass. The target task is to carry an item from the current position to the other side of the warehouse. The target position is at the far end of the warehouse. The embodied intelligent robot needs to bypass the obstacles and carry the item smoothly. The decision-making action path is that the embodied intelligent robot starts from the current position, moves forward, and when encountering an obstacle, it chooses to retreat and turn left to bypass the shelf. The decision-making actions include moving forward, retreating, turning, etc. The final decision result is as follows: Move forward: The robot will first try to move forward 2 meters to approach the obstacle position. Retreat: After encountering the obstacle, the robot judges that it cannot pass and immediately retreats 1 meter to prepare for the subsequent detour. Turn: The robot recognizes that there is enough space on the left side to detour, so it decides to turn left 90 degrees and continue to move forward. Continue to move forward: After the robot bypasses the obstacle, it continues to move forward along the planned path until it reaches the target position.
[0131] By analyzing the relationship between the policy score and the confidence threshold, the embodied intelligent robot can find the optimal confidence threshold, so as to be more sensitive and effective in perceiving and responding to environmental information when performing tasks. Actions greater than or equal to the optimized confidence threshold ensure that the robot has a high level of confidence when making decisions, and actions less than the optimized confidence threshold continue to be explored, thus avoiding falling into local optimal solutions. At the same time, by integrating decision-making actions, target action paths, and environmental information, the robot can generate more efficient and accurate final decision results, improve the success rate and efficiency of task completion, enable the embodied intelligent robot to make optimized decisions in complex environments, and enhance its intelligent level and execution ability.
[0132] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0133] As Figure 2 shown, it is a functional module diagram of an embodied intelligent robot perception and decision-making system provided by an embodiment of the present invention for complex scenarios.
[0134] In the embodiments of the present disclosure, an embodied intelligent robot perception and decision-making system for complex scenarios is provided. This embodied intelligent robot perception and decision-making system for complex scenarios corresponds one-to-one with the above-mentioned embodiment of an embodied intelligent robot perception and decision-making method for complex scenarios. As Figure 2As shown in the figure, the embodied intelligent robot perception and decision-making system 100 for complex scenarios includes an environmental perception module 101, an initial decision-making module 102, an initial confidence generation module 103, a confidence threshold to be screened generation module 104, a confidence threshold to be screened selection module 105, and a final decision-making module 106. The detailed description of each functional module is as follows:
[0135] The environmental perception module 101 is used to obtain the initial perception data of a preset embodied intelligent robot for a preset scenario, perform self-supervised learning on the initial perception data, and obtain the environmental information of the preset scenario;
[0136] The initial decision-making module 102 is used to obtain the target task of the preset embodied intelligent robot, perform an initial decision on the environmental information according to the target task, and obtain a state-action table;
[0137] The initial confidence generation module 103 is used to obtain the selection times of each action in the state-action table, and generate the initial confidence of each action according to the selection times;
[0138] The confidence threshold to be screened generation module 104 is used to generate a number of different random confidence thresholds to be screened by using the initial confidence and a preset change parameter;
[0139] The confidence threshold to be screened selection module 105 is used to perform policy selection on each action in the state-action table according to the confidence threshold to be screened, and obtain a number of policy scores;
[0140] The final decision-making module 106 is used to screen out the optimized confidence threshold by using the policy scores, and perform perception decision on the environmental information by using the optimized confidence threshold to obtain the final decision result.
[0141] In an embodiment, when the environmental perception module 101 performs self-supervised learning on the initial perception data to obtain the environmental information of the preset scenario, it is used for:
[0142] Perform data alignment and timestamp synchronization on the initial perception data to obtain synchronized perception data;
[0143] Obtain positive samples, negative samples, and a preset comparison factor in the known perception data, and generate a contrast loss function according to the positive samples, the negative samples, and the preset comparison factor;
[0144] Minimize the contrast loss function to obtain the minimum loss function;
[0145] Optimize a preset self-supervised learning model by using the minimum loss function to obtain a contrast learning model;
[0146] Analyze the similarity between the synchronized perception data and the positive samples using the contrastive learning model to obtain environmental similarity features;
[0147] Generate environmental information of the preset scenario according to the environmental similarity features.
[0148] In one embodiment, when the environmental perception module 101 executes analyzing the similarity between the synchronized perception data and the positive samples using the contrastive learning model to obtain environmental similarity features, it is used for:
[0149] Compare the synchronized perception data and the positive samples using the contrastive learning model to obtain environmental contrast features;
[0150] Calculate the similarity between each environmental contrast feature and the positive samples one by one;
[0151] Select the positive samples whose similarity is greater than the preset similarity threshold;
[0152] Extract the environmental similarity features similar to the synchronized perception data from the selected positive samples.
[0153] In one embodiment, when the initial decision-making module 102 executes making an initial decision on the environmental information according to the target task to obtain a state-action table, it is used for:
[0154] Obtain the initial state and initial actions of the preset embodied intelligent robot;
[0155] Randomly combine the initial state and the initial actions, and summarize each combination into an initial combination table;
[0156] Generate a target action path according to the target task and the environmental information;
[0157] Update the initial state according to the target action path to obtain an updated state;
[0158] Generate an initial target action of the target action path according to the updated state;
[0159] Obtain a preset learning rate, a preset discount factor, and the reward value of each initial target action;
[0160] Obtain the update probability of obtaining the updated state, and calculate the state-action value of each combination in the initial combination table according to the update probability;
[0161] Update the state-action value using the preset learning rate, the preset discount factor, and the reward value to obtain a state-action table.
[0162] In one embodiment, the initial confidence generation module 103 performs obtaining the number of selections of each action in the state-action table, and generating the initial confidence of each action according to the number of selections, for:
[0163] Randomly select a state in the state-action table as the target state;
[0164] Count the number of state selections of the target state;
[0165] Count the number of action selections of the actions corresponding to the target state in the state-action table;
[0166] Obtain the state-action values of the target state and the actions corresponding to the target state, and calculate the initial confidence of each action in the state-action table according to the state-action values, the number of state selections, and the number of action selections.
[0167] In one embodiment, the confidence threshold to be filtered generation module 104 performs generating a number of different random confidence thresholds to be filtered by using the initial confidence and a preset variation parameter, for:
[0168] Calculate the average confidence of the initial confidence;
[0169] Take the preset variation parameter as the standard deviation of the initial confidence;
[0170] Generate a number of different random confidence thresholds to be filtered by using the average confidence and the standard deviation.
[0171] In one embodiment, the confidence threshold to be filtered selection module 105 performs making a policy selection for each action in the state-action table according to the confidence threshold to be filtered, to obtain a number of policy scores, for:
[0172] Obtain the screening task of the preset embodied intelligent robot and the screening action path of the screening task;
[0173] Select one of the confidence thresholds to be filtered one by one as the target confidence threshold;
[0174] Select one of the actions in the state-action table one by one as the target screening action;
[0175] Obtain the initial confidence of each target screening action, and determine whether the initial confidence of the target screening action is greater than or equal to the target confidence threshold;
[0176] If the initial confidence of the target screening action is less than the target confidence threshold, randomly select an action in the state-action table for exploration;
[0177] If the initial confidence of the target screening action is greater than or equal to the target confidence threshold, calculate the state-action update value of the target screening action, and screen out the action corresponding to the maximum state-action update value to obtain the selected action;
[0178] Use the selected action to obtain the completion time of the screening action path;
[0179] Obtain the length of the screening action path, and calculate the policy score of each target confidence threshold in the screening action path according to the selected action, the completion time, and the length.
[0180] In one embodiment, the final decision-making module 106 executes to use the policy score to screen out the optimized confidence threshold for:
[0181] Generate a scatter plot of the policy score and the confidence threshold to be screened corresponding to the policy score;
[0182] Use the confidence threshold to be screened corresponding to the point with the highest vertical axis in the scatter plot as the optimized confidence threshold.
[0183] In one embodiment, the final decision-making module 106 executes to perform perception decision on the environmental information using the optimized confidence threshold to obtain the final decision result for:
[0184] Generate the decision action path of the preset embodied intelligent robot according to the environmental information and the target task;
[0185] Obtain the action to be screened according to the decision action path, and calculate the action confidence of the action to be screened;
[0186] Screen out the actions to be screened whose action confidence is greater than or equal to the optimized confidence threshold as the decision actions;
[0187] Generate the final decision result of the preset embodied intelligent robot according to the decision action, the decision action path, and the environmental information.
[0188] In the present invention, for an embodied intelligent robot perception and decision-making method for complex scenarios, by obtaining the initial perception data of a preset embodied intelligent robot for a preset scenario, performing self-supervised learning on the initial perception data to obtain the environmental information of the preset scenario, obtaining the target task of the preset embodied intelligent robot, making an initial decision on the environmental information according to the target task to obtain a state-action table, obtaining the selection times of each action in the state-action table, and generating an initial confidence level for each action according to the selection times, generating a number of different random confidence threshold candidates to be screened using the initial confidence level and a preset variation parameter, making a policy selection for each action in the state-action table according to the confidence threshold candidates to be screened to obtain a number of policy scores, screening out an optimized confidence threshold using the policy scores, and making a perception decision on the environmental information using the optimized confidence threshold to obtain a final decision result, effectively solving the problem that the embodied intelligent robot overly relies on known optimal actions and thus falls into a local optimal solution. For the specific limitations of an embodied intelligent robot perception and decision-making system for complex scenarios, reference can be made to the limitations of an embodied intelligent robot perception and decision-making method for complex scenarios in the above text, which will not be elaborated here. Each module in the above-mentioned embodied intelligent robot perception and decision-making system for complex scenarios can be implemented in whole or in part by software, hardware, or a combination thereof. The above-mentioned modules can be embedded in or independent of a processor in a computer device in the form of hardware, or stored in a memory in a computer device in the form of software, so as to facilitate the processor to call and execute the operations corresponding to the above-mentioned modules.
[0189] In the embodiments provided by the present invention, it should be understood that the disclosed system can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation.
[0190] In addition, in each embodiment of the present invention, the functional modules can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware, or in the form of a combination of hardware and software functional modules.
[0191] Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention. Any reference signs associated with the claims should not be regarded as limiting the claimed rights.
[0192] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and the present invention can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention.
[0193] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0194] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules as needed, that is, the internal structure of the system can be divided into different functional units or modules to complete all or part of the functions described above.
[0195] In the embodiments provided by the present disclosure, it should be understood that the disclosed systems and methods can also be implemented in other ways. The system embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of systems, methods, and computer program products according to multiple embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above-mentioned module, program segment, or part of code includes one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0196] It should be noted that in the present disclosure, the terms "include", "comprise", or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device that includes a series of elements includes not only those elements but also other elements not expressly listed, or elements that are inherent to such process, method, article, or device. Without further limitation, an element limited by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article, or device that includes the element.
[0197] The above-described embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments or perform equivalent replacements for some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention and should all be included in the protection scope of the present invention.
Claims
1. An embodied intelligent robot perception and decision-making method for complex scenarios, characterized in that The method includes: Obtain the initial perception data of a preset embodied intelligent robot for a preset scenario, perform self-supervised learning on the initial perception data to obtain the environmental information of the preset scenario; Obtain the target task of the preset embodied intelligent robot, and make an initial decision on the environmental information according to the target task to obtain a state-action table; Obtain the selection times of each action in the state-action table, and generate an initial confidence level for each action according to the selection times; Generate a number of different random candidate confidence thresholds by using the initial confidence levels and preset variation parameters; Obtain the screening task of the preset embodied intelligent robot and the screening action path of the screening task; Select one of the candidate confidence thresholds one by one as the target confidence threshold; Select one action in the state-action table one by one as the target screening action; Obtain the initial confidence level of each target screening action, and determine whether the initial confidence level of the target screening action is greater than or equal to the target confidence threshold; If the initial confidence level of the target screening action is less than the target confidence threshold, randomly select actions in the state-action table for exploration; If the initial confidence level of the target screening action is greater than or equal to the target confidence threshold, calculate the state-action update value of the target screening action, and screen out the action corresponding to the maximum state-action update value to obtain the selected action; Use the selected action to obtain the completion time of the screening action path; Obtain the length of the screening action path, and calculate the policy score of each target confidence threshold in the screening action path according to the selected action, the completion time, and the length; Use the policy scores to screen out the optimized confidence threshold, and use the optimized confidence threshold to perform perception decision on the environmental information to obtain the final decision result.
2. The method for the embodied intelligent robot perception and decision-making facing complex scenarios according to claim 1, wherein The performing self-supervised learning on the initial perception data to obtain the environmental information of the preset scenario includes: Perform data alignment and timestamp synchronization on the initial perception data to obtain synchronized perception data; Obtain positive samples, negative samples, and a preset comparison factor in the known perception data, and generate a contrast loss function according to the positive samples, the negative samples, and the preset comparison factor; Minimize the contrast loss function to obtain the minimum loss function; Optimize a preset self-supervised learning model by using the minimum loss function to obtain a contrast learning model; Analyze the similarity between the synchronized perception data and the positive samples by using the contrast learning model to obtain environmental similarity features; Generate the environmental information of the preset scenario according to the environmental similarity features.
3. The embodied intelligent robot perception and decision-making method for complex scenarios according to claim 2, wherein, The analyzing the similarity between the synchronized perception data and the positive samples by using the contrast learning model to obtain environmental similarity features includes: Compare the synchronized perception data and the positive samples by using the contrast learning model to obtain environmental contrast features; Calculate the similarity between each environmental contrast feature and the positive samples one by one; Screen out the positive samples whose similarity is greater than a preset similarity threshold; Extract the environmental similarity features similar to the synchronized perception data from the selected positive samples.
4. The embodied intelligent robot perception and decision-making method for complex scenarios according to claim 1, wherein The initial decision on the environmental information according to the target task to obtain a state-action table includes: Obtain the initial state and initial actions of the preset embodied intelligent robot; Randomly combine the initial state and the initial actions, and summarize each combination into an initial combination table; Generate a target action path according to the target task and the environmental information; Update the initial state according to the target action path to obtain an updated state; Generate an initial target action of the target action path according to the updated state; Obtain a preset learning rate, a preset discount factor, and the reward value of each initial target action; Obtain the update probability of obtaining the updated state, and calculate the state-action value of each combination in the initial combination table according to the update probability; Update the state-action value by using the preset learning rate, the preset discount factor, and the reward value to obtain a state-action table.
5. The embodied intelligent robot perception and decision-making method for complex scenarios according to claim 4, wherein The obtaining the selection times of each action in the state-action table and generating an initial confidence level for each action according to the selection times includes: Randomly select a state in the state-action table as the target state; Count the state selection times of the target state; Count the action selection times of the actions corresponding to the target state in the state-action table; Obtain the state-action value of the target state and the actions corresponding to the target state, and calculate the initial confidence level of each action in the state-action table according to the state-action value, the state selection times, and the action selection times.
6. The method for embodied intelligent robot perception and decision-making for complex scenarios according to claim 1, wherein, The generating a number of different random candidate confidence thresholds by using the initial confidence level and a preset variation parameter includes: Calculate the average confidence level of the initial confidence level; Take the preset variation parameter as the standard deviation of the initial confidence level; Generate a number of different random candidate confidence thresholds by using the average confidence level and the standard deviation.
7. The embodied intelligent robot perception and decision-making method for complex scenarios according to claim 1, wherein The screening out an optimized confidence threshold by using the policy score includes: Generate a scatter plot of the policy score and the candidate confidence threshold corresponding to the policy score; Take the candidate confidence threshold corresponding to the point with the highest vertical axis in the scatter plot as the optimized confidence threshold.
8. The embodied intelligent robot perception and decision-making method for complex scenarios according to claim 1, characterized in that The making a perception decision on the environmental information by using the optimized confidence threshold to obtain a final decision result includes: Generate a decision action path of the preset embodied intelligent robot according to the environmental information and the target task; Obtain candidate actions according to the decision action path, and calculate the action confidence level of the candidate actions; Screen out the candidate actions whose action confidence level is greater than or equal to the optimized confidence threshold as decision actions; Generate the final decision result of the preset embodied intelligent robot according to the decision actions, the decision action path, and the environmental information.
9. An embodied intelligent robot perception and decision-making system for complex scenarios, characterized in that, The system includes: An environmental perception module, configured to obtain initial perception data of a preset embodied intelligent robot for a preset scenario, and perform self-supervised learning on the initial perception data to obtain environmental information of the preset scenario; An initial decision-making module, configured to obtain the target task of the preset embodied intelligent robot, and make an initial decision on the environmental information according to the target task to obtain a state-action table; An initial confidence generation module, configured to obtain the selection times of each action in the state-action table, and generate the initial confidence of each action according to the selection times; A to-be-screened confidence threshold generation module, configured to generate a plurality of different random to-be-screened confidence thresholds by using the initial confidence and a preset change parameter; A to-be-screened confidence threshold selection module, configured to perform policy selection on each action in the state-action table according to the to-be-screened confidence threshold to obtain a plurality of policy scores, including: Obtaining the screening task of the preset embodied intelligent robot and the screening action path of the screening task; Selecting one of the to-be-screened confidence thresholds one by one as the target confidence threshold; Selecting one of the actions in the state-action table one by one as the target screening action; Obtaining the initial confidence of each target screening action, and determining whether the initial confidence of the target screening action is greater than or equal to the target confidence threshold; If the initial confidence of the target screening action is less than the target confidence threshold, randomly select an action in the state-action table for exploration; If the initial confidence of the target screening action is greater than or equal to the target confidence threshold, calculate the state-action update value of the target screening action, and screen out the action corresponding to the maximum state-action update value to obtain the selected action; Using the selected action to obtain the completion time of the screening action path; Obtaining the length of the screening action path, and calculating the policy score of each target confidence threshold in the screening action path according to the selected action, the completion time, and the length; A final decision-making module, configured to screen out the optimized confidence threshold by using the policy score, and perform perceptual decision-making on the environmental information by using the optimized confidence threshold to obtain the final decision result.
Citation Information
Patent Citations
Object recognition method in dynamic scene and applicable intelligent robot thereof
CN119380252A
Robot control method and device and robot
CN119526399A