Complex scene-oriented intelligent robot perception decision-making method and system

Through self-supervised learning and confidence threshold screening methods, the problem of embodied intelligent robots over-relying on known actions in complex scenarios is solved, and more efficient and diversified decision-making capabilities are achieved.

CN120103719AActive Publication Date: 2025-06-06SANY INTELLIGENT MFG (SHENZHEN) CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510600685.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-06-06
Estimated Expiration
2045-05-12

AI Technical Summary

Technical Problem

In the prior art, embodied intelligent robots are prone to over-rely relying on known optimal actions in complex scenarios, resulting in falling into local optimal solutions and difficult to effectively respond to dynamically changing environments and real-time decision-making needs.

Method used

By obtaining the preset embossed intelligent robot to self-supervise the initial perception data of the preset scene, generate environmental information; make initial decisions based on the target task, obtain a state action table; calculate the number of selections of each action and generate initial confidence; use confidence and change parameters to generate a random confidence threshold; make strategy selections for actions based on the threshold to obtain policy scores; filter out an optimization confidence threshold for perceived decisions.

Benefits of technology

Effectively avoid embodied intelligent robots from over-rely relying on known actions, improve the diversity of decisions and the ability to implement global optimal solutions, and improve adaptability and task execution efficiency in complex and dynamic environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120103719A_ABST
    Figure CN120103719A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent decision making, and discloses a complex scene-oriented intelligent robot perception decision making method and system, and the method comprises the steps: obtaining the initial perception data of a preset intelligent robot for a preset scene, carrying out the self-supervised learning of the initial perception data, and obtaining environment information, performing initial decision on the environment information according to an obtained target task of a preset intelligent robot with a body to obtain a state action table, obtaining the number of selection times of each action in the state action table, generating a plurality of confidence threshold values to be screened according to the initial confidence corresponding to the number of selection times and a preset change parameter, and performing strategy selection on each action in the state action table according to a confidence threshold to be screened to obtain a plurality of strategy scores, screening out an optimized confidence threshold, and performing perception decision on the environment information by using the optimized confidence threshold to obtain a final decision result. According to the method, the problem that the intelligent robot with the body falls into the decision-making local optimal solution is effectively optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent decision-making technology, and in particular to a perception decision-making method and system for an embodied intelligent robot for complex scenarios. Background Art

[0002] In the current existing technologies, the perception and decision-making capabilities of an embodied intelligent robot for complex scenarios mainly rely on traditional perception and decision-making algorithms. These methods are usually analyzed through static data models, which are difficult to effectively cope with the changing environment and real-time decision-making needs in complex scenarios. The shortcomings of the existing technologies include: insufficient response to dynamic changes, poor real-time performance, and limited processing capabilities. Especially in the face of rapidly changing environments, decision lags or insufficient accuracy are prone to occur. In addition, these technologies usually lack the ability to comprehensively process multimodal information, making it difficult to accurately capture and understand complex perception information, resulting in the robot's adaptability and decision-making accuracy being limited in complex scenarios.

[0003] Traditional Q-learning mainly selects actions by maximizing the current state-action value. This method easily causes the embodied intelligent robot to over-rely on the known optimal action, thus falling into the local optimal solution. This limitation stems from the fact that Q-learning pays too much attention to the current reward information and ignores the exploration of other possible actions. Therefore, the embodied intelligent robot may miss those actions with lower initial returns but greater long-term potential, limiting the diversity of strategies and the realization of the global optimal solution.

[0004] The existing technology has the problem of over-reliance on known optimal actions for embodied intelligent robots, thus falling into the problem of local optimal solutions. Summary of the invention

[0005] The present invention provides a perception decision-making method for an embodied intelligent robot for complex scenarios, the main purpose of which is to solve the problem of the embodied intelligent robot over-relying on known optimal actions and thus falling into a local optimal solution.

[0006] In a first aspect, to achieve the above-mentioned purpose, the present invention provides a perception decision-making method of an embodied intelligent robot for complex scenarios, comprising: Acquire initial perception data of a preset embodied intelligent robot for a preset scene, perform self-supervised learning on the initial perception data, and obtain environmental information of the preset scene; Obtaining the target task of the preset embodied intelligent robot, making an initial decision on the environmental information according to the target task, and obtaining a state-action table; Obtaining the number of selections of each action in the state-action table, and generating an initial confidence of each action according to the number of selections; Using the initial confidence and preset variation parameters to generate a number of different random confidence thresholds to be screened; Selecting a strategy for each action in the state-action table according to the confidence threshold to be screened, and obtaining a number of strategy scores; The strategy score is used to filter out an optimized confidence threshold, and the optimized confidence threshold is used to make a perception decision on the environmental information to obtain a final decision result.

[0007] In a second aspect, the present invention further provides an embodied intelligent robot perception and decision-making system for complex scenarios, the system comprising: An environmental perception module is used to obtain initial perception data of a preset embodied intelligent robot for a preset scene, perform self-supervised learning on the initial perception data, and obtain environmental information of the preset scene; An initial decision module is used to obtain the target task of the preset embodied intelligent robot, make an initial decision on the environmental information according to the target task, and obtain a state-action table; An initial confidence generation module, used to obtain the number of selections of each action in the state-action table, and generate an initial confidence of each action according to the number of selections; A confidence threshold generation module for screening, used to generate a number of different random confidence thresholds for screening using the initial confidence and preset variation parameters; A to-be-screened confidence threshold selection module, used to select a strategy for each action in the state-action table according to the to-be-screened confidence threshold, and obtain a number of strategy scores; The final decision module is used to use the strategy score to filter out the optimized confidence threshold, and use the optimized confidence threshold to make a perception decision on the environmental information to obtain a final decision result.

[0008] The present invention obtains initial perception data of a preset scene by a preset embodied intelligent robot, performs self-supervised learning on the initial perception data, obtains environmental information of the preset scene, obtains a target task of the preset embodied intelligent robot, makes an initial decision on the environmental information according to the target task, obtains a state-action table, obtains the number of selections of each action in the state-action table, generates an initial confidence for each action according to the number of selections, generates a number of different random confidence thresholds to be screened using the initial confidence and preset change parameters, performs strategy selection on each action in the state-action table according to the confidence thresholds to be screened, obtains a number of strategy scores, uses the strategy scores to screen out an optimized confidence threshold, and uses the optimized confidence threshold to perform perception decision on the environmental information to obtain a final decision result, thereby effectively solving the problem of the embodied intelligent robot over-relying on known optimal actions and thus falling into a local optimal solution. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative labor.

[0010] Figure 1 A schematic diagram of a flow chart of a perception and decision-making method of an embodied intelligent robot for complex scenarios provided by an embodiment of the present invention; Figure 2 A schematic diagram of a module of an embodied intelligent robot perception and decision-making system for complex scenarios provided by an embodiment of the present invention; The realization of the purpose, functional features and advantages of the present invention will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0011] In order to enable those skilled in the art to better understand the technical solution of the present disclosure, and to fully understand and implement how the present disclosure applies technical means to solve technical problems and achieve the corresponding technical effects, the technical solution in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only embodiments of a part of the present disclosure, not all of the embodiments. The embodiments of the present disclosure and the various features in the embodiments can be combined with each other without conflict, and the technical solutions formed are all within the scope of protection of the present disclosure. Based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in this field without making creative work should fall within the scope of protection of the present disclosure.

[0012] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product, or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products, or devices.

[0013] The embodiment of the present application provides a perception and decision-making method for an embodied intelligent robot for complex scenarios, which can be executed by software or hardware installed on a terminal device or a server device. The server includes but is not limited to: a single server, a server cluster, a cloud server or a cloud server cluster, etc. The server can be an independent server, or it can be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.

[0014] Reference Figure 1 FIG. 1 is a flow chart of a perception and decision-making method of an embodied intelligent robot for complex scenes provided by an embodiment of the present invention. In this embodiment, the perception and decision-making method of an embodied intelligent robot for complex scenes includes: S1. Obtain initial perception data of a preset embodied intelligent robot for a preset scene, perform self-supervised learning on the initial perception data, and obtain environmental information of the preset scene.

[0015] In the embodiment of the present invention, the embodied intelligent robot has perception capabilities, wherein the perception devices possessed by the embodied intelligent robot, such as visual sensors, auditory sensors, tactile sensors, position sensors, etc., collect initial perception data in a complex environment through these perception devices, and assist the embodied intelligent robot in perceiving objects, obstacles, people, etc. in the space. By learning the similarities and differences of the initial perception data, the embodied intelligent robot can automatically obtain effective features from the unlabeled initial perception data and obtain environmental information.

[0016] In detail, the self-supervised learning of the initial perception data to obtain the environmental information of the preset scene includes: Performing data alignment and timestamp synchronization on the initial perception data to obtain synchronized perception data; Acquire positive samples, negative samples and preset contrast factors in known perception data, and generate a contrast loss function according to the positive samples, the negative samples and the preset contrast factors; Minimizing the contrast loss function to obtain a minimum loss function; Optimizing the preset self-supervised learning model using the minimum loss function to obtain a comparative learning model; Analyzing the similarity between the synchronous perception data and the positive sample using the contrastive learning model to obtain environmental similarity features; The environmental information of the preset scene is generated according to the environmental similarity features.

[0017] In detail, the use of the contrastive learning model to analyze the similarity between the synchronous perception data and the positive sample to obtain environmental similarity features includes: Using the contrastive learning model to compare the synchronous perception data with the positive sample to obtain an environmental contrast feature; Calculating the similarity between the environmental contrast features and the positive samples one by one; Filter out positive samples whose similarity is greater than a preset similarity threshold; Extracting environmental similarity features similar to the synchronous perception data from the screened positive samples.

[0018] In detail, initial perception data in a complex environment is collected through perception devices, and data alignment and timestamp synchronization are performed on the initial perception data to obtain synchronized perception data. The specific operation steps are: obtain the timestamp of each initial perception data, summarize the data with the same timestamp, and sort each initial perception data in the order of timestamps to obtain synchronized perception data.

[0019] Get positive samples, negative samples and preset contrast factors from known perception data, and generate a contrast loss function based on the positive samples, negative samples and preset contrast factors, where positive samples represent data from the same data source, different perspectives or different timestamps of the same instance. For example, images from the same object captured at different times can be considered positive samples. Negative samples represent perception data from different objects, different classes, or different environments, and are samples that are significantly different from the current initial perception data. The calculation formula is as follows: in, Indicates Positive samples, Indicates Positive samples, Indicates negative samples, Indicates the preset contrast factor, Represents the total number of negative samples.

[0020] The contrast loss function is minimized to obtain a minimum loss function, and the preset self-supervised learning model is optimized using the minimum loss function to obtain a contrast learning model. The calculation formula is as follows: in, Indicates Positive samples, Indicates Positive samples, Indicates negative samples, Indicates the preset contrast factor, represents the total number of negative samples, Represents a preset self-supervised learning model.

[0021] The contrast learning model is used to analyze the similarity between the synchronous perception data and the positive sample to obtain the environmental similarity feature, and the environmental information of the preset scene is generated according to the environmental similarity feature. The specific operation is: input the synchronous perception data into the contrast learning model to obtain the environmental contrast feature, calculate the similarity between the environmental contrast feature and each feature of the positive sample, extract the features whose similarity is greater than the preset similarity threshold, and obtain the environmental similarity feature of the synchronous perception data. The calculation formula of the similarity is as follows: in, represents a positive sample, Representing contrastive features of the environment for synchronous perceptual data.

[0022] The environmental information of the preset scene is generated according to the environmental similarity features, including: environmental status: such as maps, obstacles, paths, terrain, etc. State changes: such as whether the path is blocked, the appearance and movement of obstacles, etc.

[0023] By effectively processing the initial perception data through self-supervised learning, the embodied intelligent robot can learn similar features of the environment from the synchronized perception data without label data. By optimizing the comparative learning model, the robot can more accurately identify and analyze important features in the environment, thereby improving the accuracy and robustness of perception decisions. It can not only enable the embodied intelligent robot to self-adjust the perception model in a dynamic environment, but also enhance its adaptability in different scenarios, enabling it to perform tasks and make decisions more efficiently in complex environments.

[0024] S2. Obtain the target task of the preset embodied intelligent robot, make an initial decision on the environmental information according to the target task, and obtain a state-action table.

[0025] In an embodiment of the present invention, the target task of the embodied intelligent robot is obtained, and the environmental information is initially decided based on the target task based on the Q-learning algorithm to obtain a state-action table. The state-action table includes: the state of the embodied intelligent robot, such as the position and direction of the embodied intelligent robot; the action of the embodied intelligent robot, such as the choice of the embodied intelligent robot to move forward, backward, turn, etc.

[0026] In detail, the initial decision is made on the environmental information according to the target task to obtain a state action table, including: Obtaining the initial state and initial action of the preset embodied intelligent robot; Randomly combining the initial state and the initial action, and summarizing each combination into an initial combination table; Generate a target action path according to the target task and the environmental information; Update the initial state according to the target action path to obtain an updated state; generating an initial target action of the target action path according to the updated state; Obtaining a preset learning rate, a preset discount factor, and a reward value for each of the initial target actions; Obtaining an update probability of the update state, and calculating a state action value of each combination in the initial combination table according to the update probability; The state-action value is updated using the preset learning rate, the preset discount factor and the reward value to obtain a state-action table.

[0027] In detail, the target task of the preset embodied intelligent robot is obtained, wherein the target task is to allow the embodied intelligent robot to interact with a preset environment, such as completing a certain path planning, reaching a certain specified location, or completing a certain specific operation.

[0028] A target action path is generated according to the target task and the environmental information, and the action corresponding to the target action path in the initial combination table is used as the initial target action. The specific operation steps are: determine the starting point and target point of the embodied robot, analyze each action and state on each path, consider the feasibility and efficiency of the path, select the path with the highest feasibility as the target action path, and use the action on the target action path as the initial target action.

[0029] The initial combination table is a two-dimensional matrix, in which the elements of the two-dimensional matrix are composed of each state and action of the embodied intelligent robot. The update probability of the updated state is obtained, and the state action value of each combination in the initial combination table is calculated according to the update probability. The calculation formula is as follows: in, Indicates the initial state Take initial action When the update state The probability of represents the initial state, Indicates the initial action, Indicates the update status. represents the initial target action, represents the preset discount factor, Represents the reward value, Indicates updating status The maximum state-action value of all initial target actions under .

[0030] The state-action value is updated using the preset learning rate, the preset discount factor and the reward value, and the final state-action table is obtained by continuously interacting with the environmental information and updating the state-action value. The calculation formula is as follows: in, Indicates the initial state Take initial action The state action value, represents the initial state, Indicates the initial action, Indicates the update status. represents the initial target action, represents the preset discount factor, Represents the reward value, Indicates updating status The maximum state-action value of all initial target actions under Represents the preset learning rate.

[0031] By randomly combining initial states and actions, and gradually selecting the most appropriate actions based on the target task and environmental information, the embodied intelligent robot can make autonomous decisions in a dynamic environment. By updating the initial combination table using learning rate, discount factor, and reward value, the robot can continuously optimize its action strategy, gradually improve the efficiency and accuracy of task completion, and enhance the adaptive ability of the embodied intelligent robot, enabling it to effectively perceive and make decisions in a constantly changing environment, thereby better performing complex tasks.

[0032] S3. Obtain the number of selections of each action in the state-action table, and generate an initial confidence of each action according to the number of selections.

[0033] In an embodiment of the present invention, an initial confidence level is generated by the number of selections of each action in the state-action table, indicating the degree of trust in each action, thereby providing data support for better decision-making of subsequent embodied intelligent robots.

[0034] In detail, obtaining the number of selections of each action in the state-action table and generating an initial confidence of each action according to the number of selections includes: Randomly select a state in the state-action table as the target state; Counting the number of state selections of the target state; Counting the number of action selections for the action corresponding to the target state in the state-action table; The target state and the state action value of the action corresponding to the target state are obtained, and the initial confidence of each action in the state action table is calculated according to the state action value, the state selection number and the action selection number.

[0035] In detail, the confidence of each action is calculated through the UCB algorithm to provide a data basis for the subsequent embodied intelligent robot in the exploration stage. When an action is selected in small quantities, the confidence will be adjusted higher to encourage more exploration; and when an action is selected multiple times, the confidence will gradually decrease to avoid local optimality.

[0036] The target state and the state action value of the action corresponding to the target state are obtained, and the initial confidence of each action in the state action table is calculated according to the state action value, the state selection number and the action selection number. The calculation formula is as follows: in, In the target state Take the action corresponding to the target state The state action value, represents the target state, Indicates the target state corresponding to the action, Indicates the number of state selections, Indicates the number of action selections, In the target state Take the action corresponding to the target state The initial confidence level.

[0037] By calculating the number of times each action is selected and its correlation with the target state, the embodied intelligent robot can evaluate and update the confidence of the action. By combining the state action value, the number of state selections, and the number of action selections, the robot can optimize the decision-making process based on historical experience and action feedback, so that it can more accurately select the best action when facing complex tasks, improve the robot's learning efficiency and adaptability in a dynamic environment, and better make autonomous decisions in changing scenarios and improve the success rate of task execution.

[0038] S4. Generate a number of different random confidence thresholds to be screened using the initial confidence and preset variation parameters.

[0039] In the embodiment of the present invention, a number of different random confidence thresholds to be screened are generated by using the initial confidence and preset variation parameters through normal distribution, so as to provide a data basis for subsequent screening of the optimal confidence threshold for decision making.

[0040] In detail, the initial confidence and the preset variation parameter are used to generate a number of different random confidence thresholds to be screened, including: Calculating an average confidence of the initial confidences; Using the preset variation parameter as the standard deviation of the initial confidence; A plurality of different random confidence thresholds to be screened are generated using the average confidence and the standard deviation.

[0041] In detail, the average confidence of the initial confidence is calculated, the preset change parameter is used as the standard deviation of the initial confidence, and a number of different random confidence thresholds to be screened are generated by using the average confidence and the standard deviation through normal distribution. The calculation formula is as follows: in, In the target state Take the action corresponding to the target state The initial confidence level of It means the mean , the standard deviation is The confidence range is described by using the average confidence and standard deviation of the normal distribution. Multiple values ​​are randomly selected from this confidence range as the confidence threshold to be screened. The confidence threshold to be screened is used for further decision-making to increase the diversity of decision-making.

[0042] By using the average and standard deviation of the initial confidence level, combined with preset variation parameters, multiple random confidence thresholds to be screened are generated, so that the embodied intelligent robot can introduce a certain degree of randomness and diversity in the decision-making process, which helps the robot explore and optimize the decision-making strategy under different environmental conditions and avoid falling into the local optimal solution. By adjusting the confidence threshold, the robot can respond to environmental changes more flexibly, improve the robustness and adaptability of decision-making, and thus make more appropriate action choices in complex or uncertain task scenarios.

[0043] S5. Perform strategy selection for each action in the state-action table according to the confidence threshold to be screened, and obtain a number of strategy scores.

[0044] In the embodiment of the present invention, a strategy is selected for each action in the state action table according to the confidence threshold to be screened, and a number of strategy scores are obtained, which provide a data analysis basis for subsequent screening of optimized confidences.

[0045] In detail, the strategy selection is performed on each action in the state action table according to the confidence threshold to be screened, and several strategy scores are obtained, including: Obtaining a screening task of the preset embodied intelligent robot and a screening action path of the screening task; Selecting one of the confidence thresholds to be screened one by one as the target confidence threshold; Select one action from the state action table one by one as the target screening action; Obtaining an initial confidence of each target screening action, and determining whether the initial confidence of the target screening action is greater than or equal to the target confidence threshold; If the initial confidence of the target screening action is less than the target confidence threshold, randomly selecting an action from the state action table for exploration; If the initial confidence of the target screening action is greater than or equal to the target confidence threshold, then calculating the state action update value of the target screening action, and screening out the action corresponding to the maximum state action update value to obtain the selected action; Using the selection action to obtain the completion time of the screening action path; The length of the screening action path is obtained, and the strategy score of each target confidence threshold in the screening action path is calculated according to the selection action, the completion time and the length.

[0046] In detail, the state action update value of each action in the state action table is calculated using the formula in S2 to obtain the selected action. The calculation formula is as follows: in, Indicates that the filter is in progress Take target screening actions When the state corresponding to the selected action is transferred The probability of Indicates the filtering status. Indicates the target filtering action, Indicates the state corresponding to the selected action. Indicates the selection action. represents the preset discount factor, represents the selection reward value of the selected action, Indicates the state corresponding to the selected action The state action updates the value of all selected actions.

[0047] The length of the screening action path is obtained, and the strategy score of each target confidence threshold in the screening action path is calculated according to the selection action, the completion time and the length. The calculation formula is as follows: in, Indicates the completion time. Indicates length, represents the discount factor for completion time, Indicates the state action update value corresponding to the selection action.

[0048] By integrating the target task, confidence threshold, action selection in the state-action table, and strategy evaluation, the embodied intelligent robot can perform complex tasks more accurately and efficiently. The embodied intelligent robot selects different actions for different confidence thresholds, calculates the completion time and length of the target action path, and further helps evaluate the actual effect of strategies with different confidence thresholds. The optimized confidence threshold is selected based on the strategy score to achieve a better task completion path and a higher success rate.

[0049] S6. Use the strategy score to filter out an optimized confidence threshold, and use the optimized confidence threshold to make a perception decision on the environmental information to obtain a final decision result.

[0050] In an embodiment of the present invention, an optimized confidence threshold is screened out from a number of confidence thresholds to be screened through a scatter plot, and the optimized confidence threshold is used to help the embodied intelligent robot obtain decision-making actions, better perceive and make decisions on environmental information, and obtain the final decision result.

[0051] In detail, the step of using the strategy score to select an optimized confidence threshold includes: Generating a scatter plot of the strategy scores and the confidence thresholds to be screened corresponding to the strategy scores; The confidence threshold to be screened corresponding to the point with the highest vertical axis in the scatter plot is used as the optimized confidence threshold.

[0052] In detail, the using the optimized confidence threshold to make a perception decision on the environmental information to obtain a final decision result includes: Generate a decision-making action path of the preset embodied intelligent robot according to the environmental information and the target task; Acquire the action to be screened according to the decision action path, and calculate the action confidence of the action to be screened; Screening out the to-be-screened actions whose action confidence is greater than or equal to the optimized confidence threshold as decision actions; A final decision result of the preset embodied intelligent robot is generated according to the decision action, the decision action path and the environmental information.

[0053] In detail, the horizontal axis of the scatter plot represents the number of each confidence threshold to be screened, and the vertical axis represents the strategy score corresponding to each confidence threshold to be screened, so that the strategy scores corresponding to each confidence threshold to be screened can be directly compared, and the confidence threshold to be screened corresponding to the highest point on the vertical axis can be selected as the optimized confidence threshold.

[0054] The action confidence of the action to be screened is obtained according to the initial confidence calculation formula in S3, and the action confidence is compared with the screened optimization confidence threshold. The actions to be screened whose action confidence is greater than or equal to the optimization confidence threshold are screened out as decision actions, and the initial target actions whose initial confidence is less than the optimization confidence threshold are explored.

[0055] The final decision result of the preset embodied intelligent robot is generated according to the decision action, the decision action path and the environmental information. In a specific embodiment, environmental information: the embodied intelligent robot is in a warehouse with a flat ground but multiple obstacles, such as stacked goods, shelves, etc. Perception data: the sensor detects the location where the goods are stacked, which is about 2 meters away from the embodied intelligent robot. There is a shelf in front of the embodied intelligent robot about 3 meters away from the embodied intelligent robot, which cannot pass directly. There is enough space on the left side to bypass it. Target task: transport an item from the current position to the other side of the warehouse. The target position is at the far end of the warehouse. The embodied intelligent robot needs to bypass the obstacle and transport the item smoothly. Decision action path: the embodied intelligent robot starts from the current position and moves forward. When encountering an obstacle, it chooses to retreat and turn left to bypass the shelf. Decision action: forward, backward, turn, etc. Final decision result: forward: the robot will first try to move forward 2 meters to approach the obstacle. Backward: after encountering an obstacle, the robot determines that it cannot pass through and immediately retreats 1 meter to prepare for the subsequent detour. Turn: The robot recognizes that there is enough space on the left to bypass, so it decides to turn left 90 degrees and continue moving forward. Continue moving forward: After bypassing the obstacle, the robot continues to move forward along the planned path until it reaches the target location.

[0056] By analyzing the relationship between the strategy score and the confidence threshold, the embodied intelligent robot can find the optimal confidence threshold, so that the perception and response to environmental information when performing tasks is more sensitive and effective. Actions greater than or equal to the optimized confidence threshold ensure that the robot has a high degree of confidence when making decisions, and actions less than the optimized confidence threshold continue to explore, thereby avoiding falling into the local optimal solution. At the same time, by integrating decision actions, target action paths and environmental information, the robot can generate more efficient and accurate final decision results, improve the success rate and efficiency of task completion, and enable the embodied intelligent robot to make optimized decisions in complex environments, improving its intelligence level and execution capabilities.

[0057] It should be understood that the order of execution of the steps in the above embodiment does not necessarily mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present invention.

[0058] like Figure 2The figure shows a functional module diagram of an embodied intelligent robot perception and decision-making system for complex scenarios provided by one embodiment of the present invention.

[0059] In the embodiment of the present disclosure, a perception and decision-making system of an embodied intelligent robot for complex scenes is provided, and the perception and decision-making system of an embodied intelligent robot for complex scenes corresponds one to one with the perception and decision-making method of an embodied intelligent robot for complex scenes in the above embodiment. Figure 2 As shown, the embodied intelligent robot perception and decision-making system 100 for complex scenes includes an environment perception module 101, an initial decision module 102, an initial confidence generation module 103, a confidence threshold generation module to be screened 104, a confidence threshold selection module to be screened 105, and a final decision module 106. The functional modules are described in detail as follows: The environment perception module 101 is used to obtain initial perception data of a preset embodied intelligent robot for a preset scene, and perform self-supervised learning on the initial perception data to obtain environment information of the preset scene; An initial decision module 102 is used to obtain a target task of the preset embodied intelligent robot, make an initial decision on the environmental information according to the target task, and obtain a state-action table; An initial confidence generating module 103, used to obtain the number of selections of each action in the state action table, and generate an initial confidence of each action according to the number of selections; The confidence threshold generation module 104 is used to generate a plurality of different random confidence thresholds to be screened using the initial confidence and a preset variation parameter; A to-be-screened confidence threshold selection module 105, configured to select a strategy for each action in the state-action table according to the to-be-screened confidence threshold, and obtain a number of strategy scores; The final decision module 106 is used to use the strategy score to filter out an optimized confidence threshold, and use the optimized confidence threshold to make a perception decision on the environmental information to obtain a final decision result.

[0060] In one embodiment, the environment perception module 101 performs self-supervised learning on the initial perception data to obtain the environment information of the preset scene, which is used to: Performing data alignment and timestamp synchronization on the initial perception data to obtain synchronized perception data; Acquire positive samples, negative samples and preset contrast factors in known perception data, and generate a contrast loss function according to the positive samples, the negative samples and the preset contrast factors; Minimizing the contrast loss function to obtain a minimum loss function; Optimizing the preset self-supervised learning model using the minimum loss function to obtain a comparative learning model; Analyzing the similarity between the synchronous perception data and the positive sample using the contrastive learning model to obtain environmental similarity features; The environmental information of the preset scene is generated according to the environmental similarity features.

[0061] In one embodiment, the environment perception module 101 analyzes the similarity between the synchronous perception data and the positive sample using the contrastive learning model to obtain environment similarity features for: Using the contrastive learning model to compare the synchronous perception data with the positive sample to obtain an environmental contrast feature; Calculating the similarity between the environmental contrast features and the positive samples one by one; Filter out positive samples whose similarity is greater than a preset similarity threshold; Extracting environmental similarity features similar to the synchronous perception data from the screened positive samples.

[0062] In one embodiment, the initial decision module 102 performs an initial decision on the environmental information according to the target task to obtain a state action table for: Obtaining the initial state and initial action of the preset embodied intelligent robot; Randomly combining the initial state and the initial action, and summarizing each combination into an initial combination table; Generate a target action path according to the target task and the environmental information; Update the initial state according to the target action path to obtain an updated state; generating an initial target action of the target action path according to the updated state; Obtaining a preset learning rate, a preset discount factor, and a reward value for each of the initial target actions; Obtaining an update probability of the update state, and calculating a state action value of each combination in the initial combination table according to the update probability; The state-action value is updated using the preset learning rate, the preset discount factor and the reward value to obtain a state-action table.

[0063] In one embodiment, the initial confidence generating module 103 obtains the number of selections of each action in the state action table and generates the initial confidence of each action according to the number of selections, for: Randomly select a state in the state-action table as the target state; Counting the number of state selections of the target state; Counting the number of action selections for the action corresponding to the target state in the state-action table; The target state and the state action value of the action corresponding to the target state are obtained, and the initial confidence of each action in the state action table is calculated according to the state action value, the state selection number and the action selection number.

[0064] In one embodiment, the to-be-screened confidence threshold generation module 104 generates a plurality of different random to-be-screened confidence thresholds using the initial confidence and the preset variation parameter, for: Calculating an average confidence of the initial confidences; Using the preset variation parameter as the standard deviation of the initial confidence; A plurality of different random confidence thresholds to be screened are generated using the average confidence and the standard deviation.

[0065] In one embodiment, the to-be-screened confidence threshold selection module 105 performs strategy selection for each action in the state-action table according to the to-be-screened confidence threshold, and obtains a number of strategy scores for: Obtaining a screening task of the preset embodied intelligent robot and a screening action path of the screening task; Selecting one of the confidence thresholds to be screened one by one as the target confidence threshold; Select one action from the state action table one by one as the target screening action; Obtaining an initial confidence of each target screening action, and determining whether the initial confidence of the target screening action is greater than or equal to the target confidence threshold; If the initial confidence of the target screening action is less than the target confidence threshold, randomly selecting an action from the state action table for exploration; If the initial confidence of the target screening action is greater than or equal to the target confidence threshold, then calculating the state action update value of the target screening action, and screening out the action corresponding to the maximum state action update value to obtain the selected action; Using the selection action to obtain the completion time of the screening action path; The length of the screening action path is obtained, and the strategy score of each target confidence threshold in the screening action path is calculated according to the selection action, the completion time and the length.

[0066] In one embodiment, the final decision module 106 uses the strategy score to filter out an optimized confidence threshold for: Generating a scatter plot of the strategy scores and the confidence thresholds to be screened corresponding to the strategy scores; The confidence threshold to be screened corresponding to the point with the highest vertical axis in the scatter plot is used as the optimized confidence threshold.

[0067] In one embodiment, the final decision module 106 performs the perception decision on the environmental information using the optimized confidence threshold to obtain a final decision result, which is used to: Generate a decision-making action path of the preset embodied intelligent robot according to the environmental information and the target task; Acquire the action to be screened according to the decision action path, and calculate the action confidence of the action to be screened; Screening out the to-be-screened actions whose action confidence is greater than or equal to the optimized confidence threshold as decision actions; A final decision result of the preset embodied intelligent robot is generated according to the decision action, the decision action path and the environmental information.

[0068] In the present invention, for a perception and decision-making method of an embodied intelligent robot for complex scenes, by obtaining the initial perception data of the preset embodied intelligent robot for the preset scene, self-supervised learning is performed on the initial perception data to obtain the environmental information of the preset scene, the target task of the preset embodied intelligent robot is obtained, and the environmental information is initially decided according to the target task to obtain a state action table, the number of selections of each action in the state action table is obtained, and the initial confidence of each action is generated according to the number of selections, and the initial confidence and preset change parameters are used to generate a number of different random confidence thresholds to be screened, and a strategy is selected for each action in the state action table according to the confidence threshold to be screened to obtain a number of strategy scores, and the strategy score is used to screen out the optimized confidence threshold, and the optimized confidence threshold is used to perform perception and decision on the environmental information to obtain the final decision result, which effectively solves the problem of the embodied intelligent robot over-relying on the known optimal action, thereby falling into the local optimal solution. The specific definition of a perception and decision-making system of an embodied intelligent robot for complex scenes can be found in the above definition of a perception and decision-making method of an embodied intelligent robot for complex scenes, which will not be repeated here. Each module in the above-mentioned complex scene-oriented embodied intelligent robot perception and decision-making system can be implemented in whole or in part by software, hardware, and a combination thereof. Each of the above-mentioned modules can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in a computer device in the form of software, so that the processor can call and execute the operations corresponding to each of the above modules.

[0069] In the embodiments provided by the present invention, it should be understood that the disclosed system can be implemented in other ways. For example, the system embodiments described above are only illustrative, for example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation.

[0070] In addition, each functional module in each embodiment of the present invention may be integrated into one processing unit, each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of hardware plus software functional modules.

[0071] Therefore, no matter from which point of view, the embodiments should be regarded as illustrative and non-restrictive, and the scope of the present invention is limited by the appended claims rather than the above description, so it is intended that all changes falling within the meaning and scope of the equivalent elements of the claims are included in the present invention. Any attached figure mark in the claims should not be regarded as limiting the claims involved.

[0072] It is obvious to those skilled in the art that the present invention is not limited to the details of the above exemplary embodiments, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.

[0073] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0074] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the system can be divided into different functional units or modules to complete all or part of the functions described above.

[0075] In the embodiments provided in the present disclosure, it should be understood that the disclosed systems and methods can also be implemented in other ways. The system embodiments described above are merely schematic. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to the multiple embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of the code, and the above-mentioned module, program segment or a part of the code contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart can be implemented with a dedicated hardware-based system that performs a specified function or action, or can be implemented with a combination of dedicated hardware and computer instructions.

[0076] It should be noted that in the present disclosure, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element limited by the sentence "includes a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.

[0077] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.

Claims

1. A perception and decision-making method for embodied intelligent robots in complex scenarios, characterized in that: The method comprises: Acquire initial perception data of a preset embodied intelligent robot for a preset scene, perform self-supervised learning on the initial perception data, and obtain environmental information of the preset scene; Obtaining the target task of the preset embodied intelligent robot, making an initial decision on the environmental information according to the target task, and obtaining a state-action table; Obtaining the number of selections of each action in the state-action table, and generating an initial confidence of each action according to the number of selections; Using the initial confidence and preset variation parameters to generate a number of different random confidence thresholds to be screened; Selecting a strategy for each action in the state-action table according to the confidence threshold to be screened, and obtaining a number of strategy scores; The strategy score is used to filter out an optimized confidence threshold, and the optimized confidence threshold is used to make a perception decision on the environmental information to obtain a final decision result.

2. The perception and decision-making method of an embodied intelligent robot for complex scenes as claimed in claim 1, characterized in that: The performing self-supervised learning on the initial perception data to obtain the environmental information of the preset scene includes: Performing data alignment and timestamp synchronization on the initial perception data to obtain synchronized perception data; Acquire positive samples, negative samples and preset contrast factors in known perception data, and generate a contrast loss function according to the positive samples, the negative samples and the preset contrast factors; Minimizing the contrast loss function to obtain a minimum loss function; Optimizing the preset self-supervised learning model using the minimum loss function to obtain a comparative learning model; Analyzing the similarity between the synchronous perception data and the positive sample using the contrastive learning model to obtain environmental similarity features; The environmental information of the preset scene is generated according to the environmental similarity features.

3. The perception and decision-making method of an embodied intelligent robot for complex scenes as claimed in claim 2, characterized in that: The using the contrastive learning model to analyze the similarity between the synchronous perception data and the positive sample to obtain environmental similarity features includes: Using the contrastive learning model to compare the synchronous perception data with the positive sample to obtain an environmental contrast feature; Calculating the similarity between the environmental contrast features and the positive samples one by one; Filter out positive samples whose similarity is greater than a preset similarity threshold; Extracting environmental similarity features similar to the synchronous perception data from the screened positive samples.

4. The perception and decision-making method of an embodied intelligent robot for complex scenes as claimed in claim 1, characterized in that: The initial decision is made on the environmental information according to the target task to obtain a state action table, including: Obtaining the initial state and initial action of the preset embodied intelligent robot; Randomly combining the initial state and the initial action, and summarizing each combination into an initial combination table; Generate a target action path according to the target task and the environmental information; Update the initial state according to the target action path to obtain an updated state; generating an initial target action of the target action path according to the updated state; Obtaining a preset learning rate, a preset discount factor, and a reward value for each of the initial target actions; Obtaining an update probability of the update state, and calculating a state action value of each combination in the initial combination table according to the update probability; The state-action value is updated using the preset learning rate, the preset discount factor and the reward value to obtain a state-action table.

5. The perception and decision-making method of an embodied intelligent robot for complex scenarios as claimed in claim 4, characterized in that: The obtaining of the number of selections of each action in the state-action table and generating an initial confidence of each action according to the number of selections includes: Randomly select a state in the state-action table as the target state; Counting the number of state selections of the target state; Counting the number of action selections for the action corresponding to the target state in the state-action table; The target state and the state action value of the action corresponding to the target state are obtained, and the initial confidence of each action in the state action table is calculated according to the state action value, the state selection number and the action selection number.

6. The perception and decision-making method of an embodied intelligent robot for complex scenes as claimed in claim 1, characterized in that: The method of using the initial confidence and the preset variation parameters to generate a plurality of different random confidence thresholds to be screened includes: Calculating an average confidence of the initial confidences; Using the preset variation parameter as the standard deviation of the initial confidence; A plurality of different random confidence thresholds to be screened are generated using the average confidence and the standard deviation.

7. The embodied intelligent robot perception and decision-making method for complex scenarios as claimed in claim 1, characterized in that: The strategy selection is performed on each action in the state action table according to the confidence threshold to be screened, and a plurality of strategy scores are obtained, including: Obtaining a screening task of the preset embodied intelligent robot and a screening action path of the screening task; Selecting one of the confidence thresholds to be screened one by one as the target confidence threshold; Select one action from the state action table one by one as the target screening action; Obtaining an initial confidence of each target screening action, and determining whether the initial confidence of the target screening action is greater than or equal to the target confidence threshold; If the initial confidence of the target screening action is less than the target confidence threshold, randomly selecting an action from the state action table for exploration; If the initial confidence of the target screening action is greater than or equal to the target confidence threshold, then calculating the state action update value of the target screening action, and screening out the action corresponding to the maximum state action update value to obtain the selected action; Using the selection action to obtain the completion time of the screening action path; The length of the screening action path is obtained, and the strategy score of each target confidence threshold in the screening action path is calculated according to the selection action, the completion time and the length.

8. The perception and decision-making method of an embodied intelligent robot for complex scenes as claimed in claim 1, characterized in that: The step of using the strategy score to select an optimized confidence threshold comprises: Generating a scatter plot of the strategy scores and the confidence thresholds to be screened corresponding to the strategy scores; The confidence threshold to be screened corresponding to the point with the highest vertical axis in the scatter plot is used as the optimized confidence threshold.

9. The perception and decision-making method of an embodied intelligent robot for complex scenes as claimed in claim 1, characterized in that: The using the optimized confidence threshold to make a perception decision on the environmental information to obtain a final decision result includes: Generate a decision-making action path of the preset embodied intelligent robot according to the environmental information and the target task; Acquire the action to be screened according to the decision action path, and calculate the action confidence of the action to be screened; Screening out the to-be-screened actions whose action confidence is greater than or equal to the optimized confidence threshold as decision actions; A final decision result of the preset embodied intelligent robot is generated according to the decision action, the decision action path and the environmental information.

10. An embodied intelligent robot perception and decision-making system for complex scenarios, characterized in that: The system comprises: An environmental perception module is used to obtain initial perception data of a preset embodied intelligent robot for a preset scene, perform self-supervised learning on the initial perception data, and obtain environmental information of the preset scene; An initial decision module, used to obtain the target task of the preset embodied intelligent robot, make an initial decision on the environmental information according to the target task, and obtain a state-action table; An initial confidence generation module, used to obtain the number of selections of each action in the state-action table, and generate an initial confidence of each action according to the number of selections; A confidence threshold generation module for screening, used to generate a number of different random confidence thresholds for screening using the initial confidence and preset variation parameters; A to-be-screened confidence threshold selection module, used to select a strategy for each action in the state-action table according to the to-be-screened confidence threshold, and obtain a number of strategy scores; The final decision module is used to use the strategy score to filter out the optimized confidence threshold, and use the optimized confidence threshold to make a perception decision on the environmental information to obtain a final decision result.

Citation Information

Patent Citations

  • Object recognition method in dynamic scene and applicable intelligent robot thereof

    CN119380252A

  • Robot control method and device and robot

    CN119526399A

  • Interactive dialogue control method and system for intelligent robot with body

    CN119597880A

  • Method and system for data-driven and modular decision making and trajectory generation of an autonomous agent

    US11124204B1

  • Confidence interval-based deep reinforcement learning action decision-making method

    WO2024221817A1