Multi-modal perception humanoid robot behavior test method and related equipment
By processing multimodal perception data and joint state data, the dynamic adaptability and scene adaptability of humanoid robots in complex environments are realized, and detailed test reports are generated, solving the problem of poor dynamic adaptability and scene adaptability in existing technologies.
Patent Information
- Application Number
- CN202610036655.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-13
- Publication Date
- 2026-02-10
- Estimated Expiration
- 2046-01-13
AI Technical Summary
Existing humanoid robots have poor dynamic adaptability and scene adaptability in complex environments, making it difficult to effectively adjust their behavior according to actual conditions.
By acquiring multimodal perception data and current joint state data, the data is stitched together and features are extracted. A trained processing model is used for action sampling and discrimination calculations to determine the target behavior state. Scene testing and evaluation are then performed to generate a test report.
The dynamic adaptability and scene adaptability of the humanoid robot were improved, ensuring the safety and adaptability of its movements, and a detailed test report was generated.
Smart Images

Figure CN121492069A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of humanoid robot technology, and in particular to a behavior testing method and related equipment for a multimodal perception humanoid robot. Background Technology
[0002] With the rapid development of robotics technology, humanoid robots are increasingly being used in industrial and household settings to perform tasks in complex environments, such as industrial assembly and household cleaning. However, existing technologies pre-program humanoid robots according to the pre-set program, resulting in poor dynamic adaptability and poor scene adaptation. Summary of the Invention
[0003] The main objective of this application is to propose a multimodal perception method and related equipment for testing the behavior of humanoid robots, which can improve testing efficiency and enhance the safety of simulated humanoid robot behavior.
[0004] To achieve the above objectives, one aspect of this application proposes a behavior testing method for a multimodal perceptual humanoid robot, the method comprising: Acquire multimodal sensing data and current joint state data, and concatenate the multimodal sensing data and joint state data to determine the current state vector; Action sampling is performed based on the current state vector and the trained processing model to determine several sampled actions. Then, based on the several sampled actions, the current state vector, and the preset motion model, several candidate action states are determined. Based on several candidate action states and a preset neural network model, an identification calculation is performed to determine several total identification scores. Based on the several total identification scores and several candidate action states, an analysis is performed to determine the target behavior state. Based on the target behavior state and preset scenario test cases, a test evaluation is conducted, and a test report is determined.
[0005] In some embodiments, features are extracted from the current state vector according to the trained processing model to determine multimodal features, and the multimodal features are fused to determine fused feature data. Based on the preset shared weights of the robot's two arms and the fused feature data, a fully connected processing is performed to determine the weight feature data; The action space is determined by calculation based on the weight feature data and the preset activation function, and the action space and preset degree of freedom parameters are used to sample actions to determine a number of sampled actions.
[0006] In some embodiments, several sampling actions are analyzed to determine several motor parameters; wherein, the motor parameters include the motor current and reduction ratio corresponding to the sampling action; Based on several motor parameters and a first preset formula, several first torque vectors are determined, and several first torque vectors are adjusted according to preset torque thresholds to determine several joint torque vectors. The current state vector is parsed to determine the current joint angle vector and the current joint angular velocity vector. Based on the current joint angle vector, the current joint angular velocity vector, several joint torque vectors, and the preset motion model, several joint angle vectors for the next moment are calculated to determine. Constraints are applied based on several next-moment joint angle vectors and preset joint limit thresholds to determine several candidate action states.
[0007] In some embodiments, several candidate action states are analyzed to determine several torque data, several visual data, and several tactile data; Based on several torque data and several visual data, a fully connected processing is performed to determine the first output data, and activation calculation is performed on the first output data to determine the safety sub-score; The aptamer score is determined by calculating based on several visual data points and several tactile data points. The total discrimination score is determined by calculating the security sub-score, the adaptation sub-score, and the preset weight.
[0008] In some embodiments, several total discrimination scores are compared with preset scene thresholds; If the total identification score is greater than or equal to the preset scene threshold, the candidate action state corresponding to the total identification score is taken as the target behavior state; If the total discrimination score is less than the preset scene threshold, return to the process of sampling actions based on the current state vector and the trained processing model to determine several sampling actions until the total discrimination score is greater than or equal to the preset scene threshold.
[0009] In some embodiments, behavioral simulation is performed based on the preset scenario test cases and the target behavioral state to determine the simulation results; The simulation results and preset evaluation indicators are used to calculate and determine the quality of the action output, and the test report is determined based on the quality of the action output.
[0010] To achieve the above objectives, another aspect of this application proposes a behavior testing system for a multimodal perception humanoid robot, the system comprising: The data acquisition module is used to acquire multimodal sensing data and current joint state data, and to concatenate the multimodal sensing data and the joint state data to determine the current state vector. The action sampling module is used to sample actions based on the current state vector and the trained processing model, determine several sampled actions, and calculate based on the several sampled actions, the current state vector and the preset motion model to determine several candidate action states. The identification calculation module is used to perform identification calculations based on several candidate action states and a preset neural network model, determine several total identification scores, and analyze the several total identification scores and several candidate action states to determine the target behavior state. The test evaluation module is used to evaluate the test based on the target behavior state and preset scenario test cases, and to determine the test report.
[0011] To achieve the above objectives, another aspect of this application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the method described above.
[0012] To achieve the above objectives, another aspect of the embodiments of this application proposes a computer-readable storage medium storing a computer program that, when executed by a processor, implements the methods described above.
[0013] To achieve the above objectives, another aspect of the embodiments of this application proposes a computer program product, including a computer program that, when executed by a processor, implements the aforementioned method.
[0014] The embodiments of this application include at least the following beneficial effects: This application provides a behavior testing method, system, electronic device, storage medium, and program product for a multimodal perception humanoid robot. This solution acquires multimodal perception data and the current joint state data of the humanoid robot, and concatenates the multimodal data and the current joint state data to determine the current state vector; it then samples actions based on the current state vector and a trained processing model to obtain several sampled actions; further, it calculates based on the obtained sampled actions, the concatenated current state vector, and a preset running model, selecting several candidate action states from the obtained sampled actions; and finally, it calculates based on a preset... The network model identifies and calculates several candidate action states to obtain a total identification score, and determines the target action state from the candidate action states based on the total identification score. The humanoid robot is simulated and evaluated according to the preset scenario test cases and the determined target action state, and a corresponding test report is generated. By collecting multimodal perception data and combining it with the current joint state data of the humanoid robot, the dynamic adaptability of the humanoid robot's actions is improved. Based on the feedback of the real-time environment from the multimodal perception data, action sampling is performed based on the multimodal perception data and the current joint state data to determine the target behavior state for simulation, providing scenario adaptability. Attached Figure Description
[0015] Figure 1 This is a flowchart of a behavior testing method for a multimodal perceptive humanoid robot provided in an embodiment of this application; Figure 2 yes Figure 1 The flowchart of step S102 in the document; Figure 3 yes Figure 1 Another flowchart of step S102 in the process; Figure 4 yes Figure 1 The flowchart of step S103 in the process; Figure 5 yes Figure 1 Another flowchart of step S103 in the process; Figure 6 yes Figure 1 The flowchart of step S104 in the process; Figure 7 This is a flowchart illustrating a test performed in a specific embodiment provided in this application. Figure 8 This is a schematic diagram of the workflow of a multimodal sensing module in a specific embodiment provided in this application. Figure 9 This is a schematic diagram of the workflow of a policy neural network in a specific embodiment provided in this application. Figure 10This is a flowchart illustrating the calculation of the total discrimination score by the discrimination module in a specific embodiment provided in this application. Figure 11 This is a flowchart of a simulation test in a specific embodiment provided in this application; Figure 12 This is a schematic diagram of the structure of a behavior testing system for a multimodal humanoid robot provided in an embodiment of this application; Figure 13 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0016] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit it. In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with those of this application; they are merely examples of apparatuses and methods consistent with some aspects of the embodiments of this application as detailed in the appended claims.
[0017] It is understood that the terms “first,” “second,” etc., used in this application may be used herein to describe various concepts, but unless otherwise stated, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the words “if,” “when,” or “in response to a determination” as used herein may be interpreted as “when…” or “when…” or “in response to a determination.”
[0018] As used in this application, the terms "at least one", "multiple", "each", "any", etc., "at least one" includes one, two or more, "multiple" includes two or more, "each" refers to each of the corresponding multiples, and "any" refers to any one of the multiples.
[0019] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0020] This application provides a method for testing the behavior of a multimodal humanoid robot, relating to the field of information technology. This method can be applied to a terminal, a server, or software running on either a terminal or a server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, or in-vehicle terminal, but is not limited to these. The server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network. The software can be an application implementing the multimodal humanoid robot behavior testing method, but is not limited to the above forms.
[0021] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0022] It should be noted that in all specific embodiments of this application, when processing data related to user identity or characteristics, such as user information, user behavior data, user historical data, and user location information, user permission or consent is obtained first. Furthermore, the collection, use, and processing of this data comply with relevant laws, regulations, and standards. In addition, when embodiments of this application require access to sensitive personal information of users, separate permission or consent from the user is obtained through pop-ups or redirection to confirmation pages. Only after obtaining the user's separate permission or consent is the necessary user-related data required for the proper functioning of these embodiments acquired.
[0023] Figure 1This is an optional flowchart of a behavior testing method for a multimodal perceptual humanoid robot provided in an embodiment of this application. Figure 1 The method may include, but is not limited to, steps S101 to S104.
[0024] Step S101: Obtain multimodal sensing data and current joint state data, and concatenate the multimodal sensing data and joint state data to determine the current state vector; Step S102: Perform action sampling based on the current state vector and the trained processing model to determine several sampled actions, and calculate and determine several candidate action states based on the several sampled actions, the current state vector and the preset motion model. Step S103: Based on several candidate action states and a preset neural network model, perform identification calculations to determine several total identification scores, and analyze the several total identification scores and several candidate action states to determine the target behavior state. Step S104: Conduct test evaluation based on the target behavior state and preset scenario test cases, and determine the test report.
[0025] Steps S101 to S104, as illustrated in this embodiment, involve real-time data acquisition using data acquisition devices mounted on the bipedal humanoid robot. This includes real-time acquisition of scene images via binocular cameras, real-time acquisition of torque data for each joint of the humanoid robot via joint torque sensors, real-time acquisition of pressure distribution via tactile sensors, and acquisition of joint state data for each joint of the humanoid robot, including the current angle and angular velocity of each joint. The acquired real-time data undergoes corresponding preprocessing to obtain multimodal perception data and the robot's current joint data. In this embodiment, the acquired visual data undergoes noise reduction processing, and feature processing is performed using a target monitoring model to extract scene feature data. The acquired joint torque data undergoes low-pass filtering to reduce noise interference and normalization processing to obtain multi-dimensional joint torque data. The dimensions of the joint torque data are adaptively adjusted according to the joint degrees of freedom of the humanoid robot. The acquired joint state data is also normalized to obtain multi-dimensional joint state data. The dimensions of the tactile data are adaptively adjusted according to the joint degrees of freedom. The pressure center is calculated and normalized. The preprocessed data is then concatenated to obtain the current state vector representing the current state of the bipedal humanoid robot, the dimensions of which are equal to the sum of the dimensions of the collected data. Then, the current state vector is sampled using a pre-built and trained processing model, and processed in conjunction with a pre-set dual-arm collaboration mechanism to obtain sampled actions that the humanoid robot can execute. Then, a high-degree-of-freedom motion model is established based on the robot's current state vector and the sampled action inputs for analysis. The sampled actions that meet the robot's execution requirements are determined, and the action state at the next moment is calculated based on the sampled actions as candidate action states. This allows for subsequent discrimination score calculation of the selected actions. The total discrimination score is used to determine whether the safety and scene adaptability of the actions meet the corresponding requirements, thereby determining whether to conduct a simulation test on the humanoid robot based on the selected actions and generating a corresponding test report.
[0026] Please see Figure 2 In some embodiments, step S102 may include, but is not limited to, steps S201 to S203: Step S201: Extract features from the current state vector based on the trained processing model, determine multimodal features, and fuse the multimodal features to determine the fused feature data; Step S202: Perform full-connection processing based on the preset shared weights and fused feature data of the robot's two arms to determine the weight feature data; Step S203: Calculate the action space based on the weight feature data and the preset activation function, and sample actions based on the action space and the preset degree of freedom parameters to determine several sampled actions.
[0027] In step S201 of some embodiments, the current state vector obtained by data concatenation is input into the constructed and trained processing model for feature extraction, resulting in multimodal feature data including visual features, torque features, and tactile features. Feature fusion is then performed on the feature data of different modalities to enhance feature representation and obtain fused feature data. In this embodiment, a policy neural network is constructed to extract and fuse features from the current state vector. The policy neural network is configured with three convolutional layers to extract features from the input current state vector, and the extracted feature data is input into two fully connected layers for feature fusion to obtain fused feature data.
[0028] In step S202 of some embodiments, when performing feature fusion on the extracted multimodal feature data, a dual-arm weight sharing mechanism is introduced to improve the coordination of the humanoid robot's two arms, and the weights of the humanoid robot's two arms are shared; then, the obtained fused feature data is processed by full connection to obtain weighted feature data; in this embodiment, the sharing range of the dual-arm weights in the dual-arm sharing mechanism of the humanoid robot is adaptively adjusted according to the degrees of freedom of the two arms.
[0029] In step S203 of some embodiments, the processing model outputs the weight feature data output by the fully connected layer to the output layer for processing, calculates the weight feature data through the activation function, and outputs the probability distribution result of the action space; then, it performs action sampling based on the obtained probability distribution result of the action space and the preset degree of freedom parameters to obtain the sampled actions; in this embodiment, each degree of freedom of the humanoid robot is set with three sampled actions, that is, three actions are sampled for each degree of freedom in the action space.
[0030] Please see Figure 3 In some embodiments, step S102 may also include, but is not limited to, steps S301 to S304: Step S301: Analyze several sampling actions to determine several action motor parameters; wherein, the action motor parameters include the motor current and reduction ratio corresponding to the sampling action; Step S302: Calculate and determine several first torque vectors based on several motion motor parameters and a first preset formula, and adjust several first torque vectors according to preset torque thresholds to determine several joint torque vectors. Step S303: Analyze the current state vector to determine the current joint angle vector and the current joint angular velocity vector, and calculate and determine several joint angle vectors for the next moment based on the current joint angle vector, the current joint angular velocity vector, several joint torque vectors and the preset motion model. Step S304: Perform constraint processing based on several next-time joint angle vectors and preset joint limit thresholds to determine several candidate action states.
[0031] In step S301 of some embodiments, motion sampling is performed on the motion space to obtain the sampled motions of each degree of freedom of the humanoid robot. The sampled motions are then analyzed to determine the motor current and deceleration ratio corresponding to the sampled motions. This is so that the sampled motions can be analyzed to determine whether they satisfy the humanoid robot's ability to perform the motions. The motion state at the next moment is calculated based on the sampled motions and used as a candidate motion state for subsequent simulation testing of the humanoid robot.
[0032] In step S302 of some embodiments, the torque vector of the humanoid robot for that degree of freedom is calculated based on a preset formula and the motor parameters corresponding to the sampled actions of each degree of freedom. It is then determined whether the torque vector meets the torque threshold set for the corresponding application scenario, and the calculated torque vector is adjusted according to the set torque threshold to facilitate the subsequent calculation of the joint state data for each degree of freedom at the next moment. In this embodiment, the torque vector is calculated using the following formula: , in, For sampling action The corresponding joint torque vector, Let be the torque constant of the motor for this joint. For sampling action The corresponding motor current, This represents the reduction ratio of the motor.
[0033] In step S303 of some embodiments, the current state vector obtained by data splicing is parsed to determine the current joint angular velocity vector and joint velocity vector of each degree of freedom of the humanoid robot; then, the current joint angular velocity vector and joint velocity vector of each degree of freedom, as well as the torque vector calculated previously, are input into the high degree of freedom motion model for calculation to calculate the joint angle vector of each degree of freedom at the next moment; in this embodiment, the high degree of freedom motion model is: , in, This is the joint angle vector for the next moment. This is the joint angle vector at the current moment. For time step, This is the joint angular velocity vector at the current moment. for The diagonal moment coefficient matrix is set according to the differences in joint load. The total number of degrees of freedom for a humanoid robot. This is the joint torque vector corresponding to the current sampling action. for This is the diagonal damping coefficient matrix. for The vector of static friction coefficients, This is the sign function for angular velocity.
[0034] In step S304 of some embodiments, the joint angle vector at the next moment is analyzed based on the calculated joint angle vector and the preset joint angle threshold to determine whether the calculated joint angle vector at the next moment exceeds the angle limit of the current degree of freedom, and the joint angle vector that exceeds the angle limit is truncated to the limit value of the angle limit as a candidate action state for the next moment.
[0035] Please see Figure 4 In some embodiments, step S103 may include, but is not limited to, steps S401 to S404: Step S401: Analyze several candidate action states to determine several torque data, several visual data and several tactile data; Step S402: Perform fully connected processing based on several torque data and several visual data to determine the first output data, and perform activation calculation on the first output data to determine the safety sub-score; Step S403: Calculate and determine the aptamer score based on several visual data and several tactile data; Step S404: Calculate and determine the total discrimination score based on the security sub-score, the adaptation sub-score, and the preset weights.
[0036] In step S401 of some embodiments, a total discrimination score is calculated based on the calculated candidate action state. The total discrimination score is used to determine whether the candidate action state meets the safety requirements in the corresponding scenario and whether the candidate action is suitable for the corresponding scenario. In order to select the optimal action from several candidate actions calculated for each degree of freedom based on the total discrimination score for simulation testing. In this embodiment, by analyzing several candidate actions for each degree of freedom, the corresponding torque data, visual data, and tactile data are determined so that the total discrimination score can be calculated based on the torque data, visual data, and tactile data.
[0037] In step S402 of some embodiments, the total discrimination score of each candidate action is calculated by a constructed multimodal discriminator neural network. This multimodal discriminator neural network has two branches, which calculate the total discrimination score based on the input multimodal data, including calculating the safety sub-score and the fit sub-score. The multimodal discriminator neural network extracts the torque vector from the input torque data and the obstacle distance data from the visual data. It processes the extracted torque vector and obstacle distance data through a fully connected layer and calculates the corresponding safety sub-score through an activation function. At the same time, it makes a comprehensive judgment based on the torque vector and the set scene threshold, and based on the obstacle distance data and the set distance threshold, to determine the corresponding safety sub-score range. In this embodiment, when the torque vector is less than or equal to the set scene threshold (e.g., 30 N·m for industrial scenes and 20 N·m for household scenes), and the obstacle distance is greater than or equal to 5 cm, the safety sub-score is greater than or equal to 0.8; otherwise, it is less than or equal to 0.3.
[0038] In step S403 of some embodiments, another branch of the multimodal discriminator neural network processes visual and tactile data to determine assembly accuracy in an industrial setting or motion compliance in a home setting; determines the corresponding accuracy threshold based on the degrees of freedom of the humanoid robot's arms to ensure that it matches the degree of freedom capability; and determines the corresponding fitter score based on the assembly accuracy or motion compliance.
[0039] In step S404 of some embodiments, the calculated security sub-score and adaptation sub-score are used to calculate the total identification score by dynamically adjusting the weights according to the scenario risk level and performing a weighted sum of the security sub-score and adaptation score. , in, The total discrimination score, As weight, For safe sub-fractions, To adapt the sub-scores; in this embodiment, in industrial scenarios, the weight ranges from 0.7 to 0.8 to prioritize safety in high-risk scenarios, while in home scenarios, the weight ranges from 0.5 to 0.6 to balance safety and flexibility in medium-risk scenarios.
[0040] Please see Figure 5 In some embodiments, step S103 may also include, but is not limited to, steps S501 to S503: Step S501: Compare several total discrimination scores with preset scene thresholds respectively; Step S502: If the total identification score is greater than or equal to the preset scene threshold, the candidate action state corresponding to the total identification score is taken as the target behavior state. Step S503: If the total discrimination score is less than the preset scene threshold, return to execute the action sampling based on the current state vector and the trained processing model, determine several sampling actions, until the total discrimination score is greater than or equal to the preset scene threshold.
[0041] In step S501 of some embodiments, the total discrimination score corresponding to each candidate action is determined through the preceding calculation process, and the action state at the next moment is selected by comparing the total discrimination score corresponding to each candidate action with the corresponding scene threshold.
[0042] In step S502 of some embodiments, if the total discrimination score corresponding to the candidate action is greater than or equal to the corresponding scene threshold, the candidate action is taken as the target behavior state so that the humanoid robot can conduct simulation tests in the future. In this embodiment, the candidate actions can also be sorted according to the total discrimination scores corresponding to each degree of freedom, and the candidate action with the highest total discrimination score is selected as the target action state.
[0043] In step S503 of some embodiments, if the total discrimination score corresponding to the candidate action is less than the corresponding scene threshold, the action space is returned to perform action sampling, the number of candidate actions corresponding to the degree of freedom is re-determined, and the corresponding total discrimination score is recalculated for judgment to determine the target action state.
[0044] Please see Figure 6 In some embodiments, step S104 may also include, but is not limited to, steps S601 to S602: Step S601: Perform behavioral simulation based on preset scenario test cases and target behavioral states, and determine the simulation results; Step S602: Calculate the motion output quality based on the simulation results and preset evaluation indicators, and determine the test report based on the motion output quality.
[0045] In step S601 of some embodiments, after obtaining the target action state corresponding to each degree of freedom, the target action state of each degree of freedom of the humanoid robot is tested according to the pre-set scenario test cases, and the corresponding test results are determined, including action execution quality, action safety, etc. In this embodiment, the scenario test cases include industrial assembly part offset, household obstacle occlusion, etc., and the action output quality of the humanoid robot is evaluated through behavior simulation of multiple time steps.
[0046] In step S602 of some embodiments, the target action states of each degree of freedom of the humanoid robot are tested according to the scenario test cases. The number of actions that meet the scenario requirements or test requirements in the test scenario is counted. The corresponding evaluation indicators, such as success rate and safety violation rate, are calculated based on the count and the total number of test simulations. Then, the action output quality is determined and a corresponding test report is generated.
[0047] The following is a detailed description and explanation of the solutions in the embodiments of the present invention, using specific application examples: Please see Figure 7 , Figure 7 This is a system flowchart illustrating a behavior testing method for a multimodal perception humanoid robot provided in this application, applied in a specific embodiment. A multimodal perception module is set up for data acquisition, including acquiring visual data via a binocular camera; acquiring corresponding torque vectors based on the humanoid robot's degrees of freedom (e.g., in an industrial scenario of bearing assembly with 28 degrees of freedom, 14 degrees of freedom for the two arms, an assembly precision less than or equal to 0.1 mm, and a joint torque threshold of 30 N·m); acquiring 28-dimensional torque vector data via joint torque sensors; acquiring tactile data from the humanoid robot via tactile sensors to calculate the robot's pressure center of gravity offset data; and processing the acquired multimodal perception data through a data preprocessing unit. The workflow of the multimodal perception module is as follows: Figure 8 As shown, the system adapts torque data or joint data of the corresponding degree of freedom, and then the data fusion unit fuses the output data of the data preprocessing unit to output a 2136-dimensional state vector. The behavior generation module processes the 2136-dimensional state vector through a policy neural network to generate the corresponding action space. The workflow of the policy neural network is as follows: Figure 9 As shown, convolution processing is performed through three convolutional layers, employing a dual-arm degree-of-freedom weight sharing mechanism and fully connected layers. The motion space is output through a softmax output layer, and motion sampling is performed within this space. The adopted motions of the humanoid robot are: left arm shoulder joint -2°, right arm shoulder joint +1°, torso pitch ±10°, lateral tilt ±5°, and lower limb hip joint angle error less than or equal to 2°. The next motion state is calculated using a high-degree-of-freedom motion model, yielding 28.1° (left arm) and 31.2° (right arm) as candidate states. The discrimination module uses... Figure 10The process shown calculates the total discrimination score. The multimodal discrimination neural network uses a branch network to evaluate candidate states, obtaining a safety sub-score of 0.92 and an adaptation sub-score of 0.88. Based on the weight of 0.75 configured in the score combination unit, the total discrimination score for each candidate state is calculated to be 0.91. The multimodal discrimination neural network is trained using real robot safety behavior data from different scenarios, and cross-entropy loss is used as the training loss function. The cross-entropy loss is as follows: , in, For cross-entropy loss, The model outputs a total discrimination score, where y=1 corresponds to truly safe data and y=0 corresponds to unsafe data. Based on preset test case input units and degree-of-freedom adaptation rules, the humanoid robot is input into the control test module to conduct simulated tests. Figure 11 As shown, the assembly of a 50mm diameter bearing and a 50.05mm inner diameter bushing was performed, including gripping, alignment, and pressing. The test results were evaluated by the performance evaluation unit. The force feedback was determined to be 28-dimensional data, the torque vector was 25 N·m, which is less than the scene threshold, the assembly accuracy was 0.1mm, the success rate of the humanoid robot was determined to be 92%, and the safety violation rate was 0.8%, which met the test requirements.
[0048] The embodiments of this application include at least the following beneficial effects: This application provides a behavior testing method, system, electronic device, storage medium, and program product for a multimodal perception humanoid robot. This solution acquires multimodal perception data and the current joint state data of the humanoid robot, and concatenates the multimodal data and the current joint state data to determine the current state vector; it then samples actions based on the current state vector and a trained processing model to obtain several sampled actions; further, it calculates based on the obtained sampled actions, the concatenated current state vector, and a preset running model, selecting several candidate action states from the obtained sampled actions; and finally, it calculates based on a preset... The network model identifies and calculates several candidate action states to obtain a total identification score, and determines the target action state from the candidate action states based on the total identification score. The humanoid robot is simulated and evaluated according to the preset scenario test cases and the determined target action state, and a corresponding test report is generated. By collecting multimodal perception data and combining it with the current joint state data of the humanoid robot, the dynamic adaptability of the humanoid robot's actions is improved. Based on the feedback of the real-time environment from the multimodal perception data, action sampling is performed based on the multimodal perception data and the current joint state data to determine the target behavior state for simulation, providing scenario adaptability.
[0049] Please see Figure 12This application also provides a behavior testing system for a multimodal perception humanoid robot, which can implement the above-described method. The system includes: The data acquisition module is used to acquire multimodal sensing data and current joint state data, and to concatenate the multimodal sensing data and the joint state data to determine the current state vector. The action sampling module is used to sample actions based on the current state vector and the trained processing model, determine several sampled actions, and calculate based on the several sampled actions, the current state vector and the preset motion model to determine several candidate action states. The identification calculation module is used to perform identification calculations based on several candidate action states and a preset neural network model, determine several total identification scores, and analyze the several total identification scores and several candidate action states to determine the target behavior state. The test evaluation module is used to evaluate the test based on the target behavior state and preset scenario test cases, and to determine the test report.
[0050] It is understood that the content of the above method embodiments is applicable to the present device embodiments. The specific functions implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0051] This application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described method. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.
[0052] It is understood that the content of the above method embodiments is applicable to this device embodiment. The specific functions implemented by this device embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0053] Please see Figure 13 , Figure 13 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes: The processor 1301 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application. The memory 1302 can be implemented as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 1302 can store the operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1302 and is called and executed by the processor 1301 using the methods described in the embodiments of this application. The input / output interface 1303 is used to implement information input and output; The communication interface 1304 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.). Bus 1305 transmits information between various components of the device (e.g., processor 1301, memory 1302, input / output interface 1303, and communication interface 1304); The processor 1301, memory 1302, input / output interface 1303 and communication interface 1304 are connected to each other within the device via bus 1305.
[0054] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method.
[0055] It is understood that the content of the above method embodiments is applicable to this storage medium embodiment. The specific functions implemented in this storage medium embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.
[0056] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.
[0057] It is understood that the content of the above method embodiments is applicable to the embodiments of this program product. The specific functions implemented by the embodiments of this program product are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0058] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0059] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.
[0060] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.
[0061] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0062] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.
[0063] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0064] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0065] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0066] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0067] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0068] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0069] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.
Claims
1. A method for testing the behavior of a multimodal perception humanoid robot, characterized in that, The method includes: Acquire multimodal sensing data and current joint state data, and concatenate the multimodal sensing data and joint state data to determine the current state vector; Action sampling is performed based on the current state vector and the trained processing model to determine several sampled actions. Then, based on the several sampled actions, the current state vector, and the preset motion model, several candidate action states are determined. Based on several candidate action states and a preset neural network model, an identification calculation is performed to determine several total identification scores. Based on the several total identification scores and several candidate action states, an analysis is performed to determine the target behavior state. Based on the target behavior state and preset scenario test cases, a test evaluation is conducted, and a test report is determined.
2. The method according to claim 1, characterized in that, The step of sampling actions based on the current state vector and the trained processing model to determine several sampling actions specifically includes: Based on the trained processing model, feature extraction is performed on the current state vector to determine multimodal features, and feature fusion is performed on the multimodal features to determine fused feature data; Based on the preset shared weights of the robot's two arms and the fused feature data, a fully connected processing is performed to determine the weight feature data; The action space is determined by calculation based on the weight feature data and the preset activation function, and the action space and preset degree of freedom parameters are used to sample actions to determine a number of sampled actions.
3. The method according to claim 1, characterized in that, The step of determining several candidate action states based on several sampled actions, the current state vector, and a preset motion model specifically includes: The sampling actions are analyzed to determine several motor parameters; wherein, the motor parameters include the motor current and reduction ratio corresponding to the sampling action. Based on several motor parameters and a first preset formula, several first torque vectors are determined, and several first torque vectors are adjusted according to preset torque thresholds to determine several joint torque vectors. The current state vector is parsed to determine the current joint angle vector and the current joint angular velocity vector. Based on the current joint angle vector, the current joint angular velocity vector, several joint torque vectors, and the preset motion model, several joint angle vectors for the next moment are calculated to determine. Constraints are applied based on several next-moment joint angle vectors and preset joint limit thresholds to determine several candidate action states.
4. The method according to claim 1, characterized in that, The step of performing identification calculations based on several candidate action states and a preset neural network model to determine several total identification scores specifically includes: The candidate action states are analyzed to determine several torque data, several visual data, and several tactile data. Based on several torque data and several visual data, a fully connected processing is performed to determine the first output data, and activation calculation is performed on the first output data to determine the safety sub-score; The aptamer score is determined by calculating based on several visual data points and several tactile data points. The total discrimination score is determined by calculating the security sub-score, the adaptation sub-score, and the preset weight.
5. The method according to claim 1, characterized in that, The step of analyzing several total discrimination scores and several candidate action states to determine the target behavior state specifically includes: Each of the total identification scores is compared with a preset scene threshold. If the total identification score is greater than or equal to the preset scene threshold, the candidate action state corresponding to the total identification score is taken as the target behavior state; If the total discrimination score is less than the preset scene threshold, return to the process of sampling actions based on the current state vector and the trained processing model to determine several sampling actions until the total discrimination score is greater than or equal to the preset scene threshold.
6. The method according to claim 1, characterized in that, The step of evaluating the test based on the target behavior state and preset scenario test cases, and determining the test report, specifically includes: Behavioral simulation is performed based on the preset scenario test cases and the target behavioral state, and the simulation results are determined. The simulation results and preset evaluation indicators are used to calculate and determine the quality of the action output, and the test report is determined based on the quality of the action output.
7. A behavior testing system for a multimodal perception humanoid robot, characterized in that, The system includes: The data acquisition module is used to acquire multimodal sensing data and current joint state data, and to concatenate the multimodal sensing data and the joint state data to determine the current state vector. The action sampling module is used to sample actions based on the current state vector and the trained processing model, determine several sampled actions, and calculate based on the several sampled actions, the current state vector and the preset motion model to determine several candidate action states. The identification calculation module is used to perform identification calculations based on several candidate action states and a preset neural network model, determine several total identification scores, and analyze the several total identification scores and several candidate action states to determine the target behavior state. The test evaluation module is used to evaluate the test based on the target behavior state and preset scenario test cases, and to determine the test report.
8. An electronic device, characterized in that, include: At least one processor; At least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the method as described in any one of claims 1 to 6.
9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 6.
Citation Information
Patent Citations
Human body action recognition method and device based on multi-modal fusion
CN116758628A
Mechanical arm control method and system based on multi-mode driving and storage medium
CN118752495A
Multi-modal sensing humanoid robot action self-adaptive control method and multi-modal sensing humanoid robot action self-adaptive control system
CN119610112A
Humanoid robot motion training method and device based on real scene
CN120116237A
Systems and Methods Automatic Anomaly Detection in Mixed Human-Robot Manufacturing Processes
US20210170590A1