End-to-end control method, system, equipment and product of humanoid robot
By using an end-to-end control method and leveraging a multimodal perception dataset and scene adaptation model, optimal control commands are generated, solving the problem of high error in multi-module collaborative control of humanoid robots and achieving efficient and safe robot operation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-31
- Publication Date
- 2026-04-14
AI Technical Summary
In existing technologies, humanoid robots have relatively high errors in the process of multi-module collaborative control, which affects work efficiency, especially in industrial assembly and home scenarios where it is difficult to meet the requirements for precision and safety.
An end-to-end control approach is adopted. By acquiring multimodal perception datasets and task instructions, a scene adaptation model is constructed, and a scene-based reward function is introduced to generate control instructions, thereby achieving direct mapping from perception to actuation and eliminating modular collaborative errors.
It improves the control efficiency of humanoid robots, reduces multimodal data timing synchronization errors, and enhances the collaborative efficiency and safety of robots in different scenarios.
Smart Images

Figure CN121857447A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of robotics technology, and in particular to an end-to-end control method, system, device and product for a humanoid robot. Background Technology
[0002] In related technologies, some robots employ modular control, where each module relies on manually defined data interfaces for collaborative control. However, in practical applications, it has been found that the coordinates output by the robot's vision module require multiple levels of data transformation before being transmitted to the control module. This results in high errors in the collaborative control of multiple modules, impacting the working efficiency of the humanoid robot.
[0003] In summary, the technical problems existing in the relevant technologies need to be improved. Summary of the Invention
[0004] The main objective of this application is to provide an end-to-end control method, system, device, and product for humanoid robots, which can improve the working efficiency of humanoid robots.
[0005] To achieve the above objectives, one aspect of this application proposes an end-to-end control method for a humanoid robot, the method comprising: Acquire multimodal perception datasets and task instructions for humanoid robots in target scenarios; A scene adaptation model is constructed based on the target scene and the multimodal perception dataset; The scene adaptation model is guided by a preset scene-based reward function to generate control commands. A control interface is generated based on the multimodal perception dataset, the task instructions, and the control instructions, and the humanoid robot is controlled end-to-end based on the control interface.
[0006] In some embodiments, the construction of the scene adaptation model based on the target scene and the multimodal perception dataset includes: Based on the target scene, a set of scene modal features is extracted from the multimodal perception dataset; Noise optimization is performed on each feature in the scene modal feature set to obtain an optimized feature set; The fused features are obtained by fusing all features in the optimized feature set; Based on the target scene, the scene modal feature set is constrained to generate scene constraint terms; The scene adaptation model is obtained by training a preset reinforcement learning model based on the fusion features and the scene constraints.
[0007] In some embodiments, the step of generating scene constraint terms by constraining the scene modal feature set based on the target scene includes: When the target scenario is an industrial scenario, robot force perception features and workpiece position features are extracted from the scenario modal feature set; The workpiece position features are multiplied based on preset mapping coefficients, and the multiplication results are subtracted based on the robot force perception features to obtain the scene constraint terms.
[0008] In some embodiments, the step of generating scene constraint terms by constraining the scene modal feature set based on the target scene further includes: When the target scenario is a home scenario, target semantic features and target behavioral features are extracted from the scenario modal feature set; The scene constraint term is obtained by calculating the similarity between the target semantic features and the target behavioral features.
[0009] In some embodiments, the process of guiding the scene adaptation model to generate control instructions based on a preset scene-based reward function includes: The scene adaptation model outputs multiple candidate instructions in the target scene. The control command is obtained by selecting the optimal command from the candidate commands based on a preset scenario-based reward function.
[0010] In some embodiments, the step of selecting the optimal control instruction from the candidate instructions based on a preset scenario-based reward function includes: Select the corresponding scenario-based reward function based on the target scenario; The candidate instruction is input into the task simulation environment, the state parameters after the candidate instruction is executed are collected, and the cumulative reward value corresponding to the candidate instruction is calculated by substituting them into the scenario-based reward function. The accumulated reward values are sorted, and the candidate instruction with the largest accumulated reward value is selected as the control instruction.
[0011] In some embodiments, generating a control interface based on the multimodal perception dataset, the task instructions, and the control instructions, and performing end-to-end control of the humanoid robot based on the control interface, includes: The multimodal perception dataset, the task instructions, and the control instructions are subjected to feature concatenation and encapsulation processing to obtain the control interface; Based on the control interface, joint drive signals are generated through the underlying controller of the humanoid robot; The humanoid robot is controlled end-to-end via the joint drive signals.
[0012] To achieve the above objectives, another aspect of this application proposes an end-to-end control system for a humanoid robot, the system comprising: The data acquisition module is used to acquire multimodal perception datasets and task instructions of the humanoid robot in the target scene; The model building module is used to build a scene adaptation model based on the target scene and the multimodal perception dataset; The instruction generation module is used to guide the scene adaptation model to generate control instructions based on a preset scene-based reward function. An end-to-end control module is used to generate a control interface based on the multimodal perception dataset, the task instructions, and the control instructions, and to perform end-to-end control of the humanoid robot based on the control interface.
[0013] To achieve the above objectives, another aspect of this application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the method described above.
[0014] To achieve the above objectives, another aspect of the embodiments of this application proposes a computer-readable storage medium storing a computer program that, when executed by a processor, implements the methods described above.
[0015] To achieve the above objectives, another aspect of this application provides a computer program product, including a computer program that, when executed by a processor, implements the methods described above. The embodiments of this application include at least the following beneficial effects: This application provides an end-to-end control method, system, device, and product for a humanoid robot. This solution acquires a multimodal perception dataset and task instructions for the humanoid robot in a target scene, and constructs a scene adaptation model based on the target scene and the multimodal perception dataset. This allows the introduction of scene constraints on the model, improving the model's accurate cognitive ability for different scenes. Furthermore, this solution guides the scene adaptation model to generate control instructions based on a preset scene-based reward function, enabling the model to generate optimal control instructions and improving the control efficiency of the humanoid robot. Moreover, this solution generates a control interface based on the multimodal perception dataset, task instructions, and control instructions, and performs end-to-end control of the humanoid robot based on the control interface. This enables the construction of a low-latency, highly compatible multimodal unified interface, achieving direct mapping from perception to actuation, eliminating modular collaborative errors, and improving the robot's working efficiency. Attached Figure Description
[0016] Figure 1This is a schematic diagram of an implementation environment provided in an embodiment of this application; Figure 2 This is a flowchart of an end-to-end control method for a humanoid robot provided in an embodiment of this application; Figure 3 This is a schematic diagram of the structure of an end-to-end control system for a humanoid robot provided in an embodiment of this application; Figure 4 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0017] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit it. In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with those of this application; they are merely examples of systems and methods consistent with some aspects of the embodiments of this application as detailed in the appended claims.
[0018] It is understood that the terms “first,” “second,” etc., used in this application may be used herein to describe various concepts, but unless otherwise stated, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the words “if,” “when,” or “in response to a determination” as used herein may be interpreted as “when…” or “when…” or “in response to a determination.”
[0019] As used in this application, the terms "at least one", "multiple", "each", "any", etc., "at least one" includes one, two or more, "multiple" includes two or more, "each" refers to each of the corresponding multiples, and "any" refers to any one of the multiples.
[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0021] Before providing a detailed description of the embodiments of this application, some of the nouns and terms involved in the embodiments of this application will be explained first. The nouns and terms involved in the embodiments of this application are subject to the following interpretations.
[0022] 1) Humanoid robots, also known as bionic robots, are robots designed to mimic human appearance and behavior, especially those with similar physiques to humans. The structural design of humanoid robots represents a remarkable reshaping of the human body, requiring not only interdisciplinary integration but also the culmination of cutting-edge technologies. Their design principles primarily include the following aspects: the organic integration of bionics and mechanical engineering, breakthroughs in the integration of sensing technology and control theory, and precise coordination between drive mechanisms and execution actions.
[0023] 2) Reinforcement Learning (RL) is a machine learning method. Its fundamental framework is the Markov Decision Process, which allows an agent to learn optimal policies through trial and error in its interactions with the environment. The agent performs actions in the environment and receives feedback, or rewards, based on the outcomes of those actions. These reward signals guide the agent to adjust its policy to maximize long-term cumulative rewards.
[0024] In related technologies, the control methods for humanoid robots generally adopt a split architecture of "perception-decision-control," with each module relying on manually defined data interfaces, resulting in significant coordination errors. For example, in industrial assembly control methods for humanoid robots, the workpiece coordinates output by the vision module need to undergo three levels of data conversion before being transmitted to the control module. The coordination error between multiple modules reaches 50 to 80 ms, which cannot meet the real-time requirements of precision assembly, such as the M5 bolt assembly which requires a positional deviation of ≤0.5 mm. In home scenarios, the interactive versions of related products rely on fixed voice command libraries, and the timing synchronization error between the voice and motion control modules is large, resulting in a significant delay in the robot's response after a "take a cup" command is given. For example, in human-robot collaborative assembly, sudden adjustments to human movements, such as changes in the angle of part grip, can lead to collision risks for the robot. Furthermore, preset trajectories cannot balance task accuracy and safe distance; for example, if the robot continuously approaches the target position when tightening bolts, it may conflict with the hands of production line workers, lacking comfortable control over human-robot contact forces. The relevant control methods are prone to damage to parts or discomfort to production line workers due to improper force application. They do not have game strategies designed for the high safety and high precision requirements of industrial scenarios, making them difficult to apply directly to human-machine hybrid operation scenarios.
[0025] In view of this, this application provides an end-to-end control method, system, device, and product for a humanoid robot, which can be applied to human-computer interaction application scenarios. Specifically, the end-to-end control method for a humanoid robot provided in this application can be applied to the controller of the humanoid robot or to the server controlling the humanoid robot. Taking its application to the controller as an example, it can generate specific control programs and execute specific action generation steps based on the controller. This application constructs a scene adaptation model based on the target scene and multimodal perception dataset. Differentiated constraints can be designed for industrial and home scenarios, enabling the model to adapt to multiple application scenarios. Furthermore, by constructing a unified multimodal interface to achieve direct mapping from perception to actuation, it can reduce multimodal data timing synchronization errors, improve the collaborative efficiency of the humanoid robot, and thus improve the robot's control efficiency.
[0026] It should be noted that in all specific embodiments of this application, when processing data related to user identity or characteristics, such as user information, user behavior data, user historical data, and user location information, user permission or consent is obtained first. Furthermore, the collection, use, and processing of this data comply with relevant laws, regulations, and standards. In addition, when embodiments of this application require access to sensitive personal information of users, separate permission or consent from the user is obtained through pop-ups or redirection to confirmation pages. Only after obtaining the user's separate permission or consent is the necessary user-related data required for the proper functioning of these embodiments acquired.
[0027] Figure 1 This is a schematic diagram illustrating the implementation environment of a method provided in an embodiment of this application. (Refer to...) Figure 1 The main hardware and software components of this implementation environment include a humanoid robot 101 and a server 102, with the humanoid robot 101 and server 102 communicating with each other. The method can be executed based on the interaction between the humanoid robot 101 and server 102. Furthermore, the humanoid robot 101 and server 102 can be nodes in a blockchain, but this embodiment does not specifically limit this.
[0028] Figure 2 This is an optional flowchart of an end-to-end control method for a humanoid robot provided in an embodiment of this application. Figure 2 The method may include, but is not limited to, steps S102 to S204.
[0029] Step S201: Obtain the multimodal perception dataset and task instructions of the humanoid robot in the target scene; Step S202: Construct a scene adaptation model based on the target scene and the multimodal perception dataset; Step S203: Based on a preset scenario-based reward function, guide the scenario adaptation model to generate control instructions; Step S204: Generate a control interface based on the multimodal perception dataset, the task instructions, and the control instructions, and perform end-to-end control of the humanoid robot based on the control interface.
[0030] Steps S201 to S204 of this embodiment involve acquiring a multimodal perception dataset and task instructions for the humanoid robot in a target scenario. The target scenario is the environment in which the humanoid robot is used, which may include industrial scenarios, home scenarios, etc. Industrial scenarios may include scenarios such as precision assembly of automotive parts and intelligent handling of heavy workpieces, while home scenarios may include scenarios such as retrieving and placing household items and interaction for elderly / child care. The multimodal perception dataset may include data collected by the humanoid robot through various sensors, including visual data, voice data, and environmental data. The task instructions are the operation commands given by the user to the robot, such as instructing the robot to process parts. Then, based on the target scenario, corresponding data is selected from the multimodal perception dataset, processed, and a scenario adaptation model is constructed. This scenario adaptation model is used to generate control instructions for the humanoid robot according to different scenarios and can be constructed using a neural network model. Then, a preset scenario-based reward function guides the scenario adaptation model to generate control instructions. This scenario-based reward function is used to quantitatively evaluate the scenario adaptability of candidate instructions, enabling the model to generate the optimal control instructions. Finally, based on the multimodal perception dataset, task instructions, and control instructions, data fusion and encapsulation are performed to generate a control interface. Based on the control interface, the humanoid robot performs end-to-end control to complete the task instructions issued by the user.
[0031] One of the above technical solutions has the following advantages or beneficial effects: This application embodiment collects multi-dimensional perception data and task instructions from a humanoid robot in a target scene, associating the data modalities with scene requirements. Furthermore, a scene adaptation model is constructed using the multi-modal data, introducing a scene-specific constraint model for accurate scene recognition. This application embodiment also designs a differentiated scene reward function, which can quantitatively evaluate the scene adaptability of candidate instructions generated by the model, selecting the optimal control instruction. Finally, the multi-modal data and control instructions are used to construct a low-latency, highly compatible multi-modal unified interface, enabling direct mapping from perception to actuation of the robot, eliminating modular collaborative errors, and achieving end-to-end control of the humanoid robot.
[0032] In step S201 of some embodiments, the humanoid robot acquires a multimodal perception dataset and task instructions in the target scene; In this embodiment, multimodal datasets for different target scenarios can be acquired using multiple sensing devices or data acquisition devices of the humanoid robot. These target scenarios can include industrial and residential scenarios. For example, in an industrial scenario, point clouds of the workstation environment can be acquired using 3D LiDAR, detailed images of the workpiece can be acquired using a camera, torque values of each joint can be acquired using joint torque sensors, and environmental temperature and humidity can be collected using temperature and humidity sensors. The task instructions are structured instructions issued by an industrial host computer. In a residential scenario, user behavior and home layout images can be acquired using the humanoid robot's torso RGB-D camera, user voice commands can be acquired using a head microphone array, grasping pressure can be acquired using a head tactile sensor, and furniture positions can be acquired using LiDAR.
[0033] One of the above technical solutions has the following advantages or beneficial effects: This application embodiment collects multi-dimensional perception data and task instructions of a humanoid robot in a target scene, so that the data modality is associated with the scene requirements, which can provide a data foundation for subsequent generation of control instructions.
[0034] In step S202 of some embodiments, constructing a scene adaptation model based on the target scene and the multimodal perception dataset includes: Based on the target scene, a set of scene modal features is extracted from the multimodal perception dataset; Noise optimization is performed on each feature in the scene modal feature set to obtain an optimized feature set; The fused features are obtained by fusing all features in the optimized feature set; Based on the target scene, the scene modal feature set is constrained to generate scene constraint terms; The scene adaptation model is obtained by training a preset reinforcement learning model based on the fusion features and the scene constraints.
[0035] In this embodiment, different data are extracted from the multimodal perception dataset according to different target scenarios to obtain a scene modal feature set. For example, in an industrial scenario, a convolutional encoder is used to perform multimodal feature extraction processing on RGB-D visual images, joint torque time-series data, environmental sensing data, etc., in the multimodal perception dataset. In a home scenario, RGB-D visual images, hand tactile pressure matrices, furniture position point clouds, speech waveforms, etc., are extracted, and corresponding feature data are obtained based on the encoder, thereby obtaining the scene modal feature set.
[0036] For example, the convolutional encoder includes three convolutional layers, two batch normalization layers, and one fully connected layer. The convolutional layers have 3×3 kernels, a stride of 2, and ReLU activation to achieve feature downsampling. The batch normalization layers are located after the first and second convolutional layers, which can accelerate convergence and improve generalization. The fully connected layer is used to map multimodal data into low-dimensional feature vectors.
[0037] In some embodiments, this application further obtains an optimized feature set by performing noise optimization on each feature in the scene modal feature set. The noise optimization process includes random noise addition and denoising. Specifically, this application adds Gaussian noise to each single-modal feature in the scene modal feature set to expand the dataset and avoid model overfitting. The formula for random noise addition is shown below: ; In the formula, Represents the original features of a single mode. Represents the noise figure, where, in industrial scenarios, it is set... At that time, the model's adaptability to workpiece tolerances improved; in home scenarios, settings were made... At the same time, the recognition rate of ambiguous voice commands is improved; This represents Gaussian noise, used to balance noise intensity and characteristic fidelity.
[0038] Then, in this embodiment, noise is eliminated using the UNet denoising network. This denoising network can employ 4 downsampling layers with parameters set to 3×3 convolutional kernels and a stride of 2; and 4 upsampling layers set to transposed convolutions; it also includes 2 attention modules in between. Downsampling is used to progressively compress the feature dimension; upsampling is used for transposed convolution operations to progressively restore the feature dimension; and the attention modules are used to capture long-distance feature dependencies, improving denoising accuracy.
[0039] In some embodiments, all features in the optimized feature set are fused to obtain fused features; Specifically, in this embodiment, after noise optimization processing, an optimized feature set is obtained, and all features in the optimized feature set are fused. Temporary features can be obtained by concatenating the features in the feature set along the channel dimension. Then, the weights of each modality feature are learned through a multi-head self-attention layer, and weighted features are output. Finally, the weighted features are mapped to fused features of a unified dimension through a fully connected layer.
[0040] In some embodiments, the step of generating scene constraint terms by constraining the scene modal feature set based on the target scene includes: When the target scenario is an industrial scenario, robot force perception features and workpiece position features are extracted from the scenario modal feature set; The workpiece position features are multiplied based on preset mapping coefficients, and the multiplication results are subtracted based on the robot force perception features to obtain the scene constraint terms.
[0041] In this embodiment, different scene constraints are set for different target scenarios. When the target scenario is an industrial scenario, robot force features and workpiece position features can be extracted from the scene modal feature set. Then, the workpiece position features are multiplied based on a preset mapping coefficient to obtain the multiplication result. This mapping coefficient is a force-position mapping coefficient, which can be adjusted according to the workpiece material and is used to quantify the mapping relationship between torque and position. The robot force features can be obtained by encoding joint torque sensor data through an encoder. The scene constraints are obtained by subtracting the multiplication result of the robot force features. Then, in this embodiment, the fused features and scene constraints are added together, and the resulting data is input into a pre-set reinforcement learning model for training to obtain a scene adaptation model. The expression of the feature vector input to the scene adaptation model is as follows: ; In the formula, This represents the input feature vector, which is used for the generation of subsequent control commands. Indicates fusion characteristics; This represents the force-position co-constraint weighting coefficient, used to control the degree of influence of the constraint terms on the model; Indicates the robot's force perception characteristics, Indicates the force-position mapping coefficient. This indicates the positional characteristics of the workpiece.
[0042] In some embodiments, the step of generating scene constraint terms by constraining the scene modal feature set based on the target scene further includes: When the target scenario is a home scenario, target semantic features and target behavioral features are extracted from the scenario modal feature set; The scene constraint term is obtained by calculating the similarity between the target semantic features and the target behavioral features.
[0043] In this embodiment, when the target scene is a home scene, target semantic features and target behavioral features are extracted from the scene modal feature set. Then, the feature similarity between the target semantic features and the target behavioral features is calculated, and the similarity calculation result is used as a scene constraint term under the home scene. The fused features and scene constraint term are added together and then input into a pre-set reinforcement learning model for training, thereby obtaining a scene adaptation model. The expression for the feature vector input to the scene adaptation model is as follows: ; In the formula, This represents the feature vector input in a home environment. Indicates fusion characteristics; This represents the weight coefficient for intent matching constraints, used to control the degree of influence of intent terms on the model; The semantic features of the target can be obtained by encoding microphone array speech data using a pre-trained speech encoder. The target behavior features can be represented by extracting user behavior features based on the 3D skeleton.
[0044] One of the above technical solutions has the following advantages or beneficial effects: By constructing different differentiated constraint terms for different target scenarios, the embodiments of this application can improve the model's ability to recognize different scenarios, reduce the robot's collaborative error, and enable the model to adapt to multiple scenarios.
[0045] In some embodiments, the process of guiding the scene adaptation model to generate control instructions based on a preset scene-based reward function includes: The scene adaptation model outputs multiple candidate instructions in the target scene. The control command is obtained by selecting the optimal command from the candidate commands based on a preset scenario-based reward function.
[0046] In this embodiment, the scene adaptation model outputs multiple candidate instructions corresponding to the target scene. When the target scene is an industrial scene, the scene adaptation model can output the current workstation state, such as the 3D position of the workpiece and the distance to obstacles, to determine the mechanically feasible domains of each joint angle, such as the shoulder joint angle range and the technologically feasible domain of the torque threshold. The Latin hypercube sampling method is used to generate 3 to 5 sets of differentiated joint angle sequences within the joint angle feasible domain, and simultaneously, corresponding torque thresholds are generated within the torque threshold feasible domain, thus obtaining the corresponding candidate instructions. It is conceivable that this embodiment can also perform feasibility verification on the candidate instructions. Feasibility verification includes mechanical constraint verification and technological constraint verification. Mechanical constraint verification checks whether the joint angle sequence exceeds the robot's mechanical limits; if it does, the candidate instruction is excluded. Technological constraint verification checks whether the torque threshold exceeds the workpiece's load-bearing limit; if it does, the candidate instruction is excluded. Finally, valid candidate instructions that simultaneously satisfy both mechanical and technological constraints are retained.
[0047] For example, in a home setting, the feasible domains for action speed and speech speed are determined by outputting user features and home scene features through a scene adaptation model. Then, the target task, such as "taking a glass," is decomposed into joint action units such as "shoulder abduction → elbow flexion → hand grasping." By adjusting the execution order of these action units (e.g., abduction before flexion or flexion before abduction), and by adjusting the change in joint angle, a set of differentiated action sequences is generated. Within the feasible domains of speech speed and action speed, corresponding interaction parameters are generated using random sampling to obtain corresponding candidate instructions. This embodiment can also simulate the interaction effect of candidate instructions, calculate the intent matching degree to select candidate instructions, and retain valid candidate instructions that meet the intent matching degree criteria.
[0048] In some embodiments, the step of selecting the optimal control instruction from the candidate instructions based on a preset scenario-based reward function includes: Select the corresponding scenario-based reward function based on the target scenario; The candidate instruction is input into the task simulation environment, the state parameters after the candidate instruction is executed are collected, and the cumulative reward value corresponding to the candidate instruction is calculated by substituting them into the scenario-based reward function. The accumulated reward values are sorted, and the candidate instruction with the largest accumulated reward value is selected as the control instruction.
[0049] In this embodiment, a corresponding scenario-based reward function is selected based on different target scenarios. The scenario-based reward function includes an industrial scenario-based reward function and a household scenario-based reward function. The expression for the industrial scenario-based reward function is shown below: ; In the formula, This represents the reward value for industrial scenarios; Representing task-precision weights is a core requirement in industrial scenarios; Indicates the task completion rate; Indicates positional precision; Indicates flexible weights; Indicates force-controlled flexibility; Indicates safety weight; This indicates safety. Specifically, the industrial-scenario-based reward function emphasizes the humanoid robot's precision, flexibility, and safety in manipulating industrial parts.
[0050] The expression for the home-based reward function is as follows: ; In the formula, This represents the reward value for home-based scenarios; Indicates the intention to identify weights; Indicates the degree of intent matching; Indicates the comfort level weight; Indicates comfort level; Indicates the scene adaptation weight; This indicates scene adaptability. Among them, the industrial scene-specific reward function focuses on the humanoid robot's intent recognition, comfort, and scene adaptability.
[0051] In some embodiments, a control interface is generated based on the multimodal perception dataset, the task instructions, and the control instructions, and end-to-end control of the humanoid robot is performed based on the control interface, including: The multimodal perception dataset, the task instructions, and the control instructions are subjected to feature concatenation and encapsulation processing to obtain the control interface; Based on the control interface, joint drive signals are generated through the underlying controller of the humanoid robot; The humanoid robot is controlled end-to-end via the joint drive signals.
[0052] In this embodiment, a unified control interface is generated by encoding and splicing the multimodal perception dataset, task instructions, and control instructions. The control interface is then input into the robot's underlying controller to generate joint drive signals, and feedback data, such as actual joint angles and user voice feedback, is collected in real time to update the feature weights of the scene adaptation model.
[0053] This application can be specifically applied to industrial or home scenarios. For example, in an industrial setting, it can be used to assemble automotive bolts. Multimodal data is collected using LiDAR, RGB-D cameras, and torque sensors, and user-issued task commands are received. A humanoid robot performs feature extraction, noise processing, and fusion optimization on the data. Based on a scenario adaptation model, it outputs optimal control commands. These commands, along with the corresponding multimodal data and task commands, are fused to generate a unified control interface. This interface enables the robot's controller to generate corresponding drive signals for end-to-end control.
[0054] Please see Figure 3 This application also provides an end-to-end control system for a humanoid robot, which can implement the above-described method. The system includes: Data acquisition module 301 is used to acquire multimodal perception datasets and task instructions of the humanoid robot in the target scene; Model building module 302 is used to build a scene adaptation model based on the target scene and the multimodal perception dataset; The instruction generation module 303 is used to guide the scene adaptation model to generate control instructions based on a preset scene-based reward function. The end-to-end control module 304 is used to generate a control interface based on the multimodal perception dataset, the task instructions and the control instructions, and to perform end-to-end control of the humanoid robot based on the control interface.
[0055] It is understood that the content of the above method embodiments is applicable to this system embodiment. The specific functions implemented in this system embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.
[0056] This application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described method. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.
[0057] It is understood that the content of the above method embodiments is applicable to this device embodiment. The specific functions implemented by this device embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0058] Please see Figure 4 , Figure 4 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes: The processor 401 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application. The memory 402 can be implemented as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 402 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 402 and is called and executed by the processor 401 using the methods described in the embodiments of this application. Input / output interface 403 is used to implement information input and output; The communication interface 404 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.). Bus 405 transmits information between various components of the device (e.g., processor 401, memory 402, input / output interface 403, and communication interface 404); The processor 401, memory 402, input / output interface 403 and communication interface 404 are connected to each other within the device via bus 405.
[0059] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method.
[0060] It is understood that the content of the above method embodiments is applicable to this storage medium embodiment. The specific functions implemented in this storage medium embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.
[0061] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.
[0062] It is understood that the content of the above method embodiments is applicable to the embodiments of this program product. The specific functions implemented by the embodiments of this program product are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0063] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0064] This application provides an end-to-end control method, system, device, and product for a humanoid robot. This solution obtains a multi-joint dynamic model by performing dynamic modeling on the multi-joint dynamic characteristics of the humanoid robot. Based on this model, the robot's inertia, gravity, and other dynamic characteristics can be used as constraints to limit the robot's output, improving the accuracy of dynamic action generation. Furthermore, this solution acquires the robot's state, the operator's operating state, and environmental task information, and integrates these with the multi-joint dynamic model to construct a state parameter set. This set can fuse the robot's state, the operator's operating state, and task environment information to form a high-dimensional state vector, comprehensively describing the human-robot collaboration scenario and improving the data comprehensiveness of human-robot collaboration. Moreover, this solution constructs a multi-dimensional reward function for the humanoid robot based on the multi-joint dynamic model, which can improve the safety of human-robot collaboration, enhance interactive adaptability, and improve task completion accuracy. The solution also performs strategy optimization processing on the humanoid robot based on the multi-joint dynamic model and the multi-dimensional reward function, outputting dynamic game actions that conform to physical laws. This can shorten the humanoid robot's response time to operating actions and improve collaboration efficiency.
[0065] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.
[0066] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.
[0067] The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0068] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.
[0069] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0070] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0071] In the embodiments provided in this application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between systems or units may be electrical, mechanical, or other forms.
[0072] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0073] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0074] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0075] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.
Claims
1. An end-to-end control method for a humanoid robot, characterized in that, The method includes the following steps: Acquire multimodal perception datasets and task instructions for humanoid robots in target scenarios; A scene adaptation model is constructed based on the target scene and the multimodal perception dataset; The scene adaptation model is guided by a preset scene-based reward function to generate control commands. A control interface is generated based on the multimodal perception dataset, the task instructions, and the control instructions, and the humanoid robot is controlled end-to-end based on the control interface.
2. The method according to claim 1, characterized in that, The scene adaptation model constructed based on the target scene and the multimodal perception dataset includes: Based on the target scene, a set of scene modal features is extracted from the multimodal perception dataset; Noise optimization is performed on each feature in the scene modal feature set to obtain an optimized feature set; The fused features are obtained by fusing all features in the optimized feature set; Based on the target scene, the scene modal feature set is constrained to generate scene constraint terms; The scene adaptation model is obtained by training a preset reinforcement learning model based on the fusion features and the scene constraints.
3. The method according to claim 2, characterized in that, The process of generating scene constraint terms by constraining the scene modal feature set based on the target scene includes: When the target scenario is an industrial scenario, robot force perception features and workpiece position features are extracted from the scenario modal feature set; The workpiece position features are multiplied based on preset mapping coefficients, and the multiplication results are subtracted based on the robot force perception features to obtain the scene constraint terms.
4. The method according to claim 2, characterized in that, The step of generating scene constraint terms by constraining the scene modal feature set based on the target scene further includes: When the target scenario is a home scenario, target semantic features and target behavioral features are extracted from the scenario modal feature set; The scene constraint term is obtained by calculating the similarity between the target semantic features and the target behavioral features.
5. The method according to claim 1, characterized in that, The control instructions generated by the scenario adaptation model based on the preset scenario-based reward function include: The scene adaptation model outputs multiple candidate instructions in the target scene. The control command is obtained by selecting the optimal command from the candidate commands based on a preset scenario-based reward function.
6. The method according to claim 5, characterized in that, The process of selecting the optimal control instruction from the candidate instructions based on a preset scenario-based reward function includes: Select the corresponding scenario-based reward function based on the target scenario; The candidate instruction is input into the task simulation environment, the state parameters after the candidate instruction is executed are collected, and the cumulative reward value corresponding to the candidate instruction is calculated by substituting them into the scenario-based reward function. The accumulated reward values are sorted, and the candidate instruction with the largest accumulated reward value is selected as the control instruction.
7. The method according to any one of claims 1 to 6, characterized in that, The step of generating a control interface based on the multimodal perception dataset, the task instructions, and the control instructions, and performing end-to-end control of the humanoid robot based on the control interface, includes: The multimodal perception dataset, the task instructions, and the control instructions are subjected to feature concatenation and encapsulation processing to obtain the control interface; Based on the control interface, joint drive signals are generated through the underlying controller of the humanoid robot; The humanoid robot is controlled end-to-end via the joint drive signals.
8. An end-to-end control system for a humanoid robot, characterized in that, The system includes: The data acquisition module is used to acquire multimodal perception datasets and task instructions of the humanoid robot in the target scene; The model building module is used to build a scene adaptation model based on the target scene and the multimodal perception dataset; The instruction generation module is used to guide the scene adaptation model to generate control instructions based on a preset scene-based reward function. An end-to-end control module is used to generate a control interface based on the multimodal perception dataset, the task instructions, and the control instructions, and to perform end-to-end control of the humanoid robot based on the control interface.
9. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method according to any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 7.
Citation Information
Patent Citations
Robot control system and method, storage medium, controller and robot
CN118927246A
Multi-modal collaborative decision-making method and device for humanoid robot in industrial scene
CN120422253A
Medical robot control method and device, electronic equipment and storage medium
CN121043119A