Humanoid robot control method and related equipment

By acquiring sensor information from humanoid robots, fault identification and task reassignment are performed. A lightweight decision model is used to optimize multi-robot collaborative tasks, which solves the problem of insufficient adaptability of traditional humanoid robots in complex production tasks and improves the robustness of the production line and the flexibility of task execution.

CN121552328APending Publication Date: 2026-02-24广州里工实业有限公司
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511278061.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-09
Publication Date
2026-02-24

AI Technical Summary

Technical Problem

Traditional humanoid robots are not adaptable enough to complex production tasks involving multiple varieties and small batches. They suffer from limitations in data processing and low quality of motion samples, which affects the stability and flexibility of intelligent industrial production.

Method used

By acquiring sensor information from multiple humanoid robots, fault identification is performed, and a lightweight decision model is used to calculate task redistribution. By combining convolutional neural networks and attention networks to extract action features, dynamic scheduling and optimization of multi-robot collaborative tasks are achieved.

Benefits of technology

It improves the robustness of the production line and the flexibility of task execution, ensures the smooth and continuous operation of the system in the event of failure, and enhances the response speed and overall stability of the multi-robot system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121552328A_ABST
    Figure CN121552328A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a control method of a humanoid robot and related equipment, and belongs to the technical field of industrial automation control. The method comprises the following steps: acquiring sensor information corresponding to a plurality of human-shaped robots respectively; carrying out fault identification on the sensor information corresponding to the plurality of human-shaped robots; if it is determined that the first humanoid robot breaks down, task redistribution calculation is conducted according to sensor information corresponding to the humanoid robots through a lightweight decision model, and target task distribution information corresponding to all the humanoid robots is obtained; the first humanoid robot is any humanoid robot, and the target task allocation information comprises allocation task data, path planning data and target action data corresponding to each humanoid robot. By implementing the embodiment of the invention, the whole system can still operate stably and continuously when a part of units have faults, and the robustness of a production line and the flexibility of task execution are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of industrial automation control technology, and in particular to a control method and related equipment for a humanoid robot. Background Technology

[0002] In current industrial production environments, traditional humanoid robots typically employ fixed-path robotic arm structures, relying on pre-programmed procedures to complete repetitive tasks. They exhibit limitations in adaptability when faced with complex production tasks involving diverse products and small batches. As the manufacturing industry continues its transformation towards flexibility and intelligence, higher demands are placed on the responsiveness and task switching capabilities of production systems. Against this backdrop, humanoid robots, with their human-like morphology, highly biomimetic multi-degree-of-freedom movement capabilities, and excellent environmental interaction potential, are gradually demonstrating their unique value in replacing manual labor for high-risk, high-precision tasks, becoming one of the important technological pathways driving the upgrade of intelligent manufacturing. However, humanoid robots in industrial production face limitations in data processing and low-quality motion samples, which affect the stability and flexibility of intelligent industrial production. Summary of the Invention

[0003] The main objective of this application is to propose a control method and related equipment for a humanoid robot, which can improve the robustness of the production line and the flexibility of task execution.

[0004] To achieve the above objectives, one aspect of this application proposes a control method for a humanoid robot, applied to an electronic device. The electronic device is communicatively connected to multiple humanoid robots, each equipped with various sensors. The method includes the following steps:

[0005] Acquire sensor information corresponding to each of the multiple humanoid robots; the sensor information consists of data collected by various sensors installed on each of the humanoid robots.

[0006] Fault identification is performed on the sensor information corresponding to the multiple humanoid robots;

[0007] If the first humanoid robot is determined to be faulty, a lightweight decision model is used to perform task redistribution calculations based on the sensor information corresponding to each of the multiple humanoid robots, thereby obtaining the target task allocation information for each humanoid robot. The first humanoid robot can be any humanoid robot, and the target task allocation information includes the assigned task data, path planning data, and target action data corresponding to each humanoid robot.

[0008] In some embodiments, before acquiring the sensor information corresponding to the plurality of humanoid robots respectively, the method further includes:

[0009] Obtain a sample action dataset; the sample action dataset contains multiple sample action data.

[0010] The sample action dataset is input into the lightweight decision model to be trained. The lightweight decision model to be trained performs task redistribution calculation based on each sample action data to obtain the initial task allocation information corresponding to each sample action data.

[0011] The model parameters of the lightweight decision model to be trained are adjusted according to the initial task allocation information. If the difference between the initial task allocation information obtained by multiple iterations is less than the preset error threshold, the trained lightweight decision model is obtained.

[0012] In some embodiments, prior to obtaining the sample action dataset, the method further includes:

[0013] Acquire raw sensor information; the raw sensor information includes raw motion data;

[0014] The raw sensor information is fused and preprocessed to obtain standardized raw sensor information;

[0015] The original motion data is corrected based on the standardized original sensor information to obtain sample motion data.

[0016] In some embodiments, after obtaining the sample action dataset, the method further includes:

[0017] The multiple sample motion data are feature-extracted by convolutional neural networks and attention networks to obtain multiple sample motion features; the attention network assigns high weights to the motion features of key joints.

[0018] The scale difference is eliminated by processing the multiple sample action features to obtain multiple standardized sample action features;

[0019] The step involves inputting the sample action dataset into a lightweight decision model to be trained. The lightweight decision model then performs task redistribution calculations based on each sample action data point to obtain initial task allocation information corresponding to each sample action data point, including:

[0020] The standardized sample action features are input into the lightweight decision model to be trained. The lightweight decision model to be trained performs task redistribution calculations based on each sample action data to obtain the initial task allocation information corresponding to each sample action data.

[0021] In some embodiments, the step of performing task redistribution calculations based on the sensor information corresponding to the plurality of humanoid robots using a lightweight decision model to obtain target task allocation information for each humanoid robot includes:

[0022] The initial task allocation information for each humanoid robot is obtained by calculating based on the sensor information corresponding to each humanoid robot using a lightweight decision model.

[0023] The initial task allocation information corresponding to each humanoid robot is collaboratively calculated to obtain the target task allocation information.

[0024] In some embodiments, the sensor information includes real-time position data and raw motion data; based on the sensor information corresponding to each humanoid robot, initial task allocation information for each humanoid robot is calculated, including:

[0025] Based on the real-time location data corresponding to the first humanoid robot, the distance between the first humanoid robot and other humanoid robots is calculated to detect the second humanoid robot that is closest to the first humanoid robot.

[0026] The working status of the second humanoid robot is detected. If the working status of the second humanoid robot meets the working standards, the task allocation data corresponding to the first humanoid robot is used as the task allocation data corresponding to the second humanoid robot.

[0027] Based on the real-time location data corresponding to the second humanoid robot, the path planning data from the second humanoid robot to the first humanoid robot is calculated;

[0028] Based on the pre-stored motion allocation parameters in the lightweight decision model and the assigned task data corresponding to the second humanoid robot, the original motion data corresponding to the second robot is allocated to obtain the initial motion data.

[0029] In some embodiments, the step of collaboratively calculating the initial task allocation information corresponding to each humanoid robot to obtain target task allocation information includes:

[0030] According to formula R 协同 =ω1·(1-|e p |)+ω2·v s +ω3·(1-t / T)+ω4·(1-|e sync |)+ω5·(1-δ load )-ω6·c conflict The initial task allocation information corresponding to each humanoid robot is collaboratively calculated to obtain the target task allocation information, wherein e pv is the normalized value of the end effector position error. s Here, t is the motion smoothness coefficient, t is the actual operation time, T is the standard operation time threshold, and e is the motion smoothness coefficient. sync δ represents the normalized value of the motion synchronization error for each humanoid robot. load Let c be the load difference coefficient for each humanoid robot. conflict ω1, ω2, ω3, ω4, ω5, and ω6 are the task conflict coefficients, and ω1+ω2+ω3+ω4+ω5+ω6=1.

[0031] To achieve the above objectives, another aspect of this application proposes a control device for a humanoid robot, applied to an electronic device. The electronic device is communicatively connected to multiple humanoid robots, each equipped with various sensors. The device includes:

[0032] The information acquisition module is used to acquire sensor information corresponding to the multiple humanoid robots; the sensor information is data collected by various sensors installed on each of the humanoid robots.

[0033] The fault identification module is used to identify faults in the sensor information corresponding to the multiple humanoid robots.

[0034] The task allocation module is used to calculate the task redistribution based on the sensor information corresponding to each of the multiple humanoid robots by using a lightweight decision model if the first humanoid robot is determined to be faulty, so as to obtain the target task allocation information corresponding to each humanoid robot. The first humanoid robot is any humanoid robot, and the target task allocation information includes the allocated task data, path planning data and target action data corresponding to each humanoid robot.

[0035] To achieve the above objectives, another aspect of this application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the method described above.

[0036] To achieve the above objectives, another aspect of the embodiments of this application proposes a computer-readable storage medium storing a computer program that, when executed by a processor, implements the methods described above.

[0037] To achieve the above objectives, another aspect of this application provides a computer program product, including a computer program that, when executed by a processor, implements the methods described above.

[0038] The embodiments of this application include at least the following beneficial effects: This application provides a control method, device, electronic device, storage medium, and program product for a humanoid robot. This solution acquires sensor information corresponding to multiple humanoid robots; the sensor information consists of data collected by various sensors set on each humanoid robot; fault identification is performed on the sensor information corresponding to the multiple humanoid robots; if a fault is determined in the first humanoid robot, a lightweight decision model is used to perform task redistribution calculation based on the sensor information corresponding to the multiple humanoid robots to obtain the target task allocation information corresponding to each humanoid robot; the first humanoid robot is any humanoid robot, and the target task allocation information includes the allocated task data, path planning data, and target action data corresponding to each humanoid robot. By implementing the embodiments of this application, the operating status of each robot is monitored, and fault identification is performed based on the sensor information. Once a fault is detected in a humanoid robot, the lightweight decision model is immediately activated. Based on the real-time sensor data of all humanoid robots, the production tasks are quickly redistributed. Moreover, the lightweight decision model is a computationally efficient and responsive intelligent algorithm that can achieve dynamic scheduling and optimization of multi-robot collaborative tasks in a limited computing environment. This ensures that the overall system can still operate smoothly and continuously when some units fail, thereby improving the robustness of the production line and the flexibility of task execution. Attached Figure Description

[0039] Figure 1 This is a schematic diagram of an implementation environment provided in an embodiment of this application;

[0040] Figure 2 This is a flowchart of the control method for a humanoid robot provided in an embodiment of this application;

[0041] Figure 3 This is a timing diagram of multi-machine collaborative operation in one embodiment;

[0042] Figure 4 This is a flowchart of sensor information fusion in one embodiment;

[0043] Figure 5 This is a flowchart of training a lightweight decision model in one embodiment;

[0044] Figure 6 This is a schematic diagram of the structure of the control device for the humanoid robot provided in the embodiments of this application;

[0045] Figure 7 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0046] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit it. In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with those of this application; they are merely examples of apparatuses and methods consistent with some aspects of the embodiments of this application as detailed in the appended claims.

[0047] It is understood that the terms “first,” “second,” etc., used in this application may be used herein to describe various concepts, but unless otherwise stated, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the words “if,” “when,” or “in response to a determination” as used herein may be interpreted as “when…” or “when…” or “in response to a determination.”

[0048] As used in this application, the terms "at least one", "multiple", "each", "any", etc., "at least one" includes one, two or more, "multiple" includes two or more, "each" refers to each of the corresponding multiples, and "any" refers to any one of the multiples.

[0049] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0050] In current industrial production, most traditional humanoid robots still employ fixed-path robotic arm structures, relying on pre-programmed instructions for operation. This results in significant limitations in adaptability to unstructured environments or changing process requirements. As the manufacturing industry continues its transformation towards flexibility and intelligence, humanoid robots, with their humanoid form and multi-degree-of-freedom movement capabilities, are increasingly seen as a key vehicle for replacing manual labor in high-risk, high-precision, and complex tasks. However, these robots still face numerous technical challenges in practical industrial applications.

[0051] First, the multi-source sensor data (visual, force, tactile, etc.) generated in industrial settings are often heterogeneous in format and difficult to process uniformly. The lack of efficient data fusion mechanisms limits the robot's overall environmental perception. Second, motion generation and annotation still heavily rely on manual teaching, which is not only inefficient but also susceptible to the influence of human experience on annotation quality, hindering the construction and application of standardized motion sample libraries. Third, existing control models are mostly trained based on specific scenarios. When faced with changes in workpiece specifications or ambient lighting, their generalization ability is significantly insufficient, often requiring frequent manual intervention and parameter readjustment. Furthermore, in multi-robot collaborative operations, there are generally problems such as unreasonable task allocation strategies and large system communication delays, which can easily lead to motion conflicts or a decrease in overall operational efficiency. Finally, industrial scenarios have extremely high requirements for system real-time performance and reliability, but existing model deployment architectures often cannot guarantee millisecond-level response times and lack efficient fault recovery mechanisms, severely restricting the robustness of practical applications. Therefore, there is an urgent need to build an integrated system that can integrate multi-source sensing data fusion, intelligent annotation and training, and edge-side collaborative deployment to comprehensively improve the autonomous operation capability and collaborative control level of humanoid robots in complex industrial scenarios.

[0052] In view of this, this application provides a control method for a humanoid robot, which improves the robustness of the production line and the flexibility of task execution.

[0053] The humanoid robot control method provided in this application relates to the field of industrial automation control. The humanoid robot control method provided in this application can be applied to a terminal, a server, or software running on a terminal or server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, or vehicle terminal, but is not limited to these. The server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network. The software can be an application that implements the humanoid robot control method, but is not limited to the above forms.

[0054] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics devices, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0055] It should be noted that in all specific embodiments of this application, when processing data related to user identity or characteristics, such as user information, user behavior data, user historical data, and user location information, user permission or consent is obtained first. Furthermore, the collection, use, and processing of this data comply with relevant laws, regulations, and standards. In addition, when embodiments of this application require access to sensitive personal information of users, separate permission or consent from the user is obtained through pop-ups or redirection to confirmation pages. Only after obtaining the user's separate permission or consent is the necessary user-related data required for the proper functioning of these embodiments acquired.

[0056] like Figure 1 The diagram shown is a schematic representation of an implementation environment provided in an embodiment of this application. (Refer to...) Figure 1 The environment includes electronic equipment 101 and multiple humanoid robots 102. The electronic equipment is communicatively connected to the humanoid robots, and each humanoid robot is equipped with various sensors, such as position sensors, motion sensors, and temperature sensors, which are not limited here. The position sensors are used to acquire real-time position data of the humanoid robots, the motion sensors are used to acquire raw motion data of the humanoid robots, and the temperature sensors are used to acquire temperature data of the environment in which the humanoid robots are located.

[0057] Electronic device 101 can acquire sensor information corresponding to multiple humanoid robots 102 respectively; the sensor information is data collected by various sensors set on each humanoid robot 102; fault identification is performed on the sensor information corresponding to multiple humanoid robots 102 respectively; if the first humanoid robot is determined to be faulty, the lightweight decision model is used to perform task redistribution calculation based on the sensor information corresponding to multiple humanoid robots to obtain the target task allocation information corresponding to each humanoid robot 102; the first humanoid robot is any humanoid robot, and the target task allocation information includes the allocated task data, path planning data and target action data corresponding to each humanoid robot 102.

[0058] Figure 2 This is a flowchart of a control method for a humanoid robot provided in an embodiment of this application. The execution subject of this control method can be any of the aforementioned electronic devices (including servers or terminals). Figure 2 The method may include, but is not limited to, steps S201 to S203.

[0059] Step S201: Obtain sensor information corresponding to multiple humanoid robots; the sensor information consists of data collected by various sensors set on each humanoid robot.

[0060] In some embodiments, real-time data collected by various sensors on multiple humanoid robots, such as force, vision, position, and environmental sensors, are acquired to construct a complete sensing information system. This relies on multi-source sensor data fusion technology, which involves the unified acquisition, alignment, and integration of heterogeneous sensor signals from different physical quantities, sampling frequencies, and formats to form a comprehensive perception of the environment and the robot's own state.

[0061] Step S202: Fault identification is performed on the sensor information corresponding to the multiple humanoid robots.

[0062] Step S203: If the first humanoid robot is determined to be faulty, the lightweight decision model is used to perform task redistribution calculation based on the sensor information corresponding to each of the multiple humanoid robots to obtain the target task allocation information corresponding to each humanoid robot; the first humanoid robot is any humanoid robot, and the target task allocation information includes the allocated task data, path planning data and target action data corresponding to each humanoid robot.

[0063] In some embodiments, before acquiring the sensor information corresponding to the plurality of humanoid robots, a sample action dataset can be acquired; the sample action dataset contains multiple sample action data; the sample action dataset is input into a lightweight decision model to be trained, and the lightweight decision model to be trained performs task redistribution calculation based on each sample action data to obtain initial task allocation information corresponding to each sample action data; the model parameters of the lightweight decision model to be trained are adjusted according to each initial task allocation information; if the difference between the initial task allocation information obtained from multiple iterations is less than a preset error threshold, then the trained lightweight decision model is obtained.

[0064] A sample dataset containing various typical actions is collected, consisting of a series of pre-recorded or generated sample action data. This dataset is then input into a lightweight decision model to be trained. This lightweight decision model, a high-efficiency computing architecture designed specifically for resource-constrained industrial environments, can perform task redistribution simulations based on the input sample action sequences and output corresponding initial task allocation information. By comparing the initial allocation information generated through multiple iterations and adjusting the model's internal parameters accordingly, the model reaches a stable convergence state and training is complete when the difference between the iteration results is below a preset error tolerance. This enables the lightweight decision model to quickly and reliably achieve dynamic task allocation and coordination among multiple robots in practical applications, significantly improving the system's response speed and overall stability under complex working conditions.

[0065] In some embodiments, raw sensor information is acquired; the raw sensor information includes raw motion data; the raw sensor information is fused and preprocessed to obtain standardized raw sensor information; the raw motion data is corrected based on the standardized raw sensor information to obtain sample motion data.

[0066] Raw information from multiple sensors is collected, including raw motion data of the humanoid robot. Then, using multi-source sensor information fusion technology and data preprocessing methods, the raw information undergoes time synchronization, noise filtering, and format standardization to generate a standardized sensor data stream. Based on this, the raw motion data is compared and corrected using the sensor fusion results, for example, by using motion reconstruction or state estimation methods to eliminate biases and outliers, ultimately generating high-quality, highly consistent sample motion data.

[0067] Furthermore, multimodal data from vision sensors, six-axis force sensors, tactile arrays, and environmental sensors can be accessed via industrial bus and wireless sensor network. The original sensor information can be deduplicated and standardized in format. Based on timestamp alignment and spatial coordinate system transformation, the data from different sensors can be fused into unified three-dimensional work scene data. Synthetic data under extreme working conditions can be generated through a physical simulation engine and fused with real data to expand sample diversity, thereby obtaining standardized original sensor information.

[0068] At the same time, the collected classified industrial data can be anonymized, and data tampering logs can be recorded through blockchain technology to ensure data integrity. Furthermore, based on a pre-trained motion recognition model, robot joint angles and end effector poses in standardized raw sensor information can be automatically labeled to generate coarse-labeled trajectories. These coarse-labeled trajectories are then distributed to engineers for accuracy correction. A 3D visualization interface is used to compare label differences and confirm consensus. High-quality samples are selected based on label consistency indicators to construct sample motion data containing work scenario labels.

[0069] By fusing and preprocessing the raw sensor information to obtain standardized raw sensor information, and then correcting the raw motion data based on the standardized raw sensor information to obtain sample motion data, the reliability and usability of the data can be effectively improved, providing an accurate and stable input source for subsequent model training and decision-making, thereby enhancing the system's perception and control capabilities in actual industrial scenarios.

[0070] In some embodiments, features are extracted from multiple sample action data using a convolutional neural network and an attention network to obtain multiple sample action features; the attention network assigns high weights to the action features of key joints; the multiple sample action features are processed to eliminate scale differences to obtain multiple standardized sample action features; the multiple standardized sample action features are input into a lightweight decision model to be trained, and the lightweight decision model to be trained performs task redistribution calculations based on each of the sample action data to obtain the initial task allocation information corresponding to each of the sample action data.

[0071] By employing convolutional neural networks (CNNs) and attention networks, feature extraction is performed on multiple sample action data to obtain sample action features rich in spatiotemporal information. Specifically, the CNN extracts local and global motion patterns from the action sequence layer by layer, while the attention network automatically focuses on key joint movements that significantly impact the task by weighting the importance of different joints. Subsequently, these features are scale-normalized to eliminate differences introduced by variations in sensor units, dimensions, or motion amplitude, thus obtaining standardized sample action features.

[0072] Furthermore, motion trajectory data for single robots and multi-robot collaboration, such as joint angle sequences, end effector pose changes, and operation timestamps, are extracted from the motion sample library and standardized into fixed-length temporal vectors, such as 1000 frames / action. A temporal convolutional network is used to convolve these vectors to capture local temporal features of the motion, such as the speed change trend during welding. Simultaneously, a spatial attention mechanism is introduced to assign higher weights to the motion features of key joints, strengthening the features of parts that significantly affect operational accuracy. Key joints can be the elbows or wrists of a humanoid robot's robotic arm, and the specific joints can be adjusted according to actual conditions. The extracted spatiotemporal features are then processed using Layer Normalization to eliminate scale differences between different motion samples, resulting in a uniform-dimensional motion feature vector.

[0073] The feature vectors from the action sample library are used as pre-training data to initialize the base model of the adaptive training module, such as the feature extraction layer of Humanoid-CLIP. Domain adversarial networks are used to reduce the difference in feature distribution between general scenarios and industrial scenarios, enabling the model to quickly adapt to industrial action patterns. At the same time, typical cooperative actions in the action sample library (such as the division of labor sequence in multi-machine assembly) are used as "expert demonstrations" for reinforcement learning and introduced in the early stage of PPO algorithm (Proximal Policy Optimization) training to accelerate the model's learning of cooperative rules. For example, by imitating "conflict-free path planning" samples in the sample library, invalid trial and error in the exploration phase is reduced.

[0074] By using convolutional neural networks and attention networks to extract features from multiple sample action data, multiple sample action features are obtained. The attention network assigns high weights to the action features of key joints and performs scale difference elimination processing on multiple sample action features to obtain multiple standardized sample action features. This not only enhances the ability to represent multi-source heterogeneous action data, but also improves the generalization and stability of the task allocation model in real industrial scenarios.

[0075] As an optional implementation, a lightweight decision-making model calculates initial task allocation information for each humanoid robot based on sensor information. This initial task allocation information is then collaboratively calculated to obtain target task allocation information. Real-time processing and analysis of multi-source sensor information from multiple humanoid robots generates preliminary initial task allocation information for each robot. Furthermore, a multi-agent collaborative optimization mechanism is introduced. By coordinating and integrating the initial allocation schemes, and comprehensively balancing factors such as efficiency, conflict avoidance, and overall operational consistency, the target task allocation information is finally generated.

[0076] The lightweight decision model can optimize the generation of action decision strategies based on the industrial operation reward function of the Markov decision process, combined with the PPO near-end policy optimization algorithm, and compress the initial control model into a lightweight action decision model adapted to edge computing through knowledge distillation and quantization techniques, ensuring that the inference latency is less than 50ms.

[0077] The PPO algorithm is an optimization tool for multi-robot collaborative operations. Its core principle is to integrate multi-robot collaborative constraints into the single-robot operational metrics. During training, the lightweight decision model simultaneously learns the single-robot motion accuracy and multi-robot collaborative rules. Furthermore, by quantifying the time difference and spatial overlap of multi-robot actions, it avoids collaborative failures caused by motion misalignment. When multi-robot operational paths intersect or task priorities conflict, a penalty mechanism reduces the probability of conflict; for example, the penalty coefficient increases by 0.2 for every 10% increase in the overlapping area. Based on the real-time load of each humanoid robot, additional rewards are given for states where the load difference is less than 10%, improving overall collaborative efficiency. The real-time load can include CPU utilization, remaining battery power, etc.

[0078] By using a lightweight decision model to calculate the initial task allocation information for each humanoid robot based on the sensor information of each robot, the initial task allocation information for each humanoid robot is obtained through collaborative calculation. This fully leverages the real-time reasoning capability of the lightweight model in the edge computing environment. At the same time, collaborative calculation improves the overall collaboration level and task execution efficiency of the multi-robot system, significantly enhancing the adaptability and stability of the intelligent manufacturing system under dynamic working conditions.

[0079] Optionally, based on the real-time position data of the first humanoid robot, the distance between the first humanoid robot and other humanoid robots is calculated to detect the second humanoid robot that is closest to the first humanoid robot; the working status of the second humanoid robot is detected, and if the working status of the second humanoid robot meets the working standards, the task allocation data corresponding to the first humanoid robot is used as the task allocation data corresponding to the second humanoid robot; based on the real-time position data of the second humanoid robot, path planning data from the second humanoid robot to the first humanoid robot is calculated; based on the pre-stored motion allocation parameters in the lightweight decision model and the task allocation data corresponding to the second humanoid robot, the original motion data corresponding to the second robot is allocated to obtain initial motion data.

[0080] Based on the real-time position data of the first humanoid robot, a spatial distance calculation algorithm is used to calculate its relative distance to other robots in real time, thereby identifying the closest second humanoid robot. Subsequently, the working status of the second humanoid robot is detected. If it is determined to be in a working state that meets the standards, such as idle, ready, or low load, the task data originally assigned to the first robot is dynamically assigned to the second robot. Further, based on the real-time position of the second robot, an algorithm based on graph search or real-time motion planning is used to generate optimal path planning data from its current position to the first robot. Using pre-set motion allocation parameters in a lightweight decision model (which can be obtained by encoding and optimizing typical task actions), combined with the newly assigned task content, the original motion sequence of the second robot is reconstructed and adjusted in real time to generate initial motion data that can be executed immediately. This effectively improves the task takeover and collaborative adaptability of the multi-robot system in dynamic environments, enhancing the overall system's robustness and response efficiency.

[0081] For example, suppose the force sensor of humanoid robot A malfunctions, detecting welding pressure fluctuations > ±2N, exceeding the threshold three times consecutively. Assume humanoid robot B is closest to humanoid robot A's work area (< 5m). The fault self-healing unit of the edge-deployed collaborative module triggers an alarm. Based on real-time position data, the lightweight decision model outputs "Robot B pauses welding in area 2 and takes over operation in area 1." It generates an obstacle avoidance path for humanoid robot B from area 2 to area 1, planning a speed of 0.5m / s to ensure arrival within 10 seconds. For the welding parameters in area 1, the lightweight decision model calls the "alternative robot welding adaptation action" feature from the action sample library, outputting an adjusted joint control sequence. This could be an increase of 10% in elbow joint movement amplitude to compensate for positional deviations. Welding parameters could be, for example, a current of 200A and a trajectory radius of 5mm. If the planned path of humanoid robot B passes through area 1 where humanoid robot C is located, then humanoid robot C can be controlled to extend its welding time in area 3 by 10 seconds, waiting for humanoid robot B to complete its work in area 1 before performing an overall weld inspection, thus avoiding process disconnect.

[0082] Furthermore, it can be determined according to formula R 协同 =ω1·(1-|e p |)+ω2·v s +ω3·(1-t / T)+ω4·(1-|e sync |)+ω5·(1-δ load )-ω6·c conflict The initial task allocation information for each humanoid robot is collaboratively calculated to obtain the target task allocation information, where e p v is the normalized value of the end effector position error. sHere, t is the motion smoothness coefficient, t is the actual operation time, T is the standard operation time threshold, and e is the motion smoothness coefficient. sync δ represents the normalized value of the motion synchronization error for each humanoid robot. load Let c be the load difference coefficient for each humanoid robot. conflict Let ω1, ω2, ω3, ω4, ω5, and ω6 be the task conflict coefficients, and let ω1 + ω2 + ω3 + ω4 + ω5 + ω6 = 1.

[0083] Among them, ω4 is the synchronization weight, ω5 is the load weight, and ω6 is the conflict penalty weight, which can be dynamically adjusted according to the scenario. For example, ω4 = 0.2 in high-precision assembly scenario and ω5 = 0.2 in large-scale sorting scenario.

[0084] In some embodiments, the operational errors of the humanoid robot are monitored in real time, and the PID parameters and learning rate of the control model are dynamically adjusted based on a Bayesian optimization algorithm. Furthermore, multi-dimensional operational errors are collected and quantified, including the end effector position accuracy error e. p Force control error e f motion smoothness error e s and multi-machine collaborative synchronization error e sync Calculate the overall error E = ω p ·e p +ω f ·e f +ω s ·e s +ω sync ·e sync (ω p ,ω f ,ω s ,ω syn2 (where K is the weight coefficient and its sum is 1); define the parameter space to be optimized (PID parameter K). p ,K i ,K d Based on Gaussian process regression, a parameter-error probabilistic model is constructed using the model learning rate η. The optimal parameter combination is selected through the expected improvement function. Error evaluation and parameter iteration are performed every 100ms, with the single parameter adjustment range limited to ±20% to ensure the stability of the adjustment. The adjusted parameters are written to the runtime memory of the lightweight action decision model in real time.

[0085] In some embodiments, a lightweight decision model is encapsulated as a container image and deployed to the robot's local computing unit through an edge computing framework. Based on an improved Hungarian algorithm, dynamic allocation of multi-robot tasks is achieved. Task priorities are adjusted in conjunction with real-time load feedback. Through heartbeat detection and model version backtracking mechanisms, task reassignment is triggered and the system automatically switches to a backup model when the humanoid robot experiences computing overload or decision abnormalities.

[0086] Furthermore, the triggering condition for task reassignment is that when the end effector position error is detected to be >5mm, CPU load >90%, force control deviation >20%, or work progress deviation >20% for three consecutive sampling cycles, and the overall error decreases by <10% after dynamic parameter adjustment, task reassignment is triggered. The robot's state parameters (load, health indicators), task characteristic parameters (type, priority, spatial coordinates), and environmental cooperation parameters (interference intensity, historical success rate of cooperation) are aggregated in real time, and the original task is split into independent sub-tasks based on the features of the action sample library. This is achieved through a multi-objective optimization formula S... i,j =α·R i,j +β·(1-D i,j / D max )+γ·(1-C i )+δ·P j ·(1-ΔT j Calculate the fit score between the humanoid robot and the sub-task.

[0087] Among them, R i,j D represents the historical success rate of humanoid robot i performing subtask j. i,j D is the actual distance from humanoid robot i to the work area of ​​subtask j. max C represents the maximum distance from all humanoid robots to the work area of ​​this subtask. i P represents the real-time CPU load rate of the humanoid robot i. j Priority of subtask j, ranging from 1 to 5, with 1 being the highest, preset by the MES system or task scheduling center. ΔT j Let be the schedule deviation rate for subtask j. The formula for calculating the schedule deviation rate is: Schedule Deviation Rate = |Actual Progress - Planned Progress| / Planned Progress. α, β, γ, and δ are weighting coefficients, summing to 1. They can be dynamically adjusted. For example, in high-precision task scenarios, α can be increased to 0.4 to prioritize historical success rates and balance the impact of different parameters on adaptability.

[0088] Furthermore, initial matching is achieved based on the Hungarian algorithm, prioritizing assignment to humanoid robots with higher sensor configuration matching degrees. The fitness scores S for all subtasks are then calculated. i,j Organized into a "Robot-Subtask" score matrix (as shown in the table below, where rows = robots and columns = subtasks):

[0089]

[0090] To align with the objective of "maximizing fit score", the score matrix is ​​converted into a cost matrix, using the formula Cost. i,j =KS i,jK represents the maximum score. If K = 1, the problem of "finding the maximum fit" is transformed into "finding the minimum cost". Taking the table above as an example, the cost matrix is ​​as follows:

[0091]

[0092] By marking rows and columns as "matched," the minimum-cost matching solution is found: For each row (robot), mark the minimum-cost element and cross out the corresponding column (subtask); for unmatched rows / columns, construct augmenting paths by covering all 0 elements with lines; iteratively adjust the matching relationships until all robots correspond one-to-one with subtasks. The final matching result directly corresponds to the optimal "subtask-robot" allocation, as shown in the table below:

[0093] 1. Subtask 1 (Welding) → Robot A (Cost 0.15, original score 0.85 highest) 2. Subtask 2 (Assembly) → Robot B (Cost 0.1, original score 0.9 highest) 3. Subtask 3 (Transportation) → Robot C (Cost 0.2, original score 0.8 highest)

[0094] In some embodiments, the "alternative execution" feature of the motion sample library can be invoked to generate obstacle avoidance paths and joint compensation control sequences. The "alternative execution" feature is a collection of historical success cases stored in the motion sample library, including: joint motion compensation parameters when different robots perform the same task (e.g., when robot B replaces robot A to perform a welding task, the wrist joint angle needs to be compensated by +5°); obstacle avoidance path planning templates when switching tasks (e.g., when switching from "transportation path" to "assembly path", the coordinate sequence of the area X to be bypassed); and timing compensation rules in multi-robot collaborative scenarios (e.g., when robot C replaces faulty robot D, it needs to start with a 0.5-second delay to match the rhythm of other robots).

[0095] When task reassignment is triggered (e.g., humanoid robot A malfunctions and humanoid robot B needs to replace it to perform the welding task), the following related logic is executed: Search the action sample library for keywords: task type (e.g., "welding") + original performing robot (e.g., "robot A") + replacement robot (e.g., "robot B"), and match the "replacement execution" feature template.

[0096] Extract joint angle compensation values ​​(such as Δθ) from the template. 腕关节 = +5°), combined with the current joint state of humanoid robot B (such as the current wrist joint angle θ). 当前 =30°), calculate the target angle:

[0097] θ 目标 =θ 当前 +Δθ 腕关节 =35°

[0098] Extract the path detour rules from the template (such as "avoid area X (coordinates 10, 20-15, 25)"), combine them with the current position of humanoid robot B (such as map road network coordinates (5, 10)), and plan a new path in real time using the A* algorithm: new path = current position → detour point (8, 18) → work area (12, 22).

[0099] Extract the collaborative timing rules from the template (such as "delay 0.5 seconds to start"), combine them with the original timing of the multi-machine collaborative task (such as robot C should start at T=10 seconds), and adjust the start time of the replacement humanoid robot B: new start time = T+0.5 seconds = 10.5 seconds.

[0100] After reallocation, the error is continuously monitored. If it does not improve, the dynamic parameter tuning unit is triggered to perform secondary optimization of the PID parameters, or a secondary reallocation is initiated. Details are shown in the table below.

[0101]

[0102] Humanoid robot A is performing a welding task when a positional error e occurs due to a malfunction in the force sensor. p =6mm (first-time error), the e p The formula for representing the position accuracy error of the end effector is:

[0103]

[0104] The x 实际 y 实际 z 实际 The three-dimensional coordinates of the end effector are collected in real time by the visual sensors (such as 3D cameras) or laser positioning system on the robot.

[0105] The x 目标 y 目标 z 目标 It is the theoretical target coordinates (such as weld center, bolt hole position) planned by the lightweight motion decision model based on the task.

[0106] Dynamic parameter tuning unit adjusts position loop K p From 1.0 to 1.2, the error decreased to e. p =4mm (a 33% decrease, effective), continue with the original task. If parameter adjustment is ineffective, and the error remains e after adjustment... p =5.5mm, a decrease of <5% can trigger task reassignment. Humanoid robot B takes over the execution, invoking the "alternate execution" feature, controlling the wrist joint angle +5°, and after adjustment, the error is reduced to e. p =3mm, which meets the accuracy requirements. If the problem persists after replacement, such as the error e during execution by humanoid robot B. p =4.5mm, then the secondary redistribution is activated for robot C, and the dynamic parameter tuning unit customizes the parameters for robot C (e.g., K). p =1.3), and the error eventually decreased to e. p =2mm. Figure 3 This is a timing diagram of multi-machine collaborative operation in one embodiment, such as... Figure 3As shown, this illustrates the task allocation and action synchronization time points for three humanoid robots.

[0107] Figure 4 A flowchart of sensor information fusion in one embodiment, such as Figure 4 As shown, the system starts with the multi-source data acquisition and processing module, which is responsible for collecting and processing multi-source data. The data then flows to the action annotation and review module for action annotation and review to ensure data accuracy. Next, the data is passed to the adaptive training module, where algorithms and models are used for analysis and learning to improve system performance. Finally, the trained lightweight decision model is deployed to the light blue edge deployment and collaboration module to realize practical applications. At the same time, the collected feedback information is returned to the action annotation and review module through the feedback optimization path, forming a closed-loop optimization mechanism to continuously improve system performance.

[0108] Figure 5 A flowchart of training a lightweight decision model in one embodiment, such as Figure 5 As shown, the humanoid robot's movements are first simulated in a virtual simulation environment. The humanoid robot performs movements according to the set joint angles and end-effector poses. Then, the state monitoring module tracks key indicators such as position error, movement smoothness, and operation time in real time. These data are input into the Reward function for comprehensive evaluation, calculating a comprehensive score of accuracy, efficiency, and safety. Next, the PPO algorithm adjusts the joint control force parameters based on the Reward value to optimize the strategy. At the same time, the system continuously iterates and optimizes through two feedback paths: direct feedback of state monitoring data and feedback of Reward value, forming a complete closed loop from movement execution to monitoring and evaluation to parameter adjustment. Ultimately, this achieves continuous performance improvement of the robot's movements in the virtual environment.

[0109] In this embodiment, sensor information corresponding to multiple humanoid robots is acquired. The sensor information consists of data collected by various sensors on each humanoid robot. Fault identification is performed on the sensor information corresponding to each of the multiple humanoid robots. If a fault is determined in the first humanoid robot, a lightweight decision model is used to perform task redistribution calculations based on the sensor information corresponding to each of the multiple humanoid robots, obtaining target task allocation information for each humanoid robot. The first humanoid robot can be any humanoid robot, and the target task allocation information includes the assigned task data, path planning data, and target action data for each humanoid robot. The operating status of each robot is monitored, and fault identification is performed based on this sensor information. Once a fault is detected in a humanoid robot, the lightweight decision model is immediately activated. Based on the real-time sensor data of all humanoid robots, a rapid redistribution calculation of production tasks is performed. This lightweight decision model is a computationally efficient and responsive intelligent algorithm that can achieve dynamic scheduling and optimization of multi-robot collaborative tasks under limited computing power, thereby ensuring that the overall system can still operate smoothly and continuously when some units fail, improving the robustness of the production line and the flexibility of task execution.

[0110] Please see Figure 6 This application also provides a control device for a humanoid robot that can implement the above-described method. The device includes:

[0111] The information acquisition module 601 is used to acquire sensor information corresponding to multiple humanoid robots; the sensor information is data collected by various sensors installed on each of the humanoid robots.

[0112] Fault identification module 602 is used to identify faults in the sensor information corresponding to multiple humanoid robots.

[0113] The task allocation module 603 is used to calculate the task redistribution based on the sensor information corresponding to each of the multiple humanoid robots by using a lightweight decision model if the first humanoid robot is determined to be faulty, so as to obtain the target task allocation information corresponding to each humanoid robot. The first humanoid robot is any humanoid robot, and the target task allocation information includes the allocated task data, path planning data and target action data corresponding to each humanoid robot.

[0114] It is understood that the content of the above method embodiments is applicable to the present device embodiments. The specific functions implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0115] This application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described method. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.

[0116] It is understood that the content of the above method embodiments is applicable to this device embodiment. The specific functions implemented by this device embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0117] Please see Figure 7 , Figure 7 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes:

[0118] The processor 701 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application.

[0119] The memory 702 can be implemented as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 702 can store the operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 702 and is called and executed by the processor 701 using the methods described in the embodiments of this application.

[0120] The input / output interface 703 is used to implement information input and output;

[0121] The communication interface 704 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0122] Bus 705 transmits information between various components of the device (e.g., processor 701, memory 702, input / output interface 703, and communication interface 704);

[0123] The processor 701, memory 702, input / output interface 703, and communication interface 704 are connected to each other within the device via bus 705.

[0124] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method.

[0125] It is understood that the content of the above method embodiments is applicable to this storage medium embodiment. The specific functions implemented in this storage medium embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.

[0126] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.

[0127] It is understood that the content of the above method embodiments is applicable to the embodiments of this program product. The specific functions implemented by the embodiments of this program product are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0128] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0129] The control method, apparatus, electronic device, storage medium, and program product for humanoid robots provided in this application improve the robustness of the production line and the flexibility of task execution.

[0130] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.

[0131] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.

[0132] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0133] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.

[0134] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0135] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0136] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0137] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0138] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0139] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0140] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.

Claims

1. A control method for a humanoid robot, characterized in that, The method, applied to an electronic device that is communicatively connected to multiple humanoid robots, wherein each humanoid robot is equipped with various sensors, includes the following steps: Acquire sensor information corresponding to each of the multiple humanoid robots; the sensor information consists of data collected by various sensors installed on each of the humanoid robots. Fault identification is performed on the sensor information corresponding to the multiple humanoid robots; If the first humanoid robot is determined to be faulty, a lightweight decision model is used to perform task redistribution calculations based on the sensor information corresponding to each of the multiple humanoid robots, thereby obtaining the target task allocation information for each humanoid robot. The first humanoid robot can be any humanoid robot, and the target task allocation information includes the assigned task data, path planning data, and target action data corresponding to each humanoid robot.

2. The method according to claim 1, characterized in that, Before acquiring the sensor information corresponding to the plurality of humanoid robots, the method further includes: Obtain a sample action dataset; the sample action dataset contains multiple sample action data. The sample action dataset is input into the lightweight decision model to be trained. The lightweight decision model to be trained performs task redistribution calculation based on each sample action data to obtain the initial task allocation information corresponding to each sample action data. The model parameters of the lightweight decision model to be trained are adjusted according to the initial task allocation information. If the difference between the initial task allocation information obtained by multiple iterations is less than the preset error threshold, the trained lightweight decision model is obtained.

3. The method according to claim 2, characterized in that, Before obtaining the sample action dataset, the method further includes: Acquire raw sensor information; the raw sensor information includes raw motion data; The raw sensor information is fused and preprocessed to obtain standardized raw sensor information; The original motion data is corrected based on the standardized original sensor information to obtain sample motion data.

4. The method according to claim 2, characterized in that, After obtaining the sample action dataset, the method further includes: The multiple sample motion data are feature-extracted by convolutional neural networks and attention networks to obtain multiple sample motion features; the attention network assigns high weights to the motion features of key joints. The scale difference is eliminated by processing the multiple sample action features to obtain multiple standardized sample action features; The step involves inputting the sample action dataset into a lightweight decision model to be trained. The lightweight decision model then performs task redistribution calculations based on each sample action data point to obtain initial task allocation information corresponding to each sample action data point, including: The standardized sample action features are input into the lightweight decision model to be trained. The lightweight decision model to be trained performs task redistribution calculations based on each sample action data to obtain the initial task allocation information corresponding to each sample action data.

5. The method according to claim 1, characterized in that, The step involves using a lightweight decision-making model to perform task redistribution calculations based on the sensor information corresponding to each of the multiple humanoid robots, thereby obtaining target task allocation information for each humanoid robot, including: The initial task allocation information for each humanoid robot is obtained by calculating based on the sensor information corresponding to each humanoid robot using a lightweight decision model. The initial task allocation information corresponding to each humanoid robot is collaboratively calculated to obtain the target task allocation information.

6. The method according to claim 5, characterized in that, The sensor information includes real-time position data and raw motion data; based on the sensor information corresponding to each humanoid robot, initial task allocation information for each humanoid robot is calculated, including: Based on the real-time location data corresponding to the first humanoid robot, the distance between the first humanoid robot and other humanoid robots is calculated to detect the second humanoid robot that is closest to the first humanoid robot; The working status of the second humanoid robot is detected. If the working status of the second humanoid robot meets the working standards, the task allocation data corresponding to the first humanoid robot is used as the task allocation data corresponding to the second humanoid robot. Based on the real-time location data corresponding to the second humanoid robot, the path planning data from the second humanoid robot to the first humanoid robot is calculated; Based on the pre-stored motion allocation parameters in the lightweight decision model and the assigned task data corresponding to the second humanoid robot, the original motion data corresponding to the second robot is allocated to obtain the initial motion data.

7. The method according to claim 5, characterized in that, The step of collaboratively calculating the initial task allocation information corresponding to each humanoid robot to obtain the target task allocation information includes: According to formula R 协同 =ω1·(1-|e p |)+ω2·v s +ω3·(1-t / T)+ω4·(1-|e sync |)+ω5·(1-δ load )-ω6·c conflict The initial task allocation information corresponding to each humanoid robot is collaboratively calculated to obtain the target task allocation information, wherein e p v is the normalized value of the end effector position error. s Here, t is the motion smoothness coefficient, t is the actual operation time, T is the standard operation time threshold, and e is the motion smoothness coefficient. sync δ represents the normalized value of the motion synchronization error for each humanoid robot. load Let c be the load difference coefficient for each humanoid robot. conflict ω1, ω2, ω3, ω4, ω5, and ω6 are the task conflict coefficients, and ω1+ω2+ω3+ω4+ω5+ω6=1.

8. A control device for a humanoid robot, characterized in that, Applied to an electronic device, the electronic device is communicatively connected to multiple humanoid robots, each of which is equipped with various sensors; the device includes: The information acquisition module is used to acquire sensor information corresponding to the multiple humanoid robots; the sensor information is data collected by various sensors installed on each of the humanoid robots. The fault identification module is used to identify faults in the sensor information corresponding to the multiple humanoid robots. The task allocation module is used to calculate the task redistribution based on the sensor information corresponding to each of the multiple humanoid robots by using a lightweight decision model if the first humanoid robot is determined to be faulty, so as to obtain the target task allocation information corresponding to each humanoid robot. The first humanoid robot is any humanoid robot, and the target task allocation information includes the allocated task data, path planning data and target action data corresponding to each humanoid robot.

9. A computer device, characterized in that, The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 7.

Citation Information

Patent Citations

  • Robot control method and system, electronic equipment and storage medium

    CN118952221A

  • Multi-modal sensing humanoid robot action self-adaptive control method and multi-modal sensing humanoid robot action self-adaptive control system

    CN119610112A

  • Factory environment operation system based on humanoid robot, humanoid robot and method

    CN119820584A

  • Power industry robot collaborative inspection and fault self-diagnosis system and method

    CN120233762A

  • Electronic apparatus for allocating task to multiple robots and operation method thereof

    KR102508952B1