Humanoid robot real-time cooperation decision-making method based on multi-modal perception fusion
By building a hierarchical multimodal perception fusion architecture and real-time decision-making system, the problem of low human-computer collaboration in complex scenarios is solved, and the secure collaboration of multi-dimensional information collaborative reasoning and instant response is realized, the human-computer collaboration efficiency is improved, and the evolution of human-computer collaboration to group intelligence is promoted.
Patent Information
- Application Number
- CN202510925069.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-04
- Publication Date
- 2025-08-29
AI Technical Summary
In the prior art, humanoid robots have problems such as difficulty in fusion of multimodal information, lagging real-time response, insufficient security control and lack of group collaboration capabilities in the field of human-computer collaboration, especially in complex scenarios, human-computer collaboration has low efficiency.
Build a hierarchical multimodal perception fusion architecture, adopt dynamic deployment of sensor arrays, physiological signal fusion processing and multimodal feature enhancement, use a cross-modal semantic inference system supported by large models, and combines a distributed incremental learning mechanism with double buffering and coordination to realize multi-dimensional information collaborative reasoning and real-time collaborative decision-making.
It realizes multi-dimensional information collaborative reasoning in complex scenarios, instant response in dynamic environments, and accurate control of secure collaboration, improves human-machine collaboration efficiency, supports multi-robot collaborative operations, and promotes human-machine collaboration to group intelligence.
Smart Images

Figure CN120552071A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of robotics, and in particular to a real-time collaborative decision-making method for humanoid robots based on multimodal perception fusion. Background Art
[0002] In the existing technology, humanoid robots in the field of human-machine collaboration mainly rely on single-modal perception systems (such as vision or force control sensors) and discrete decision-making architectures, which lead to problems such as difficulty in multimodal information fusion, delayed real-time response, insufficient safety control, and lack of group collaboration capabilities. Although the latest research has proposed a multimodal fusion method and online learning algorithm based on Transformer, there are still technical bottlenecks: traditional fusion models cannot achieve dynamic alignment of spatiotemporal features, and the accuracy of intention understanding in complex scenarios is limited; cloud-based decision-making architectures have network delays and are difficult to meet the needs of instant collaboration; the latest force control system has large contact force control errors, and the emergency braking performance has not yet reached safety standards; the emotional computing model only supports two-dimensional analysis and cannot fuse physiological signals to achieve cognitive load assessment. This patent is aimed at the industrial manufacturing field. By constructing a hierarchical multimodal perception fusion architecture and a real-time decision-making system, especially relying on an adaptive multi-dimensional collaborative perception architecture, a cross-modal semantic reasoning system supported by a large model, and a double-buffered collaborative mechanism, it breaks through the existing technical bottlenecks and realizes multi-dimensional information collaborative reasoning, instant response to dynamic environments, and precise control of safe collaboration.
[0003] To address the above-mentioned problems, no effective solutions have been proposed so far. Summary of the Invention
[0004] An embodiment of the present invention provides a real-time collaborative decision-making method for a humanoid robot based on multimodal perception fusion, so as to at least solve the technical problem of low human-machine collaboration efficiency of humanoid robots in complex scenarios.
[0005] According to one aspect of an embodiment of the present invention, a real-time collaborative decision-making method for a humanoid robot based on multimodal perception fusion is provided, comprising: dynamically deploying a sensor array, collecting multimodal raw perception data of the humanoid robot through the deployed sensor array, and performing physiological signal fusion processing and multimodal feature enhancement processing on the multimodal raw perception data to obtain a multi-level feature set; utilizing a cross-modal semantic reasoning system supported by a large model to perform cross-modal semantic mapping and intention reasoning on the multi-level feature set to obtain a collaborative execution instruction containing a target action and force control parameters; based on the collaborative execution instruction, utilizing a distributed incremental learning mechanism with double buffering collaboration to dynamically update the behavioral decision model of the humanoid robot, and based on the behavioral decision model, obtain the real-time collaborative decision of the humanoid robot.
[0006] According to another aspect of an embodiment of the present invention, a real-time collaborative decision-making device for a humanoid robot based on multimodal perception fusion is provided, comprising: a deployment module configured to dynamically deploy a sensor array, collect multimodal original perception data of the humanoid robot through the deployed sensor array, and perform physiological signal fusion processing and multimodal feature enhancement processing on the multimodal original perception data to obtain a multi-level feature set; an inference module configured to utilize a cross-modal semantic reasoning system supported by a large model to perform cross-modal semantic mapping and intention reasoning on the multi-level feature set to obtain a collaborative execution instruction containing target actions and force control parameters; a decision module configured to dynamically update the behavior decision model of the humanoid robot based on the collaborative execution instruction and utilize a distributed incremental learning mechanism with double buffering collaboration, and obtain a real-time collaborative decision of the humanoid robot based on the behavior decision model.
[0007] In an embodiment of the present invention, a sensor array is dynamically deployed to collect the multimodal raw perception data of the humanoid robot. The multimodal raw perception data is then subjected to physiological signal fusion processing and multimodal feature enhancement processing to obtain a multi-level feature set. A cross-modal semantic reasoning system supported by a large model is used to perform cross-modal semantic mapping and intention reasoning on the multi-level feature set to obtain collaborative execution instructions containing target actions and force control parameters. Based on the collaborative execution instructions, a distributed incremental learning mechanism with double buffering collaboration is used to dynamically update the humanoid robot's behavioral decision model, and based on the behavioral decision model, the humanoid robot's real-time collaborative decision is obtained. The above solution solves the technical problem of low human-machine collaboration efficiency of humanoid robots in complex scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0009] Figure 1 A method for real-time collaborative decision-making of a humanoid robot based on multimodal perception fusion according to an embodiment of the present invention;
[0010] Figure 2A Another method of a humanoid robot real-time collaborative decision-making method based on multimodal perception fusion according to an embodiment of the present invention;
[0011] Figure 2B is a structural diagram of a humanoid robot real-time collaborative decision-making system based on multimodal perception fusion according to an embodiment of the present invention;
[0012] Figure 3is a flowchart of an optional method for collecting multimodal raw perception data according to an embodiment of the present invention;
[0013] Figure 4 is a flowchart of an optional intention reasoning method according to an embodiment of the present invention;
[0014] Figure 5 is a flow chart of an optional distributed incremental learning method according to an embodiment of the present invention;
[0015] Figure 6 1 is a schematic structural diagram of an optional humanoid robot real-time collaborative decision-making device based on multimodal perception fusion according to an embodiment of the present invention;
[0016] Figure 7 A schematic structural diagram of an electronic device suitable for implementing the embodiments of the present disclosure is shown. DETAILED DESCRIPTION
[0017] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0018] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0019] According to an embodiment of the present invention, a method embodiment of a humanoid robot real-time collaborative decision-making method based on multimodal perception fusion is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0020] Figure 1is a method of a humanoid robot real-time collaborative decision-making method based on multimodal perception fusion according to an embodiment of the present invention, such as Figure 1 As shown, the method includes the following steps:
[0021] Step S102 : dynamically deploying a sensor array, collecting multimodal raw perception data of the humanoid robot through the deployed sensor array, and performing physiological signal fusion processing and multimodal feature enhancement processing on the multimodal raw perception data to obtain a multi-level feature set.
[0022] First, dynamically deploy the sensor array. For example, a distributed sensor node architecture is adopted to integrate multiple sensors at each node, wherein the multiple sensors include at least one of the following: a 3D structured light camera, a flexible tactile sensor, and a micro inertial measurement unit; a sensor angle adjustment algorithm based on particle swarm optimization is used to adjust the angle of one or more of the multiple sensors. In this embodiment, by constructing a distributed sensor node architecture, 3D structured light cameras, flexible tactile sensors, and micro inertial measurement units are integrated into each node to achieve all-round collection of multimodal data such as human motion trajectory, environmental semantics, and physical interaction forces. Using a sensor angle adjustment algorithm based on particle swarm optimization (PSO), the sensor's perception angle is dynamically optimized, effectively solving the problem of perception blind spots, improving the comprehensiveness and accuracy of data collection, and providing a solid data foundation for subsequent multimodal fusion.
[0023] Next, physiological signal fusion processing is performed. For example, a dry electrode array is used to collect the user's EEG and EMG mixed signal; the EEG and EMG mixed signal is blindly separated using an independent component analysis algorithm to obtain a plurality of independent component signals; the plurality of independent component signals are subjected to spatiotemporal alignment processing to obtain a physiological fusion signal. In some embodiments, blind source separation processing is performed, including: whitening the EEG and EMG mixed signal to obtain a whitening matrix; initializing a separation matrix based on the whitening matrix, and iteratively updating the separation matrix to obtain a plurality of independent component signals. In some embodiments, spatiotemporal alignment processing is performed, including: calculating the delay value between the plurality of independent component signals based on a cross-correlation function, and performing time synchronization on the plurality of independent component signals based on the delay value.
[0024] In step S104, a cross-modal semantic reasoning system supported by a large model is used to perform cross-modal semantic mapping and intention reasoning on the multi-level feature set to obtain a collaborative execution instruction including a target action and force control parameters.
[0025] First, semantic mapping is performed. Based on a three-level mapping network, a hierarchical parser is used to perform cross-modal semantic mapping on the multi-level feature set.
[0026] Next, online reasoning is performed. Based on the cross-modal semantic mapping, online intent reasoning is performed to obtain a resource allocation strategy.
[0027] Finally, closed-loop correction is performed. Tactile signals are detected for anomalies. When the tactile signal exceeds a set threshold, visual re-identification is triggered to obtain the tracked target. Based on the tracked target, a reinforcement learning mechanism is used to optimize the force control parameters of the resource allocation strategy. This optimization includes constructing a state space, an action space based on the force control gain coefficient, and a reward function based on force error and position error. The state space includes the current force value, position deviation, and contact time. A proximal policy optimization algorithm is then used for optimization training to optimize the force control parameters.
[0028] This embodiment builds a multimodal reasoning engine with a large language model (LLM) at its core, mapping multi-dimensional information such as vision, language, force, and touch into a unified semantic space. By designing encoding and projection processes specific to vision, language, and touch, the fusion of multimodal information is achieved. A multimodal command parsing system is developed, leveraging syntax analysis, intent inference, and parameter generation to enable robots to accurately understand and execute complex commands.
[0029] Step S106, based on the collaborative execution instruction, using the distributed incremental learning mechanism of double buffer collaboration, dynamically update the behavior decision model of the humanoid robot, and obtain the real-time collaborative decision of the humanoid robot based on the behavior decision model.
[0030] Performing knowledge memory management on the behavior decision model, constructing a dual-buffer storage structure including a recent memory buffer and a long-term memory buffer to store recent data and historical samples respectively, and dynamically adjusting the forgetting gating parameters of the behavior decision model through the Lagrangian optimization method to balance new knowledge learning and old knowledge retention;
[0031] Performing distributed learning on the behavior decision model, wherein the distributed learning includes: using federated learning to encrypt and aggregate model parameters uploaded by multiple robots, and using transfer learning to download a pre-trained model from a group knowledge base and perform local fine-tuning; constructing a reward function based on individual rewards and collaborative rewards, and optimizing the reward alignment strategy of the behavior decision model;
[0032] The behavioral decision model is continuously evaluated for performance, and the prediction variance and the lower bound of evidence are calculated to form a comprehensive confidence. When the comprehensive confidence drops below a set threshold for three consecutive cycles, the behavioral decision model is triggered to roll back and the model with the best performance in the most recent version is restored.
[0033] This embodiment employs a dual-buffered storage structure, storing recent data and historical representative samples separately. Sample priority is determined based on predicted entropy. A forgetfulness control algorithm based on Lagrangian optimization is used to dynamically update model parameters. A system integrating federated learning and transfer learning is built, enabling multi-robot collaborative learning and optimization. A reward alignment algorithm is also employed to enhance collaboration among multiple robots.
[0034] Figure 2A Another method of a humanoid robot real-time collaborative decision-making method based on multimodal perception fusion according to an embodiment of the present invention is applied in Figure 2B As shown in the overall system architecture, Figure 2A As shown, the method includes the following steps:
[0035] Step S202: Collect multimodal raw perception data using an adaptive multi-dimensional collaborative perception architecture.
[0036] like Figure 3 As shown, the method for collecting multimodal raw perception data includes the following steps:
[0037] Step S2022: dynamic sensor array deployment.
[0038] A distributed sensor node architecture is adopted, with each node integrating a 3D structured light camera, a flexible tactile sensor and a micro inertial measurement unit; the 3D structured light camera is used to collect three-dimensional spatial information of the environment, with a resolution of 1280×720 pixels and a frame rate of 30fps; the flexible tactile sensor adopts the piezoresistive sensing principle, with a spatial resolution of 10×10 array, each sensor unit is 2mm×2mm in size, and has a sampling frequency of 1kHz; the micro inertial measurement unit includes a three-axis accelerometer and a three-axis gyroscope, with an accelerometer range of ±16g, a gyroscope range of ±2000° / s, and a sampling frequency of 100Hz.
[0039] The sensor angle is adjusted using a sensor angle adjustment algorithm based on particle swarm optimization (PSO).
[0040] Specifically, first, the particle swarm position is initialized, and N particles (N=50) are randomly generated within the adjustable angle range of the sensor (horizontally 0°-180°, vertically -90°-90°), and the position of each particle represents the horizontal and vertical angles of the sensor.
[0041] Next, the fitness value of each particle is calculated (based on the coverage of the perception blind area). The fitness function is:
[0042]
[0043] Among them, BlindAreaCoverage is the area of the perception blind area, and TotalArea is the total area of the target perception area.
[0044] Then, update the particle speed and position. The speed update formula is:
[0045]
[0046] The position update formula is:
[0047]
[0048] in, is the d-th dimension velocity of the i-th particle in the k-th iteration, is the d-dimensional position of the i-th particle in the k-th iteration, ω is the inertia weight (initial value 0.7, linearly decreasing to 0.4), c1, c2 are learning factors (both 2.0), r1, r2 are random numbers in the interval [0, 1], is the individual optimal position of the i-th particle, is the global optimal position.
[0049] Finally, repeat the iteration until convergence (maximum number of iterations is 100 or the fitness value changes less than 1e-4)
[0050] Step S2024: perform physiological signal fusion processing.
[0051] Use a dry electrode array to collect EEG / EMG signals: Design a dedicated signal preprocessing circuit. The dry electrodes are made of Ag / AgCl material, have an impedance of <50 kΩ, and are sampled at a frequency of 1024 Hz. There are 16 channels in total, arranged according to the international 10-20 system. Use a fusion algorithm, such as a blind source separation algorithm based on independent component analysis (ICA), to whiten the mixed signal. For example, the whitening formula can be:
[0052] Z=WX,
[0053] Where W is the whitening matrix. After that, the separation matrix is initialized and iteratively updated to output the independent component signal. Then, the time-space alignment algorithm is used to synchronize the signals based on the cross-correlation function. The synchronization formula is:
[0054]
[0055] Where x(t) and y(t) are the two signals to be synchronized, τ is the time delay, and N is the signal length. The delay is determined by τ opt =argmax τ CCF(τ).
[0056] Step S2026: multimodal feature enhancement.
[0057] First, a spatiotemporal attention pyramid structure is constructed: the bottom layer is CNN to extract local features, using a ResNet-18 network to output a 512-dimensional feature vector; the middle layer is a spatiotemporal Transformer to capture temporal dependencies, consisting of 6 layers of Transformer encoders, each with 8 attention heads, a hidden layer dimension of 1024, and a feedforward neural network dimension of 4096; the top layer is global feature aggregation, and 256 global features are obtained through global average pooling and a fully connected layer.
[0058] Next, adversarial enhancement training is performed, using a conditional generative adversarial network (cGAN) to enhance low-quality tactile signals: the generator uses a U-Net structure, inputting a noise vector + tactile signal label; the discriminator uses a three-layer CNN to distinguish between real and generated signals; finally, the training objective function is: L = L adv +λ·L L1 , where L adv represents the adversarial loss, usually from the feedback of the discriminator, L L1 represents the L1 loss, and λ represents the balance coefficient, which is used to control the relative importance of the two losses.
[0059] Step S204: perform intention reasoning using a cross-modal semantic reasoning system supported by a large model.
[0060] like Figure 4 As shown, the method of intention reasoning includes the following steps:
[0061] Step S2042: cross-modal semantic mapping.
[0062] First, a three-level mapping network is used for visual encoding, language encoding, and tactile encoding. Visual encoding: ResNet-50 extracts visual features (2048 dimensions) → Transformer encoding (1024 dimensions); Language encoding: LLaMA-2 processes text input (768 dimensions) → Linear projection (1024 dimensions); Tactile encoding: Tactile-BERT extracts tactile features (768 dimensions) → Linear projection (1024 dimensions).
[0063] Next, a hierarchical parser is used to parse multimodal instructions. 1) Syntax analysis: Construct an abstract syntax tree using a top-down recursive descent analysis method. 2) Intention reasoning: Intention resolution based on a knowledge graph. The knowledge graph contains entities (such as objects, actions, and environments) and relationships (such as "has property," "is a," and "causes"). Graph neural networks (GNNs) are used for reasoning, with a node feature dimension of 1024 and a layer count of 3. 3) Parameter generation: Generate an execution plan containing force control parameters. The force control parameter calculation formula is F = k·F desired +b·v, where Fdesired is the expected force, v is the running speed, k and b are gain coefficients (optimized by reinforcement learning).
[0064] Step S2044: online inference optimization.
[0065] First, model distillation compression is performed, and knowledge distillation technology is used to compress the multimodal large model. Among them, the teacher model is the original multimodal large model, which contains a 12-layer Transformer encoder, a hidden layer dimension of 1024, and 8 attention heads; the student model is a lightweight model, which contains a 6-layer Transformer encoder, a hidden layer dimension of 512, and 4 attention heads.
[0066] Among them, the distillation loss function is L distill =α·L KL +(1-α)·L MSE , the KL divergence loss is Where T is the temperature parameter; mean square error loss σ is the weight coefficient, α is the fusion weight coefficient of distillation loss, KL is the divergence calculation function, logits student is the student model log odds, logits teacher is the log-odds of the teacher model.
[0067] Then, dynamic resource allocation is performed, and the computational power scheduling algorithm based on task complexity is used to evaluate the complexity and calculate the characteristic entropy value (H) of the current task. where p i is the probability distribution of task categories, and n is the total number of task categories.
[0068] The resource allocation strategy is as follows: when H < 0.5, the student model is used; when 0.5 ≤ H < 0.8, hybrid model inference is used (the student and teacher models run simultaneously, and the results are weighted averaged); when H ≥ 0.8, the teacher model is activated. Switching latency is less than 50ms, based on caching the intermediate results of the last 100 inferences on edge devices.
[0069] Step S2046: perform closed-loop correction.
[0070] First, tactile re-identification is triggered, and a visual re-identification module triggered by tactile feedback is built. Tactile signal anomaly detection (threshold ±3σ) is performed; visual re-identification (YOLOv8 real-time detection) is performed; and target tracking (DeepSORT algorithm) is performed. The motion model uses Kalman filtering, and the state vector includes position, scale, and velocity.
[0071] Next, the force control parameters are optimized, for example, using a force control parameter adjustment algorithm based on reinforcement learning, where: 1) State space: current force value, position deviation, contact time, the formula is S = (fcurrent ,e position ,t contact ); 2) Action space: A = k gain Force control gain coefficient k gain Range 0.1-5.0); 3) Reward function: -|force error|-0.5×|position error|; 4) Training algorithm: PPO (Proximal Policy Optimization), using Generalized Advantage Estimation (GAE), γ is the discount factor, and λ is the GAE parameter.
[0072] Step S206: utilizing a distributed incremental learning mechanism with double buffering collaboration.
[0073] like Figure 5 As shown, the distributed incremental learning method includes the following steps:
[0074] Step S2062, knowledge memory management.
[0075] A double buffer storage structure is used, for example, a double buffer storage based on priority, wherein the recent memory buffer (capacity 2000 samples) stores the data of the last 5 hours; the long-term memory buffer (capacity 3000 samples) stores historical typical samples. After that, priority calculation is performed based on the predicted entropy value (H = -∑p i logp i ), where p i Represents the probability distribution of the i-th class sample in the task.
[0076] Adopt dynamic forget gating, for example, the forget control algorithm based on Lagrangian optimization, where the objective function is minL(α)=L new +λ·(R old -0.98) 2 , L new is the new knowledge learning loss, R old is the old knowledge retention rate, λ is the Lagrange multiplier; the constraint condition is R old ≥98%, L new ≤0.1, where R old is the old knowledge retention rate, L new Learn the loss for the new knowledge. Finally, optimize the variable forget gate parameter α (0≤α≤1).
[0077] Step S2064: distributed learning strategy.
[0078] First, a hybrid learning architecture is adopted. A system combining federated learning and transfer learning is constructed. The federated learning layer allows multiple robots to upload model parameters (encrypted aggregation). The transfer learning layer downloads pre-trained models from a swarm knowledge base and uses domain adaptation techniques (such as adversarial domain adaptation) to adjust the models to the local scenario. Local fine-tuning is then performed, for example, by performing incremental learning based on the current scenario data using online gradient descent.
[0079] Afterwards, we used reward function alignment to construct a reward alignment algorithm for multi-robot collaboration, where individual rewards are based on local task completion (0-1); collaborative rewards are based on group knowledge gain (0-0.5); total reward = individual reward + collaborative reward; and the final reward alignment effect is a reward standard deviation from 0.3 to 0.15.
[0080] Step S2066: Continue performance evaluation.
[0081] First, a multi-dimensional evaluation system is designed using confidence evaluation indicators, where the prediction variance (σ 2 ) reflects the model uncertainty, where f k is the k-th prediction result, is the predicted mean, K is the number of predictions. The evidence lower bound (ELBO) is used to quantify the degree of data fit; the overall confidence level is 0.6σ 2 +0.4ELBO (weights are dynamically adjusted based on the task type), where ELBO represents the lower bound of evidence and is used to quantify how well the model fits the data.
[0082] Next, we implemented an early warning and rollback mechanism. We built a model degradation detection system that checks whether the overall confidence level of the metrics has dropped by >15% for three consecutive cycles. If so, a rollback is triggered, restoring the optimal model from the three most recent versions (based on model snapshots). The rollback response time is <30 seconds (relying on an incremental storage mechanism).
[0083] This application addresses the following technical issues: 1) Multimodal perception fusion: Overcoming the information limitations of traditional single-modal systems to achieve collaborative reasoning and efficient fusion of multi-dimensional data such as image, semantics, and force perception; 2) Real-time collaborative decision latency: Through a low-latency control architecture and online incremental learning algorithms, the decision-making response cycle in dynamic environments is significantly shortened. 3) Deepening the level of cognitive interaction: Integrating vision-language models with physiological signal analysis to achieve a leap from functional instruction reception to cognitive intent understanding.
[0084] This application significantly improves the human-machine collaboration efficiency of humanoid robots in complex scenarios by constructing a hierarchical multimodal perception fusion architecture and real-time decision-making system. The innovative design of spatiotemporal feature alignment and cross-modal attention mechanism breaks through the information limitations of traditional single-modal systems, enabling robots to integrate multi-dimensional data such as images, semantics, and force perception for collaborative reasoning, and achieve accurate understanding and execution of abstract instructions. Based on the online incremental learning algorithm, it supports robots to continuously optimize decision-making models in dynamic environments, significantly shortens the scene adaptation cycle, and enhances the task generalization ability in unknown environments. Combining real-time operating systems with nanosecond interrupt response technology, the system decision delay is reduced, ensuring the immediacy and smoothness of human-machine collaboration. The multimodal reasoning large model and swarm intelligence architecture introduced in this application realize high-dimensional decision-making and distributed learning for multi-robot collaborative operations, and promote the evolution of human-machine collaboration from single-machine autonomy to swarm intelligence. Through the deep integration of visual-language models and tactile feedback technology, robots can accurately identify object attributes (such as material and shape) and dynamically adjust clamping strategies. These technological breakthroughs have enabled humanoid robots to demonstrate stronger environmental adaptability and intelligent interaction in various fields, providing key technical support for the realization of "AI physicalization" and "embodied intelligence", and promoting the industry's paradigm shift from preset instruction execution to autonomous decision-making and collaboration.
[0085] This application also provides a humanoid robot real-time collaborative decision-making device based on multimodal perception fusion, such as Figure 6 As shown, it includes: a deployment module 62, which is configured to dynamically deploy a sensor array, collect multimodal raw perception data of the humanoid robot through the deployed sensor array, and perform physiological signal fusion processing and multimodal feature enhancement processing on the multimodal raw perception data to obtain a multi-level feature set; an inference module 64, which is configured to use a cross-modal semantic reasoning system supported by a large model to perform cross-modal semantic mapping and intention reasoning on the multi-level feature set to obtain a collaborative execution instruction containing target actions and force control parameters; a decision module 66, which is configured to dynamically update the behavior decision model of the humanoid robot based on the collaborative execution instruction and utilize a distributed incremental learning mechanism with double buffering collaboration, and obtain real-time collaborative decision of the humanoid robot based on the behavior decision model.
[0086] It should be noted that the humanoid robot real-time collaborative decision-making device based on multimodal perception fusion provided in the above embodiment is only illustrated by the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the humanoid robot real-time collaborative decision-making device based on multimodal perception fusion provided in the above embodiment and the humanoid robot real-time collaborative decision-making method embodiment based on multimodal perception fusion belong to the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.
[0087] Figure 7 Schematic diagram of the structure of an electronic device suitable for implementing the embodiment of the present disclosure is shown. Figure 7 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0088] like Figure 7 As shown, the electronic device includes a central processing unit (CPU) 1001, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1002 or the program loaded from the storage part 1008 into the random access memory (RAM) 1003. Various programs and data required for system operation are also stored in the RAM 1003. The CPU 1001, ROM 1002 and RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.
[0089] The following components are connected to the I / O interface 1005: an input section 1006 including a keyboard, a mouse, and the like; an output section 1007 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), and a speaker; a storage section 1008 including a hard disk; and a communication section 1009 including a network interface card such as a LAN card or a modem. The communication section 1009 performs communication processing via a network such as the Internet. A drive 1010 is also connected to the I / O interface 1005 as needed. A removable medium 1011, such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory, is installed in the drive 1010 as needed, so that computer programs read therefrom can be installed into the storage section 1008 as needed.
[0090] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A real-time collaborative decision-making method for humanoid robots based on multimodal perception fusion, characterized in that: include: Dynamically deploying a sensor array to collect multimodal raw perception data of the humanoid robot through the deployed sensor array, and performing physiological signal fusion processing and multimodal feature enhancement processing on the multimodal raw perception data to obtain a multi-level feature set; Using a cross-modal semantic reasoning system supported by a large model, cross-modal semantic mapping and intention reasoning are performed on the multi-level feature set to obtain collaborative execution instructions including target actions and force control parameters; Based on the collaborative execution instruction, a distributed incremental learning mechanism with double buffering collaboration is used to dynamically update the behavior decision model of the humanoid robot, and based on the behavior decision model, a real-time collaborative decision of the humanoid robot is obtained.
2. The method according to claim 1, characterized in that Dynamically deploy sensor arrays, including: A distributed sensor node architecture is adopted, and multiple sensors are integrated in each node, wherein the multiple sensors include at least one of the following: a 3D structured light camera, a flexible tactile sensor, and a micro inertial measurement unit; A sensor angle adjustment algorithm based on particle swarm optimization is used to adjust the angles of one or more sensors.
3. The method according to claim 1, characterized in that Perform physiological signal fusion processing, including: A dry electrode array is used to collect the user's EEG and EMG mixed signals; Performing blind source separation on the EEG and EMG mixed signals using an independent component analysis algorithm to obtain multiple independent component signals; A plurality of the independent component signals are subjected to spatiotemporal alignment processing to obtain a physiological fusion signal.
4. The method according to claim 3, characterized in that Performing blind source separation processing, including: whitening the EEG and EMG mixed signal to obtain a whitening matrix; initializing a separation matrix based on the whitening matrix, and iteratively updating the separation matrix to obtain multiple independent component signals; and / or Performing spatiotemporal alignment processing includes: calculating delay values between the plurality of independent component signals based on a cross-correlation function, and performing time synchronization on the plurality of independent component signals based on the delay values.
5. The method according to claim 1, wherein Using a cross-modal semantic reasoning system supported by a large model, cross-modal semantic mapping and intention reasoning are performed on the multi-level feature set to obtain collaborative execution instructions containing target actions and force control parameters, including: Based on a three-level mapping network, a hierarchical parser is used to perform cross-modal semantic mapping on the multi-level feature set; Based on the cross-modal semantic mapping, online intention reasoning is performed to obtain a resource allocation strategy; The resource allocation strategy is modified in a closed loop to obtain the collaborative execution instruction including the target action and force control parameters.
6. The method according to claim 5, characterized in that Performing closed-loop correction on the resource allocation strategy to obtain the collaborative execution instruction including the target action and force control parameters, including: Perform anomaly detection on the tactile signal. When the tactile signal exceeds a set threshold range, visual re-identification is triggered to obtain the tracked target; Based on the tracking target, the force control parameters of the resource allocation strategy are optimized using a reinforcement learning mechanism, wherein the optimization includes: constructing a state space, an action space based on a force control gain coefficient, and a reward function based on force error and position error, wherein the state space includes the current force value, position deviation, and contact time; and using a proximal policy optimization algorithm for optimization training to optimize the force control parameters.
7. The method according to claim 1, characterized in that The humanoid robot's behavior decision model is dynamically updated using a distributed incremental learning mechanism with double buffering collaboration, including: Performing knowledge memory management on the behavior decision model, constructing a dual-buffer storage structure including a recent memory buffer and a long-term memory buffer to store recent data and historical samples respectively, and dynamically adjusting the forgetting gating parameters of the behavior decision model through the Lagrangian optimization method to balance new knowledge learning and old knowledge retention; Performing distributed learning on the behavior decision model, wherein the distributed learning includes: using federated learning to encrypt and aggregate model parameters uploaded by multiple robots, and using transfer learning to download a pre-trained model from a group knowledge base and perform local fine-tuning; constructing a reward function based on individual rewards and collaborative rewards, and optimizing the reward alignment strategy of the behavior decision model; The behavioral decision model is continuously evaluated for performance, and the prediction variance and the lower bound of evidence are calculated to form a comprehensive confidence. When the comprehensive confidence drops below a set threshold for three consecutive cycles, the behavioral decision model is triggered to roll back and the model with the best performance in the most recent version is restored.
8. A humanoid robot real-time collaborative decision-making device based on multimodal perception fusion, characterized in that: include: a deployment module configured to dynamically deploy a sensor array, collect multimodal raw perception data of the humanoid robot through the deployed sensor array, and perform physiological signal fusion processing and multimodal feature enhancement processing on the multimodal raw perception data to obtain a multi-level feature set; a reasoning module configured to utilize a cross-modal semantic reasoning system supported by a large model to perform cross-modal semantic mapping and intention reasoning on the multi-level feature set to obtain a collaborative execution instruction including a target action and force control parameters; The decision module is configured to dynamically update the behavior decision model of the humanoid robot based on the collaborative execution instruction and utilize a distributed incremental learning mechanism with double buffering collaboration, and obtain the real-time collaborative decision of the humanoid robot based on the behavior decision model.
9. A computer device, characterized in that: include: memory and processor, The memory stores a computer program; The processor is configured to execute a computer program stored in the memory, wherein the computer program enables the processor to execute the method according to any one of claims 1 to 7 when the program is executed.
10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Humanoid robot control method and system and related equipment
CN120755894A
Method and system for edge-end multi-mode perception and decision collaboration
CN120873530A
Human-computer interaction method and system based on intelligent sensor
CN121093971A
Humanoid robot teleoperation method based on multi-modal data fusion and related equipment
CN121132639A
Robot decision control method based on gradient rarefaction and robot
CN121157058A