Environment interaction method, system, equipment and product of humanoid robot
Through multimodal sensor data processing and brain-like neural circuits, robots can adapt to environmental changes, improving the interaction capabilities and work efficiency of humanoid robots.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-31
- Publication Date
- 2026-04-07
AI Technical Summary
In existing technologies, robots cannot adapt to changes in dynamic environments, resulting in poor interaction capabilities and affecting the work efficiency of humanoid robots.
By acquiring multimodal sensor data from a humanoid robot, feature extraction and fusion are performed. Then, brain-like neural circuits are used to extract environmental dynamics information, generate interactive control commands, and drive the robot to complete environmental interactive actions.
It enables robots to adapt and interact in dynamic environments, improves scene migration capabilities and decision-making intelligence, and increases work efficiency.
Smart Images

Figure CN121809538A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of robots, and in particular to an environment interaction method, system, device and product of a humanoid robot. BACKGROUND
[0002] In the related art, there is an environment interaction mode of robots using perception separation, that is, the robots manually configure sensor parameters, decision rules and execution programs for specific scenes. However, in actual application, it is found that when the scene changes, the robots need to redesign algorithms, which cannot adapt to dynamic environments, resulting in poor interaction ability of the robots and affecting the working efficiency of the humanoid robot.
[0003] To sum up, the technical problems existing in the related art need to be improved. SUMMARY
[0004] The main purpose of the embodiments of the present application is to provide an environment interaction method, system, device and product of a humanoid robot, which can improve the working efficiency of the humanoid robot.
[0005] To achieve the above-mentioned purpose, one aspect of the embodiments of the present application provides an environment interaction method of a humanoid robot, which comprises: obtaining multi-modal sensor data of the humanoid robot; performing feature extraction processing on the multi-modal sensor data to obtain global environment features; extracting environment dynamics information from the global environment features based on the time sequence features of the humanoid robot; inputting the environment dynamics information into a brain-like neural circuit network to output interaction control instructions; parsing the interaction control instructions into control signals to drive the humanoid robot to complete environment interaction actions.
[0006] In some embodiments, the feature extraction processing on the multi-modal sensor data to obtain global environment features comprises: performing feature encoding processing on the multi-modal sensor data to obtain a plurality of sensor features; performing feature fusion processing on the plurality of sensor features to obtain fusion features; mapping the fusion features to a global environment coordinate system to obtain the global environment features.
[0007] In some embodiments, the extracting environment dynamics information from the global environment features based on the time sequence features of the humanoid robot comprises: determining a previous time execution action signal and historical time environment information based on the time sequence features; fitting the global environment feature and the previous time action signal fitting to obtain an environment dynamics posterior distribution; fitting the historical time environment information and the previous time action signal fitting to obtain an environment dynamics prior distribution; generating the current time environment dynamics information based on the environment dynamics posterior distribution and the environment dynamics prior distribution.
[0008] In some embodiments, the generating the current time environment dynamics information based on the environment dynamics posterior distribution and the environment dynamics prior distribution further comprises: inputting the environment dynamics posterior distribution and the environment dynamics prior distribution into a pre-constructed recurrent neural network; training the recurrent neural network by taking the relative entropy minimization of the environment dynamics posterior distribution and the environment dynamics prior distribution as the target, and outputting the current time environment dynamics information.
[0009] In some embodiments, the inputting the environment dynamics information into a brain-like neural circuit network to output an interaction control instruction comprises: the brain-like neural circuit network comprises a perception neuron layer, an internal neuron layer, an instruction neuron layer, and an execution neuron layer; performing visual edge extraction on the environment dynamics information through the perception neuron layer to obtain visual edge features; performing multi-source information integration processing on the visual edge features through the internal neuron layer to obtain an environment state integrated feature vector; performing recurrent decision processing on the environment state integrated feature vector through the instruction neuron layer to obtain the interaction decision instruction; performing motion control simulation processing on the interaction decision instruction through the execution neuron layer to obtain the interaction control instruction.
[0010] In some embodiments, the performing visual edge extraction on the environment dynamics information through the perception neuron layer to obtain visual edge features comprises: performing convolution operation with lateral inhibition on the environment dynamics information to filter redundant visual information to obtain visual information; performing edge feature extraction processing on the visual information to obtain the visual edge features.
[0011] In some embodiments, the performing multi-source information integration processing on the visual edge features through the internal neuron layer to obtain an environment state integrated feature vector comprises: performing attention weighted fusion processing on the visual edge features and the environment dynamics information to obtain a fusion feature; The fused features are subjected to information redundancy elimination and dimensionality reduction processing to obtain the integrated feature vector of the environmental state.
[0012] To achieve the above objectives, another aspect of this application proposes an environmental interaction system for a humanoid robot, the system comprising: The data acquisition module is used to acquire multimodal sensor data of the humanoid robot; The feature extraction module is used to perform feature extraction processing on the multimodal sensing data to obtain global environmental features; An information extraction module is used to extract environmental dynamics information from the global environmental features based on the temporal characteristics of the humanoid robot; The network processing module is used to input the environmental dynamics information into the brain-like neural circuit network and output interactive decision instructions. The environment interaction module is used to parse the interaction decision instructions into control signals to drive the humanoid robot to complete the environment interaction actions.
[0013] To achieve the above objectives, another aspect of this application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the method described above.
[0014] To achieve the above objectives, another aspect of the embodiments of this application proposes a computer-readable storage medium storing a computer program that, when executed by a processor, implements the methods described above.
[0015] To achieve the above objectives, another aspect of this application provides a computer program product, including a computer program that, when executed by a processor, implements the methods described above. The embodiments of this application include at least the following beneficial effects: This application provides a method, system, device, and product for environmental interaction of a humanoid robot. This solution acquires multimodal sensing data from the humanoid robot; performs feature extraction processing on the multimodal sensing data to obtain global environmental features; extracts environmental dynamics information from the global environmental features based on the temporal features of the humanoid robot; inputs the environmental dynamics information into a brain-like neural circuit network to output interactive control commands; and parses the interactive control commands into control signals to drive the humanoid robot to complete environmental interaction actions. This application processes data and outputs corresponding interactive control commands through a brain-like neural circuit network, eliminating the need for extensive manual engineering configuration. It can dynamically adapt to environmental changes, improving the humanoid robot's scene migration capability and decision-making intelligence level, and increasing the robot's working efficiency. Attached Figure Description
[0016] Figure 1 This is a schematic diagram of an implementation environment provided in an embodiment of this application; Figure 2 This is a flowchart of an environmental interaction method for a humanoid robot provided in an embodiment of this application; Figure 3 This is a schematic diagram of the structure of an environmental interaction system for a humanoid robot provided in an embodiment of this application; Figure 4 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0017] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit it. In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with those of this application; they are merely examples of systems and methods consistent with some aspects of the embodiments of this application as detailed in the appended claims.
[0018] It is understood that the terms “first,” “second,” etc., used in this application may be used herein to describe various concepts, but unless otherwise stated, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the words “if,” “when,” or “in response to a determination” as used herein may be interpreted as “when…” or “when…” or “in response to a determination.”
[0019] As used in this application, the terms "at least one", "multiple", "each", "any", etc., "at least one" includes one, two or more, "multiple" includes two or more, "each" refers to each of the corresponding multiples, and "any" refers to any one of the multiples.
[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0021] Before providing a detailed description of the embodiments of this application, some of the nouns and terms involved in the embodiments of this application will be explained first. The nouns and terms involved in the embodiments of this application are subject to the following interpretations.
[0022] 1) Humanoid robots, also known as bionic robots, are robots designed to mimic human appearance and behavior, especially those with similar physiques to humans. The structural design of humanoid robots represents a remarkable reshaping of the human body, requiring not only interdisciplinary integration but also the culmination of cutting-edge technologies. Their design principles primarily include the following aspects: the organic integration of bionics and mechanical engineering, breakthroughs in the integration of sensing technology and control theory, and precise coordination between drive mechanisms and execution actions.
[0023] 2) Reinforcement Learning (RL) is a machine learning method. Its fundamental framework is the Markov Decision Process, which allows an agent to learn optimal policies through trial and error in its interactions with the environment. The agent performs actions in the environment and receives feedback, or rewards, based on the outcomes of those actions. These reward signals guide the agent to adjust its policy to maximize long-term cumulative rewards.
[0024] In related technologies, humanoid robots mostly adopt a modular design that separates "perception-decision-execution". They need to manually configure sensor parameters, decision rules and execution programs for specific scenarios, such as industrial assembly and household sweeping. When the scenario changes, such as industrial workpieces being replaced with PCB boards or furniture being moved in a household environment, the algorithm needs to be redesigned, and they cannot adapt to dynamic environments.
[0025] In addition, humanoid robots need to rely on multiple sensors such as vision, touch, and force to obtain environmental information. However, existing technologies mostly process single-modal data separately, such as relying solely on vision to locate objects or relying solely on force to control grasping. They lack the ability to fuse cross-modal information, resulting in a one-sided understanding of the environment. For example, when a home robot grasps a glass of water using only vision, it is prone to slipping because it does not perceive the smoothness of the surface, resulting in a high failure rate.
[0026] Furthermore, the decision-making modules in related technologies are mostly based on rule bases or traditional deep learning models, lacking the end-to-end "perception-integration-decision-execution" capabilities similar to biological nervous systems, and lacking a direct visual-motor coordination mechanism, thus failing to meet the rapid interaction needs of humanoid robots.
[0027] In view of this, this application provides a method, system, device, and product for environmental interaction of a humanoid robot, which can be applied to human-computer interaction application scenarios. Specifically, the environmental interaction method for a humanoid robot provided in this application can be applied to the controller of the humanoid robot, or to the server controlling the humanoid robot. Taking its application to the controller of the humanoid robot as an example, it can generate specific control programs and execute specific action generation steps based on the controller.
[0028] This application embodiment uses multimodal sensor data from a humanoid robot as input, and obtains environmental features from a global perspective through multimodal feature encoding and fusion; and extracts and memorizes dynamic environmental dynamic information based on historical environmental information and real-time features; converts the environmental dynamic information into decision commands through a brain-like neural circuit network to obtain corresponding interactive control commands; and parses the commands into joint control signals and end effector action signals to achieve adaptive interaction in industrial and home scenarios.
[0029] It should be noted that in all specific embodiments of this application, when processing data related to user identity or characteristics, such as user information, user behavior data, user historical data, and user location information, user permission or consent is obtained first. Furthermore, the collection, use, and processing of this data comply with relevant laws, regulations, and standards. In addition, when embodiments of this application require access to sensitive personal information of users, separate permission or consent from the user is obtained through pop-ups or redirection to confirmation pages. Only after obtaining the user's separate permission or consent is the necessary user-related data required for the proper functioning of these embodiments acquired.
[0030] Figure 1 This is a schematic diagram illustrating the implementation environment of a method provided in an embodiment of this application. (Refer to...) Figure 1 The main hardware and software components of this implementation environment include a humanoid robot 101 and a server 102, with the humanoid robot 101 and server 102 communicating with each other. The method can be executed based on the interaction between the humanoid robot 101 and server 102. Furthermore, the humanoid robot 101 and server 102 can be nodes in a blockchain, but this embodiment does not specifically limit this.
[0031] Figure 2 This is an optional flowchart of an environmental interaction method for a humanoid robot provided in an embodiment of this application. Figure 2 The method may include, but is not limited to, steps S102 to S205.
[0032] Step S201: Acquire multimodal sensing data of the humanoid robot; Step S202: Perform feature extraction processing on the multimodal sensing data to obtain global environmental features; Step S203: Extract environmental dynamics information from the global environmental features based on the temporal characteristics of the humanoid robot; Step S204: Input the environmental dynamics information into the brain-like neural circuit network and output interactive control commands; Step S205: The interactive control command is parsed into a control signal to drive the humanoid robot to complete the environmental interaction action.
[0033] Steps S201 to S205, as illustrated in this embodiment, involve using multimodal sensor data from a humanoid robot as input. The input data undergoes multimodal feature encoding, cross-modal fusion, and global perspective mapping to obtain environmental features from a global perspective. Then, based on these global environmental features, the current environmental dynamics information is obtained. This environmental dynamics information includes the positional changes of dynamic targets in the environment, the physical properties of objects, and their interaction states. The environmental dynamics information is then output to a brain-like neural circuit network, which can employ a fruit fly visual-motor coordinating brain-like neural circuit decision module. This fruit fly visual-motor coordinating brain-like neural circuit decision module simulates the fruit fly visual-motor coordinating neural system, including the direct projection mechanism between the visual signal processing unit of the optic lobe and the motion control unit of the thoracic ganglion, establishing a four-layer brain-like neural circuit network. The environmental dynamics information is then input into the brain-like neural circuit network to generate interactive control commands for the humanoid robot. Finally, the interactive control commands are parsed into joint control signals and end effector motion signals. The joint control signals are used to control the joint angles of the humanoid robot's torso and arms; the end effector motion signals are used to control the hand's grasping force and gripping angle, driving the humanoid robot to complete environmental interactive actions.
[0034] One of the above technical solutions has the following advantages or beneficial effects: the embodiments of this application do not require a large amount of manual engineering configuration, can dynamically adapt to environmental changes, improve the scene migration ability and decision-making intelligence level of humanoid robots, and compared with the related nematode neural circuit model, the interaction response speed is significantly improved, the industrial assembly precision reaches the sub-millimeter level, and the household grasping success rate is high.
[0035] In step S201 of some embodiments, multimodal sensing data of the humanoid robot is acquired; This application relates to environmental perception, dynamic memory, and brain-like decision-making technologies for humanoid robots, particularly suitable for adaptive interaction and intelligent decision-making in industrial and home scenarios. Industrial scenarios may include precision electronic component assembly and heavy workpiece handling, while home scenarios may include household item services and elderly / child assisted care. Multimodal sensing data, including visual sensor data, tactile sensor data, and force sensor data, can be acquired through various sensors pre-installed on the humanoid robot.
[0036] In step S202 of some embodiments, the step of performing feature extraction processing on the multimodal sensing data to obtain global environmental features includes: The multimodal sensing data is subjected to feature encoding processing to obtain multiple sensing features; The multiple sensing features are fused to obtain fused features; The fused features are mapped to the global environment coordinate system to obtain the global environment features.
[0037] In this embodiment, visual features, including object contours and position coordinates, are extracted using ResNet50. Tactile features, including contact pressure distribution and surface hardness, are extracted using a 3-layer MLP. Force features, including joint torque and end-effector force, are extracted using a 2-layer MLP, resulting in multiple sensor features. Then, an attention mechanism is used to assign different weights to the different sensor features for fusion processing to obtain fused features. For example, in an industrial scenario, visual features can be assigned a weight of 0.6, tactile features a weight of 0.2, and force features a weight of 0.2. The fused features are then mapped to the global environmental coordinate system of "robot base coordinate system - environmental target" to eliminate the perspective bias of local sensors, thus obtaining global environmental features.
[0038] One of the above technical solutions has the following advantages or beneficial effects: By fusing visual, tactile, and force data through an attention mechanism and setting a corresponding global environmental coordinate system, the embodiment of this application can eliminate the perspective deviation of local sensors and provide a data basis for subsequent control command generation.
[0039] In step S203 of some embodiments, extracting environmental dynamics information from the global environmental features based on the temporal features of the humanoid robot includes: Based on the aforementioned temporal characteristics, determine the action signal executed at the previous moment and the environmental information at the historical moment; The posterior distribution of environmental dynamics is obtained by fitting the global environmental features and the action signal executed at the previous moment. The prior distribution of environmental dynamics is obtained by fitting the historical environmental information and the action signal executed at the previous moment. The environmental dynamics information for the current moment is generated based on the posterior distribution and the prior distribution of environmental dynamics.
[0040] In this embodiment, the posterior distribution of environmental dynamics is obtained by fitting global environmental features and the action signal executed at the previous moment, and the prior distribution of environmental dynamics is obtained by fitting historical environmental information and the action signal executed at the previous moment.
[0041] Among them, posterior features Hidden features containing historical moment information The action signal executed at the previous moment and global environmental perspective features Sampling generation; prior features Hidden features containing historical moment information and the action signal executed at the previous moment Sample generation; Hypothetical posterior features with prior features All conform to a normal distribution, and their generation process is represented as follows: ; In the formula, This represents the global environmental perspective features after multimodal fusion at time k; Represents a multimodal feature fusion function; This represents the visual sensor input at time k, such as an environmental image or object outline. This represents the tactile sensor input at time k, such as contact pressure or the surface hardness of the object. This represents the force sensor input at time k, such as the force on the end effector or the joint torque. Represents the probability distribution of posterior features; Indicates a normal distribution. The mean, For variance; Let these represent the mean and variance parameters of the posterior distribution, respectively (derived from the parameters). control); This indicates the hidden features at time k, including storing historical environmental dynamics information. In industrial scenarios, the focus is on the workpiece pressure-location correlation, while in home scenarios, the focus is on the item location-user habit correlation. This indicates the execution action signal output by the control module at time k-1; Represents the probability distribution of prior features; Let these represent the mean and variance parameters of the prior distribution, respectively (derived from the parameters). control); Represents the hidden features at time k+1; This represents the feature update function of a recurrent neural network (RNN), used for the transmission of temporal information, and employs an LSTM structure to adapt to the long-term memory requirements.
[0042] The generation of environmental dynamics information at the current moment based on the environmental dynamics posterior distribution and the environmental dynamics prior distribution includes: The posterior distribution and the prior distribution of environmental dynamics are input into a pre-constructed recurrent neural network; The recurrent neural network is trained with the goal of minimizing the relative entropy of the posterior and prior distributions of environmental dynamics, and the environmental dynamics information at the current moment is output.
[0043] This embodiment trains the model with the objective of minimizing the relative entropy between the posterior distribution and the prior distribution of environmental dynamics, and generates environmental dynamics information for the current moment based on the posterior distribution and hidden features. Simultaneously, it updates the hidden features using the current environmental dynamics information as the hidden features for the next moment. These hidden features are temporal features containing environmental information from historical moments.
[0044] For example, embodiments of this application perform posterior distribution fitting based on global environmental perspective features and temporal features. It can also combine real-time sensor data, such as chip position deviation in industrial scenarios and water cup contact pressure in household scenarios, with prior distribution fitting. By combining historical information, such as chip movement patterns and user water cup habitual positions, environmental dynamics information can be extracted. At the same time, embodiments of this application can also update historical environmental information through LSTM to achieve long-term memory of dynamic environment.
[0045] In step S204 of some embodiments, inputting the environmental dynamics information into a brain-like neural circuit network and outputting interactive control commands includes: The brain-like neural circuit network includes a sensory neuron layer, an internal neuron layer, a command neuron layer, and an execution neuron layer; Visual edge features are obtained by extracting visual edges from the environmental dynamics information through the perceptual neuron layer. The visual edge features are integrated through the internal neuron layer to obtain an integrated feature vector of environmental state. The interactive decision-making instruction is obtained by performing cyclic decision processing on the integrated feature vector of the environmental state through the instruction neuron layer. The interactive control instructions are obtained by performing motion control simulation processing on the interactive decision instructions through the execution neuron layer.
[0046] In this embodiment, a brain-like neural circuit network is used to simulate the visual-motor synergistic neural system of a fruit fly, including the direct projection mechanism between the visual signal processing unit of the optic lobe and the motor control unit of the thoracic ganglion, establishing a four-layer brain-like neural circuit network; and inputting environmental dynamics information into the brain-like neural circuit network to generate interactive decision-making instructions for the humanoid robot; the four-layer brain-like neural circuit network includes: One sensory neuron is used to correspond to the L2 / L4 neurons of the optic nerve lobe in fruit flies, and is responsible for visual edge feature extraction; One internal neuron, used to correspond to the integrated neurons of the Drosophila central nervous system; One instruction neuron corresponds to the decision neurons in the fruit fly's brain center; and Each neuron is an executive neuron, corresponding to a neuron in the thoracic ganglion of the fruit fly, responsible for adjusting movement posture.
[0047] In some embodiments, the step of extracting visual edges from the environmental dynamics information through the perceptual neuron layer to obtain visual edge features includes: The environmental dynamics information is processed by convolution with side suppression to filter out redundant visual information and obtain the visual information. The visual information is processed by edge feature extraction to obtain the visual edge features.
[0048] In this embodiment, the input to the perceptual neuron layer is environmental dynamics information, which includes global environmental features fused from multiple modalities. For example, in an industrial scenario, this includes "positional deviation of the PCB board-chip and distribution of clamp contact pressure," and in a home scenario, it includes "layout features of objects and furniture and temporal features of user interaction habits," generated by attention fusion of visual, tactile, and force features. By applying convolutional operations with lateral inhibition to the input global environmental features, redundant visual information, such as PCB board background noise and irrelevant furniture outlines in a home scenario, is filtered out. Key visual edge features, such as the outline edges of chip pads and the geometric outline of a water cup, are extracted to generate visual edge feature vectors, thus obtaining visual edge features. These visual edge feature vectors are then passed to the internal neuron layer; and through the sub-output, a shortcut connection is used to directly pass the raw visual fast features, which have not undergone extensive modal fusion, to the execution neuron layer.
[0049] The embodiments of this application can simulate the rapid projection mechanism of "visual signal → motion signal" in fruit flies, meeting the millisecond-level response requirements for emergency obstacle avoidance and precision assembly.
[0050] In some embodiments, the step of integrating multi-source information on the visual edge features through the internal neuron layer to obtain an integrated environmental state feature vector includes: The visual edge features and the environmental dynamics information are subjected to attention-weighted fusion processing to obtain fused features; The fused features are subjected to information redundancy elimination and dimensionality reduction processing to obtain the integrated feature vector of the environmental state.
[0051] In this embodiment, environmental dynamics information is obtained by acquiring the visual edge feature vector output from the perceptual neuron layer and the hidden features of the environmental memory module. Then, simulating the multi-source information integration mechanism of the Drosophila central neurons, attention-weighted fusion is used to sum the weighted values of "visual edge features (weight 0.7)" and "hidden features (weight 0.3)". Information redundancy is eliminated through a fully connected network (2 layers, ReLU activation function), generating an integrated environmental state feature (dimensional compressed to 48, containing fused information of "current environmental dynamics + historical patterns"). The output is an integrated environmental state feature vector (48 dimensions), which is then passed to the instruction neuron layer.
[0052] Specifically, the instruction neuron layer acquires the integrated environmental state features from the output of the internal neuron layer and introduces them into a recurrent neural network (LSTM structure) to memorize historical decision sequences, such as "joint adjustments in the first 3 assembly attempts" in an industrial scenario and "finger bending angles in the first 5 grasping attempts" in a home scenario. The integrated features of the current environmental state are fused with the historical decision sequences through a gating mechanism to generate interactive decision instructions, such as "shoulder joint angle +0.5°, clamp pressure -0.2N" in an industrial scenario and "finger bending 30°, grasping force 0.3N" in a home scenario. The output is a standardized decision instruction, which includes quantified parameters such as joint angles and end effector force control, and is then passed to the execution neuron layer.
[0053] For example, the execution neuron layer acquires the standardized decision instructions output by the instruction neuron layer and the visual fast features of the perception neuron layer. The standardized decision instructions are then mapped to scene-specific parameters by simulating the motion control mechanism of the Drosophila thoracic ganglion neurons. By combining the emergency response requirements of the shortcut visual fast features, such as the sudden appearance of obstacles, the mapped control signals are dynamically compensated. For example, the joint angle adjustment is increased by 0.1° in an industrial scenario and the grasping force is reduced by 0.05N in a household scenario. Finally, the execution control signal is output, which includes joint control signals and end effector action signals, such as "shoulder joint angle 0.5°, grasping force 0.3N".
[0054] This application embodiment fully realizes end-to-end brain-like collaboration of "environmental perception-memory-decision-execution" through a four-layer brain-like neural network, which extracts edge features through the perception layer, integrates multi-source information through the internal layer, generates decision sequences through the instruction layer, and parses control signals through the execution layer. Moreover, the input, processing, and output logic of each layer corresponds one-to-one with the physiological mechanism of the Drosophila visual-motor coordinating nervous system.
[0055] The four layers of neurons in the brain-like neural circuit network satisfy the following connection rules: A new shortcut connection is added between the sensory neurons and the executive neurons (simulating the direct projection mechanism from the optic lobe to the thoracic ganglion in fruit flies, enabling rapid transmission of visual signals to motor signals). Each sensory neuron inserts a connection into the executive neuron. One synapse, and ( To optimize the number of neurons, it is designed to adapt to emergency interaction scenarios for humanoid robots (such as obstacle avoidance in home scenarios and collision avoidance in industrial scenarios). Between any two consecutive layers (from sensory neuron to inner neuron, from inner neuron to instruction neuron, from instruction neuron to execution neuron), insert [the following] into each source neuron. One synapse, and ( (This represents the number of target neurons); synaptic polarity (excitability / inhibition) follows a Bernoulli distribution (probability p=0.5). Each target neuron is randomly selected using a binomial distribution; In any two consecutive layers, for the target neuron j that is not connected to a synapse, insert additional... One synapse, and ( (This represents the number of synaptic connections in target neuron i); synaptic polarity follows a Bernoulli distribution. Each source neuron is randomly selected using a binomial distribution; The instruction neurons are connected in a loop, and an instruction neuron is inserted into each instruction neuron. One synapse, and ( (This refers to the number of instruction neurons); synaptic polarity follows a Bernoulli distribution. Each source neuron is randomly selected using a binomial distribution.
[0056] Specifically, in this embodiment, the sensory neuron layer inputs rapid visual features, such as edge abrupt changes of a chip in an industrial scene or the visual contours of obstacles in a home scene (32-dimensional), and the execution neuron layer inputs standardized decision-making instructions. By inputting the rapid visual features of the sensory neuron, the direct projection mechanism from the optic lobe to the thoracic ganglion in a fruit fly is simulated. Through rapid synaptic transmission without intermediate integration (delay ≤ 0.01s), visual emergency signals, such as suddenly shifting chips or sudden obstacles, are directly transmitted to the execution neuron. The output of the execution neuron is then fused with the received decision-making instructions to provide emergency compensation for joint control signals or end effector signals. In this embodiment, other source neurons that have not yet established a connection with the target neuron j are randomly selected using a binomial distribution. There are 1 neuron, and the synaptic polarity follows a Bernoulli distribution (p=0.5). This allows the neuron to complete the input signal of the target neuron j, enabling it to activate normally and participate in subsequent information processing, such as in the inner neuron layer where the neuron was not activated. Neuron j connected via Synapses receive sensory features to avoid decision-making biases.
[0057] In this embodiment of the application, the dynamic model of each neuron in the brain-like neural circuit network is represented as follows: ; In the formula, This represents the synaptic current of the neuron at time t (reflecting the neuron's activation state). Represents the time constant (controlling the current decay rate, in industrial scenarios). Corresponding to the visual-motor response of fruit flies Physiological characteristics; Home use (Adapting to the slow-response requirements of flexible interaction). This represents the neuron activation function (using the ReLU function). ); t represents the external input of the neuron at time t (the sum of synaptic signals from the previous layer of neurons, including fast signals from shortcut connections); A represents the deviation matrix (in industrial scenarios, A=0.8, corresponding to the neuron activation threshold when fruit flies process precise visual signals (such as chip pad positioning) to ensure decision accuracy; in household scenarios, A=0.6, corresponding to the low activation threshold when fruit flies process flexible grasping signals (such as water cup contact) to ensure flexible movements).
[0058] In this embodiment, environmental dynamics information is converted into interactive decision instructions through function g, and the process is represented as follows: ; In the formula, represents the interactive decision instruction at time k (historical time); g represents the decision mapping function of the Drosophila visual-motor coordinating brain-like neural circuit; Represents the visual signal suppression subfunction in fruit flies. It is used to suppress weak visual signals, mimicking the signal filtering mechanism of the optic nerve lobe of fruit flies, and reducing interference from redundant information. This indicates the hidden feature at time k; Represents the posterior features at time k; This represents the interactive decision-making instruction at time k+T (a future time, without real-time sensor input); Represents the hidden features at time k+T; This represents the prior features at time k+T.
[0059] Specifically, decision-making instructions It is a composite instruction that includes joint angle adjustment coefficients and end effector force control coefficients. The execution control module executes the instruction through functions. Two-way analysis: ; In the formula, This represents the execution action signal (including joint control signals) at time k. With end effector signal ); This indicates the execution signal parsing function (including linear scaling and force control compensation terms, adapted to the robot's motion accuracy requirements). Represents the interactive decision instruction at time k; Represents the joint stiffness coefficient (in industrial applications). Suitable for rigid movements in precision assembly; suitable for home use. (Adapts to flexible grasping elastic movements). Indicates the force control coefficient of the end effector (in industrial scenarios) This corresponds to the contact force threshold of the footpads of fruit flies when grasping hard objects; for household scenarios. (corresponding to the low contact force threshold for fruit flies to grasp soft objects). The symbol function is used to ensure that the force control compensation direction is consistent with the decision command direction.
[0060] For example, embodiments of this application select training data and time... to The data from multiple sensors and the execution signals are used as historical training data. The industrial data comes from a chip of a certain specification, and the household data comes from three typical household scenarios of grasping wooden / glass / plastic water cups. The predicted data is used as future training data; the training objective is set to maximize the execution of the action sequence. Interaction sequence with the environment joint probability Variational inference is used to obtain the lower bound of the joint probability, expressed by the following formula: ; In the formula, This represents the expectation operation; This represents the probability of executing the action signal at time t; This represents the probability of the environmental interaction effect at time t (in an industrial scenario). Quantified as "1 (assembly accuracy ≤ 0.02mm), 0.5 (0.02mm < accuracy ≤ 0.05mm), 0 (accuracy > 0.05mm)", for home use scenarios. Quantified as "1 (successful grab without slippage), 0 (failed grab or object deformation)"); Describe the posterior distribution With prior distribution The relative entropy (a measure of distributional disparity; the training objective is to make this value ≤ 0.1).
[0061] This application relates to environmental perception, dynamic memory, and brain-like decision-making technologies for humanoid robots. These technologies are particularly suitable for adaptive interaction and intelligent decision-making in industrial scenarios (such as precision electronic component assembly and heavy workpiece handling) and home scenarios (such as home furnishing services and elderly / child care assistance). The core technology relies on the fruit fly vision-motor collaborative brain-like neural circuit to achieve rapid coordination of "perception-action," solving the problems of insufficient multimodal fusion and poor scene transferability of robots in related technologies.
[0062] This application can be specifically applied to industrial scenarios, such as PCB assembly. First, images of the PCB board and the assembly station are acquired using a visual sensor. The contact pressure distribution between the fixture and the assembly station is obtained using a tactile sensor. The torques of the shoulder and elbow joints during assembly are obtained using a force sensor. Then, the visual features are processed using ResNet50 to extract the PCB board center coordinates and chip outline. The tactile features are processed using a 3-layer MLP to extract the average pressure distribution. The force features are processed using a 2-layer MLP to extract the joint torque change rate. An attention mechanism is used to assign weights of 0.6 to the visual features, 0.2 to the tactile features, and 0.2 to the force features, resulting in a 64+16+16=96-dimensional multimodal fusion feature. This fusion feature is mapped to the global coordinate system of "PCB board-humanoid robot arm," with the origin at the lower left corner of the PCB board, to obtain the global environmental perspective features.
[0063] This application's embodiments combine global environmental perspective features and the previous moment's fixture pressure signal with a posterior distribution fitted using a two-layer fully connected network to generate posterior features. These posterior features reflect the dynamic correlation between the current chip position and pressure. A prior distribution is fitted using a two-layer fully connected network by fitting the chip offset patterns from 10 historical assembly data sets and the previous moment's fixture pressure signal. The training unit aims to minimize the relative entropy between the posterior and prior distributions, updating hidden features through LSTM to memorize the correlation between the current chip pressure and offset.
[0064] Then, the environmental dynamics information is input into a brain-like neural circuit network. Perceptual neurons suppress weak visual signals, such as PCB background noise, through lateral inhibition functions, outputting effective features to internal neurons. Internal neurons integrate the data and eliminate positional deviation noise. Command neurons generate assembly decision commands through cyclic connections, such as: "Shoulder joint angle +0.5°, elbow joint angle -0.3°, clamp pressure reduced to 0.6N". Execution neurons receive command signals and shortcut signals from perceptual neurons, outputting standardized control signals. Finally, joint control signals and end effector signals are generated to control the robot, ultimately completing the precision assembly of the PCB board at the workstation, improving assembly accuracy and reducing assembly time.
[0065] Please see Figure 3This application also provides an environmental interaction system for a humanoid robot, which can implement the above-described method. The system includes: Data acquisition module 301 is used to acquire multimodal sensing data of the humanoid robot; Feature extraction module 302 is used to perform feature extraction processing on the multimodal sensing data to obtain global environmental features; The information extraction module 303 is used to extract environmental dynamics information from the global environmental features based on the temporal features of the humanoid robot; Network processing module 304 is used to input the environmental dynamics information into a brain-like neural circuit network and output interactive decision instructions; The environment interaction module 305 is used to parse the interaction decision instructions into control signals to drive the humanoid robot to complete the environment interaction actions.
[0066] It is understood that the content of the above method embodiments is applicable to this system embodiment. The specific functions implemented in this system embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.
[0067] This application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described method. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.
[0068] It is understood that the content of the above method embodiments is applicable to this device embodiment. The specific functions implemented by this device embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0069] Please see Figure 4 , Figure 4 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes: The processor 401 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application. The memory 402 can be implemented as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 402 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 402 and is called and executed by the processor 401 using the methods described in the embodiments of this application. Input / output interface 403 is used to implement information input and output; The communication interface 404 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.). Bus 405 transmits information between various components of the device (e.g., processor 401, memory 402, input / output interface 403, and communication interface 404); The processor 401, memory 402, input / output interface 403 and communication interface 404 are connected to each other within the device via bus 405.
[0070] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method.
[0071] It is understood that the content of the above method embodiments is applicable to this storage medium embodiment. The specific functions implemented in this storage medium embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.
[0072] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.
[0073] It is understood that the content of the above method embodiments is applicable to the embodiments of this program product. The specific functions implemented by the embodiments of this program product are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0074] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0075] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.
[0076] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.
[0077] The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0078] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.
[0079] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0080] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0081] In the embodiments provided in this application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between systems or units may be electrical, mechanical, or other forms.
[0082] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0083] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0084] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0085] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.
Claims
1. A method for environmental interaction of a humanoid robot, characterized in that, The method includes the following steps: Acquire multimodal sensor data from humanoid robots; The multimodal sensing data is processed to extract features to obtain global environmental features; Environmental dynamics information is extracted from the global environmental features based on the temporal characteristics of the humanoid robot. The environmental dynamics information is input into a brain-like neural circuit network, and interactive control commands are output. The interactive control commands are parsed into control signals to drive the humanoid robot to complete environmental interaction actions.
2. The method according to claim 1, characterized in that, The step of performing feature extraction processing on the multimodal sensing data to obtain global environmental features includes: The multimodal sensing data is subjected to feature encoding processing to obtain multiple sensing features; The multiple sensing features are fused to obtain fused features; The fused features are mapped to the global environment coordinate system to obtain the global environment features.
3. The method according to claim 1, characterized in that, The extraction of environmental dynamics information from the global environmental features based on the temporal features of the humanoid robot includes: Based on the aforementioned temporal characteristics, determine the action signal executed at the previous moment and the environmental information at the historical moment; The posterior distribution of environmental dynamics is obtained by fitting the global environmental features and the action signal executed at the previous moment. The prior distribution of environmental dynamics is obtained by fitting the historical environmental information and the action signal executed at the previous moment. The environmental dynamics information for the current moment is generated based on the posterior distribution and the prior distribution of environmental dynamics.
4. The method according to claim 3, characterized in that, The generation of environmental dynamics information at the current moment based on the environmental dynamics posterior distribution and the environmental dynamics prior distribution includes: The posterior distribution and the prior distribution of environmental dynamics are input into a pre-constructed recurrent neural network; The recurrent neural network is trained with the goal of minimizing the relative entropy of the posterior and prior distributions of environmental dynamics, and the environmental dynamics information at the current moment is output.
5. The method according to claim 1, characterized in that, The process of inputting the environmental dynamics information into a brain-like neural circuit network and outputting interactive control commands includes: The brain-like neural circuit network includes a sensory neuron layer, an internal neuron layer, a command neuron layer, and an execution neuron layer; Visual edge features are obtained by extracting visual edges from the environmental dynamics information through the perceptual neuron layer. The visual edge features are integrated through the internal neuron layer to obtain an integrated feature vector of environmental state. The interactive decision-making instruction is obtained by performing cyclic decision processing on the integrated feature vector of the environmental state through the instruction neuron layer. The interactive control instructions are obtained by performing motion control simulation processing on the interactive decision instructions through the execution neuron layer.
6. The method according to claim 5, characterized in that, The step of extracting visual edge features from the environmental dynamics information through the perceptual neuron layer includes: The environmental dynamics information is processed by convolution with side suppression to filter out redundant visual information and obtain the visual information. The visual information is processed by edge feature extraction to obtain the visual edge features.
7. The method according to claim 5, characterized in that, The step of integrating multi-source information from the visual edge features through the internal neuron layer to obtain an integrated environmental state feature vector includes: The visual edge features and the environmental dynamics information are subjected to attention-weighted fusion processing to obtain fused features; The fused features are subjected to information redundancy elimination and dimensionality reduction processing to obtain the integrated feature vector of the environmental state.
8. An environmental interaction system for a humanoid robot, characterized in that, The system includes: The data acquisition module is used to acquire multimodal sensor data of the humanoid robot; The feature extraction module is used to perform feature extraction processing on the multimodal sensing data to obtain global environmental features; An information extraction module is used to extract environmental dynamics information from the global environmental features based on the temporal characteristics of the humanoid robot; The network processing module is used to input the environmental dynamics information into the brain-like neural circuit network and output interactive decision instructions. The environment interaction module is used to parse the interaction decision instructions into control signals to drive the humanoid robot to complete the environment interaction actions.
9. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method according to any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 7.
Citation Information
Patent Citations
Brain-like decision and motion control system
CN110427536A
Unmanned aerial vehicle automatic driving method and system, medium and equipment
CN118915787A
Humanoid robot decision interaction method, system and equipment and medium
CN120974368A