Self-adaptive adjustment system for behavior mode of intelligent robot with body

Through technical means such as neuromorphic perception modules and cognitive constraint decision-making modules, the problems of unbalanced multimodal data processing and waste of computing resources in the intelligent robot behavior pattern adaptive adjustment system are solved, efficient environmental perception and behavior optimization are achieved, and the system's adaptability and resource utilization efficiency are improved.

CN120773064AActive Publication Date: 2025-10-14SHENZHEN QIANHAI GEZHI TECH CO LTD

Patent Information

Application Number
CN202511246234.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-02
Publication Date
2025-10-14
Estimated Expiration
2045-09-02

AI Technical Summary

Technical Problem

Existing intelligent robot behavior pattern adaptive adjustment systems lack an effective multimodal data alignment and normalized mapping mechanism, resulting in information loss or bias, unbalanced multi-channel processing, lack of sparse activation capabilities, inability to respond to complex environments in real time, and serious waste of computing resources.

Method used

A neuromorphic perception module is used to perform spatiotemporal compression of multimodal environmental perception information, and a sparse activation sequence is generated by drawing on the sparse coding principle of the brain's visual cortex. Combined with a cognitive constraint decision module and a variable impedance execution module, dynamic adjustment and optimization of robot behavior are achieved through a cross-modal continuous learning module and an embodied behavior inversion module.

Benefits of technology

It improves the accuracy and responsiveness of multimodal environmental perception, reduces computing resource consumption, enhances understanding and adaptability to complex environments, has the ability to migrate across scenarios, and reduces energy consumption and computing redundancy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120773064A_ABST
    Figure CN120773064A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of intelligent control of robots, and discloses a self-adaptive adjustment system for behavior modes of an intelligent robot with a body, which comprises a neuromorphic sensing module for acquiring and fusing multi-mode environment sensing information, performing space-time compression on the multi-mode environment sensing information by referring to a brain visual cortex sparse coding principle, and generating a neural network; generating an environment state representation vector, and constructing a neuromorphic perception path; the cognitive constraint decision module is used for generating an action strategy by adopting a neural symbol hybrid architecture based on the environment state representation vector; through a cognitive momentum optimization mechanism, a strategy inertia item is introduced to improve an updating process of an action strategy parameter, and a behavior decision vector is obtained; the variable impedance execution module is used for mapping the behavior decision vector into a plurality of virtual muscle cooperation elements through a muscle cooperation mapping rule, and solving an expected movement track of each execution joint of the robot; and the intelligent robot with the body has higher adaptability, stability and execution capability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of robot intelligent control technology, and more specifically, to a system for adaptively adjusting the behavior pattern of an embodied intelligent robot. Background Art

[0002] Patent publication number CN119849301A discloses a quadruped robot motion control method and system based on reinforcement learning. By constructing a quadruped robot simulation model, its motion behavior can be simulated and analyzed in a virtual environment, significantly reducing R&D costs and cycles. The balance control problem is converted into a Markov decision process, and the state, action and reward functions are defined, which helps to achieve a more intelligent and adaptive balance strategy. The deterministic policy gradient algorithm is used for interactive training with the simulation model to obtain an efficient balance control strategy, and the robot behavior is adjusted in real time to maintain balance. Using the hierarchical control idea, the walking task is decomposed into two sub-tasks: foothold selection and gait generation. Through reinforcement learning, they are respectively implemented to improve the robot's walking flexibility and stability on irregular terrain. Irregular terrain is built on the simulation platform for experimental verification, which can safely and repeatedly evaluate the robot's performance and further optimize the strategy.

[0003] The existing intelligent robot behavior mode adaptive adjustment system has the following main problems: In existing systems, different types of perception modalities often lack effective modal alignment and normalization mapping mechanisms, resulting in information loss or bias after multimodal data fusion, which in turn affects the accuracy and real-time performance of environmental state recognition. In existing neural perception mechanisms, multi-channel processing methods are prone to the phenomenon that some channels are always active and some channels are long-term silent. There is a lack of dynamic threshold adjustment methods based on competition and regulation mechanisms, which leads to unbalanced system responses. Most state representation methods ignore the temporal evolution characteristics of perception signals and only use static feature vectors for expression. They are unable to express dynamic characteristics such as sudden events and intensity changes, which limits the system's ability to understand and respond to complex scenarios. Most traditional feature encoding methods do not have sparse activation capabilities, resulting in a large number of channels being calculated at every moment, which is not suitable for embedded or edge deployment.

[0004] Traditional robot motion strategies often rely on offline training or rule-setting, are unable to cope with perception changes and task disturbances in complex environments, and lack adaptability and transfer capabilities. Existing behavioral learning or strategy optimization often cannot utilize feedback signals from task execution to make strategy corrections in real time or within an appropriate time window, resulting in low feedback utilization. High energy consumption or redundant activation are prominent issues. In existing robot behavior pattern adaptive adjustment systems, behavioral decision paths are in a state of widespread activation, resulting in a large amount of computing resources being consumed on non-critical paths, causing increased energy consumption, delayed responses, and redundant computation during execution.

[0005] In view of this, the present invention proposes an embodied intelligent robot behavior mode adaptive adjustment system to solve the above problems. Summary of the Invention

[0006] In order to overcome the above-mentioned defects of the prior art and to achieve the above-mentioned objectives, the present invention provides the following technical solution: a system for adaptively adjusting the behavior pattern of an embodied intelligent robot, comprising: The neuromorphic perception module collects and integrates multimodal environmental perception information. Drawing on the sparse coding principle of the brain's visual cortex, it compresses the multimodal environmental perception information in time and space, generates an environmental state representation vector, and constructs a neuromorphic perception path. The cognitive constraint decision module uses a neural-symbolic hybrid architecture to generate action strategies based on the environmental state representation vector. Through the cognitive momentum optimization mechanism, the strategy inertia term is introduced to improve the update process of the action strategy parameters and obtain the behavior decision vector. The variable impedance execution module maps the action decision vector into multiple virtual muscle synergies through muscle synergy mapping rules, and calculates the expected motion trajectory of each execution joint of the robot. It introduces a nonlinear stiffness field mapping mechanism and combines it with the robot's contact feedback data to construct the Riemann stiffness tensor field of the action contact area and adjust the parameters of each execution joint. The cross-modal continuous learning module draws on the memory consolidation mechanism of the hippocampus and adopts the synaptic timing-dependent plasticity rule to adaptively update the synaptic connection weights in the neuromorphic perception pathway and generate action strategy generalization patches; The embodied behavior inversion module receives robot execution feedback data in real time, combines it with the action strategy generalization patch, dynamically optimizes the robot's behavior pattern, corrects the robot's expected motion trajectory and action strategy, and adaptively executes the corresponding task.

[0007] Preferably, the method for collecting and fusing multimodal environmental perception information includes: \Deploy a multimodal sensor array to collect multimodal environmental perception information in the environment where the intelligent robot is located. The sensors in the multimodal sensor array include visual sensors, auditory sensors, somatosensory modal sensors, temperature and humidity sensors, and light sensors; multimodal environmental perception information includes visual modal information, auditory modal information, somatosensory posture information, ambient temperature, ambient humidity, and light intensity.

[0008] Preferably, the method for generating the environmental state representation vector includes: The multimodal environment perception information is normalized and modally aligned, uniformly mapped to a shared feature space, and a multimodal environment feature vector is obtained. A sparse dictionary is constructed based on the receptive field mechanism of the visual cortex. The multimodal environment feature vector is sparsely represented as an atomic linear combination through the sparse dictionary to obtain a sparse activation vector. For the activation value of each sparse pulse channel in the sparse activation vector, a dynamic threshold adaptive adjustment mechanism is designed. , preset The activation value of the sparse spike channel is , No. The pulse triggering threshold of a sparse pulse channel is According to the The activation value of the sparse spike channel The ratio of the total activation value of all sparse pulse channels is used to dynamically adjust the pulse triggering threshold of the sparse pulse channel; In the preset time window Internal The activation values ​​of the sparse pulse channels are integrated and accumulated to obtain the A sparse pulse channel at time The cumulative activation value of ; Use the pulse emission discrimination function to determine whether the neuron is triggered to emit a pulse. When the cumulative activation amount of a sparse pulse channel is greater than or equal to the preset pulse trigger threshold, a pulse event is triggered, and a sparse pulse sequence is obtained; the sparse pulse sequence is input into the synaptic mapping function, and time smoothing, synaptic weighting and semantic aggregation are performed to generate an environmental state representation vector.

[0009] Preferably, the method for constructing the neuromorphic sensing pathway includes: The activated sparse pulse channels in the environmental state representation vector are regarded as the output of the primary perception element, and are routed to the initial node of the neuromorphic perception path through the index binding mechanism of the sparse coding dictionary to form the multi-channel pulse trigger source of the input layer; Based on the principle of biological synaptic plasticity, a synaptic mapping function connection map is constructed, and the synaptic connection weights are defined for each activated sparse pulse channel according to the STDP rule; synaptic propagation delay is modeled and the synaptic propagation delay is set according to the timing alignment error between different modalities; A connection sparsity constraint is introduced, and a preset activation frequency threshold for the synaptic path is set. The activation frequency of any synaptic path is counted. If the activation frequency of the synaptic path is greater than the preset activation frequency threshold, the target task corresponding to the synaptic path is determined to be strongly correlated with the preset target task. Only synaptic pathways that are strongly related to the preset target task are retained to form a dynamic and sparse pulse propagation pathway network. Pulse propagation pathways are selected based on the synaptic mapping function connection map, and the WTA mechanism is used to inhibit non-dominant pathways. For the same type of target task, only propagation pathways with pulse response strength greater than the preset pulse response strength threshold are retained. By setting a dynamic activation window, only a limited number of m effective channels are activated in any time period to control temporal sparsity. A multi-level neuromorphic network is constructed to receive and propagate sparse pulse signals, complete state aggregation, and obtain the final neuromorphic perception path.

[0010] Preferably, the method for generating the action strategy includes: Inputting the environmental state representation vector into a multi-layer neural network, which can be any form of a feedforward network, a convolutional network, or a spiking neural network; outputting a neural feature representation vector through the multi-layer neural network; A symbolic knowledge base is pre-set, including action-effect rules, state transition rules, and action constraint boundary information. A symbolic parser is used to convert the symbolic knowledge base into a computable structure to form a symbolic rule template. The neural feature representation vector output by the neural network is semantically aligned with the symbolic rule template, and matching is completed through an interface mechanism. The matched symbolic rule fragments are structured and combined to form a set of candidate action paths. The set of matched candidate action paths is optimized to generate action strategies. A scoring function is used to evaluate the execution cost, expected benefit, and task fit of each candidate action path, and different candidate action paths are ranked. The candidate path with the highest score is selected as the current action strategy. If there are multiple equivalent candidate paths, a neural evaluator can be introduced for further discrimination to output the final action strategy.

[0011] Preferably, the method for obtaining the behavior decision vector includes: To introduce the policy inertia term, a momentum term is constructed to record and accumulate the changing trend of the action policy parameters at each historical moment. This momentum term decays according to a preset ratio each time the action policy is updated and is superimposed on the current action policy gradient direction to perform temporal smoothing of the action policy evolution direction. The current gradient is weightedly fused with the momentum term of the previous moment, and the fused momentum term is used as the update direction of the action strategy parameters. The fused momentum term is used to update and adjust the action strategy parameters. This update and adjustment process is performed offline or asynchronously during the training phase of the policy network, so that the latest action strategy parameters are used in the next round of behavior generation. After the action strategy parameters are updated, the new environment state representation vector is input into the updated strategy network to generate a behavior decision vector, which includes the robot motion control parameters, action selection probability, control instruction encoding, adjustment feedback parameters and multimodal response output instructions.

[0012] Preferably, the method for obtaining the expected motion trajectory of each execution joint of the robot includes: The muscle synergy mapping rule is used to map the behavior decision vector into n virtual muscle synergists. Each muscle synergist corresponds to a group of virtual muscle units with biological synergy characteristics. The muscle synergy mapping rule is based on the physiological mechanism of human muscle control and predefines a group of muscle synergist combinations. Each combination controls the movement pattern of one or q joint movements. The virtual muscle synergists are activated and scheduled, and the expected tension output of each virtual muscle unit is calculated by combining the preset target action direction and the preset target force output. A mapping relationship is established between the robot body structure and the muscle synergists, and the expected tension output of each virtual muscle unit is applied to the corresponding pre-built bone-joint structure model. Based on the bone-joint structure model, the system mechanical response under muscle drive is solved to obtain the expected motion trajectory of each execution joint of the robot.

[0013] Preferably, the method for adjusting the parameters of each execution joint includes: The robot's contact feedback data includes the spatial position coordinates of the contact point, the contact normal component, the contact tangential component, the contact force magnitude, the contact force direction, and the duration of action. Based on the robot's contact feedback data, nonlinear stiffness modeling is performed on the robot's action contact area. The elastic response to external forces in different directions in space is described in the form of Riemann stiffness tensor field. In the contact space, a symmetrical positive stiffness tensor is assigned to each coordinate point, which describes the deformation resistance of the point in different directions. The tensor field has spatial inhomogeneity and directional correlation. The contact characteristics are mapped to stiffness responses by introducing a nonlinear mapping function. Based on the constructed Riemann stiffness tensor field, the parameters of each executive joint are adjusted. The influence range of each executive joint in the contact space and the corresponding contact area tensor field value are determined. For each executive joint, the adjustment coefficient is calculated according to its corresponding contact point stiffness tensor, and the expected motion trajectory is adjusted based on the adjustment coefficient.

[0014] Preferably, the method for generating an action strategy generalization patch includes: During the robot's intervals between tasks, the state-action sequence within h periods is internally replayed based on the hippocampal memory consolidation mechanism. For each sequence segment, the spiking neurons in the neuromorphic perception pathway are stimulated in sequence, causing the corresponding synaptic pathways to activate sequentially. During memory consolidation, presynaptic neurons The postsynaptic neuron fires a spike at time Release pulse; record ;in, Represents the firing time difference between the postsynaptic neuron and the presynaptic neuron; updates the weight of each synaptic connection according to the synaptic timing-dependent plasticity rule; After each update behavior is executed, the error between the current action strategy output and the expected behavior is calculated based on the execution result. Feedback the error to the memory consolidation unit, and update the corresponding synaptic connection increment in the sequence that triggers playback Make adjustments; The synaptic connection weight update and error feedback process is repeated until a preset number of iterations is reached, and the weight-enhanced subgraphs in the neuromorphic perception path are scanned. A weight threshold is preset, and paths with weights greater than the preset weight threshold are recorded as high-weight paths. For each high-weight path, the state-action pair it represents is extracted, and all state-action pairs are combined to form an action strategy generalization patch.

[0015] Preferably, the method for correcting the desired motion trajectory and action strategy of the robot includes: Receive robot execution feedback data in real time, preprocess the robot execution feedback data into a standardized execution state sequence, and construct an expected state sequence based on the current robot's expected motion trajectory and dynamic parameters. Compare the execution state sequence with the expected state sequence, calculate the robot's execution deviation through the error function, and simultaneously construct a deviation mapping function to infer the action adjustment amount that caused the execution deviation. The action adjustment amount and the action strategy generalization patch are jointly modeled to calculate the behavior correction vector, and the robot's behavior action is optimized according to the behavior correction vector. The optimized behavior action is embedded and represented using a sparse coding mechanism. The behavior correction vector is stored in a preset strategy memory buffer to correct the robot's expected motion trajectory and action strategy.

[0016] Compared with the prior art, the present invention has the following beneficial effects: Leveraging the sparse coding principles of the visual cortex, this system achieves spatiotemporal compression and sparse coding of multimodal environmental perception information, generating a sparse spike activation sequence that reflects the environmental state. This sequence not only reduces data dimensionality and redundancy but also enhances the temporal dynamic expression of information, facilitating efficient and accurate state inference and behavior generation in subsequent cognitive decision-making modules. The overall process mimics the sparse coding and spiking neuron firing mechanisms of the visual cortex. Combining the heterogeneous nature and spatiotemporal correlations of multimodal perceptual information, it achieves deep fusion and efficient compression of multimodal environmental information, significantly enhancing the system's perception and response capabilities to complex environments. Drawing on the receptive field mechanism of the visual cortex to construct a sparse dictionary, each frame of input signal activates only a small number of atomic channels, achieving sparse representation. This improves the discriminability and controllability of perceptual representation while reducing computational redundancy, making it suitable for resource-constrained devices. The system automatically adjusts the triggering difficulty based on the relative activation level of the current channel within the overall system, achieving dynamic inhibition and competition, effectively alleviating channel activation bias and more closely resembling the synaptic competition and inhibition mechanisms of biological neural systems.

[0017] By constructing action policy generalization patches for state-action pairs extracted from high-weight paths, the system can reuse existing experience in similar tasks, reducing relearning costs and enabling a certain degree of cross-scenario transferability. Automatically replaying state-action sequences and updating synaptic pathways during inter-execution intervals aligns policy evolution with the biological brain's memory consolidation process, helping to maintain the long-term availability of key experiences. Accurately simulating the impact of pre- and post-synaptic spike timing on synaptic weights enables dynamic enhancement / inhibition, improving the system's selectivity for task-related pathways and reducing interference. Incorporating action error feedback into the memory replay phase regulates update increments, ensuring that the learning direction is consistently aligned with task performance improvements and enhancing the interpretability and effectiveness of behavior correction. Using preset thresholds, key high-weight pathways are extracted from neuromorphic pathways to form a compact, sparsely activated action policy subgraph, reducing energy consumption and improving inference speed. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 This is a schematic diagram of the structure of a system for adaptively adjusting the behavior pattern of an embodied intelligent robot according to the present invention; Figure 2 A schematic flow chart of a method for constructing a neuromorphic sensing pathway provided by the present invention; Figure 3 Schematic diagram of the flow of a method for adaptively adjusting the behavior pattern of an embodied intelligent robot according to the present invention. DETAILED DESCRIPTION

[0019] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0020] Example 1 See also Figure 1 and Figure 2 As shown, embodiment 1 further illustrates a system for adaptively adjusting the behavior pattern of an embodied intelligent robot proposed by the present invention, including: With the continuous development of intelligent robotics, the ability of embodied intelligent robots to autonomously perceive and adjust their behavior in complex dynamic environments has become an increasingly research hotspot. Currently, adaptive adjustment systems for robotic behavior patterns generally face numerous challenges, which restrict their performance in complex application scenarios.

[0021] Multimodal perception technology, a crucial means for robots to understand their environment, encompasses multiple sensory channels, including vision, touch, hearing, and inertia. However, existing systems generally lack effective multimodal data alignment and normalized mapping mechanisms, leading to biases or loss in the fused information, which in turn affects the ability to accurately identify environmental states and respond in real time. Furthermore, multi-channel processing within neuromorphic perception mechanisms is prone to imbalances, with some channels experiencing sustained high activity and others remaining silent for extended periods. Existing systems lack dynamic threshold adjustment strategies based on competition and regulation, resulting in uneven perceptual responses and reduced information utilization efficiency.

[0022] Most state representation methods rely solely on static feature vectors, ignoring the temporal dynamics of sensory signals. This makes it difficult to capture dynamic behaviors such as sudden events and intensity changes, limiting the robot's ability to understand and quickly respond to sudden changes in complex environments. Furthermore, traditional feature encoding methods lack sparse activation capabilities, resulting in a large number of channels being activated for computation at any given moment. This leads to high system energy consumption and low computational efficiency, making them unsuitable for embedded or edge computing deployments with stringent energy efficiency requirements.

[0023] Existing robotic motion strategies often rely on offline training or preset rules, lacking the ability to adapt and migrate online in dynamically changing environments and task disturbances. During behavioral learning and strategy optimization, feedback signals are not fully and promptly utilized, resulting in low feedback utilization, which leads to delayed strategy adjustments and insufficient learning efficiency and stability. Currently widely used error backpropagation and reinforcement learning methods suffer from slow learning processes and poor stability. Furthermore, they suffer from high energy consumption and redundant activation of behavioral decision paths. The consumption of large amounts of non-critical computing resources leads to slow responses and limited overall system performance.

[0024] In order to effectively solve the above problems, the present invention proposes a system for adaptively adjusting the behavior pattern of an embodied intelligent robot, comprising: The neuromorphic perception module collects and integrates multimodal environmental perception information. Drawing on the sparse coding principle of the brain's visual cortex, it compresses the multimodal environmental perception information in time and space, generates an environmental state representation vector, and constructs a neuromorphic perception path. The cognitive constraint decision module uses a neural-symbolic hybrid architecture to generate action strategies based on the environmental state representation vector. Through the cognitive momentum optimization mechanism, the strategy inertia term is introduced to improve the update process of the action strategy parameters and obtain the behavior decision vector. The variable impedance execution module maps the action decision vector into multiple virtual muscle synergies through muscle synergy mapping rules, and calculates the expected motion trajectory of each execution joint of the robot. It introduces a nonlinear stiffness field mapping mechanism and combines it with the robot's contact feedback data to construct the Riemann stiffness tensor field of the action contact area and adjust the parameters of each execution joint. The cross-modal continuous learning module draws on the memory consolidation mechanism of the hippocampus and adopts the synaptic timing-dependent plasticity rule to adaptively update the synaptic connection weights in the neuromorphic perception pathway and generate action strategy generalization patches; The embodied behavior inversion module receives robot execution feedback data in real time, combines it with the action strategy generalization patch, dynamically optimizes the robot's behavior pattern, corrects the robot's expected motion trajectory and action strategy, and adaptively executes the corresponding task.

[0025] Methods for collecting and fusing multimodal environmental perception information include: Deploy a multimodal sensor array to collect multimodal environmental perception information in the environment where the intelligent robot is located. The sensors in the multimodal sensor array include visual sensors, auditory sensors, somatosensory modal sensors, temperature and humidity sensors, and light sensors; multimodal environmental perception information includes visual modal information, auditory modal information, somatosensory posture information, ambient temperature, ambient humidity, and light intensity.

[0026] Visual modality information includes RGB images, depth images, and event camera data; auditory modality information includes audio waveform data and sound source orientation data; and somatosensory posture information includes the robot's acceleration, angular velocity, and posture angle.

[0027] The generation method of the environment state representation vector includes: Normalize and align the multimodal environmental perception information, uniformly map it to the shared feature space, and obtain the multimodal environmental feature vector. In the actual production process, this process can be directly processed by PyTorch (tool) to obtain the expected effect of the data. For example, through PyTorch's built-in functions (such as torch.nn.functional.normalize function, which can be used for normalization) and custom multimodal fusion models (such as Transformer-based encoder function, which can be used for fusion), efficient modal alignment and feature mapping are achieved to ensure that different modalities (such as visual, auditory and somatosensory data) maintain semantic consistency and low-dimensional representation in the shared space, thereby improving the accuracy and efficiency of subsequent sparse coding; PyTorch can be processed on externally deployed devices (such as independent edge computing servers or GPU clusters) and then transmitted through external interfaces (such as REST The system is connected to the system through API, gRPC, or ROS message queues, and then transmitted to the system for use, reducing the system's computing power while improving the flexibility and scalability of the overall deployment and avoiding overloading the robot's hardware due to complex calculations. A sparse dictionary is constructed by drawing on the receptive field mechanism of the visual cortex. The sparse dictionary is used to sparsely approximate the multimodal environment feature vector as an atomic linear combination to obtain a sparse activation vector. Since the sparse pulse triggering mechanism depends on the threshold setting, an unreasonable threshold will lead to excessive pulses or loss of key information; a fixed threshold cannot adapt to input fluctuations and requires an adaptive adjustment mechanism; noise falsely triggers pulses, causing state vector pollution; insufficient sparsity cannot achieve effective compression, and excessive sparsity will result in the loss of key environmental dynamics.

[0028] For the activation value of each sparse pulse channel in the sparse activation vector, a dynamic threshold adaptive adjustment mechanism is designed. , preset The activation value of the sparse spike channel is , No. The pulse triggering threshold of a sparse pulse channel is According to the The activation value of the sparse spike channel The ratio of the total activation value of all sparse pulse channels is used to dynamically adjust the pulse trigger threshold of the sparse pulse channel; the pulse trigger threshold is ;in, Indicates the A sparse pulse channel at time The pulse trigger threshold; represents the index of each sparse spike channel, ; Indicates the index of the current moment; Indicates that all sparse pulse channels are at the current moment The sum of the activation values ​​of ; The adjustment factor for threshold adjustment controls the amplitude and speed of threshold change. A larger value means a faster adjustment. The adjustment factor can be pre-set based on experimental data analysis, historical data analysis, or expert experience. represents a sparse pulse channel index variable; represents the total number of sparse spike channels in the sparse activation vector; It should be noted that neurons in biological nervous systems have a competitive inhibition mechanism when perceiving information, that is, neurons with high activity inhibit neurons with low activity; this mechanism helps to improve sparsity, improve information discrimination ability, and save computing resources; the formula for adjusting the pulse-triggered threshold introduces a similar adjustment strategy, that is, the higher the activation value of the perception channel, the lower its threshold; the lower the activation value of the perception channel, the higher its threshold; it is easier to make active ones emit, and more difficult to make inactive ones emit.

[0029] The core idea is to compare the proportion of the current channel activation value in the overall activation level by normalization and dynamically adjust the threshold; if A ratio above the average indicates that the channel is dominant in the competition; a ratio below the average indicates that the channel is suppressed. The dynamic threshold can be continuously adjusted based on sensory input, ensuring the system's adaptability. The system always maintains a certain sparse firing trend, preventing all channels from being activated or silenced simultaneously. Stronger activation values ​​and lower thresholds lead to easier reactivation, consistent with the neuronal facilitation effect. The formula does not rely on complex derivations or backpropagation mechanisms, making it suitable for implementation in low-power embedded systems or neuromorphic hardware. It can be dynamically updated in real time, without relying on historical windows or large-scale caches.

[0030] In the preset time window Internal The activation values ​​of the sparse pulse channels are integrated and accumulated to obtain the A sparse pulse channel at time The cumulative activation value of ; ;in, Indicates the A sparse pulse channel at time The cumulative activation value of Indicates the At any time, a sparse pulse channel The instantaneous activation value of represents the time variable within the integration interval, ; The pulse emission discriminant function is used to determine whether the neuron is triggered to emit a pulse. When the cumulative activation amount of a sparse pulse channel is greater than or equal to the preset pulse trigger threshold, a pulse event is triggered, and a sparse pulse sequence is obtained; the sparse pulse sequence is input into the synaptic mapping function, and time smoothing, synaptic weighting and semantic aggregation are performed to generate an environmental state representation vector.

[0031] The pulse emission discrimination function is: ;in, Indicates the sparse pulse channels at the current moment Whether to generate a pulse when the binary state; Indicates the release of pulses; Indicates that no pulse is emitted; For example, in an automated inspection scenario at a petrochemical plant, a mobile robot with multimodal perception capabilities was deployed to identify potential safety risks, including gas leaks, temperature anomalies, and sudden noise changes. Equipped with an infrared thermal imaging sensor, a highly sensitive gas sensor, a microphone array, and a visual optical flow sensor, the robot's perception information needed to be uniformly encoded into sparse pulse signals, which were then used to generate environmental state representation vectors for use by the neuromorphic cognitive module.

[0032] Description of sparse pulse coding process: For the The sampling frequency of each modal channel (for example, a high-sensitivity methane gas sensor channel) is 200 Hz. In ms, the instantaneous activation values ​​of the continuous input are integrated to calculate the cumulative activation value of the channel at time t: , suppose that during a certain inspection process, A channel detects a rapidly increasing concentration of methane gas, and the instantaneous activation value rises rapidly to about 0.6 ppm / s within 0.5 seconds. The integral result within this time window is approximately: ; Set the pulse trigger threshold is 0.25, then according to the pulse emission discriminant function , at this time The channel is marked as active and triggers a sparse pulse. Multiple channels, such as thermal imaging detecting a temperature rise exceeding 2°C / s or an audio channel detecting a sudden change of more than 30dB, can also trigger pulses in parallel to form a sparse pulse sequence.

[0033] The problems existing in the prior art are solved: in the existing system, different types of sensing modalities often lack effective modal alignment and normalization mapping mechanisms; after multi-modal data fusion, information is lost or biased, which affects the accuracy and real-time performance of environment state recognition. In the existing neural sensing mechanism, the multi-channel processing method is prone to the phenomenon that some channels are always active and some channels are always dormant; there is a lack of threshold dynamic adjustment method based on competition and regulation mechanism, which leads to unbalanced system response and low information utilization rate; most state representation methods ignore the time dimension evolution characteristics of sensing signals and only use static feature vectors for expression; they cannot express dynamic characteristics such as event burst and intensity change, which limits the understanding and response ability of the system to complex scenes. Most traditional feature encoding methods do not have sparse activation ability; a large number of channels are calculated at each moment, which is high in energy consumption, low in efficiency and not suitable for embedded or edge deployment.

[0034] The beneficial effects of the prior art are that through the mechanism, the system realizes the spatio-temporal compression and sparse coding of multi-modal environment sensing information, and generates a sparse pulse activation sequence reflecting the environment state. This sequence not only reduces the data dimension and redundancy, but also enhances the time sequence dynamic expression ability of information, which is convenient for the subsequent cognitive decision module to efficiently and accurately perform state inference and behavior generation.

[0035] The overall process simulates the sparse coding and pulse neuron firing mechanism of the brain visual cortex, combines the heterogeneous characteristics and spatio-temporal correlation of multi-modal sensing information, realizes the deep fusion and efficient compression of multi-modal environment information, and significantly improves the perception and response ability of the system to complex environments. By referring to the receptive field mechanism of the visual cortex to construct a sparse dictionary, only a small number of atomic channels are activated for each input signal, realizing sparse representation; The discriminability and controllability of perception expression are improved, and the calculation redundancy is reduced, which is suitable for resource-constrained devices. The trigger difficulty can be automatically adjusted according to the relative activation degree of the current channel in the whole, realizing dynamic inhibition and competition, effectively alleviating the channel activation bias problem, and being closer to the synaptic competition and inhibition mechanism in the biological nervous system.

[0036] The construction method of the neuromorphic perception path comprises: The activated sparse pulse channel in the environment state representation vector is regarded as the output of the primary perception element, and through the index binding mechanism of the sparse coding dictionary, it is routed to the initial node of the neuromorphic perception path to form a multi-channel pulse trigger source of the input layer; Based on the principle of biological synaptic plasticity, a synaptic mapping function connection graph is constructed, and synaptic connection weights are defined for each activated sparse pulse channel according to the STDP rule; the synaptic propagation delay is modeled, and the synaptic propagation delay is set according to the timing alignment error between different modalities; A connection sparsity constraint is introduced, and a preset activation frequency threshold for the synaptic path is set. The activation frequency of any synaptic path is counted. If the activation frequency of the synaptic path is greater than the preset activation frequency threshold, the target task corresponding to the synaptic path is determined to be strongly correlated with the preset target task. Only synaptic pathways that are strongly related to the preset target task are retained to form a dynamic and sparse pulse propagation pathway network. Pulse propagation pathways are selected based on the synaptic mapping function connection map, and the WTA mechanism is used to inhibit non-dominant pathways. For the same type of target task, only propagation pathways with pulse response strength greater than the preset pulse response strength threshold are retained. By setting a dynamic activation window, only a limited number of m effective channels are activated in any time period to control temporal sparsity. A multi-level neuromorphic network is constructed to receive and propagate sparse pulse signals, complete state aggregation, and obtain the final neuromorphic perception path.

[0037] The methods for generating action strategies include: Inputting the environmental state representation vector into a multi-layer neural network, which can be any form of a feedforward network, a convolutional network, or a spiking neural network; outputting a neural feature representation vector through the multi-layer neural network; A symbolic knowledge base is pre-set, including action-effect rules, state transition rules, and action constraint boundary information. A symbolic parser is used to convert the symbolic knowledge base into a computable structure to form a symbolic rule template. The neural feature representation vector output by the neural network is semantically aligned with the symbolic rule template, and matching is completed through an interface mechanism. The matched symbolic rule fragments are structured and combined to form a set of candidate action paths. The set of matched candidate action paths is optimized to generate action strategies. A scoring function is used to evaluate the execution cost, expected benefit, and task fit of each candidate action path, and different candidate action paths are ranked. The candidate path with the highest score is selected as the current action strategy. If there are multiple equivalent candidate paths, a neural evaluator can be introduced for further discrimination to output the final action strategy.

[0038] The method for obtaining the behavior decision vector includes: To introduce the policy inertia term, a momentum term is constructed to record and accumulate the changing trend of the action policy parameters at each historical moment. This momentum term decays according to a preset ratio each time the action policy is updated and is superimposed on the current action policy gradient direction to perform temporal smoothing of the action policy evolution direction. The current gradient is weightedly fused with the momentum term of the previous moment, and the fused momentum term is used as the update direction of the action strategy parameters. The fused momentum term is used to update and adjust the action strategy parameters. This update and adjustment process is performed offline or asynchronously during the training phase of the policy network, so that the latest action strategy parameters are used in the next round of behavior generation. After the action strategy parameters are updated, the new environment state representation vector is input into the updated strategy network to generate a behavior decision vector, which includes the robot motion control parameters, action selection probability, control instruction encoding, adjustment feedback parameters and multimodal response output instructions.

[0039] The robot motion control parameters include position offset, posture angle, velocity, acceleration, joint angle and angular velocity; the action selection probability includes the behavior option probability vector and the strategy distribution function parameters; the control instruction encoding includes the robot behavior pattern label, action strategy ID and trigger signal; the adjustment feedback parameters include the expected reward value, action strategy confidence score and resource constraint parameters; the multimodal response output instructions include language output instructions, tactile drive parameters and expression drive parameters.

[0040] The method for obtaining the expected motion trajectory of each execution joint of the robot includes: The muscle synergy mapping rule is used to map the behavior decision vector into n virtual muscle synergists. Each muscle synergist corresponds to a group of virtual muscle units with biological synergy characteristics. The muscle synergy mapping rule is based on the physiological mechanism of human muscle control and predefines a group of muscle synergist combinations. Each combination controls the movement pattern of one or q joint movements. The virtual muscle synergists are activated and scheduled, and the expected tension output of each virtual muscle unit is calculated by combining the preset target action direction and the preset target force output. A mapping relationship is established between the robot body structure and the muscle synergists, and the expected tension output of each virtual muscle unit is applied to the corresponding pre-built bone-joint structure model. Based on the bone-joint structure model, the system mechanical response under muscle drive is solved to obtain the expected motion trajectory of each execution joint of the robot.

[0041] Methods for adjusting parameters of each execution joint include: The robot's contact feedback data includes the spatial position coordinates of the contact point, the contact normal component, the contact tangential component, the contact force magnitude, the contact force direction, and the duration of action. Based on the robot's contact feedback data, nonlinear stiffness modeling is performed on the robot's action contact area. The elastic response to external forces in different directions in space is described in the form of Riemann stiffness tensor field. In the contact space, a symmetrical positive stiffness tensor is assigned to each coordinate point, which describes the deformation resistance of the point in different directions. For any contact point, its stiffness response can be described as ;in, It represents the force response caused by unit deformation; Indicates at the contact point The stiffness tensor at ; represents the displacement vector; The tensor field has spatial inhomogeneity and directional correlation. The contact characteristics are mapped to stiffness responses by introducing a nonlinear mapping function. Based on the constructed Riemann stiffness tensor field, the parameters of each executive joint are adjusted. The influence range of each executive joint in the contact space and the corresponding contact area tensor field value are determined. For each executive joint, the adjustment coefficient is calculated according to its corresponding contact point stiffness tensor, and the expected motion trajectory is adjusted based on the adjustment coefficient.

[0042] Adjustment coefficient ;in, Represents a nonlinear mapping function; Indicates the The adjustment coefficient of each execution joint; Represents the modulus of the stiffness tensor, which is the length of the robot's execution joint Corresponding contact points At , the strength value of the constructed Riemann stiffness tensor field; Indicates at the contact point The stiffness tensor at ; Indicates the The spatial position vector of the contact point corresponding to the execution joint or action contact area; Indicates the The contact force vector of the area where the execution joint is located.

[0043] Methods for generating action policy generalization patches include: During the robot's intermission period (such as the system's "sleep" or low-activity state), the state-action sequence (including multimodal perception data such as vision, touch, and inertia, and corresponding behavioral instructions) within the h-period is internally replayed based on the hippocampal memory consolidation mechanism. For each sequence segment, the spiking neurons in the neuromorphic perception pathway are stimulated in turn, causing the corresponding synaptic pathways to generate sequential activation. During memory consolidation, presynaptic neurons The postsynaptic neuron fires a spike at time Release pulse; record ;in, Represents the firing time difference between the postsynaptic neuron and the presynaptic neuron; according to the synaptic timing-dependent plasticity rule, the weight of each synaptic connection is updated ;in, is the forward learning rate constant; is the reverse learning rate constant; and is the time decay constant; all activated synaptic connections are updated synchronously through the synaptic timing-dependent plasticity rule, so that the weight of the path that often sends the first pulse is enhanced and the weight of the lagging path is weakened; It should be noted that updating the weight of each synaptic connection is a classic and widely recognized bioinspired learning rule in neuromorphic computing and biological neural learning mechanisms. Its core principle is to adjust synaptic strength based on the difference in firing time between pre- and post-synaptic neurons, thereby simulating the learning process of a real neural system. Learning can be performed solely based on the timing differences in activity between neurons, making it suitable for unsupervised or minimally supervised learning scenarios. It can achieve adaptive optimization of synaptic weights without relying on external supervisory signals, effectively strengthening synaptic connections with causal activation relationships and suppressing redundant paths, thus providing a neural dynamics foundation for the self-correction and generalization of embodied intelligent robot behavior patterns.

[0044] After each update behavior is executed, the error between the current action strategy output and the expected behavior is calculated based on the execution results (such as end posture error, trajectory deviation, contact force change, etc.) Feedback the error to the memory consolidation unit, and update the corresponding synaptic connection increment in the sequence that triggers playback Make adjustments to ensure that the update direction is consistent with task performance improvements; The synaptic connection weight update and error feedback process is repeated until the preset number of iterations is reached. The weight-enhanced subgraphs (subnetworks) in the neuromorphic perception path are scanned. A weight threshold is preset, and paths with weights greater than the preset weight threshold are recorded as high-weight paths. For each high-weight path, the state-action pair it represents is extracted, and all state-action pairs are combined to form an action strategy generalization patch.

[0045] The following problems existing in the existing technology are solved: Traditional robot action strategies often rely on offline training or rule setting, cannot cope with perception changes and task disturbances in complex environments, and lack adaptability and migration capabilities. Existing behavior learning or strategy optimization often cannot use feedback signals from task execution to correct strategies in real time or within a suitable time window, and the feedback utilization rate is low. Most existing models use simple error backpropagation or reinforcement learning methods, which have low learning efficiency and poor stability. High energy consumption or redundant activation problems are prominent; in the existing robot behavior pattern adaptive adjustment system, the behavior decision path is in a large-scale activation state, resulting in a large amount of computing resources being consumed on non-critical paths, causing high energy consumption, response lag and computational redundancy during execution.

[0046] Advantages over existing technologies: By constructing action policy generalization patches for state-action pairs extracted from high-weight paths, the system can reuse existing experience in similar tasks, reducing relearning costs and enabling a certain degree of cross-scenario transfer capabilities. Automatically replaying state-action sequences and updating synaptic pathways during inter-execution intervals coordinates policy evolution with the biological brain's memory consolidation process, helping to maintain the long-term availability of key experiences. Accurately simulating the impact of the time difference between pre- and post-synaptic spikes on synaptic weights enables dynamic enhancement / inhibition, improving the system's selectivity for task-related pathways and reducing interference. Incorporating action error feedback into the memory replay phase adjusts the update increment to ensure that the learning direction is always consistent with task performance improvement, enhancing the interpretability and effectiveness of behavior correction. Using preset thresholds, key high-weight paths are extracted from neuromorphic pathways to form a compact, sparsely activated action policy subgraph, reducing energy consumption and improving inference speed.

[0047] Methods for correcting the robot's desired motion trajectory and action strategy include: Receive robot execution feedback data in real time. The robot execution feedback data includes multimodal feedback signals from various joint drive units, end effectors, vision / tactile / inertial sensors, etc., including actual joint angles and angular velocities, end effector trajectory paths, external touch forces / contact surface states, internal motor loads, currents, and temperature changes. Preprocess the robot execution feedback data into a standardized execution state sequence, and construct an expected state sequence based on the current robot's expected motion trajectory and dynamic parameters (target position, velocity curve, and contact posture). Compare the execution state sequence with the expected state sequence, calculate the robot's execution deviation through an error function, and construct a deviation mapping function to infer the motion adjustment amount that caused the execution deviation. The action adjustment amount and the action strategy generalization patch are jointly modeled, the behavior correction vector is calculated, and the robot's behavior is optimized according to the behavior correction vector. The optimized behavior is embedded and represented using a sparse coding mechanism to ensure that the corrected robot behavior does not cause abnormal activation or unexpected energy consumption surges. The behavior correction vector is stored in a preset strategy memory buffer to correct the robot's expected motion trajectory and action strategy.

[0048] The activation frequency threshold of the preset synaptic path is set by the staff. By collecting the activation frequencies of different synaptic paths, the average activation frequencies of multiple synaptic paths are taken as the activation frequency threshold of the preset synaptic path; similarly, the preset pulse response intensity threshold and the preset weight threshold are set.

[0049] In this embodiment, the system leverages the sparse coding principles of the visual cortex to achieve spatiotemporal compression and sparse encoding of multimodal environmental perception information, generating a sparse spike activation sequence that reflects the environmental state. This sequence not only reduces data dimensionality and redundancy but also enhances the temporal dynamic representation of information, facilitating efficient and accurate state inference and behavior generation in subsequent cognitive decision-making modules. The overall process mimics the sparse coding and spiking neuron firing mechanisms of the visual cortex. Combining the heterogeneous nature and spatiotemporal correlation of multimodal perceptual information, it achieves deep fusion and efficient compression of multimodal environmental information, significantly enhancing the system's perception and responsiveness to complex environments. By incorporating the receptive field mechanism of the visual cortex into the construction of a sparse dictionary, each frame of input signal activates only a small number of atomic channels, achieving sparse representation. This improves the discriminability and controllability of perceptual representation while reducing computational redundancy, making it suitable for resource-constrained devices. The system automatically adjusts the triggering difficulty based on the relative activation level of the current channel within the population, achieving dynamic inhibition and competition, effectively alleviating channel activation bias and more closely resembling the synaptic competition and inhibition mechanisms of biological neural systems.

[0050] By constructing action policy generalization patches for state-action pairs extracted from high-weight paths, the system can reuse existing experience in similar tasks, reducing relearning costs and enabling a certain degree of cross-scenario transferability. Automatically replaying state-action sequences and updating synaptic pathways during inter-execution intervals aligns policy evolution with the biological brain's memory consolidation process, helping to maintain the long-term availability of key experiences. Accurately simulating the impact of pre- and post-synaptic spike timing on synaptic weights enables dynamic enhancement / inhibition, improving the system's selectivity for task-related pathways and reducing interference. Incorporating action error feedback into the memory replay phase regulates update increments, ensuring that the learning direction is consistently aligned with task performance improvements and enhancing the interpretability and effectiveness of behavior correction. Using preset thresholds, key high-weight pathways are extracted from neuromorphic pathways to form a compact, sparsely activated action policy subgraph, reducing energy consumption and improving inference speed.

[0051] Example 2 See also Figure 3 As shown, for the parts not described in detail in this embodiment, please refer to the description of Example 1. A method for adaptively adjusting the behavior mode of an embodied intelligent robot is provided, comprising: S1. Collect and fuse multimodal environmental perception information, draw on the sparse coding principle of the brain's visual cortex, compress the multimodal environmental perception information in time and space, generate environmental state representation vectors, and construct a neuromorphic perception path; S2. Based on the environmental state representation vector, a neural-symbolic hybrid architecture is used to generate the action strategy. Through the cognitive momentum optimization mechanism, the strategy inertia term is introduced to improve the update process of the action strategy parameters and obtain the behavior decision vector. S3. Map the behavior decision vector into multiple virtual muscle synergies through muscle synergy mapping rules to calculate the expected motion trajectory of each robot's execution joints. Introduce a nonlinear stiffness field mapping mechanism, combine the robot's contact feedback data, construct the Riemann stiffness tensor field of the action contact area, and adjust the parameters of each execution joint. S4. Drawing on the memory consolidation mechanism of the hippocampus, we use synaptic timing-dependent plasticity rules to adaptively update the synaptic connection weights in the neuromorphic perception pathway and generate action strategy generalization patches. S5. Receive robot execution feedback data in real time, combine it with the action strategy generalization patch, dynamically optimize the robot behavior pattern, correct the robot's expected motion trajectory and action strategy, and adaptively execute the corresponding task.

[0052] Since the electronic device introduced in this embodiment is an electronic device used to implement the embodiment of this application based on an embodied intelligent robot behavior mode adaptive adjustment system, based on the embodiment of this application, a person skilled in the art can understand the specific implementation of the electronic device of this embodiment and its various variations, so how the electronic device implements the method in the embodiment of this application will not be described in detail here. As long as a person skilled in the art implements the electronic device used by the embodiment of this application based on an embodied intelligent robot behavior mode adaptive adjustment system, it falls within the scope of protection of this application.

[0053] The above formulas are all dimensionless and numerical calculations. The formulas are obtained by collecting a large amount of data and performing software simulation to obtain the most recent real situation. The preset parameters and thresholds in the formulas are set by technicians in this field according to actual conditions.

[0054] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiment. All technical solutions based on the principles of the present invention are within the scope of protection of the present invention. It should be noted that for users of ordinary skill in the art, various improvements and modifications that do not depart from the principles of the present invention should also be considered within the scope of protection of the present invention.

Claims

1. A system for adaptively adjusting the behavior pattern of an embodied intelligent robot, characterized in that: include: The neuromorphic perception module collects and integrates multimodal environmental perception information, uses the sparse coding principle of the brain's visual cortex to perform spatiotemporal compression, generates an environmental state representation vector, and constructs a neuromorphic perception path; The cognitive constraint decision module uses a neural-symbolic hybrid architecture to generate action strategies based on the environmental state representation vector. It improves the update process of action strategy parameters through the cognitive momentum optimization mechanism and obtains the behavior decision vector. The variable impedance execution module maps the behavior decision vector into multiple virtual muscle synergies through muscle synergy mapping rules, and calculates the expected motion trajectory of each execution joint of the robot. It also introduces a nonlinear stiffness field mapping mechanism to construct the Riemann stiffness tensor field of the action contact area and adjust the parameters of each execution joint. The cross-modal continuous learning module draws on the memory consolidation mechanism of the hippocampus and adopts the synaptic timing-dependent plasticity rule to adaptively update the synaptic connection weights and generate action strategy generalization patches; The embodied behavior inversion module receives robot execution feedback data in real time, combines it with the action strategy generalization patch, dynamically optimizes the robot's behavior pattern, corrects the robot's expected motion trajectory and action strategy, and adaptively executes the corresponding task.

2. The embodied intelligent robot behavior mode adaptive adjustment system according to claim 1, characterized in that: The method for collecting and fusing multimodal environmental perception information includes: Deploy a multimodal sensor array to collect multimodal environmental perception information in the environment where the intelligent robot is located.

3. The embodied intelligent robot behavior mode adaptive adjustment system according to claim 2, characterized in that: The method for generating the environmental state representation vector includes: The multimodal environment perception information is normalized and modally aligned to obtain a multimodal environment feature vector. A sparse dictionary is constructed to sparsely represent the multimodal environment feature vector as an atomic linear combination to obtain a sparse activation vector. A dynamic threshold adaptive adjustment mechanism is designed to dynamically adjust the pulse triggering threshold of the sparse pulse channel. The pulse emission discriminant function is used to determine whether the neuron is triggered to emit a pulse. When the cumulative activation amount of a sparse pulse channel is greater than or equal to the preset pulse trigger threshold, a pulse event is triggered, and a sparse pulse sequence is obtained; the sparse pulse sequence is input into the synaptic mapping function, and time smoothing, synaptic weighting and semantic aggregation are performed to generate an environmental state representation vector.

4. The embodied intelligent robot behavior mode adaptive adjustment system according to claim 3, characterized in that: The method for constructing the neuromorphic sensing path includes: The activated sparse pulse channels in the environmental state representation vector are regarded as the output of the primary sensor element, forming the multi-channel pulse trigger source of the input layer; synaptic propagation delay is modeled and set according to the timing alignment error between different modalities; The activation frequency of any synaptic path is counted. If the activation frequency of the synaptic path is greater than the preset activation frequency threshold of the synaptic path, it is determined that the target task corresponding to the synaptic path is strongly correlated with the preset target task; a multi-level neuromorphic network is constructed to complete state aggregation and obtain the final neuromorphic perception path.

5. The embodied intelligent robot behavior mode adaptive adjustment system according to claim 4, characterized in that: The method for generating the action strategy includes: The environmental state representation vector is input into a multi-layer neural network, and a neural feature representation vector is output; a symbolic knowledge base is preset, and the symbolic knowledge base is converted into a computable structure through a symbolic parser to form a symbolic rule template; matching is completed through an interface mechanism to form a set of candidate action paths; the matched candidate action path set is optimized and an action strategy is generated. If there are multiple equivalent candidate paths, a neural evaluator is introduced for further judgment and the final action strategy is output.

6. The embodied intelligent robot behavior mode adaptive adjustment system according to claim 5, characterized in that: The method for obtaining the behavior decision vector includes: A momentum term is constructed to record and accumulate the changing trends of the action strategy parameters at each historical moment, and to perform temporal smoothing of the evolution direction of the action strategy. The current gradient is weightedly fused with the momentum term of the previous moment, and the fused momentum term is used to update and adjust the action strategy parameters. After the action strategy parameters are updated, the new environment state representation vector is input into the updated strategy network to generate a behavior decision vector.

7. The embodied intelligent robot behavior mode adaptive adjustment system according to claim 6, characterized in that: The method for obtaining the expected motion trajectory of each execution joint of the robot includes: The muscle synergy mapping rule is used to map the behavior decision vector into n virtual muscle synergists. A set of muscle synergist combinations is predefined to construct a mapping relationship between the robot's body structure and the muscle synergists. The expected tension output of each virtual muscle unit is applied to the corresponding pre-built bone-joint structure model. The system mechanical response under muscle drive is solved to obtain the expected motion trajectory of each robot's execution joint.

8. The embodied intelligent robot behavior mode adaptive adjustment system according to claim 7, characterized in that: The method for adjusting the parameters of each execution joint includes: Based on the robot's contact feedback data, the nonlinear stiffness modeling of the robot's motion contact area is performed; within the contact space area, a symmetric positive stiffness tensor is assigned to each coordinate point; the contact characteristics are mapped to the stiffness response by introducing a nonlinear mapping function; the parameters of each execution joint are adjusted; for each execution joint, the adjustment coefficient is calculated according to the corresponding contact point stiffness tensor, and the desired motion trajectory is adjusted based on the adjustment coefficient.

9. The embodied intelligent robot behavior mode adaptive adjustment system according to claim 8, characterized in that: The method for generating an action strategy generalization patch includes: During the interval of the robot's task execution, the state-action sequence within the h period is replayed internally; the presynaptic neuron is The postsynaptic neuron fires a spike at time Send out pulses; after each update behavior is executed, calculate the error amount based on the execution result Feedback the error to the memory consolidation unit, and update the corresponding synaptic connection increment in the sequence that triggers playback Make adjustments; By scanning the weight-enhanced subgraphs in the neuromorphic perception path; presetting a weight threshold, the paths with weights greater than the preset weight threshold are recorded as high-weight paths; for each high-weight path, the state-action pair it represents is extracted, and all state-action pairs are combined to form an action strategy generalization patch.

10. The embodied intelligent robot behavior mode adaptive adjustment system according to claim 9, characterized in that: The method for correcting the desired motion trajectory and action strategy of the robot includes: Receive the robot's execution feedback data to form the expected state sequence; calculate the robot's execution deviation through the error function, and at the same time construct a deviation mapping function to infer the action adjustment amount that causes the execution deviation; jointly model the action adjustment amount and the action strategy generalization patch, calculate the behavior correction vector, and correct the robot's expected motion trajectory and action strategy.

Citation Information

Patent Citations

  • Quadruped robot motion control method and system based on reinforcement learning

    CN119849301A

  • Pre-estimation model training method, data processing method and device and electronic equipment

    CN118779686A

  • Enhancing perceptual data using large language models in environment reconstruction systems and applications

    CN119151006A

  • Upper limb rehabilitation training method based on emotion feedback, storage medium and product

    CN119763770A

  • Indoor robot navigation method based on multi-modal feature fusion

    CN120313600A

Cited By

  • Self-adaptive strategy optimization method and system for robot with body based on interactive feedback

    CN121223793A

  • Big reasoning model-based cooperative control method and system for intelligent agent with body

    CN121276986A

  • Wire harness assembly robot control method and system based on machine learning

    CN121315980A

  • A machine learning-based wire harness assembly robot control method and system

    CN121315980B

  • Humanoid robot multi-modal sensing fusion method based on dynamic sparse activation

    CN121552445A