A body-possessed intelligent robot behavior mode self-adaptive adjustment system
By employing sparse coding and cross-modal learning mechanisms, the problems of inconsistent multimodal data fusion and computational redundancy in the adaptive adjustment system of intelligent robot behavior patterns are solved, achieving efficient environmental perception and behavior optimization, and improving the system's adaptability and resource utilization efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-02
- Publication Date
- 2026-04-10
AI Technical Summary
In existing intelligent robot behavior pattern adaptive adjustment systems, multimodal data fusion lacks an effective modal alignment and normalization mapping mechanism, resulting in information loss or bias, unbalanced response of neural perception mechanism, lack of sparse activation capability, computational redundancy of traditional feature encoding methods, inability of action strategy to adapt to complex environments due to reliance on offline training, low feedback utilization, resulting in high energy consumption and response lag.
A neuromorphic perception module is used to sparsely encode multimodal environmental perception information, and sparse pulse activation sequences are generated by drawing on the sparse coding principle of the brain's visual cortex to construct a neuromorphic perception path; a cognitive constraint decision-making module generates action strategies through cognitive momentum optimization and introduces a strategy inertia term to adjust parameters; a variable impedance execution module is mapped to virtual muscle co-element and joint parameters are adjusted by combining Riemann stiffness field; a cross-modal continuous learning module updates synaptic weights through the hippocampal memory consolidation mechanism; and an embodied behavior inversion module optimizes behavior patterns in real time.
It achieves efficient compression and deep fusion of multimodal environmental information, improves perception and response capabilities, reduces computational redundancy, has cross-scene transfer capabilities, reduces energy consumption, and improves learning efficiency and response speed.
Smart Images

Figure CN120773064B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of robot intelligent control, more specifically, the present application relates to a kind of embodied intelligent robot behavior mode adaptive adjustment system. BACKGROUND
[0002] The patent with the patent publication number CN119849301A discloses a four-legged robot motion control method and system based on reinforcement learning. By constructing a four-legged robot simulation model, the motion behavior of the robot can be simulated and analyzed in a virtual environment, significantly reducing the development cost and cycle. The balance control problem is converted into a Markov decision process, and the state, action and reward functions are defined, which helps to achieve a more intelligent and adaptive balance strategy. The deterministic policy gradient algorithm is used to interact with the simulation model for training, and an efficient balance control strategy can be obtained. The robot behavior is adjusted in real time to maintain balance. The hierarchical control idea is used to divide the walking task into two sub-tasks: foot point selection and gait generation, which are realized through reinforcement learning, respectively, improving the walking flexibility and stability of the robot on irregular terrain. The performance of the robot can be safely and repeatedly evaluated on the simulation platform to further optimize the strategy.
[0003] The existing intelligent robot behavior mode adaptive adjustment system mainly has the following problems:
[0004] In the existing system, different types of perception modalities often lack effective modal alignment and normalization mapping mechanisms, resulting in information loss or bias after multi-modal data fusion, which affects the accuracy and real-time performance of environment state recognition. In the existing neural perception mechanism, the multi-channel processing method may cause some channels to be always active and some channels to be dormant for a long time. There is a lack of threshold dynamic adjustment method based on competition and regulation mechanism, which leads to unbalanced system response. Most state representation methods ignore the time dimension evolution characteristics of perception signals and only use static feature vectors for representation. They cannot express dynamic characteristics such as event burst and intensity change, limiting the system's understanding and response ability to complex scenes. Traditional feature encoding methods mostly do not have sparse activation ability, which leads to a large number of channels being calculated at each time, making them unsuitable for embedded or edge deployment.
[0005] Traditional robot action strategies often rely on offline training or rule setting and cannot adapt to changes in perception and task disturbances in complex environments, lacking adaptability and migration ability. Existing behavior learning or strategy optimization often cannot utilize feedback signals in task execution to correct the strategy in real time or within a suitable time window, resulting in low feedback utilization. High energy consumption or redundant activation problems are prominent. In the existing robot behavior mode adaptive adjustment system, the behavior decision path is in a wide range of active state, which leads to a large amount of computing resources being consumed on non-critical paths, causing increased energy consumption, response lag and computational redundancy in the execution process.
[0006] In view of this, the present application proposes a body-possessed intelligent robot behavior mode adaptive adjustment system to solve the above problems. SUMMARY
[0007] In order to overcome the above-mentioned defects of the prior art, in order to achieve the above-mentioned purpose, the present application provides the following technical scheme: a body-possessed intelligent robot behavior mode adaptive adjustment system, comprising:
[0008] A neuromorphic perception module collects and fuses multi-modal environment perception information, learns from the sparse coding principle of the brain visual cortex, performs spatio-temporal compression on the multi-modal environment perception information, generates an environment state representation vector, and constructs a neuromorphic perception path.
[0009] A cognitive constraint decision module generates an action strategy based on the environment state representation vector using a neural-symbol hybrid architecture; through a cognitive momentum optimization mechanism, a strategy inertia term is introduced to improve the update process of the action strategy parameters, and a behavior decision vector is obtained.
[0010] A variable impedance execution module maps the behavior decision vector to a multi-virtual muscle synergy element through a muscle synergy mapping rule, calculates the expected motion trajectory of each execution joint of the robot, introduces a nonlinear stiffness field mapping mechanism, combines the contact feedback data of the robot, constructs a Riemann stiffness tensor field of the action contact area, and adjusts the parameters of each execution joint.
[0011] A cross-modal continuous learning module learns from the hippocampal memory consolidation mechanism, uses synaptic timing-dependent plasticity rules to adaptively update the synaptic connection weights in the neuromorphic perception path, and generates an action strategy generalization patch.
[0012] A body-possessed behavior inversion module receives real-time robot execution feedback data, dynamically optimizes the robot behavior mode in combination with the action strategy generalization patch, corrects the expected motion trajectory and action strategy of the robot, and adaptively executes the corresponding task.
[0013] Preferably, the method of collecting and fusing multi-modal environment perception information comprises:
[0014] A multi-modal sensor array is deployed to collect multi-modal environment perception information in the environment where the body-possessed intelligent robot is located, and the sensors in the multi-modal sensor array include visual sensors, auditory sensors, somatosensory modality sensors, temperature and humidity sensors, and light sensors. The multi-modal environment perception information includes visual modality information, auditory modality information, somatosensory posture information, environmental temperature, environmental humidity, and light intensity.
[0015] Preferably, the method of generating the environment state representation vector comprises:
[0016] Multimodal environment perception information is normalized and modally aligned, and uniformly mapped to a shared feature space to obtain multimodal environment feature vectors; a sparse dictionary is constructed by drawing on the receptive field mechanism of the visual cortex; the multimodal environment feature vectors are sparsely approximated as a linear combination of atoms through the sparse dictionary to obtain sparse activation vectors;
[0017] For the activation value of each sparse impulse channel in the sparse activation vector, a dynamic threshold adaptive adjustment mechanism is designed at the current time. Preset number The activation value of each sparse pulse channel is , No. The pulse trigger threshold for each sparse pulse channel is According to the first Activation value of sparse pulse channels The pulse trigger threshold of the sparse pulse channel is dynamically adjusted by the ratio of the sum of the activation values of all sparse pulse channels.
[0018] Within the preset time window Internal to the first Integrating and accumulating the activation values of the sparse pulse channels, we obtain the th . A sparse pulse channel at time Cumulative activation value The impulse firing discriminant function determines whether a neuron is triggered to fire a impulse. When the cumulative activation of a sparse pulse channel is greater than or equal to a preset pulse trigger threshold, a pulse event is triggered, resulting in a sparse pulse sequence. The sparse pulse sequence is then input into a synaptic mapping function for time smoothing, synaptic weighting, and semantic aggregation to generate an environmental state representation vector.
[0019] Preferably, the method for constructing the neuromorphic perception pathway includes:
[0020] The activated sparse pulse channels in the environmental state representation vector are regarded as the outputs of the primary sensing elements, and are routed to the initial node of the neuromorphic sensing path through the index binding mechanism of the sparse coding dictionary, thus forming a multi-channel pulse trigger source for the input layer.
[0021] Based on the principle of biological synaptic plasticity, a synaptic mapping function connectivity map is constructed, and synaptic connection weights are defined for each activated sparse pulse channel according to the STDP rule. Synaptic propagation delay modeling is performed, and the synaptic propagation delay is set according to the temporal alignment error between different modes.
[0022] The connection sparsity constraint is introduced, a preset activation frequency threshold of a synaptic path is set, the activation frequency of any synaptic path is counted, and if the activation frequency of the synaptic path is greater than the preset activation frequency threshold of the synaptic path, it is determined that the target task corresponding to the synaptic path is strongly related to the preset target task;
[0023] Only the synaptic path strongly related to the preset target task is reserved to form a dynamically sparse pulse propagation path network; on the basis of the synaptic mapping function connection graph, the pulse propagation path is selected, the WTA mechanism is used to suppress the non-dominant path, and for the same type of target task, only the propagation path with a pulse response intensity greater than a preset pulse response intensity threshold is reserved;
[0024] By setting a dynamic activation window, only a limited number of effective channels are activated at any time period for time sparsity control; a multi-level neuromorphic network is constructed to receive and propagate sparse pulse signals, complete state aggregation, and obtain the final neuromorphic perception path.
[0025] Preferably, the action policy generation method comprises:
[0026] The environment state representation vector is input into a multi-layer neural network, which can be any one of a feedforward network, a convolutional network or a spiking neural network; and a neural feature representation vector is output by the multi-layer neural network;
[0027] A preset symbolic knowledge base is provided, the symbolic knowledge base comprises action-effect rules, state transition rules and action constraint boundary information; the symbolic knowledge base is converted into a computable structure by a symbolic parser to form a symbolic rule template; the neural feature representation vector output by the neural network is semantically aligned with the symbolic rule template, matching is completed through an interface mechanism, the matched symbolic rule fragments are structured and combined to form a candidate action path set;
[0028] The matched candidate action path set is optimized to generate an action policy; a scoring function is used to evaluate the execution cost, expected return and task fit degree of each candidate action path, different candidate action paths are sorted, and the candidate path with the highest score is selected as the current action policy; if there are multiple equivalent candidate paths, a neural evaluator can be introduced for further discrimination to output the final action policy.
[0029] Preferably, the behavior decision vector acquisition method comprises:
[0030] To introduce a strategy inertia term, a momentum term is constructed to record and accumulate the change trend of the action policy parameters at each time in history, and the momentum term is decayed by a preset proportion and superimposed with the current action policy gradient direction at each action policy update to perform time smoothing of the evolution direction of the action policy;
[0031] The current gradient is weighted and fused with the momentum item of the previous moment, and the fused momentum item is used as the update direction of the action policy parameter; the action policy parameter is updated and adjusted using the fused momentum item, and the update and adjustment process is performed in an offline or asynchronous manner in the training link of the policy network, so that the latest action policy parameter is used in the next round of behavior generation;
[0032] After the action policy parameter is updated, the new environment state representation vector is input into the updated policy network to generate a behavior decision vector, which includes a robot motion control parameter, an action selection probability, a control instruction code, an adjustment feedback parameter and a multi-modal response output instruction.
[0033] Preferably, the method for obtaining the expected motion trajectory of each actuating joint of the robot comprises:
[0034] The behavior decision vector is mapped into n virtual muscle synergy units using a muscle synergy mapping rule, each muscle synergy unit corresponding to a group of virtual muscle units with biological synergy characteristics; the muscle synergy mapping rule is constructed based on human muscle control physiological mechanisms, and a group of muscle synergy unit combinations are predefined, each combination controlling the motion mode of one or q joints;
[0035] The virtual muscle synergy units are activated and scheduled, and the expected tension output of each virtual muscle unit is calculated in combination with the preset target action direction and the preset target force output; a mapping relationship between the robot body structure and the muscle synergy units is constructed, and the expected tension output of each virtual muscle unit is applied to the corresponding pre-constructed skeletal-joint structure model; based on the skeletal-joint structure model, the system mechanics response under muscle driving is solved to obtain the expected motion trajectory of each actuating joint of the robot.
[0036] Preferably, the method for adjusting each actuating joint parameter comprises:
[0037] The robot contact feedback data includes the spatial position coordinates of the contact point, the contact normal component, the contact tangent component, the contact force size, the contact force direction and the action duration; the action contact area of the robot is nonlinearly stiffness modeled based on the robot contact feedback data;
[0038] The elastic response ability of the space in different directions to the external force is described in the form of a Riemann stiffness tensor field; a symmetric positive definite stiffness tensor is assigned to each coordinate point in the contact space area, which describes the deformation resistance ability of the point in different directions;
[0039] The tensor field has spatial non-uniformity and direction correlation, and a nonlinear mapping function is introduced to map the contact characteristics to the stiffness response; based on the constructed Riemann stiffness tensor field, the parameters of each execution joint are adjusted; the influence range of each execution joint in the contact space and the corresponding contact area tensor field value are determined; for each execution joint, the adjustment coefficient is calculated according to the corresponding contact point stiffness tensor, and the expected motion trajectory is adjusted based on the adjustment coefficient.
[0040] Preferably, the method for generating the action strategy generalization patch comprises:
[0041] During the intermittent period of the robot performing the task, the state-action sequence in the h period is internally replayed based on the hippocampal memory consolidation mechanism; for each sequence segment, the pulse neurons in the neuromorphic perception path are stimulated in turn, so that the corresponding synaptic pathway generates sequence activation;
[0042] During the memory consolidation process, the presynaptic neuron fires a pulse at time , and the postsynaptic neuron fires a pulse at time ; the memory ; wherein, represents the firing time difference between the postsynaptic neuron and the presynaptic neuron; according to the synaptic timing-dependent plasticity rule, the weight of each synaptic connection is updated;
[0043] After each update of the behavior execution, the error amount between the current action strategy output and the expected behavior is calculated according to the execution result; The error amount is fed back to the memory consolidation unit to adjust the incremental update amount of the corresponding synaptic connection in the sequence triggered by the playback;
[0044] The synaptic connection weight update and error feedback process is repeated, and the iteration is stopped after a preset number of iterations are reached, the subgraph with weight enhancement in the neuromorphic perception path is scanned, a preset weight threshold is set, and the path with weight greater than the preset weight threshold is recorded as a high weight path; for each high weight path, the state-action pair represented by it is extracted, and all state-action pairs are combined to form an action strategy generalization patch.
[0045] Preferably, the method for correcting the expected motion trajectory and the action strategy of the robot comprises:
[0046] Real-time receiving of robot execution feedback data, preprocessing of the robot execution feedback data into standardized execution state sequences, construction of expected state sequences based on the current expected action trajectory and dynamic parameters of the robot, comparison of the execution state sequences and the expected state sequences, calculation of the execution deviation of the robot through an error function, and construction of a deviation mapping function to deduce the action adjustment amount causing the execution deviation.
[0047] The action adjustment amount is combined with the action strategy generalization patch to model a behavior correction vector, the behavior of the robot is optimized according to the behavior correction vector, and the optimized behavior is embedded by using a sparse coding mechanism; the behavior correction vector is stored in a preset strategy memory buffer, and the expected motion trajectory of the robot and the action strategy are corrected.
[0048] Compared with the prior art, the present application has the following beneficial effects:
[0049] The present application realizes the space-time compression and sparse coding of multi-modal environment perception information through the principle of brain visual cortex sparse coding, and generates a sparse pulse activation sequence reflecting the state of the environment. This sequence not only reduces the data dimension and redundancy, but also enhances the time sequence dynamic expression ability of the information, facilitating the subsequent cognitive decision module to efficiently and accurately perform state inference and behavior generation. The overall process simulates the sparse coding and pulse neuron firing mechanism of the brain visual cortex, combines the heterogeneous characteristics and space-time correlation of multi-modal perception information, realizes the deep fusion and efficient compression of multi-modal environment information, and significantly improves the perception and response ability of the system to complex environments. By referring to the receptive field mechanism of the visual cortex to construct a sparse dictionary, only a small number of atomic channels are activated for each input signal, achieving sparse representation; the discriminability and controllability of perception expression are improved, while the computational redundancy is reduced, making it suitable for resource-constrained devices. The difficulty of triggering can be automatically adjusted according to the relative activation degree of the current channel in the whole, realizing dynamic inhibition and competition, effectively alleviating the channel activation bias problem, and being more close to the synaptic competition and inhibition mechanism in the biological nervous system.
[0050] By constructing an action strategy generalization patch from the state-action pairs extracted in the high-weight path, the system can reuse existing experience in similar tasks, reduce the cost of relearning, and have certain cross-scene migration ability. The replay of the state-action sequence and the update of the synaptic path are automatically performed during the execution interval, making the strategy evolution consistent with the memory consolidation process of the biological brain, which helps to maintain the long-term availability of key experiences. Precise simulation of the effect of the time difference between presynaptic / postsynaptic pulses on synaptic weights realizes dynamic enhancement / inhibition, improves the selectivity of the system to task-related paths, and reduces interference. The action error feedback is incorporated into the memory replay stage to adjust the update increment, ensuring that the learning direction is always consistent with the improvement of task performance, and strengthening the explainability and effectiveness of behavior correction. The preset threshold is used to extract key high-weight paths from the neuromorphic path to form a compact and sparse activated action strategy subgraph, reducing energy consumption and improving inference speed. BRIEF DESCRIPTION OF DRAWINGS
[0051] Figure 1 It is a structure schematic diagram of a body intelligent robot behavior mode adaptive adjustment system of the present application;
[0052] Figure 2 A flowchart of a construction method of a neuromorphic perception path provided by the present application is shown in the figure;
[0053] Figure 3 A flowchart of a behavior pattern adaptive adjustment method of a body-equipped intelligent robot of the present application is shown in the figure. DETAILED DESCRIPTION
[0054] The technical solutions in the embodiments of the present application will be clearly and completely described with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of the present application.
[0055] Embodiment one
[0056] Please refer to Figure 1 and Figure 2 Embodiment one further describes a behavior pattern adaptive adjustment system of a body-equipped intelligent robot proposed by the present application, which includes:
[0057] With the continuous development of intelligent robot technology, the autonomous perception and behavior adjustment capability of body-equipped intelligent robots in complex dynamic environments has increasingly become a research hotspot. Currently, the robot behavior pattern adaptive adjustment system generally faces many challenges, which restricts its performance in complex application scenarios.
[0058] Multi-modal perception technology, as an important means for robots to understand the environment, covers multiple sensing channels such as vision, touch, hearing, and inertia. However, existing systems generally lack effective multi-modal data alignment and normalization mapping mechanisms, resulting in deviations or information loss in the fused information, which further affects the accurate identification of the environment state and the real-time response capability. In addition, the multi-channel processing in the neuromorphic perception mechanism is prone to the imbalance phenomenon of continuous high activity of some channels and long-term dormancy of some channels. Existing systems lack threshold dynamic adjustment strategies based on competition and regulation, which leads to unbalanced perception response and reduces information utilization efficiency.
[0059] Most state representation methods only use static feature vectors, ignoring the time dynamic evolution characteristics of perception signals, making it difficult to capture dynamic behaviors such as event bursts and intensity changes, and limiting the understanding and rapid response of robots to sudden situations in complex environments. At the same time, traditional feature encoding methods lack sparse activation ability, resulting in a large number of channels being activated and calculated at any time, causing high system energy consumption and low computing efficiency, which is extremely unsuitable for embedded or edge computing deployment environments with strict energy efficiency requirements.
[0060] The robot action strategy in the prior art depends on offline training or preset rules, and lacks online adaptation and migration ability under dynamic environment and task disturbance conditions. In the behavior learning and strategy optimization process, the feedback signal cannot be fully and timely utilized, the feedback utilization rate is low, the strategy adjustment is lagged, and the learning efficiency and stability are insufficient. The current commonly used error back propagation and reinforcement learning method has defects such as slow learning process and poor stability, and the problems of high energy consumption and redundant activation of behavior decision path are prominent, and a large number of non-key computing resources are consumed, which leads to slow response and limited overall system performance.
[0061] In order to effectively solve the above problems, the application provides a body intelligent robot behavior mode adaptive adjustment system, which comprises:
[0062] The neuromorphic perception module collects and fuses multi-modal environment perception information, and generates an environment state representation vector by compressing the multi-modal environment perception information in space and time based on the sparse coding principle of the visual cortex of the brain, and constructs a neuromorphic perception path.
[0063] The cognitive constraint decision module generates an action strategy based on the environment state representation vector using a neural-symbol hybrid architecture; an action strategy parameter update process is improved by introducing a strategy inertia term through a cognitive momentum optimization mechanism to obtain a behavior decision vector.
[0064] The variable impedance execution module maps the behavior decision vector to multiple virtual muscle coordination units through muscle coordination mapping rules to solve the expected motion trajectory of each execution joint of the robot; a nonlinear stiffness field mapping mechanism is introduced to combine the contact feedback data of the robot to construct a Riemann stiffness tensor field of the action contact area, and adjust the parameters of each execution joint.
[0065] The cross-modal continuous learning module updates the synaptic connection weight in the neuromorphic perception path adaptively by using synaptic timing-dependent plasticity rules based on the hippocampal memory consolidation mechanism, and generates an action strategy generalization patch.
[0066] The embodied behavior inversion module receives the robot execution feedback data in real time, dynamically optimizes the robot behavior mode in combination with the action strategy generalization patch, corrects the expected motion trajectory and action strategy of the robot, and adaptively executes the corresponding task.
[0067] The method for collecting and fusing multi-modal environment perception information comprises:
[0068] A multi-modal sensor array is deployed to collect multi-modal environment perception information in the environment where the intelligent robot is located. The sensors in the multi-modal sensor array include visual sensors, auditory sensors, somatosensory modality sensors, temperature and humidity sensors, and illumination sensors. The multi-modal environment perception information includes visual modality information, auditory modality information, somatosensory posture information, environmental temperature, environmental humidity, and illumination intensity.
[0069] The visual modality information includes RGB images, depth images, and event camera data. The auditory modality information includes audio waveform data and sound source orientation data. The somatosensory posture information includes the robot's acceleration, angular velocity, and attitude angle.
[0070] The method for generating the environment state representation vector includes:
[0071] The multi-modal environment perception information is normalized and modality-aligned, and is uniformly mapped to a shared feature space to obtain a multi-modal environment feature vector. In actual production, the expected results of the data can be directly processed using PyTorch (a tool). For example, efficient modality alignment and feature mapping can be achieved by using built-in functions of PyTorch (such as the torch.nn.functional.normalize function, which can be used for normalization) and custom multi-modal fusion models (such as the Transformer-based encoder function, which can be used for fusion), to ensure that different modalities (such as visual, auditory, and somatosensory data) maintain semantic consistency and low-dimensional representation in the shared space, thereby improving the accuracy and efficiency of subsequent sparse coding. PyTorch can be processed externally on devices such as independent edge computing servers or GPU clusters, and then connected to the system through external interfaces such as REST API, gRPC, or ROS message queues, and then transmitted to the system for use, reducing the system's algorithmic load while improving the overall deployment flexibility and scalability, and avoiding overloading of the robot's hardware due to complex calculations. By drawing on the receptive field mechanism of the visual cortex, a sparse dictionary is constructed. The multi-modal environment feature vector is sparsely approximated as an atomic linear combination through the sparse dictionary, and a sparse activation vector is obtained.
[0072] Since the sparse pulse trigger mechanism relies on the setting of a threshold, an unreasonable threshold will result in too many pulses or loss of key information. A fixed threshold cannot adapt to input fluctuations and requires an adaptive adjustment mechanism. Noise can trigger pulses, causing state vector pollution. Insufficient sparsity cannot achieve effective compression, and excessive sparsity can lose key environmental dynamics.
[0073] For the activation value of each sparse pulse channel in the sparse activation vector, a dynamic threshold adaptive adjustment mechanism is designed. At the current time , the activation value of the th sparse pulse channel is , No. The pulse trigger threshold for each sparse pulse channel is According to the first Activation value of sparse pulse channels The pulse trigger threshold of the sparse pulse channel is dynamically adjusted based on the ratio of the sum of the activation values of all sparse pulse channels; the pulse trigger threshold is... ;in, Indicates the first A sparse pulse channel at time The pulse trigger threshold; This represents the index of each sparse pulse channel. ; Indicates the index at the current time; This indicates that all sparse pulse channels are at the current time. The sum of activation values; The adjustment factor represents the threshold adjustment and controls the magnitude and speed of threshold change. A larger value results in faster adjustment. The adjustment factor can be preset based on experimental data analysis, historical data analysis, or expert experience. Represents a sparse pulse channel index variable; This represents the total number of sparse pulse channels in the sparse activation vector;
[0074] It should be noted that neurons in biological nervous systems exhibit a competitive inhibition mechanism when sensing information, namely, highly active neurons inhibit less active neurons. This mechanism helps improve sparsity, enhance information discrimination ability, and conserve computational resources. The formula for adjusting the pulse trigger threshold introduces a similar adjustment strategy: the higher the activation value of the sensory channel, the lower its threshold; the lower the activation value of the sensory channel, the higher its threshold; thus, active neurons are easier to fire, while inactive neurons are more difficult to fire.
[0075] The core idea is to dynamically adjust the threshold by comparing the proportion of the current channel activation value in the overall activation level through normalization; if A ratio higher than average indicates that the channel is dominant in the competition; a ratio lower than average indicates that the channel is suppressed. The dynamic threshold can continuously adjust according to the perceived input, ensuring the system's adaptability. The system always maintains a certain sparse firing trend, avoiding the problem of all channels being activated simultaneously or completely silenced. The stronger the activation value, the lower the threshold, and the easier it is to reactivate, consistent with the neuron facilitation effect. The formula does not rely on complex differentiation or backpropagation mechanisms, making it suitable for implementation in low-power embedded systems or neuromorphic hardware; it can be dynamically updated in real time, without relying on historical windows or large-scale caches.
[0076] Within the preset time window Internal to the first Integrating and accumulating the activation values of the sparse pulse channels, we obtain the th . A sparse pulse channel at time Cumulative activation value ; ;in, Indicates the first A sparse pulse channel at time The cumulative activation value; Indicates the first A sparse pulse channel at any time The instantaneous activation value; Represents the time variable within the integration interval. ;
[0077] The impulse firing discriminant function determines whether a neuron has fired a impulse. When the cumulative activation of a sparse pulse channel is greater than or equal to a preset pulse trigger threshold, a pulse event is triggered, resulting in a sparse pulse sequence. The sparse pulse sequence is then input into a synaptic mapping function for temporal smoothing, synaptic weighting, and semantic aggregation to generate an environmental state representation vector.
[0078] The pulse firing discrimination function is: ;in, Indicates the first A sparse pulse channel at the current moment A binary state indicating whether a pulse is generated at any given time; Indicates the emission of a pulse; This indicates that no pulse will be emitted;
[0079] For example, in an automated inspection scenario at a petrochemical plant, a mobile robot with multimodal perception capabilities was deployed to identify potential safety risks, including gas leaks, abnormal temperatures, and sudden noise changes. This robot is equipped with an infrared thermal imaging sensor, a high-sensitivity gas sensor, a microphone array, and a visual optical flow sensor. Its perceived information needs to be uniformly encoded as sparse pulse signals and further used to generate an environmental state representation vector that can be used by the neuromorphic cognitive module.
[0080] Explanation of the sparse pulse coding process: For the first... Each modal channel (e.g., a high-sensitivity methane gas sensor channel) is sampled at a frequency of 200Hz, within each sliding time window. Within ms, integrate the instantaneous activation values of the continuous inputs to calculate the cumulative activation value of the channel at time t: Suppose that during a certain inspection process, the first... One channel detected a rapidly increasing concentration of methane gas, with the instantaneous activation value rising rapidly to approximately 0.6 ppm / s within 0.5 seconds. The integral result within this time window is approximately: ;
[0081] Set the pulse trigger threshold 0.25, then according to the pulse firing discrimination function , the first channel is marked as active, and a sparse pulse is triggered. Multiple channels such as thermal imaging detect temperature rise exceeding 2℃ / s, audio channel detects more than 30dB mutation, etc. can also trigger pulses in parallel to form a sparse pulse sequence.
[0082] The existing problems in the prior art are solved: in the existing system, different types of perception modalities often lack effective modal alignment and normalization mapping mechanism; leading to information loss or bias after multi-modal data fusion, thereby affecting the accuracy and real-time performance of environmental state recognition. In the existing neural perception mechanism, the multi-channel processing method is prone to the phenomenon that some channels are always active and some channels are always dormant; there is a lack of threshold dynamic adjustment method based on competition and regulation mechanism, resulting in unbalanced system response and low information utilization rate; most state representation methods ignore the time dimension evolution characteristics of the perception signal, and only use static feature vectors for expression; unable to express dynamic characteristics such as event burst and intensity change, limiting the system's understanding and response ability to complex scenes. Most traditional feature encoding methods do not have sparse activation ability; leading to a large number of channels computing at each moment, high energy consumption, low efficiency, and not suitable for embedded or edge deployment.
[0083] The beneficial effects of the prior art are as follows: through the mechanism, the system realizes the spatio-temporal compression and sparse coding of multi-modal environmental perception information, and generates a sparse pulse activation sequence reflecting the environmental state. This sequence not only reduces the data dimension and redundancy, but also enhances the time sequence dynamic expression ability of the information, facilitating the subsequent cognitive decision module to efficiently and accurately infer the state and generate behavior.
[0084] The overall process simulates the sparse coding and pulse neuron firing mechanism of the brain visual cortex, combines the heterogeneous characteristics and spatio-temporal correlation of multi-modal perception information, realizes the deep fusion and efficient compression of multi-modal environmental information, and significantly improves the system's perception and response ability to complex environments. By referring to the receptive field mechanism of the visual cortex to construct a sparse dictionary, only a small number of atomic channels are activated for each input signal, realizing sparse representation;
[0085] The discriminability and controllability of perception expression are improved, and the computational redundancy is reduced, which is suitable for resource-constrained devices. The trigger difficulty can be automatically adjusted according to the relative activation degree of the current channel in the whole, realizing dynamic inhibition and competition, effectively alleviating the channel activation bias problem, and being more close to the synaptic competition and inhibition mechanism in the biological nervous system.
[0086] The construction method of the neuromorphic perception path comprises:
[0087] The activated sparse pulse channel in the environmental state representation vector is regarded as the output of the primary perception element, and is routed to the initial node of the neuromorphic perception path through the index binding mechanism of the sparse coding dictionary, to constitute a multi-channel pulse trigger source of the input layer;
[0088] A synaptic mapping function connection graph is constructed based on the principle of biological synaptic plasticity, and synaptic connection weights are defined for each activated sparse pulse channel according to the STDP rule; synaptic propagation delay modeling is performed, and synaptic propagation delay is set according to the timing alignment error between different modalities;
[0089] A connection sparsity constraint is introduced, a preset activation frequency threshold of the synaptic path is set, the activation frequency of any synaptic path is counted, and if the activation frequency of the synaptic path is greater than the preset activation frequency threshold of the synaptic path, it is determined that the target task corresponding to the synaptic path is strongly related to the preset target task;
[0090] Only the synaptic path strongly related to the preset target task is retained to form a dynamically sparse pulse propagation path network; pulse propagation path selection is performed based on the synaptic mapping function connection graph, and a WTA mechanism is used to suppress non-dominant paths; for the same type of target task, only the propagation path with a pulse response intensity greater than a preset pulse response intensity threshold is retained;
[0091] By setting a dynamic activation window, only a limited number of m effective channels are activated at any time period, and time sparsity control is performed; a multi-level neuromorphic network is constructed to receive and propagate sparse pulse signals, complete state aggregation, and obtain the final neuromorphic perception path.
[0092] The method for generating an action policy comprises:
[0093] An environmental state representation vector is input into a multi-layer neural network, which can be any one of a feedforward network, a convolutional network, or a spiking neural network; a neural feature representation vector is output through the multi-layer neural network;
[0094] A symbol knowledge base is preset, which includes action-effect rules, state transition rules, and action constraint boundary information; the symbol knowledge base is converted into a computable structure through a symbol parser to form a symbol rule template; the neural feature representation vector output by the neural network is semantically aligned with the symbol rule template, and matching is completed through an interface mechanism; the matched symbol rule fragments are structured and combined to form a candidate action path set;
[0095] The matched candidate action path set is optimized, and an action policy is generated; a scoring function is used to evaluate the execution cost, expected return and task fitness of each candidate action path, and different candidate action paths are sorted, and the candidate path with the highest score is selected as the current action policy; if there are multiple equivalent candidate paths, a neural evaluator can be introduced for further discrimination, and the final action policy is output.
[0096] The method for obtaining the behavior decision vector comprises:
[0097] In order to introduce a strategy inertia term, a momentum term is constructed to record and accumulate the change trend of the action policy parameters at each time in history. The momentum term is attenuated by a preset proportion and superimposed with the current action policy gradient direction at each action policy update, and time smoothing of the evolution direction of the action policy is performed.
[0098] The current gradient and the momentum term at the previous time are weighted and fused, and the fused momentum term is used as the update direction of the action policy parameters; the fused momentum term is used to update and adjust the action policy parameters, and the update and adjustment process is performed in an offline or asynchronous manner in the training process of the policy network, so that the latest action policy parameters are used in the next round of behavior generation.
[0099] After updating the action policy parameters, the new environment state representation vector is input into the updated policy network to generate a behavior decision vector, which includes robot motion control parameters, action selection probability, control instruction code, adjustment feedback parameters and multi-modal response output instructions.
[0100] The robot motion control parameters include position offset, attitude angle, speed, acceleration, joint angle and angular velocity; the action selection probability includes a behavior option probability vector and a strategy distribution function parameter; the control instruction code includes a robot behavior mode label, an action policy ID and a trigger signal; the adjustment feedback parameters include an expected reward value, an action policy confidence score and a resource constraint parameter; and the multi-modal response output instructions include language output instructions, haptic driving parameters and expression driving parameters.
[0101] The method for obtaining the expected motion trajectory of each execution joint of the robot comprises:
[0102] The behavior decision vector is mapped to n virtual muscle synergy units using a muscle synergy mapping rule, each muscle synergy unit corresponding to a group of virtual muscle units with biological synergy characteristics; the muscle synergy mapping rule is constructed based on human muscle control physiological mechanisms, and a group of muscle synergy unit combinations are predefined, each combination controlling the motion mode of one or q joints;
[0103] The virtual muscle coordinators are activated and scheduled. Combined with the preset target motion direction and preset target force output, the expected tension output of each virtual muscle unit is calculated. The mapping relationship between the robot body structure and the muscle coordinators is constructed, and the expected tension output of each virtual muscle unit is applied to the corresponding pre-constructed skeletal-joint structure model. Based on the skeletal-joint structure model, the mechanical response of the system driven by the muscle is solved to obtain the expected motion trajectory of each joint of the robot.
[0104] Methods for adjusting the parameters of each joint include:
[0105] The robot contact feedback data includes the spatial coordinates of the contact point, the normal component of the contact, the tangential component of the contact, the magnitude of the contact force, the direction of the contact force, and the duration of the action; based on the robot contact feedback data, a nonlinear stiffness model is performed on the robot's motion contact area;
[0106] The elastic response to external forces in different directions in space is described using a Riemann stiffness tensor field. Within the contact space region, a symmetric positive definite stiffness tensor is assigned to each coordinate point, which describes the deformation resistance of that point in different directions. For any contact point, its stiffness response can be described as... ;in, This represents the force response caused by a unit deformation. Indicates at the contact point Stiffness tensor at the point; Represents the displacement vector;
[0107] The tensor field exhibits spatial non-uniformity and direction correlation. By introducing a nonlinear mapping function, contact characteristics are mapped to stiffness response. Based on the constructed Riemann stiffness tensor field, the parameters of each execution joint are adjusted. The influence range of each execution joint in the contact space and the corresponding tensor field value of the contact region are determined. For each execution joint, an adjustment coefficient is calculated based on its corresponding contact point stiffness tensor, and the desired motion trajectory is adjusted based on the adjustment coefficient.
[0108] Adjustment coefficient ;in, Represents a nonlinear mapping function; Indicates the first Adjustment coefficients for each execution joint; The magnitude of the stiffness tensor is used by the joints of a robot. corresponding contact points The strength value of the constructed Riemann stiffness tensor field at that location; Indicates at the contact point Stiffness tensor at the point; Indicates the first a spatial position vector of a contact point corresponding to the joint or action contact area; a contact force vector representing a region where the first execution joint is located.
[0109] The method for generating an action strategy generalization patch comprises:
[0110] During the intermittent period of the robot performing a task (such as a system "sleep" state or a low activity state), based on the hippocampal memory consolidation mechanism, the state-action sequence (including visual, tactile, inertial and other multi-modal perception data and corresponding behavior instructions) within h time is internally replayed; for each sequence segment, the pulse neurons in the neuromorphic perception path are stimulated in turn, so that the corresponding synaptic pathways generate sequence activation;
[0111] During the memory consolidation process, the presynaptic neuron fires a pulse at time , and the postsynaptic neuron fires a pulse at time ; the memory ; wherein, represents the firing time difference between the postsynaptic neuron and the presynaptic neuron; according to the synaptic timing-dependent plasticity rule, the weight of each synaptic connection is updated ; wherein, is a positive learning rate constant; is a negative learning rate constant; and are time decay constants; through the synaptic timing-dependent plasticity rule, all activated synaptic connections are updated synchronously, so that the path weight of the frequently first firing pulse is enhanced and the lag path weight is weakened;
[0112] It should be noted that updating the weight of each synaptic connection is a very classical and widely recognized biological heuristic learning rule in neuromorphic computing and biological neural learning mechanism, and the core is to adjust the synaptic strength according to the firing time difference between the presynaptic / postsynaptic neurons, thereby simulating the learning process of the real nervous system. Only the activity timing difference between neurons can be used for learning, which is suitable for unsupervised or less supervised learning situations. It can realize the adaptive optimization of synaptic weights without relying on external supervision signals, effectively strengthen the synaptic connections with causal activation relationship, and suppress redundant paths, thereby providing a neurodynamic basis for self-correction and generalization of embodied intelligent robot behavior patterns.
[0113] After each update of the behavior execution, the error amount between the current action strategy output and the expected behavior is calculated according to the execution result (such as end pose error, trajectory deviation, contact force change, etc.); the error amount is fed back to the memory consolidation unit, and the synaptic connection increment update amount Adjusting to ensure that the update direction is consistent with the improvement of task performance;
[0114] The synaptic connection weight update and error feedback process are repeated and iterated until a preset number of iterations is reached, and by scanning the subgraph (subnetwork) of the weight-enhanced neural morphological perception path, a preset weight threshold is set, and paths with weights greater than the preset weight threshold are recorded as high weight paths; for each high weight path, the state-action pair represented thereby is extracted, and all state-action pairs are combined to form an action strategy generalization patch.
[0115] The problems existing in the prior art are solved: the traditional robot action strategy often relies on offline training or rule setting, and cannot cope with perception changes and task disturbances in complex environments, lacking adaptability and migration ability. Existing behavior learning or strategy optimization often cannot utilize feedback signals in task execution in real time or within a suitable time window for strategy correction, and the feedback utilization rate is low. Most existing models use simple error backpropagation or reinforcement learning methods, which have low learning efficiency and poor stability. High energy consumption or redundant activation problems are prominent; in the existing robot behavior mode adaptive adjustment system, the behavior decision path is in a wide range of activation state, causing a large amount of computing resources to be consumed on non-critical paths, resulting in high energy consumption, response lag and calculation redundancy in the execution process.
[0116] The beneficial effects of the prior art are: by constructing an action strategy generalization patch from the state-action pairs extracted from the high weight path, the system can reuse existing experience in similar tasks, reducing the cost of relearning and having certain cross-scene migration ability. Replay of the state-action sequence and synaptic path update are automatically performed during the execution interval, making the strategy evolution consistent with the memory consolidation process of the biological brain, which helps to maintain the long-term availability of key experiences. Precise simulation of the effect of the time difference between presynaptic and postsynaptic pulses on synaptic weights, dynamic enhancement / inhibition, improves the selectivity of the system to task-related paths, and reduces interference. Incorporating action error feedback into the memory playback phase, adjusting the update increment, ensures that the learning direction is always consistent with the improvement of task performance, and strengthens the explainability and effectiveness of behavior correction. Using a preset threshold to extract key high weight paths from the neural morphological path, a compact and sparse action strategy subgraph is formed, reducing energy consumption and improving inference speed.
[0117] The method for correcting the expected motion trajectory and the action strategy of the robot comprises:
[0118] Real-time receive robot execution feedback data, robot execution feedback data includes multi-modal feedback signals from each joint drive unit, end effector, vision / tactile / inertial sensor, etc., including actual joint angle and angular velocity, end effector trajectory path, external touch force / contact surface state, internal motor load, current and temperature change; The robot execution feedback data is pre-processed into a standardized execution state sequence, and based on the expected action trajectory and dynamic parameters (target position, speed curve and contact posture) of the current robot, an expected state sequence is formed; Compare the execution state sequence with the expected state sequence, calculate the execution deviation of the robot through the error function, and at the same time build a deviation mapping function, and back-propagate the action adjustment amount that causes the execution deviation;
[0119] Joint modeling of action adjustment amount and action strategy generalization patch, calculation of behavior correction vector, optimization of robot behavior action according to behavior correction vector; and using sparse coding mechanism, embedding representation of optimized behavior action is carried out to ensure that the corrected robot behavior action will not cause abnormal activation or unexpected energy consumption surge; The behavior correction vector is stored in the preset strategy memory buffer, and the expected motion trajectory and action strategy of the robot are corrected.
[0120] The activation frequency threshold of the preset synapse path is set by the staff, and the average value of the activation frequencies of multiple synapse paths is taken as the activation frequency threshold of the preset synapse path by collecting the activation frequencies of different synapse paths; Similarly, the preset pulse response intensity threshold and the preset weight threshold are set.
[0121] In this embodiment, through the principle of sparse coding of the visual cortex of the brain, the system realizes the space-time compression and sparse coding of multi-modal environmental perception information, and generates a sparse pulse activation sequence reflecting the state of the environment. This sequence not only reduces the data dimension and redundancy, but also enhances the time sequence dynamic expression ability of information, facilitating the subsequent cognitive decision module to efficiently and accurately infer the state and generate behavior. The overall process simulates the sparse coding and pulse neuron firing mechanism of the visual cortex of the brain, combines the heterogeneous characteristics and space-time correlation of multi-modal perception information, realizes the deep fusion and efficient compression of multi-modal environmental information, and significantly improves the perception and response ability of the system to complex environments. By referring to the receptive field mechanism of the visual cortex, a sparse dictionary is constructed, so that only a small number of atomic channels are activated for each input signal, realizing sparse representation; improve the discriminability and controllability of perception expression, while reducing computational redundancy, suitable for resource-constrained devices. The difficulty of triggering can be automatically adjusted according to the relative activation degree of the current channel in the whole, realizing dynamic inhibition and competition, effectively alleviating the channel activation bias problem, and being more close to the synaptic competition and inhibition mechanism in the biological nervous system.
[0122] The state-action pairs extracted from the high-weight path are used to build an action policy generalization patch, so that the system can reuse existing experience in similar tasks, reduce the cost of relearning, and have certain cross-scene migration ability. The replay of the state-action sequence and the update of the synaptic path are automatically performed during the execution interval, which coordinates the policy evolution with the memory consolidation process of the biological brain, and helps to maintain the long-term availability of key experiences. Precise simulation of the effect of the time difference between pre- and post-pulse on synaptic weight, dynamic enhancement / inhibition, improved system selectivity for task-related paths, and reduced interference. Incorporate action error feedback into the memory playback phase to adjust the update increment, ensuring that the learning direction is always consistent with task performance improvement, and strengthening the explainability and effectiveness of behavior correction. Use a pre-set threshold to extract key high-weight paths from the neural morphological path to form a compact and sparse action policy subgraph, reducing energy consumption and improving inference speed.
[0123] Embodiment Two
[0124] Please refer to Figure 3 The embodiment does not describe some parts in detail, see the description of embodiment 1, providing a behavior pattern adaptive adjustment method for embodied intelligent robots, comprising:
[0125] S1, collect and fuse multi-modal environment perception information, learn from the sparse coding principle of the visual cortex of the brain, and perform spatio-temporal compression on the multi-modal environment perception information to generate an environment state representation vector and build a neural morphological perception path;
[0126] S2, based on the environment state representation vector, generate an action policy using a neural-symbol hybrid architecture; through a cognitive momentum optimization mechanism, introduce a policy inertia term to improve the update process of the action policy parameters, and obtain a behavior decision vector;
[0127] S3, map the behavior decision vector to multiple virtual muscle coordination units through muscle coordination mapping rules to solve the expected motion trajectory of each execution joint of the robot; introduce a nonlinear stiffness field mapping mechanism, combine the contact feedback data of the robot, construct a Riemannian stiffness tensor field of the action contact area, and adjust the parameters of each execution joint;
[0128] S4, learn from the hippocampal memory consolidation mechanism, and use synaptic time-dependent plasticity rules to adaptively update the synaptic connection weights in the neural morphological perception path and generate an action policy generalization patch;
[0129] S5, real-time receive robot execution feedback data, combine the action policy generalization patch, dynamically optimize the robot behavior pattern, correct the expected motion trajectory and action policy of the robot, and adaptively execute the corresponding task.
[0130] Since the electronic device introduced in the embodiment is the electronic device used in the embodiment of the application based on the adaptive adjustment system of the embodied intelligent robot behavior mode, the specific implementation of the electronic device of the embodiment and its various forms can be understood by those skilled in the art based on the adaptive adjustment system of the embodied intelligent robot behavior mode in the embodiment of the application, so the implementation of the electronic device in the method of the embodiment of the application will not be introduced in detail. As long as the electronic device used in the embodiment of the application based on the adaptive adjustment system of the embodied intelligent robot behavior mode is implemented by those skilled in the art, it belongs to the scope of protection of the application.
[0131] The above formulas are dimensionless values calculated, the formula is obtained by collecting a large amount of data to simulate the formula of the nearest real situation, and the preset parameters and threshold values in the formula are set by those skilled in the art according to the actual situation.
[0132] The above is only the preferred embodiment of the application, the protection scope of the application is not limited to the above-mentioned embodiments, and any technical solution belonging to the idea of the application is within the protection scope of the application. It should be noted that for ordinary technical users in the technical field, some improvements and decorations without departing from the principles of the application are also considered as the protection scope of the application.
Claims
1. A body-aware intelligent robot behavior pattern adaptive adjustment system, characterized in that, The method comprises the following steps: The neuromorphic perception module collects and fuses multi-modal environment perception information, performs spatio-temporal compression by referring to the sparse coding principle of the visual cortex of the brain, generates an environment state representation vector, and constructs a neuromorphic perception path. The method for generating the environment state representation vector comprises the following steps: The multi-modal environment perception information is normalized and modality-aligned to obtain a multi-modal environment feature vector; a sparse dictionary is constructed, and the multi-modal environment feature vector is sparsely approximated as an atomic linear combination to obtain a sparse activation vector; and a dynamic threshold self-adaptive adjustment mechanism is designed to dynamically adjust the pulse trigger threshold of the sparse pulse channel. The impulse firing discriminant function determines whether a neuron has fired a impulse. When the cumulative activation of a sparse pulse channel is greater than or equal to a preset pulse trigger threshold, a pulse event is triggered, resulting in a sparse pulse sequence. The sparse pulse sequence is then input into a synaptic mapping function for temporal smoothing, synaptic weighting, and semantic aggregation to generate an environmental state representation vector. The cognitive constraint decision module generates an action policy based on the environment state representation vector by using a neural-symbol hybrid architecture; and the updating process of the action policy parameters is improved by using a cognitive momentum optimization mechanism to obtain a behavior decision vector. The variable impedance execution module maps the behavior decision vector to a multi-virtual muscle synergy element by using a muscle synergy mapping rule, calculates the expected motion trajectory of each execution joint of the robot, introduces a nonlinear stiffness field mapping mechanism, constructs a Riemannian stiffness tensor field of the action contact area, and adjusts the parameters of each execution joint. The cross-modal continuous learning module generates an action policy generalization patch by using a synaptic time-dependent plasticity rule to adaptively update the synaptic connection weight by referring to the hippocampal memory consolidation mechanism. The embodied behavior inversion module receives the execution feedback data of the robot in real time, dynamically optimizes the behavior mode of the robot by combining the action policy generalization patch, corrects the expected motion trajectory and the action policy of the robot, and adaptively executes the corresponding task.
2. The body-aware robot behavior pattern adaptive adjustment system of claim 1, wherein, The method for collecting and fusing multi-modal environment perception information comprises the following steps: A multi-modal sensor array is deployed to collect multi-modal environment perception information in the environment in which the intelligent robot is located.
3. The body-aware robot behavior pattern adaptive adjustment system of claim 2, wherein, The method for constructing the neuromorphic perception path comprises the following steps: The activated sparse pulse channel in the environment state representation vector is regarded as the output of the primary perception element to constitute a multi-channel pulse trigger source of the input layer; a synaptic transmission delay model is established, and the synaptic transmission delay is set according to the timing alignment error between different modalities; The activation frequency of any synaptic path is counted, if the activation frequency of the synaptic path is greater than a preset activation frequency threshold of the synaptic path, it is determined that the target task corresponding to the synaptic path is strongly related to the preset target task; and a multi-level neuromorphic network is constructed to complete state aggregation to obtain the final neuromorphic perception path.
4. The body-aware robot behavior pattern adaptive adjustment system of claim 3, wherein, The method for generating the action policy comprises the following steps: The environment state representation vector is input into a multi-layer neural network to output a neural feature representation vector; a preset symbolic knowledge base is converted into a calculable structure by a symbolic parser to form a symbolic rule template; matching is completed by an interface mechanism to form a candidate action path set; the matched candidate action path set is optimized to generate an action policy, if there are multiple equivalent candidate paths, a neural evaluator is introduced for further discrimination to output the final action policy.
5. The embodied intelligent robot behavior pattern self-adaptive adjustment system according to claim 4, wherein, The method for obtaining the behavior decision vector comprises the following steps: A momentum term is constructed to record and accumulate the change trend of the action policy parameters at each time in history, and time smoothing of the evolution direction of the action policy is performed; the current gradient is weighted and fused with the momentum term at the previous time, and the momentum term after fusion is used to update and adjust the action policy parameters; after the action policy parameters are updated, the new environment state representation vector is input into the updated policy network to generate a behavior decision vector.
6. The body-aware robot behavior pattern adaptive adjustment system of claim 5, wherein, The method for obtaining the expected motion trajectory of each execution joint of the robot comprises: The behavior decision vector is mapped into n virtual muscle synergy units by using a muscle synergy mapping rule, a group of muscle synergy unit combinations is predefined, a mapping relationship between the robot body structure and the muscle synergy units is constructed, and the expected tension output of each virtual muscle unit is applied to the corresponding pre-constructed skeletal-joint structure model; the mechanical response of the system driven by the muscle is calculated to obtain the expected motion trajectory of each execution joint of the robot.
7. The body-aware robot behavior pattern adaptive adjustment system of claim 6, wherein, The method for adjusting the parameters of each execution joint comprises: Nonlinear stiffness modeling of the action contact area of the robot is performed based on the contact feedback data of the robot; a symmetric positive definite stiffness tensor is assigned to each coordinate point in the contact space area; the contact feature is mapped to the stiffness response by introducing a nonlinear mapping function; the parameters of each execution joint are adjusted; for each execution joint, an adjustment coefficient is calculated according to the corresponding contact point stiffness tensor, and the expected motion trajectory is adjusted based on the adjustment coefficient.
8. The body-aware intelligent robot behavior pattern adaptive adjustment system of claim 7, wherein, The method for generating the action policy generalization patch comprises: During the intermittent period of the robot performing the task, the state-action sequence in the h period is internally replayed; the presynaptic neuron fires a pulse at time , and the postsynaptic neuron fires a pulse at time ; after each update of the behavior execution, the error amount is calculated according to the execution result ; the error amount is fed back to the memory consolidation unit to adjust the incremental update amount of the corresponding synaptic connection in the sequence triggered by the replay ; The subgraph with weight enhancement in the neuromorphic perception path is scanned; a preset weight threshold is set, and the path with a weight greater than the preset weight threshold is recorded as a high-weight path; for each high-weight path, a state-action pair represented thereby is extracted, and all state-action pairs are combined to form an action policy generalization patch.
9. The body-aware intelligent robot behavior pattern adaptive adjustment system of claim 8, wherein, The method for correcting the expected motion trajectory and the action policy of the robot comprises: Robot execution feedback data is received to form an expected state sequence; an error function is used to calculate the execution deviation of the robot, and a deviation mapping function is constructed to back-propagate the action adjustment amount causing the execution deviation; the action adjustment amount and the action policy generalization patch are jointly modeled to calculate a behavior correction vector, and the expected motion trajectory and the action policy of the robot are corrected.
Citation Information
Patent Citations
Quadruped robot motion control method and system based on reinforcement learning
CN119849301A
Pre-estimation model training method, data processing method and device and electronic equipment
CN118779686A
Enhancing perceptual data using large language models in environment reconstruction systems and applications
CN119151006A