Sensitive active control method and system for body intelligence
By optimizing the multimodal data fusion of the embodied intelligence system through an adaptive coding module and a dynamic response mechanism, the problems of information distortion, delay, and interference are solved, thereby improving the system's perception and control accuracy and stability.
Patent Information
- Application Number
- CN202510905598.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-02
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-07-02
AI Technical Summary
Existing embodied intelligence systems suffer from information distortion, redundancy, dynamic response delay differences, signal synchronization time conflicts, and spectrum interference in multimodal data fusion, affecting the accuracy and stability of environmental perception and control.
An adaptive coding module and dynamic response mechanism are adopted to optimize the feature fusion path through spatiotemporal correlation index, monitor signal synchronization time and phase conflict in real time, dynamically adjust the execution order of control commands, and detect spectral interference in real time to generate a feature fusion path with minimal distortion and high efficiency.
It improves the accuracy of multimodal data fusion and the environmental adaptability of the system, ensures the stable operation of intelligent agents in complex environments, and achieves efficient perception and control coordination.
Smart Images

Figure CN120406171A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of embodied intelligence technology, and particularly to a sensing and active control method and system for embodied intelligence. Background Art
[0002] As a cutting-edge direction in the field of artificial intelligence, embodied intelligence aims to endow an intelligent agent with the ability to interact with the physical world through perception and action. Its core challenge lies in how to efficiently process multi-modal perception data and achieve precise coordination between perception and control. Traditional embodied intelligence systems generally face the following problems when processing multi-modal data such as vision, touch, and inertia: Existing heterogeneous data fusion frameworks usually adopt encoding modules with fixed structures and lack the ability to adaptively adjust to the characteristics of different perception modalities. For example, the image data output by a visual sensor has the characteristics of high spatial resolution and fast dynamic changes, while the tactile data collected by a flexible electronic skin pays more attention to the time series characteristics of local pressure distribution. The fixed encoding module cannot optimize for the spatio-temporal correlations of different modalities, resulting in information distortion or redundancy during the data fusion process, which affects the intelligent agent's real-time response ability to the environment.
[0003] During the transmission process of multi-modal data, the dynamic response delays of each modality are significantly different. For example, the update frequency of the acceleration data of an inertial measurement unit (IMU) can reach the kilohertz level, while the image frame rate of a visual sensor is usually dozens of hertz. This difference will cause time misalignment of different modality signals during the fusion process. Traditional synchronization mechanisms rely on fixed clock synchronization strategies and cannot adjust the signal synchronization moment according to the real-time dynamic response delay, thus triggering phase conflicts, resulting in chaotic control instructions, and affecting the coordination and stability of the intelligent agent's actions.
[0004] When generating the feature fusion path, existing technologies often adopt preset paths or simple heuristic search algorithms, without fully considering the spatio-temporal correlation metrics (such as signal fidelity, phase consistency, etc.) between encoding modules. This makes the feature fusion path may contain encoding modules with high distortion, resulting in attenuation or distortion of the perception features during transmission, seriously affecting the intelligent agent's accurate perception and understanding of environmental information.
[0005] When multiple perception modalities trigger control instructions simultaneously, traditional priority scheduling strategies cannot dynamically adjust the execution order of control instructions according to the real-time phase conflict results. For example, in an emergency obstacle avoidance scenario, the collision warning signal of tactile perception and the obstacle image signal of visual perception may conflict due to overlapping synchronization moments. If the response priority cannot be quickly determined and the control instructions cannot be adjusted, the intelligent agent will not be able to make a correct reaction in time, and even lead to safety accidents.
[0006] In a distributed heterogeneous sensor array, different sensors (such as the electromagnetic signals of vision sensors and the electronic signals of inertial measurement units) may generate spectral interference. The prior art lacks the ability of real-time detection of spectral interference and dynamic path planning. When the interference intensity exceeds the threshold, the feature fusion path cannot be adjusted in time, resulting in a decline in the quality of perceived data and affecting the reliability and robustness of the embodied intelligent system.
[0007] Therefore, how to improve the accuracy of multimodal data fusion, signal synchronization control and multimodal coordinated decision-making ability, and enhance the environmental adaptability and robustness of the embodied intelligent system is an urgent problem to be solved. Summary of the Invention
[0008] The purpose of the present invention is to provide a perception and active control method and system for embodied intelligence to solve the problems proposed in the above background technology.
[0009] To achieve the above purpose, the present invention provides the following technical solutions: A perception and active control method and system for embodied intelligence, the method includes: Obtain multimodal perception data, and determine the feature extraction unit and dynamic control center node of each perception modality from a preset heterogeneous data fusion framework, wherein the heterogeneous data fusion framework includes a plurality of adaptive coding modules located between the feature extraction unit and the control center node; For each perception modality, select a plurality of target coding modules with the optimal spatio-temporal correlation index based on the spatio-temporal correlation indexes of each of the adaptive coding modules relative to the feature extraction unit and the control center node, and generate a feature fusion path of the perception modality according to the plurality of target coding modules; For each perception modality, obtain the dynamic response delay of the perception modality, and determine the signal synchronization moment when the perception modality passes through a plurality of target coding modules based on the sub-layer coupling degree index between the target coding modules included in the feature fusion path and the dynamic response delay; Mark the signal synchronization moment for each of the target coding modules in the heterogeneous data fusion framework to obtain a real-time fusion topology map; For each of the target coding modules, detect the signal phase conflict according to the signal synchronization moment to obtain a phase conflict result; Based on the phase conflict result, perform multimodal control instruction planning to obtain a dynamic control coordination decision result.
[0010] Preferably, the spatio-temporal correlation index is a joint evaluation value of signal fidelity and phase consistency transmitted from the feature extraction unit to each adaptive coding module and from each adaptive coding module to the control center node for each adaptive coding module; the feature fusion path is an optimized path from the feature extraction unit to the control center node and is the minimum distortion link composed of the selected multiple target coding modules.
[0011] Preferably, the multi-modal control instruction planning based on the phase conflict result to obtain the dynamic control coordination decision result includes: When the phase conflict result indicates that there is an overlap in the multiple signal synchronization times included in the target coding module, the target coding module is determined as the phase conflict module; Obtain the overlapping signal synchronization times of the phase conflict module as the conflict window, and determine the dynamic response priorities of the multiple sensing modalities corresponding to the conflict window; Based on the dynamic response priorities, perform control instruction planning on the sensing modalities to obtain the dynamic control coordination decision result.
[0012] Preferably, the sensing modalities include a first response level modality and a second response level modality. The dynamic response priority of the first response level modality is in milliseconds, and the dynamic response priority of the second response level modality is in seconds. The control instruction planning for the sensing modalities based on the dynamic response priorities to obtain the dynamic control coordination decision result includes: Based on the order of milliseconds and seconds, determine the feature fusion path corresponding to the second response level modality; Based on the real-time fusion topology map, obtain the signal synchronization times of each target coding module of the second response level modality in the corresponding feature fusion path; Based on the signal synchronization times, combine the real-time fusion topology map to perform control instruction planning on the second response level modality to obtain the dynamic control coordination decision result.
[0013] Preferably, for each sensing modality, based on the spatio-temporal correlation indexes of each adaptive coding module relative to the feature extraction unit and the control center node, select multiple target coding modules with the optimal spatio-temporal correlation index, and generate the feature fusion path of the sensing modality according to the multiple target coding modules, including: For each sensing modality, determine multiple adjacent coding modules adjacent to the feature extraction unit; Calculate the spatio-temporal correlation metrics of each adjacent coding module with respect to the feature extraction unit and the control center node, and determine the first target coding module with the optimal spatio-temporal correlation metric from the multiple adjacent coding modules based on the multiple correlation metrics; Taking the first target coding module as the starting point, determine multiple adjacent coding modules adjacent to the target coding module, calculate the spatio-temporal correlation metrics of each adjacent coding module with respect to the feature extraction unit and the control center node, and determine the next target coding module with the optimal spatio-temporal correlation metric from the multiple adjacent coding modules adjacent to the target coding module based on the multiple correlation metrics; Repeat the steps of determining multiple adjacent coding modules adjacent to the target coding module, calculating the spatio-temporal correlation metrics of each adjacent coding module with respect to the feature extraction unit and the control center node, and determining the next target coding module with the optimal spatio-temporal correlation metric from the multiple adjacent coding modules adjacent to the target coding module until the next target coding module is determined to be the control center node, so as to obtain multiple target coding modules between the feature extraction unit and the control center node; Generate a feature fusion path for the perception modality based on the multiple target coding modules.
[0014] Preferably, the calculating the spatio-temporal correlation metrics of each adjacent coding module with respect to the feature extraction unit and the control center node includes: Calculate the first signal fidelity metric of each adjacent coding module from the feature extraction unit and the second phase consistency metric of each adjacent coding module from the control center node respectively; For each adjacent coding module, use the normalized product of the first signal fidelity metric and the second phase consistency metric as the spatio-temporal correlation metric of the adjacent coding module with respect to the feature extraction unit and the control center node.
[0015] Preferably, the determining the first target coding module with the optimal spatio-temporal correlation metric from the multiple adjacent coding modules based on the multiple correlation metrics includes: Store the multiple correlation metrics corresponding to the multiple adjacent coding modules into a correlation evaluation queue, and sort the correlation metrics in ascending order in the correlation evaluation queue to obtain a sorted sequence; Based on the sorted sequence, determine the adjacent coding module corresponding to the optimal correlation metric as the first target coding module.
[0016] Preferably, after planning multi-modal control instructions based on the phase conflict result to obtain a dynamic control coordination decision result, the method further includes: After each of the sensing modalities starts to execute according to the dynamic control coordination decision result, for each of the sensing modalities, the spectral interference intensity between the sensing modality and at least one adjacent sensing modality is detected in real time; When the spectral interference intensity between the sensing modality and the adjacent sensing modality exceeds a preset tolerance threshold, the adjacent sensing modality is used as an interference source, and a dynamic feature fusion path planning is performed on the sensing modality based on the interference source to obtain an updated fusion path of the sensing modality.
[0017] Preferably, the obtaining of the multi-modal sensing data includes: Collecting an original sensing stream through a distributed heterogeneous sensor array, where the distributed heterogeneous sensor array includes a vision sensor, a flexible electronic skin, and an inertial measurement unit; Performing spatio-temporal stamp marking on the original sensing stream to generate a modal data stream with time stamps; Inputting the modal data stream with time stamps into a preprocessing buffer of the heterogeneous data fusion framework, and performing modal feature decoupling operations through a programmable photonic chip to separate independent feature vectors of each sensing modality.
[0018] Preferably, the performing of the spatio-temporal stamp marking on the original sensing stream includes: Receiving an optical clock reference signal output by the programmable photonic chip and converting it into an electrical synchronous trigger signal; Deploying a time stamp embedding circuit at the sensor data acquisition end, and generating an absolute time tag accurate to the microsecond level according to the electrical synchronous trigger signal; Performing frame structure encapsulation on the absolute time tag and the original data packet collected by the corresponding sensor, where the time tag is used as a data frame header identifier; Transmitting the encapsulated data stream with time stamps to the preprocessing buffer through a high-speed serial interface.
[0019] Preferably, the present invention further includes a sensing and dynamic control system for embodied intelligence, and the system includes: A data acquisition and modal processing module, configured to acquire multi-modal sensing data, and determine a feature extraction unit and a dynamic control center node of each sensing modality from a preset heterogeneous data fusion framework based on the multi-modal sensing data, where the heterogeneous data fusion framework includes a plurality of adaptive coding modules located between the feature extraction unit and the control center node; The encoding module selection and path generation module is used to, for each perception modality, select multiple target encoding modules with the optimal spatio-temporal correlation metrics based on the spatio-temporal correlation metrics of each of the adaptive encoding modules relative to the feature extraction unit and the control center node, and generate a feature fusion path for the perception modality according to the multiple target encoding modules; The signal synchronization determination module is used to, for each perception modality, obtain the dynamic response delay of the perception modality, and determine the signal synchronization moment when the perception modality passes through multiple target encoding modules based on the sub-layer coupling degree metrics between the target encoding modules included in the feature fusion path and the dynamic response delay; The topology graph generation module is used to mark the signal synchronization moment for each of the target encoding modules in the heterogeneous data fusion framework to obtain a real-time fusion topology graph; The phase conflict detection module is used to, for each of the target encoding modules, detect signal phase conflicts according to the signal synchronization moment to obtain a phase conflict result; The motion control instruction planning module is used to perform multi-modal control instruction planning based on the phase conflict result to obtain a motion control coordination decision result.
[0020] Compared with the prior art, the beneficial effects of the present invention are: Through the feature extraction unit and the dynamic control center node of each perception modality in the preset heterogeneous data fusion framework, combined with multiple adaptive encoding modules, the optimal target encoding modules are dynamically selected according to the spatio-temporal correlation metrics (the joint evaluation value of signal fidelity and phase consistency), and a feature fusion path with the minimum distortion is generated. This mechanism can adaptively optimize according to the spatio-temporal characteristics of different perception modalities (such as vision, touch, inertia), effectively reducing information distortion and redundancy in the data fusion process. For example, in the processing of visual modality data, encoding modules that better preserve spatial features are preferentially selected to ensure the high-fidelity transmission of key information such as image edges and textures; in the touch modality, encoding modules that are sensitive to time series changes are mainly selected to accurately capture the dynamic features of pressure changes, thereby improving the overall accuracy of multi-modal data fusion.
[0021] Based on the sub-layer coupling degree index and dynamic response delay of the feature fusion path, accurately calculate the signal synchronization moment when each perceptual modality passes through the target encoding module, and mark the synchronization moment through the real-time fusion topology map to achieve real-time monitoring of the phases of multi-modal signals. When phase conflicts are detected (such as overlapping signal synchronization moments), dynamically adjust the execution order of control instructions according to the dynamic response priorities of perceptual modalities (classified into millisecond-level and second-level). For example, in an emergency obstacle avoidance scenario, the collision warning signal of tactile perception (millisecond-level priority) can trigger the braking instruction first, while the obstacle image analysis of the visual modality (second-level priority) synchronously conducts path planning to ensure that the intelligent agent makes a safe response in the shortest time and avoid action lags or conflicts caused by signal synchronization errors.
[0022] Dynamically identify the conflict modules and conflict windows through the phase conflict results, and combine the priority differences of perceptual modalities to achieve intelligent planning of control instructions. For the first-response-level modalities with millisecond-level responses (such as touch and inertia), ensure the real-time nature of their control instructions; for the second-response-level modalities with second-level responses (such as vision and audition), perform global path optimization based on the real-time fusion topology map. In addition, during the execution of instructions, continuously detect the spectrum interference intensity. When the interference exceeds the threshold, dynamically adjust the feature fusion path to avoid the interference source and ensure the stability of perceptual data. For example, when the visual sensor and the inertial measurement unit have abnormal data due to spectrum interference, the system can automatically switch to the backup encoding module path to maintain the continuity of perceptual data and improve the reliability of the embodied intelligent system in complex electromagnetic environments.
[0023] Collect multi-modal data through a distributed heterogeneous sensor array (including visual sensors, flexible electronic skins, and inertial measurement units), and use a programmable photonic chip to decouple modal features to achieve microsecond-level timestamp marking and high-precision synchronous triggering. This technical solution not only enhances the system's compatibility with multi-source heterogeneous data but also enables the embodied intelligent system to operate stably in a dynamically changing physical environment (such as strong interference and multi-modal data burst conflict scenarios) through dynamic feature fusion path planning and spectrum interference suppression mechanisms. For example, in an industrial robotic arm operation scenario, the system can process the target positioning data guided by vision and the force control data fed back by touch in real time, dynamically adjust the grasping posture, and at the same time avoid electromagnetic interference between sensors to ensure the stability and safety of the robotic arm in high-precision operations. Brief Description of the Drawings
[0024] Figure 1 It is the working principle diagram of the perception-action control method and system for embodied intelligence according to the present invention; Figure 2 It is the flowchart of multi-modal control instruction planning based on phase conflict results of the perception-action control method and system for embodied intelligence according to the present invention; Figure 3 Flow chart for generating the feature fusion path of the sensing and active control method and system for embodied intelligence according to the present invention; Figure 4 Flow chart for the original sensing stream spatio-temporal stamp marking and data stream encapsulation transmission of the sensing and active control method and system for embodied intelligence according to the present invention. Detailed implementation manners
[0025] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0026] Please refer to Figures 1 - 4 , a sensing and active control method and system for embodied intelligence according to the present invention, and the specific implementation steps are as follows: Obtain multimodal perception data: collect the original sensing stream through a distributed heterogeneous sensor array (including visual sensors, flexible electronic skin, and inertial measurement units), and mark the spatio-temporal stamps of the original sensing stream. The specific process is as follows: receive the optical clock reference signal output by the programmable photon chip and convert it into an electrical synchronous trigger signal, generate a microsecond-level absolute time tag through the timestamp embedding circuit, encapsulate the time tag and the original data packet into a timestamped modal data stream, transmit it to the preprocessing buffer through the high-speed serial interface, and perform modal feature decoupling operations by the programmable photon chip to separate the independent feature vectors of each perception modality. Determine the feature extraction unit and the dynamic control center node for each perception modality from the preset heterogeneous data fusion framework, and this framework includes multiple adaptive coding modules located between the feature extraction unit and the control center node.
[0027] Generate the feature fusion path: for each perception modality, based on the spatio-temporal correlation indexes of each adaptive coding module relative to the feature extraction unit and the control center node respectively, select multiple target coding modules with the optimal correlation indexes. The specific steps are as follows: first, determine multiple adjacent coding modules adjacent to the feature extraction unit, calculate the spatio-temporal correlation indexes of each adjacent module (that is, the normalized product of the first signal fidelity index and the second phase consistency index), store the indexes in the correlation evaluation queue and sort them in ascending order, and select the module corresponding to the optimal index as the first target coding module; starting from the first target module, repeat the above calculation and screening steps until the next target module is the control center node, so as to obtain multiple target coding modules between the feature extraction unit and the control center node, and generate the feature fusion path based on these modules, and this path is the minimum distortion link composed of the target coding modules.
[0028] Determine the signal synchronization moment: For each sensing modality, obtain its dynamic response delay, and based on the sub-layer coupling degree index and dynamic response delay between the target encoding modules included in the feature fusion path, determine the signal synchronization moment when the sensing modality passes through multiple target encoding modules.
[0029] Construct a real-time fusion topology graph: Mark the signal synchronization moment for each target encoding module in the heterogeneous data fusion framework to obtain a real-time fusion topology graph.
[0030] Detect phase conflicts: For each target encoding module, detect the signal phase conflicts according to the signal synchronization moment to obtain the phase conflict result.
[0031] Generate a dynamic control coordination decision result: Plan multi-modal control instructions based on the phase conflict result. If the phase conflict result indicates that there are overlapping signal synchronization moments included in a target encoding module, determine this module as a phase conflict module, obtain the overlapping signal synchronization moments as the conflict window, determine the dynamic response priorities of multiple sensing modalities corresponding to the conflict window (the sensing modalities include the first response-level modality in milliseconds and the second response-level modality in seconds), and plan control instructions based on the priorities to obtain a dynamic control coordination decision result.
[0032] The present invention will be further described below in conjunction with Embodiments 1 to 5: Embodiment 1: The specific calculation method of the spatio-temporal correlation index and the selection process of the target encoding module are as follows: For each adaptive encoding module in the heterogeneous data fusion framework corresponding to each sensing modality, it is necessary to quantitatively evaluate its spatio-temporal correlation with respect to the feature extraction unit and the control center node. The calculation of the spatio-temporal correlation index needs to be carried out from two dimensions: signal fidelity and phase consistency. For each adaptive encoding module (including the encoding module adjacent to the feature extraction unit in the initial stage and the encoding module adjacent to the current target encoding module in the subsequent iterative process), first calculate its first signal fidelity index with respect to the feature extraction unit. This index is comprehensively obtained by analyzing parameters such as energy loss, noise interference degree, and signal integrity during the transmission of the signal from the feature extraction unit to this encoding module, and is used to measure the signal retention ability during the transmission process. For example, the lower the energy loss, the smaller the noise interference, and the higher the signal integrity, the larger the value of the first signal fidelity index, indicating that this module has a stronger ability to retain the output signal of the feature extraction unit.
[0033] Calculate the second phase consistency index of the adaptive coding module from the control center node. This index is determined by evaluating parameters such as the phase offset, timing alignment error, and signal synchronization when the signal is transmitted from this module to the control center node, and is used to reflect the phase stability of the signal during transmission. The smaller the phase offset, the lower the timing alignment error, and the higher the signal synchronization, the larger the value of the second phase consistency index, indicating better phase consistency in the signal transmission between this module and the control center node.
[0034] After obtaining the first signal fidelity index and the second phase consistency index of each adjacent coding module, these two indexes need to be normalized. The purpose of normalization is to eliminate the influence of different indexes on the comprehensive evaluation result due to dimensional differences, making the indexes of different modules comparable. The specific normalization method is as follows: map the value of each index to the interval [0,1]. For example, through a linear transformation formula, subtract the minimum value of the index from the index value and then divide it by the difference between the maximum value and the minimum value to obtain the normalized value. After normalization, multiply the normalized value of the first signal fidelity index of each adjacent coding module by the normalized value of the second phase consistency index to obtain the spatio-temporal correlation index of this module relative to the feature extraction unit and the control center node. This index comprehensively reflects the comprehensive performance of the module in two dimensions of signal fidelity and phase consistency. The larger the value, the more suitable the module is as a node of the feature fusion path in the spatio-temporal dimension.
[0035] When selecting the first target coding module, for multiple adjacent coding modules adjacent to the feature extraction unit, store the spatio-temporal correlation indexes corresponding to each module in the correlation evaluation queue. The correlation evaluation queue is a data storage structure used to temporarily store the index values of each module for subsequent processing. After storage, sort the correlation indexes in the queue in ascending order to obtain a sorted sequence arranged from small to large in terms of value. The purpose of sorting is to clearly distinguish the advantages and disadvantages of the indexes of each module, facilitating the subsequent selection of the optimal module. Based on this sorted sequence, select the adjacent coding module corresponding to the largest (i.e., the optimal) correlation index value as the first target coding module. This process ensures that the starting node of the feature fusion path is the node with the best spatio-temporal correlation among all adjacent modules, which can minimize the initial distortion of feature transmission and lay a foundation for the optimization of subsequent paths.
[0036] After determining the first target encoding module, the process of iteratively selecting subsequent target encoding modules is entered. Taking the first target encoding module as a new starting point, multiple encoding modules that are adjacent to it and not selected into the path are determined (i.e., adjacent modules excluding the determined feature extraction units and the first target encoding module). For these new adjacent encoding modules, repeat the above steps of calculating the spatio-temporal correlation index: calculate the first signal fidelity index and the second phase consistency index for each module respectively, and multiply them after normalization to obtain the spatio-temporal correlation index. Subsequently, store these indices in the correlation evaluation queue, sort them in ascending order again, and select the module corresponding to the optimal index as the next target encoding module. This process continues, that is, each time taking the current target encoding module as the starting point, looking for adjacent unselected modules, calculating indices, sorting, and selecting the optimal module until the next target encoding module selected is the dynamic control center node.
[0037] During the entire iterative process, each step of the decision is based on the set of adjacent modules of the current node. Through the quantitative evaluation and sorting of the spatio-temporal correlation index, it is ensured that the target encoding module selected at each step is the current local optimal solution. Although this greedy algorithm strategy is a local optimal choice, through gradual accumulation, it can construct a path composed of multiple target encoding modules between the feature extraction unit and the control center node. The characteristic of this path is that each node has the optimal spatio-temporal correlation relative to its front and rear nodes, so that the signal transmission distortion of the entire path is minimized, forming a minimum distortion link from the feature extraction unit to the control center node, that is, the feature fusion path.
[0038] It should be noted that the calculation of the spatio-temporal correlation index does not depend on specific hardware parameters or experimental data, but is based on the basic principles and parameter definitions of signal transmission, and is obtained through logical derivation and mathematical operations (such as normalization, multiplication, etc.). In practical applications, the first signal fidelity index and the second phase consistency index of each module can be calculated by the sensor monitoring the relevant parameters (such as voltage fluctuation, phase difference, etc.) in the signal transmission process in real time to ensure the real-time and accuracy of the index. In addition, the sorting operation of the correlation evaluation queue can be implemented by software algorithms (such as quick sort, bubble sort, etc.) to ensure the sorting efficiency and accuracy, so as to meet the real-time requirements of the embodied intelligent system.
[0039] Embodiment 2: The processing flow of the phase conflict result and the control instruction planning method based on the dynamic response priority are as follows: After the system completes the construction of the real-time fusion topology map, it is necessary to detect phase conflicts for the signal synchronization times stored in each target coding module. The core logic of phase conflict detection is to determine whether there are overlapping signal synchronization times of multiple sensing modalities within the same target coding module. Specifically, each target coding module corresponds to a set of signal synchronization times in the real-time fusion topology map, and these times represent the time points when signals of different sensing modalities flow through the module. When the signal synchronization times of two or more sensing modalities have an intersection on the time axis (i.e., the time intervals overlap), it is determined that there is a phase conflict in this target coding module.
[0040] Once a phase conflict is detected, the system first marks this target coding module as a phase conflict module. Next, it is necessary to further determine the specific time range of the conflict and the sensing modalities involved. The specific operations are as follows: Extract all the overlapping signal synchronization times in this phase conflict module, and merge these overlapping time intervals into one or more conflict windows. Each conflict window corresponds to a specific time period, during which at least two sensing modalities' signals flow through the module simultaneously, which may cause signal interference or conflict in the execution of control instructions.
[0041] After determining the conflict window, the system needs to determine the dynamic response priorities for the multiple sensing modalities involved in the conflict window. According to the type of sensing modality, the system presets two types of response levels: the first response level modality and the second response level modality. Among them, the dynamic response priority of the first response level modality is in milliseconds, mainly including sensing tasks with extremely high real-time requirements, such as tactile feedback (e.g., instant force feedback when a robotic arm touches an obstacle), emergency obstacle avoidance (e.g., rapid response when a vision sensor detects a sudden obstacle), motion balance (e.g., real-time adjustment when an inertial measurement unit monitors an attitude imbalance), etc. Delays in these tasks may lead to system failures or safety risks. The dynamic response priority of the second response level modality is in seconds, including sensing tasks with relatively low real-time requirements, such as environmental modeling (e.g., constructing a three-dimensional scene model through a vision sensor), path planning (e.g., calculating a moving path based on a global map), target recognition (e.g., performing deep learning inference on image data), etc. These tasks allow a certain processing delay and can be gradually completed in non-emergency situations.
[0042] When planning control instructions based on dynamic response priorities, the system follows the order principle of "millisecond level prior to second level". The specific implementation steps are as follows: First, process the control instructions of the first response level mode to ensure that they obtain the priority execution right within the conflict window. For the first response level mode, the system generates control instructions to be executed immediately according to the feature fusion path and signal synchronization moment of this mode in the real-time fusion topology map, such as triggering the emergency braking of the robotic arm, adjusting the traveling direction of the mobile robot, etc. During the process of processing the first response level mode, the system will temporarily shelve the instructions of the second response level mode until the tasks of the first response level mode are completed or the conflict window ends.
[0043] For the second response level mode, after processing the first response level mode, the system obtains the signal synchronization moments of each target coding module in its corresponding feature fusion path based on the real-time fusion topology map. Specifically, the feature fusion path of each second response level mode is composed of multiple target coding modules connected in series, and each module corresponds to a signal synchronization moment. These moments form the time sequence of signal transmission of this mode. The system calculates the execution timing of the control instructions of this mode according to this time sequence and the connection relationship of each module in the real-time fusion topology map. For example, if the feature fusion path of a certain environmental modeling mode passes through module A, module B, and module C, and their signal synchronization moments are t1, t2, and t3 respectively, the control instructions need to trigger the processing operations of each module in the order of t1, t2, and t3 in turn to ensure that the data flows in order in the path and avoid feature distortion or calculation errors caused by chaotic timing.
[0044] When planning the control instructions of the second response level mode, the system also needs to consider the impact of the conflict window on it. If the conflict window covers a signal synchronization moment of this mode, the system will make adjustments according to the priority and conflict degree of this mode. For example, if the conflict window causes the signal synchronization moment of a certain second level mode to overlap with the millisecond level mode, the system will delay the corresponding operation of this second level mode until after the conflict window ends to ensure the real-time requirements of the millisecond level mode. This priority-driven timing scheduling mechanism can effectively avoid phase conflicts of multi-modal signals when sharing coding modules, ensure the immediate execution of high-priority tasks, and reasonably arrange the execution order of low-priority tasks to maintain the overall coordination and stability of the system.
[0045] It should be noted that the division of dynamic response priorities is not fixed and can be adjusted according to the specific application scenarios of the embodied intelligent system. For example, in the industrial robotic arm scenario, tactile feedback and emergency stop signals can be set to millisecond-level priorities, while workpiece recognition and trajectory planning can be set to second-level priorities; in the service robot scenario, obstacle detection (visual / inertial modality) can be set to millisecond-level, and user voice interaction (audio modality) can be set to second-level. This configurable priority mechanism enhances the flexibility and adaptability of the system, enabling it to adapt to different task requirements.
[0046] Throughout the entire process of phase conflict handling, the system forms a complete closed-loop feedback mechanism by continuously monitoring the signal synchronization moment, dynamically identifying conflict modules and conflict windows, and scheduling control instructions based on priorities. This mechanism does not rely on preset experimental data or empirical parameters, but makes autonomous decisions based on the real-time state of the perceptual modality and preset priority rules, ensuring that during the multi-modal data fusion process, phase conflict problems can be promptly detected and resolved, avoiding control instruction conflicts caused by signal timing chaos, thereby improving the response accuracy and task execution efficiency of the embodied intelligent system in a dynamic environment. Through this refined phase conflict management and priority-driven control instruction planning, the system can operate stably in complex multi-modal interaction scenarios and achieve efficient coordination between perception and control.
[0047] Embodiment 3: The generation process of the feature fusion path adopts an iterative optimization strategy, and the specific implementation method is as follows: For each perceptual modality, the construction of the feature fusion path starts from the feature extraction unit and ends at the dynamic control center node. The entire process determines the target encoding modules in the path through a node-by-node screening method. In the initial stage, multiple adjacent encoding modules directly adjacent to the feature extraction unit are determined, and these modules form the initial candidate set for path construction. There is a direct signal transmission link between each adjacent encoding module and the feature extraction unit, which is the first-level node for feature data to enter the fusion framework from the perceptual layer.
[0048] For each adjacent encoding module, calculate its spatio-temporal correlation index with respect to the feature extraction unit and the control center node. The calculation of the spatio-temporal correlation index needs to comprehensively consider signal fidelity and phase consistency. The specific steps are as follows: First, calculate the first signal fidelity index of each adjacent encoding module from the feature extraction unit , which is evaluated through parameters such as energy loss and noise interference during signal transmission. The larger the value, the stronger the signal fidelity ability; then calculate the second phase consistency index of each adjacent encoding module from the control center node , which is evaluated through parameters such as signal phase offset and timing alignment error. The larger the value, the higher the phase stability. Subsequently, for and perform normalization to obtain the normalized metrics and , and the normalization formula is:
[0049] where represents or , and are the minimum and maximum values of the corresponding metrics respectively, is the value of the normalized metric, and the value range is . The purpose of normalization is to eliminate the dimensional differences of different metrics and make the metrics of different modules comparable. Finally, multiply the two normalized metrics to obtain the spatio-temporal correlation metric , which comprehensively reflects the transmission performance of the module in the spatio-temporal dimension. The larger the value, the better the module.
[0050] Based on the calculated spatio-temporal correlation metric, select the first target coding module from multiple adjacent coding modules. The specific operation is to store the of each module into the correlation evaluation queue, sort the metrics in the queue in ascending order, and select the module corresponding to the largest metric value as the first target coding module. This module serves as the starting point of the feature fusion path to ensure that after the feature data is output from the feature extraction unit, it first passes through the node with the optimal transmission performance, reducing the initial transmission distortion.
[0051] Determine the first target coding module After that, enter the stage of iteratively screening subsequent nodes. Take as the current node, determine multiple unselected coding modules adjacent to it (i.e., excluding other adjacent modules of the feature extraction unit and the selected nodes ), and denote them as the set . For each module in the set , repeat the above calculation process: first calculate the first signal fidelity metric of the distance from the feature extraction unit and the second phase consistency metric of the distance from the control center node , obtain and through normalization, and then calculate the spatio-temporal correlation metric . Deposit into the correlation evaluation queue and sort it, and select the module corresponding to the optimal metric as the next target coding module and add it to the feature fusion path.
[0052] Repeat the above iterative process, that is, each time starting from the current target encoding module as the starting point, determine the set of its adjacent unselected modules , calculate for each module in the set , select the optimal module after sorting , until is the dynamic control center node . At this time, the target encoding module sequence between the feature extraction unit and the control center node is all determined, where
[0053] During the iterative process, the module screening at each step follows the "local optimal" principle, that is, only the adjacent modules of the current node are considered each time, and the module with the highest spatio-temporal correlation index among them is selected as the next-hop node. Although this greedy algorithm strategy does not guarantee global optimality, by optimizing node by node, a path composed of multiple "local optimal" nodes connected in series can be constructed between the feature extraction unit and the control center node. Since the selection of each node is based on the comprehensive evaluation of signal fidelity and phase consistency, this path can overall achieve the minimum distortion of signal transmission, so it is defined as the minimum distortion link.
[0054] It should be noted that the definition of adjacent modules during the iterative process is related to the topological structure of the heterogeneous data fusion framework. The adaptive encoding modules in the framework are connected in a graph structure, and each module can be connected to multiple other modules to form a mesh topology. Adjacent modules refer to the modules that have a direct link with the current node, and the transmission characteristics of the link (such as delay, bandwidth) will affect the calculation of signal fidelity and phase consistency indicators. For example, if there is high noise interference in the link between a certain module and the current node, its first signal fidelity index will decrease accordingly, resulting in the spatio-temporal correlation index dropping, so it will be preferentially excluded in the screening.
[0055] In addition, the sorting operation of the correlation evaluation queue can be implemented by software algorithms, such as quicksort or heapsort, to ensure that a large amount of module index data can be efficiently processed in a real-time system. The storage structure of the queue needs to support dynamic insertion and fast query to adapt to the continuously updated module set during the iterative process.
[0056] The feature fusion path generated by the above iterative optimization strategy has the following characteristics: First, each link of the path (the connection between adjacent target encoding modules) is the optimal choice for the current node, which can maximize the preservation of the integrity and temporal consistency of feature data; Second, the length of the path (i.e., the number of included modules) is jointly determined by the framework topology and module metrics. On the premise of ensuring transmission performance, the signal transmission path is shortened as much as possible to reduce latency; Third, the path has dynamic adaptability. When the module state in the framework changes (such as a module failure or transmission performance fluctuation), the iterative process can be re-triggered to generate a new feature fusion path to ensure the robustness of the system.
[0057] In an embodied intelligent system, different perception modalities (such as vision, touch, and inertia) can independently execute the above iterative process to generate their respective feature fusion paths. Due to the different characteristics of the perception data of each modality (such as a large amount of image data and high real-time requirements for touch data), the calculation parameters of the corresponding spatio-temporal correlation metrics (such as energy loss threshold, phase offset tolerance) can be adjusted according to the modality requirements, so as to achieve differential path optimization. For example, the vision modality can focus on the first signal fidelity metric (to retain image details), and the touch modality can focus on the second phase consistency metric (to ensure accurate timing of real-time responses).
[0058] The generation process of the feature fusion path constructs the minimum distortion link from the feature extraction unit to the control center node by iteratively calculating the spatio-temporal correlation metrics and screening the optimal modules node by node. This process does not rely on prior knowledge or preset path templates, but is dynamically optimized based on the real-time monitored signal transmission parameters, and can adapt to the complex topology of the heterogeneous data fusion framework and the differential requirements of multimodal data, providing key support for the efficient fusion of perception data and the precise planning of control instructions in the embodied intelligent system.
[0059] Example 4: After the multimodal control instruction planning is completed and each perception modality is triggered to execute, the system enters the real-time operation stage, and at this time, it is necessary to dynamically monitor and process the spectrum interference between the perception modalities.
[0060] Taking an embodied intelligent robotic arm with visual perception, tactile perception, and motion control as an example, its distributed heterogeneous sensor array includes a visual camera (collecting image data), a flexible electronic skin (collecting touch pressure data), and an inertial measurement unit (IMU, collecting attitude acceleration data). When the robotic arm performs an assembly task, the vision modality is responsible for identifying the workpiece position, the touch modality is responsible for the force feedback during grasping, and the inertial modality is responsible for monitoring the joint motion state of the robotic arm. The data of each modality is transmitted to the control center node through the feature fusion path, forming a collaborative control link of "visual guidance → tactile adjustment → inertial stabilization".
[0061] After the robotic arm starts to execute a task, the system uses a spectrum monitoring module to collect the signal spectrum data of each sensing modality in real time. For example, the image data of the visual modality is transmitted in the form of high-frequency signals, the pressure signal of the tactile modality is transmitted in the form of medium-frequency signals, and the motion data of the inertial modality is transmitted in the form of low-frequency signals. The spectrum monitoring module continuously calculates the signal spectrum overlap degree and power interference ratio between any two adjacent modalities to evaluate the spectrum interference intensity. Suppose that at a certain moment, the spectrum overlap degree of the signals of the visual modality and the tactile modality at a certain target coding module exceeds the preset tolerance threshold (such as the proportion of the overlapping frequency band exceeds 30%), the system determines that the tactile modality is the interference source and the visual modality is the interfered modality.
[0062] At this time, the system triggers the dynamic feature fusion path planning mechanism. First, locate the node sequence in the current feature fusion path of the interfered visual modality. Suppose the original path is "feature extraction unit → module A → module B → module C → control center node". Since spectrum interference occurs at module B, the system needs to exclude the interfering node (module B) in this path and regenerate the updated fusion path from the feature extraction unit to the control center node.
[0063] The process of re-planning the path is as follows: Starting from the feature extraction unit of the visual modality and ending at the control center node, traverse all non-interfered adaptive coding modules in the heterogeneous data fusion framework. For each candidate module, evaluate its spatio-temporal correlation with respect to the feature extraction unit and the control center node, specifically including the fidelity of signal transmission (such as energy loss, noise level) and phase consistency (such as timing alignment error). For example, module D is connected to the feature extraction unit, with relatively high signal fidelity but medium phase consistency; module E is adjacent to the control center node, with excellent phase consistency but relatively low signal fidelity. Through node-by-node comparison, the system preferentially selects modules that satisfy both high fidelity and phase consistency, and finally generates a new path "feature extraction unit → module D → module E → control center node".
[0064] During the path switching process, the system needs to ensure that the data transmission of the visual modality is not interrupted. By caching the unprocessed feature vectors and adjusting the timing synchronization signal, the signal synchronization moment of the new path is matched with the execution rhythm of the original control instruction. For example, the signal synchronization moment of module B in the original path is t = 10ms, and the synchronization moments of modules D and E in the new path are t = 8ms and t = 12ms respectively. The system delays the processing operation of module E to make the overall transmission delay consistent with the original path and avoid misalignment of control instruction execution.
[0065] In another scenario, during the movement of the robotic arm, if the inertial modality (high-frequency motion data) and the visual modality (mid-frequency image data) generate electromagnetic coupling interference between signal cables due to the close proximity of sensor deployment locations, and the spectral interference intensity exceeds the threshold. At this time, the inertial modality is identified as the interference source, and the visual modality needs to re-plan the path. The system detects that module F in the original path is a shared interference node, so it excludes module F and selects a detour path through module G and module H. Although the new path adds one hop node, it avoids the high-frequency interference area and ensures the feature integrity of the visual data.
[0066] The core logic of dynamic feature fusion path planning includes: Interference source localization: Determine the sensing modality that generates interference and its shared target coding module through spectral analysis. For example, when the signals of two modalities are transmitted within the same time period of a certain module and the spectral overlap exceeds the standard, this module is the interference node.
[0067] Path reconstruction scope: Only re-plan the path for the interfered sensing modality, and keep the paths of other modalities unchanged to reduce system resource consumption. In the above scenario where touch interferes with vision, the path of the inertial modality does not need to be adjusted.
[0068] Real-time guarantee: The path reconstruction algorithm needs to complete the evaluation of candidate modules within microseconds to ensure that the execution delay of control instructions is lower than the system response threshold. This is achieved through a hardware-accelerated metric calculation unit and a pre-stored module connection relationship table.
[0069] Multi-path redundancy: In the system initialization stage, pre-compute multiple alternative paths for each sensing modality. When interference is detected, the alternative paths can be directly called to further shorten the reconstruction time. For example, the visual modality can pre-store a "high-speed path" and an "anti-interference path" and dynamically switch according to the real-time interference situation.
[0070] During the precise assembly process of the robotic arm, multiple spectral interference events may occur. For example, the first interference occurs when grasping the workpiece (conflict between touch and vision modalities), and the system switches the visual path to give priority to ensuring the real-time nature of touch feedback; subsequent interference occurs when moving the robotic arm (conflict between inertia and touch modalities), and the system then adjusts the path of the inertial modality to ensure the stable transmission of touch signals. After each path adjustment, the system automatically updates the node markings in the real-time fusion topology map, enabling subsequent phase conflict detection and control instruction planning to be based on the latest path structure.
[0071] It should be noted that dynamic path planning does not change the priority attributes of the perception modalities. For example, even if the visual modality switches paths due to interference, its priority as the second response-level modality (second-level) is still lower than that of the tactile modality (millisecond-level), ensuring that the handling of emergency tasks is not affected by path adjustment. In addition, the valid nodes in the original path will be retained during the path reconstruction process. For example, if module A has been verified as an efficient transmission node before the interference occurs and it is not in the interference link, it will continue to be retained in the new path to avoid duplicate calculation overhead.
[0072] Through this dynamic adjustment mechanism based on real-time spectrum interference monitoring, the embodied intelligent system can maintain stable operation in complex electromagnetic environments or high-density data transmission scenarios. For example, in a factory environment with multi-robot collaborative operations, the sensing signals of different devices may cause cross-interference. Through dynamic path planning, each robot arm can autonomously avoid frequency band conflicts to ensure the precise execution of assembly tasks. Another example is when a service robot passes through a crowd, the high-frequency signals of the visual sensor may interfere with the wireless signals of surrounding devices. By switching the feature fusion path, the system can maintain the continuity of environmental perception and the stability of navigation control.
[0073] Identify the interference source through real-time spectrum monitoring. For the affected perception modalities, re-screen the target coding modules based on spatio-temporal correlation to generate an updated fusion path that bypasses the interference nodes. This process combines the physical deployment characteristics of the embodied intelligent agent (such as sensor location, link topology) and signal transmission characteristics (such as spectrum distribution, timing requirements) to achieve a closed-loop adaptive adjustment from interference detection to path reconstruction, improving the robustness and task execution reliability of the system in a dynamic interference environment.
[0074] Example 5: The acquisition and preprocessing process of multi-modal perception data is achieved through the collaborative work of a distributed heterogeneous sensor array, spatio-temporal stamp marking, a preprocessing buffer, and a programmable photonic chip.
[0075] Taking an embodied intelligent robot with environmental perception and autonomous operation capabilities as an example, the distributed heterogeneous sensor array it carries includes: a visual sensor (such as an RGB camera) deployed on the head for collecting environmental image data; a flexible electronic skin covering the surface of the robot arm for sensing contact pressure and object texture; and an inertial measurement unit (IMU) installed on the body for monitoring the robot's motion parameters such as acceleration and angular velocity. These sensors are deployed in a distributed architecture to collect multi-modal raw sensing streams such as vision, touch, and inertia, providing rich environmental and body state information for the system.
[0076] When the robot starts working, each sensor synchronously collects raw data. For example, the vision sensor captures environmental images at a rate of 30 frames per second, the flexible electronic skin samples the contact pressure value in real time, and the IMU outputs triaxial acceleration and angular velocity data at a high frequency. These raw sensing streams contain mixed signals of different modalities and need to be processed with spatio-temporal stamp marking and feature decoupling before being input into the heterogeneous data fusion framework.
[0077] The core of the spatio-temporal stamp marking process is to achieve the spatio-temporal alignment of multi-modal data. The programmable photon chip generates a high-precision optical clock reference signal, which is converted into an electrical synchronization trigger signal through an optoelectronic conversion module and transmitted to the data acquisition ends of each sensor. Taking the flexible electronic skin as an example, after the built-in timestamp embedding circuit receives the electrical synchronization trigger signal, it generates an absolute time tag accurate to the microsecond level. This time tag is frame-structured and encapsulated with the raw data packet collected by the pressure sensor, where the time tag serves as a fixed field (such as the first 16 bytes) at the head of the data frame, ensuring that each tactile data point carries the timestamp of the acquisition moment. Similarly, each frame of image data collected by the vision sensor and each motion data point collected by the IMU are encapsulated with microsecond-level time tags through the same mechanism, forming a modal data stream with timestamps.
[0078] The encapsulated data stream is transmitted to the preprocessing buffer through a high-speed serial interface (such as USB3.2 or Gigabit Ethernet). This buffer has a distributed storage structure and can receive multi-modal data streams simultaneously and temporarily store the data to be processed. Taking the vision modality as an example, after each frame of image data with a timestamp enters the buffer, the programmable photon chip performs modal feature decoupling operations. The photon chip utilizes the parallel processing characteristics of optical signals to perform spectral analysis on the image data stream, separating visual feature vectors such as brightness, color, and edges; for the tactile modality, the photon chip extracts tactile feature vectors such as pressure amplitude and acting direction through filtering and time-domain analysis; inertial modality data is separated into independent feature components such as translational acceleration and rotational angular velocity through frequency-domain conversion.
[0079] In specific implementations, the accuracy of spatio-temporal stamp marking directly affects the synchronization of multi-modal data. For example, when the robot grasps an object, the moment (t1) when the vision sensor detects the object contour, the moment (t2) when the flexible electronic skin senses the contact force, and the moment (t3) when the IMU monitors the change in the robotic arm posture need to be strictly aligned to ensure that subsequent fusion algorithms can accurately associate events of different modalities. Through the optical clock reference signal and microsecond-level time tags, the system can control the time deviation of multi-modal data within the microsecond level, meeting the timing requirements of real-time fusion.
[0080] The design of the preprocessing buffer needs to balance data throughput and low-latency characteristics. For example, a vision sensor generates hundreds of megabytes of data per second, and the buffer needs to have caching capabilities to avoid data congestion. The modal feature decoupling operation of the programmable photonic chip is based on parallel photonic links and can process multi-modal data streams simultaneously. For example, when decoupling the edge features of a visual image, the tactile pressure signal can be denoised synchronously, significantly improving the efficiency of data preprocessing.
[0081] Taking the robot obstacle avoidance scenario as an example, when the vision sensor detects an obstacle ahead, the original image data flows through the timestamp marking and enters the buffer. The photonic chip quickly extracts the contour and distance features of the obstacle; at the same time, the inertial data of the IMU is decoupled into the translational speed and rotational angle features of the robot. These independent feature vectors are synchronously transmitted to the heterogeneous data fusion framework for subsequent path planning and motion control. If feature decoupling is not performed, the mixed raw data may cause the fusion algorithm to misassociate visual noise with inertial interference, resulting in control instruction deviation.
[0082] The processing of tactile data from flexible electronic skin also demonstrates another application value of timestamp marking. When the robotic arm touches an object, the tactile sensors at different positions generate pressure signals in spatial order. The time tag not only records the signal acquisition time but also implicitly contains the spatial position information of the sensor (through the predefined mapping relationship during deployment). By analyzing the timestamp sequence, the system can restore the spatial distribution and timing of the touch event, assisting in judging the shape and contact direction of the object.
[0083] It is worth noting that the preprocessing buffer and the programmable photonic chip are connected by a high-speed bus to ensure that the data transmission delay is below the microsecond level. The modal feature decoupling algorithm of the photonic chip can be dynamically adjusted through firmware upgrades. For example, optimizing the visual feature extraction parameters in different lighting environments or enhancing the sensitivity threshold of tactile signals in high-precision operation scenarios.
[0084] In the scenario of multi-sensor collaborative work, timestamp marking can also be used to detect sensor failures. For example, if the timestamp of a certain vision sensor shows jumps or disordered arrangements, the system can determine that the sensor communication is abnormal and automatically switch to a redundant sensor or trigger a fault alarm. This health monitoring mechanism based on time tags enhances the reliability of the system.
[0085] Multimodal raw data is collected through a distributed sensor array. The optical clock synchronization mechanism is used to generate microsecond-level time tags and encapsulate them into data frames, which are then transmitted to the buffer through a high-speed interface. The programmable photonic chip performs feature decoupling and outputs independent feature vectors for each modality. This process realizes the spatio-temporal alignment and pure separation of multi-source data, providing a basis for subsequent processes such as feature fusion path generation and signal synchronization moment calculation. In the practical application of embodied intelligent robots, this process ensures the precise coordination of multimodal data such as visual guidance, tactile feedback, and inertial navigation, enabling the robot to perform refined operation tasks in complex environments, such as assembling parts, grasping fragile objects, or dynamically avoiding obstacles. Through the time synchronization design at the hardware level and the parallel processing ability of the photonic chip, the system lays an efficient and reliable perception foundation in the data acquisition and preprocessing stage.
[0086] It should be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device.
[0087] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. An embodied intelligence-oriented sensing and active control method, characterized in that, The method includes: Obtaining multi-modal perception data, and determining a feature extraction unit and a dynamic control center node for each perception modality from a preset heterogeneous data fusion framework, where the heterogeneous data fusion framework includes a plurality of adaptive coding modules located between the feature extraction unit and the control center node; For each perception modality, based on the spatio-temporal correlation indexes of each of the adaptive coding modules with respect to the feature extraction unit and the control center node, selecting a plurality of target coding modules with the optimal spatio-temporal correlation indexes, and generating a feature fusion path for the perception modality according to the plurality of target coding modules; For each perception modality, obtaining the dynamic response delay of the perception modality, and determining the signal synchronization moment when the perception modality passes through a plurality of target coding modules based on the sub-layer coupling degree index between the target coding modules included in the feature fusion path and the dynamic response delay; Marking the signal synchronization moment for each of the target coding modules in the heterogeneous data fusion framework to obtain a real-time fusion topology graph; For each of the target coding modules, detecting signal phase conflicts according to the signal synchronization moment to obtain a phase conflict result; Performing multi-modal control instruction planning based on the phase conflict result to obtain a dynamic control coordination decision result.
2. The sensor-actuator control method for embodied intelligence according to claim 1, wherein The spatio-temporal correlation index is a joint evaluation value of signal fidelity and phase consistency for each adaptive coding module, from the feature extraction unit to each of the adaptive coding modules, and from each of the adaptive coding modules to the control center node; The feature fusion path is an optimized path from the feature extraction unit to the control center node, and is the minimum distortion link composed of the selected plurality of target coding modules.
3. The sensorimotor control method for embodied intelligence according to claim 1, wherein The performing multi-modal control instruction planning based on the phase conflict result to obtain a dynamic control coordination decision result includes: When the phase conflict result indicates that there is an overlap in the plurality of signal synchronization moments included in the target coding module, determining the target coding module as a phase conflict module; Obtaining the overlapping signal synchronization moments of the phase conflict module as a conflict window, and determining the dynamic response priorities of the plurality of perception modalities corresponding to the conflict window; Performing control instruction planning on the perception modalities based on the dynamic response priorities to obtain a dynamic control coordination decision result.
4. The active control method for embodied intelligence according to claim 3, characterized in that The perception modalities include a first response level modality and a second response level modality, the dynamic response priority of the first response level modality is in milliseconds, and the dynamic response priority of the second response level modality is in seconds; The performing control instruction planning on the perception modalities based on the dynamic response priorities to obtain a dynamic control coordination decision result includes: Determining the feature fusion path corresponding to the second response level modality based on the order of milliseconds and seconds; Based on the real-time fusion topology graph, obtaining the signal synchronization moments of the second response level modality at each target coding module in the corresponding feature fusion path; Based on the signal synchronization moment, combined with the real-time fusion topology map, control instruction planning is performed on the second response-level mode to obtain the dynamic control coordination decision result.
5. The active control method for embodied intelligence according to claim 1, wherein, For each sensing mode, based on the spatio-temporal correlation indexes of each of the adaptive coding modules relative to the feature extraction unit and the control center node, a plurality of target coding modules with the optimal spatio-temporal correlation indexes are selected, and a feature fusion path of the sensing mode is generated according to the plurality of target coding modules, including: For each sensing mode, determine a plurality of adjacent coding modules adjacent to the feature extraction unit; Calculate the spatio-temporal correlation indexes of each adjacent coding module relative to the feature extraction unit and the control center node, and determine the first target coding module with the optimal spatio-temporal correlation index from the plurality of adjacent coding modules based on the plurality of correlation indexes; Taking the first target coding module as the starting point, determine a plurality of adjacent coding modules adjacent to the target coding module, calculate the spatio-temporal correlation indexes of each adjacent coding module relative to the feature extraction unit and the control center node, and determine the next target coding module with the optimal spatio-temporal correlation index from the plurality of adjacent coding modules adjacent to the target coding module based on the plurality of correlation indexes; Repeat the steps of determining a plurality of adjacent coding modules adjacent to the target coding module, calculating the spatio-temporal correlation indexes of each adjacent coding module relative to the feature extraction unit and the control center node, and determining the next target coding module with the optimal spatio-temporal correlation index from the plurality of adjacent coding modules adjacent to the target coding module until the next target coding module is determined to be the control center node to obtain a plurality of target coding modules between the feature extraction unit and the control center node; Generate a feature fusion path of the sensing mode based on the plurality of target coding modules.
6. The active control method for embodied intelligence according to claim 5, characterized in that The calculation of the spatio-temporal correlation indexes of each adjacent coding module relative to the feature extraction unit and the control center node includes: Calculate the first signal fidelity index of each adjacent coding module from the feature extraction unit and the second phase consistency index of each adjacent coding module from the control center node respectively; For each adjacent coding module, take the normalized product of the first signal fidelity index and the second phase consistency index as the spatio-temporal correlation index of the adjacent coding module relative to the feature extraction unit and the control center node.
7. The active control method for embodied intelligence according to claim 5, characterized in that The determination of the first target coding module with the optimal spatio-temporal correlation index from the plurality of adjacent coding modules based on the plurality of correlation indexes includes: Store the plurality of correlation indexes corresponding to the plurality of adjacent coding modules into a correlation evaluation queue, and perform ascending sorting on the correlation indexes in the correlation evaluation queue to obtain a sorting sequence; Based on the sorting sequence, determine the adjacent coding module corresponding to the optimal correlation index as the first target coding module.
8. The active control method for embodied intelligence according to claim 6, characterized in that After performing multi-modal control instruction planning based on the phase conflict result and obtaining the dynamic control coordination decision result, it further includes: After each of the sensing modalities starts to execute according to the dynamic control coordination decision result, for each of the sensing modalities, the spectral interference intensity between the sensing modality and at least one adjacent sensing modality is detected in real time; When the spectral interference intensity between the sensing modality and the adjacent sensing modality exceeds a preset tolerance threshold, the adjacent sensing modality is regarded as an interference source, and based on the interference source, a dynamic feature fusion path planning is performed on the sensing modality to obtain an updated fusion path of the sensing modality.
9. The active control method for embodied intelligence according to claim 1, characterized in that, The obtaining of multi-modal sensing data includes: Collecting the original sensing stream through a distributed heterogeneous sensor array, where the distributed heterogeneous sensor array includes a vision sensor, a flexible electronic skin, and an inertial measurement unit; Performing spatio-temporal stamp marking on the original sensing stream to generate a modal data stream with time stamps; Inputting the modal data stream with time stamps into the preprocessing buffer of the heterogeneous data fusion framework, and performing modal feature decoupling operations through a programmable photon chip to separate out the independent feature vectors of each sensing modality; The performing of spatio-temporal stamp marking on the original sensing stream includes: Receiving the optical clock reference signal output by the programmable photon chip and converting it into an electrical synchronous trigger signal; Deploying a time stamp embedding circuit at the sensor data acquisition end, and generating an absolute time tag accurate to the microsecond level according to the electrical synchronous trigger signal; Performing frame structure encapsulation on the absolute time tag and the original data packet collected by the corresponding sensor, where the time tag is used as the data frame header identifier; Transmitting the encapsulated data stream with time stamps to the preprocessing buffer through a high-speed serial interface.
10. An active control system for embodied intelligence, characterized in that, The system includes: A data acquisition and modal processing module, configured to acquire multi-modal sensing data, and determine the feature extraction unit and the dynamic control center node of each sensing modality from a preset heterogeneous data fusion framework based on the multi-modal sensing data, where the heterogeneous data fusion framework includes a plurality of adaptive coding modules located between the feature extraction unit and the control center node; An encoding module selection and path generation module, configured to, for each sensing modality, select a plurality of target encoding modules with the optimal spatio-temporal correlation index based on the spatio-temporal correlation indexes of each of the adaptive coding modules with respect to the feature extraction unit and the control center node, and generate a feature fusion path of the sensing modality according to the plurality of target encoding modules; A signal synchronization determination module, configured to, for each sensing modality, obtain the dynamic response delay of the sensing modality, and determine the signal synchronization moment when the sensing modality passes through a plurality of target encoding modules based on the sub-layer coupling degree index between the target encoding modules included in the feature fusion path and the dynamic response delay; A topology graph generation module, configured to mark the signal synchronization moment on each of the target encoding modules in the heterogeneous data fusion framework to obtain a real-time fusion topology graph; A phase conflict detection module, which is used to detect signal phase conflicts for each of the target encoding modules according to the signal synchronization time to obtain a phase conflict result; A motion control instruction planning module, which is used to plan multi-modal control instructions based on the phase conflict result to obtain a motion control coordination decision result.
Citation Information
Patent Citations
New energy control execution system and method based on multi-market coupling mechanism
CN120103717A
Intelligent fish blocking and ship passing system based on multi-mode AI fusion and self-adaptive control
CN120233674A
Multi-modal, multi-disciplinary feature discovery to detect cyber threats in electric power grid
US20180262525A1
Multi-agent beyond-visual-range networked collaborative perception and dynamic decision-making method and related device
WO2024098438A1
Cited By
Intelligent hierarchical memory perception method and system for body and electronic equipment
CN120724133A
Body intelligence layered memory perception method, system and electronic device
CN120724133B
Social robot account detection method and system oriented to time evolution
CN120930117A