Self-adaptive multi-mode integrated sensing communication method for intelligent agent with body

By adopting an adaptive multimodal fusion architecture and a neighborhood feature enhancement strategy, the problems of modal dynamism and lack in embodied agents are solved, enabling efficient collaborative perception and communication between agents and improving the robustness and recognition accuracy of the system.

CN122069487APending Publication Date: 2026-05-19NANJING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NANJING UNIV
Filing Date
2026-02-04
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing multimodal integrated sensing and communication methods suffer from problems such as modal dynamism, modal absence and inconsistency, lack of inter-agent collaboration mechanisms, and communication link interference in embodied agents, resulting in limited sensing capabilities and insufficient information sharing.

Method used

An adaptive multimodal fusion architecture and a neighborhood feature enhancement strategy are adopted. The fusion features of neighboring agents are obtained through wireless links, and noise reduction and fusion are performed in combination with channel information to generate a collaborative enhancement representation, thereby achieving efficient collaborative perception among agents.

Benefits of technology

It significantly improves the reliability of embodied intelligent agents in perception and communication decision-making in complex environments, enhances the robustness and recognition accuracy of the system, and supports generalization capabilities in diverse scenarios and tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122069487A_ABST
    Figure CN122069487A_ABST
Patent Text Reader

Abstract

The invention discloses a self-adaptive multi-mode integrated sensing communication method for an intelligent body, and belongs to the technical field of wireless communication and artificial intelligence fusion. The self-adaptive multi-mode integrated sensing communication method is provided with a self-adaptive multi-mode fusion framework and a neighborhood feature enhancement strategy, and the self-adaptive multi-mode fusion framework is based on a hybrid expert mechanism; the most relevant expert network can be dynamically activated to perform feature fusion according to the currently available heterogeneous sensing mode of the body agent, robust multi-mode representation is generated, the neighborhood feature enhancement strategy realizes collaborative perception among the body agents through a wireless link, the local body agent is allowed to utilize the fusion feature of the neighbor body agent, and the robustness of the body agents is enhanced. And de-noising and attention fusion are carried out in combination with channel state information so as to compensate for self mode loss or degradation, and the perception and communication decision reliability in complex scenes such as shielding and illumination variation is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of wireless communication and artificial intelligence fusion technology, specifically relating to an adaptive multimodal integrated sensing communication method for embodied intelligent agents. Background Technology

[0002] Embodied intelligent agents, as physical entities within a network, achieve real-time environmental perception by integrating multiple sensors (such as cameras, LiDAR, and millimeter-wave radar), and make decisions and communicate accordingly. Integrated sensing and communication technology unifies sensing and communication functions on the same spectrum and hardware platform, and is one of the key technologies of 6G. In recent years, multimodal sensing and communication (ISAC) has become a research hotspot by fusing sensor data from different modalities and leveraging their complementarity to overcome the limitations of single-modal systems (such as the susceptibility of vision to lighting conditions and the inability of radar to acquire textures).

[0003] However, most existing MISC solutions are designed for fixed single-modal devices, and they face significant challenges when applied to embodied intelligent agents equipped with multiple heterogeneous sensors:

[0004] 1. Modal dynamism: The combination of modalities available to an agent may change dynamically due to sensor failure, energy management, or environmental changes (such as changes in visual quality caused by day and night), and fixed fusion patterns cannot adapt.

[0005] 2. Modal missing and inconsistency: Different agents may be equipped with different types and numbers of sensors, which may cause a single agent to lack modalities that are crucial to the current task.

[0006] 3. Lack of effective collaboration mechanism: Existing methods mainly focus on the fusion within a single agent, without making full use of the complementary modal information that neighboring agents may hold to improve their own perception robustness.

[0007] 4. Impact of communication links: When sharing features across agents, noise and fading in the wireless channel will degrade the quality of transmitted features, and existing methods lack effective feature-level compensation mechanisms.

[0008] Therefore, there is an urgent need for a novel MISC method that can adaptively combine dynamic modalities and support efficient and reliable collaboration among agents, so as to fully realize the potential of embodied agents in multimodal perception and communication. Summary of the Invention

[0009] To address the problems mentioned in the background art, this invention provides an adaptive multimodal integrated sensing and communication method for embodied intelligent agents, which solves the problems of lack of dynamic adaptability due to fixed modal combinations in the prior art, limited perception capabilities of single intelligent agents, and difficulty in effective coordination of cross-intelligent agent information.

[0010] To achieve the above objectives, the present invention provides the following technical solution:

[0011] An adaptive multimodal integrated sensing and communication method for embodied intelligent agents includes the following steps:

[0012] Multiple embodied agents that are neighbors are deployed, and each embodied agent synchronously collects multimodal raw data of the target area. The multimodal raw data is mapped into deep feature representations. Each embodied agent fuses the deep feature representations through an adaptive multimodal fusion architecture to generate a local fusion representation. When any embodied agent has a missing modality or its perception quality does not meet the preset requirements, a neighborhood feature enhancement strategy is activated. The local fusion representations of its neighboring embodied agents are obtained through wireless links, and denoising and fusion are performed in combination with channel information to generate a collaboratively enhanced representation for downstream tasks.

[0013] Compared with the prior art, the beneficial effects of the present invention are:

[0014] This application sets up an adaptive multimodal fusion architecture and a neighborhood feature enhancement strategy. The adaptive multimodal fusion architecture can perform feature fusion based on the heterogeneous sensing modalities currently available to the embodied agent to generate multimodal representations. The neighborhood feature enhancement strategy enables collaborative perception between agents through wireless links, allowing local agents to utilize the fused features of neighboring embodied agents and combine them with channel state information for noise reduction and attention fusion to compensate for their own modal loss or degradation, significantly improving the reliability of perception and communication decisions in complex scenarios such as occlusion and changes in illumination. Attached Figure Description

[0015] Figure 1 This is a flowchart illustrating the process of this application;

[0016] Figure 2 This is a schematic diagram of the specific architecture of an embodied intelligent agent. Detailed Implementation

[0017] To facilitate understanding of the technical content of this invention by those skilled in the art, the invention will be further described in detail below with reference to the accompanying drawings and specific examples. It should be understood that the specific examples described herein are merely illustrative and not intended to limit the scope of the invention.

[0018] Example 1

[0019] An adaptive multimodal integrated sensing and communication method for embodied intelligent agents, such as Figure 1 and Figure 2 As shown, it includes the following steps:

[0020] S1. Multi-agent multimodal data acquisition: Deploy multiple embodied agents, each of which synchronously acquires multimodal raw data such as visual images, millimeter-wave power sequences, 3D point clouds, and radio frequency signals of the target area based on its own configured sensors.

[0021] S2. Modality-Specific Deep Feature Extraction: Construct a dedicated feature extraction network for each sensing modality to map the original data into high-dimensional feature representations.

[0022] S3. Adaptive Multimodal Fusion Based on Hybrid Experts: Through an adaptive multimodal fusion architecture, all modal features currently available to the embodied agent are dynamically fused to generate a unified and adaptive local fusion representation.

[0023] S4. Neighborhood Cooperative Feature Enhancement: When the local embodied agent has modal missing or poor perception quality, the neighborhood feature enhancement strategy is activated. The fusion features of the neighboring embodied agents are obtained through the wireless link and combined with channel information for denoising and fusion to generate a cooperative enhanced representation for downstream tasks.

[0024] Sections S2 to S4 describe the core processes from feature extraction and dynamic fusion to collaborative enhancement. Section S2 includes the following sub-steps:

[0025] S21. For each modality of data, a dedicated feature extraction network is used. Original input Mapped to deep feature representation:

[0026] ;

[0027] in, The input is the raw data for each modality. For the corresponding feature extraction network, Let k be the number of modalities available to the k-th embodied intelligent agent. For the corresponding feature extraction network The parameters are set as follows: CNN or ViT for visual images, PointNet for point clouds, and MLP for power sequences.

[0028] S22. For complex radio frequency signals, a complex convolutional network is used for processing. The complex signal... The real and imaginary parts are input into the convolutional layer respectively, and the operation follows the rules of complex convolution:

[0029] ;

[0030] in This represents the convolution operation. These represent the real and imaginary parts of the convolution kernel, respectively. Then, through activation functions and pooling layers, the final output is the radio frequency (RF) feature. .

[0031] The adaptive multimodal fusion architecture described in S3 is based on a hybrid expert model and includes the following steps:

[0032] S31. Concatenate all available modal features into The weights of each expert are calculated using a gating network:

[0033] ;

[0034] ;

[0035] in, The total number of experts, Let j be the probability of choosing expert j. Here is the weight matrix of the gated network. is the bias vector of the gated network.

[0036] S32. Select the top S experts with the highest probabilities to form the activation set. Each expert j has a parameter of... The subnetwork (such as the cross-attention module) outputs:

[0037] ;

[0038] in, is the mapping function (or nonlinear transformation function) of the j-th expert network, used to process the spliced ​​multimodal features;

[0039] S33. Sum the expert outputs according to probability weights to obtain adaptive fusion features:

[0040] .

[0041] The neighborhood feature enhancement strategy described in S4 includes the following steps:

[0042] S41, Neighbor Embodied Intelligent Agent Its fusion features Send to the embodied intelligent agent k via wireless channel:

[0043] ;

[0044] in For the channel matrix, It is additive noise.

[0045] S42. Utilizing channel state information Through denoising network Recovery characteristics:

[0046] ;

[0047] in For noise reduction network Learnable parameters;

[0048] S43, Based on local characteristics For querying neighborhood features For keys and values, attention is computed and weighted fusion is performed to generate collaboratively enhanced representations. Used for downstream communication sensing tasks:

[0049] ;

[0050] ;

[0051] Among them, W Q W K and W V K represents the linear projection weight matrices of the query, key, and value, respectively. T Let K be the transpose of the key matrix. The weights are fused, and d is the feature dimension;

[0052] S44. Employ an end-to-end multi-task learning strategy for joint optimization. Specifically, collaboratively enhance features... Input prediction network Perform communication sensing tasks such as beam prediction and output optimal decisions. :

[0053] ;

[0054] in For beamcodebook, Let b be the learnable parameters of the prediction network G, and b be the candidate beam index.

[0055] The model's total loss function includes task loss and expert load balancing regularization:

[0056] ;

[0057] in The loss function for downstream communication sensing tasks is typically cross-entropy loss, π. k This represents the expert probability distribution output by the gating network in S31. Let u be a hyperparameter, and D be a uniform distribution. KL KL divergence is used to measure the difference between the gated probability distribution and the uniform distribution, serving as a load balancer. By minimizing the total loss function, all network parameters for feature extraction, adaptive fusion, neighborhood enhancement, and task prediction are jointly optimized.

[0058] S5. Joint parameter optimization and convergence determination based on the total loss function;

[0059] Using the total loss function calculated in S4 Guide the end-to-end training of the model and determine whether the model meets the preset convergence conditions (the loss function value no longer decreases significantly or reaches the preset number of training epochs):

[0060] If convergence is not achieved, then based on the total loss function... The gradient is calculated using the backpropagation algorithm, and the parameters α of the feature extraction network, β of the adaptive fusion architecture (including gating and expert network parameters), γ of the denoising network, and σ of the prediction network are jointly updated. After the parameters are updated, the process returns to S1, and the next batch of multimodal raw data (such as batch data from the DeepSense 6G dataset) is input to enter the next round of iterative training. Otherwise, the current model parameters are determined to be the optimal solution, the training process ends, and the finally optimized model is deployed on the embodied agent for real-time integrated perception and communication tasks.

[0061] In this embodiment, this application proposes an adaptive radio frequency-visual fusion perception enhancement method for embodied agents. Based on the penetrability of radio frequency signals and the high-resolution structural perception capability of visual images, this method achieves complementary enhancement of target information under complex occlusion, low light, and adverse weather conditions through a cross-modal dynamic fusion mechanism. Simultaneously, it introduces an adaptive fusion strategy based on hybrid experts to support real-time perception inference under arbitrary modal combinations, significantly improving the system's robustness and recognition accuracy in modality-deficient scenarios. Furthermore, our method integrates a multi-head decoder and a multi-task learning mechanism, supporting parallel output of tasks such as velocity estimation, image restoration, and beam prediction, achieving feature sharing and task collaboration, and enhancing the model's generalization ability and learning efficiency in diverse scenarios and tasks. This method possesses good scalability and deployment flexibility, providing a feasible path and theoretical support for building a robust, cross-scenario, and cross-agent collaborative next-generation multimodal embodied perception system.

[0062] This application designs an adaptive multimodal fusion architecture and a neighborhood feature enhancement strategy, realizing intelligent collaboration and robust decision-making in multimodal perception and communication under dynamic and uncertain environments, significantly improving the system's perception accuracy, communication reliability, and environmental adaptability. Finally, the entire method is guided by communication tasks such as beam prediction, and adopts end-to-end multi-task learning for joint optimization. While ensuring high accuracy, its modular design combines high computational efficiency and good scalability, enabling effective deployment on various resource-constrained embodied agents and edge devices, providing a practical solution for collaborative perception and communication in next-generation 6G networks.

[0063] Those skilled in the art will recognize that the embodiments, drawings, and specific steps described herein are merely for the purpose of helping to understand the technical principles, core architecture, and implementation methods of the present invention, and should not be construed as limiting the scope of protection of the present invention. Any obvious modifications, variations, substitutions, or combinations made to the module composition, network structure, fusion strategy, training method, or application scenario of the present invention based on the core ideas of adaptive multimodal fusion and collaborative enhancement disclosed in this invention, without departing from the essence of the present invention, such as modifying the specific implementation of the expert network, adjusting the gating mechanism, adopting different denoising network structures, or applying it to other embodied intelligence scenarios such as robot collaboration, vehicle networking, and drone formation, should be considered to be included within the scope of protection of the present invention.

Claims

1. An adaptive multimodal integrated sensing and communication method for embodied intelligent agents, characterized in that, Includes the following steps: Multiple embodied agents that are neighbors are deployed, and each embodied agent synchronously collects multimodal raw data of the target area. The multimodal raw data is mapped into deep feature representations. Each embodied agent fuses the deep feature representations through an adaptive multimodal fusion architecture to generate a local fusion representation. When any embodied agent has a missing modality or its perception quality does not meet the preset requirements, a neighborhood feature enhancement strategy is activated. The local fusion representations of its neighboring embodied agents are obtained through wireless links, and denoising and fusion are performed in combination with channel information to generate a collaboratively enhanced representation for downstream tasks.

2. The adaptive multimodal integrated sensing and communication method for embodied intelligent agents according to claim 1, characterized in that, The multimodal raw data includes visual images, millimeter-wave power sequences, 3D point clouds, and radio frequency signals.

3. The adaptive multimodal integrated sensing and communication method for embodied intelligent agents according to claim 2, characterized in that, Dedicated feature extraction networks are constructed for the raw data of each modality, thereby mapping the multimodal raw data into deep feature representations. Specifically, convolutional neural networks or visual Transformers are used for visual images, PointNet-based networks are used for 3D point clouds, multilayer perceptrons are used for millimeter-wave power sequences, and complex convolutional networks are used for complex radio frequency signals to extract features respectively, thus obtaining deep feature representations of the raw data of each modality.

4. The adaptive multimodal integrated sensing and communication method for embodied intelligent agents according to claim 3, characterized in that, Deep feature representation Expressed as: ; in, The input is the raw data for each modality. For the corresponding feature extraction network, Let k be the number of modalities available to the k-th embodied intelligent agent. For the corresponding feature extraction network Parameters; When the original data is a complex radio frequency signal, a complex convolutional network is used for processing: complex radio frequency signals The real and imaginary parts are input into the convolutional layer respectively, and the operation follows the rules of complex convolution: ; in This represents the convolution operation. These represent the real and imaginary parts of the convolution kernel, respectively. Then, through activation functions and pooling layers, the final output is a deep feature representation of radio frequency. .

5. The adaptive multimodal integrated sensing and communication method for embodied intelligent agents according to claim 4, characterized in that, The adaptive multimodal fusion architecture includes an input module, a gating network, multiple expert networks, a weighting module, and an output module. The input module receives and concatenates the deep feature representations of each modality, and then inputs them into the gating network. The gating network calculates the activation probabilities of each expert network and dynamically selects and activates the most relevant expert networks for processing based on the activation probabilities. The weighting module sums the outputs of the selected expert networks according to their activation probabilities to obtain the final adaptive fusion features, which are then sent to the output module.

6. The adaptive multimodal integrated sensing and communication method for embodied intelligent agents according to claim 5, characterized in that, The specific steps for generating local fusion representations are as follows: Concatenate the deep feature representations of all available modalities into ; The weights of each expert network are calculated using a gating network: ; ; in, The total number of experts, Let j be the probability of choosing expert j. Here is the weight matrix of the gated network. is the bias vector of the gated network; The activation set is formed by selecting the top S experts with the highest probabilities. Each expert j has a parameter of The subnetwork, whose output is: ; in, is the mapping function (or nonlinear transformation function) of the j-th expert network, used to process the spliced ​​multimodal features; The expert outputs are summed in probability-weighted order to obtain the adaptive fusion features: 。 7. The adaptive multimodal integrated sensing and communication method for embodied intelligent agents according to claim 6, characterized in that, The neighborhood feature enhancement strategy includes the following steps: Neighbor Embodied Intelligent Agent Its fusion features Send to the local embodied intelligent agent k via wireless channel: ; in For the channel matrix, It is additive noise; Utilizing channel state information Through denoising network Recovery characteristics: ; in For noise reduction network Learnable parameters; With local characteristics of the local embodied intelligent agent k For query, For keys and values, attention is computed and weighted fusion is performed to generate collaboratively enhanced representations. Used for downstream communication sensing tasks: ; ; Among them, W Q W K and W V K represents the linear projection weight matrices of the query, key, and value, respectively. T Let K be the transpose of the key matrix. The weights are fused, and d is the feature dimension; An end-to-end multi-task learning strategy is employed for joint optimization, specifically, collaboratively enhancing features. Input prediction network Perform communication sensing tasks such as beam prediction and output optimal decisions. : ; in For beamcodebook, Let b be the learnable parameters of the prediction network G, and b be the candidate beam index. Total loss function Includes task loss and expert load balancing regularization terms: ; in For the loss function of the downstream communication sensing task, π k This represents the expert probability distribution output by the gating network. Let u be a hyperparameter, and D be a uniform distribution. KL KL divergence is used to measure the difference between the gated probability distribution and the uniform distribution, and plays a role in load balancing. By minimizing the total loss function, it jointly optimizes all network parameters for feature extraction, adaptive fusion, neighborhood enhancement, and task prediction.