A multi-source collaborative perception and fusion transmission method based on semantic coding

CN122554056APending Publication Date: 2026-08-11THE CHINESE UNIV OF HONG KONG (SHENZHEN) +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-09
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

因此,若仍采用通用压缩或固定融合方式,容易出现“传输了大量数据但对感知结果贡献有限”的问题

Benefits of technology

[0012]本发明的有益效果是:本发明针对多设备协同感知系统中原始数据传输开销大、中心节点计算负担重、异构模态信息难以有效融合的问题,传统压缩编码未面向感知任务优化,导致关键语义信息保留不足,以及融合权重难以随设备可靠性、信道质量和任务贡献度自适应调整的问题,提出了基于语义编码的多源协同感知与融合传输方法,可以在降低传输开销的同时,提高感知精度、增强多模态兼容性、提升系统鲁棒性,推动了该方面研究向动态网络适应性、工程部署简单性等方向上的发展。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122554056A_ABST
    Figure CN122554056A_ABST
Patent Text Reader

Abstract

This invention discloses a multi-source collaborative sensing and fusion transmission method based on semantic coding, characterized by the following steps: Step S1: Determine the type of sensing task to be performed and record the target state parameter corresponding to the sensing task as ; Step S2: Obtain the local sensing information set of each sensing device; Step S3: Perform feature extraction to obtain low-dimensional features related to the sensing task, map the low-dimensional features to semantic codewords corresponding to the sensing task, and generate transmission symbols; Step S4: After receiving semantic information from K sensing devices, the central node performs normalization processing and calculates importance weights for weighted fusion; Step S5: Recover the target state estimation result. This invention improves sensing accuracy, enhances multimodal compatibility, and improves system robustness while reducing transmission overhead.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of semantic communication, and in particular to a multi-source collaborative sensing and fusion transmission method based on semantic coding. Background Technology

[0002] In scenarios such as intelligent transportation, vehicle-to-everything (V2X) communication, low-altitude target monitoring, industrial internet, and smart cities, network nodes not only need to transmit data but also need to perceive the surrounding environment or target status, such as estimating the target's position, speed, angle, distance, motion trend, or category. In these target scenarios, individual sensing devices are limited by deployment location, obstruction, field of view, channel fading, and noise interference, making it difficult to consistently obtain stable and comprehensive sensing results. Therefore, to improve sensing coverage and estimation accuracy, a multi-device collaborative sensing approach is usually required. This involves multiple devices acquiring target information from different locations or modalities and then uploading the data to a central node for fusion.

[0003] In practical systems, data collected by multiple devices often exhibits characteristics such as high dimensionality, significant repetition, inconsistent sampling frequencies, and large differences in channel conditions. Directly uploading the original signal or high-dimensional features not only incurs substantial communication overhead but also increases the computational burden on the central node. Furthermore, traditional compression coding typically aims to reduce reconstruction errors, focusing on recovering the original signal itself; however, in collaborative sensing tasks, the system truly needs effective semantic information relevant to the sensing target. Therefore, if a general compression or fixed fusion method is still used, the problem of "transmitting a large amount of data but contributing little to the sensing results" easily arises. In summary, it is essential to propose a novel multi-source collaborative sensing transmission scheme that enables devices to extract and transmit key semantic information around specific sensing tasks, and allows the central node to adaptively fuse data based on the contributions of different devices and modalities. This approach reduces transmission overhead while improving sensing accuracy and system robustness. Summary of the Invention

[0004] The purpose of this invention is to overcome the shortcomings of the prior art and provide a multi-source collaborative sensing and fusion transmission method based on semantic coding, which can improve sensing accuracy, enhance multimodal compatibility, and improve system robustness while reducing transmission overhead.

[0005] The objective of this invention is achieved through the following technical solution: a multi-source collaborative sensing and fusion transmission method based on semantic coding, used for information fusion transmission in a multi-source collaborative sensing and fusion transmission system, wherein the multi-source collaborative sensing and fusion transmission system includes a target / sensing scene, multi-source sensing devices, a transmission link, and a central node; the multi-source sensing devices include K sensing devices;

[0006] The method includes the following steps:

[0007] Step S1: Determine the type of perception task to be performed, and record the target state parameter corresponding to the perception task as... ;

[0008] Step S2: Each sensing device collects target-related information from different spatial locations, different observation angles, or different modalities, and obtains the local sensing information set of each sensing device. ;

[0009] Step S3: Each sensing device extracts network features using local semantic features. Local sensing information set Feature extraction is performed to obtain low-dimensional features relevant to the perception task. , low-dimensional features Mapped to semantic codewords corresponding to the perception task The semantic codewords are then jointly source-channel coded according to the current link state, and transmission symbols are generated. ;

[0010] Step S4: After receiving semantic information from K sensing devices, the central node first performs normalization processing on this information to convert the received semantic information into normalized features in a unified feature space, and calculates importance weights for weighted fusion.

[0011] Step S5: The central node inputs the fused semantic features into the perceptual estimation network to recover the target state estimation result.

[0012] The beneficial effects of this invention are as follows: This invention addresses the problems of high overhead in original data transmission, heavy computational burden on central nodes, and difficulty in effectively fusing heterogeneous modal information in multi-device collaborative sensing systems. Traditional compression coding is not optimized for sensing tasks, resulting in insufficient retention of key semantic information and difficulty in adaptively adjusting fusion weights according to device reliability, channel quality, and task contribution. This invention proposes a multi-source collaborative sensing and fusion transmission method based on semantic coding, which can improve sensing accuracy, enhance multimodal compatibility, and improve system robustness while reducing transmission overhead. This promotes the development of research in this area towards dynamic network adaptability and simplified engineering deployment. Attached Figure Description

[0013] Figure 1 This is a schematic diagram of the system principle involved in the present invention;

[0014] Figure 2 This is a flowchart of the method of the present invention. Detailed Implementation

[0015] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings, but the scope of protection of the present invention is not limited to the following description.

[0016] This invention mainly consists of three core parts: (1) First, the perception task is determined, and a semantic codebook and encoding network are constructed according to the characteristics of the task; (2) Second, the perception device compresses the local perception information into transmittable semantic symbols according to the semantic codebook; (3) Finally, the central node unifies and weights the received semantic symbols, and outputs information such as the target state or category estimation result. Based on the above ideas, this invention introduces a semantic compression mechanism for perception tasks at the device end, so that each device only transmits low-dimensional semantic features that are highly related to the current perception task, and introduces codeword normalization, importance assessment and multi-source weighted fusion mechanism at the central node, so that information from different devices, different modalities and different channel conditions can be uniformly processed and used for target state estimation.

[0017] like Figure 1 The diagram illustrates the multi-source collaborative sensing and fusion transmission system based on semantic coding constructed in this invention. This system comprises a target or sensing scene, multi-source sensing devices, a transmission link, and a central fusion node. Specifically, the multi-source sensing devices include K sensing devices, which can be radar devices, camera devices, communication transceivers, roadside units, drones, vehicle terminals, or industrial sensor nodes, etc. The central node can be a base station, edge server, fusion processor, or cloud server, etc. Furthermore, it should be noted that each sensing device includes at least a sensing information acquisition module, a semantic feature extraction module, a codebook matching module, and a joint source-channel coding module. The central node includes at least a receiving module, a codeword normalization module, an importance assessment module, a multi-source fusion module, and a sensing result output module.

[0018] like Figure 2 As shown, a multi-source collaborative sensing and fusion transmission method based on semantic coding includes the following steps:

[0019] Step S1: Determine the type of perception task to be performed, and record the target state parameter corresponding to the perception task as... ;

[0020] The system determines the type of perception task to be performed, such as target localization, distance estimation, angle estimation, velocity estimation, trajectory prediction, target recognition, or multi-task joint perception, and records the target state parameters corresponding to that task as follows: The target state parameter can be a continuous variable, such as distance, speed, and angle, or a discrete variable, such as target category or behavior state.

[0021] Step S2: Each sensing device collects target-related information from different spatial locations, different observation angles, or different modalities, and obtains the local sensing information set of each sensing device. ;

[0022] Its expression is:

[0023]

[0024] in, Indicates the number of observation time slots. Indicates the first The sensing information acquisition module of the sensing device is in the first In this application's embodiments, the sensing data includes one or more of the following: radar echo data, image frame or video frame data, wireless channel status information, positioning measurement information, inertial measurement data, lidar point cloud data, and acoustic sensing data.

[0025] Step S3: Each sensing device extracts network features using local semantic features. Local sensing information set Feature extraction is performed to obtain low-dimensional features relevant to the perception task. , low-dimensional features Mapped to semantic codewords corresponding to the perception task The semantic codewords are then jointly source-channel coded according to the current link state, and transmission symbols are generated. ;

[0026] Compared to traditional multi-source collaborative methods that directly transmit collected data, in this invention, the semantic feature extraction module utilizes a local semantic feature extraction network. Local sensing information set Feature extraction is performed to obtain low-dimensional features relevant to the perception task. Its expression is:

[0027]

[0028] in, Indicates the sensory task identifier. Indicates the first The semantic feature extraction network parameters for each sensing device. One point that needs special explanation is the sensing task identifier. The introduction of this feature allows the device to retain different semantic information for different tasks during feature extraction. For example, motion change information can be prioritized in velocity estimation tasks, while structural or texture information can be prioritized in category recognition tasks.

[0029] In the embodiments of this application, the local semantic feature extraction network can employ convolutional neural networks, visual Transformers, temporal convolutional networks, recurrent neural networks, graph neural networks, multilayer perceptrons, or combinations thereof, depending on the type of perceptual data. Specifically, for image or video frame data, convolutional neural networks or visual Transformers can be used; for radar echoes, wireless channel state information, or temporal measurement data, one-dimensional convolutional networks, temporal convolutional networks, recurrent neural networks, or Transformer networks can be used.

[0030] In this embodiment, pre-training can be performed offline based on labeled samples or historical sensing data, and fine-tuning can be performed online based on sensing error feedback; the training samples include a set of local sensing information. This includes task annotations such as target state parameters and target category corresponding to the perception task. During training, the local semantic feature network... by and As input, output low-dimensional features relevant to the current perception task. And it is trained using target loss, classification loss, detection loss, or a combination thereof, so that... Semantic information that contributes to the current perception task is retained. In other embodiments, if there are sufficient training samples, end-to-end joint training can also be performed with the semantic codebook, the joint source-channel coding network, and the perception estimation network.

[0031] Extracting low-dimensional features for different perceptual tasks Next, the codebook matching module establishes semantic codebooks for different sensing tasks / sensing devices, which are sets of codewords associated with their task performance. Then, the sensing device uses low-dimensional features... The expression for mapping one or more codewords in the semantic codebook to achieve codeword matching is as follows:

[0032]

[0033] in, Representation and low-dimensional features Matching semantic codewords Indicates the first The semantic codebook corresponding to each sensing device. This represents the matching function corresponding to the matching method, such as nearest neighbor matching, learnable codebook matching, and attention matching. Codebook matching encoding can convert continuous or high-dimensional features into semantic symbols that are easy to transmit, thereby reducing the transmission burden on the link.

[0034] In obtaining the semantic codewords corresponding to the perception task Subsequently, the joint source-channel coding module performs joint source-channel coding on the semantic codewords according to the current link state and generates transmission symbols. Unlike traditional methods that first compress and then independently correct errors, this invention incorporates information such as the relevance of the sensing task, the importance of semantic codewords, and the channel state as encoding inputs. Its expression is:

[0035]

[0036] in, Indicates the first The joint source-channel coding network used by each sensing device Indicates the first Symbols transmitted by a sensing device Indicates the first The link status from each sensing device to the central node Indicates the first The joint source channel coding network parameters used by each sensing device.

[0037] Step S4: After receiving semantic information from K sensing devices, the central node first performs normalization processing on this information to convert the received semantic information into normalized features in a unified feature space, and calculates importance weights for weighted fusion.

[0038] In the central node, the receiving module first receives semantic information from K sensing devices;

[0039] Since different devices and modalities may have different codeword dimensions, scales, and statistical distributions, direct fusion may lead to over-amplification or neglect of certain modalities. Therefore, after receiving semantic information from K sensing devices, the central node's receiving module first normalizes this information through a codeword normalization module to convert the received semantic information into normalized features in a unified feature space. Specifically, the codewords from the Kth sensing devices... The standardized characteristics of a sensing device can be represented as:

[0040]

[0041] in, Indicates the relationship with the first A standardized network corresponding to each sensing device. Indicates the first Transmission symbols sent by a sensing device The semantic information is received and recovered by the central node after passing through the transmission link. Indicates to Features obtained through normalization The parameters represent the normalized network. The normalized network is used to map semantic information uploaded from different sensing devices or different modalities to a feature space with a unified dimension and a unified numerical range. Specifically, it includes linear mapping layers, batch normalization layers, layer normalization layers, multilayer perceptrons, attention mapping networks, autoencoders, or combinations thereof. In embodiments of this application, the normalized network can be pre-trained offline using multi-device, multimodal sample data to learn scale alignment and distribution alignment relationships between different semantic codewords or semantic features. Specifically, training samples can include received semantic information obtained by different sensing devices under the same sensing task, the same target state, or adjacent observation time slots. Standardized networks Input, output normalized features When reference normalization features exist, the normalization network is trained using mean squared error, cosine distance, or contrastive learning loss. When explicit reference normalization features are absent, the normalization network can be connected to an importance evaluation network, a multi-source fusion module, and a perceptual estimation network, and then jointly trained end-to-end using the final perceptual estimation error. Thus, the normalization network can map semantic information uploaded from different devices or modalities to a feature space with a unified dimension and numerical range.

[0042] After converting semantic information from all sensing devices into normalized features in a unified feature space, the importance assessment module of the central node evaluates the reliability of the devices (by...). , Characterization), modal contribution (by) , Representation), perceptual task type (by) Characterization) and historical estimation error (by (Characteristics) Calculate importance weights and other factors for the first... Weights are assigned to each sensing device. Its expression is:

[0043]

[0044] in, The importance weight calculation network can be implemented using a weighted scoring function, a multilayer perceptron, an attention network, or a gating network. Indicates the first The historical estimation error or device confidence level of each sensing device is updated using a moving average method, that is, based on the historical estimation error or device confidence level of the current time slot. The error between the estimation result and the reference result corresponding to each sensing device is used to update the historical error from the previous moment using weighted average. The parameters represent the importance assessment network. During offline training, the training samples for the importance assessment network can consist of historical sensing data from multiple devices and multiple modalities. For the ... Each sensing device, with normalized features as input to each training sample. Link status Perceive task identifiers And historical estimation errors or equipment confidence levels The training labels for the importance assessment network can be generated using leave-one-out method, teacher fusion model, or expert rules. In the absence of explicit importance labels, the importance assessment network can be jointly trained end-to-end with the multi-source fusion module and the perception estimation network, or incrementally updated based on feedback errors during online operation. This step ensures that the entire system can reduce its fusion weights when a device has poor channel quality, a poor observation angle, or insufficient modal information.

[0045] After obtaining the standardized features and corresponding importance weights of each sensing device, the central node performs weighted fusion of multi-source semantic features through a multi-source fusion module to form a fused semantic representation for the current sensing task. The fusion process can employ methods such as weighted summation, attention fusion, gating fusion, and hierarchical fusion. This invention uses weighted summation as an example to illustrate the fusion process. Specifically, the fused semantic features can be calculated using the following formula:

[0046]

[0047] in, This represents the semantic features after fusion. Indicates the first The first sensing device or the first Importance weights corresponding to class modalities.

[0048] This step ensures that the central node can suppress the impact of low-quality links or low-contribution modes and highlight information sources that contribute significantly to the current sensing task.

[0049] Step S5: The perception result output module of the central node inputs the fused semantic features into the perception estimation network to recover the target state estimation result, namely the target distance, speed, angle, position, category, trajectory or combination thereof.

[0050] Specifically, the estimation of the target state parameters can be expressed as:

[0051] ,

[0052] in, Represents a perception estimation network, This represents the parameters of the perception estimation network. This represents an estimate of the state of the target parameters;

[0053] The perception estimation network is trained offline based on multi-device, multi-modal sample data and perception errors. The perception estimation network is implemented using a neural network. During offline training, the system collects historical perception samples from multiple devices and multiple modalities, and obtains fused semantic features after semantic feature extraction, codebook matching, normalization, and multi-source fusion. Simultaneously, obtain the true target state parameters corresponding to the sample. Or task labels. During training, the perceptual estimation network fuses semantic features. And optional perception task identifiers As input, output the estimated results of the target state parameters. and according to and The differences between these parameters are used to construct the loss function. For example, for continuous state parameters such as distance, velocity, angle, or position, mean squared error, mean absolute error, or a combination thereof can be used; for discrete state parameters such as target category or behavioral state, cross-entropy loss can be used; for trajectory prediction tasks, trajectory error loss can be used, and then the perceptual estimation network parameters are updated through gradient descent or backpropagation. Incremental updates can also be performed based on feedback errors during the online operation phase.

[0054] One point that needs special explanation is that the central node also synchronously outputs an estimate of the target parameter state. Information such as confidence level, error estimation, and effective device set is collected to facilitate subsequent feedback and updates.

[0055] In embodiments of this application, the fusion transmission method further includes a feedback update step:

[0056] The central node obtains an estimate of the target parameter state. Subsequently, based on real annotations, historical trajectories, prior constraints, consistency verification results among multiple devices, or subsequent observation results, the current perception error is calculated and used as a feedback signal to update the semantic feature extraction network, semantic codebook, joint source-channel coding network, codeword normalization network, importance weight calculation network, and perception estimation network. Specifically, during the offline training phase, the system jointly trains the above modules based on sample data; during the online operation phase, when the perception error, link quality degradation, or target scene change exceeds a preset threshold for multiple consecutive time slots, the system triggers incremental updates to adjust the codebook, coding parameters, or fusion weights.

[0057] In addition to the steps mentioned above, the construction method of the semantic codebook is also explained. Specifically, the semantic codebook can be constructed in two ways: offline training and online updating. In the offline training phase, the system collects sample data from different devices, modalities, channel conditions, and target states, and trains the codebook with the performance of the perception task as the optimization objective. In the online operation phase, the system can periodically update the codebook based on recent changes in channel conditions, target type, or perception errors.

[0058] The foregoing description illustrates and describes a preferred embodiment of the present invention. However, as previously stated, it should be understood that the present invention is not limited to the forms disclosed herein and should not be construed as excluding other embodiments. It can be used in various other combinations, modifications, and environments, and can be altered within the scope of the inventive concept described herein through the foregoing teachings or techniques or knowledge in related fields. Any modifications and variations made by those skilled in the art that do not depart from the spirit and scope of the present invention should be within the protection scope of the appended claims.

Claims

1. A multi-source collaborative sensing and fusion transmission method based on semantic coding, used for information fusion transmission in a multi-source collaborative sensing and fusion transmission system, characterized in that: The multi-source collaborative sensing and fusion transmission system includes a target / sensing scene, multi-source sensing devices, a transmission link, and a central node; the multi-source sensing devices include K sensing devices. The method includes the following steps: Step S1: Determine the type of perception task to be performed, and record the target state parameter corresponding to the perception task as... ; Step S2: Each sensing device collects target-related information from different spatial locations, different observation angles, or different modalities, and obtains the local sensing information set of each sensing device. ; Step S3: Each sensing device extracts network features using local semantic features. Local sensing information set Feature extraction is performed to obtain low-dimensional features relevant to the perception task. , low-dimensional features Mapped to semantic codewords corresponding to the perception task The semantic codewords are then jointly source-channel coded according to the current link state, and transmission symbols are generated. ; Step S4: After receiving semantic information from K sensing devices, the central node first performs normalization processing on this information to convert the received semantic information into normalized features in a unified feature space, and calculates importance weights for weighted fusion. Step S5: The central node inputs the fused semantic features into the perceptual estimation network to recover the target state estimation result.

2. The multi-source collaborative sensing and fusion transmission method based on semantic coding according to claim 1, characterized in that: The sensing devices include radar equipment, camera equipment, communication transceiver equipment, roadside units, drones, vehicle terminals, or industrial sensing nodes. The central node is a base station, edge server, fusion processor, or cloud server.

3. The multi-source collaborative sensing and fusion transmission method based on semantic coding according to claim 1, characterized in that: The types of perception tasks include target localization, distance estimation, angle estimation, velocity estimation, trajectory prediction, target recognition, or multi-task joint perception; The target state parameters are either continuous or discrete variables; the continuous variables include distance, speed, and angle; the discrete variables include target category or behavioral state.

4. The multi-source collaborative sensing and fusion transmission method based on semantic coding according to claim 1, characterized in that: In step S2, the first of the multi-source sensing devices The set of local sensing information obtained by a sensing device is denoted as . : ; in, Indicates the number of observation time slots. Indicates the first The first sensing device in the Each time slot contains sensory data observed from different spatial locations, different observation angles, or in different modalities.

5. The multi-source collaborative sensing and fusion transmission method based on semantic coding according to claim 1, characterized in that: Step S3 includes: S301. Sensing devices utilize local semantic features to extract network information. Local sensing information set Feature extraction is performed to obtain low-dimensional features relevant to the perception task. The expression is: ; in, Indicates the sensory task identifier. Indicates the first Semantic feature extraction network parameters for each sensing device; S302. Extracting low-dimensional features for different perceptual tasks Then, semantic codebooks are established for different sensing tasks / sensing devices, which are sets of codewords associated with their task performance; S303. Sensing devices use low-dimensional features The expression for mapping one or more codewords in the semantic codebook to achieve codeword matching is as follows: ; in, Representation and low-dimensional features Matching semantic codewords Indicates the first The semantic codebook corresponding to each sensing device. The matching function represents the matching method corresponding to the matching method, which includes nearest neighbor matching, learnable codebook matching, or attention matching. S304. The sensing device performs joint source-channel coding on the semantic codewords based on the current link state and generates transmission symbols. Its expression is: ; in, Indicates the first The joint source-channel coding network used by each sensing device Indicates the first Symbols transmitted by a sensing device Indicates the first The link status from each sensing device to the central node Indicates the first The joint source channel coding network parameters used by each sensing device.

6. The multi-source collaborative sensing and fusion transmission method based on semantic coding according to claim 1, characterized in that: Step S4 includes: S401. After receiving semantic information from K sensing devices, the central node must first normalize this information to convert the received semantic information into normalized features in a unified feature space, which comes from the Kth sensing device. The normalized characteristics of a sensing device are represented as follows: ; in, Indicates the relationship with the first A standardized network corresponding to each sensing device. Indicates the first Transmission symbols sent by a sensing device The semantic information is received and recovered by the central node after passing through the transmission link. Indicates to Features obtained through standardization The parameters represent the planned network; S402. The central node calculates the importance weight based on the device's reliability, modal contribution, sensing task type, and historical estimation error, and assigns it to the [number]th [node]. Weights are assigned to each sensing device. Its expression is: ; in, The importance weight calculation network can be implemented using a weighted scoring function, a multilayer perceptron, an attention network, or a gating network. Indicates the first The historical estimation error or device confidence level of each sensing device is updated using a moving average method, that is, based on the historical estimation error or device confidence level of the current time slot. The error between the estimation result and the reference result corresponding to each sensing device is used to update the historical error from the previous moment using weighted average. The parameters represent the importance assessment network; S403. After obtaining the standardized features and corresponding importance weights of each sensing device, the central node performs weighted fusion of multi-source semantic features to form a fused semantic representation oriented towards the current sensing task. The fusion methods include one of the following: weighted summation, attention fusion, gating fusion, and hierarchical fusion.

7. The multi-source collaborative sensing and fusion transmission method based on semantic coding according to claim 6, characterized in that: In step S5, the estimation of the target state parameters is expressed as follows: , in, Represents a perception estimation network, This represents the parameters of the perception estimation network. This represents an estimate of the state of the target parameters.

8. A multi-source collaborative sensing and fusion transmission method based on semantic coding according to any one of claims 1 to 7, characterized in that: The fusion transmission method further includes a feedback update step: The central node obtains an estimate of the target parameter state. Then, based on the real annotations, historical trajectories, prior constraints, consistency test results among multiple devices, or subsequent observation results, the current perception error is calculated, and this error is used as a feedback signal to update the semantic feature extraction network, semantic codebook, joint source-channel coding network, normalization network, importance weight calculation network, and perception estimation network.