Channel polarization oriented multi-modal data semantic coding and reliable transmission method

By extracting anchor points and target modal features from multimodal data, and adjusting the path metric by combining channel polarization effects and cross-modal semantic distortion, the problem of pruning errors in traditional decoders under poor channel conditions is solved, thus achieving reliable transmission and semantic coherence of multimodal data.

CN122339629APending Publication Date: 2026-07-03NANJING LUKOU INT AIRPORT AIRPORT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NANJING LUKOU INT AIRPORT AIRPORT TECH CO LTD
Filing Date
2026-04-01
Publication Date
2026-07-03

AI Technical Summary

Technical Problem

In existing communication systems, under harsh channel polarization conditions or deep fading environments, traditional SCL decoders are susceptible to noise interference, leading to erroneous pruning and making it impossible to achieve reliable transmission of multimodal data. Furthermore, the CRC check mechanism results in a high communication interruption rate, failing to meet the requirements for semantic coherence and reliable recovery.

Method used

By extracting anchor points and target modal semantic features from multimodal data, calculating the reliability of polarized subchannels based on channel polarization effects, prioritizing the decoding of anchor point features, and adjusting path metrics in conjunction with cross-modal semantic distortion, polarization coding and reliable transmission are achieved.

Benefits of technology

Under deep fading limits, ensure the preservation of core semantics, reduce false pruning rate, achieve reliable communication in complex scenarios, and meet the needs of continuous semantic transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122339629A_ABST
    Figure CN122339629A_ABST
Patent Text Reader

Abstract

This invention discloses a method for semantic encoding / decoding and reliable transmission of multimodal data oriented towards channel polarization. The method includes: acquiring anchor points and target modal semantic features of multimodal data; calculating the reliability of polarization subchannels; asymmetrically mapping the two types of features to different reliability intervals according to priority and encoding and transmitting them; the receiver preferentially decodes anchor point features to reconstruct cross-modal semantic prior vectors; using the SCL algorithm to decode and split the target features, reconstructing the local semantic features of the current split path and calculating the cross-modal semantic distortion with the prior vector; and dynamically generating a joint path metric based on polarization feature parameters for list pruning, breaking the cyclic redundancy check red line, blocking retransmissions, and directly outputting the path with the minimum semantic distortion. This invention breaks down the separation between the physical and semantic layers, avoiding erroneous pruning and retransmissions caused by local noise in deep fading channels, and achieving semantic coherence and reliable transmission in complex environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of communication technology and channel coding and decoding technology, and in particular to a method for multimodal data semantic coding and decoding and reliable transmission oriented to channel polarization. Background Technology

[0002] With the gradual rollout of sixth-generation (6G) mobile communication technology, the communication paradigm is shifting from traditional symbol-oriented precise transmission to semantic feature-oriented transmission. In the field of physical layer channel coding and decoding, polar codes, as the only channel coding scheme rigorously proven to reach the Shannon limit, have been widely applied in modern mobile communication standards. Currently, the mainstream decoding algorithm for polar codes at the receiver is the Cyclic Redundancy Check-assisted Continuous Cancellation List (CA-SCL) decoding algorithm, which significantly improves the error correction performance of short and medium-length codes by retaining multiple candidate paths in the search tree. Meanwhile, at the application level, the demand for semantic-level transmission of multimodal data (such as strongly correlated images and text) is surging. Most existing semantic communication systems focus on deep learning feature extraction at the source end, aiming to filter out redundant pixel information to reduce transmission bandwidth, thereby achieving high data compression efficiency at the application level.

[0003] However, existing communication system architectures generally suffer from a technical bottleneck where the physical layer and higher-level semantics are disconnected. In environments with severe channel polarization or deep fading and ultra-low signal-to-noise ratios, traditional SCL decoders blindly rely on microscopic algebraic log-likelihood ratios (LLRs) for path measurement and pruning. Lacking guidance from macroscopic semantic prior knowledge, they are highly susceptible to sudden noise interference, leading to the incorrect pruning of correct paths. Furthermore, existing CA-SCL decoding mechanisms heavily rely on stringent cyclic redundancy check (CRC) for final hard decision-making. If sporadic bit flips at the lower level cause CRC failure, the system directly triggers a retransmission mechanism or declares packet loss. This all-or-nothing rigid physical layer verification mechanism completely ignores the high tolerance of multimodal semantic data to local pixel-level errors, resulting in persistently high communication interruption rates under adverse channel conditions, severely wasting bandwidth resources, and failing to achieve semantic coherence and reliable recovery under extreme deep fading environments. Summary of the Invention

[0004] The purpose of this section is to outline some aspects of embodiments of the present invention and to briefly describe some preferred embodiments. Simplifications or omissions may be made in this section, as well as in the abstract and title of this application, to avoid obscuring the purpose of these documents; however, such simplifications or omissions should not be construed as limiting the scope of the invention.

[0005] In view of the aforementioned existing problems, this invention is proposed. Therefore, this invention provides a channel-polarization-oriented multimodal data semantic encoding / decoding and reliable transmission method to solve the problems mentioned in the background art.

[0006] To address the aforementioned technical problems, this invention provides the following technical solution: a channel-polarization-oriented multimodal data semantic encoding / decoding and reliable transmission method, comprising: Acquire multimodal source data, and extract anchor modal semantic features and target modal semantic features from the multimodal source data; The reliability of polarized sub-channels is calculated based on the channel polarization effect. The polarized sub-channels are sorted according to the reliability values. The anchor modal semantic features are mapped to polarized sub-channels that satisfy the first preset sorting interval. The target modal semantic features are mapped to polarized sub-channels that satisfy the second preset sorting interval. Polarization coding and transmission are then performed. The receiver acquires the channel output signal and calculates the initial log-likelihood ratio of the polarization sub-channel. It then prioritizes decoding the polarization sub-channel that maps the anchor modal semantic features to reconstruct the cross-modal semantic prior vector. The polarization subchannel containing the target modal semantic features is decoded and split using a continuous cancellation list decoding algorithm, and the algebraic path metric of each split path is calculated cumulatively. During the decoding splitting process, the corresponding target semantic features are reconstructed based on the decoded bit sequence on the current splitting path, and the cross-modal semantic distortion degree between the target semantic features and the cross-modal semantic prior vector is calculated. By combining the polarization feature parameters of the current polarization subchannel and the cross-modal semantic distortion, the algebraic path metric is adjusted to generate a joint path metric. Based on the joint path metric, the split path is pruned to complete the polar code decoding and output the final decoded sequence.

[0007] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. Multimodal data is decoupled into anchor features, which are extremely sensitive to bit errors, and target features, which are redundant and fault-tolerant. A Gaussian approximation algorithm is used to bind the core anchor features to a polarized sub-channel with mutually information capacity limits, and the target features are dynamically mapped according to semantic entropy weights. This ensures that even under extremely harsh deep fading channels, the core semantic content is preserved despite the loss of massive amounts of edge details, avoiding catastrophic communication interruptions.

[0008] 2. Traditional SCL decoders rely solely on algebraic path metrics (PM) for blind pruning, which easily leads to the incorrect rejection of correct paths under low signal-to-noise ratio (SNR). This invention prioritizes locking the absolutely correct anchor point prior vector, reconstructs the local target semantics midway through polarization tree splitting, and generates a joint path metric by combining the confidence adjustment factor of the current polarization sub-channel. When the physical channel quality is extremely poor, the algorithm automatically reduces the underlying physical confidence and instead relies on multimodal macroscopic semantics for path preservation, thus reducing the false pruning rate caused by burst noise in the channel.

[0009] 3. Given the inherent high tolerance of multimodal data to local pixel-level errors, this invention forcibly intervenes and blocks the MAC layer's ARQ retransmission command in the extreme case where all surviving paths fail the Cyclic Redundancy Check (CRC). By comparing the final cross-modal semantic distortion, flawed paths with macroscopic logical coherence are extracted and directly delivered to the application layer, meeting the continuous and reliable communication requirements in complex scenarios. Attached Figure Description

[0010] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Wherein: Figure 1 This is a flowchart illustrating the overall process of a channel-polarization-oriented multimodal data semantic encoding / decoding and reliable transmission method according to an embodiment of the present invention. Detailed Implementation

[0011] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of the present invention.

[0012] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0013] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.

[0014] This invention is described in detail with reference to the schematic diagrams. When detailing the embodiments of this invention, for ease of explanation, the cross-sectional views illustrating the device structure may be partially enlarged, not adhering to the usual scale. Furthermore, the schematic diagrams are merely examples and should not be construed as limiting the scope of protection of this invention. In actual fabrication, the three-dimensional spatial dimensions of length, width, and depth should be included.

[0015] Furthermore, in the description of this invention, it should be noted that the terms "upper," "lower," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. These terms are used solely for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. In addition, the terms "first," "second," or "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0016] Unless otherwise explicitly specified and limited, the terms "installation," "connection," and "joining" in this invention should be interpreted broadly. For example, they can refer to fixed connections, detachable connections, or integral connections; similarly, they can refer to mechanical connections, electrical connections, or direct connections, or indirect connections through an intermediate medium, or internal connections between two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances. Example 1

[0017] Reference Figure 1 This is the first embodiment of the present invention, which provides a method for multimodal data semantic encoding and decoding and reliable transmission oriented to channel polarization, including: S1. Obtain multimodal source data and extract anchor modal semantic features and target modal semantic features from the multimodal source data.

[0018] It should be noted that, in order to break the status quo of the separation between the information source and the channel in traditional communication systems, this invention constructs a joint multimodal semantic feature extraction model oriented towards the transmission characteristics of polarized channels, transforming macroscopic heterogeneous data into a binary semantic feature stream with significant importance gradients, providing a priori basis for the subsequent unequal weight mapping of polarized sub-channels.

[0019] Furthermore, in practical communication scenarios, acquiring highly correlated multimodal source data is crucial. .in, Defined as anchor modal data (such as text description, core skeleton signaling, etc.), this modal data is small in volume but highly generalized in logic, and is extremely sensitive to bit errors. Loss of this data will lead to the collapse of the overall semantics. Defined as target modal data (such as high-resolution images, dense point clouds, etc.), this modal data is large in volume and highly redundant, and has a natural tolerance for local pixel-level distortion.

[0020] Furthermore, a pre-trained multimodal joint semantic coding network is deployed at the source. Structurally, this network includes anchor encoder branches and target encoder branches.

[0021] Specifically, in this embodiment, the pre-trained multimodal joint semantic coding network is input to a training set containing pairs of anchor-target data (e.g., text-image pairs). During the training phase, target modality reconstruction loss (mean squared error MSE), cross-modal contrastive loss (InfoNCE Loss), and a total loss function constraining the feature distribution are defined. The AdamW optimizer is employed, combined with a channel simulator (incorporating random additive white Gaussian noise) for end-to-end backpropagation training.

[0022] Specifically, regarding anchor point modal data A Transformer network architecture based on a multi-layer bidirectional self-attention mechanism is used as the anchor encoder. After word embedding and positional encoding, the input data sequence undergoes global self-attention operations in each hidden layer, outputting a continuous-state anchor feature vector that aggregates global logic. ,in This represents the anchor feature dimension. It's important to note that, because it carries absolute prior knowledge across modalities, this branch does not perform downsampling to preserve maximum semantic accuracy.

[0023] Specifically, for target modal data A hybrid network architecture combining depthwise residual convolution and a visual transformer (ViT) is employed as the target encoder. Input data is cascaded through multi-scale receptive fields of convolutional layers and then divided into... Each local feature patch is input into a self-attention layer, which outputs a continuous-state target feature tensor. ,in For target feature dimensions.

[0024] Furthermore, since the physical layer polarization encoder can only process binary bitstreams, the aforementioned continuous-state floating-point features need to be discretized and mapped. In this embodiment, a nonlinear quantization activation function with a learnable step size is used to perform dimension-preserving binarization compression on the aforementioned continuous-state anchor feature vector and continuous-state target feature tensor, generating anchor bitstreams respectively. With the target bitstream sequence set ,in It is the total length of the anchor bitstream, that is, the specific number of binary bits (0 or 1) contained in the anchor bitstream sequence.

[0025] It should be noted that since the derivative of the standard binarized function is zero at all non-zero points, direct use will cause gradient vanishing during backpropagation, making the encoding network untrainable. Therefore, this invention uses hard quantization to generate a 0 or 1 bitstream during forward propagation, while employing a pass-through estimator algorithm during backpropagation gradient calculation. This algorithm uses an identity mapping or a smoothed Tanh derivative to approximate the derivative of the step function, allowing the loss gradient of the multimodal network to seamlessly penetrate the nonlinear quantization layer, achieving joint optimization of source semantic extraction and underlying physical polarization mapping.

[0026] Furthermore, in order to map target bits with different redundancies to polarization sub-channels with varying reliability, it is necessary to process each local feature sub-block in the target bit stream set. ( The importance assessment is performed. This invention introduces a semantic entropy weighting evaluation index based on the coupling of energy distribution and information gain to calculate the importance of the target modality. Semantic entropy weights of local feature sub-blocks : in, Represents the continuous-state target feature tensor The Middle Sub-block, the first The activation response value of each dimension. Indicates the first The total number of feature dimensions contained in each sub-block. The hyperparameter representing the adjustment that balances global energy and local attention dependence (empirical range taken as...) ). A scaling factor set to prevent gradient vanishing due to excessively small activation values. : indicates the first The sum of normalized attention weights (i.e., cross-block connection tightness) obtained by each sub-block in the self-attention layer of the target encoder.

[0027] It should be noted that dimensions with larger activation response values ​​often represent high-frequency sensitive information such as image edges and subject contours. When the semantic entropy weights are calculated comprehensively... The larger the value, the more severe the macroscopic semantic distortion will be due to the loss of feature sub-blocks when reconstructing the target modality.

[0028] Furthermore, based on the calculated semantic entropy weight set Arranging the target bitstream sequence set in descending order yields a target modal semantic feature sequence with gradients, thus providing accurate numerical basis for unequal weight mapping of the mutual information capacity of polarization sub-channels in decreasing order.

[0029] S2. Calculate the reliability of polarized sub-channels based on channel polarization effect, sort the polarized sub-channels according to the reliability values, map the anchor modal semantic features to the polarized sub-channels that satisfy the first preset sorting interval, map the target modal semantic features to the polarized sub-channels that satisfy the second preset sorting interval, and perform polarization coding and transmission.

[0030] It should be noted that traditional polar code coding schemes often treat the source as an indiscriminate random bit stream, following a single criterion of transmitting information bits in the best channel and freezing bits in the worst channel. Based on this, this invention constructs a semantic-aware, channel-capacity-adaptive asymmetric mapping mechanism, which tilts limited high-quality physical channel resources toward key features that determine the success or failure of macroscopic semantics, thereby achieving degradation rather than catastrophic outages under poor channel conditions.

[0031] Furthermore, before performing physical layer mapping, the receiver or base station estimates the current physical channel state using pilot signals and obtains the noise variance of additive white Gaussian noise. Based on this channel state information, the Gaussian Approximation (GA) algorithm is used to track the mean evolution of the log-likelihood ratio (LLR) in the polar code decoding tree, thereby calculating the mother code length. ( The reliability of all polarized subchannels under ( ).

[0032] Specifically, assuming the mean LLR of the initial physical channel is For the first polarization tree Sub-channels ( The mean LLR at different polarization levels is recursively calculated according to the following rules: (i) For the left node of the polarization tree (corresponding to the modulo-2 addition operation), its mean evolution formula is: (ii) For the right node of the polarization tree (corresponding to the bit accumulation operation), its mean evolution formula is: In the above formula, and These represent the average LLR values ​​of the two branches fed into the current polarization unit from the previous polarization level, respectively; It is a nonlinear capacity evolution kernel function under a continuous additive white Gaussian noise channel (which can be quickly approximated by a preset parameterized piecewise function).

[0033] Specifically, after completing the recursion at all levels, we obtain the first... The final LLR mean of each polarization sub-channel Furthermore, this is converted into the mutual information capacity of the sub-channel. : in, For mutual information conversion functions, These are the standard preset fitting constants.

[0034] It is important to emphasize that the calculated mutual information capacity This represents the absolute reliability of the polarization sub-channel. The closer the value is to 1, the stronger the channel's ability to transmit bits without errors.

[0035] Furthermore, based on the mutual information capacity calculated above, the total... The polarization sub-channels are arranged in descending order. The length generated in S1 is... anchor bitstream Directly allocated to mutual information capacity Greater than the preset capacity threshold (The threshold value range is) In the polarization sub-channels of ), the set of polarization sub-channels that meet the threshold conditions constitutes the first preset sorting interval.

[0036] It should be noted that anchor point modal data (such as text labels, target skeleton coordinates, etc.) is the foundation of multimodal reconstruction. Once a single bit flip occurs, it will cause severe semantic drift in the cross-modal prior vector reconstructed by the receiver. Therefore, by setting a capacity threshold, absolute protection at the channel limit level is provided for the anchor point bit stream to ensure that even in a deep fading environment, even if a large amount of image details are lost, the receiver can still decode the main logic of the target.

[0037] Furthermore, after removing the polarization sub-channels occupied by the anchor mode, the set of remaining available polarization sub-channels is obtained, and they are strictly selected according to their mutual information capacity. The values ​​are sorted in descending order, and this sorted sequence constitutes the second preset sorting interval. Subsequently, the calculated semantic entropy weight set of the target feature sub-blocks is extracted. ,in accordance with Arrange the corresponding target bitstream in descending order. The target information bits are mapped sequentially and bit by bit into the polarization sub-channels within the second preset sorting interval. After all target information bits have been mapped, the remaining unmapped low-reliability polarization sub-channels are uniformly set to frozen bits (usually assigned a fixed value of 0).

[0038] It should be noted that the semantic value of different regions within the target modality varies drastically (e.g., the features of foreground faces are far more important than those of the background environment). Through the above operations, the communication system can identify key feature blocks (high local attention and high energy) that contain high local attention. Automatically matched to physical layer with relatively high capacity (high) The system automatically sacrifices the fidelity of redundant background features to ensure the clarity of the core target under sudden severe channel conditions where the signal-to-noise ratio deteriorates rapidly. This is achieved by aligning the physical channel capacity with the source semantic entropy.

[0039] Furthermore, after all mapping operations are completed, a length of [length missing] is generated. Mixed bit vector This includes anchor information bits, target information bits, and freeze bits. Subsequently, a polar code generation matrix is ​​used. Perform physical layer encoding. Finally, the encoded physical layer codeword sequence... After being modulated by baseband such as BPSK or QAM, the signal is sent to the radio frequency front end and transmitted to the receiving end via a wireless physical channel.

[0040] Specifically, the polar code generator matrix The process of performing physical layer coding is represented as follows: in, , This is a bit reversal permutation matrix. Polarization kernel matrix of Cronek's product.

[0041] S3. The receiver acquires the channel output signal and calculates the initial log-likelihood ratio of the polarization sub-channel. It prioritizes decoding the polarization sub-channels mapped with anchor modal semantic features to reconstruct the cross-modal semantic prior vector.

[0042] It should be noted that traditional polar code receivers perform an indiscriminate search of the entire codeword with equal weight during decoding. Based on this, we introduce an anchor-priority asymmetric decoding mechanism. Furthermore, since the anchor features determining the macroscopic logic are already bound to the physically reliable polar sub-channel in S2, at the receiver, the communication system first extracts the correct semantic skeleton with extremely low computational complexity, using this as the direction for splitting and pruning the massive target modal data in the polar decoding tree.

[0043] Furthermore, after receiving the baseband analog signal that has been fading and interfered with by the physical channel, the RF front-end at the receiving end converts it into a discrete signal sequence through sampling and decision-making. Assuming the current channel is an additive white Gaussian noise channel, the mathematical expression of the received signal is: ,in For the encoded bits to be sent, This is channel noise. The receiver uses this to calculate the initial log-likelihood ratio (LLR) sequence for each bit at the root node of the polar code decoding tree. The calculation formula is: in, and These respectively represent the signals received. Under the condition that the original bit at the sending end is 0 or 1, the posterior probability. The noise variance of the current physical channel can be obtained in real time through the channel estimation module. The sign of the value represents the tendency to make a hard decision (0 for positive and 1 for negative), and the magnitude of its absolute value represents the confidence or reliability of the decision.

[0044] Furthermore, based on the aforementioned initial log-likelihood ratio, the receiver initiates the polar code decoding process. Since the polar sub-channels mapped with anchor mode semantic features possess extremely high mutual information capacity (approaching a noiseless channel), their channel polarization effect has amplified their confidence to its limit. Therefore, for this sub-channel, there is no need to initiate the highly complex Successive Cancellation List (SCL) decoding split; instead, low-latency single-path hard decisions are made sequentially according to the inherent Successive Cancellation (SC) decoding timing of the polar codes. For the first... For each polarization subchannel with anchor point characteristics, the updated posterior log-likelihood ratio at the receiver is: The anchor point recovery bit is obtained by directly performing a hard decision operation. : It should be noted that by peeling the bits sequentially in this order, the anchor bit stream at the receiving end can eventually be obtained. The anchor bitstream is a one-dimensional binary sequence.

[0045] Furthermore, to obtain Subsequently, it cannot be directly used for backend path measurement; it must be projected back from the discrete binary physical space to a high-dimensional continuous semantic space. Therefore, this embodiment deploys a multimodal semantic decoder at the receiving end that matches the sending end. The core module of this decoder processes the following: Inverse quantization and embedding layers: The vectors are divided into fixed-width groups and then transformed into continuous-state floating-point feature tensors using an inverse mapping function and a learnable embedding dictionary. .

[0046] Multi-head self-attention reconstruction layer: The input is a deep network containing multiple Transformer decoding blocks. This network captures the global contextual dependencies within the anchor data through the core computational formula of the attention mechanism, repairing minor quantization distortions that may exist under extreme channel conditions.

[0047] Cross-modal semantic projection layer: Dimension alignment is performed through a fully connected linear layer, outputting a deterministic cross-modal semantic prior vector. : in, This represents the feature tensor after processing by a multi-head self-attention network. and These are the learnable weight matrix and bias vector of the projection layer, respectively; A smooth activation function (GELU in this embodiment) is used to ensure the continuity of the output vector space; To unify the cross-modal feature alignment dimension, this vector can be in the same mathematical metric space as the features generated by the subsequent target modality.

[0048] Furthermore, the aforementioned deterministic cross-modal semantic prior vectors It is written to and latched in the high-speed SRAM decoding register dedicated to the polar code decoder.

[0049] It should be noted that latching operations are crucial in the underlying hardware of communication receivers. Because the decoding tree generates numerous split paths during SCL decoding of the target modality features using polar codes, computations are extremely frequent. By pre-generating and storing these paths in a high-speed register, any decoding split path related to the target modality can be retrieved from the register at any time with low latency for semantic cosine similarity comparison. This allows for microscopic physical layer path pruning based on macroscopic semantics without increasing additional network forward propagation overhead.

[0050] S4. The polarization sub-channel containing the target modal semantic features is decoded and split using a continuous cancellation list decoding algorithm, and the algebraic path metric of each split path is accumulated and calculated.

[0051] It should be noted that in communication S3, the system has already stripped and latched the crucial anchor point prior knowledge with extremely low latency. However, due to the large amount of target modal data (such as image pixel sequences) and the inconsistent reliability of the polarization sub-channels it resides in, it is susceptible to local bit flips caused by burst noise. If a misjudgment occurs at any step in traditional Successive Cancellation (SC) decoding, the error will propagate avalanche along the polarization tree. Therefore, it is necessary to activate the Successive Cancellation List (SCL) decoding algorithm for the polarization sub-channel where the target modal features are located.

[0052] Furthermore, before polarization tree decoding begins, the full polarization information bit sequence corresponding to the target modality semantic features to be decoded is divided into... A fixed-length sub-block. Assume the total length of the target information bits is... The set sub-block length is (Usually the value is taken as) or Then the partitioned set of sequences can be represented as , of which Sub-block .

[0053] It should be noted that the fixed-length sub-block partitioning is used because in a multimodal semantic encoder, a single isolated bit (0 or 1) has no macroscopic semantic meaning and cannot be input into the neural network for feature extraction. If semantic evaluation is performed only after the entire codeword is decoded, the opportunity for timely pruning during the decoding tree search is lost, leading to severe waste of computing power and high latency. By setting the sub-block length, the minimum input dimension (i.e., local receptive field) of the neural network's forward propagation is matched, while ensuring that the decoder can accumulate sufficient local features at the physical layer before performing cross-modal semantic verification with actual physical meaning.

[0054] Furthermore, during bit-by-bit decoding within each sub-block, the receiver needs to determine the current splitting path based on the butterfly signal flow graph (i.e., polarization tree structure) of the polar code. For each node in the polarization tree, recursively calculate the posterior LLR. Assume the current processing reaches the [number]th [node] of the polarization tree. Layer, node length is The LLRs input to the upper and lower branches of this layer are respectively and ,in This represents the soft information (i.e., LLR value) that is transmitted from the previous level to the branch on the current butterfly computing unit. This represents soft information that is transmitted from the previous level to the current branch of the butterfly computing unit.

[0055] Specifically, for the left node of the polarization tree, its physical essence is a modulo-2 addition operation, equivalent to the cascading of two parallel channels, and its reliability is limited by the weaker one. To reduce hardware computational complexity, the system performs a sign-minimum approximation operation, and the update formula is: Specifically, for the right node of the polarization tree, its physical essence is an information accumulation operation after knowing the decision result of the left side, equivalent to diversity reception. The communication system combines the decided bit values ​​(i.e., local sums) Perform the accumulation operation and update the formula as follows: in, For symbolic functions, To find the minimum function. Through the above... Nodes and By recursively alternating nodes, the current path can be obtained when the decoding depth reaches a leaf node of the polarization tree. Next, the The posterior log-likelihood ratio of the bits to be decoded .

[0056] Furthermore, when decoding reaches the unfrozen information bits of the target mode, the SCL algorithm no longer retains only a single hard decision result like the SC algorithm, but instead retains both branches for decision 0 and decision 1, thus doubling the number of currently surviving paths (i.e., decoding splits). To quantify the reliability of each split path at the physical layer, an algebraic path metric (PM) must be calculated.

[0057] Specifically, the definition of the first The path in the decoding The algebraic path metric after bits is For the current bit to be decoded, according to its a posteriori LLR... The calculated ideal hard decision symbol is When the algorithm attempts to split the current bit (i.e., guesses) it... At that time, the update rule for its algebraic path metric is as follows: It should be noted that, due to The magnitude of this value represents the confidence level of the physical channel in determining the polarity of that bit. When the guess value of the split path... When the hard decision sign aligns with the channel observations, it indicates that the path conforms to the physical layer observations and requires no penalty (PM does not increase). However, when a split path, in order to explore potential noise flips, forcibly makes a guess contrary to the channel observations (i.e., opposite sign), a cost must be paid. In this case, the absolute value of its posterior log-likelihood ratio is... The penalty is added to the PM value of the path. The smaller the PM value of the algebraic path metric, the higher the likelihood probability of the path at the purely physical level. However, it should be noted that in some severe fading channels, relying solely on this low-level algebraic addition method can easily lead to local optima, causing paths with correct semantic macroscopic logic to be mistakenly identified as high-penalty paths and prematurely eliminated.

[0058] S5. During the decoding and splitting process, the target semantic features are reconstructed based on the decoded bit sequences on the current splitting path, and the cross-modal semantic distortion between the target semantic features and the cross-modal semantic prior vector is calculated.

[0059] It should be noted that in S4 above, the communication system relies solely on the underlying algebraic path metric (PM) to evaluate the reliability of the physical layer. However, in deep fading channels with ultra-low signal-to-noise ratios or sudden interference, the log-likelihood ratio (LLR) of the pure physical layer is prone to failure, causing split paths with truly correct macroscopic logic to be penalized with high penalties due to local noise. Therefore, this invention uses cross-modal semantic priors as the direction of decoding, mapping the messy bitstream at the bottom layer back to a high-dimensional semantic space during the polarization tree splitting process for semantic verification.

[0060] Furthermore, the target semantic features are reconstructed based on the decoded bit sequences on the current splitting path.

[0061] Specifically, during the continuous cancel list (SCL) decoding process, the communication system's built-in counter monitors the current bit index in real time. Since isolated, sporadic bits lack semantic representation, the communication system only checks the bit index when decoding progresses to any preset sub-block boundary (i.e., the current bit index = m × B, where...). The semantic reconstruction mechanism is triggered only when the current sub-block number is reached. At this point, the current sub-block number is extracted. On each split path, all bits that have completed hard decisions, from the start bit of the polar code target information (i.e., the 1st bit) to the end bit of the current sub-block boundary (i.e., the mBth bit), constitute a cumulative local decoding bit vector. : It should be noted that the truncation here does not only truncate the current number. Within each sub-block Instead of individual bits, it extracts the entire history of decided bits from the beginning.

[0062] Furthermore, since the SCL decoding tree simultaneously retains dozens or even hundreds of candidate paths (determined by the list capacity), calling a massive deep neural network (such as ResNet or Transformer) for feature reconstruction at the arrival of each sub-block would result in catastrophic computational latency and overhead, making it practically impossible to implement in engineering. Therefore, this invention embeds a set of pre-trained offline linear semantic projection matrices within the decoder at the receiving end. The extracted local decoding bit vector Directly input into the corresponding preset linear semantic projection matrix Perform matrix multiplication and append bias vectors. Then, a nonlinear activation function is applied for mapping, outputting a local target semantic feature tensor corresponding to the decoding progress of the current sub-block. (i.e., target semantic features): Wherein, projection matrix , To unify the alignment dimension for multimodal applications. It is a nonlinear activation function with saturation characteristics (the Tanh function is selected in this embodiment).

[0063] Specifically, the training process of this pre-trained offline set of linear semantic projection matrices is as follows: Under the condition of freezing the feature extraction network at the transmitter, an autoencoder with local decoded bit input and global semantic feature output is constructed. During training, random flip noise conforming to the channel polarization characteristics is injected into the complete polar code information bits to simulate SCL decoding truncation states at different depths, resulting in noisy local decoded bit vectors. Mean square error loss is used to force the local target semantic feature tensor mapped by the linear projection matrix to approximate the complete target features output by the transmitter in the noiseless state.

[0064] It should be noted that the linear semantic projection matrix was obtained through end-to-end joint training of an autoencoder architecture consisting of a target encoder at the transmitting end and a linear decoder at the receiving end on a massive multimodal dataset before the communication system was deployed.

[0065] Furthermore, after obtaining the local target semantic feature tensor of the current path, the communication system immediately retrieves the absolutely correct cross-modal semantic prior vector from the high-speed decoding register pre-latched in S3 via the underlying hardware bus. (That is, the feature skeleton reconstructed from anchor data). Subsequently, the dot product of the two data points and their respective Euclidean norms (L2 norms) are calculated, and the dot product is divided by the product of the two norms to obtain the cross-modal cosine similarity of the current splitting path at the semantic level. : Furthermore, by subtracting the cross-modal cosine similarity from the numerical value, the final quantized cross-modal semantic distortion can be obtained. : It should be noted that in high-dimensional manifold spaces, the absolute amplitude of feature vectors often fluctuates due to signal fading or quantization precision, and the angle between their orientations determines their true semantic category. Using cosine similarity instead of mean squared error (MSE) can perfectly filter out feature amplitude scaling caused by some unimportant bit flips, measuring only the semantic angular deviation between the target image fragment parsed from the current path and the anchor text / skeleton.

[0066] S6. Combining the polarization feature parameters of the current polarization subchannel and the cross-modal semantic distortion, adjust the algebraic path metric to generate a joint path metric. Based on the joint path metric, prune the split path to complete the polarization code decoding and output the final decoding sequence.

[0067] It should be noted that in traditional Sequential Cancellation List (SCL) decoding, the accumulation of algebraic path metric (PM) is based on the assumption that the channel is memoryless and stationary. However, in real-world bursty and severe channels, one or two instantaneous noise spikes caused by deep fading can lead to correctly semantically correct decoding paths being wrongly penalized with extremely high PM penalties. Therefore, a dynamic penalty mechanism that is aware of the channel state is needed: when the physical channel quality is excellent, the underlying algebraic decision is trusted; when the physical channel is under extreme fading, the weight of the underlying physical decision is automatically reduced, and path preservation is instead relied upon based on multimodal semantic prior knowledge.

[0068] Furthermore, by combining the polarization feature parameters of the current decoding sub-channel with the cross-modal semantic distortion, the algebraic path metric is adjusted to generate a joint path metric, as follows: When decoding progresses to the... When defining the boundary of a sub-block, first extract the set of all polarization sub-channels within the current sub-block that have been assigned as target information bits. Then calculate the average mutual information capacity of all polarization sub-channels within this set. This is used as a polarization feature parameter to characterize the quality of the current local decoding environment: in, Indicates the first The set of information bit indexes contained within each sub-block This represents the total number of information bits within the sub-block. The first one calculated based on the Gaussian approximation algorithm Mutual information capacity of each polarization subchannel.

[0069] Subsequently, the average mutual information capacity was calculated. As the independent variable of the negative exponential term, it is input into the preset natural exponential decay function to calculate the polarization confidence adjustment factor. : in, The preset channel sensitivity control parameter (which is a constant, usually taken as...) ).when When the value approaches 1, it means that the physical channel in which the current sub-block is located is extremely reliable and the signal-to-noise ratio is extremely high. It decreases exponentially and rapidly, approaching 0; while when When it approaches 0 (meaning the current sub-block has fallen into the deep fading range, and the physical layer signal is almost completely submerged by noise), The rapid rebound approaches the maximum value of 1. This adjustment factor allows the system to adaptively quantify the degree to which the underlying physical signal should be suspected.

[0070] Next, obtain the current number calculated in S5. Cross-modal semantic distortion of split paths Combine it with the polarization confidence modulator. and the preset global semantic weight constant The three factors are multiplied consecutively to generate a dynamic semantic penalty term. : Finally, the algebraic path metric generated from the bottom layer of the polarization decoding tree in S4 is used. The dynamic semantic penalty term is added as a scalar to generate the final joint path metric used to guide list pruning. : It should be noted that traditional decoders rely solely on Selecting a path to survival. In this solution, This achieves a perfect cross-boundary coupling between Shannon's information theory foundations and deep learning semantic features. This means that even if a splitting path encounters sudden noise leading to several incorrect guesses at the physical layer, causing its algebraic penalty... The value is abnormally high, but as long as the reconstructed image features are highly consistent with the anchor semantics (such as the text skeleton) (i.e. ), and its superimposed semantic penalty terms It will be extremely small. In extremely poor polarization channels ( In the case of this, the semantically correct flawed path... The total value remains low, thus being preserved by the communication system as valuable correct logic, thereby avoiding a semantic cliff-like avalanche caused by blind pruning at the physical layer.

[0071] Furthermore, after the current sub-block is decoded and split, all sub-blocks in the current decoding tree are collected. Candidate splitting paths ( The joint path metric (based on the preset maximum list capacity). .in accordance with The numerical values ​​of all candidate splitting paths are strictly ordered in ascending order.

[0072] Specifically, the order of the truncations is based on the maximum capacity of the list. Ahead (i.e.) The smallest value Candidate splitting paths (those with certain patterns) are selected and retained for the next polarization decoding sub-block search. Meanwhile, other paths appearing later in the sequence are discarded. A number of candidate splitting paths are generated, freeing up hardware memory resources.

[0073] Furthermore, the polar code decoding is completed and the final decoded sequence is output, including the following two decision mechanisms based on the physical layer verification state.

[0074] Specifically, the first decision-making mechanism involves extracting the current reserved list after the last polarization sub-channel (i.e., the final leaf node of the polarization tree) containing the target modality semantic features has been decoded. The full information bit sequence corresponding to each surviving path is processed, and Cyclic Redundancy Check (CRC) is performed using the check bits at the end of the polar code. If multiple surviving paths pass the CRC check (although the probability is low, it is possible under short codes or special polar tree structures), the path with the smallest algebraic path metric is not blindly output. Instead, the full information bit sequences corresponding to each surviving path that passes the check are re-inputted into the target modality decoder at the receiver, and the continuous-state global target semantic features are decoded. The final cross-modal semantic distortion is calculated with respect to the pre-stored cross-modal semantic prior vector according to the same cosine similarity rule as S5. The surviving path with the smallest cross-modal semantic distortion is selected, and its corresponding full information bit sequence is output as the final decoded sequence.

[0075] Specifically, the second decision mechanism. If all surviving paths in the current reserve list fail the cyclic redundancy check, according to traditional communication protocols (such as 5G NR / LTE), the MAC layer will immediately discard the data packet and trigger a HARQ (Hybrid Automatic Repeat Request) instruction. However, for multimodal data, sporadic pixel-level errors often do not impair the human eye's or encoder's macroscopic perception of the image. Therefore, it is necessary to forcibly intervene in the physical layer's rigid verification mechanism to directly block the NACK (Non-Accept) retransmission request instruction fed back to the sender. Subsequently, the full information bit sequences corresponding to all surviving paths that failed the verification are extracted and input into the target modality decoder at the receiver to decode and generate global target semantic features, and the final cross-modal semantic distortion of each failed path is calculated. The surviving path with the absolute smallest cross-modal semantic distortion value (i.e., the most coherent macroscopic logic) is selected, and under the condition of completely ignoring and violating the cyclic redundancy check failure state, the full information bit sequence corresponding to the selected surviving path is directly forcibly output to the application layer as the final decoding sequence.

[0076] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A method for multimodal data semantic encoding / decoding and reliable transmission oriented to channel polarization, characterized in that, include: Acquire multimodal source data, and extract anchor modal semantic features and target modal semantic features from the multimodal source data; The reliability of polarized sub-channels is calculated based on the channel polarization effect. The polarized sub-channels are sorted according to the reliability values. The anchor modal semantic features are mapped to polarized sub-channels that satisfy the first preset sorting interval. The target modal semantic features are mapped to polarized sub-channels that satisfy the second preset sorting interval. Polarization coding and transmission are then performed. The receiver acquires the channel output signal and calculates the initial log-likelihood ratio of the polarization sub-channel. It then prioritizes decoding the polarization sub-channel that maps the anchor modal semantic features to reconstruct the cross-modal semantic prior vector. The polarization subchannel containing the target modal semantic features is decoded and split using a continuous cancellation list decoding algorithm, and the algebraic path metric of each split path is calculated cumulatively. During the decoding splitting process, the corresponding target semantic features are reconstructed based on the decoded bit sequence on the current splitting path, and the cross-modal semantic distortion degree between the target semantic features and the cross-modal semantic prior vector is calculated. By combining the polarization feature parameters of the current polarization subchannel and the cross-modal semantic distortion, the algebraic path metric is adjusted to generate a joint path metric. Based on the joint path metric, the split path is pruned to complete the polar code decoding and output the final decoded sequence.

2. The method for multimodal data semantic encoding / decoding and reliable transmission oriented to channel polarization as described in claim 1, characterized in that, The reliability of polarized sub-channels is calculated based on channel polarization effects, and the anchor point modal semantic features are mapped, including: Based on the noise variance of the current channel state, the mutual information capacity of all polar sub-channels under the polar code mother code length is recursively calculated using the Gaussian approximation algorithm. The anchor modal semantic features are converted into anchor bit streams, and the anchor bit streams are allocated to polarization sub-channels whose mutual information capacity is greater than a preset capacity threshold. Calculate the semantic entropy weights of each dimension in the target modal semantic features, convert the target modal semantic features into a target bit stream, and map the target bit stream sequentially to the remaining available polarization sub-channels with decreasing mutual information capacity according to the descending order of the semantic entropy weights. Set a freeze bit for the unmapped polarization sub-channels.

3. The method for multimodal data semantic encoding / decoding and reliable transmission oriented to channel polarization as described in claim 1, characterized in that, Prioritize decoding of polarimetric subchannels mapped with the anchor modal semantic features, including: Based on the initial log-likelihood ratio, the bits in the polar sub-channels mapped with the anchor modal semantic features are sequentially determined according to the serial cancellation decoding timing of the polar codes. The determined bitstream is input into the multimodal decoder to generate the deterministic cross-modal semantic prior vector, and the cross-modal semantic prior vector is latched in the decoding register.

4. The method for multimodal data semantic encoding / decoding and reliable transmission oriented to channel polarization as described in claim 1, characterized in that, The polarization sub-channel containing the target modal semantic features is decoded and split using continuous cancellation list decoding, and the algebraic path metric of each split path is calculated cumulatively, including: The polarization information bit sequence corresponding to the target modality semantic features is divided into multiple fixed-length sub-blocks; During the bit-by-bit decoding process within each sub-block, a symbol min-sum operation corresponding to the left node of the polar code tree is performed according to the polar code tree structure, or a combined decided bit accumulation operation corresponding to the right node of the polar code is performed to update the posterior log-likelihood ratio of the current bit. When the sign of the hard decision value of the current bit to be decoded is opposite to the sign of the posterior log-likelihood ratio, the absolute value of the posterior log-likelihood ratio is added to the algebraic path metric of the previous time step.

5. The method for multimodal data semantic encoding / decoding and reliable transmission oriented to channel polarization as described in claim 4, characterized in that, Reconstructing the target semantic features based on the decoded bit sequences on the current splitting path includes: When the continuous cancellation list decoding progresses to the boundary of any of the sub-blocks, all decided bits from the polar code start bit to the end bit of the current sub-block boundary of the current split path are extracted to form a local decoding bit vector; The local decoded bit vector is input into a preset linear semantic projection matrix to perform matrix multiplication, and a nonlinear activation function is applied for mapping. The output is a local target semantic feature tensor corresponding to the current sub-block decoding progress as the target semantic feature.

6. The method for multimodal data semantic encoding / decoding and reliable transmission oriented to channel polarization as described in claim 5, characterized in that, Calculate the cross-modal semantic distortion of the target semantic features and the cross-modal semantic prior vector, including: Calculate the dot product between the local target semantic feature tensor and the cross-modal semantic prior vector; Calculate the norm of the local target semantic feature tensor and the norm of the cross-modal semantic prior vector, respectively; Dividing the dot product result by the product of the two norms yields the cross-modal cosine similarity. The quantized cross-modal semantic distortion is obtained by subtracting the cross-modal cosine similarity from the numerical value.

7. The channel-polarization-oriented multimodal data semantic encoding / decoding and reliable transmission method as described in claim 1 or 4, characterized in that, The algebraic path metric is adjusted by combining the polarization feature parameters of the current polarization subchannel and the cross-modal semantic distortion to generate the joint path metric, including: The average mutual information capacity of all polarization sub-channels containing information bits within the sub-block where the current decoding is located is calculated as the polarization feature parameter; The mean mutual information capacity is used as the independent variable of the negative exponential term and input into the preset natural exponential decay function to calculate the polarization confidence adjustment factor. The polarization confidence adjustment factor, the preset global semantic weight constant, and the cross-modal semantic distortion are multiplied consecutively to generate a dynamic semantic penalty term; The algebraic path metric of the current split path is scalarly added to the dynamic semantic penalty term to generate the joint path metric.

8. The method for multimodal data semantic encoding / decoding and reliable transmission oriented to channel polarization as described in claim 1, characterized in that, Pruning the split path based on the joint path metric includes: Collect the joint path metric for all current candidate split paths; All candidate split paths are sorted in ascending order based on the numerical value of the joint path metric; Candidate split paths whose order is before the maximum capacity value of the preset list are selected as retained paths and entered into the subsequent polarization decoding tree search stage, while candidate split paths whose order is after the maximum capacity value of the preset list are discarded.

9. The method for multimodal data semantic encoding / decoding and reliable transmission oriented to channel polarization as described in claim 1, characterized in that, Complete the polar code decoding and output the final decoded sequence, including: After the last polarization subchannel containing the target modal semantic feature is decoded, the full information bit sequence corresponding to all live paths in the current retention list is extracted and cyclic redundancy check is performed. If multiple survival paths pass the cyclic redundancy check, the final cross-modal semantic distortion of the global target semantic features and the cross-modal semantic prior vector corresponding to the survival paths that pass the check is calculated respectively. The survival path with the smallest final cross-modal semantic distortion value is selected, and its corresponding full information bit sequence is output as the final decoding sequence.

10. The method for multimodal data semantic encoding / decoding and reliable transmission oriented to channel polarization as described in claim 9, characterized in that, Also includes: If none of the live paths in the current retention list pass the cyclic redundancy check, then the retransmission request instruction is blocked. Extract the global target semantic features corresponding to all surviving paths that fail the verification, and calculate the final cross-modal semantic distortion of each surviving path and the cross-modal semantic prior vector; The survival path with the smallest final cross-modal semantic distortion value is selected. If the cyclic redundancy check fails, the full information bit sequence corresponding to the selected survival path is output as the final decoding sequence.