Dynamic adaptive image semantic transmission method based on state space model

The dynamic adaptive image semantic transmission method, which uses a state-space model and adaptive modulation of signal-to-noise ratio sequence, solves the bottleneck of transmission efficiency and robustness of the end-to-end JSCC method in the vehicle-road cooperative environment, and realizes efficient and robust image transmission.

CN120769302BActive Publication Date: 2025-12-26CHINA JILIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511261333.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-05
Publication Date
2025-12-26
Estimated Expiration
2045-09-05

AI Technical Summary

Technical Problem

Existing end-to-end JSCC methods suffer from insufficient transmission efficiency and robustness when faced with drastic channel state fluctuations and bandwidth constraints in vehicle-road cooperative environments, and cannot effectively balance high dynamism and semantic information fidelity.

Method used

A dynamic adaptive image semantic transmission method based on a state-space model is adopted. Through state-space coding, signal-to-noise ratio sequence adaptive modulation, semantically aware bandwidth selection, and reverse decoding, adaptive processing of channel time-varying characteristics and selection and compression of key semantic information are achieved.

Benefits of technology

It improves the efficiency and robustness of image transmission, reduces processing latency, optimizes the utilization of limited bandwidth resources, and is suitable for dynamic vehicle-road cooperative environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120769302B_ABST
    Figure CN120769302B_ABST
Patent Text Reader

Abstract

The application discloses a dynamic adaptive image semantic transmission method based on a state space model, inputs an image into a state space backbone network, constructs an encoder based on a multistage Mamba block, and efficiently models long-distance dependency by using the linear calculation complexity characteristics of SSM; then, the encoded features are input into an SNR sequence adaptive modulator, a modulation vector is generated by dynamically analyzing a historical SNR sequence through a lightweight Mamba network, and the modulation vector is multiplied with a feature map at a channel level to fuse the modulation vector with the feature map, so that the feature expression is adaptive to the time-varying characteristics of a channel; the modulated features are input into a semantic perception bandwidth selector, the semantic importance weight of the features is learned, an importance mask is dynamically generated according to a target channel bandwidth ratio to perform channel-level screening, key semantic information is reserved, and redundant data is compressed; and the screened features are transmitted to a receiving end through a channel, are calibrated by an SNR sequence adaptive demodulator, and are reconstructed into a high-quality image by a decoder of the backbone network.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of image transmission and semantic communication, and particularly relates to a dynamic adaptive image semantic transmission method based on a state space model. BACKGROUND

[0002] In the process of promoting the evolution of automatic driving and intelligent transportation systems by wireless communication technology, efficient and reliable information interaction in a vehicle-road cooperative environment has become a key enabling factor. Among them, the real-time transmission of high-fidelity image data is crucial to improving the environmental perception and decision-making ability of automatic driving vehicles. In recent years, semantic communication as a breakthrough communication paradigm is attracting widespread attention. In this field, end-to-end joint source-channel coding technology based on deep learning represents an important development direction. Existing end-to-end JSCC methods optimize source compression and channel transmission strategies jointly. In relatively stable and slowly changing channel conditions, static or quasi-static scenarios, these methods have preliminarily verified their significant potential in bandwidth utilization and reconstruction quality. However, in the face of highly dynamic vehicle-road cooperative application environments, these end-to-end JSCC methods that perform well under ideal conditions face severe real-world challenges, with the core being the dual bottleneck of transmission efficiency and robust adaptive ability.

[0003] First, the high-speed mobility of vehicles, complex multipath effects, random obstacle occlusion, and Doppler shift factors all contribute to the dramatic and unpredictable instantaneous fluctuations of the channel state. However, most existing JSCC architectures are trained and deployed based on specific or fixed channel condition assumptions. This makes them significantly worse when faced with rapid and large fluctuations in real channel conditions beyond the training assumptions. Second, existing methods have structural defects in managing semantic information fidelity under bandwidth constraints. To meet strict wireless channel bandwidth limitations, existing solutions usually employ some form of hard compression strategy to control the amount of data transmitted. However, this "one-size-fits-all" compression paradigm has major drawbacks: it often treats all semantic features to be transmitted without discrimination.

[0004] Therefore, breaking through the two core limitations of the existing end-to-end JSCC framework, which are the lack of adaptability when facing channel time-varying and the insufficient semantic optimization when dealing with bandwidth constraints, exploring a new approach to end-to-end semantic information transmission that can effectively balance transmission efficiency and robustness under high dynamic and high complexity vehicle-road cooperative wireless channel conditions, is particularly urgent and of great significance for truly enabling real-time and reliable perception of the next generation of intelligent transportation systems. SUMMARY

[0005] In order to solve the problems in the prior art, improve the adaptive ability of environment perception of the autonomous vehicle, and effectively balance the transmission efficiency and robustness of the end-to-end semantic information transmission, the present application adopts the following technical solutions:

[0006] The dynamic adaptive image semantic transmission method based on the state space model comprises the following steps:

[0007] Step 1: Obtain an image and perform state space encoding, utilize the linear calculation complexity of the state space model to extract features, efficiently model long-range dependencies, realize fast encoding of a high-resolution image, and obtain a feature map;

[0008] Step 2: Perform signal-to-noise ratio sequence adaptive modulation on the feature map, dynamically analyze historical signal-to-noise ratio sequences, generate a multi-dimensional modulation vector, fuse the multi-dimensional modulation vector with the feature map, so that the modulated features have adaptive channel time-varying characteristics;

[0009] Step 3: Perform semantic perception bandwidth selection on the modulated features, learn the semantic importance of feature channels, dynamically generate a channel importance mask according to a target channel bandwidth, filter key semantic information and compress redundant data according to the channel importance mask, and obtain filtered features;

[0010] Step 4: The filtered features are transmitted to a receiving end through a wireless channel, the receiving end reconstructs the features based on the channel importance mask, performs signal-to-noise ratio sequence adaptive demodulation to compensate for channel distortion, obtains calibrated features, and then performs state space decoding on the calibrated features to gradually reconstruct the image by using reverse state space modeling.

[0011] Further, the step 1 comprises the following steps:

[0012] Step 1.1: Perform two-dimensional block embedding on the image, divide the image into a plurality of non-overlapping two-dimensional blocks, flatten each block and map it to a high-dimensional feature space through linear projection to form an initial block sequence, this process converts the spatial structure of the image into a feature representation, and the output resolution is halved and the number of channels is multiplied;

[0013] Step 1.2: The initial block sequence is scanned through a visual state space block in two dimensions, the two-dimensional image features are unfolded into a one-dimensional sequence through a plurality of scanning paths in different spatial directions; a selective state space model is constructed to calculate the sequence generated by each path in parallel, the weights are dynamically adjusted to capture long-range relationships in the sequence, semantic features are obtained, and the semantic features of the path sequence are integrated to obtain a two-dimensional feature map;

[0014] Step 1.3: merging adjacent two-dimensional blocks, deep feature extraction by stacking multiple visual state space blocks to obtain the feature map, the hierarchical structure gradually improves the semantic abstraction ability of the feature by balancing spatial downsampling and channel expansion.

[0015] Further, in step 1.2, the two-dimensional selective scanning is performed from the top left corner to the bottom right corner, from the bottom right corner to the top left corner, from the top right corner to the bottom left corner, and from the bottom left corner to the top right corner, traversing the initial block in four different directions, so that each pixel block aggregates context information from four complementary directions, covering all spatial position relationships, laying the foundation for capturing global dependencies.

[0016] Further, in step 1.2, the selective state space model generates dynamic parameters for sequence elements in the scanning path, including state evolution rate control parameters, dynamic weighted input weight vectors, and dynamic weighted output weight vectors. The state evolution rate control parameters are used to adjust the projection intensity, and the discrete state transition matrix is generated based on the state evolution rate control parameters and the continuous system matrix to control the state decay rate. The dynamic weighted input weight vector is used to select key channels, and the discrete input projection matrix is generated based on the dynamic weighted input weight vector and the state evolution rate control parameter to dynamically scale the influence intensity of the input signal. The result of multiplying the hidden state at the previous time by the discrete state transition matrix is summed with the result of multiplying the discrete input projection matrix by the sequence element to obtain the hidden state at the current time. The hidden state at the current time is multiplied by the dynamic weighted output weight vector to obtain the semantic feature.

[0017] Further, in step 1.2, the semantic feature integration of the path sequence is to reshape the output of each path sequence to the original spatial dimension, and then merge the output feature maps corresponding to the same spatial position from multiple paths respectively. This step reconstructs the spatial structure while preserving the global context accumulated from the scanning path. The semantic feature is mapped back through the scanning path to restore the one-dimensional sequence to a two-dimensional image, and the output features of the pixel points in multiple scanning paths are weighted and fused to obtain a two-dimensional feature map.

[0018] Further, in step 1.3, a plurality of units composed of a merging operation and a plurality of stacked visual state space blocks are connected in sequence to obtain the two-dimensional feature. The merging operation is used to concatenate adjacent two-dimensional blocks along the channel dimension, which halves the resolution and doubles the number of channels. Then, linear projection is used for dimension reduction. Multiple visual state space blocks are used for deep feature extraction, maintaining the resolution. The hierarchical structure gradually improves the semantic abstraction ability of the feature by balancing spatial downsampling and channel expansion.

[0019] Further, the step 2 comprises the following steps:

[0020] Step 2.1: Calculate embedding vectors for the signal-to-noise ratio values of the current time and historical time through the feature map, and obtain an embedding sequence of the signal-to-noise ratio sequence;

[0021] Step 2.2: The embedding sequence is dynamically analyzed by a lightweight Mamba network to perform time series modeling, and a multi-dimensional modulation vector is obtained after the generated lightweight output matrix passes through a modulation layer. The multi-dimensional modulation vector is fused with the feature map through channel-level multiplication to obtain a feature with real-time adaptive signal channel time-varying characteristics.

[0022] Further, the dynamic analysis in step 2.2 is to multiply the hidden state of the previous time with the state transition matrix, sum the result with the result of multiplying the embedding vector with the output matrix, and obtain the hidden state of the current time. Then, the output vector of the end state is obtained by multiplying the hidden state of the current time with the output matrix to generate the final output matrix.

[0023] Further, the step 3 comprises the following steps:

[0024] Step 3.1: Generate a modulation vector based on the target channel bandwidth ratio;

[0025] Step 3.2: Perform coarse-to-fine progressive modulation by multiplying a group of modulation vectors with the modulated features channel by channel, adjust the channel weights according to the target bandwidth ratio, and obtain a modulated feature map;

[0026] Step 3.3: Perform global average pooling along the spatial dimension to generate an importance score vector (spatial average of each channel of the feature map) for the modulated feature map, then perform channel importance sorting, combine the target channel bandwidth ratio to generate an importance mask, and multiply the importance mask with the modulated feature map element by element to retain important channels that need to be transmitted, achieve further compression, and obtain the screened features.

[0027] Further, the step 4 comprises the following steps:

[0028] Step 4.1: The screened features are transmitted to the receiving end through the wireless channel, and the receiving end reconstructs the features based on the importance mask, restores the features of the importance mask, and eliminates the noise generated by the non-importance mask;

[0029] Step 4.2: Perform signal-to-noise ratio sequence adaptive demodulation to compensate for channel distortion in the reverse process, and obtain calibrated features;

[0030] Step 4.3: Perform state space decoding on the calibrated features, and gradually reconstruct the image through reverse state space modeling.

[0031] The advantages and beneficial effects of the present application are that:

[0032] The present application proposes a multi-level Mamba structure to build a coding and decoding backbone, uses the long-distance dependence modeling capability of its linear complexity, and solves the calculation delay bottleneck in real-time processing of high-resolution images; the present application designs an SNR sequence adaptive modulator, dynamically analyzes the historical channel state sequence through a lightweight Mamba network, generates a refined modulation vector and performs channel-level fusion with the feature map, and solves the problem of response lag of traditional methods to instantaneous channel fluctuations; the present application introduces a semantic-aware bandwidth selection mechanism, dynamically evaluates the semantic importance of feature channels through learnable weights, realizes soft mask adaptive screening under bandwidth constraints, and significantly optimizes the utilization efficiency of limited bandwidth resources. BRIEF DESCRIPTION OF DRAWINGS

[0033] Figure 1 is a method flowchart of an embodiment of the present application.

[0034] Figure 2 is a structural schematic diagram of a visual state space block in an embodiment of the present application.

[0035] Figure 3 is a structural schematic diagram of an SNR sequence adaptive modulator / demodulator in an embodiment of the present application.

[0036] Figure 4 is a structural schematic diagram of a semantic-aware bandwidth selector in an embodiment of the present application.

[0037] Figure 5 is a structural schematic diagram of a target channel bandwidth ratio modulator in an embodiment of the present application.

[0038] Figure 6 is a structural schematic diagram of an importance mask generation module in an embodiment of the present application. DETAILED DESCRIPTION

[0039] The specific embodiments of the present application are described in detail below in conjunction with the accompanying drawings. It should be understood that the specific embodiments described herein are only used to illustrate and explain the present application, and are not used to limit the present application.

[0040] As shown in Figure 1 , the dynamic adaptive image semantic transmission method based on state space model uses state space model, channel adaptation, semantic communication and other technologies to improve the end-to-end semantic information transmission efficiency and robustness in dynamic driving environment, including the following steps:

[0041] Step 1: input visible light image into state space backbone network, use linear computational complexity of state space model (SSM) to efficiently model long-range dependencies, realize fast encoding of high-resolution image. By selecting state space mechanism to capture key spatial features, reduce processing delay, output multi-scale feature map, including the following steps:

[0042] Step 1.1: input visible light image I with size HxWx3 into encoder network. First, 2D block embedding is performed on the image: the input image is divided into multiple non-overlapping 2D blocks, each block is flattened and mapped to a high-dimensional feature space through linear projection to form an initial block sequence. This process converts the spatial structure of the image into a feature representation, which reduces the resolution by half and doubles the number of channels.

[0043] Step 1.2: input the initial block sequence into the visual state space block, and use the linear computational complexity of the state space model to efficiently extract features. As shown in Figure 2 The core of the visual state space block is 2D selective scanning, which expands 2D image features into 1D sequences through four scanning paths in the spatial diagonal direction for subsequent selective state space model (S6 module) processing.

[0044] The implementation process of the visual state space block is as follows: the input features are first normalized to stabilize the distribution, then linearly projected to expand the dimension, local spatial features are extracted through a deep convolutional layer, and nonlinearity is introduced through a SiLU activation function. The core module 2D selective scanning captures global long-range dependencies through four-way scanning; the output of 2D selective scanning is stabilized by layer normalization, then compressed to the original size by linear projection, and finally mixed by a feedforward neural network to enhance the expression ability. The entire process includes residual connection to preserve the original information and prevent network degradation.

[0045] The 2D selective scanning process is as follows:

[0046] First, traverse the blocks along four different directions: from the top left to the bottom right, from the bottom right to the top left, from the top right to the bottom left, and from the bottom left to the top right. Each path expands the two-dimensional structure into a one-dimensional sequence. This design enables each pixel block to aggregate context information from four complementary directions, covering all spatial position relationships and laying the foundation for capturing global dependencies. The feature map M is converted into four independent 1D sequences, each with a length of T = HxW, as follows:

[0047] Path 1 (top left → bottom right): ;

[0048] Path 2 (bottom right → top left): ;

[0049] Path 3 (top right → bottom left): ;

[0050] Path 4 (bottom left → top right): .

[0051] in, , Indicates the first Scan path, Represents the sequence elements, where Is the pixel in the path The Middle The position of the step.

[0052] Subsequently, the sequences generated for each path in the selective state-space model are processed in parallel by independent S6 modules, dynamically adjusting weights to capture long-range relationships within the sequences. The S6 modules, based on a structured state-space model, map the feature vectors of each input block to hidden states through a dynamic parameterization mechanism, enabling the model to adaptively filter important information while eliminating noise. Discretized state-space equations are applied to each sequence. Calculate independently; the specific formula is as follows:

[0053]

[0054] (State transition matrix)

[0055] (Input projection matrix)

[0056]

[0057]

[0058] in, Represents the input sequence elements (each sequence) elements in Directly used as input to the S6 module ), From the current input Dynamically generated (selective mechanism) dynamic parameters. Controlling the rate of state evolution ( Large → Long memory (Small → Short Memory) to adjust projection intensity. , These represent the dynamically weighted input and output weight vectors, respectively, allowing the model to focus on the relevant context. Used to select key channels, A represents the continuous system matrix, which provides the underlying physical laws of state evolution (such as the continuity prior of feature propagation). This represents the discrete state transition matrix and the actual control state decay rate. denotes the discrete input projection matrix, which scales the influence strength of the input signal dynamically, denotes the hidden state at time t, which is used to carry the semantic information flow across time, denotes the final output semantic feature.

[0059] Finally, the four-way merging stage re-integrates the processed four sequences into a 2D feature map. The output of each sequence is reshaped to the original spatial dimension, and then the four output feature maps from the four paths respectively corresponding to the same spatial location are merged. This step reconstructs the spatial structure while preserving the global context accumulated from the scanning paths. This process goes through sequence-to-image reorganization and weighted fusion, and the specific formula is as follows:

[0060]

[0061]

[0062] denotes the inverse mapping that restores the one-dimensional sequence to a two-dimensional image. denotes the learnable channel weight vector, denotes element-wise multiplication, In the first output feature of the pixel point in the first scanning path, the final output feature of each pixel point is the weighted sum of the outputs of the four scanning paths.

[0063] Step 1.3: The features output by the first stage (steps 1.1 and 1.2 constitute the first stage processing) enter the subsequent second, third, and fourth stages. The second stage (the structure is the same as the third and fourth stages) contains block merging operations and multiple stacked visual state space blocks: block merging concatenates adjacent blocks along the channel dimension (resulting in halving of resolution and doubling of channel number), and then linear projection is performed for dimension reduction; multiple visual state space blocks are stacked for deep feature extraction (maintaining resolution). This hierarchical structure balances spatial downsampling and channel expansion, gradually improving the semantic abstraction ability of features.

[0064] Step 2: Input the feature map output by step 1 into the SNR sequence adaptive modulator, as shown in Figure 3 , dynamically analyze the historical signal-to-noise ratio sequence (the training stage generates simulated sequences from the physical channel model, and the test stage uses the cached real-time historical SNR) through a lightweight Mamba network, generate a multi-dimensional modulation vector, and perform channel-level multiplication fusion with the feature map to make the feature expression adapt to the real-time channel time-varying characteristics, which includes the following steps:

[0065] Step 2.1: The features after the state space backbone network Input SNR sequence adaptive modulator. First, SNR is generated by SNR sequence generator to generate SNR sequence, and becomes learnable features through embedding layer. The specific formula can be expressed as:

[0066]

[0067]

[0068]

[0069] wherein, n represents the SNR sequence, the SNR value scalar history 6, current time 1, represents the embedding vector of the i-th SNR value, , is the embedding layer parameter, is the embedding dimension, represents the embedding sequence after embedding of the entire SNR sequence.

[0070] Step 2.2: The above generated learnable features are dynamically analyzed by light Mamba network to model the time sequence, and generate multi-dimensional modulation vector after modulation layer. The vector is fused with the feature by channel-level multiplication, and the processed feature is output, so that the feature expression can adapt to the real-time channel time-varying characteristics. The specific formula of light Mamba network can be expressed as:

[0071]

[0072] (state equation)

[0073] (output equation)

[0074] wherein, is the hidden state at time t, is the output vector at time t, is the light Mamba network function. is the state transition matrix, and a diagonalization matrix is used to reduce parameters, and the output matrix / Output matrix uses grouped convolution, and the output takes the end state: lightweight, represents the final output matrix of light Mamba network.

[0075] In summary, the entire process of steps 2.1 and 2.2 can be expressed as:

[0076]

[0077] where, denotes a Sigmoid function, ensuring the output value in the interval [0, 1], denotes a channel-wise multiplication.

[0078] Step 3: Modulated feature input semantic perception bandwidth selector. As shown in Figure 4 , this module learns the semantic importance of feature channels, dynamically generates a channel importance mask according to the target channel bandwidth ratio, filters key semantic information and compresses redundant data, which includes the following steps:

[0079] Step 3.1: The features output by the SNR sequence adaptive modulator are input into the semantic perception bandwidth selector. First, the target channel bandwidth ratio is generated through a three-layer fully connected layer by the target channel bandwidth ratio modulator, as shown in . Specifically, it can be represented as: Figure 5

[0080]

[0081]

[0082]

[0083] where, is the target channel bandwidth ratio modulation vector, CBR represents the target bandwidth ratio, the ReLU activation function introduces nonlinearity, and the Sigmoid compresses the output to the interval [0, 1], realizing the channel attention mechanism.

[0084] Step 3.2: Seven target channel bandwidth ratio modulators are connected in series to realize progressive modulation from coarse to fine, and the generated modulation vector is multiplied with the feature map channel by channel, adjusting the feature map channel weight according to the target bandwidth ratio.

[0085] Step 3.3: The modulated feature map generates a channel importance mask through the importance mask generation module, retaining important channels and further reducing the transmission bandwidth, as shown in Figure 6 . Specifically, the modulated feature map first performs global average pooling along the spatial dimension to generate an importance score vector (the spatial average of each channel of the feature map) S, and then performs channel importance sorting combined with the target channel bandwidth ratio to generate an importance binary mask , finally transmitting only important channels to achieve further compression. The formula can be expressed as:

[0086]

[0087] where, is the feature after progressive modulation, ​importance binary mask, final output of the semantic-aware bandwidth selector.

[0088] Step 4: The screened features are transmitted to the receiving end through the wireless channel. The receiving end first reconstructs the features based on the importance mask, recovers the features with the mask being 1, and eliminates the noise generated in the area with the mask being 0. Step 4.2: The SNR sequence adaptive demodulator is used to compensate for channel distortion. Step 4.3: The calibrated features are input into the state space decoder, and the high-quality visible light image is gradually reconstructed through reverse state space modeling.

[0089] Step 4.1: The screened features are transmitted to the receiving end through the wireless channel. The receiving end first reconstructs the features based on the importance mask, recovers the features with the mask being 1, and eliminates the noise generated in the area with the mask being 0.

[0090] Step 4.2: The SNR sequence adaptive demodulator is used to compensate for channel distortion.

[0091] Step 4.3: The calibrated features are input into the state space decoder, and the high-quality visible light image is gradually reconstructed through reverse state space modeling.

[0092] In summary, the dynamic adaptive image semantic transmission method based on the state space model of the application includes four key steps: first, input the visible light image into the state space backbone network, use the linear calculation complexity of the state space model (SSM) to efficiently model the long-distance dependence relationship, and realize the fast coding of the high-resolution image. Through the selective state space mechanism, the key spatial features are captured, the processing delay is reduced, and the multi-scale feature map is output; then, the feature map output in step 1 is input into the SNR sequence adaptive modulator, a multi-dimensional modulation vector is generated through the lightweight Mamba network dynamically analyzing the historical signal-to-noise ratio sequence (the simulation sequence is generated by the physical channel model in the training stage, and the real-time historical SNR is used in the test stage), and the vector is fused with the feature map through channel-level multiplication, so that the feature expression is real-time adaptive to the channel time-varying characteristics; then, the modulated features are input into the semantic-aware bandwidth selector. This module learns the semantic importance of the feature channels, dynamically generates a channel importance mask according to the target channel bandwidth ratio, screens the key semantic information and compresses the redundant data; finally, the screened features are transmitted to the receiving end through the wireless channel. The receiving end first reconstructs the features based on the importance mask, and then compensates for the channel distortion through the SNR sequence adaptive demodulator; the calibrated features are input into the state space decoder, and the high-quality visible light image is gradually reconstructed through reverse state space modeling.

[0093] ​The application significantly improves the transmission robustness in a dynamic channel, reduces the end-to-end processing delay, effectively improves the transmission efficiency and robustness of the semantic information of the image, breaks through the calculation bottleneck of the traditional CNN / Transformer architecture, is particularly suitable for a dynamic vehicle-road environment, provides an efficient solution for real-time image transmission scenes such as automatic driving and remote medical treatment, and has important significance for improving the reliability and safety of an automatic driving system.

[0094] The above examples are only used to illustrate the technical solutions of the present application, but not limit them; although the present application has been described in detail with reference to the foregoing examples, those skilled in the art should understand that the technical solutions recorded in the foregoing examples can be modified, or some or all of the technical features can be replaced by equivalents; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A dynamic adaptive image semantic transmission method based on a state space model, characterized in that The method comprises the following steps: Step 1: obtaining an image and performing state space coding, extracting features by using the linear computational complexity of the state space model, and obtaining a feature map; Step 2: performing signal-to-noise ratio sequence adaptive modulation on the feature map, dynamically analyzing historical signal-to-noise ratio sequences, generating a multi-dimensional modulation vector, and fusing the multi-dimensional modulation vector with the feature map, so that the modulated features have adaptive channel time-varying characteristics; Step 3: performing semantic perception bandwidth selection on the modulated features, dynamically generating a channel importance mask according to the target channel bandwidth, and screening key semantic information and compressing redundant data according to the channel importance mask to obtain screened features; Step 4: transmitting the screened features to a receiving end through a wireless channel, and after the receiving end reconstructs the features based on the channel importance mask, performing signal-to-noise ratio sequence adaptive demodulation to obtain calibrated features, and then performing state space decoding on the calibrated features to gradually reconstruct the image by using reverse state space modeling.

2. The dynamic adaptive image semantic transmission method based on state space model according to claim 1, characterized in that: The step 1 comprises the following steps: Step 1.1: performing two-dimensional block embedding on the image, dividing it into a plurality of non-overlapping two-dimensional blocks, flattening each block and mapping it to a high-dimensional feature space through linear projection to form an initial block sequence; Step 1.2: performing two-dimensional selective scanning on the initial block sequence through a visual state space block, expanding the two-dimensional image features into a one-dimensional sequence through a plurality of scanning paths in different spatial directions; constructing a selective state space model to calculate the sequence generated by each path in parallel to capture the long-range relationship in the sequence, obtaining semantic features, and integrating the semantic features of the path sequence to obtain a two-dimensional feature map; Step 1.3: merging adjacent two-dimensional blocks to extract deep features through a plurality of visual state space blocks stacked to obtain the feature map.

3. The dynamic adaptive image semantic transmission method based on state space model according to claim 2, characterized in that: In the step 1.2, the two-dimensional selective scanning traverses the initial block from the top left corner to the bottom right corner, from the bottom right corner to the top left corner, from the top right corner to the bottom left corner, and from the bottom left corner to the top right corner along four different directions to aggregate context information from four complementary directions for each pixel block.

4. The dynamic adaptive image semantic transmission method based on state space model according to claim 2, characterized in that: In the step 1.2, the selective state space model generates dynamic parameters for the sequence elements in the scanning path, including a state evolution rate control parameter, a dynamic weighting input weight vector, and a dynamic weighting output weight vector. The state evolution rate control parameter is used to adjust the projection intensity, a discrete state transition matrix is generated based on the state evolution rate control parameter and a continuous system matrix to control the state decay rate, the dynamic weighting input weight vector is used to select key channels, a discrete input projection matrix is generated based on the dynamic weighting input weight vector and the state evolution rate control parameter to dynamically scale the influence intensity of the input signal, the result of multiplying the discrete state transition matrix with the hidden state at the previous time is summed with the result of multiplying the discrete input projection matrix with the sequence element element by element to obtain the hidden state at the current time, and the hidden state at the current time is multiplied with the dynamic weighting output weight vector element by element to obtain the semantic features.

5. The dynamic adaptive image semantic transmission method based on state space model according to claim 2, characterized in that: The semantic feature integration of the path sequence in step 1.2 is to reshape the output of each path sequence into the original spatial dimension, and then merge the output feature maps from multiple paths corresponding to the same spatial position; the semantic feature is mapped back through the scanning path to restore the one-dimensional sequence to a two-dimensional image, and the output features of the pixel points in multiple scanning paths are weighted and fused to obtain a two-dimensional feature map.

6. The dynamic adaptive image semantic transmission method based on state space model according to claim 2, characterized in that: In step 1.3, a plurality of units composed of a plurality of stacked visual state space blocks connected in sequence are constructed to obtain the two-dimensional features, the adjacent two-dimensional blocks are spliced along the channel dimension through the merging operation, and then the linear projection dimension reduction is performed, and the plurality of visual state space blocks perform deep feature extraction.

7. The dynamic adaptive image semantic transmission method based on state space model according to claim 1, characterized in that: The step 2 includes the following steps: Step 2.1: Calculate the embedding vector for the signal-to-noise ratio value of the current time and the historical time through the feature map to obtain the embedding sequence of the signal-to-noise ratio sequence; Step 2.2: The embedding sequence is time-series modeled through dynamic analysis, and the output matrix generated after the modulation layer obtains a multi-dimensional modulation vector; the multi-dimensional modulation vector is multiplied with the feature map at the channel level to obtain a feature with real-time adaptive signal channel time-varying characteristics.

8. The dynamic adaptive image semantic transmission method based on state space model according to claim 7, characterized in that: The dynamic analysis in step 2.2 is to multiply the hidden state of the previous time with the state transition matrix, sum the result with the result of multiplying the embedding vector with the output matrix to obtain the hidden state of the current time, multiply the hidden state of the current time with the output matrix to obtain the output vector of the current time, and take the output vector of the end state to generate the final output matrix.

9. The dynamic adaptive image semantic transmission method based on state space model according to claim 1, characterized in that: The step 3 includes the following steps: Step 3.1: Generate a modulation vector based on the target channel bandwidth ratio; Step 3.2: Multiply a group of modulation vectors with the modulated features channel by channel to perform coarse-to-fine progressive modulation, adjust the channel weights according to the target bandwidth ratio, and obtain the modulated feature map; Step 3.3: The modulated feature map is globally averaged pooled along the spatial dimension to generate an importance score vector, and then the channel importance is sorted to generate an importance mask combined with the target channel bandwidth ratio; the importance mask is multiplied with the modulated feature map element by element to retain the important channels to be transmitted, and the filtered features are obtained.

10. The dynamic adaptive image semantic transmission method based on state space model according to claim 1, characterized in that: The step 4 includes the following steps: Step 4.1: The filtered features are transmitted to the receiving end through the wireless channel, and the receiving end reconstructs the features based on the importance mask to recover the features of the importance mask and eliminate the noise generated by the non-importance mask; Step 4.2: The signal-to-noise ratio sequence is adaptively demodulated to compensate for channel distortion in the reverse process to obtain the calibrated features; Step 4.3: The calibrated features are decoded in the state space to gradually reconstruct the image through the reverse state space modeling.

Citation Information

Patent Citations

  • Semantic transmission method for visible light and infrared fusion image of night vehicle road environment

    CN119360341A

  • Target detection method based on Mama feature fusion

    CN120298667A