Dynamic adaptive image semantic transmission method based on state space model
Through the state space model and signal-to-noise ratio sequence adaptive modulation, the problems of transmission efficiency and robustness of the end-to-end JSCC method in the vehicle-road cooperative environment are solved, and efficient image transmission and semantic information reconstruction are achieved, which is suitable for scenarios such as autonomous driving and telemedicine.
Patent Information
- Application Number
- CN202511261333.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-05
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-09-05
AI Technical Summary
Existing end-to-end JSCC methods lack transmission efficiency and robustness when faced with drastic fluctuations in channel states and bandwidth constraints in vehicle-road cooperative environments, and are unable to effectively balance high dynamics and semantic information fidelity.
A dynamic adaptive image semantic transmission method based on the state-space model is adopted to achieve fast image encoding, adaptive modulation, and screening and reconstruction of key semantic information through state-space coding, signal-to-noise ratio sequence adaptive modulation, semantic-aware bandwidth selection, and state-space decoding.
It improves the efficiency and robustness of image transmission, reduces processing delay, optimizes the utilization of limited bandwidth resources, and is suitable for dynamic vehicle-road collaborative environments.
Smart Images

Figure CN120769302A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image transmission and semantic communication, and in particular relates to a dynamic adaptive image semantic transmission method based on a state space model. Background Art
[0002] As wireless communication technologies drive the evolution of autonomous driving and intelligent transportation systems, efficient and reliable information exchange in vehicle-infrastructure cooperative environments has become a key enabling factor. Real-time transmission of high-fidelity image data is crucial for improving the environmental perception and decision-making capabilities of autonomous vehicles. In recent years, semantic communication, as a groundbreaking communication paradigm, has garnered widespread attention. Deep learning-based end-to-end joint source-channel coding (JSCC) represents a significant advancement in this field. Existing end-to-end JSCC methods, by jointly optimizing source compression and channel transmission strategies, have demonstrated significant potential in improving bandwidth utilization and reconstruction quality in relatively stable, static or quasi-static scenarios with slowly changing channel conditions. However, in the highly dynamic real-world application of vehicle-infrastructure cooperative systems, these end-to-end JSCC methods, which perform well under ideal conditions, face significant challenges, primarily due to the dual bottlenecks of transmission efficiency and robust adaptability.
[0003] First, factors such as the high-speed mobility of vehicles, complex multipath effects, random obstacle obstructions, and Doppler shift together cause the channel state to exhibit violent and unpredictable instantaneous fluctuations. However, most existing JSCC architectures are trained and deployed based on specific or fixed channel condition assumptions. This results in significant performance degradation when faced with rapid and large fluctuations in real channel conditions that exceed the training assumptions. Second, existing methods have structural defects in managing the fidelity of semantic information under bandwidth constraints. To meet the strict bandwidth limitations of wireless channels, existing solutions typically adopt some form of hard compression strategy to control the amount of data ultimately transmitted. However, this "one-size-fits-all" compression paradigm has a major drawback: it often treats all semantic features to be transmitted indiscriminately.
[0004] Therefore, it is urgent and of great significance to break through the two core limitations of the existing end-to-end JSCC framework: the lack of adaptability when facing channel time-variability and insufficient semantic optimization when dealing with bandwidth constraints. Exploring a new end-to-end semantic information transmission approach that can effectively balance transmission efficiency and robustness under highly dynamic and complex vehicle-road cooperative wireless channel conditions is particularly urgent and of great significance for truly enabling real-time and reliable perception of the next generation of intelligent transportation systems. Summary of the Invention
[0005] To address the shortcomings of existing technologies, improve the environmental perception and adaptability of autonomous vehicles, and effectively balance transmission efficiency and robustness for end-to-end semantic information transmission, the present invention adopts the following technical solutions:
[0006] The dynamic adaptive image semantic transmission method based on the state space model includes the following steps:
[0007] Step 1: Acquire the image and perform state-space encoding. Utilize the linear computational complexity of the state-space model to extract features, efficiently model long-distance dependencies, and achieve fast encoding of high-resolution images to obtain feature maps.
[0008] Step 2: Adaptively modulate the feature map using a signal-to-noise ratio sequence, dynamically analyze the historical signal-to-noise ratio sequence, generate a multidimensional modulation vector, and fuse the multidimensional modulation vector with the feature map so that the modulated feature has adaptive channel time-varying characteristics;
[0009] Step 3: Performing semantic-aware bandwidth selection on the modulated features to learn the semantic importance of feature channels, dynamically generating a channel importance mask based on the target channel bandwidth ratio, filtering key semantic information based on the channel importance mask and compressing redundant data to obtain filtered features;
[0010] Step 4: The screened features are transmitted to the receiving end via a wireless channel. The receiving end reconstructs the features based on the channel importance mask, and then adaptively demodulates the features through the signal-to-noise ratio sequence to compensate for channel distortion to obtain calibration features. The calibration features are then decoded in the state space, and the image is gradually reconstructed using inverse state space modeling.
[0011] Furthermore, the step 1 includes the following steps:
[0012] Step 1.1: Perform 2D block embedding on the image, dividing it into multiple non-overlapping 2D blocks. Each block is flattened and mapped to a high-dimensional feature space via linear projection to form an initial block sequence. This process transforms the spatial structure of the image into a feature representation, halving the output resolution and doubling the number of channels.
[0013] Step 1.2: The initial block sequence is subjected to a two-dimensional selective scan through the visual state space block, and the two-dimensional image features are expanded into a one-dimensional sequence through scanning paths in multiple spatial directions. A selective state space model is constructed to parallelly calculate the sequence generated by each path, and the weights are dynamically adjusted to capture long-range relationships in the sequence to obtain semantic features. The semantic features of the path sequence are integrated to obtain a two-dimensional feature map.
[0014] Step 1.3: merging adjacent two-dimensional blocks, deep feature extraction by stacking multiple visual state space blocks to obtain the feature map, the hierarchical structure gradually improves the semantic abstraction ability of the feature by balancing spatial downsampling and channel expansion.
[0015] Further, in step 1.2, the two-dimensional selective scanning is performed from the top left corner to the bottom right corner, from the bottom right corner to the top left corner, from the top right corner to the bottom left corner, and from the bottom left corner to the top right corner, traversing the initial block in four different directions, so that each pixel block aggregates context information from four complementary directions, covering all spatial position relationships, laying the foundation for capturing global dependencies.
[0016] Further, in step 1.2, the selective state space model generates dynamic parameters for sequence elements in the scanning path, including state evolution rate control parameters, dynamic weighted input weight vectors, and dynamic weighted output weight vectors. The state evolution rate control parameters are used to adjust the projection intensity, and the discrete state transition matrix is generated based on the state evolution rate control parameters and the continuous system matrix to control the state decay rate. The dynamic weighted input weight vector is used to select key channels, and the discrete input projection matrix is generated based on the dynamic weighted input weight vector and the state evolution rate control parameter to dynamically scale the influence intensity of the input signal. The result of multiplying the hidden state at the previous time by the discrete state transition matrix is summed with the result of multiplying the discrete input projection matrix by the sequence element to obtain the hidden state at the current time. The hidden state at the current time is multiplied by the dynamic weighted output weight vector to obtain the semantic feature.
[0017] Further, in step 1.2, the semantic feature integration of the path sequence is to reshape the output of each path sequence to the original spatial dimension, and then merge the output feature maps corresponding to the same spatial position from multiple paths respectively. This step reconstructs the spatial structure while preserving the global context accumulated from the scanning path. The semantic feature is mapped back through the scanning path to restore the one-dimensional sequence to a two-dimensional image, and the output features of the pixel points in multiple scanning paths are weighted and fused to obtain a two-dimensional feature map.
[0018] Further, in step 1.3, a plurality of units composed of a merging operation and a plurality of stacked visual state space blocks are connected in sequence to obtain the two-dimensional feature. The merging operation is used to concatenate adjacent two-dimensional blocks along the channel dimension, which halves the resolution and doubles the number of channels. Then, linear projection is used for dimension reduction. Multiple visual state space blocks are used for deep feature extraction, maintaining the resolution. The hierarchical structure gradually improves the semantic abstraction ability of the feature by balancing spatial downsampling and channel expansion.
[0019] Furthermore, the step 2 includes the following steps:
[0020] Step 2.1: Using the feature graph, calculate embedding vectors for the signal-to-noise ratio values at the current moment and the historical moment to obtain an embedding sequence of the signal-to-noise ratio sequence;
[0021] Step 2.2: The embedded sequence is dynamically parsed through a lightweight Mamba network to perform time series modeling. The generated lightweight output matrix is modulated to obtain a multidimensional modulation vector. The multidimensional modulation vector is fused with the feature map by channel-level multiplication to obtain a feature with real-time adaptive channel time-varying characteristics.
[0022] Furthermore, the dynamic analysis in step 2.2 is to sum the result of multiplying the hidden state at the previous moment by the state transfer matrix and the result of multiplying the embedding vector by the output matrix to obtain the hidden state at the current moment, and then multiply the hidden state at the current moment by the output matrix to obtain the output vector at the current moment, and take the output vector of the terminal state to generate the final output matrix.
[0023] Furthermore, the step 3 includes the following steps:
[0024] Step 3.1: Generate a modulation vector based on the target channel bandwidth ratio;
[0025] Step 3.2: Perform progressive modulation from coarse to fine by multiplying a set of modulation vectors with the modulated features channel by channel, and adjust the weight of each channel according to the target bandwidth ratio to obtain a modulated feature map;
[0026] Step 3.3: The modulated feature map is globally averaged pooled along the spatial dimension to generate an importance score vector (the spatial average of each channel of the feature map), and then the channel importance is sorted. The importance mask is generated in combination with the target channel bandwidth ratio. The importance mask is multiplied element by element with the modulated feature map to retain the important channels to be transmitted, achieve further compression, and obtain the screening features.
[0027] Furthermore, step 4 includes the following steps:
[0028] Step 4.1: The screened features are transmitted to a receiving end via a wireless channel. The receiving end reconstructs features based on the importance mask, restores the features of the importance mask, and eliminates noise generated by the non-importance mask.
[0029] Step 4.2: Adaptively demodulate the signal-to-noise ratio sequence and compensate for channel distortion through the inverse process to obtain the calibrated features.
[0030] Step 4.3: The calibrated features are decoded in state space, and the image is gradually reconstructed through inverse state space modeling.
[0031] The advantages and beneficial effects of the present invention are:
[0032] The present invention proposes a multi-level Mamba structure to construct a coding and decoding backbone, and utilizes its linear complexity long-distance dependency modeling capability to solve the computational delay bottleneck in real-time processing of high-resolution images. The present invention designs an SNR sequence adaptive modulator, dynamically analyzes historical channel state sequences through a lightweight Mamba network, generates refined modulation vectors, and performs channel-level fusion with feature maps, overcoming the problem of delayed response to instantaneous channel fluctuations in traditional methods. The present invention introduces a semantically aware bandwidth selection mechanism, which dynamically evaluates the semantic importance of feature channels through learnable weights, realizes soft mask adaptive screening under bandwidth constraints, and significantly optimizes the utilization efficiency of limited bandwidth resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 It is a flow chart of a method according to an embodiment of the present invention.
[0034] Figure 2 2 is a schematic diagram of the structure of the visual state space block in an embodiment of the present invention.
[0035] Figure 3 4 is a schematic structural diagram of an SNR sequence adaptive modulator / demodulator in an embodiment of the present invention.
[0036] Figure 4 4 is a schematic diagram of the structure of the semantic-aware bandwidth selector in an embodiment of the present invention.
[0037] Figure 5 4 is a schematic structural diagram of a target channel bandwidth ratio modulator in an embodiment of the present invention.
[0038] Figure 6 2 is a schematic structural diagram of an importance mask generation module in an embodiment of the present invention. DETAILED DESCRIPTION
[0039] The following describes the specific embodiments of the present invention in detail with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only used to illustrate and explain the present invention and are not intended to limit the present invention.
[0040] like Figure 1 As shown in FIG, a dynamic adaptive image semantic transmission method based on a state-space model utilizes state-space models, channel adaptation, semantic communication and other technologies to improve the efficiency and robustness of image transmission in dynamic driving environments. The method includes the following steps:
[0041] Step 1: Input the visible light image into the state space backbone network, and use the linear computational complexity of the state space model (SSM) to efficiently model long-range dependencies, achieving fast encoding of high-resolution images. The selective state space mechanism captures key spatial features, reduces processing latency, and outputs multi-scale feature maps. The specific steps include the following:
[0042] Step 1.1: Input a visible light image I of size H × W × 3 to the encoder network. First, perform 2D block embedding on the image: split the input image into multiple non-overlapping 2D blocks. Each block is flattened and mapped to a high-dimensional feature space via linear projection, forming an initial block sequence. This process transforms the spatial structure of the image into a feature representation, halving the output resolution and doubling the number of channels.
[0043] Step 1.2: Initial block sequence inputs visual state space block, and uses the linear computational complexity of the state space model to efficiently extract features. Figure 2 As shown in Figure 2, the core of the visual state space block is the 2D selective scan, which expands the 2D image features into a 1D sequence through scanning paths in four spatial diagonal directions for subsequent selective state space model (S6 module) processing.
[0044] The visual state space block implements the following process: input features are first stabilized by layer normalization, followed by linear projection to expand their dimensionality. Deep convolutional layers extract local spatial features, and nonlinearity is introduced via the SiLU activation function. The core module, 2D Selective Scan, captures global long-range dependencies through four-way scanning. The 2D Selective Scan output is stabilized by layer normalization, then linearly projected to compress the channel dimensions back to their original size. Finally, a feedforward neural network performs feature blending to enhance expressiveness. The entire process incorporates residual connections to preserve original information and prevent network degradation.
[0045] The 2D selective scanning process is as follows:
[0046] First, we traverse these blocks along four different directions: from the top left corner to the bottom right corner, the bottom right corner to the top left corner, the top right corner to the bottom left corner, and the bottom left corner to the top right corner. Each path expands the two-dimensional structure into a one-dimensional sequence. This design enables each pixel block to aggregate contextual information from four complementary directions, covering all spatial position relationships and laying the foundation for capturing global dependencies. The feature map M is converted into four independent 1D sequences, each of length T = H × W. The specific formula is as follows:
[0047] Path 1 (upper left → lower right): ;
[0048] Path 2 (lower right → upper left): ;
[0049] Path 3 (upper right → lower left): ;
[0050] Path 4 (lower left → upper right): .
[0051] in, , Indicates the Scan paths, Represents a sequence element, where is the pixel in the path Middle The position of the step.
[0052] Subsequently, the sequence generated by each path is processed in parallel by an independent S6 module, and the weights are dynamically adjusted to capture the long-range relationship in the sequence. The S6 module is based on a structured state space model and maps the feature vector of each input block to a hidden state through a dynamic parameterization mechanism, so that the model can adaptively screen important information while filtering noise. The discretized state space equation is used for each sequence. Independent calculation, the specific formula is as follows:
[0053]
[0054] (State transition matrix)
[0055] (Input projection matrix)
[0056]
[0057]
[0058] in, Represents the input sequence elements (each sequence Elements in Directly as input to the S6 module ), By current input Dynamic parameters for dynamic generation (selectivity mechanism), Control state evolution rate ( Big → long memory, Small → Short Memory) to adjust the projection intensity, 、 Represent the dynamically weighted input weight vector and output weight vector, respectively, so that the model focuses on relevant context, Used to select key channels, A represents the continuous system matrix, providing the underlying physical laws of state evolution (such as the continuity prior of feature propagation), Represents the discrete state transfer matrix, which actually controls the state decay rate. Represents a discrete input projection matrix that dynamically scales the impact of the input signal. Represents the hidden state at time t, which is used to carry the semantic information flow across time. Represents the final output semantic features.
[0059] Finally, the four-way merging stage reassembles the four processed sequences into 2D feature maps. The output of each sequence is reshaped to the original spatial dimensions, and then the four output feature maps corresponding to the same spatial position from the four paths are merged. This step reconstructs the spatial structure while preserving the global context accumulated from the scanning path. This process goes through sequence-to-image reorganization and weighted fusion, and the specific formula is as follows:
[0060]
[0061]
[0062] The inverse mapping of , restores the one-dimensional sequence to a two-dimensional image. represents the learnable channel weight vector, represents element-wise multiplication, In the Pixel points in the scan path The output features of each pixel The final output features is the weighted sum of the outputs of the four scan paths.
[0063] Step 1.3: The features output from the first stage (steps 1.1 and 1.2 constitute the first stage) are processed in the subsequent second, third, and fourth stages. The second stage (structured similarly to stages 3 and 4) consists of a block merging operation and multiple stacked visual state space blocks. Block merging concatenates adjacent blocks along the channel dimension (halving the resolution and doubling the number of channels), followed by dimensionality reduction via linear projection. Multiple visual state space blocks are then stacked for deep feature extraction (preserving resolution). This hierarchical structure balances spatial downsampling with channel expansion to gradually improve the semantic abstraction of features.
[0064] Step 2: Input the feature map output from step 1 into the SNR sequence adaptive modulator, as Figure 3 As shown in the figure, a lightweight Mamba network dynamically analyzes the historical signal-to-noise ratio sequence (the physical channel model generates a simulated sequence in the training phase, and the cached real instantaneous historical SNR is used in the testing phase) to generate a multidimensional modulation vector. This vector is then fused with the feature map through channel-level multiplication, making the feature expression adaptive to the time-varying characteristics of the channel in real time. The specific steps include the following:
[0065] Step 2.1: Features after passing through the state-space backbone network Input SNR sequence adaptive modulator. First, SNR is generated by SNR sequence generator to generate SNR sequence, which is then transformed into learnable features by embedding layer. The specific formula can be expressed as:
[0066]
[0067]
[0068]
[0069] in, n represents the SNR sequence, with 6 SNR value scalars at historical moments and 1 at the current moment. represents the embedding vector of the i-th SNR value, , are the embedding layer parameters, is the embedding dimension, Represents the embedded sequence after the entire SNR sequence is embedded.
[0070] Step 2.2: The learnable features generated above are dynamically parsed by the lightweight Mamba network for temporal modeling. After passing through the modulation layer, a multi-dimensional modulation vector is generated. This vector is combined with the feature vector. Perform channel-level multiplication fusion and output the processed features , so that the feature expression is real-time adaptive to the time-varying characteristics of the channel. The specific formula of the lightweight Mamba network can be expressed as:
[0071]
[0072] (Equation of State)
[0073] (Output equation)
[0074] in, is the hidden state at time t, is the output vector at time t, ( ) is a lightweight Mamba network function. is the state transfer matrix, and the diagonal matrix is used to reduce the parameters. Output matrix Grouped convolution is used, and the output takes the terminal state: Achieve lightweight, Represents the final output matrix of the lightweight Mamba network.
[0075] In summary, the entire process of steps 2.1 and 2.2 can be expressed as:
[0076]
[0077] in, Represents the Sigmoid function, ensuring that the output value is in the range [0,1], Represents channel-by-channel multiplication.
[0078] Step 3: The modulated features are input into the semantic-aware bandwidth selector. Figure 4 As shown in the figure, this module learns the semantic importance of feature channels, dynamically generates channel importance masks based on the target channel bandwidth ratio, filters key semantic information and compresses redundant data. Specifically, it includes the following steps:
[0079] Step 3.1: Characteristics of the output of the SNR sequence adaptive modulator Input to the semantic-aware bandwidth selector. First, the target channel bandwidth ratio passes through the target channel bandwidth ratio modulator through three fully connected layers to generate the corresponding modulation vector, such as Figure 5 Specifically, it can be expressed as:
[0080]
[0081]
[0082]
[0083] in, is the target channel bandwidth ratio modulation vector, CBR represents the target bandwidth ratio, the ReLU activation function introduces nonlinearity, and Sigmoid compresses the output to the [0,1] interval to implement the channel attention mechanism.
[0084] Step 3.2: Connect 7 target channel bandwidth ratio modulators in series to achieve coarse-to-fine progressive modulation. The generated modulation vector is multiplied by the feature map channel by channel. The weight of each channel of the feature map is adjusted according to the target bandwidth ratio.
[0085] Step 3.3: After the modulated feature map, the importance mask generation module is used to generate a channel importance mask, retaining important channels and further reducing the transmission bandwidth, such as Figure 6 Specifically, the modulated feature map is first globally averaged pooled along the spatial dimension to generate an importance score vector S (the spatial average of each channel of the feature map), and then the channel importance is sorted and combined with the target channel bandwidth ratio to generate an importance binary mask. , and finally only important channels are retained for transmission to achieve further compression. The formula can be expressed as:
[0086]
[0087] in, It is a characteristic after gradual modulation. represents the importance binary mask, Represents the final output of the semantic-aware bandwidth selector.
[0088] Step 4: The selected features are transmitted to the receiver via the wireless channel. The receiver first reconstructs the features based on the importance mask. , and then compensate the channel distortion through the SNR sequence adaptive demodulator to obtain the calibrated feature Input state space decoder, and gradually reconstruct high-quality visible light image through inverse state space modeling , specifically including the following steps:
[0089] Step 4.1: The selected features are transmitted to the receiver via the wireless channel. The receiver first reconstructs the features based on the importance mask, recovering the features with mask 1 and eliminating the noise generated by the mask 0 area.
[0090] Step 4.2: Compensate for channel distortion through the inverse process of the SNR sequence adaptive demodulator;
[0091] Step 4.3: The calibrated features are input into the state space decoder, and the high-quality visible light image is gradually reconstructed through inverse state space modeling.
[0092] In summary, the state-space model-based dynamic adaptive image semantic transmission method of the present invention includes four key steps: First, a visible light image is input into a state-space backbone network. The linear computational complexity of the state-space model (SSM) is leveraged to efficiently model long-range dependencies, enabling fast encoding of high-resolution images. A selective state-space mechanism is used to capture key spatial features, reduce processing latency, and output a multi-scale feature map. Next, the feature map output from step 1 is input into an SNR sequence adaptive modulator. A lightweight Mamba network dynamically analyzes historical signal-to-noise ratio (SNR) sequences (simulated sequences generated by a physical channel model during training, and cached real instantaneous historical SNRs during testing) to generate a multidimensional modulation vector. This vector is then fused with the feature map through channel-level multiplication, enabling the feature representation to adapt to the time-varying characteristics of the channel in real time. The modulated features are then input into a semantic-aware bandwidth selector. This module learns the semantic importance of feature channels, dynamically generates channel importance masks based on the target channel bandwidth ratio, filters key semantic information, and compresses redundant data. Finally, the selected features are transmitted to the receiver via a wireless channel. The receiver first reconstructs features based on the importance mask and then compensates for channel distortion through the SNR sequence adaptive demodulator; the calibrated features are input into the state space decoder, and the high-quality visible light image is gradually reconstructed through inverse state space modeling.
[0093] The application significantly improves the transmission robustness in a dynamic channel, reduces the end-to-end processing delay, effectively improves the transmission efficiency and robustness of the semantic information of the image, breaks through the calculation bottleneck of the traditional CNN / Transformer architecture, is particularly suitable for a dynamic vehicle-road environment, provides an efficient solution for real-time image transmission scenes such as automatic driving and remote medical treatment, and has important significance for improving the reliability and safety of an automatic driving system.
[0094] The above examples are only used to illustrate the technical solutions of the present application, but not limit them; although the present application has been described in detail with reference to the foregoing examples, those skilled in the art should understand that the technical solutions recorded in the foregoing examples can be modified, or some or all of the technical features can be replaced by equivalents; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A dynamic adaptive image semantic transmission method based on a state-space model, characterized by The steps include: Step 1: Acquire the image and perform state-space encoding, extract features using the linear computational complexity of the state-space model, and obtain a feature map; Step 2: Adaptively modulate the feature map using a signal-to-noise ratio sequence, dynamically analyze the historical signal-to-noise ratio sequence, generate a multidimensional modulation vector, and fuse the multidimensional modulation vector with the feature map so that the modulated feature has adaptive channel time-varying characteristics; Step 3: Perform semantic-aware bandwidth selection on the modulated features, dynamically generate a channel importance mask based on the target channel bandwidth ratio, filter key semantic information based on the channel importance mask and compress redundant data to obtain filtered features; Step 4: The screened features are transmitted to the receiving end via a wireless channel. The receiving end reconstructs the features based on the channel importance mask, and then adaptively demodulates the signal-to-noise ratio sequence to obtain the calibration features. The calibration features are then decoded in the state space, and the image is gradually reconstructed using inverse state space modeling.
2. The dynamic adaptive image semantic transmission method based on the state space model according to claim 1, characterized in that: The step 1 comprises the following steps: Step 1.1: Perform 2D block embedding on the image, splitting it into multiple non-overlapping 2D blocks. Each block is flattened and mapped to a high-dimensional feature space through linear projection to form an initial block sequence. Step 1.2: The initial block sequence is subjected to a two-dimensional selective scan through a visual state space block, and the two-dimensional image features are expanded into a one-dimensional sequence through scanning paths in multiple spatial directions. A selective state space model is constructed to parallelly calculate the sequence generated by each path to capture long-range relationships in the sequence and obtain semantic features. The semantic features of the path sequence are integrated to obtain a two-dimensional feature map. Step 1.3: Merge adjacent two-dimensional blocks and perform deep feature extraction by stacking multiple visual state space blocks to obtain the feature map.
3. The dynamic adaptive image semantic transmission method based on the state space model according to claim 2, characterized in that: The two-dimensional selective scanning in step 1.2 traverses the initial block along four different directions, from the upper left corner to the lower right corner, the lower right corner to the upper left corner, the upper right corner to the lower left corner, and the lower left corner to the upper right corner, so that each pixel block aggregates context information from four complementary directions.
4. The dynamic adaptive image semantic transmission method based on the state space model according to claim 2, characterized in that: The selective state space model in step 1.2 generates dynamic parameters for the sequence elements in the scanning path through a dynamic parameterization mechanism, including a state evolution rate control parameter, a dynamic weighted input weight vector and a dynamic weighted output weight vector. The state evolution rate control parameter is used to adjust the projection intensity. A discrete state transfer matrix is generated based on the state evolution rate control parameter and the continuous system matrix to control the state attenuation rate. The dynamic weighted input weight vector is used to select key channels. A discrete input projection matrix is generated based on the dynamic weighted input weight vector and the state evolution rate control parameter to dynamically scale the influence intensity of the input signal. The result of the element-by-element multiplication of the hidden state at the previous moment with the discrete state transfer matrix is summed with the result of the element-by-element multiplication of the discrete input projection matrix with the sequence element to obtain the hidden state at the current moment. The hidden state at the current moment is multiplied element-by-element by the dynamic weighted output weight vector to obtain the semantic feature.
5. The dynamic adaptive image semantic transmission method based on the state space model according to claim 2, characterized in that: The semantic feature integration of the path sequence in step 1.2 is to reshape the output of each path sequence into the original spatial dimension, and then merge the output feature maps corresponding to the same spatial position from multiple paths; the semantic features are inversely mapped through the scanning path to restore the one-dimensional sequence to a two-dimensional image, and the output features of the pixel points in multiple scanning paths are weightedly fused to obtain a two-dimensional feature map.
6. The dynamic adaptive image semantic transmission method based on the state space model according to claim 2, characterized in that: In step 1.3, a plurality of sequentially connected units consisting of merging operations and a plurality of stacked visual state space blocks are constructed to obtain the two-dimensional features. Adjacent two-dimensional blocks are spliced along the channel dimension through the merging operation, and then the dimensionality is reduced by linear projection, and deep features are extracted from the plurality of visual state space blocks.
7. The dynamic adaptive image semantic transmission method based on the state space model according to claim 1, characterized in that: The step 2 comprises the following steps: Step 2.1: Using the feature graph, calculate embedding vectors for the signal-to-noise ratio values at the current moment and the historical moment to obtain an embedding sequence of the signal-to-noise ratio sequence; Step 2.2: The embedded sequence is subjected to time series modeling through dynamic analysis, and the generated output matrix is modulated to obtain a multidimensional modulation vector. The multidimensional modulation vector is fused with the feature map by channel-level multiplication to obtain a feature with real-time adaptive channel time-varying characteristics.
8. The dynamic adaptive image semantic transmission method based on the state space model according to claim 7, characterized in that: The dynamic analysis in step 2.2 is to sum the result of multiplying the hidden state at the previous moment by the state transfer matrix and the result of multiplying the embedding vector by the output matrix to obtain the hidden state at the current moment, then multiply the hidden state at the current moment by the output matrix to obtain the output vector at the current moment, and take the output vector of the terminal state to generate the final output matrix.
9. The dynamic adaptive image semantic transmission method based on the state space model according to claim 1, characterized in that: The step 3 comprises the following steps: Step 3.1: Generate a modulation vector based on the target channel bandwidth ratio; Step 3.2: Perform progressive modulation from coarse to fine by multiplying a set of modulation vectors with the modulated features channel by channel, and adjust the weight of each channel according to the target bandwidth ratio to obtain a modulated feature map; Step 3.3: The modulated feature map is globally averaged pooled along the spatial dimension to generate an importance score vector, and then the channel importance is sorted. The importance mask is generated in combination with the target channel bandwidth ratio. The importance mask is multiplied element by element with the modulated feature map to retain the important channels to be transmitted to obtain the screening features.
10. The dynamic adaptive image semantic transmission method based on the state space model according to claim 1, characterized in that: The step 4 comprises the following steps: Step 4.1: The screened features are transmitted to a receiving end via a wireless channel. The receiving end reconstructs features based on the importance mask, restores the features of the importance mask, and eliminates noise generated by the non-importance mask. Step 4.2: Adaptively demodulate the signal-to-noise ratio sequence and compensate for channel distortion through the inverse process to obtain the calibrated features. Step 4.3: The calibrated features are decoded in state space, and the image is gradually reconstructed through inverse state space modeling.
Citation Information
Patent Citations
Semantic transmission method for visible light and infrared fusion image of night vehicle road environment
CN119360341A
Target detection method based on Mama feature fusion
CN120298667A
Cited By
Data compression method and system based on neural network
CN121239233A