Low-delay high-performance image semantic-channel joint coding system

Through the joint coding system of the neural analyzer and the quadtree partitioning module, the high latency problem caused by the fixed coding strategy in the existing JSCC system is solved, low-latency, high-performance image semantic-channel joint coding is achieved, and the system's adaptability and anti-interference ability are improved.

CN120729477AActive Publication Date: 2025-09-30TSINGHUA UNIVERSITY

Patent Information

Application Number
CN202511141493.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-15
Publication Date
2025-09-30
Estimated Expiration
2045-08-15

AI Technical Summary

Technical Problem

The existing JSCC system has fixed coding strategies, lacks flexibility and adaptability in image and video transmission, resulting in high latency and insufficient anti-interference ability, making it difficult to meet the needs of low-latency and high-quality real-time communication.

Method used

A neural analyzer is used to extract semantic features and divide them into multiple partitions through a quadtree partitioning module. Context-aware encoding is performed using a feature encoder and mapped into complex-valued channel symbols. The receiving end reconstructs the semantic features and generates images through a feature decoder, and optimizes the symbol length factor by combining super-prior modeling and entropy estimation.

Benefits of technology

It achieves low-latency, high-performance image semantic-channel joint coding, improves the system's adaptability and anti-interference capabilities, reduces transmission delay and improves image reconstruction quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120729477A_ABST
    Figure CN120729477A_ABST
Patent Text Reader

Abstract

The invention provides a low-delay high-performance image semantic-channel joint coding system, and relates to the technical field of communication and image processing, and the method comprises the steps that a sending end extracts semantic features through a neural analyzer, and divides the semantic features into a plurality of partitions through a quadtree structure; and coding each partition step by step by using a context-aware feature coding strategy, and finally mapping to a complex value channel symbol to be sent to a receiving end. And a receiving end reconstructs the semantic features through a corresponding decoding and synthesizing module and generates a reconstructed image, so that end-to-end efficient image transmission is realized. According to the method, a context-based partition coding mechanism and a quadtree partition module strategy are introduced into a coding structure, so that the method has higher context modeling capability and parallel processing capability, and the adaptive capability to an image structure and semantic content can be effectively improved; transmission delay can be reduced, and image reconstruction quality and channel interference resistance can be enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of communication and image processing technology, and in particular to a low-latency, high-performance image semantic-channel joint coding system. Background Art

[0002] As research into 6G communication technologies continues to deepen, semantic communication is rapidly developing towards semantic perception and intelligent fusion. To effectively alleviate bandwidth bottlenecks and channel noise interference in image and video transmission, deep semantic communication has become a research hotspot. The JSCC (Joint Semantic-Channel Coding) framework, with its significant advantages in transmission efficiency and interference resistance, is considered a key component of next-generation communication architectures.

[0003] However, in existing JSCC systems, the encoding of semantic features and the channel transmission process are mostly separated from each other. The encoding strategy often adopts fixed modules and static parameter configurations, and fails to flexibly adapt to different image characteristics, task requirements and channel states, resulting in limited practicality of the overall system in low-latency, large-scale image communication.

[0004] In particular, in semantic encoding modules, current mainstream methods rely on entropy estimation strategies based on autoregressive modeling. While these methods offer good compression performance, their strong sequential dependencies significantly inhibit parallel processing capabilities, leading to significant encoding delays and making it difficult to meet the stringent low-latency requirements of real-time semantic communication.

[0005] Furthermore, existing JSCC systems generally lack flexibility and contextual adaptability in key aspects such as feature transformation, symbol mapping, and side information construction. For example, symbol dimensions are typically fixed, making it impossible to dynamically adjust the transmission bit rate based on the local complexity of the image. This wastes bandwidth and affects image quality reconstruction. In high-resolution or complex semantic scenarios, the system is highly sensitive to channel errors, and even minor decoding errors can lead to semantic restoration failures, severely impacting transmission stability and user experience.

[0006] Therefore, there is an urgent need for a low-latency, high-performance image semantic-channel joint coding system. Summary of the Invention

[0007] In view of the above problems, an embodiment of the present application provides a low-latency, high-performance image semantics-channel joint coding system to overcome the above problems or at least partially solve the above problems.

[0008] In a first aspect, an embodiment of the present application provides a low-latency, high-performance image semantic-channel joint coding system, including: The sender in the semantic communication system extracts semantic features from the source image through a neural analyzer; The sending end divides the semantic features into the sending end 0th partition semantic features to the sending end 3rd partition semantic features through a quadtree partition module; The transmitting end encodes the semantic feature of the 0th partition of the transmitting end through a feature encoder to obtain a semantic feature encoding result of the 0th partition; The transmitting end encodes the semantic feature of the nth partition of the transmitting end by the feature encoder, using the encoding results of the semantic features of the 0th to n-1th partitions as context, to obtain the encoding result of the semantic feature of the nth partition, where n is an integer between 1 and 3; The transmitting end maps the semantic feature encoding results of the 0th to nth partitions into complex-valued channel symbols through the feature encoder, and sends the complex-valued channel symbols to the receiving end in the semantic communication system; The receiving end obtains a reconstructed semantic feature by using the received complex-valued channel symbols through a feature decoder; The receiving end obtains a reconstructed image by using the reconstructed semantic features through a neural synthesizer.

[0009] Optionally, the receiving end obtains the reconstructed semantic features by using the received complex-valued channel symbols through a feature decoder, including: The receiving end determines the semantic features of the receiving end partition 0 to the receiving end partition 3 by using the received complex-valued channel symbols through a feature decoder; The transmitting end encodes the semantic feature of the 0th partition of the transmitting end through a feature encoder to obtain a decoding result of the semantic feature of the 0th partition; The transmitting end decodes the semantic feature of the nth partition of the receiving end through the feature decoder, using the decoding results of the semantic features of the 0th to n-1th partitions as context, to obtain a decoding result of the semantic feature of the nth partition, where n is an integer between 1 and 3; The transmitting end obtains the reconstructed semantic feature according to the decoding results of the semantic features of the 0th to the third partitions.

[0010] Optionally, it also includes: The transmitting end inputs the semantic feature into a super a priori encoder to obtain a super a priori latent variable, and obtains a super a priori latent variable quantization result based on the super a priori latent variable; The transmitting end obtains the 0th mean and the 0th variance by using the quantization result of the super priori latent variable through the super priori decoder; The transmitting end obtains an nth mean and an nth variance by using an entropy estimator, using the 0th mean and the 0th variance, and the nth context, where n is an integer between 1 and 3; The transmitting end calculates a corresponding symbol length factor vector for each element in the nth partition semantic feature according to the nth mean and the nth variance, and provides the vector to the feature encoder and the feature decoder.

[0011] Optionally, the sending end divides the semantic features into sending end partition 0 semantic features to sending end partition 3 semantic features through a quadtree partition module, including: The sending end divides the semantic features into four parts along the channel dimension through a quadtree partitioning module; The sending end adds the elements indexed by n in the four parts to obtain the semantic feature of the nth partition, where n is an integer from 0 to 3; Also includes: The sending end connects the elements with index 0 in the four parts to obtain a first context; The sending end connects the element with index 0 and the element with index 1 in the four parts to obtain a second context; The sending end connects the element with index 0, the element with index 1, and the element with index 2 in the four parts to obtain a third context.

[0012] Optionally, the feature encoder includes at least: a first visual encoding module, a first contextual visual encoding module, a second contextual visual encoding module, and a third contextual visual encoding module; the semantic feature encoding results of the 0th to nth partitions are obtained according to the following steps: The transmitting end performs a first encoding on the semantic features of the 0th to 3rd partitions through the first visual encoding module to obtain an encoding result of the semantic features of the 0th partition and first intermediate encoding results of the 1st to 3rd partitions; The transmitting end performs cross-attention calculation using the first contextual visual encoding module, with the semantic feature encoding result of the 0th partition as the key and value, and the first intermediate encoding result of the 1st partition as the query, to obtain the semantic feature encoding result of the 1st partition; The transmitting end performs cross-attention calculation using the second context visual coding module, with the semantic feature coding results of the 0th to 1st partitions as keys and values, and the first intermediate coding result of the second partition as a query, to obtain the semantic feature coding result of the second partition; The sending end uses the second context visual encoding module to perform cross-attention calculations using the semantic feature encoding results of the 0th to 2nd partitions as keys and values ​​and the first intermediate encoding result of the 3rd partition as a query to obtain the semantic feature encoding result of the 3rd partition.

[0013] Optionally, the feature decoder includes at least: a zeroth visual decoding module, a first visual decoding module, a first context visual decoding module, a second visual decoding module, a second context visual decoding module, a third visual decoding module, and a third context visual decoding module; the semantic feature decoding results of the 0th to nth partitions are obtained according to the following steps: The receiving end decodes the semantic features of the 0th partition of the sending end through the 0th visual decoding module to obtain a decoding result of the semantic features of the 0th partition; The receiving end performs cross-attention calculation by using the first visual decoding module and the first contextual visual decoding module, with the decoding result of the semantic feature of the 0th partition as the key and value and the semantic feature of the first partition of the receiving end as the query, to obtain the decoding result of the semantic feature of the first partition; The receiving end performs cross-attention calculation using the second visual decoding module and the second contextual visual decoding module, with the decoding results of the semantic features of the 0th to 1st partitions as keys and values, and the semantic features of the second partition of the receiving end as a query, to obtain a decoding result of the semantic features of the second partition; The receiving end performs cross-attention calculation through the third visual decoding module and the third context visual decoding module, using the decoding results of the semantic features of the 0th to 2nd partitions as keys and values, and the semantic features of the 3rd partition of the receiving end as a query to obtain the decoding results of the semantic features of the 3rd partition.

[0014] Optionally, the neural analyzer includes: 1st to Nth downsampling modules and at least two attention modules, The neural synthesizer includes: first to Nth upsampling modules, at least two attention modules and a feature refinement module; the feature refinement module is connected to the first upsampling module; The first to Nth upsampling modules correspond one to one to the first to Nth downsampling modules.

[0015] Optionally, the method performed by the low-latency, high-performance image semantics-channel joint coding system is implemented by a low-latency, high-performance image semantics-channel joint coding model, and the training process of the low-latency, high-performance image semantics-channel joint coding model includes at least a first stage and a second stage: Under the condition of freezing the model parameters of the feature encoder to be trained and the feature decoder to be trained, the image semantic encoder based on the quadtree partition module to be trained is trained through the first stage and the second stage to obtain a trained image semantic encoder based on the quadtree partition module; the image semantic encoder based on the quadtree partition module to be trained includes: a neural analyzer to be trained, a neural synthesizer to be trained, a super-prior encoder to be trained, a super-prior decoder to be trained, and an entropy estimator to be trained; analog quantization is used in the first stage, and quantization based on a straight-through estimator is used in the second stage.

[0016] Optionally, the training process of the low-latency, high-performance image semantic-channel joint coding model further includes a third stage and a fourth stage; Unfreezing the model parameters of the feature encoder and the feature decoder to be trained, and integrating them with the trained image semantic encoder based on the quadtree partitioning module to obtain a low-latency, high-performance image semantic-channel joint coding model to be trained, wherein the feature encoder to be trained is connected to the neural analyzer trained in the second stage, and the feature decoder to be trained is connected to the neural synthesizer trained in the second stage; Through the third stage and the fourth stage, the low-latency high-performance image semantic-channel joint coding model to be trained is trained to obtain a trained low-latency high-performance image semantic-channel joint coding model; In the third stage, from the initial rate set Select the symbol length factor vector in ; In the fourth stage, from the target rate set Select the symbol length factor vector in .

[0017] Optionally, in the first stage and the second stage, the image semantic encoder based on the quadtree partition module to be trained is trained using a first loss function value, where the first loss function value is determined based on a first difference between a sample original image and a corresponding sample semantically encoded image, where the sample semantically encoded image is output by the image semantic encoder based on the quadtree partition module to be trained for the sample original image; In the third stage and the fourth stage, the low-latency high-performance image semantic-channel joint coding model to be trained is trained using a second loss function value, wherein the second loss function value is determined based on the second difference between the sample original image and the corresponding sample reconstructed image, and the first difference, and the sample reconstructed image is output by the low-latency high-performance image semantic-channel joint coding model to be trained for the sample original image.

[0018] Beneficial effects of this application: The present application proposes a low-latency, high-performance image semantic-channel joint coding system, comprising: a transmitter in a semantic communication system extracts semantic features from a source image through a neural analyzer; the transmitter divides the semantic features into semantic features of the transmitter's 0th partition to semantic features of the transmitter's 3rd partition through a quadtree partitioning module; the transmitter encodes the semantic features of the transmitter's 0th partition through a feature encoder to obtain a 0th partition semantic feature coding result; the transmitter encodes the semantic features of the transmitter's nth partition through the feature encoder using the 0th to n-1th partition semantic feature coding results as context to obtain a nth partition semantic feature coding result, where n is an integer between 1 and 3; the transmitter maps the 0th to nth partition semantic feature coding results into complex-valued channel symbols through the feature encoder and sends them to a receiver in the semantic communication system; the receiver obtains reconstructed semantic features using the received complex-valued channel symbols through a feature decoder; and the receiver obtains a reconstructed image using the reconstructed semantic features through a neural synthesizer.

[0019] This application proposes a low-latency, high-performance image semantic-channel joint coding system. At the transmitter, a neural analyzer extracts semantic features and divides them into multiple partitions using a quadtree structure. Each partition is then progressively encoded using a context-aware feature coding strategy, ultimately mapped to complex-valued channel symbols and transmitted to the receiver. The receiver then reconstructs the semantic features and generates a reconstructed image using corresponding decoding and synthesis modules, achieving efficient end-to-end image transmission.

[0020] This application introduces a context-based partition coding mechanism and a quadtree partition module strategy in the coding structure, which has stronger context modeling and parallel processing capabilities, can effectively improve the adaptability to image structure and semantic content, and help reduce transmission delays, enhance image reconstruction quality and anti-channel interference capabilities. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments of the present application. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0022] Figure 1 This is a low-latency, high-performance image semantic-channel joint coding system provided by an embodiment of the present application; Figure 2 This is a flowchart of a process for generating auxiliary information based on hyper-prior modeling provided by an embodiment of the present application; Figure 3 This is a structural block diagram of a feature encoder provided in an embodiment of the present application; Figure 4 This is a structural block diagram of a feature decoder provided in an embodiment of the present application; Figure 5 This is a schematic diagram of the complete architecture of a low-latency, high-performance image semantic-channel joint coding system provided in an embodiment of the present application. DETAILED DESCRIPTION

[0023] The exemplary embodiments of the present application will be described in more detail below in conjunction with the accompanying drawings in the embodiments of the present application. Although the accompanying drawings show exemplary embodiments of the present application, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to enable a more thorough understanding of the present application and to fully convey the scope of the present application to those skilled in the art.

[0024] In a first aspect of the present application, a low-latency, high-performance image semantic-channel joint coding system is provided. Figure 1 As shown, the system includes: a transmitting end and a receiving end, wherein the transmitting end includes: a neural analyzer, a quadtree partitioning module, and a feature encoder, and the receiving end includes: a feature decoder and a neural synthesizer.

[0025] The sender in the semantic communication system extracts semantic features from the source image through a neural analyzer; The sending end divides the semantic features into the sending end 0th partition semantic features to the sending end 3rd partition semantic features through a quadtree partition module; The transmitting end encodes the semantic feature of the 0th partition of the transmitting end through a feature encoder to obtain a semantic feature encoding result of the 0th partition; The transmitting end encodes the semantic feature of the nth partition of the transmitting end by the feature encoder, using the encoding results of the semantic features of the 0th to n-1th partitions as context, to obtain the encoding result of the semantic feature of the nth partition, where n is an integer between 1 and 3; The transmitting end maps the semantic feature encoding results of the 0th to nth partitions into complex-valued channel symbols through the feature encoder, and sends the complex-valued channel symbols to the receiving end in the semantic communication system; The receiving end obtains a reconstructed semantic feature by using the received complex-valued channel symbols through a feature decoder; The receiving end obtains a reconstructed image by using the reconstructed semantic features through a neural synthesizer.

[0026] Specifically, if Figure 1As shown, the system described in the present application includes: a transmitting end and a receiving end, wherein the transmitting end includes: a neural analyzer, a quadtree partitioning module, and a feature encoder, and the receiving end includes: a feature decoder and a neural synthesizer.

[0027] The transmitting end of the system described in this application first uses a neural analyzer to extract rich semantic features from the source image. To improve compression and coding efficiency, the system further uses a quadtree partitioning strategy to divide the semantic features into four partitions (partitions 0 to 3). This strategy can perform structural division based on the spatial distribution of image semantic information, taking into account both local relevance and global context. Next, the system adopts a step-by-step context-aware encoding method: first, the 0th partition is encoded, and then when encoding the nth partition (n=1 to 3), the encoding results of the first n-1 partitions are introduced as context information. This design effectively improves the targeted feature compression and the completeness of semantic expression, while reducing redundant information.

[0028] After encoding, the encoding results of each partition are mapped into complex-valued channel symbols via a feature encoder for direct transmission over the communication channel, compatible with wireless physical layer modulation requirements. Upon receiving these channel symbols, the receiver reconstructs semantic features through a feature decoder and restores these semantic features to a complete image using a neural synthesizer. This synthesis process is based on end-to-end training and maximizes image quality under constrained bitrate and channel conditions.

[0029] Through the above structural design, this embodiment achieves a highly integrated semantic communication process. Compared to existing SC systems based on autoregressive structures, it uses partitioned context modeling instead of full-sequence autoregressive entropy estimation, significantly reducing overall system latency and enhancing parallel computing capabilities. Furthermore, this structure offers considerable flexibility, adapting to the semantic compression and symbol mapping requirements under varying channel conditions. It is a highly practical solution for 6G semantic communication scenarios.

[0030] In one embodiment, the receiving end obtains the reconstructed semantic features by using the received complex-valued channel symbols through a feature decoder, including: The receiving end determines the semantic features of the receiving end partition 0 to the receiving end partition 3 by using the received complex-valued channel symbols through a feature decoder; The transmitting end encodes the semantic feature of the 0th partition of the transmitting end through a feature encoder to obtain a decoding result of the semantic feature of the 0th partition; The transmitting end decodes the semantic feature of the nth partition of the receiving end through the feature decoder, using the decoding results of the semantic features of the 0th to n-1th partitions as context, to obtain a decoding result of the semantic feature of the nth partition, where n is an integer between 1 and 3; The transmitting end obtains the reconstructed semantic feature according to the decoding results of the semantic features of the 0th to the third partitions.

[0031] Specifically, in this embodiment, the receiving end receives the complex-valued channel symbols encoded by the transmitting end and inputs them into a feature decoder. The feature decoder is configured to parse the signal to extract semantic features of each partition, including semantic features of the receiving end partition 0, semantic features of the receiving end partition 1, semantic features of the receiving end partition 2, and semantic features of the receiving end partition 3.

[0032] In this process, the receiving end first decodes the semantic features of the receiving end partition 0 to obtain the decoding result of the semantic features of partition 0. Since this partition is the decoding starting area, the decoding process does not rely on other context information and can be completed directly.

[0033] Subsequently, for the semantic features of the nth partition (where n is an integer between 1 and 3), the receiver introduces a context modeling mechanism during the decoding process. This mechanism uses the decoding results of the previous n-1 partitions as context input to assist in the decoding of the current nth partition. Specifically, the feature decoder uses the semantic feature decoding results of partitions 0 to n-1 as a reference and uses a context-aware mechanism to predict and reconstruct the semantic features of the nth partition, thereby obtaining the corresponding semantic feature decoding result for the nth partition.

[0034] After the receiver completes decoding the semantic features of partitions 0 to 3, it fuses the decoding results of the four partitions to obtain a complete reconstructed semantic feature. This reconstructed semantic feature will serve as the input of the subsequent neural synthesizer to further generate the corresponding reconstructed image.

[0035] The context-aware decoding strategy employed in this embodiment leverages the guidance of prior partition features to enhance the ability to reconstruct high-dimensional semantic representations. Its structural design avoids the long-chain dependencies and decoding delays associated with traditional autoregressive entropy estimation. By optimizing partition-level context modeling and decoding in parallel, this embodiment significantly reduces decoding latency, improving decoding robustness and overall image reconstruction quality.

[0036] In one embodiment, the transmitting end inputs the semantic feature into a super a priori encoder to obtain a super a priori latent variable, and obtains a super a priori latent variable quantization result based on the super a priori latent variable; The transmitting end obtains the 0th mean and the 0th variance by using the quantization result of the super priori latent variable through the super priori decoder; The transmitting end obtains an nth mean and an nth variance by using an entropy estimator, using the 0th mean and the 0th variance, and the nth context, where n is an integer between 1 and 3; The transmitting end calculates a corresponding symbol length factor vector for each element in the nth partition semantic feature according to the nth mean and the nth variance, and provides the vector to the feature encoder and the feature decoder.

[0037] In this embodiment, Figure 2 As shown in the figure, in order to further improve the modeling accuracy of semantic feature distribution in the encoding stage and enhance the system's adaptability to different spatial semantic complexities, the transmitter introduces an auxiliary information generation process based on super-prior modeling before partitioning the semantic features to guide the accurate estimation of the symbol length factor. The specific steps include: First, the transmitter inputs the extracted semantic features into a super-prior encoder, which generates latent variables for modeling the feature distribution, denoted as super-prior latent variables. This super-prior latent variable captures the overall statistical structure of the semantic features and helps to more accurately characterize their probability distribution.

[0038] Next, the transmitter quantizes the super-prior latent variable to obtain a quantized result of the super-prior latent variable, reducing the bit overhead of this information during transmission and ensuring the repeatability of subsequent decoding. This quantized result is then input into the super-prior decoder to generate statistical parameters for initializing the probability model, including the mean (0th mean) and variance (0th variance) of the semantic features of the 0th partition.

[0039] Furthermore, the transmitter combines these initial statistical parameters with context modeling information to calculate the probability estimation parameters required for the remaining partitions (nth partition, where n is 1 to 3). Specifically, the transmitter uses an entropy estimator based on the 0th mean and 0th variance, combined with the nth context information (i.e., the semantic feature encoding results of the first n-1 partitions), to obtain the mean (nth mean) and variance (nth variance) of the semantic features of the nth partition. This context-driven estimation method ensures that the statistical model underlying the encoding of the partition features is highly correlated, thereby improving the accuracy of entropy modeling.

[0040] After obtaining the statistical parameters for the nth partition, the transmitter further calculates the symbol length factor vector corresponding to each element in the semantic feature of the nth partition based on the nth mean and nth variance. This vector represents the bit contribution weight of each position in the semantic feature and is used to guide the feature encoder to implement differentiated compression for regions of different importance. This factor is also synchronously provided to the feature decoder to restore the entropy modeling consistency in the decoding path.

[0041] This embodiment constructs a precise, dynamic, and context-aware symbol length prediction framework by introducing a mechanism that combines super-prior modeling and context entropy estimation before feature encoding. This embodiment significantly improves parallelism while improving the accuracy of entropy estimation and reduces system coding delay. Especially in compression scenarios of high-resolution or complex-structured images, this mechanism can significantly alleviate the contradiction between bit rate and reconstruction quality, achieve better compression rate control and robustness, and is one of the important supporting means for achieving low-latency and large-scale image transmission in 6G semantic communication.

[0042] In one embodiment, the sending end divides the semantic features into sending end partition 0 semantic features to sending end partition 3 semantic features through a quadtree partition module, including: The sending end divides the semantic features into four parts along the channel dimension through a quadtree partitioning module; The sending end adds the elements indexed by n in the four parts to obtain the semantic feature of the nth partition, where n is an integer from 0 to 3; Also includes: The sending end connects the elements with index 0 in the four parts to obtain a first context; The sending end connects the element with index 0 and the element with index 1 in the four parts to obtain a second context; The sending end connects the element with index 0, the element with index 1, and the element with index 2 in the four parts to obtain a third context.

[0043] In this implementation, in order to improve the structural modeling capability and context controllability of semantic features during the encoding process, the sending end performs a partitioning operation on the extracted semantic features through a quadtree partitioning module, dividing the overall semantic features into multiple partitions to achieve local refinement modeling and step-by-step context construction, specifically including the following steps: First, the transmitter uses a quadtree partitioning module to partition the original semantic features along the channel dimension. Specifically, the semantic features are divided into four parts, corresponding to the four initial sub-regions in the subsequent encoding process.

[0044] After partitioning, the sender aggregates the elements indexed by n in each partition. Specifically, for each integer n (n ranges from 0 to 3), the sender sums the elements indexed by n in the four partitions (or performs an equivalent aggregation operation) to obtain the semantic features of the corresponding nth partition. This operation ensures a certain degree of spatial semantic fusion of the semantic features within the partition structure, helping to improve the semantic integrity and decoding stability of each partition.

[0045] To support context modeling in subsequent partition encoding, the embodiment further constructs multi-level context information to guide the encoding and entropy estimation process of the nth partition (n = 1 to 3): The sending end connects all elements with index 0 in the four parts to form a first context, which is used as a reference for encoding semantic features of the first partition; The sending end concatenates the element with index 0 and the element with index 1 to form the second context, which serves as the context input for the semantic feature encoding of the second partition; Similarly, the sender connects the elements indexed 0, 1, and 2 in sequence to form the third context, which is provided for the encoding and entropy estimation of the third partition.

[0046] The quadtree partitioning strategy in this embodiment not only achieves effective semantic feature segmentation, but also provides progressive semantic reference information for the encoding of different partitions through a gradually expanding context construction mechanism. This embodiment constructs context through structured connections, significantly improving context controllability and parallel computing capabilities, significantly reducing overall system latency while maintaining encoding accuracy.

[0047] In one embodiment, the feature encoder includes at least: a first visual encoding module, a first contextual visual encoding module, a second contextual visual encoding module, and a third contextual visual encoding module; the semantic feature encoding results of the 0th to nth partitions are obtained according to the following steps: The transmitting end performs a first encoding on the semantic features of the 0th to 3rd partitions through the first visual encoding module to obtain an encoding result of the semantic features of the 0th partition and first intermediate encoding results of the 1st to 3rd partitions; The transmitting end performs cross-attention calculation using the first contextual visual encoding module, with the semantic feature encoding result of the 0th partition as the key and value, and the first intermediate encoding result of the 1st partition as the query, to obtain the semantic feature encoding result of the 1st partition; The transmitting end performs cross-attention calculation using the second context visual coding module, with the semantic feature coding results of the 0th to 1st partitions as keys and values, and the first intermediate coding result of the second partition as a query, to obtain the semantic feature coding result of the second partition; The sending end uses the second context visual encoding module to perform cross-attention calculations using the semantic feature encoding results of the 0th to 2nd partitions as keys and values ​​and the first intermediate encoding result of the 3rd partition as a query to obtain the semantic feature encoding result of the 3rd partition.

[0048] In this embodiment, in order to further improve the information extraction capability and context modeling accuracy in the semantic feature encoding process, the feature encoder is designed to include multiple submodules with cross-attention mechanisms, which are used to construct progressive contextual relationships for different partitioned semantic features, such as Figure 3 The structural block diagram of the feature encoder shown specifically includes the following sub-modules: a first visual encoding module, a first contextual visual encoding module, a second contextual visual encoding module, and a third contextual visual encoding module.

[0049] The semantic feature encoding results of the 0th to nth partitions are obtained according to the following steps: First, the transmitter sequentially inputs the semantic features of partitions 0 through 3 into the first visual encoding module, which performs preliminary encoding on each partition feature. This module outputs the semantic feature encoding result for partition 0 and the first intermediate encoding results for partitions 1, 2, and 3. These intermediate results serve as the input for subsequent context-aware encoding.

[0050] Next, the sending end performs contextual enhancement on the semantic features of partition 1 through the first contextual visual encoding module. Specifically, this module uses the semantic feature encoding results of partition 0 as the "key" and value inputs, and the first intermediate encoding result of partition 1 as the "query". It performs cross-attention calculations and ultimately outputs the semantic feature encoding results of partition 1. This mechanism models the dependency of partition 1 on the semantic information of partition 0, thereby improving semantic continuity and compressed expressiveness.

[0051] Subsequently, the sending end performs similar processing on the semantic features of the second partition through the second contextual visual encoding module. This module uses the semantic feature encoding results of partitions 0 to 1 as keys and values, and the first intermediate encoding result of partition 2 as a query. It also performs cross-attention calculations to obtain the semantic feature encoding result of partition 2, ensuring that the second partition fully perceives the encoded information of the first two partitions.

[0052] Furthermore, in order to process the semantic features of the third partition, the sending end again utilizes the third context visual encoding module, which uses the semantic feature encoding results of the 0th to 2nd partitions as keys and values, and the first intermediate encoding result of the third partition as a query, performs cross-attention calculations, and outputs the semantic feature encoding results of the third partition.

[0053] Through the above progressive encoding path, the encoding results of the semantic features of the 0th to 3rd partitions are finally obtained.

[0054] The cross-attention context modeling mechanism adopted in this embodiment can significantly improve the efficiency of contextual interaction between different semantic areas, breaking the traditional context estimation mode that relies on serial processing and supporting a higher degree of parallel execution. Based on feature partitioning, the attention enhancement path of the key-value-query structure is introduced, combined with the representation weights of different semantic areas, which effectively enhances the expressive power of the semantic encoder and the modeling ability of spatial dependency structures in complex scenarios. While maintaining the coding accuracy, this embodiment reduces the overall delay and can effectively improve the real-time image transmission quality of the 6G semantic communication system.

[0055] In one embodiment, the feature decoder includes at least: a zeroth visual decoding module, a first visual decoding module, a first context visual decoding module, a second visual decoding module, a second context visual decoding module, a third visual decoding module, and a third context visual decoding module; the semantic feature decoding results of the 0th to nth partitions are obtained according to the following steps: The receiving end decodes the semantic features of the 0th partition of the sending end through the 0th visual decoding module to obtain a decoding result of the semantic features of the 0th partition; The receiving end performs cross-attention calculation by using the first visual decoding module and the first contextual visual decoding module, with the decoding result of the semantic feature of the 0th partition as the key and value and the semantic feature of the first partition of the receiving end as the query, to obtain the decoding result of the semantic feature of the first partition; The receiving end performs cross-attention calculation using the second visual decoding module and the second contextual visual decoding module, with the decoding results of the semantic features of the 0th to 1st partitions as keys and values, and the semantic features of the second partition of the receiving end as a query, to obtain a decoding result of the semantic features of the second partition; The receiving end performs cross-attention calculation through the third visual decoding module and the third context visual decoding module, using the decoding results of the semantic features of the 0th to 2nd partitions as keys and values, and the semantic features of the 3rd partition of the receiving end as a query to obtain the decoding results of the semantic features of the 3rd partition.

[0056] In this embodiment, in order to ensure that the semantic features of the complex-valued channel symbols after decoding can be accurately restored layer by layer at the receiving end, and to make full use of the context information between the semantic partitions to improve the decoding accuracy, as shown in FIG. Figure 4 The structural block diagram of the feature decoder shown in the figure is designed to have a multi-level attention enhancement structure, including the following sub-modules: a zeroth visual decoding module, a first visual decoding module, a first context visual decoding module, a second visual decoding module, a second context visual decoding module, a third visual decoding module, and a third context visual decoding module; The semantic feature decoding results of the 0th to nth partitions are obtained according to the following steps: First, the receiver receives the complex-valued channel symbols transmitted by the transmitter and maps them back into multiple partition semantic features, including the semantic features of the receiver partitions 0 to 3. The receiver first decodes the semantic features of the receiver partition 0 through the zeroth visual decoding module to obtain the decoding results of the semantic features of the zeroth partition.

[0057] Next, the receiving end uses the first visual decoding module and the first contextual visual decoding module to collaboratively perform cross-attention decoding on the semantic features of the receiving end's first partition. Specifically, this module uses the decoded results of the semantic features of partition 0 as the key and value, and the semantic features of the receiving end's first partition as the query. By performing cross-attention calculations, it generates the decoding results of the semantic features of partition 1. This step implements context-aware modeling of the first partition with respect to the previous region.

[0058] Furthermore, the receiving end decodes the semantic features of the second partition through the second visual decoding module and the second context visual decoding module. The decoding operation uses the decoding results of the semantic features of the 0th to 1st partitions as keys and values, and the semantic features of the second partition of the receiving end as queries to complete the cross-attention calculation and obtain the decoding results of the semantic features of the second partition.

[0059] Furthermore, the receiving end uses the third visual decoding module and the third context visual decoding module to decode the semantic features of the third partition in a similar manner: using the decoding results of the semantic features of the 0th to 2nd partitions as keys and values, and using the semantic features of the third partition as a query to perform cross-attention calculations, and outputting the decoding results of the semantic features of the third partition.

[0060] Through the above decoding process, the receiving end completes the gradual restoration process from the complex-valued symbols received from the channel to the semantic features of the four partitions.

[0061] This embodiment utilizes a cross-attention decoding structure with layer-by-layer context enhancement. While restoring semantic information, it effectively models the contextual relationships between semantic regions, improving the semantic accuracy of decoding and the quality of image reconstruction. Compared to traditional decoding methods that rely on static mapping or fixed parameters, this embodiment can dynamically cross-decode based on the order of actual transmitted semantic features and contextual dependencies, significantly enhancing the system's performance in complex image restoration and adapting to variable channel environments. It is an important technical support for building low-latency, high-fidelity image semantic communication.

[0062] In one embodiment, the neural analyzer includes: 1st to Nth downsampling modules and at least two attention modules, The neural synthesizer includes: first to Nth upsampling modules, at least two attention modules and a feature refinement module; the feature refinement module is connected to the first upsampling module; The first to Nth upsampling modules correspond one to one to the first to Nth downsampling modules.

[0063] In this embodiment, to accurately extract multi-level semantic features from source images and reconstruct high-quality images at the receiving end, the neural analyzer and neural synthesizer in the semantic communication system adopt symmetrical structural designs and introduce an attention mechanism and feature refinement module. The neural analyzer is used to extract semantic features of the image at the sending end. The neural analyzer includes: 1st to Nth downsampling modules and at least two attention modules. The downsampling modules are used to extract spatial hierarchical features of the image layer by layer, and the attention modules are used to enhance the response strength of key areas or task-related semantic information in the image, thereby constructing a high-level semantic representation with discriminative capabilities.

[0064] The neural synthesizer is used to convert the reconstructed semantic features into the final image result at the receiving end. The neural synthesizer includes: 1st to Nth upsampling modules, at least two attention modules, and a feature refinement module. Specifically, the 1st to Nth upsampling modules correspond one-to-one with the 1st to Nth downsampling modules in the neural analyzer, and are used to gradually upsample the semantic features to achieve a gradual restoration of spatial resolution. In addition, at least two attention modules are embedded in the reconstruction path to strengthen the modeling of inter-context associations and improve the accuracy of semantic restoration. The feature refinement module is set at the initial stage of the neural synthesizer and connected to the 1st upsampling module to perform fine-grained restoration and texture completion of the lowest-level semantic features.

[0065] The structural design of the neural analyzer and neural synthesizer described in this embodiment has the following advantages: on the one hand, the downsampling-upsampling symmetric structure can effectively maintain the continuity and stability of the semantic space, avoiding the loss of feature information during the encoding and decoding process; on the other hand, the introduction of the attention mechanism enables the model to have stronger regional discrimination and context integration capabilities; the addition of the feature refinement module further enhances the ability to restore local details and texture information, which helps to improve the image reconstruction quality and perceptual consistency.

[0066] In summary, the neural analyzer and neural synthesizer structure provided in this embodiment not only enhances the end-to-end communication model's ability to express and reconstruct semantic information, but also improves the system's generalization and stability in complex scenarios, laying an important foundation for achieving high-fidelity image transmission in semantic communication.

[0067] In one embodiment, the method performed by the low-latency, high-performance image semantics-channel joint coding system is implemented by a low-latency, high-performance image semantics-channel joint coding model. The training process of the low-latency, high-performance image semantics-channel joint coding model includes at least a first stage and a second stage: Under the condition of freezing the model parameters of the feature encoder to be trained and the feature decoder to be trained, the image semantic encoder based on the quadtree partition module to be trained is trained through the first stage and the second stage to obtain a trained image semantic encoder based on the quadtree partition module; the image semantic encoder based on the quadtree partition module to be trained includes: a neural analyzer to be trained, a neural synthesizer to be trained, a super-prior encoder to be trained, a super-prior decoder to be trained, and an entropy estimator to be trained; analog quantization is used in the first stage, and quantization based on a straight-through estimator is used in the second stage.

[0068] To build a low-latency, high-efficiency, and robust image semantic-channel joint coding system, this embodiment provides a method for training the core model in the coding system. This method is based on a low-latency, high-performance image semantic-channel joint coding model. The model training process includes at least two phases: a first phase and a second phase.

[0069] The training process specifically includes the following steps: First, during training, the network parameters of the feature encoder and feature decoder remain frozen, meaning their model structures do not participate in gradient updates during training. On this basis, only the image semantic encoder, built using the quadtree partitioning module, is trained.

[0070] The image semantic encoder includes the following modules to be trained: a neural analyzer to be trained; a neural synthesizer to be trained; a super-prior encoder to be trained; a super-prior decoder to be trained; and an entropy estimator to be trained.

[0071] In the first stage of training, in order to achieve a differentiable approximation of the quantization process, the encoder output is perturbed by the additive uniform noise method to simulate the quantization behavior while maintaining the differentiability of the training process to facilitate backpropagation optimization.

[0072] In the second stage of training, we switch to using a straight-through estimator for quantization approximation, that is, performing non-differentiable quantization operations during forward propagation, and directly passing gradients during backpropagation, thereby more realistically reflecting the operating behavior of the model in the actual deployment stage.

[0073] The advantages of this two-stage training strategy are: in the first stage, additive noise is used to maintain stability in the initial training of the network, preventing gradient oscillations; in the second stage, a straight-through estimator is used to improve the practical performance of the final model and enhance its adaptability to real-world quantization processes. This phased switching strategy achieves a balance between model trainability and deployment adaptability.

[0074] In summary, the training method described in this embodiment can significantly improve the performance of the semantic encoder in tasks such as multi-partition semantic modeling, low-latency mapping, and context feature estimation, and provides an efficient and practical training paradigm for the construction of high-performance encoding modules in semantic communication systems.

[0075] In one embodiment, the training process of the low-latency, high-performance image semantic-channel joint coding model further includes a third stage and a fourth stage; Unfreezing the model parameters of the feature encoder and the feature decoder to be trained, and integrating them with the trained image semantic encoder based on the quadtree partitioning module to obtain a low-latency, high-performance image semantic-channel joint coding model to be trained, wherein the feature encoder to be trained is connected to the neural analyzer trained in the second stage, and the feature decoder to be trained is connected to the neural synthesizer trained in the second stage; Through the third stage and the fourth stage, the low-latency high-performance image semantic-channel joint coding model to be trained is trained to obtain a trained low-latency high-performance image semantic-channel joint coding model; In the third stage, from the initial rate set Select the symbol length factor vector in ; In the fourth stage, from the target rate set Select the symbol length factor vector in .

[0076] In this embodiment, in order to further improve the overall coding quality and channel adaptability of the low-latency, high-performance image semantic-channel joint coding model, based on the previous embodiment, this embodiment introduces the third and fourth stages to the model training process to achieve end-to-end joint tuning.

[0077] Specifically, after completing the first and second stage training and obtaining the trained image semantic encoder based on the quadtree partitioning module, the original frozen modules, including the feature encoder and feature decoder, are unfrozen, that is, their parameters are enabled to participate in the training process.

[0078] Next, the unfrozen feature encoder to be trained is connected to the trained neural analyzer, and the feature decoder to be trained is connected to the trained neural synthesizer, thereby constructing a complete, end-to-end optimized, low-latency, high-performance image semantic-channel joint coding model to be trained.

[0079] The model is trained in the third and fourth phases as follows: The third stage: with the goal of improving the model's adaptability under different channel bandwidth conditions, starting from a set of initial rate sets The symbol length factor vector is selected for training. This stage of training emphasizes the model's response characteristics to various basic channel capacity conditions, and builds a robust representation capability for different compression ratios and symbol length distributions.

[0080] Phase 4: Focus on more refined target bandwidth control, starting from the target rate set The symbol length factor vector is selected for training. This stage of training guides the model to more accurately map semantic content to the symbol length distribution at different target bit rates to adapt to the transmission requirements of different application scenarios.

[0081] Through the joint training of the third and fourth stages mentioned above, not only the fine-tuning of the feature encoder and feature decoder is achieved, but also the model's rate control capability and semantic preservation capability under different channel conditions are further improved.

[0082] In summary, this embodiment significantly enhances the transmission performance of the semantic communication system in complex and dynamic channel environments through the phased training process of "freeze-train-thaw-joint tuning", ensuring high-quality image semantic reconstruction and precise bit rate matching while maintaining low latency.

[0083] In one embodiment, in the first stage and the second stage, the image semantic encoder based on the quadtree partition module to be trained is trained using a first loss function value, where the first loss function value is determined based on a first difference between a sample original image and a corresponding sample semantically encoded image, where the sample semantically encoded image is output by the image semantic encoder based on the quadtree partition module to be trained for the sample original image; In the third stage and the fourth stage, the low-latency high-performance image semantic-channel joint coding model to be trained is trained using a second loss function value, wherein the second loss function value is determined based on the second difference between the sample original image and the corresponding sample reconstructed image, and the first difference, and the sample reconstructed image is output by the low-latency high-performance image semantic-channel joint coding model to be trained for the sample original image.

[0084] In this embodiment, in order to further improve the training accuracy and convergence stability of the low-latency high-performance image semantic-channel joint coding model, this embodiment introduces a difference-aware multi-stage loss function design during the training process, using the first loss function in the first and second stages, and the second loss function in the third and fourth stages for training.

[0085] Specifically: In the first and second phases, the model only trains the image semantic encoder based on the quadtree partitioning module, and the loss function used in this phase is the first loss function value. The first loss function value is calculated based on the following difference: The input is the sample original image and the sample semantically encoded image after the original image is output by the image semantic encoder to be trained. The first loss function is used to measure the restoration accuracy of the image semantic features during the encoding process. Perceptual loss, structural similarity loss or L2 norm are often used as measurement standards. This design ensures that the model has good semantic extraction and reconstruction capabilities in the early stages of training, laying a solid foundation for subsequent end-to-end training.

[0086] In the third and fourth stages, the model introduces the trained semantic encoding module and performs end-to-end training on the feature encoder and feature decoder. The loss function used in this stage is the second loss function value, which has a more comprehensive definition and combines the following two items: a second difference between the sample original image and the sample reconstructed image generated by the complete model (including channel mapping and decoding); The first difference value calculated in the first stage is used as an auxiliary item for training semantic consistency.

[0087] The second loss function combines image reconstruction error and semantic feature fidelity, and has stronger global consistency constraint capabilities. It not only optimizes the final image quality, but also ensures semantic stability and structural fidelity during transmission.

[0088] In summary, this embodiment strengthens semantic modeling capabilities in the early stages of training through the design of a staged loss function, and achieves end-to-end semantic and reconstruction collaborative optimization in the later stages, which helps to significantly improve the model's generalization ability and robustness under dynamic channels. It is an important optimization strategy for building a low-latency, high-performance image semantic communication system.

[0089] In one embodiment, Figure 5 Schematic diagram of the complete architecture of a low-latency, high-performance image semantic-channel joint coding system provided for this application: The system as a whole consists of a transmitter and a receiver, adopts an end-to-end structured coding method, and integrates neural networks and visual attention mechanisms. While maintaining high semantic expression accuracy, it significantly reduces coding and transmission delays, thereby improving the system's reliability and reconstruction performance under real channel conditions.

[0090] In terms of the system architecture, a neural analyzer first processes the image (original image) to extract semantic features with deep semantic representation capabilities. The semantic features are then partitioned into partitions 0 to 3 along the channel dimension using a quadtree partitioning module. Multi-level contextual information is then constructed using index aggregation and concatenation. During feature encoding, the transmitter encodes the semantic features of partition 0 through a feature encoder, which then serves as context key-value pairs to guide cross-attention modeling and context-aware encoding of subsequent partitions, resulting in the encoding of semantic features for all four partitions. To improve channel compression and mapping capabilities, the system introduces a super-prior encoder and decoder to extract latent variables and generate mean and variance predictions. This assists the entropy estimator in estimating the symbol length factor vector for each partition's semantic features, enabling dynamic compression rate adjustment and symbol-to-complex value mapping. At the receiver, the system gradually recovers the semantic features of each partition using a feature decoder combined with a contextual vision module. Finally, a neural synthesizer completes image reconstruction and outputs the reconstructed image. Notably, this architecture achieves a good balance between semantic context representation and module parallelism by avoiding the latency bottleneck associated with traditional autoregressive coding.

[0091] In terms of model training, this application adopts a four-stage step-by-step training strategy to improve model stability and multi-objective optimization capabilities. In the first two stages, the feature encoder and decoder are frozen, and only the semantic encoder based on the quadtree partition module is trained. The additive noise and pass-through estimator are gradually used to complete the semantic feature learning, and the semantic image reconstruction accuracy is optimized by the first loss function. In the last two stages, all model parameters are unfrozen, and the trained semantic encoder is integrated with the feature encoder and decoder to form a complete semantic-channel joint model, and combined with different sets of symbol length factor vectors to adapt to a variety of bit rate conditions. Finally, by introducing a second loss function that combines image reconstruction error and semantic coding error, collaborative optimization is achieved at both the visual quality and transmission efficiency levels.

[0092] Overall, the present invention constructs an image semantic-channel joint coding system with a clear structure, strong semantic expression, and efficient transmission. While maintaining good image reconstruction performance, it achieves high compression rate, low transmission latency, and excellent channel adaptability. It is suitable for key application scenarios in low-latency intelligent semantic communication systems, such as autonomous driving, telemedicine, and immersive XR. This technical solution not only solves the problems of strong coding dependency, poor parallelism, and fixed symbol granularity in existing JSCC solutions, but also provides important technical support for the development of semantic communication systems towards multi-scale, adaptive, and highly robust directions.

[0093] Each embodiment in this specification focuses on the differences from other embodiments, and the same or similar parts between the embodiments can be referenced to each other.

[0094] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, devices, or computer program products. Therefore, the embodiments of the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the embodiments of the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0095] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0096] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0097] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable terminal device to implement the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0098] Although preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they become aware of the basic inventive concepts. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present invention.

[0099] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or terminal device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or terminal device that includes the element.

[0100] The above is a detailed introduction to a low-latency, high-performance image semantic-channel joint coding system provided. This article uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method of this application and its core idea; at the same time, for general technical personnel in this field, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on this application.

Claims

1. A low-latency, high-performance image semantic-channel joint coding system, characterized in that: include: The sender in the semantic communication system extracts semantic features from the source image through a neural analyzer; The sending end divides the semantic features into the sending end 0th partition semantic features to the sending end 3rd partition semantic features through a quadtree partition module; The transmitting end encodes the semantic feature of the 0th partition of the transmitting end through a feature encoder to obtain a semantic feature encoding result of the 0th partition; The transmitting end encodes the semantic feature of the nth partition of the transmitting end by the feature encoder, using the encoding results of the semantic features of the 0th to n-1th partitions as context, to obtain the encoding result of the semantic feature of the nth partition, where n is an integer between 1 and 3; The transmitting end maps the semantic feature encoding results of the 0th to nth partitions into complex-valued channel symbols through the feature encoder, and sends the complex-valued channel symbols to the receiving end in the semantic communication system; The receiving end obtains a reconstructed semantic feature by using the received complex-valued channel symbols through a feature decoder; The receiving end obtains a reconstructed image by using the reconstructed semantic features through a neural synthesizer.

2. The low-latency, high-performance image semantic-channel joint coding system according to claim 1, characterized in that: The receiving end obtains a reconstructed semantic feature by using the received complex-valued channel symbols through a feature decoder, including: The receiving end determines the semantic features of the receiving end partition 0 to the receiving end partition 3 by using the received complex-valued channel symbols through a feature decoder; The transmitting end encodes the semantic feature of the 0th partition of the transmitting end through a feature encoder to obtain a decoding result of the semantic feature of the 0th partition; The transmitting end decodes the semantic feature of the nth partition of the receiving end through the feature decoder, using the decoding results of the semantic features of the 0th to n-1th partitions as context, to obtain a decoding result of the semantic feature of the nth partition, where n is an integer between 1 and 3; The transmitting end obtains the reconstructed semantic feature according to the decoding results of the semantic features of the 0th to the third partitions.

3. The low-latency, high-performance image semantic-channel joint coding system according to claim 1, characterized in that: Also includes: The transmitting end inputs the semantic feature into a super a priori encoder to obtain a super a priori latent variable, and obtains a super a priori latent variable quantization result based on the super a priori latent variable; The transmitting end obtains the 0th mean and the 0th variance by using the quantization result of the super priori latent variable through the super priori decoder; The transmitting end obtains an nth mean and an nth variance by using an entropy estimator, using the 0th mean and the 0th variance, and the nth context, where n is an integer between 1 and 3; The transmitting end calculates a corresponding symbol length factor vector for each element in the nth partition semantic feature according to the nth mean and the nth variance, and provides the vector to the feature encoder and the feature decoder.

4. The low-latency, high-performance image semantic-channel joint coding system according to claim 3, characterized in that: The sending end divides the semantic features into the sending end 0th partition semantic features to the sending end 3rd partition semantic features through a quadtree partition module, including: The sending end divides the semantic features into four parts along the channel dimension through a quadtree partitioning module; The sending end adds the elements indexed by n in the four parts to obtain the semantic feature of the nth partition, where n is an integer from 0 to 3; Also includes: The sending end connects the elements with index 0 in the four parts to obtain a first context; The sending end connects the element with index 0 and the element with index 1 in the four parts to obtain a second context; The sending end connects the element with index 0, the element with index 1, and the element with index 2 in the four parts to obtain a third context.

5. The low-latency, high-performance image semantic-channel joint coding system according to claim 1, characterized in that: The feature encoder includes at least: a first visual encoding module, a first contextual visual encoding module, a second contextual visual encoding module, and a third contextual visual encoding module; the semantic feature encoding results of the 0th to nth partitions are obtained according to the following steps: The transmitting end performs a first encoding on the semantic features of the 0th to 3rd partitions through the first visual encoding module to obtain an encoding result of the semantic features of the 0th partition and first intermediate encoding results of the 1st to 3rd partitions; The transmitting end performs cross-attention calculation using the first contextual visual encoding module, with the semantic feature encoding result of the 0th partition as the key and value, and the first intermediate encoding result of the 1st partition as the query, to obtain the semantic feature encoding result of the 1st partition; The transmitting end performs cross-attention calculation using the second context visual coding module, with the semantic feature coding results of the 0th to 1st partitions as keys and values, and the first intermediate coding result of the second partition as a query, to obtain the semantic feature coding result of the second partition; The sending end uses the second context visual encoding module to perform cross-attention calculations using the semantic feature encoding results of the 0th to 2nd partitions as keys and values ​​and the first intermediate encoding result of the 3rd partition as a query to obtain the semantic feature encoding result of the 3rd partition.

6. The low-latency, high-performance image semantic-channel joint coding system according to claim 2, characterized in that: The feature decoder includes at least: a zeroth visual decoding module, a first visual decoding module, a first context visual decoding module, a second visual decoding module, a second context visual decoding module, a third visual decoding module, and a third context visual decoding module; the semantic feature decoding results of the 0th to nth partitions are obtained according to the following steps: The receiving end decodes the semantic features of the 0th partition of the sending end through the 0th visual decoding module to obtain a decoding result of the semantic features of the 0th partition; The receiving end performs cross-attention calculation by using the first visual decoding module and the first contextual visual decoding module, with the decoding result of the semantic feature of the 0th partition as the key and value and the semantic feature of the first partition of the receiving end as the query, to obtain the decoding result of the semantic feature of the first partition; The receiving end performs cross-attention calculation using the second visual decoding module and the second contextual visual decoding module, with the decoding results of the semantic features of the 0th to 1st partitions as keys and values, and the semantic features of the second partition of the receiving end as a query, to obtain a decoding result of the semantic features of the second partition; The receiving end performs cross-attention calculation through the third visual decoding module and the third context visual decoding module, using the decoding results of the semantic features of the 0th to 2nd partitions as keys and values, and the semantic features of the 3rd partition of the receiving end as a query to obtain the decoding results of the semantic features of the 3rd partition.

7. The low-latency, high-performance image semantic-channel joint coding system according to claim 1, characterized in that: The neural analyzer includes: 1st to Nth downsampling modules and at least two attention modules, The neural synthesizer includes: first to Nth upsampling modules, at least two attention modules and a feature refinement module; the feature refinement module is connected to the first upsampling module; The first to Nth upsampling modules correspond one to one to the first to Nth downsampling modules.

8. The low-latency, high-performance image semantic-channel joint coding system according to claim 1, characterized in that: The method performed by the low-latency, high-performance image semantics-channel joint coding system is implemented by a low-latency, high-performance image semantics-channel joint coding model. The training process of the low-latency, high-performance image semantics-channel joint coding model includes at least a first stage and a second stage: Under the condition that the model parameters of the feature encoder to be trained and the feature decoder to be trained are frozen, the image semantic encoder based on the quadtree partition module to be trained is trained through the first stage and the second stage to obtain a trained image semantic encoder based on the quadtree partition module; The image semantic encoder based on the quadtree partition module to be trained includes: a neural analyzer to be trained, a neural synthesizer to be trained, a super priori encoder to be trained, a super priori decoder to be trained, and an entropy estimator to be trained; Analog quantization is used in the first stage and straight-through estimator based quantization is used in the second stage.

9. The low-latency, high-performance image semantics-channel joint coding system according to claim 8, characterized in that: The training process of the low-latency, high-performance image semantic-channel joint coding model also includes a third stage and a fourth stage; Unfreezing the model parameters of the feature encoder and the feature decoder to be trained, and integrating them with the trained image semantic encoder based on the quadtree partitioning module to obtain a low-latency, high-performance image semantic-channel joint coding model to be trained, wherein the feature encoder to be trained is connected to the neural analyzer trained in the second stage, and the feature decoder to be trained is connected to the neural synthesizer trained in the second stage; Through the third stage and the fourth stage, the low-latency high-performance image semantic-channel joint coding model to be trained is trained to obtain a trained low-latency high-performance image semantic-channel joint coding model; In the third stage, from the initial rate set Select the symbol length factor vector in ; In the fourth stage, from the target rate set Select the symbol length factor vector in .

10. The low-latency, high-performance image semantics-channel joint coding system according to claim 9, characterized in that: In the first stage and the second stage, the image semantic encoder based on the quadtree partition module to be trained is trained using a first loss function value, wherein the first loss function value is determined according to a first difference between a sample original image and a corresponding sample semantically encoded image, wherein the sample semantically encoded image is output by the image semantic encoder based on the quadtree partition module to be trained for the sample original image; In the third stage and the fourth stage, the low-latency high-performance image semantic-channel joint coding model to be trained is trained using a second loss function value, wherein the second loss function value is determined based on the second difference between the sample original image and the corresponding sample reconstructed image, and the first difference, and the sample reconstructed image is output by the low-latency high-performance image semantic-channel joint coding model to be trained for the sample original image.

Citation Information

Patent Citations

  • Context modeling semantic communication code transmission and receiving method and related equipment

    CN116935840A

  • Semantic communication method and system for high-resolution image

    CN119383363A

  • Lightweight semantic communication image transmission system and method based on parallel mixer

    CN120416492A

  • Encoding method, decoding method, and device

    US20230388490A1

Cited By

  • Robust semantic communication method based on antagonism purification

    CN120930520A

  • A robust semantic communication method based on adversarial purification

    CN120930520B