A low-latency high-performance image semantic-channel joint coding system

By using a joint coding system of neural analyzer and quadtree partitioning module, the problem of fixed coding strategy in JSCC system is solved, realizing low-latency, high-performance image semantic-channel joint coding, and improving the stability and quality of image transmission.

CN120729477BActive Publication Date: 2025-11-25TSINGHUA UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511141493.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-15
Publication Date
2025-11-25
Estimated Expiration
2045-08-15

AI Technical Summary

Technical Problem

Existing JSCC systems suffer from fixed encoding strategies in image and video transmission, which limits their practicality for low-latency and large-scale communication. Furthermore, they are sensitive to channel errors in high-resolution or complex semantic scenarios, affecting transmission stability and user experience.

Method used

A neural analyzer is used to extract semantic features, which are then divided into multiple partitions using a quadtree partitioning module. A feature encoder is used for progressive context-aware encoding, which is finally mapped to complex-valued channel symbols. The receiver reconstructs the semantic features and images using a feature decoder.

Benefits of technology

It significantly reduces system latency, improves parallel processing capabilities and resistance to channel interference, enhances image reconstruction quality and transmission stability, and adapts to image transmission requirements under different channel conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120729477B_ABST
    Figure CN120729477B_ABST
Patent Text Reader

Abstract

The application provides a low-delay high-performance image semantic-channel joint coding system, and relates to the technical fields of communication and image processing, and comprises the following steps: a sending end extracts semantic features through a neural analyzer, and divides the semantic features into multiple partitions by using a quadtree structure; each partition is encoded step by step by using a context-aware feature coding strategy, and finally mapped into complex-valued channel symbols and sent to a receiving end. The receiving end reconstructs the semantic features and generates a reconstructed image through a corresponding decoding and synthesizing module, and realizes efficient image transmission in an end-to-end mode. The application introduces a context-based partition coding mechanism and a quadtree partition module strategy in the coding structure, has stronger context modeling capability and parallel processing capability, can effectively improve the adaptive capability for image structure and semantic content, and is helpful to reduce transmission delay, enhance image reconstruction quality and channel interference resistance.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of communication and image processing, and in particular to a low-latency high-performance image semantic-channel joint coding system. BACKGROUND

[0002] With the deepening of the research on 6G communication technology, semantic communication is rapidly developing towards semantic perception and intelligent fusion. In order to effectively alleviate the bandwidth bottleneck and channel noise interference problem in image and video transmission, deep semantic communication has gradually become a research hotspot. Among them, the JSCC (Joint Semantic-Channel Coding) framework is considered an important part of the next generation communication architecture due to its significant advantages in transmission efficiency and anti-interference.

[0003] However, in the existing JSCC system, the encoding of semantic features and the channel transmission process are mostly separated from each other, and the encoding strategy often uses fixed modules and static parameter configurations, which cannot be flexibly adapted according to different image characteristics, task requirements and channel states, resulting in limited practicability of the overall system in low-latency large-scale image communication.

[0004] Especially in the semantic encoding module, the current mainstream method mostly relies on entropy estimation strategy based on autoregressive modeling. Although this method has good compression performance, due to its strong sequence dependence, it greatly inhibits parallel processing capability, resulting in significant encoding delay, which is difficult to meet the stringent requirements of real-time semantic communication for low latency.

[0005] In addition, the existing JSCC system generally lacks flexibility and context adaptation capability in key links such as feature transformation, symbol mapping and side information construction. For example, the symbol dimension usually adopts fixed configuration, which cannot dynamically adjust the transmission bit quantity combined with the local complexity of the image, wasting bandwidth and affecting image quality reconstruction. In high-resolution or complex semantic scenarios, the system is highly sensitive to channel errors, and slight decoding errors may cause semantic restoration failure, seriously affecting transmission stability and user experience.

[0006] Therefore, there is an urgent need for a low-latency high-performance image semantic-channel joint coding system. SUMMARY

[0007] In view of the above problems, the embodiments of the present application provide a low-latency high-performance image semantic-channel joint coding system in order to overcome the above problems or at least partially solve the above problems.

[0008] In a first aspect, the embodiments of the present application provide a low-latency high-performance image semantic-channel joint coding system, comprising:

[0009] The sending end in the semantic communication system extracts semantic features from a source image through a neural analyzer;

[0010] The sending end divides the semantic features into sending end 0th partition semantic features to sending end third partition semantic features through a quadtree partitioning module;

[0011] The sending end encodes the sending end 0th partition semantic features through a feature encoder to obtain 0th partition semantic feature encoding results;

[0012] The sending end encodes the sending end nth partition semantic features through the feature encoder with the 0th to (n-1)th partition semantic feature encoding results as contexts to obtain nth partition semantic feature encoding results, n being an integer between 1 and 3;

[0013] The sending end maps the 0th to nth partition semantic feature encoding results into complex-valued channel symbols through the feature encoder and sends them to the receiving end in the semantic communication system;

[0014] The receiving end obtains reconstructed semantic features through a feature decoder using the received complex-valued channel symbols;

[0015] The receiving end obtains a reconstructed image through a neural synthesizer using the reconstructed semantic features.

[0016] Optionally, the receiving end obtains reconstructed semantic features through a feature decoder using the received complex-valued channel symbols, comprising:

[0017] The receiving end determines the receiving end 0th partition semantic features to receiving end third partition semantic features through a feature decoder using the received complex-valued channel symbols;

[0018] The sending end encodes the sending end 0th partition semantic features through a feature encoder to obtain 0th partition semantic feature decoding results;

[0019] The sending end decodes the receiving end nth partition semantic features through the feature decoder with the 0th to (n-1)th partition semantic feature decoding results as contexts to obtain nth partition semantic feature decoding results, n being an integer between 1 and 3;

[0020] The sending end obtains the reconstructed semantic features according to the 0th to third partition semantic feature decoding results.

[0021] Optionally, it further comprises:

[0022] The sending end inputs the semantic feature into a hyper-prior encoder to obtain a hyper-prior latent variable, and obtains a hyper-prior latent variable quantization result according to the hyper-prior latent variable;

[0023] The sending end obtains the 0th mean and the 0th variance by using the hyper-prior latent variable quantization result through a hyper-prior decoder;

[0024] The sending end obtains the nth mean and the nth variance by using the 0th mean and the 0th variance, and the nth context through an entropy estimator, n being an integer between 1 and 3;

[0025] The sending end calculates a corresponding symbol length factor vector for each element in the nth partition semantic feature according to the nth mean and the nth variance, and provides the feature encoder and the feature decoder with the corresponding symbol length factor vector.

[0026] Optionally, the sending end divides the semantic feature into the 0th partition semantic feature to the third partition semantic feature of the sending end through a quadtree partition module, including:

[0027] The sending end divides the semantic feature into four parts along the channel dimension through the quadtree partition module;

[0028] The sending end adds elements with index n in the four parts to obtain the nth partition semantic feature, n being an integer between 0 and 3;

[0029] Further comprising:

[0030] The sending end connects elements with index 0 in the four parts to obtain the first context;

[0031] The sending end connects elements with index 0 and elements with index 1 in the four parts to obtain the second context;

[0032] The sending end connects elements with index 0, elements with index 1 and elements with index 2 in the four parts to obtain the third context.

[0033] Optionally, the feature encoder at least includes a first visual coding module, a first context visual coding module, a second context visual coding module and a third context visual coding module; the 0th to nth partition semantic feature coding results are obtained according to the following steps:

[0034] The sending end performs first coding on the 0th to third partition semantic features through the first visual coding module to obtain the 0th partition semantic feature coding result and the first to third partition first intermediate coding results;

[0035] The sending end performs cross-attention calculation by taking the 0th partition semantic feature coding result as the key and value and taking the 1st partition first intermediate coding result as the query through the first context visual coding module, to obtain the 1st partition semantic feature coding result.

[0036] The sending end performs cross-attention calculation by taking the 0th to 1st partition semantic feature coding result as the key and value and taking the 2nd partition first intermediate coding result as the query through the second context visual coding module, to obtain the 2nd partition semantic feature coding result.

[0037] The sending end performs cross-attention calculation by taking the 0th to 2nd partition semantic feature coding result as the key and value and taking the 3rd partition first intermediate coding result as the query through the second context visual coding module, to obtain the 3rd partition semantic feature coding result.

[0038] Optionally, the feature decoder comprises: a zeroth visual decoding module, a first visual decoding module, a first context visual decoding module, a second visual decoding module, a second context visual decoding module, a third visual decoding module, and a third context visual decoding module; and the 0th to nth partition semantic feature decoding results are obtained by the following steps:

[0039] The receiving end decodes the 0th partition semantic feature of the sending end through the zeroth visual decoding module to obtain the 0th partition semantic feature decoding result.

[0040] The receiving end performs cross-attention calculation by taking the 0th partition semantic feature decoding result as the key and value and taking the 1st partition semantic feature of the receiving end as the query through the first visual decoding module and the first context visual decoding module, to obtain the 1st partition semantic feature decoding result.

[0041] The receiving end performs cross-attention calculation by taking the 0th to 1st partition semantic feature decoding result as the key and value and taking the 2nd partition semantic feature of the receiving end as the query through the second visual decoding module and the second context visual decoding module, to obtain the 2nd partition semantic feature decoding result.

[0042] The receiving end performs cross-attention calculation by taking the 0th to 2nd partition semantic feature decoding result as the key and value and taking the 3rd partition semantic feature of the receiving end as the query through the third visual decoding module and the third context visual decoding module, to obtain the 3rd partition semantic feature decoding result.

[0043] Optionally, the neural analyzer comprises: 1st to Nth down-sampling modules and at least two attention modules,

[0044] The neural synthesizer comprises: 1st to Nth up-sampling modules, at least two attention modules, and a feature refinement module; the feature refinement module is connected with the 1st up-sampling module;

[0045] The 1st to Nth up-sampling modules correspond to the 1st to Nth down-sampling modules one by one.

[0046] Optionally, the method performed by the low-latency high-performance image semantic-channel joint coding system is implemented through a low-latency high-performance image semantic-channel joint coding model, and a training process of the low-latency high-performance image semantic-channel joint coding model comprises at least a first stage and a second stage.

[0047] Under the condition of freezing the model parameters of the to-be-trained feature encoder and the to-be-trained feature decoder, the to-be-trained image semantic encoder based on the quadtree partitioning module is trained through the first stage and the second stage to obtain a trained image semantic encoder based on the quadtree partitioning module; the to-be-trained image semantic encoder based on the quadtree partitioning module comprises a to-be-trained neural analyzer, a to-be-trained neural synthesizer, a to-be-trained hyper-prior encoder, a to-be-trained hyper-prior decoder, and a to-be-trained entropy estimator; the first stage uses analog quantization, and the second stage uses quantization based on a pass-through estimator.

[0048] Optionally, the training process of the low-latency high-performance image semantic-channel joint coding model further comprises a third stage and a fourth stage.

[0049] The model parameters of the to-be-trained feature encoder and the to-be-trained feature decoder are unfrozen, and are integrated with the trained image semantic encoder based on the quadtree partitioning module to obtain a to-be-trained low-latency high-performance image semantic-channel joint coding model, wherein the to-be-trained feature encoder is connected with the neural analyzer trained in the second stage, and the to-be-trained feature decoder is connected with the neural synthesizer trained in the second stage.

[0050] The to-be-trained low-latency high-performance image semantic-channel joint coding model is trained through the third stage and the fourth stage to obtain a trained low-latency high-performance image semantic-channel joint coding model.

[0051] In the third stage, a symbol length factor vector is selected from an initial rate set .

[0052] In the fourth stage, a symbol length factor vector is selected from a target rate set .

[0053] Optionally, in the first stage and the second stage, the first loss function value is used to train the to-be-trained image semantic encoder based on the quadtree partition module, and the first loss function value is determined according to a first difference between a sample original image and a corresponding sample semantic encoded image, and the sample semantic encoded image is output by the to-be-trained image semantic encoder based on the quadtree partition module for the sample original image.

[0054] In the third stage and the fourth stage, the second loss function value is used to train the to-be-trained low-latency high-performance image semantic-channel joint encoding model, and the second loss function value is determined according to a second difference between the sample original image and a corresponding sample reconstructed image, and the first difference, and the sample reconstructed image is output by the to-be-trained low-latency high-performance image semantic-channel joint encoding model for the sample original image.

[0055] The beneficial effects of the present application are as follows:

[0056] The present application provides a low-latency high-performance image semantic-channel joint encoding system, comprising: a sending end in a semantic communication system extracts semantic features from a source image through a neural analyzer; the sending end divides the semantic features into sending end 0th partition semantic features to sending end third partition semantic features through a quadtree partition module; the sending end encodes the sending end 0th partition semantic features through a feature encoder to obtain a 0th partition semantic feature encoding result; the sending end encodes the sending end nth partition semantic features through the feature encoder using the 0th to (n-1)th partition semantic feature encoding results as context to obtain an nth partition semantic feature encoding result, n being an integer between 1 and 3; the sending end maps the 0th to nth partition semantic feature encoding results into complex-valued channel symbols through the feature encoder and sends them to a receiving end in the semantic communication system; the receiving end obtains reconstructed semantic features through a feature decoder using the received complex-valued channel symbols; and the receiving end obtains a reconstructed image through a neural synthesizer using the reconstructed semantic features.

[0057] The low-latency high-performance image semantic-channel joint encoding system provided by the present application comprises a sending end extracting semantic features through a neural analyzer and dividing the semantic features into multiple partitions using a quadtree structure; using a context-aware feature encoding strategy to gradually encode each partition, and finally mapping into complex-valued channel symbols to be sent to a receiving end. The receiving end then reconstructs the semantic features and generates a reconstructed image through a corresponding decoding and synthesizing module, realizing efficient image transmission end-to-end.

[0058] The application introduces a context-based partition coding mechanism and a four-tree partition module strategy in the coding structure, has stronger context modeling capability and parallel processing capability, can effectively improve the adaptive capability to image structure and semantic content, and helps to reduce transmission delay, enhance image reconstruction quality and channel interference resistance. BRIEF DESCRIPTION OF DRAWINGS

[0059] In order to more clearly illustrate the technical solutions of the embodiments of the application, the drawings needed to be used in the description of the embodiments of the application will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the application, and other drawings can be obtained by those skilled in the art without creative labor.

[0060] Figure 1 is a low-latency high-performance image semantic-channel joint coding system provided by the embodiments of the application;

[0061] Figure 2 is a process block diagram of introducing an auxiliary information generation process based on hyper-prior modeling provided by the embodiments of the application;

[0062] Figure 3 is a structure block diagram of a feature encoder provided by the embodiments of the application;

[0063] Figure 4 is a structure block diagram of a feature decoder provided by the embodiments of the application;

[0064] Figure 5 is a complete architecture schematic diagram of a low-latency high-performance image semantic-channel joint coding system provided by the embodiments of the application. DETAILED DESCRIPTION

[0065] The exemplary embodiments of the application will be described in more detail below with reference to the drawings of the embodiments of the application. Although the exemplary embodiments of the application are shown in the drawings, it should be understood that the application can be implemented in various forms and should not be limited by the embodiments described herein. On the contrary, these embodiments are provided to enable a more thorough understanding of the application and to fully convey the scope of the application to those skilled in the art.

[0066] The first aspect of the embodiments of the application provides a low-latency high-performance image semantic-channel joint coding system, as shown in Figure 1 The system includes a sending end and a receiving end, wherein the sending end includes a neural analyzer, a four-tree partition module and a feature encoder, and the receiving end includes a feature decoder and a neural synthesizer.

[0067] The sending end in the semantic communication system extracts semantic features from a source image through the neural analyzer;

[0068] The sending end divides the semantic features into sending end 0th partition semantic features to sending end third partition semantic features through a quadtree partitioning module.

[0069] The sending end encodes the sending end 0th partition semantic features through a feature encoder to obtain 0th partition semantic feature encoding results.

[0070] The sending end encodes the sending end nth partition semantic features through the feature encoder with the 0th to (n-1)th partition semantic feature encoding results as context to obtain nth partition semantic feature encoding results, n being an integer between 1 and 3.

[0071] The sending end maps the 0th to nth partition semantic feature encoding results into complex-valued channel symbols through the feature encoder and sends them to the receiving end in the semantic communication system.

[0072] The receiving end obtains reconstructed semantic features through a feature decoder using the received complex-valued channel symbols.

[0073] The receiving end obtains a reconstructed image through a neural synthesizer using the reconstructed semantic features.

[0074] Specifically, as shown in Figure 1 The system described in the present application includes a sending end and a receiving end, wherein the sending end includes a neural analyzer, a quadtree partitioning module, and a feature encoder, and the receiving end includes a feature decoder and a neural synthesizer.

[0075] The sending end of the system described in the present application first extracts rich semantic features from a source image using a neural analyzer. In order to improve compression and encoding efficiency, the system further adopts a quadtree partitioning strategy to divide the semantic features into four partitions (0th to 3rd partitions), which can be structurally divided according to the spatial distribution of image semantic information, taking into account both local correlation and global context. Then, the system adopts a step-by-step context-aware encoding method: first, the 0th partition is encoded, and then when the nth partition (n = 1 to 3) is encoded, the encoding results of the previous n-1 partitions are introduced as context information. This design effectively improves the relevance of feature compression and the completeness of semantic expression, while reducing redundant information.

[0076] After the encoding is completed, the encoding results of each partition are mapped into complex-valued channel symbols by the feature encoder, so as to be directly transmitted on the communication channel, compatible with the wireless physical layer modulation requirements. After receiving the channel symbols, the receiving end sequentially reconstructs the semantic features through the feature decoder, and restores the complete image through the neural synthesizer. The synthesis process is completed based on end-to-end training, and the image quality can be maximally recovered under the limited bit rate and channel conditions.

[0077] Through the above structural design, the embodiment realizes a highly integrated semantic communication process. Compared with the existing SC system based on the autoregressive structure, the embodiment replaces the full-sequence autoregressive entropy estimation with partitioned context modeling, significantly reduces the overall delay of the system, and enhances the parallel computing capability. At the same time, the structure has good flexibility, and is suitable for the semantic compression and symbol mapping requirements under different channel conditions. It is a highly practical scheme for 6G semantic communication scenarios.

[0078] In an embodiment, the receiving end obtains the reconstructed semantic features by using the received complex-valued channel symbols through the feature decoder, including:

[0079] The receiving end determines the receiving end 0th partition semantic feature to the receiving end third partition semantic feature by using the received complex-valued channel symbols through the feature decoder.

[0080] The sending end encodes the sending end 0th partition semantic feature through the feature encoder to obtain the 0th partition semantic feature decoding result.

[0081] The sending end decodes the receiving end nth partition semantic feature by using the 0th to (n-1)th partition semantic feature decoding results as the context through the feature decoder to obtain the nth partition semantic feature decoding result, where n is an integer between 1 and 3.

[0082] The sending end obtains the reconstructed semantic features according to the 0th to third partition semantic feature decoding results.

[0083] Specifically, in the embodiment, the receiving end receives the complex-valued channel symbols processed by the sending end, and inputs them into the feature decoder. The feature decoder is used to parse the semantic features of each partition from the signals, including the receiving end 0th partition semantic feature, the receiving end 1st partition semantic feature, the receiving end 2nd partition semantic feature, and the receiving end 3rd partition semantic feature.

[0084] In this process, the receiving end first decodes the receiving end 0th partition semantic feature to obtain the decoding result of the 0th partition semantic feature. Since this partition is the decoding starting area, the decoding process does not depend on other context information, and can be directly completed.

[0085] Subsequently, for the nth partition semantic feature (where n is an integer between 1 and 3), the receiving end introduces a context modeling mechanism in the decoding process, that is, the decoding results of the previous n-1 partitions are used as context input to assist the decoding operation of the current nth partition. Specifically, the feature decoder uses the semantic feature decoding results of the 0th to (n-1)th partitions as references to predict and reconstruct the nth partition semantic feature through a context-aware mechanism, thereby obtaining the corresponding nth partition semantic feature decoding result.

[0086] When the receiving end completes the decoding of the 0th to 3rd partition semantic features, it can fuse the decoding results of the four partitions to obtain the complete reconstructed semantic feature. This reconstructed semantic feature will be used as input for the subsequent neural synthesizer to further generate the corresponding reconstructed image.

[0087] The context-aware decoding strategy used in this embodiment improves the reconstruction capability of high-dimensional semantic representation with the guidance of prior partition features, and avoids the long-chain dependency and decoding delay problems caused by traditional autoregressive entropy estimation in the structure design. This embodiment significantly reduces the decoding delay through parallel optimization of partition-level context modeling and decoding, and improves the decoding robustness and overall image reconstruction quality.

[0088] In one embodiment, the sending end inputs the semantic feature into a hyper-prior encoder to obtain a hyper-prior latent variable, and obtains a hyper-prior latent variable quantization result according to the hyper-prior latent variable;

[0089] The sending end obtains the 0th mean and the 0th variance through a hyper-prior decoder using the hyper-prior latent variable quantization result;

[0090] The sending end obtains the nth mean and the nth variance through an entropy estimator using the 0th mean and the 0th variance, and the nth context, where n is an integer between 1 and 3;

[0091] The sending end calculates a corresponding sign length factor vector for each element in the nth partition semantic feature according to the nth mean and the nth variance, and provides it to the feature encoder and the feature decoder.

[0092] In this embodiment, as shown in Figure 2 To further improve the modeling accuracy of the semantic feature distribution in the encoding stage and enhance the adaptive ability of the system to different spatial semantic complexity, the sending end introduces an auxiliary information generation process based on hyper-prior modeling before partitioning the semantic feature for encoding, which is used to guide the accurate estimation of the sign length factor. Specifically, the steps include:

[0093] Firstly, the sender inputs the extracted semantic features into a hyper-prior encoder to generate latent variables for modeling the distribution of the features, denoted as hyper-prior latent variables. The hyper-prior latent variables can capture the statistical structure of the whole semantic features, which helps to more accurately depict the probability distribution thereof.

[0094] Then, the sender quantizes the hyper-prior latent variables to obtain quantization results of the hyper-prior latent variables, so as to reduce the bit overhead of the information in the transmission process and ensure the repeatability of subsequent decoding. The quantization results are then input into a hyper-prior decoder to generate statistical parameters for initializing the probability model, including the mean (0th mean) and variance (0th variance) of the 0th partition semantic features.

[0095] Further, the sender calculates the probability estimation parameters required by other partitions (nth partition, n is 1 to 3) in combination with the initial statistical parameters and context modeling information. Specifically, the sender uses an entropy estimator to obtain the mean (nth mean) and variance (nth variance) of the nth partition semantic features based on the 0th mean and 0th variance in combination with the nth context information (i.e., the semantic feature encoding results of the previous n−1 partitions). The context-driven estimation method ensures that the statistical model relied on by the partition features during encoding has high relevance, thereby improving the entropy modeling accuracy.

[0096] After obtaining the statistical parameters of the nth partition, the sender further calculates the symbol length factor vector corresponding to each element in the nth partition semantic features according to the nth mean and nth variance. The vector represents the bit contribution weight of each position in the semantic features, which is used to guide the feature encoder to implement differential compression on different importance areas, and the factor is also provided to the feature decoder simultaneously to restore the consistency of entropy modeling in the decoding path.

[0097] The embodiment introduces a mechanism combining hyper-prior modeling and context entropy estimation before feature encoding to construct a precise, dynamic, and context-aware symbol length prediction framework. The embodiment improves the entropy estimation accuracy while significantly improving parallelism and reducing system encoding delay. In particular, in the compression scene of high-resolution or complex structure images, the mechanism can significantly alleviate the contradiction between bit rate and reconstruction quality, achieve better compression rate control and robustness, and is one of the important support means for realizing 6G semantic communication low latency and large-scale image transmission.

[0098] In an embodiment, the sender divides the semantic features into 0th partition semantic features to third partition semantic features of the sender through a quadtree partitioning module, including:

[0099] The sender divides the semantic features into four parts along the channel dimension through the quadtree partitioning module.

[0100] The sending end adds the element indexed as n in the four parts to obtain the nth partition semantic feature, n being an integer from 0 to 3;

[0101] Further comprising:

[0102] The sending end concatenates the element indexed as 0 in the four parts to obtain the 1st context;

[0103] The sending end concatenates the element indexed as 0 and the element indexed as 1 in the four parts to obtain the 2nd context;

[0104] The sending end concatenates the element indexed as 0, the element indexed as 1 and the element indexed as 2 in the four parts to obtain the 3rd context.

[0105] In the embodiment, in order to improve the structure modeling capability and context controllability of the semantic feature in the encoding process, the sending end performs partitioning operation on the extracted semantic feature through the quadtree partitioning module, divides the overall semantic feature into multiple partitions to realize local refinement modeling and context construction step by step, and specifically includes the following steps:

[0106] Firstly, the sending end divides the original semantic feature along the channel dimension through the quadtree partitioning module. Specifically, the semantic feature is divided into four parts, which correspond to four initial sub-regions in subsequent encoding processing.

[0107] After the division is completed, the sending end performs summary processing on the element indexed as n in each part. Specifically, for each integer n (n takes values from 0 to 3), the sending end adds (or performs equivalent aggregation operation) the element indexed as n in the four parts, thereby obtaining the corresponding nth partition semantic feature. This operation enables the semantic feature to have a certain degree of spatial semantic fusion in the partition structure, which helps to improve the semantic integrity and decoding stability of each partition.

[0108] In order to support the context modeling in the subsequent partition encoding process, the embodiment further constructs multiple levels of context information for guiding the encoding and entropy estimation process of the nth partition (n = 1 to 3):

[0109] The sending end concatenates all the elements indexed as 0 in the four parts to form the 1st context, which is used for the encoding reference of the 1st partition semantic feature;

[0110] The sending end concatenates the element indexed as 0 and the element indexed as 1 to form the 2nd context, which is used as the context input for the encoding of the 2nd partition semantic feature;

[0111] Similarly, the sender concatenates the elements with indexes 0, 1 and 2 in sequence to form the 3rd context, which is provided to the 3rd partition for encoding and entropy estimation.

[0112] The quadtree partitioning strategy in this embodiment not only realizes effective division of semantic features, but also provides progressive semantic reference information for encoding of different partitions through the step-by-step expansion context construction mechanism. This embodiment constructs contexts through structured connection, greatly improves the controllability and parallel computing capability of contexts, and significantly reduces the overall system delay while maintaining encoding accuracy.

[0113] In an embodiment, the feature encoder at least includes: a first visual encoding module, a first context visual encoding module, a second context visual encoding module, and a third context visual encoding module; and the 0th to nth partition semantic feature encoding results are obtained according to the following steps:

[0114] The sender performs first encoding on the 0th to 3rd partition semantic features through the first visual encoding module to obtain the 0th partition semantic feature encoding result and the 1st to 3rd partition first intermediate encoding results;

[0115] The sender performs cross-attention calculation on the 1st partition first intermediate encoding result with the 0th partition semantic feature encoding result as the key and value through the first context visual encoding module to obtain the 1st partition semantic feature encoding result;

[0116] The sender performs cross-attention calculation on the 2nd partition first intermediate encoding result with the 0th to 1st partition semantic feature encoding results as the key and value through the second context visual encoding module to obtain the 2nd partition semantic feature encoding result;

[0117] The sender performs cross-attention calculation on the 3rd partition first intermediate encoding result with the 0th to 2nd partition semantic feature encoding results as the key and value through the second context visual encoding module to obtain the 3rd partition semantic feature encoding result.

[0118] In this embodiment, in order to further improve the information extraction capability and context modeling accuracy in the semantic feature encoding process, the feature encoder is designed to include multiple sub-modules with cross-attention mechanism, which are used to construct progressive context relationships for different partition semantic features, as shown in the structural block diagram of the feature encoder. Figure 3 The structural block diagram of the feature encoder specifically includes the following sub-modules: a first visual encoding module, a first context visual encoding module, a second context visual encoding module, and a third context visual encoding module.

[0119] The 0th to nth partition semantic feature encoding results are obtained according to the following steps:

[0120] Firstly, the sender inputs the 0th to 3rd partition semantic features into the first visual encoding module in sequence, and performs preliminary encoding on each partition feature. The module outputs include: the 0th partition semantic feature encoding result, and the first intermediate encoding results of the 1st, 2nd and 3rd partitions, which will be used as the input basis for subsequent context-aware encoding.

[0121] Then, the sender performs context enhancement processing on the 1st partition semantic feature through the first context visual encoding module. Specifically, the module takes the 0th partition semantic feature encoding result as the “key” and value input, takes the first intermediate encoding result of the 1st partition as the “query”, performs cross-attention calculation, and finally outputs the 1st partition semantic feature encoding result. This mechanism can model the dependence of the 1st partition on the 0th partition semantic information, thereby improving the semantic continuity and compression expression.

[0122] Subsequently, for the 2nd partition semantic feature, the sender performs similar processing through the second context visual encoding module. The module takes the 0th to 1st partition semantic feature encoding results as the key and value, takes the 2nd partition first intermediate encoding result as the query, also performs cross-attention calculation, and obtains the 2nd partition semantic feature encoding result, realizing the full perception of the 2nd partition on the encoding information of the previous two partitions.

[0123] Further, to process the 3rd partition semantic feature, the sender again uses the third context visual encoding module, which takes the 0th to 2nd partition semantic feature encoding results as the key and value, takes the 3rd partition first intermediate encoding result as the query, performs cross-attention calculation, and outputs the 3rd partition semantic feature encoding result.

[0124] Through the above progressive encoding path, the encoding results of the 0th to 3rd partition semantic features are finally obtained.

[0125] The cross-attention context modeling mechanism used in this embodiment can significantly improve the context interaction efficiency between different semantic regions, breaking the traditional context estimation mode which relies on serial processing and supporting higher degree of parallel execution. By introducing the attention enhancement path of key-value-query structure based on feature partitioning and combining the representation weights of different semantic regions, the expression ability of the semantic encoder and the modeling ability of the spatial dependency structure in complex scenarios are effectively enhanced. This embodiment reduces the overall delay while maintaining the encoding accuracy, and can effectively improve the real-time image transmission quality of the 6G semantic communication system.

[0126] In an embodiment, the feature decoder at least comprises: a zeroth visual decoding module, a first visual decoding module, a first context visual decoding module, a second visual decoding module, a second context visual decoding module, a third visual decoding module, and a third context visual decoding module; and the 0th to nth partition semantic feature decoding results are obtained by the following steps:

[0127] The receiving end decodes the 0th partition semantic feature of the sending end by the zeroth visual decoding module to obtain a 0th partition semantic feature decoding result.

[0128] The receiving end performs cross-attention calculation by taking the 0th partition semantic feature decoding result as the key and value, and taking the 1st partition semantic feature of the receiving end as the query, by the first visual decoding module and the first context visual decoding module, to obtain a 1st partition semantic feature decoding result.

[0129] The receiving end performs cross-attention calculation by taking the 0th to 1st partition semantic feature decoding results as the key and value, and taking the 2nd partition semantic feature of the receiving end as the query, by the second visual decoding module and the second context visual decoding module, to obtain a 2nd partition semantic feature decoding result.

[0130] The receiving end performs cross-attention calculation by taking the 0th to 2nd partition semantic feature decoding results as the key and value, and taking the 3rd partition semantic feature of the receiving end as the query, by the third visual decoding module and the third context visual decoding module, to obtain a 3rd partition semantic feature decoding result.

[0131] In this embodiment, in order to ensure that the semantic features decoded from the complex-valued channel symbols at the receiving end can be accurately restored layer by layer, and to fully utilize the context information between the semantic partitions to improve the decoding accuracy, as shown in the structural block diagram of the feature decoder, Figure 4 The feature decoder is designed to have a multi-level attention enhancement structure, including the following sub-modules: a zeroth visual decoding module, a first visual decoding module, a first context visual decoding module, a second visual decoding module, a second context visual decoding module, a third visual decoding module, and a third context visual decoding module.

[0132] The 0th to nth partition semantic feature decoding results are obtained by the following steps:

[0133] First, the receiving end receives the complex-valued channel symbols transmitted by the sending end and maps them to a plurality of partition semantic features, including the 0th to 3rd partition semantic features of the receiving end. The receiving end first decodes the 0th partition semantic feature by the zeroth visual decoding module to obtain a 0th partition semantic feature decoding result.

[0134] Then, the receiving end uses the first visual decoding module and the first context visual decoding module to cooperatively decode the first-part semantic feature of the receiving end. Specifically, the module takes the decoding result of the 0th-part semantic feature as the key (key) and value (value), takes the first-part semantic feature of the receiving end as the query (query), and generates the decoding result of the first-part semantic feature by performing cross-attention calculation. This step realizes the context-aware modeling of the first-part to the previous region.

[0135] Further, the receiving end decodes the second-part semantic feature through the second visual decoding module and the second context visual decoding module. The decoding operation takes the decoding results of the 0th-part to the first-part semantic features as the key and the value, takes the second-part semantic feature of the receiving end as the query, completes the cross-attention calculation, and obtains the decoding result of the second-part semantic feature.

[0136] Further, the receiving end uses the third visual decoding module and the third context visual decoding module to decode the third-part semantic feature. The way is similar: taking the decoding results of the 0th-part to the second-part semantic features as the key and the value, taking the third-part semantic feature as the query to perform cross-attention calculation, and outputting the decoding result of the third-part semantic feature.

[0137] Through the above decoding process, the receiving end completes the step-by-step restoration process from the complex-valued symbols received from the channel to the four-part semantic features.

[0138] The embodiment adopts a cross-attention decoding structure with layer-by-layer context enhancement, effectively models the context relationship between the semantic regions while realizing the restoration of semantic information, and improves the semantic accuracy and image reconstruction quality of decoding. Compared with the traditional decoding method relying on static mapping or fixed parameters, the embodiment can dynamically cross-decode according to the actual transmission order and context dependence of the semantic features, significantly enhance the performance of the system in complex image restoration, variable channel environment adaptation, and other aspects, and is an important technical support for building low-delay high-fidelity image semantic communication.

[0139] In one embodiment, the neural analyzer includes a first to an Nth down-sampling module and at least two attention modules,

[0140] The neural synthesizer includes a first to an Nth up-sampling module, at least two attention modules, and a feature refinement module; the feature refinement module is connected with the first up-sampling module.

[0141] The first to the Nth up-sampling modules correspond to the first to the Nth down-sampling modules one by one.

[0142] In this embodiment, in order to realize accurate extraction of multi-level semantic features from a source image and high-quality reconstruction of the image at a receiving end, the neural analyzer and the neural synthesizer in the semantic communication system are designed with a symmetrical structure, and an attention mechanism and a feature refinement module are introduced. The neural analyzer is used to extract semantic features of the image at the sending end, and includes a first to an Nth down-sampling module and at least two attention modules. The down-sampling module is used to extract spatial level features of the image layer by layer, and the attention module is used to enhance the response strength of key areas or task-related semantic information in the image, thereby constructing a high-level semantic representation with discrimination ability.

[0143] The neural synthesizer is used to convert the reconstructed semantic features into a final image result at the receiving end, and includes a first to an Nth up-sampling module, at least two attention modules and a feature refinement module. The first to the Nth up-sampling modules correspond one-to-one in structure to the first to the Nth down-sampling modules in the neural analyzer, are used to up-sample the semantic features step by step, and realize step-by-step recovery of spatial resolution. In addition, the at least two attention modules are embedded in the reconstruction path, are used to strengthen the association modeling between contexts, and improve the accuracy of semantic restoration. The feature refinement module is arranged at an initial stage of the neural synthesizer, is connected with the first up-sampling module, and is used to perform fine-grained restoration and texture completion on the bottom-level semantic features.

[0144] The structural design of the neural analyzer and the neural synthesizer has the following advantages: on the one hand, through the symmetrical structure of down-sampling and up-sampling, the continuity and stability of the semantic space can be effectively maintained, and the loss of feature information in the encoding and decoding process can be avoided; on the other hand, the introduction of the attention mechanism enables the model to have stronger regional discrimination ability and context integration ability; the addition of the feature refinement module further enhances the restoration ability of local details and texture information, which helps to improve the image reconstruction quality and perceptual consistency.

[0145] In summary, the structure of the neural analyzer and the neural synthesizer provided in this embodiment not only enhances the expression and reconstruction ability of the end-to-end communication model for semantic information, but also improves the generalization and stability of the system in complex scenes, thereby laying an important foundation for realizing high-fidelity image transmission in semantic communication.

[0146] In one embodiment, the method performed by the low-latency high-performance image semantic-channel joint coding system is implemented through a low-latency high-performance image semantic-channel joint coding model, and a training process of the low-latency high-performance image semantic-channel joint coding model includes at least a first stage and a second stage.

[0147] Under the condition of freezing the model parameters of the to-be-trained feature encoder and the to-be-trained feature decoder respectively, the to-be-trained image semantic encoder based on the quadtree partitioning module is trained through the first stage and the second stage to obtain a trained image semantic encoder based on the quadtree partitioning module; the to-be-trained image semantic encoder based on the quadtree partitioning module comprises a to-be-trained neural analyzer, a to-be-trained neural synthesizer, a to-be-trained hyper-prior encoder, a to-be-trained hyper-prior decoder and a to-be-trained entropy estimator; analog quantization is used in the first stage, and quantization based on a pass-through estimator is used in the second stage.

[0148] In this embodiment, in order to construct an image semantic-channel joint coding system with low latency, high transmission efficiency and strong robustness, a method for training core models in the coding system is provided. The method is based on a low-latency high-performance image semantic-channel joint coding model, and the training process of the model includes at least two stages: a first stage and a second stage.

[0149] The training process specifically includes the following steps:

[0150] First, in the training process, the network parameters of the feature encoder and the feature decoder remain in a frozen state, that is, the model structure thereof does not participate in gradient update in the training stage. On this basis, only the image semantic encoder constructed based on the quadtree partitioning module is trained.

[0151] The image semantic encoder includes the following to-be-trained modules: a to-be-trained neural analyzer, a to-be-trained neural synthesizer, a to-be-trained hyper-prior encoder, a to-be-trained hyper-prior decoder and a to-be-trained entropy estimator.

[0152] In the first stage of training, in order to realize the differentiable approximation of the quantization process, additive uniform noise method is used to disturb the encoder output to simulate the quantization behavior, while maintaining the differentiability of the training process, which is convenient for back propagation optimization.

[0153] In the second stage of training, the pass-through estimator is switched to perform quantization approximation, that is, the non-differentiable quantization operation is performed in the forward propagation, and the gradient is directly transmitted in the back propagation process, so as to more truly reflect the running behavior of the model in the actual deployment stage.

[0154] The advantage of the two-stage training strategy is that: in the first stage, the additive noise is used to maintain the stability of the initial training of the network, and to avoid gradient oscillation; in the second stage, the pass-through estimator is used to improve the practical effect of the final model, and to enhance its adaptability to the real quantization process. The balance between trainability and deployment adaptability of the model is realized through the stage switching strategy.

[0155] In summary, the training method described in the embodiment can significantly improve the performance of the semantic encoder in tasks such as multi-partition semantic modeling, low-delay mapping, and context feature estimation, providing an efficient and practical training paradigm for the construction of high-performance encoding modules in semantic communication systems.

[0156] In one embodiment, the training process of the low-latency high-performance image semantic-channel joint encoding model further includes a third stage and a fourth stage.

[0157] Unfreeze the model parameters of the feature encoder to be trained and the feature decoder to be trained, and integrate them with the trained image semantic encoder based on the quadtree partition module to obtain a low-latency high-performance image semantic-channel joint encoding model to be trained, wherein the feature encoder to be trained is connected to the neural analyzer trained in the second stage, and the feature decoder to be trained is connected to the neural synthesizer trained in the second stage.

[0158] Through the third stage and the fourth stage, the low-latency high-performance image semantic-channel joint encoding model to be trained is trained to obtain a trained low-latency high-performance image semantic-channel joint encoding model.

[0159] In the third stage, a symbol length factor vector is selected from an initial rate set .

[0160] In the fourth stage, a symbol length factor vector is selected from a target rate set .

[0161] In this embodiment, to further improve the overall encoding quality and channel adaptability of the low-latency high-performance image semantic-channel joint encoding model, based on the foregoing embodiment, the model training process is introduced into the third stage and the fourth stage to realize end-to-end joint optimization.

[0162] Specifically, after completing the first stage and the second stage training and obtaining the trained image semantic encoder based on the quadtree partition module, the original frozen modules including the feature encoder and the feature decoder are unfrozen, that is, their parameters are enabled to participate in the training process.

[0163] Then, the feature encoder to be trained after being unfrozen is connected to the neural analyzer that has been trained, and the feature decoder to be trained is connected to the neural synthesizer that has been trained, thereby constructing a complete low-latency high-performance image semantic-channel joint encoding model to be trained that can be optimized end-to-end.

[0164] The model is trained in the third stage and the fourth stage as follows:

[0165] The third stage: taking the adaptive ability of the model under different channel bandwidth conditions as the goal, a symbol length factor vector is selected from a set of initial rates The training of this stage emphasizes the response characteristics of the model to various basic channel capacity conditions, and builds a robust representation ability for different compression ratios and symbol length distributions.

[0166] The fourth stage: focusing on more refined target bandwidth control, a symbol length factor vector is selected from a set of target rates The training of this stage guides the model to more accurately map semantic content to symbol length distribution under different target code rates to adapt to transmission requirements in different application scenarios.

[0167] Through the joint training of the third and fourth stages, not only the fine tuning of the feature encoder and the feature decoder is realized, but also the code rate control ability and semantic preservation ability of the model under different channel conditions are further improved.

[0168] In summary, through the phased training process of “freezing-training-thawing-joint optimization”, the transmission performance of the semantic communication system in complex and dynamic channel environments is significantly enhanced, ensuring low delay while achieving high-quality image semantic reconstruction and accurate code rate matching.

[0169] In one embodiment, in the first stage and the second stage, the image semantic encoder based on the quadtree partitioning module to be trained is trained using a first loss function value, the first loss function value being determined according to a first difference between a sample original image and a corresponding sample semantic encoded image, the sample semantic encoded image being output by the image semantic encoder based on the quadtree partitioning module to be trained for the sample original image;

[0170] In the third stage and the fourth stage, the low-latency high-performance image semantic-channel joint encoding model to be trained is trained using a second loss function value, the second loss function value being determined according to a second difference between the sample original image and a corresponding sample reconstructed image, and the first difference, the sample reconstructed image being output by the low-latency high-performance image semantic-channel joint encoding model to be trained for the sample original image.

[0171] In this embodiment, in order to further improve the training accuracy and convergence stability of the low-latency high-performance image semantic-channel joint encoding model, a difference-aware multi-stage loss function design is introduced in the training process, a first loss function is used in the first stage and the second stage, and a second loss function is used in the third stage and the fourth stage for training.

[0172] Specifically:

[0173] In the first stage and the second stage, the model only trains an image semantic encoder based on a quadtree partitioning module, and a first loss function value is used in this stage. The first loss function value is calculated based on the following difference:

[0174] The input is a sample original image and a sample semantic encoding image output by the image semantic encoder to be trained. The first loss function is used to measure the reconstruction accuracy of the image semantic features in the encoding process, and a perception loss, a structural similarity loss or an L2 norm is often used as a measurement standard. This design ensures that the model has good semantic extraction and reconstruction ability at the beginning of training, laying a solid foundation for subsequent end-to-end training.

[0175] In the third stage and the fourth stage, the model introduces a trained semantic encoding module, and the feature encoder and the feature decoder are trained end-to-end. The second loss function value is used in this stage, which is more comprehensive and includes the following two items:

[0176] The second difference between the sample original image and the sample reconstructed image generated by the complete model (including channel mapping and decoding);

[0177] The first difference value calculated in the first stage is used as an auxiliary item for training semantic consistency.

[0178] The second loss function combines image reconstruction error and semantic feature fidelity, and has stronger global consistency constraint ability, which not only optimizes the final image quality, but also guarantees the semantic stability and structural fidelity in the transmission process.

[0179] In summary, the embodiment enhances semantic modeling ability in the early stage of training and realizes end-to-end semantic and reconstruction collaborative optimization in the later stage through the design of stage loss function, which helps to significantly improve the generalization ability and robustness of the model under dynamic channels, and is an important optimization strategy for building a low-latency high-performance image semantic communication system.

[0180] In one embodiment, Figure 5 A low-latency high-performance image semantic-channel joint encoding system complete architecture schematic diagram is provided in the present application: the system is composed of a sending end and a receiving end, adopts an end-to-end structured encoding method, and combines a neural network and a visual attention mechanism, thereby significantly reducing the encoding and transmission delay while maintaining high semantic expression accuracy, and improving the reliability and reconstruction performance of the system under real channel conditions.

[0181] On the system structure, first, the image (original image) is processed by the neural analyzer to extract semantic features with deep semantic representation ability. The semantic features are divided into the 0th to 3rd partitions by the quadtree partitioning module according to the channel dimension, and the multi-level context information is constructed based on the index aggregation and connection mode. In the feature encoding process, the sender encodes the 0th partition semantic feature through the feature encoder, and then uses it as the context key-value pair to guide the cross-attention modeling and context-aware encoding of the subsequent partitions, so as to obtain the encoding results of the four partition semantic features. In order to improve the channel compression and mapping ability, the system introduces a hyper-prior encoder and decoder to extract latent variables and generate mean and variance prediction, which helps the entropy estimator to estimate the symbol length factor vector of each partition semantic feature, so as to realize dynamic adjustment of compression rate and symbol complex value mapping. At the receiving end, the system gradually recovers each partition semantic feature through the feature decoder combined with the context vision module, and finally completes the image reconstruction by the neural synthesizer to output the reconstructed image. It is worth noting that this structure avoids the delay bottleneck caused by traditional autoregressive encoding, and achieves a good balance between semantic context expression and module parallelism.

[0182] In terms of model training, the present application adopts a four-stage step-by-step training strategy to improve the stability and multi-objective optimization capability of the model. In the first two stages, the feature encoder and decoder are frozen, only the semantic encoder based on the quadtree partitioning module is trained, the semantic feature learning is gradually completed by using additive noise and a straight-through estimator, and the semantic image reconstruction accuracy is optimized by a first loss function. In the last two stages, all model parameters are unfrozen, the trained semantic encoder and feature encoder and decoder are integrated to form a complete semantic-channel joint model, and different symbol length factor vector sets are combined to adapt to various code rate conditions. Finally, by introducing a second loss function combining image reconstruction error and semantic encoding error, the collaborative optimization from the aspects of visual quality and transmission efficiency is realized.

[0183] Overall, the present application constructs an image semantic-channel joint coding system with clear structure, strong semantic expression and efficient transmission. While maintaining good image reconstruction performance, it realizes high compression rate, low transmission delay and excellent channel adaptability, and is suitable for key application scenarios in low-latency intelligent semantic communication systems, such as autonomous driving, remote medical treatment, immersive XR, etc. This technical solution not only solves the problems of strong coding dependence, poor parallelism and fixed symbol granularity in existing JSCC schemes, but also provides important technical support for the development of semantic communication systems in the direction of multi-scale, self-adaptation and high robustness.

[0184] In the present specification, each embodiment focuses on the difference from other embodiments, and the same or similar parts between embodiments can be referred to each other.

[0185] Those skilled in the art will appreciate that embodiments of the application can be supplied as a method, a device, or a computer program product. Thus, embodiments of the application can take the form of an entirely hardware embodiment, an entirely software embodiment or an embodiment combining software and hardware aspects. Furthermore, embodiments of the application can take the form of a computer program product on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage devices, etc.) embodying computer readable program code.

[0186] Embodiments of the application are described herein with reference to the drawings, in which are shown flowcharts and / or block diagrams of methods, apparatuses (systems) and computer program products according to embodiments of the application. It will be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general purpose computer, special purpose computer, embedded processing device or other programmable data processing terminal devices to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal devices, create means for implementing the functions specified in the flowcharts and / or block diagrams block or blocks. Figure 1 Figure 1

[0187] These computer program instructions can also be stored in a computer- readable memory that can direct a computer or other programmable data processing terminal device to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instructions which implement the function specified in the flowcharts and / or block diagrams block or blocks. Figure 1 Figure 1

[0188] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device to cause a series of operational steps to be performed on the computer or other programmable terminal device to produce a computer implemented process such that the instructions which execute on the computer or other programmable terminal device provide steps for implementing the functions specified in the flowcharts and / or block diagrams block or blocks. Figure 1 Figure 1

[0189] While preferred embodiments of the application have been described, those skilled in the art will appreciate that additional modifications and variations to the preferred embodiments are possible in light of the above teachings. It is, therefore, intended that the appended claims be interpreted as including all such modifications and variations as fall within the true spirit of the application.

[0190] ​​​​​​Finally, it is to be understood that the phraseology or terminology such as "first" and "second" etc. used herein is merely intended to differentiate one entity or operation from another entity or operation, without necessarily requiring or implying any actual such relationship or order between such entities or operations. Moreover, the terms "comprises", "comprising", or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element proceeded by "comprises... a" does not, without more constraints, exclude the existence of additional identical elements in the process, method, article, or apparatus that comprises the element.

[0191] The above provides a low-latency high-performance image semantic-channel joint coding system, and the principles and implementation manners of the present application are described by using specific examples. The above description of the embodiments is only used to help understand the method and core idea of the present application. Meanwhile, for those skilled in the art, the specific implementation manners and application ranges can be changed according to the idea of the present application. In summary, the content of the specification should not be understood as a limitation of the present application.

Claims

1. A low-latency high-performance image semantic-channel joint coding system, characterized in that, The sending end in the semantic communication system extracts semantic features from a source image through a neural analyzer; The sending end divides the semantic features into sending end 0th partition semantic features to sending end 3rd partition semantic features through a quadtree partitioning module; The sending end encodes the sending end 0th partition semantic features through a feature encoder to obtain 0th partition semantic feature encoding results; The sending end encodes the sending end nth partition semantic features through the feature encoder with the 0th to (n-1)th partition semantic feature encoding results as contexts to obtain nth partition semantic feature encoding results, n being an integer between 1 and 3; The sending end maps the 0th to nth partition semantic feature encoding results into complex-valued channel symbols through the feature encoder and sends them to the receiving end in the semantic communication system; The receiving end obtains reconstructed semantic features through a feature decoder using the received complex-valued channel symbols; The receiving end obtains a reconstructed image through a neural synthesizer using the reconstructed semantic features. The receiving end obtains reconstructed semantic features through a feature decoder using the received complex-valued channel symbols, including:

2. The low-latency high-performance image semantic-channel joint coding system of claim 1, wherein, The receiving end determines the receiving end 0th partition semantic features to the receiving end 3rd partition semantic features through the feature decoder using the received complex-valued channel symbols; The receiving end decodes the receiving end 0th partition semantic features through the feature decoder to obtain 0th partition semantic feature decoding results; The receiving end decodes the receiving end nth partition semantic features through the feature decoder with the 0th to (n-1)th partition semantic feature decoding results as contexts to obtain nth partition semantic feature decoding results, n being an integer between 1 and 3; The receiving end obtains the reconstructed semantic features according to the 0th to 3rd partition semantic feature decoding results. Further comprising:

3. The low-latency high-performance image semantic-channel joint coding system of claim 1, wherein, The sending end inputs the semantic features into a hyper-prior encoder to obtain a hyper-prior latent variable and obtains a hyper-prior latent variable quantization result according to the hyper-prior latent variable; The sending end obtains a 0th mean and a 0th variance through a hyper-prior decoder using the hyper-prior latent variable quantization result; The sending end obtains an nth mean and an nth variance through an entropy estimator using the 0th mean and the 0th variance, and an nth context, n being an integer between 1 and 3; The sending end calculates a corresponding symbol length factor vector for each element in the nth partition semantic features according to the nth mean and the nth variance and provides it to the feature encoder and the feature decoder. The sending end divides the semantic features into sending end 0th partition semantic features to sending end 3rd partition semantic features through a quadtree partitioning module, including:

4. The low-latency high-performance image semantic-channel joint coding system of claim 3, wherein, The sending end divides the semantic features into four parts along the channel dimension through a quadtree partitioning module; The sending end adds elements with index n in the four parts to obtain the nth partition semantic features, n being an integer between 0 and 3; Further comprising: The sending end concatenates elements with index 0 in the four parts to obtain a 1st context; ​ The sending end connects the element with index 0 and the element with index 1 in the four parts to obtain the second context; The sending end connects the element with index 0, the element with index 1 and the element with index 2 in the four parts to obtain the third context.

5. The low-latency high-performance image semantic-channel joint coding system of claim 1, wherein, The feature encoder at least comprises a first visual encoding module, a first context visual encoding module, a second context visual encoding module and a third context visual encoding module; the 0th to nth partition semantic feature encoding results are obtained in the following steps: The sending end encodes the 0th to 3rd partition semantic features through the first visual encoding module to obtain the 0th partition semantic feature encoding result and the 1st to 3rd partition first intermediate encoding results; The sending end performs cross-attention calculation on the 0th partition semantic feature encoding result as the key and value, the 1st partition first intermediate encoding result as the query, through the first context visual encoding module to obtain the 1st partition semantic feature encoding result; The sending end performs cross-attention calculation on the 0th to 1st partition semantic feature encoding results as the key and value, the 2nd partition first intermediate encoding result as the query, through the second context visual encoding module to obtain the 2nd partition semantic feature encoding result; The sending end performs cross-attention calculation on the 0th to 2nd partition semantic feature encoding results as the key and value, the 3rd partition first intermediate encoding result as the query, through the second context visual encoding module to obtain the 3rd partition semantic feature encoding result.

6. The low-latency high-performance image semantic-channel joint coding system of claim 2, wherein, The feature decoder at least comprises a zeroth visual decoding module, a first visual decoding module, a first context visual decoding module, a second visual decoding module, a second context visual decoding module, a third visual decoding module and a third context visual decoding module; the 0th to nth partition semantic feature decoding results are obtained in the following steps: The receiving end decodes the 0th partition semantic feature of the receiving end through the zeroth visual decoding module to obtain the 0th partition semantic feature decoding result; The receiving end performs cross-attention calculation on the 0th partition semantic feature decoding result as the key and value, the 1st partition semantic feature of the receiving end as the query, through the first visual decoding module and the first context visual decoding module to obtain the 1st partition semantic feature decoding result; The receiving end performs cross-attention calculation on the 0th to 1st partition semantic feature decoding results as the key and value, the 2nd partition semantic feature of the receiving end as the query, through the second visual decoding module and the second context visual decoding module to obtain the 2nd partition semantic feature decoding result; The receiving end performs cross-attention calculation on the 0th to 2nd partition semantic feature decoding results as the key and value, the 3rd partition semantic feature of the receiving end as the query, through the third visual decoding module and the third context visual decoding module to obtain the 3rd partition semantic feature decoding result.

7. The low-latency high-performance image semantic-channel joint coding system of claim 1, wherein, The neural analyzer comprises 1st to Nth down-sampling modules and at least two attention modules, The neural synthesizer comprises: 1st to Nth up-sampling modules, at least two attention modules, and a feature refinement module; the feature refinement module is connected with the 1st up-sampling module; The 1st to Nth up-sampling modules correspond to the 1st to Nth down-sampling modules one by one.

8. The low-latency high-performance image semantic-channel joint coding system of claim 1, wherein, The method performed by the low-latency high-performance image semantic-channel joint coding system is realized through a low-latency high-performance image semantic-channel joint coding model, and a training process of the low-latency high-performance image semantic-channel joint coding model comprises at least a first stage and a second stage: Under the condition of freezing the model parameters of the to-be-trained feature encoder and the to-be-trained feature decoder, the to-be-trained image semantic encoder based on the quadtree partitioning module is trained through the first stage and the second stage, so as to obtain a trained image semantic encoder based on the quadtree partitioning module. The to-be-trained image semantic encoder based on the quadtree partitioning module comprises: a to-be-trained neural analyzer, a to-be-trained neural synthesizer, a to-be-trained hyper-prior encoder, a to-be-trained hyper-prior decoder, and a to-be-trained entropy estimator. In the first stage, analog quantization is used, and in the second stage, quantization based on a pass-through estimator is used.

9. The low-latency high-performance image semantic-channel joint coding system according to claim 8, wherein, The training process of the low-latency high-performance image semantic-channel joint coding model further comprises a third stage and a fourth stage. The model parameters of the to-be-trained feature encoder and the to-be-trained feature decoder are unfrozen, and are integrated with the trained image semantic encoder based on the quadtree partitioning module, so as to obtain a to-be-trained low-latency high-performance image semantic-channel joint coding model, wherein the to-be-trained feature encoder is connected with the neural analyzer trained in the second stage, and the to-be-trained feature decoder is connected with the neural synthesizer trained in the second stage. The to-be-trained low-latency high-performance image semantic-channel joint coding model is trained through the third stage and the fourth stage, so as to obtain a trained low-latency high-performance image semantic-channel joint coding model. In the third stage, a symbol length factor vector is selected from an initial set of rate sets ; In the fourth stage, a symbol length factor vector is selected from a set of target rate s.

10. The low-latency high-performance image semantic-channel joint coding system of claim 9, wherein, In the first stage and the second stage, the to-be-trained image semantic encoder based on the quadtree partitioning module is trained using a first loss function value, which is determined according to a first difference between a sample original image and a corresponding sample semantic encoded image, wherein the sample semantic encoded image is output by the to-be-trained image semantic encoder based on the quadtree partitioning module for the sample original image. In the third stage and the fourth stage, the to-be-trained low-latency high-performance image semantic-channel joint coding model is trained using a second loss function value, which is determined according to a second difference between the sample original image and a corresponding sample reconstructed image, and the first difference, wherein the sample reconstructed image is output by the to-be-trained low-latency high-performance image semantic-channel joint coding model for the sample original image.

Citation Information

Patent Citations

  • Context modeling semantic communication code transmission and receiving method and related equipment

    CN116935840A

  • Semantic communication method and system for high-resolution image

    CN119383363A