Synonymous variation image sending method, synonymous variation image receiving method and related model training method

By employing a synonymous variational image transmission method, which utilizes semantic equivalence constraints for synonymous parsing transformation and adaptive channel state selection, the problem of weak channel adaptability in traditional visual information transmission is solved, achieving efficient utilization of channel resources and adaptive transmission.

CN121985125APending Publication Date: 2026-05-05BEIJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING UNIV OF POSTS & TELECOMM
Filing Date
2026-02-06
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Traditional visual information transmission methods have weak channel adaptability at the syntax level and low channel resource utilization, making it difficult to adapt to the needs of intelligent transmission.

Method used

The synonym variational image transmission method performs a synonym parsing transformation with semantic equivalence constraints on the original image to generate a continuous latent feature map, determines the channel overhead sequence, and selects the target synonym level for transmission based on the channel state parameters, thus discarding redundant feature transmission.

Benefits of technology

It achieves efficient utilization of channel resources, adapts to the transmission of semantic information of different granularities, and improves the channel adaptive matching capability and resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121985125A_ABST
    Figure CN121985125A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a synonymous variational image sending method, a synonymous variational image receiving method and a related model training method. The synonymous variational image sending method comprises the following steps: carrying out synonymous analysis transformation processing on an original image on the basis of meeting semantic equivalence constraint to obtain a continuous hidden feature map; determining probability distribution of each element of the continuous hidden feature map; calculating channel overhead corresponding to each feature vector in the continuous hidden feature map according to the probability distribution, and generating a channel overhead sequence; determining a target synonymous level according to the current channel state parameter of the channel; determining a synonymous signal sequence corresponding to a synonymous representation part of the target synonymous hierarchy according to the channel overhead sequence; and transmitting the synonymous signal sequence to the channel as a target transmission signal sequence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of communication technology, and in particular to a method for transmitting and receiving synonymous variational images and a related model training method. Background Technology

[0002] Currently, among the various services offered by the mobile internet, visual information transmission, primarily in the form of images and videos, accounts for the vast majority of network bandwidth. Traditional visual information transmission relies on a separable coding framework supported by Shannon's source-channel coding theorem. This framework separates image / video compression and data transmission into two independent information processing processes. It designs distortion-limited image / video compression schemes based on syntactic symbol accuracy and transmission processes based on bit-level reliability. However, with the advancement of wireless communication technology towards intelligence, this syntactic-level visual information transmission method has revealed numerous problems, including poor transmission robustness, weak channel adaptability, a mismatch between symbol accuracy and perceived quality, and difficulty in efficiently meeting the demands of intelligent transmission.

[0003] Based on semantic information theory, the same perceptual meaning / semantic content can be expressed through different grammatical symbols / data representations, and these different representations can achieve semantic equivalence at the perceptual level. Simply put, different representations may have the same meaning. This fundamental difference breaks the traditional information processing logic of unique correspondence between grammatical level symbols and precise restoration. Against this backdrop, based on the theoretically superior joint source-channel coding framework, a semantic-level communication paradigm has been developed. This paradigm uses a deep neural network to build an end-to-end semantic communication model, achieving joint source-channel coding. Through an attention mechanism, channel states are introduced into the neural network processing flow, and implicit channel state adaptation is achieved through training by traversing the channel states. However, this approach has poor adaptability to semantic granularity, leading to a mismatch between channel resources and semantic requirements, low channel adaptation accuracy, and poor resource utilization. Summary of the Invention

[0004] This disclosure proposes a synonymous variational image transmission and reception method and a related model training method to solve or partially solve the above-mentioned problems.

[0005] This disclosure provides a synonym variational image transmission method, comprising: performing a synonym parsing transformation on an original image based on satisfying semantic equivalence constraints to obtain a continuous latent feature map; determining the semantic probability distribution of each element of the continuous latent feature map; calculating the channel overhead corresponding to each feature vector in the continuous latent feature map according to the semantic probability distribution, and generating a channel overhead sequence; determining a target synonym level according to the current channel state parameters, wherein multiple synonym levels are pre-set with serial numbers, the serial number of each synonym level corresponds to a preset synonym level partitioning rule, the synonym level partitioning rule is used to characterize the partitioning method of the synonym representation part and the detail representation part of the latent feature map in the current synonym level, and the serial numbers of the multiple synonym levels are sorted according to the semantic importance and / or semantic granularity of the latent feature map; determining a synonym signal sequence corresponding to the synonym representation part of the target synonym level according to the channel overhead sequence; and transmitting the synonym signal sequence as a target transmission signal sequence to the channel.

[0006] A second aspect of this disclosure provides a synonym variational image receiving method, comprising: receiving a signal sequence from a channel; determining a target synonym level corresponding to the signal sequence, wherein multiple synonym levels identified by serial numbers are pre-set, the serial number of each synonym level corresponds to a preset synonym level partitioning rule, the synonym level partitioning rule is used to characterize the partitioning method of the synonym representation part and detail representation part of the latent feature map in the current synonym level, the serial numbers of the multiple synonym levels are sorted according to the semantic importance and / or semantic granularity of the latent feature map; determining a continuous latent feature map corresponding to the target synonym level according to the signal sequence; and performing a synonym synthesis transformation on the continuous latent feature map that satisfies semantic equivalence constraints to obtain a reconstructed image corresponding to the signal sequence.

[0007] This disclosure provides a training method for a synonymous variational image transmission model, comprising: acquiring image samples; performing synonymous parsing transformation on the image samples based on semantic equivalence constraints to obtain a continuous latent feature map; estimating the probability distribution of each element in the continuous latent feature map based on a variational entropy model to obtain the probability distribution features of each element in the continuous latent feature map; and dividing the continuous latent feature map into multiple sets of synonymous representation parts and detail representation parts corresponding to each synonymous level based on preset sequence numbers of multiple synonymous levels, wherein the synonymous levels are ordered by the sequence number of ... The identifiers for each synonym level correspond to preset synonym level partitioning rules. These rules characterize the partitioning of the synonym representation and detail representation parts of the latent feature map within the current synonym level. The sequence numbers of the multiple synonym levels are ordered based on the semantic importance and / or semantic granularity of the latent feature map. Based on the probability distribution characteristics of each element in the continuous latent feature map, the coding rate of the synonym representation part at each synonym level is estimated. The channel overhead corresponding to each feature vector in the continuous latent feature map is determined, and a channel overhead sequence is generated. The coding rate is then used to determine the channel overhead. The process involves estimating the SNR threshold corresponding to each synonym level using the reference SNR and the channel overhead sequence; using the continuous latent feature map as input to the synonym encoder in the synonym variational image transmission model to determine the synonym signal sequence corresponding to the synonym representation part of the target synonym level selected in this training; using the target SNR threshold corresponding to the target synonym level as the SNR of the channel in this training, and transmitting the synonym signal sequence to the channel; and inputting the synonym signal sequence received from the channel into the synonym decoder in the synonym variational image transmission model to enable the synonym... The decoder determines the target synonym level and the corresponding continuous latent feature map based on the target signal-to-noise ratio threshold; performs a synonym synthesis transformation that satisfies semantic equivalence constraints on the corresponding continuous latent feature map to obtain a reconstructed image; calculates a loss value based on the reconstructed image and the image sample; updates and trains the synonym variational image transmission model according to the loss value to obtain the trained synonym variational image transmission model, wherein the loss value includes a total loss obtained by weighted summation of loss values ​​calculated based on at least two different loss functions.

[0008] This fourth aspect of the disclosure provides a computer device including one or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and executed by the one or more processors, and the one or more programs include instructions for performing the methods of the first aspect, the second aspect, and the third aspect.

[0009] The synonym variational image transmission method provided in this disclosure converts the original image into a continuous latent feature map by performing a synonym parsing transformation based on semantic equivalence constraints. This transformation shifts the original image from pixel-level syntactic representation to semantic-level feature representation, locking in the core semantic information to be transmitted and abandoning the excessive pursuit of accuracy in detailed information. Furthermore, by determining the semantic uncertainty of each element in the continuous latent feature map, the stability and importance of the element's semantic expression can be measured. Based on the uncertainty index of each element, the channel overhead corresponding to each feature vector is calculated, establishing a correlation mapping between the importance of semantic features and channel transmission resources, enabling on-demand allocation of channel overhead. Subsequently, the target synonym level is determined according to the current channel state parameters, achieving the goal of adaptive matching transmission strategies based on channel state. This allows for explicit channel state adaptation in the signal space, enabling the synonym variational image transmission method of this disclosure to adapt to the transmission of semantic information at different granularities. Simultaneously, only the synonym signal sequence corresponding to the target synonym level is transmitted on the channel, eliminating the transmission of redundant features and significantly improving the utilization efficiency of channel resources. Attached Figure Description

[0010] To more clearly illustrate the technical solutions in this disclosure or related technologies, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0011] Figure 1 A schematic flowchart of an exemplary synonymous variational image transmission method provided in an embodiment of this disclosure is shown.

[0012] Figure 2 A schematic flowchart of an exemplary method for transmitting and receiving synonymous variational images based on incremental signals, provided in an embodiment of this disclosure, is shown.

[0013] Figure 3 A schematic flowchart of an exemplary synonymous variational image receiving method provided in an embodiment of this disclosure is shown.

[0014] Figure 4 A schematic diagram of a general coding framework for an exemplary synonymous variational image transmission and reception method provided in this disclosure is shown.

[0015] Figure 5 A flowchart illustrating an exemplary training method for a synonymous variational image transfer model provided in an embodiment of this disclosure is shown. Figures 6(a)-6(f) show schematic diagrams of reconstructed image quality curves under different datasets and different wireless channel models corresponding to two exemplary embodiments provided in this disclosure. Figure 7 A schematic diagram illustrating the reconstructed image effects under different synonym levels provided by embodiments of this disclosure is shown.

[0016] Figure 8 A schematic diagram of the hardware structure of an exemplary computer device provided in an embodiment of this disclosure is shown. Detailed Implementation

[0017] To make the objectives, technical solutions, and advantages of this disclosure clearer, the following detailed description is provided in conjunction with specific embodiments and the accompanying drawings.

[0018] It should be noted that, unless otherwise defined, the technical or scientific terms used in the embodiments of this disclosure should have the ordinary meaning understood by one of ordinary skill in the art to which this disclosure pertains. The terms "first," "second," and similar terms used in the embodiments of this disclosure do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.

[0019] It is understood that before using the technical solutions of the various embodiments in this disclosure, users will be informed of the type, scope of use, and usage scenarios of the personal information involved in an appropriate manner, and user authorization will be obtained.

[0020] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose, based on the prompt message, whether to provide personal information to the software or hardware such as electronic devices, applications, servers, or storage media performing the operations of this disclosed technical solution.

[0021] As an optional but not limited implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0022] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0023] As an example, the image transmitting device preprocesses the original image and extracts semantic features. It then incorporates channel states into the end-to-end semantic communication model's processing flow through an attention mechanism. A semantically encoded signal is generated via end-to-end joint source-channel coding. This signal is trained by traversing the channel states to achieve implicit channel adaptation. Finally, the semantically encoded signal is converted into a physical layer signal and transmitted to the channel. The signal is then sent to the image receiving device, which preprocesses the signal and reconstructs the image by combining matching decoding with implicit adaptation capabilities to restore semantic features. However, this implicit channel adaptation method has poor adaptability to different semantic granularities, leading to a mismatch between channel resources and semantic requirements, low channel adaptation accuracy, and poor resource utilization.

[0024] In view of this, embodiments of the present disclosure provide a method for transmitting synonymous variational images to solve or partially solve the above-mentioned problems.

[0025] Figure 1 A schematic flowchart of an exemplary synonymous variational image transmission method 200 provided in an embodiment of this disclosure is shown. This method can be used for transmitting synonymous variational images.

[0026] In step 102, the original image is subjected to synonymous parsing transformation based on semantic equivalence constraints to obtain a continuous latent feature map.

[0027] Synonymous analytic transformation, as an implementable method, satisfies semantic equivalence constraints. It can be used to extract the original image from the data space. As input, continuous latent feature maps in the latent space are obtained. Assuming the original image Characterized by the number of channels Gao Wei Width A three-dimensional real matrix, i.e. Then the continuous hidden feature map The number of channels can be Gao Wei Width A three-dimensional real matrix, i.e. In this context, the dimension of the latent space is lower than the dimension of the data space. (Synonym parse transformation) For example, it can be implemented based on a nonlinear transformation process, which can transform the original image... Parse into the required continuous hidden feature map The ability to perform synonym analysis. It can also be used to directly distinguish, or to distinguish consecutive latent feature maps by cascading other modules. Synonym representation part and detailed characterization section Among them, the synonym representation part As an ideal set of synonyms corresponding to the original image Shared latent space features of all image samples within the quantized feature map The information is separated from the original image, while the remaining information serves as the detail representation of the original image. In one or more embodiments of this disclosure, synonym resolution transformation It can be built based on deep neural network structures, including but not limited to convolutional neural networks, recurrent neural networks, and Transformers, to achieve nonlinear transformation processing. It can contain trainable parameters, and these parameters can be optimized through model training, enabling it to parse the original image into the desired continuous latent feature map. For example, a synonym parsing transformation model can be constructed using a deep learning network integrating a synonym constraint loss function. Taking the original image pixel tensor in the data space as input data, the model undergoes multi-level processing through feature extraction layers, semantic clustering layers, and dimensionality reduction mapping layers to generate a latent feature map with continuous numerical distribution characteristics in the latent space—that is, a continuous latent feature map. This achieves the transformation of the original image from pixel-level syntactic representation to semantic-level feature representation. Through synonym parsing transformation, the extracted continuous latent feature map satisfies semantic equivalence constraints, meaning the continuous latent feature map completely preserves the core semantic information of the original image. This ensures that regardless of subsequent channel adaptation, the transmission is based on features semantically equivalent to the original image, preventing the loss of core semantics during transmission and providing a semantically faithful foundation for the transmission of the original image.

[0028] In step 104, the probability distribution of each element of the continuous hidden feature map is determined.

[0029] In step 106, the channel overhead corresponding to each feature vector in the continuous hidden feature map is calculated according to the probability distribution to generate a channel overhead sequence.

[0030] In step 108, the target synonym level is determined according to the current channel state parameters. Multiple synonym levels are pre-set and identified by serial numbers. The serial number of each synonym level corresponds to a preset synonym level division rule. The synonym level division rule is used to characterize the division method of the synonym representation part and the detail representation part of the latent feature map in the current synonym level. The serial numbers of the multiple synonym levels are sorted according to the semantic importance and / or semantic granularity of the latent feature map.

[0031] The synonym hierarchy considers the existence of multiple synonym levels. As one possible approach, based on the difference in importance of information at each synonym level, they can be labeled from highest to lowest importance as follows: The sequence number is used as the index for each synonym hierarchy. In some embodiments, highly important semantic information is also coarse-grained semantic information, and conversely, less important semantic information is also fine-grained semantic information. Therefore, the order of the synonym hierarchy indexes can also represent an order from coarse-grained synonym hierarchy to fine-grained synonym hierarchy. In other embodiments, the order of importance or granularity of synonym hierarchy can be the reverse of the above order, that is, they can be labeled from low to high importance as follows: This also indicates the order from fine-grained synonym hierarchy to coarse-grained synonym hierarchy. For example, number 1 can correspond to high-order semantics such as "object category", number 2 can correspond to mid-order features such as "edge texture", and number 3 can correspond to low-order details such as "local pixel difference".

[0032] In one or more embodiments of this disclosure, a linear or non-linear relationship between channel state parameters and the synonym level index can be preset, that is, a linear expression or a non-linear expression between channel state parameters and the synonym level index can be preset as an algorithm for identifying the target synonym level index. Based on this, the target synonym level index... Recognition algorithms can be based on linear or nonlinear methods, utilizing the current signal-to-noise ratio of the channel. (An example of channel state parameters) Determine the synonym level .

[0033] In some embodiments, the encoding rate of the synonym representation portion corresponding to each synonym level is... It can be used as a non-essential input condition for this recognition algorithm to assist in the generation of synonym hierarchical numbers. The introduction of this unnecessary input condition allows for identification at the same signal-to-noise ratio. Below are different original images. Determine different but more accurate synonym hierarchy numbers .

[0034] In step 110, a synonym signal sequence corresponding to the synonym representation portion of the target synonym level is determined based on the channel overhead sequence.

[0035] In one embodiment, determining the synonym signal sequence corresponding to the synonym representation portion of the target synonym level based on the channel overhead sequence may specifically include: The continuous latent feature map is used as input to a synonym encoder in a pre-trained synonym variational image transmission model. The synonym encoder maps the continuous latent feature map from the latent space to a signal sequence in the signal space based on an unequal-length nonlinear transformation, resulting in multiple sub-signal sequences. These sub-signal sequences are concatenated to form a signal sequence, where the channel overhead of each sub-signal sequence is determined based on the channel overhead sequence. Synonymous signal sequences corresponding to the synonym representation portion of the target synonym level are separated from the signal sequence. For example, the synonym encoder can utilize an unequal-length nonlinear transformation to map the continuous feature map in the latent space to a signal sequence in the signal space. This signal sequence can be represented as each latent space feature vector. The sub-signal sequence obtained after unequal-length nonlinear transformation The vector formed by cascading, i.e. Each feature vector By using unequal-length nonlinear transformation, sub-signal sequences of different lengths are obtained. The channel overhead it occupies is determined by the channel overhead sequence. The corresponding channel overhead Confirmed. Based on the target synonym hierarchy. In the signal sequence Separating synonymous signal sequences This synonymous signal sequence is related to the first one in the data space. Synonym hierarchy ideal synonym set Correspondingly, that is, the first Ideal Synonyms Any sample within They all have the same or similar synonymous signal sequences. It can be based on linear or nonlinear signal processing methods, according to the signal sequence. Synonyms Obtain synonym signal sequence .

[0036] In another embodiment, determining the synonym signal sequence corresponding to the synonym representation portion of the target synonym level based on the channel overhead sequence may specifically include: The continuous hidden feature map is subjected to an unequal-length nonlinear transformation to obtain the intermediate state feature representation of the continuous hidden feature map; Based on the synonym hierarchy partitioning rule, the intermediate state feature representation is mapped to multiple sets of incremental signals corresponding to the synonym hierarchy. Each set of incremental signals is the estimated difference between the synonym signal of the current synonym hierarchy and the synonym signal of the previous synonym hierarchy. For the same latent feature map, the multiple incremental signals corresponding to each synonym hierarchy have the same channel overhead. Based on the order of the sequence numbers of the multiple synonym levels, the incremental signal corresponding to the target synonym level and the incremental signals corresponding to all synonym levels before the target synonym level are superimposed to obtain the synonym signal sequence.

[0037] Figure 2 A schematic flowchart illustrating an exemplary method for transmitting and receiving synonymous variational images based on incremental signals, provided in an embodiment of this disclosure, is shown. Figure 2 As shown in the synonymous signal transmission structure section, continuous hidden feature map It is converted into an intermediate state through non-linear conversion at the encoding end. (This is an example of the intermediate state feature representation described above). This intermediate state... It can be a three-dimensional real matrix The number of channels c' of this three-dimensional real matrix does not need to be related to the continuous latent feature map. The number of channels c remains constant, and the height h and width w are continuous latent feature maps. Maintain consistency. Keep intermediate states consistent. Through multiple incremental signal formation processes, it is mapped into multiple incremental signals. The number of incremental signal formation processes and the number of incremental signals obtained are equal to the number of synonym levels, which is... Each incremental signal can be represented as a latent space feature vector. The corresponding number Synonymous hierarchical incremental sub-signal sequence The vector formed by cascading, i.e. Among them, for the same Each synonym level The corresponding multiple incremental signals With the same channel overhead . No. indivual( Incremental signals are synonymous signals from adjacent synonymous levels. and The estimation of the difference, i.e. As for the first incremental signal, by assuming Vector, incremental signal It should be noted that, due to the synonymous signal sequences at each level... These signals cannot be obtained directly; they need to be obtained through neural network training. Therefore, this processing method, which determines synonymous signals based on incremental signals, is helpful in determining the synonym hierarchy. Synonymous signal sequences are determined by signal accumulation. Specifically, when the synonym hierarchy number... At that time, the result of signal accumulation That is, the signal sequence in the above signal space. .

[0038] In some embodiments, the continuous latent feature map estimated by the variational entropy model The corresponding mean matrix Sum of standard deviation matrix It can also be introduced into the nonlinear spatial transformation process at the transmitting end to enrich the distribution characteristics.

[0039] Based on the order of the synonym levels, the incremental signals corresponding to the target synonym level and the incremental signals corresponding to all synonym levels preceding the target synonym level are superimposed to obtain the synonym signal sequence.

[0040] Separating synonymous signal sequences based on incremental signals can effectively improve the synonym resolution capability within a limited model size.

[0041] In step 112, the synonymous signal sequence is used as the target transmission signal sequence and transmitted to the channel.

[0042] The synonym variational image transmission method provided in this disclosure converts the original image into a continuous latent feature map by performing a synonym parsing transformation based on semantic equivalence constraints. This transformation shifts the original image from pixel-level syntactic representation to semantic-level feature representation, locking in the core semantic information to be transmitted and abandoning the excessive pursuit of accuracy in detailed information. Furthermore, by determining the semantic uncertainty of each element in the continuous latent feature map, the stability and importance of the element's semantic expression can be measured. Based on the uncertainty index of each element, the channel overhead corresponding to each feature vector is calculated, establishing a correlation mapping between the importance of semantic features and channel transmission resources, enabling on-demand allocation of channel overhead. Subsequently, the target synonym level is determined according to the current channel state parameters, achieving the goal of adaptive matching transmission strategies based on channel state. This allows for explicit channel state adaptation in the signal space, enabling the synonym variational image transmission method of this disclosure to adapt to the transmission of semantic information at different granularities. Simultaneously, only the synonym signal sequence corresponding to the target synonym level is transmitted on the channel, eliminating the transmission of redundant features and significantly improving the utilization efficiency of channel resources.

[0043] In one or more embodiments of this disclosure, the channel state parameter may include the signal-to-noise ratio. Based on this, the synonymous variational image transmission method of this disclosure may further include: After determining the probability distribution of each element in the continuous latent feature map, the continuous latent feature map is divided into multiple sets of synonym representation parts and detail representation parts corresponding to each synonym level based on the indices of the multiple synonym levels. The synonym representation parts under each synonym level are the shared spatial features of all image samples within the synonym set corresponding to the original image. The average distortion between any image sample in the synonym set and the original image satisfies the distortion metric corresponding to each synonym level. The distortion metric corresponding to each synonym level increases or decreases sequentially according to the indices of each synonym level. The perceptual loss between any image sample in the synonym set and the original image satisfies the perceptual loss metric corresponding to each synonym level. The perceptual loss metric corresponding to each synonym level increases or decreases sequentially according to the indices of each synonym level. Since the synonym level indices correspond to preset synonym level partitioning rules, the continuous latent feature map can be divided into multiple sets of synonym representation parts and detail representation parts corresponding to each synonym level, based on the synonym level partitioning rules corresponding to the indices of each synonym level. In some embodiments, the synonym level partitioning rules can divide the synonym levels uniformly or non-uniformly along the channel dimension. For example, when multiple synonym level indices... When sorting by importance from highest to lowest, select the synonym level. The dimension is According to the synonym hierarchy division rules, from the first channel dimension to the second... The three-dimensional matrix formed by the nth channel dimension can be represented as the nth channel dimension. Synonymous representation section And from the first The channel dimension to the first The three-dimensional matrix formed by the nth channel dimension, i.e., the remaining representation part, can be represented as the nth channel dimension. Level of detail representation Furthermore, after fully training the synonymous variational image transmission model used to perform the synonymous variational image transmission method of the present disclosure embodiments, this partitioning of the latent feature map in the channel dimension can give the determined synonymous representation parts corresponding importance or semantic information granularity.

[0044] For example, under a specific synonym hierarchy, the first Synonym representation part corresponding to the synonym hierarchy and detailed characterization section It can be used as a synonym representation part and detailed characterization section One specific manifestation of this is the synonym representation. As the first corresponding to the original image Synonym hierarchy ideal synonym set The shared latent space features of all image samples within the range, while the detail representation part This is represented as a continuous hidden feature map. relative to the synonym representation part The remaining representation portion, as the first part of the original image Detailed representation of synonym hierarchy Among them, the ideal synonym set Any image sample within With the original image The average distortion satisfies the following relationship: ; In the above formula, For distortion measurement functions, in some embodiments, MSE (Mean Squared Error) or MS-SSIM (Multi-Scale Structural Similarity Index Measure) can be selected.

[0045] Ideal Synonyms Any image sample within With the original image The perceived loss satisfies the following relationship: ; In the above formula, For the perceptual loss metric function, in some embodiments, divergence calculation methods such as KL (Kullback-Leibler Divergence), Wasserstein divergence, and JS (Jensen-Shannon Divergence) can be selected, or perceptual metrics that are close to perceptual quality, such as LPIPS (Learned Perceptual Image Patch Similarity) and DISTS (Deep Image Structure and Texture Similarity), can be used. When multiple synonymous hierarchical indices are used... When sorted by importance from highest to lowest, the above relationship needs to satisfy the following equation: and style .

[0046] Based on the probability distribution of each element of the continuous hidden feature map, the coding rate of the synonym representation part corresponding to each synonym level is estimated. Determining the target synonym level based on the current channel state parameters can specifically include: The sequence number of the target synonym level or a soft value sequence used to characterize the sequence number is determined based on the signal-to-noise ratio and the coding rate using a preset linear or nonlinear relationship.

[0047] For example, synonym hierarchy It can be used directly as a condition to guide the hard superposition process of synonymous level signals. It can also be based on a soft-value sequence. (An example of the soft-value sequence used above to characterize the synonym hierarchy) Indirectly representing the synonym hierarchy This guides the soft superposition process of synonymous level signals. In this case, It needs to be based on a synonym hierarchy Directly determined binary vector As the objective, optimization is performed using cross-entropy as the constraint loss, which can be expressed as follows.

[0048] ; When multiple synonym hierarchical numbers When sorting by importance from highest to lowest, When multiple synonym hierarchical numbers When sorting by importance from low to high, .

[0049] Finally, by using the synonym hierarchy signal superposition process, the synonym signal is output. Among them, when synonym levels When used directly as a condition, the synonymous hierarchical signal hard superposition process is performed according to the following formula: ; Specifically, when the synonym hierarchy number hour, .

[0050] And when soft value sequence When used as a condition, the soft superposition process of synonymous level signals is performed according to the following formula: ; when hour, .

[0051] Based on the above process of hard or soft overlay, and synonymous hierarchical sequence numbers The corresponding synonym signal sequence It will be transmitted as a synonym signal to the wireless channel and then transmitted to the decoding end.

[0052] In one or more embodiments of this disclosure, the channel state parameters may include the signal-to-noise ratio (SNR), and the target synonym level may be determined based on the current channel state parameters. Specifically, this may include: Based on a preset linear or nonlinear relationship, determine the sequence number of the target synonymous level or a soft value sequence used to characterize the sequence number according to the signal-to-noise ratio, and determine the weight coefficient sequence corresponding to the features of the target synonymous level. Based on the order of the sequence numbers of the multiple synonym levels, the incremental signals corresponding to the target synonym level and the incremental signals of all synonym levels preceding the target synonym level are superimposed according to the sequence number and the weight coefficient sequence to obtain the synonym signal sequence. Alternatively, the incremental signals corresponding to the target synonym level and the incremental signals of all synonym levels preceding the target synonym level are superimposed according to the soft value sequence and the weight coefficient sequence to obtain the synonym signal sequence.

[0053] The method may further include: after determining the sequence number of the target synonym level based on the signal-to-noise ratio according to a preset linear relationship or a preset nonlinear relationship, sending the sequence number to the receiving device corresponding to the original image based on the side information transmission link.

[0054] For example, the neural network size for multiple incremental signal formation processes is fixed for different source image samples. To provide adaptive adjustments for the source, specific feature information for each level, such as signal power, average signal amplitude, average coding rate, and Gaussian equivalent signal-to-noise ratio threshold, can be introduced into the synonym hierarchy recognition process. This allows the synonym hierarchy recognition process to additionally output a learnable sequence of weight coefficients. (This is an example of the weight coefficient sequence mentioned above). Under this condition, the hard superposition process of synonymous hierarchical signals can be represented as follows: The soft superposition process of synonymous hierarchical signals can be expressed as .

[0055] In one or more embodiments of this disclosure, determining the probability distribution of each element of the continuous latent feature map may specifically include: The probability distribution of each element in the continuous latent feature map is estimated based on the variational entropy model to obtain the Gaussian distribution characteristics of each element. For example, the variational entropy model The input is the continuous latent feature map in the example above. The output is a continuous latent feature map. Each element The estimated Gaussian distribution characteristics, i.e., the mean and standard deviation The mean matrix formed Sum of standard deviation matrix , where the mean matrix Sum of standard deviation matrix With continuous hidden feature maps The dimensions remain consistent, that is, the number of channels is Gao Wei Width .

[0056] Based on the probability distribution of each element in the continuous latent feature map, the encoding rate of the synonym representation part at each synonym level is estimated, which may specifically include: Based on the mean and standard deviation of the Gaussian distribution characteristics, calculate the sub-information of each element in the synonym representation part; The cumulative probability of the sub-information of each element in the synonym representation part within its preset neighborhood is calculated to obtain the encoding rate of the synonym representation part.

[0057] For example, synonym hierarchy The coding rate of the corresponding synonym representation part Based on the synonym representation part Each element Sub-information estimate Obtained by summation. That is, based on the mean. with standard deviation A Gaussian distribution characterized by elements is calculated. Nearby area The cumulative probability (as an example of the aforementioned preset neighbor interval) is used as an element. Probability estimates.

[0058] In some embodiments, when the synonym hierarchy When ranking from highest to lowest importance, the synonym representation part coding rate It can be expressed as follows:

[0059] In other embodiments, when the synonym hierarchy When ranking from low to high importance, the synonym representation part coding rate It can be expressed as follows:

[0060] In the embodiments of this disclosure, the coding rate of the synonym representation part corresponding to each synonym level can be used as a condition for determining the target synonym level. In this way, under the same signal-to-noise ratio, different but more accurate synonym level numbers can be determined for different original images, thereby improving the recognition accuracy of the synonym level numbers.

[0061] In one or more embodiments of this disclosure, determining the channel overhead corresponding to each feature vector in the continuous latent feature map may specifically include: Calculate the single-symbol signal capacity based on the current signal-to-noise ratio of the channel; The channel overhead corresponding to each feature vector is calculated based on the sub-information of each feature vector and the single-symbol channel capacity.

[0062] In one example, the uncertainty index is based on a Gaussian distribution with mean and standard deviation. First, the sub-information value of each element can be estimated based on the mean and standard deviation. Then, the channel overhead corresponding to each feature vector can be calculated based on the sub-information value. Finally, a channel overhead sequence can be constructed based on the channel overhead corresponding to each feature vector.

[0063] element Sub-information estimate It can be expressed as follows: ; That is, based on the mean with standard deviation A Gaussian distribution with uniform distribution intervals is introduced. Calculate by element Nearby area The cumulative probability as an element The probability estimate of the element The estimated sub-information content. In some embodiments, when the synonym hierarchy When ranking from highest to lowest importance, the synonym representation part coding rate It can be represented as In other embodiments, when the synonym hierarchy... When ranking from low to high importance, the synonym representation part coding rate It can be represented as .

[0064] In one example, let each feature vector in a continuous latent feature map... Corresponding wireless channel overhead (As an example of the channel overhead described above) and constitute a channel overhead sequence. , where the feature vector It is a continuous hidden feature map The first in The information content of a feature vector can be expressed by the following formula: ; Select a reference signal-to-noise ratio Its single-symbol channel capacity can be expressed by the following formula: ; Based on the information content of the feature vector With single-symbol channel capacity Determine the overhead of the feature vector in the wireless channel. .in, For discrete selection of functions, based on The calculation results are within a certain set of integer symbols for channel overhead. In this process, an appropriate integer value is selected according to certain principles (such as nearest neighbor, rounding down, and rounding up) as the wireless channel overhead of the feature vector. For a transmitted image sample, the wireless channel overhead of all feature vectors in the latent space in the signal space constitutes the channel overhead sequence. ,in, denoted as the number of feature vectors in the latent space.

[0065] In one or more embodiments of this disclosure, the probability distribution of each element in the continuous latent feature map is estimated based on a variational entropy model to obtain the Gaussian distribution feature of each element as the probability distribution of each element. Specifically, this may include: After quantization and entropy encoding of the intermediate matrix of the variational entropy model at the bottleneck layer, a first quantized intermediate encoding sequence is obtained. The first quantized intermediate encoding sequence is sent as side information to the receiving device corresponding to the original image through the side information transmission link. The receiving device then estimates the mean matrix and standard deviation matrix of the Gaussian distribution feature based on the first quantized intermediate encoding sequence through entropy decoding and the portion of the variational entropy model after the bottleneck layer, thereby obtaining the probability distribution. For example, the variational entropy model Intermediate matrix in the bottleneck layer After quantization and entropy encoding, the data can be sent as side information to the decoder via the side information transmission link. At this point, the decoder can obtain the quantization intermediate matrix through entropy decoding. In this case, the variational entropy model The part before the bottleneck layer can be called the variational entropy analytical transformation model. The part after the bottleneck layer is called the variational entropy comprehensive transformation model. At this point, the variational entropy comprehensive transformation model It can be applied simultaneously to the decoding end, for use with an intermediate matrix. As input, the mean matrix Sum of standard deviation matrix Estimation is performed. In some embodiments, the variational entropy synthesis transformation model at the decoding end... The fine-tuned variational entropy synthesis transformation model is allowed to be obtained through fine-tuning. This allows it to possess a transformation model that integrates with variational entropy. Different parameters, with the input being a quantized intermediate matrix Under these conditions, a more accurate mean matrix can be obtained. Sum of standard deviation matrix The estimated value.

[0066] Alternatively, the intermediate matrix can be quantized to obtain a second quantized intermediate matrix. The mean matrix and the standard deviation matrix can then be estimated based on the second quantized intermediate matrix in the portion of the variational entropy model after the bottleneck layer to obtain the probability distribution.

[0067] For example, the intermediate matrix of the variational entropy model at the encoding end in the bottleneck layer. First, quantization can be performed to obtain the intermediate quantization matrix. Then, the variational entropy synthesis transformation model is used. For the mean matrix Sum of standard deviation matrix Make an estimate.

[0068] In some embodiments, in order to enhance the characteristics of the Gaussian distribution, i.e. the mean... and standard deviation To improve accuracy, variational entropy models can employ entropy model structures such as those based on prior knowledge, context, and checkerboard patterns to fully utilize the correlations between various dimensions in continuous latent feature maps and enhance estimation accuracy.

[0069] In one or more embodiments of this disclosure, the synonymous variational image transmission method may further include: After generating the channel overhead sequence, the channel overhead sequence is used as side information and transmitted to the receiving device corresponding to the original image via the side information transmission link. This allows the receiving device to use the channel overhead sequence received in the side information transmission link to divide the received signal sequence from the transmitting end into multiple sets of received sub-signal sequences. The length of each received sub-signal sequence is the same as the length of the channel overhead sequence.

[0070] The synonym variational image transmission method based on embodiments of this disclosure can realize an explicit signal-to-noise ratio (SNR) adaptive process based on synonym hierarchy recognition, and can adaptively adjust the granularity of transmitted semantic information according to changes in the channel SNR value. When the SNR increases, the transmitted semantic information evolves towards finer-grained semantic information. At this time, the size of the ideal synonym set at the transmitting end shrinks, making the reconstructed image and the original image consistent with finer-grained semantics. Conversely, when the SNR decreases, the transmitted semantic information changes towards coarse-grained semantic information. At this time, the size of the ideal synonym set at the transmitting end expands, making the reconstructed image and the original image consistent with coarser-grained semantics.

[0071] Figure 3 A schematic flowchart of an exemplary synonymous variational image receiving method 300 provided in an embodiment of this disclosure is shown. Figure 3 As shown, the method may include the following processing.

[0072] In step 302, a signal sequence is received from the channel. .

[0073] In step 304, the target synonym level corresponding to the signal sequence is determined. Multiple synonym levels are pre-set and identified by serial numbers. The serial number of each synonym level corresponds to a preset synonym level division rule. The synonym level division rule is used to characterize the division method of the synonym representation part and the detail representation part of the latent feature map in the current synonym level. The serial numbers of the multiple synonym levels are sorted according to the semantic importance and / or semantic granularity of the latent feature map.

[0074] In some embodiments, determining the target synonym level corresponding to the signal sequence may specifically include: Obtain the sequence number of the target synonym level from the edge information transmission link; Alternatively, the sequence number of the target synonym level can be determined based on the current signal-to-noise ratio of the channel using a preset linear or nonlinear relationship; Alternatively, the synonym decoder in a pre-trained synonym variational image transmission model determines the index of the target synonym level or a soft value sequence to characterize the index based on the signal-to-noise ratio.

[0075] In step 306, the continuous latent feature map corresponding to the target synonym level is determined based on the signal sequence.

[0076] In one or more embodiments of this disclosure, determining the continuous latent feature map corresponding to the target synonym level based on the signal sequence may specifically include: The signal sequence is used as input to the synonym decoder in the pre-trained variational synonym image transmission model, so that the synonym decoder converts the signal sequence into multiple sets of sub-signal sequences based on the channel overhead sequence obtained from the side information transmission link. The length of each sub-signal sequence is the same as the length of the channel overhead sequence, and the channel overhead sequence includes the channel overhead corresponding to the feature vector. Each sub-signal sequence is mapped to multiple sets of latent feature vectors corresponding to the target synonym level through a nonlinear transformation. The multiple sets of latent feature vectors are merged to obtain the continuous latent feature map.

[0077] For example. Figure 2 As shown in the synonym signal receiving structure section, the decoder in the image receiving device obtains the signal sequence from the wireless channel receiving port. This sequence is used as input for the image reconstruction process. A synonym decoder is then employed. Synonym hierarchy identification is performed based on the current signal-to-noise ratio of the channel. (As an example of the channel state parameters mentioned above) Determine the target synonym level Signal sequence output via wireless channel As a synonym decoder Input, determine the receptive latent feature map of the target synonym level. As a latent space feature map Specifically, this can be based on linear or nonlinear identification algorithms, utilizing the current signal-to-noise ratio of the channel. Determine the synonym hierarchy (As an example of the aforementioned target synonym hierarchy), this recognition algorithm is allowed to have inconsistent forms, steps, and model structures with the encoding algorithm, but it needs to identify the same synonym hierarchy as the encoding algorithm. Subsequently, the decoder can utilize the channel overhead sequence received in the side information transmission link. , signal sequence Divided into Group received sub-signal sequence Among them, the received sub-signal sequence The sequence length is Subsequently, according to the sequence number of the synonym hierarchy. By utilizing the corresponding inverse unequal-length nonlinear transformation, each received sub-signal sequence in the signal space is transformed. Mapped to the latent feature vector in the latent space Finally, The received latent feature vectors are combined to obtain the received latent feature map. Each of its internal latent feature vectors receives a latent feature vector. The arrangement order and the latent feature map at the encoding end Each latent feature vector The order of arrangement is the same.

[0078] Alternatively, based on a preset nonlinear transformation of unequal length and the synonym hierarchy division rule, incremental feature detection is performed on the signal sequence, mapping the signal sequence to multiple incremental feature detection results corresponding to the synonym hierarchy; Based on the sequence number of the target synonym level or the soft value sequence used to characterize the sequence number, the incremental feature detection results corresponding to the target synonym level and all synonym levels before the target synonym level are superimposed to obtain the intermediate state feature representation of the continuous latent feature map. The intermediate state feature representation is used to obtain the continuous latent feature map based on unequal-length nonlinear transformation.

[0079] For example, a signal sequence Through multiple incremental feature detection processes, it is mapped to multiple incremental feature detection results. The number of incremental feature detection processes and the number of incremental feature detection results obtained are equal to the number of synonym levels, which is... . No. indivual( Incremental feature detection results are for incremental signals The incremental semantic features within the decoder are estimated in their intermediate state at the decoding end. Secondly, the synonym hierarchy recognition process is utilized, with channel signal-to-noise ratio as the primary factor. As input, identify the target synonym level corresponding to this signal-to-noise ratio. This information is then provided as a condition for the synonymous level signal superposition process. This signal superposition process is the same as the incremental signal superposition process at the transmitting end, and will not be described further here. Subsequently, using the synonymous level feature superposition process, the received intermediate state is output. (This is an example of the intermediate state feature representation described above). Finally, the intermediate state is received. Through a nonlinear spatial transformation process at the decoding end, it is converted into the first Synonymous hierarchical receiving hidden feature map As a received continuous latent feature map .

[0080] It should be noted that the received hidden feature map It can be separated into receiving synonym representation parts. With the receiving detail characterization section Because wireless channels introduce random noise, for each received signal sequence at the current channel signal-to-noise ratio... It can obtain the same information in the latent space through a synonym decoder. Similar or related receptive synonyms in the latent space and inconsistent reception detail representations .

[0081] In step 308, the continuous latent feature map is subjected to a synonym synthesis transformation that satisfies semantic equivalence constraints to obtain a reconstructed image corresponding to the signal sequence.

[0082] For example, the nonlinear transformation that satisfies the semantic equivalence constraint on the aforementioned continuous latent feature map can be, for instance, the aforementioned synonym synthesis transformation. This synonym synthesis transformation is based on a nonlinear transformation process, which can transform the latent space representation sequence. Converted to reconstructed image in data space and each reconstructed image Integrate into a reconstructed set of synonyms .

[0083] According to the embodiments of this disclosure, the corresponding synonym level sequence number can be determined based on different channel signal-to-noise ratios, and the semantic information of the source image and the corresponding granularity of the synonym level sequence number can be transmitted through the synonym variational wireless image transmission process.

[0084] In one or more embodiments of this disclosure, the synonymous variational image receiving method may further include: The side information receiving is based on the side information link, which contains the first quantization intermediate coding sequence. The side information containing the first quantization intermediate coding sequence is entropy decoded to obtain the first quantization intermediate coding sequence. The first quantization intermediate coding sequence is obtained by quantizing and entropy coding the intermediate matrix of the variational entropy model at the bottleneck layer. The portion following the bottleneck layer of the variational entropy model estimates the mean matrix and standard deviation matrix based on the first quantization intermediate encoding sequence. The continuous latent feature map is corrected based on the mean matrix and the standard deviation matrix to obtain the corrected continuous latent feature map.

[0085] Based on the corrected continuous hidden feature map, a synonym parsing transformation satisfying semantic equivalence constraints is performed to obtain the reconstructed image corresponding to the signal sequence.

[0086] For example, the quantization intermediate matrix is ​​obtained based on the side information link reception. (As an example of the first quantization intermediate encoded sequence mentioned above), using the variational entropy synthesis transformation model described above. For the mean matrix Sum of standard deviation matrix An estimation is then performed. Subsequently, linear or nonlinear algorithms can be used, utilizing the mean matrix. Sum of standard deviation matrix For the received latent space feature map Fine-tuning was performed to better fit the mean matrix. Sum of standard deviation matrix The constructed high-dimensional Gaussian distribution.

[0087] Figure 4 This diagram illustrates a general coding framework for an exemplary synonymous variational image transmission and reception method provided in this disclosure. The process, represented by solid lines, involves the transmission and reception of semantic information from a source image based on the aforementioned synonymous variational image transmission and reception method. At the coding end, analytical transform is first utilized... , the original image Mapped to continuous latent feature maps And use the variational entropy model to estimate the continuous latent feature map. After analyzing the distribution characteristics of each element, rate estimation is performed, and the channel overhead sequence is calculated. It is then transmitted to the receiving end via a side information transmission link, and finally utilized by a synonym encoder. Obtain synonym signal sequence As a sequence of synonymous signals for transmission And send it to the wireless channel; the decoding end will receive the signal sequence. Input to synonym decoder In the process, the latent space feature map of the receiver is obtained. Then, synonym synthesis transformation was used. The latent space feature map will be received. Converted to source image in data space Reconstruction graph with synonyms .

[0088] Figure 4 The dotted-dash line pattern represents a selectively executable side information transmission process. Specifically, the quantization intermediate matrix is ​​determined in the main link using a variational entropy analytical transformation model. Subsequently, it can be transmitted to the receiving end via a side information transmission link (which may include necessary processes such as entropy coding, channel coding, modulation, wireless channel, demodulation, channel decoding, and entropy decoding). The mean matrix is ​​estimated using the variational entropy synthesis transformation model at the decoding end. Sum of standard deviation matrix Used for receiving hidden feature maps Make minor adjustments and corrections.

[0089] Figure 4 The dashed lines in the diagram are used to support the effective training of the synonymous variational image transfer model.

[0090] Figure 5A flowchart of an exemplary training method 500 for a synonymous variational image transmission model provided in this disclosure is shown. This training method optimizes the trainable parameters in a deep neural network model constructed based on the synonymous variational image transmission and reception method of this disclosure. The optimized model enables the implementation of the synonymous variational image transmission and reception method of this disclosure. The synonymous variational image transmission model includes a synonymous encoder and a synonymous decoder. The training method utilizes original images from an image dataset to form training batches. Based on the trainable modules in the synonymous variational image transmission model, multiple sets of reconstructed images are generated for each image to establish a loss function. Gradient descent is then used to achieve optimal training of the synonymous variational image transmission model. Specifically, as shown... Figure 5 As shown, the training process for the synonymous variational image transfer model can include the following steps: In step 502, image samples are acquired.

[0091] Selecting images from the training image dataset This allows multiple images to be selected from the training image dataset at once for gradient backpropagation and parameter updates during a single model training session.

[0092] In step 504, the image samples are subjected to synonym parsing transformation to satisfy semantic equivalence constraints, resulting in a continuous latent feature map.

[0093] For example, the original image can be... Input to the above synonym parsing transformation In the process, continuous hidden feature maps are obtained. .

[0094] In step 506, the probability distribution of each element in the continuous latent feature map is estimated based on the variational entropy model to obtain the probability distribution characteristics of each element.

[0095] The estimation of the probability distribution of each element in the continuous latent feature map based on the variational entropy model is consistent with the processing method of the probability distribution of each element in the continuous latent feature map based on the variational entropy model in the above synonymous variational image transmission method, and will not be repeated here.

[0096] In step 508, based on the preset sequence numbers of multiple synonym levels, the continuous latent feature map is divided into multiple sets of synonym representation parts and detail representation parts corresponding to each synonym level. The synonym level is identified by the sequence number, and the sequence number of each synonym level corresponds to a preset synonym level division rule. The synonym level division rule is used to characterize the division method of the synonym representation part and detail representation part of the latent feature map in the current synonym level. The sequence numbers of the multiple synonym levels are sorted according to the semantic importance and / or semantic granularity of the latent feature map. The method of dividing the continuous latent feature map based on the synonym hierarchy number is the same as the method of dividing the continuous latent feature map in the above synonym variational image sending method, and will not be repeated here.

[0097] In step 510, the coding rate of the synonym representation part of each synonym level is estimated based on the probability distribution characteristics of each element of the continuous hidden feature map.

[0098] The estimation of the coding rate of the synonym representation part of each synonym level based on the probability distribution characteristics of each element of the continuous latent feature map is consistent with the processing method of estimating the coding rate of the synonym representation part of each synonym level in the above synonym variational image transmission method, and will not be repeated here.

[0099] In step 512, the channel overhead corresponding to each feature vector in the continuous hidden feature map is determined, and a channel overhead sequence is generated.

[0100] The method for generating the channel overhead sequence is the same as that for generating the channel overhead sequence in the aforementioned synonymous variational image transmission method, and will not be repeated here.

[0101] In step 514, the signal-to-noise ratio threshold corresponding to each synonymous level is estimated based on the coding rate, the reference signal-to-noise ratio, and the channel overhead sequence.

[0102] For example, the signal-to-noise ratio (SNR) thresholds for each synonymous level can be estimated based on the Gaussian equivalent SNR estimation algorithm, according to the coding rate, reference SNR, and channel overhead sequence. The coding rate for each synonymous level is used as the basis for this estimation. and reference signal-to-noise ratio and channel overhead sequence As input, a corresponding signal-to-noise ratio threshold is determined for each synonymous level. In the calculation, the Gaussian channel capacity equivalence principle is followed, and the reference signal-to-noise ratio is calculated. In the channel bandwidth overhead sequence Overall transmission bandwidth overhead Using this as a reference, the equivalent channel capacity of each synonymous level on a unit channel symbol in the wireless channel is determined, as shown in the following formula, and finally the threshold signal-to-noise ratio is determined. .

[0103] ; It should be noted that the above algorithm description is merely an implementation example of the Gaussian equivalent signal-to-noise ratio estimation algorithm. Any algorithm that can obtain the same signal-to-noise ratio threshold... All algorithms described above are considered equivalent to the aforementioned algorithms and should be protected within the scope of this disclosure.

[0104] In step 516, the continuous latent feature map is used as the input to the synonym encoder in the synonym variational image transmission model to determine the synonym signal sequence corresponding to the synonym representation part of the target synonym level selected in this training.

[0105] For example, the latent space feature map can be used first. As input, transform into an intermediate state. Subsequently, the synonym levels selected in the current training batch are used. The synonym signal sequence is determined by using a hard superposition process of synonymous level signals. Alternatively, linear or nonlinear recognition algorithms can be used, utilizing synonym hierarchies. The corresponding signal-to-noise ratio threshold The synonym signal sequence is determined by a soft superposition process of synonymous level signals. .

[0106] In step 518, the target signal-to-noise ratio threshold corresponding to the target synonym level is used as the signal-to-noise ratio of the channel in this training, and the synonym signal sequence is transmitted to the channel.

[0107] Among them, synonym signal sequence Multiple transmissions are permitted within the same wireless channel under the same channel stripe to obtain multiple received signal sequences. At this point, subsequent image reconstruction requires reconstructing each received signal sequence and calculating the average loss function for the number of samples.

[0108] In step 520, the synonym signal sequence received from the channel is input into the synonymous decoder in the synonymous variational image transmission model, so that the synonymous decoder determines the target synonymous level and the continuous latent feature map corresponding to the target synonymous level based on the target signal-to-noise ratio threshold.

[0109] Among them, the synonym level selected in the current training batch is utilized. The intermediate state can be determined by using the above-mentioned hard superposition process of synonymous hierarchical features. Alternatively, linear or nonlinear recognition algorithms can be used, utilizing synonym hierarchies. The corresponding signal-to-noise ratio threshold The aforementioned synonymous hierarchical feature soft superposition process is used to determine the intermediate receiving state. In determining the intermediate state Then, the received latent feature map of the corresponding synonym level is determined by using the nonlinear spatial transformation process at the decoding end. As a received hidden feature map .

[0110] In step 522, the corresponding continuous latent feature map is subjected to a synonym synthesis transformation that satisfies the semantic equivalence constraint to obtain the reconstructed image.

[0111] In step 524, a loss value is calculated based on the reconstructed image and the image sample.

[0112] In step 526, the synonymous variational image transmission model is updated and trained according to the loss value to obtain the trained synonymous variational image transmission model. The loss value includes the total loss obtained by weighted summation of loss values ​​calculated based on at least two different loss functions.

[0113] For example, the average distortion, average sensing loss, average latent space transmission loss, and average transmission rate upper limit are calculated, and the four calculated loss values ​​are weighted and summed to determine the overall loss, and gradient backpropagation is performed.

[0114] The optimization direction of the synonymous variational image transmission model can be characterized by the following formula: ; in, Let be a learnable parameterized probability density function, representing the probability density function derived from the original image. Determined latent space with noise With signal space signal sequence The joint probability density function; in, For the first Ideal Synonym Set Image of Synonym Hierarchy Determined latent space synonym set Synonyms for signal space The joint posterior probability, where the latent space synonym set. Synonymous representation with noisy latent space The semantic information represented corresponds to the synonym set in the signal space. Synonymous signals transmitted in signal space The semantic information represented corresponds to this.

[0115] It should be noted that this optimization direction can be equivalent to the following optimization direction constructed using synonymous variational lower bounds: ; Among them, the synonymous variational lower bound has a cross-entropy form, that is, it approximates the probability density function derived from the first... Synonym hierarchy ideal synonym set image Determined latent space synonym set Synonyms for signal space The cross-entropy function of the joint posterior probability.

[0116] Taking the above optimization directions as the starting point of the analysis, the final optimization direction is analyzed using Bayes' theorem. The analysis results can be characterized by the following formula: ; in, ; It should be noted that, since the above loss function needs to be constructed and optimized simultaneously for the reconstructed images of all synonymous layers, when the number of synonymous layers is large... When the size is large, this optimization objective can lead to very high storage and computational overhead. Therefore, during the training of each batch of the model, a subset of synonymous layers can be selected. The summation is used as the loss function for optimization. For example, only a single synonymous level can be selected. Corresponding As a loss function; two synonymous layers can also be selected. or Corresponding It is used as the loss function for optimization, where and It can be Any two distinct values ​​in the table. Alternatively, more than two synonymous levels can be selected, and the corresponding loss function can be calculated and optimized using the same processing logic; this will not be elaborated upon here.

[0117] It should be noted that, in addition to the summation form of the loss function mentioned above, some mean form loss functions can also be used to optimize the synonymous variational image transmission model, such as arithmetic mean, weighted average, and weighted summation form loss functions.

[0118] For example, the above formula It can be expressed as an equivalent arithmetic mean, as shown in the following formula.

[0119] ; For example, it can also be expressed as a weighted average: ; in, Let be any nonnegative real number. This form can also be equivalent to a weighted summation. ; in, Let be any real number between 0 and 1, and require... .

[0120] In the analysis results In the middle, the first item For average distortion, This is a distortion metric function used for the data space. In some embodiments, MSE, MS-SSIM, etc., can be selected for evaluation. (Second item) To average perceived loss, This is the perceptual loss metric function. In some embodiments, divergence calculation methods such as KL divergence, Wasserstein divergence, and JS divergence can be used, or perceptual metrics that closely approximate perceptual quality, such as learning-based perceptual image patch similarity index LPIPS and deep image structure and texture similarity index DISTS, can be used for evaluation; the third item Noisy synonym representation for latent space coding end Synonyms The corresponding received synonym representation The average latent space transmission loss between them. For distortion metrics used in latent spaces, in some embodiments, mean squared error (MSE), cosine similarity, etc., can be selected for evaluation. (Last item) To be synonymous with the hierarchy The corresponding maximum transmission rate; , , as well as These are the weighting coefficients for the four terms mentioned above. Therefore, the total loss mentioned above is a weighted loss of the four terms: average distortion, average sensing loss, average latent space transmission loss, and average transmission rate.

[0121] The gradient descent algorithm is used to update the parameters of all network structures in the synonymous variational image transmission model. In the synonymous variational image transport model, the optimal value of all parameters can be represented by the loss function. The corresponding parameter values ​​when minimized. It can use one of the gradient descent algorithms, including but not limited to stochastic gradient descent and Adam algorithm, to update all parameters inside the synonymous variational image transport model.

[0122] Iterate the above operations until the weighted loss converges, or until the required number of iterations is reached, to complete the training of the synonymous variational image transmission model and obtain the final synonymous variational image transmission model.

[0123] Based on the above optimization method, a synonymous variational image transmission model can be obtained, which can be used to implement the synonymous variational image transmission and reception method described in one or more embodiments of this disclosure. The synonymous variational image transmission and reception method of this disclosure can realize an end-to-end image transmission process with adaptive semantic information granularity of source image data according to channel state parameters, and obtain a reconstructed image with a synonymous relationship to the source image at the decoding end. Specifically, for the synonymous level... In other words, when the synonymous variational image transfer model is trained to convergence, all reconstructed images The reconstructed synonym set constituted Will be with the encoding end Synonym hierarchy ideal synonym set Overlapping allows for arbitrary image reconstruction Both can serve as ideal synonym sets. A sample, having a set of ideal synonyms The semantic information depicted.

[0124] In this embodiment of the disclosure, the terminal device that performs the transmission process at the encoding end of the synonymous variational image transmission model is called the encoding end, and the terminal device that performs the reception process of the synonymous variational image transmission model is called the decoding end.

[0125] Executing the synonymous variational image transmission method of this disclosure based on the synonymous variational image transmission model at the encoding end of this disclosure may include: Using synonymous analytic transformation, the original image in the data space is taken as input to obtain continuous latent feature maps in the latent space; The probability distribution of each element in the latent space feature map is estimated using the variational entropy model; Based on multiple pre-defined synonym hierarchical indices, the continuous hidden feature map is divided into corresponding synonym representation parts and detail representation parts; Estimate the coding rate of the synonym representation portion in each synonym level; Determine the wireless channel overhead corresponding to each feature vector in the continuous hidden feature map and construct a channel overhead sequence; Using a synonym encoder, based on the synonym level number determined by the current signal-to-noise ratio of the channel, and with the latent space feature map as input, the synonym signal sequence corresponding to the synonym representation part of the synonym level is determined. The synonymous signal sequence is transmitted as a transmission synonymous signal sequence into the wireless channel.

[0126] The synonym variational image transmission model at the decoding end based on the embodiments of this disclosure, when executing the synonym variational image reception method of the embodiments of this disclosure, may include: Obtain the received signal sequence from the wireless channel; Using a synonym decoder, the corresponding synonym level is determined based on the current signal-to-noise ratio of the channel. The received signal sequence output by the wireless channel is used as input, and the received latent feature map of the corresponding synonym level is used as the received latent feature map. By using synonym synthesis transformation as input, the reconstructed image is obtained.

[0127] The synonymous variational image transmission model at the encoding end based on embodiments of this disclosure, which generates a synonymous signal sequence based on an incremental signal, may specifically include: The synonym encoder at the image transmitter takes the latent feature map as input, learns the incremental signals between multiple synonym levels, and determines the sequence of synonym signals to be transmitted by superimposing the incremental signals between multiple synonym levels based on the target synonym level determined by the channel signal-to-noise ratio. Based on the synonymous variational image transmission model at the decoding end of this disclosure embodiment, incremental feature detection is performed on the received signal sequence to obtain a latent feature map, which may specifically include: The synonym decoder at the receiving end takes the received signal sequence as input, detects the incremental features between each synonym level, and determines the intermediate state feature representation by activating the incremental features corresponding to the target synonym level and suppressing other incremental features based on the channel signal-to-noise ratio. Then, the received latent feature map is obtained based on the intermediate state feature representation.

[0128] The verification data results of the synonymous variational image transmission and reception method of the present disclosure embodiment will be described below with reference to the accompanying drawings.

[0129] One embodiment of this disclosure provides the results of synonymous image compression encoding and decoding for natural image data. Experimental conditions include: The training dataset is selected from the publicly available OpenImages v6 dataset, consisting of 100,000 natural images. Random cropping was used to adjust the resolution of the original images used as input for model training to 256. 256.

[0130] The test dataset consists of the public datasets Kodak and the validation set of DIV2K.

[0131] Number of synonym levels in the synonym variational image transport model Set to 16. In the loss function, the distortion measure in the data space. The mean squared error function (MSE) and the perceived loss metric function are selected. The learning-based perceptual image patch similarity algorithm LPIPS is used, with a distortion measure function in the latent space. The mean square error function (MSE) is selected.

[0132] During training, the number of iterations is set to [number]. The batch size was 16, and the Adam optimizer was used for training with a learning rate set to [value missing]. The training process was conducted in an AWGN (Additive White Gaussian Noise) channel. In the test channel, an AWGN channel and a 5G NR TDL-A (5G NR TDL-A) channel were selected. thThe Generation New RadioTapped Delay Line-A (TDL-A model) was tested in a fading channel. Specifically, in the 5G NR TDL-A fading channel, after obtaining the received signal sequence, the decoder first uses MMSE equalization to minimize the adverse effects of deep fading in the fading channel, and then sends it to a synonym decoder for latent space feature recovery and subsequent image reconstruction.

[0133] During the testing process, PSNR (Peak Signal-to-Noise Ratio), which measures pixel-level accuracy, and LPIPS, a learning-based perceptual image patch similarity, were used to evaluate the perceptual quality of the reconstructed image.

[0134] Figures 6(a)-6(f) illustrate schematic diagrams of reconstructed image quality curves under different datasets and different wireless channel models for two exemplary embodiments provided in this disclosure. The two embodiments are presented in a weighted summation form. The synonym hierarchy model is optimized by traversing it. The first embodiment sets... In the result image, it is marked as "SVWIT"; another embodiment is set as follows: In the result image, it is marked as "SVWIT". "Training Model". In the architectural design of the synonym encoder and synonym decoder, the two embodiments use the incremental signal soft superposition process and the incremental feature soft superposition process for the corresponding signal processing steps.

[0135] The comparative schemes include three implicit signal-to-noise ratio (SNR) adaptive methods: ADJSCC (Adaptive Deep Joint Source-Channel Coding), NTSCC+ModNet (Nonlinear Transform Source-Channel Coding + ModNet), and NTSCC+SNR embedding. All comparative schemes have the same training process and loss function as the method in the embodiments of this disclosure.

[0136] In the experimental results Figure 6(a) , 6(b)Figures 6(d) and 6(e) show that, compared to the three implicit signal-to-noise ratio adaptive methods, the method of this disclosure has superior end-to-end transmission capability, resulting in better reconstructed image quality, both in terms of pixel-level accuracy and perceived quality. Furthermore, Figures 6(c) and 6(f) demonstrate that, under 5G NR TDL-A fading channels, the method of this disclosure exhibits stronger generalization capability in fading channels.

[0137] Figure 7 The illustration shows a schematic diagram of the reconstructed image effect under different synonym levels provided in the embodiments of this disclosure. The results show that by transmitting synonym signal sequences at different granularities corresponding to different synonym levels, reconstructed images of different perceptual qualities can be obtained. Wherein, the synonym level... At the lowest synonym level (i.e., the lowest level), a reconstructed image with high perceptual similarity to the original image can be obtained, but the consistency of texture details is weaker. This indicates that the lowest synonym level focuses on perceptual similarity across the entire image. When the synonym level... As the level increases, the reconstructed image becomes more and more similar to the original image, and the texture details become clearer. This indicates that the higher the synonym level, the more attention is paid to the perceptual similarity of the details.

[0138] Verification data results show that the synonymous variational image transmission and reception method proposed in this disclosure can effectively adjust the granularity of semantic information transmitted end-to-end according to the channel signal-to-noise ratio, thereby improving the reconstruction quality.

[0139] It should be noted that the method of this disclosure embodiment can be executed by a single device, such as a computer or server. The method of this embodiment can also be applied to a distributed scenario, where multiple devices cooperate to complete the task. In such a distributed scenario, one of these devices may execute only one or more steps of the method of this disclosure embodiment, and the multiple devices will interact with each other to complete the method described.

[0140] It should be noted that the above description describes some embodiments of this disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in a different order than that shown in the above embodiments and still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0141] For ease of description, the above computer devices are described in terms of function, divided into various modules. Of course, in implementing this disclosure, the functions of each module can be implemented in one or more software and / or hardware.

[0142] The computer device described in the above embodiments is used to implement the corresponding synonymous variational image transmission method and synonymous variational image reception method in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0143] This disclosure also provides a computer device for implementing the above-described synonymous variational image transmission method and synonymous variational image reception method. Figure 8 A schematic diagram of the hardware structure of an exemplary computer device 500 provided in an embodiment of this disclosure is shown.

[0144] like Figure 8 As shown, the computer device 500 may include: a processor 502, a memory 504, a network interface 506, a peripheral interface 508, and a bus 510. The processor 502, memory 504, network interface 506, and peripheral interface 508 are interconnected within the computer device 500 via the bus 510.

[0145] Processor 502 may be a central processing unit (CPU), image processor, neural network processor (NPU), microcontroller (MCU), programmable logic device, digital signal processor (DSP), application-specific integrated circuit (ASIC), or one or more integrated circuits. Processor 502 can be used to perform functions related to the techniques described in this disclosure. In some embodiments, processor 502 may also include multiple processors integrated as a single logic component. For example, such as... Figure 8 As shown, processor 502 may include multiple processors 502a, 502b and 502c.

[0146] Memory 504 can be configured to store data (e.g., instructions, computer code, etc.). Figure 8 As shown, the data stored in memory 504 may include program instructions (e.g., one or more programs for implementing the synonymous variational image transmission method and the synonymous variational image reception method of the embodiments of this disclosure) and data to be processed (e.g., the memory may store configuration files of other modules, etc.). Processor 502 may also access the program instructions and data stored in memory 504 and execute the program instructions to operate on the data to be processed. Memory 504 may include volatile storage devices or non-volatile storage devices. In some embodiments, memory 504 may include random access memory (RAM), read-only memory (ROM), optical disk, magnetic disk, hard disk, solid-state drive (SSD), flash memory, memory stick, etc.

[0147] Network interface 506 can be configured to provide communication with other external devices to computer device 500 via a network. This network can be any wired or wireless network capable of transmitting and receiving data. For example, the network can be a wired network, a local wireless network (e.g., Bluetooth, WiFi, Near Field Communication (NFC), etc.), a cellular network, the Internet, or a combination thereof. It is understood that the type of network is not limited to the specific examples described above.

[0148] The peripheral interface 508 can be configured to connect the computer device 500 to one or more peripheral devices to enable information input and output. For example, peripheral devices may include input devices such as keyboards, mice, touchpads, touch screens, microphones, and various sensors, as well as output devices such as displays, speakers, vibrators, and indicator lights.

[0149] Bus 510 can be configured to transfer information between various components of computer device 500 (such as processor 502, memory 504, network interface 506, and peripheral interface 508), such as internal buses (e.g., processor-memory bus), external buses (USB port, PCI-E bus), etc.

[0150] It should be noted that although the architecture of the computer device 500 described above only shows the processor 502, memory 504, network interface 506, peripheral interface 508, and bus 510, in specific implementations, the architecture of the computer device 500 may also include other components necessary for normal operation. Furthermore, those skilled in the art will understand that the architecture of the computer device 500 described above may only include the components necessary for implementing the embodiments of this disclosure, and does not necessarily include all the components shown in the figures.

[0151] Those skilled in the art should understand that the discussion of any of the above embodiments is merely exemplary and is not intended to imply that the scope of this disclosure (including the claims) is limited to these examples; within the framework of this disclosure, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations of different aspects of the embodiments of this disclosure as described above, which are not provided in detail for the sake of brevity.

[0152] Additionally, to simplify the description and discussion, and to avoid obscuring the embodiments of this disclosure, the provided drawings may or may not show well-known power / ground connections to integrated circuit (IC) chips and other components. Furthermore, the apparatus may be shown in block diagram form to avoid obscuring the embodiments of this disclosure, and this also takes into account the fact that the details of implementation of these block diagram apparatuses are highly dependent on the platform on which the embodiments of this disclosure will be implemented (i.e., these details should be fully understood by those skilled in the art). While specific details (e.g., circuits) have been set forth to describe exemplary embodiments of this disclosure, it will be apparent to those skilled in the art that the embodiments of this disclosure can be implemented without these specific details or with variations thereof. Therefore, these descriptions should be considered illustrative rather than restrictive.

[0153] Although this disclosure has been described in conjunction with specific embodiments thereof, many substitutions, modifications, and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may be used with the embodiments discussed.

[0154] This disclosure is intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A method for transmitting synonymous variational images, characterized in that, include: The original image is processed by synonymous analytic transformation based on semantic equivalence constraints to obtain a continuous latent feature map; Determine the probability distribution of each element in the continuous hidden feature map; Calculate the channel overhead corresponding to each feature vector in the continuous latent feature map based on the probability distribution, and generate a channel overhead sequence. The target synonym level is determined based on the current channel state parameters. Multiple synonym levels are pre-set with serial numbers. The serial number of each synonym level corresponds to a preset synonym level division rule. The synonym level division rule is used to characterize the division method of the synonym representation part and the detail representation part of the latent feature map in the current synonym level. The serial numbers of the multiple synonym levels are sorted according to the semantic importance and / or semantic granularity of the latent feature map. Based on the channel overhead sequence, determine the synonym signal sequence corresponding to the synonym representation part of the target synonym level; The synonymous signal sequence is used as the target transmission signal sequence and transmitted to the channel.

2. The method according to claim 1, characterized in that, Determining the synonym signal sequence corresponding to the synonym representation portion of the target synonym level based on the channel overhead sequence includes: The continuous hidden feature map is subjected to an unequal-length nonlinear transformation to obtain the intermediate state feature representation of the continuous hidden feature map; Based on the synonym hierarchy partitioning rule, the intermediate state feature representation is mapped to multiple sets of incremental signals corresponding to the synonym hierarchy. Each set of incremental signals is the estimated difference between the synonym signal of the current synonym hierarchy and the synonym signal of the previous synonym hierarchy. For the same latent feature map, the multiple incremental signals corresponding to each synonym hierarchy have the same channel overhead. Based on the order of the sequence numbers of the multiple synonym levels, the incremental signal corresponding to the target synonym level and the incremental signals corresponding to all synonym levels before the target synonym level are superimposed to obtain the synonym signal sequence. Alternatively, the continuous latent feature map can be used as the input to the synonym encoder in a pre-trained synonym variational image transmission model, so that the synonym encoder maps the continuous latent feature map from the latent space to the signal space based on unequal-length nonlinear transformation, resulting in a signal sequence, and the multiple sub-signal sequences are concatenated to form a signal sequence, wherein the channel overhead of each sub-signal sequence is determined based on the channel overhead sequence; and the synonym signal sequence corresponding to the synonym representation part of the target synonym level is separated from the signal sequence.

3. The method according to claim 2, characterized in that, The channel state parameters include the signal-to-noise ratio (SNR). The target synonym level is determined based on the current channel state parameters, including: Based on a preset linear or nonlinear relationship, the sequence number of the target synonymous level is determined according to the signal-to-noise ratio, or a soft value sequence is used to characterize the sequence number, and a weight coefficient sequence corresponding to the features of the target synonymous level is determined. Based on the order of the sequence numbers of the multiple synonym levels, the incremental signals corresponding to the target synonym level and the incremental signals of all synonym levels preceding the target synonym level are superimposed according to the sequence number and the weight coefficient sequence to obtain the synonym signal sequence; or, the incremental signals corresponding to the target synonym level and the incremental signals of all synonym levels preceding the target synonym level are superimposed according to the soft value sequence and the weight coefficient sequence to obtain the synonym signal sequence. The method further includes: after determining the sequence number of the target synonym level based on the preset linear relationship or the preset nonlinear relationship according to the signal-to-noise ratio, sending the sequence number to the receiving device corresponding to the original image based on the side information transmission link.

4. The method according to claim 1, characterized in that, The channel state parameters include the signal-to-noise ratio, and the method further includes: After determining the probability distribution of each element of the continuous latent feature map, the continuous latent feature map is divided into multiple sets of synonym representation parts and detail representation parts corresponding to each synonym level based on the index of the multiple synonym levels. The synonym representation part under each synonym level is the shared spatial feature of all image samples in the synonym set corresponding to the original image. The average distortion between any image sample in the synonym set and the original image satisfies the distortion metric corresponding to each synonym level. The distortion metric corresponding to each synonym level increases or decreases sequentially according to the index of each synonym level. The perceptual loss between any image sample in the synonym set and the original image satisfies the perceptual loss metric corresponding to each synonym level. The perceptual loss metric corresponding to each synonym level increases or decreases sequentially according to the index of each synonym level. Based on the probability distribution of each element of the continuous hidden feature map, the coding rate of the synonym representation part corresponding to each synonym level is estimated. The target synonym level is determined based on the current channel state parameters, including: The sequence number of the target synonym level or a soft value sequence used to characterize the sequence number is determined based on the signal-to-noise ratio and the coding rate.

5. A method for receiving synonymous variational images, characterized in that, include: Receive a sequence of signals from the channel; Determine the target synonym level corresponding to the signal sequence, wherein multiple synonym levels are pre-set with serial numbers, and the serial number of each synonym level corresponds to a preset synonym level division rule. The synonym level division rule is used to characterize the division method of the synonym representation part and the detail representation part of the latent feature map in the current synonym level. The serial numbers of the multiple synonym levels are sorted according to the semantic importance and / or semantic granularity of the latent feature map. Determine the continuous hidden feature map corresponding to the target synonym level based on the signal sequence; The continuous hidden feature map is subjected to a synonym synthesis transformation that satisfies semantic equivalence constraints to obtain a reconstructed image corresponding to the signal sequence.

6. The method according to claim 5, characterized in that, Determining the target synonym level corresponding to the signal sequence includes: Obtain the sequence number of the target synonym level from the edge information transmission link; Alternatively, the sequence number of the target synonym level can be determined based on the current signal-to-noise ratio of the channel using a preset linear or nonlinear relationship; Alternatively, the synonym decoder in the pre-trained synonym variational image transmission model can determine the sequence number of the target synonym level or a soft value sequence to characterize the sequence number based on the signal-to-noise ratio.

7. The method according to claim 5, characterized in that, Determining the continuous latent feature map corresponding to the target synonym level based on the signal sequence includes: The signal sequence is used as input to the synonym decoder in the pre-trained synonym variational image transmission model. The synonym decoder, based on the channel overhead sequence obtained from the side information transmission link, converts the signal sequence into multiple sets of sub-signal sequences, wherein the length of each sub-signal sequence is the same as the length of the channel overhead sequence, and the channel overhead sequence includes the channel overhead corresponding to the feature vectors. Each sub-signal sequence is mapped to multiple sets of latent feature vectors corresponding to the target synonym level through a nonlinear transformation. The multiple sets of latent feature vectors are merged to obtain the continuous latent feature map. Alternatively, based on a preset nonlinear transformation of unequal length and the synonym hierarchy division rule, incremental feature detection is performed on the signal sequence, mapping the signal sequence to multiple incremental feature detection results corresponding to the synonym hierarchy; Based on the sequence number of the target synonym level or the soft value sequence used to characterize the sequence number, the incremental feature detection results corresponding to the target synonym level and all synonym levels before the target synonym level are superimposed to obtain the intermediate state feature representation of the continuous latent feature map. The intermediate state feature representation is used to obtain the continuous latent feature map based on unequal-length nonlinear transformation.

8. The method according to any one of claims 5 to 7, characterized in that, The method further includes: The side information receiving is based on the side information link, which contains the first quantization intermediate coding sequence. The side information containing the first quantization intermediate coding sequence is entropy decoded to obtain the first quantization intermediate coding sequence. The first quantization intermediate coding sequence is obtained by quantizing and entropy coding the intermediate matrix of the variational entropy model at the bottleneck layer. The portion following the bottleneck layer of the variational entropy model estimates the mean matrix and standard deviation matrix based on the first quantization intermediate encoding sequence. The continuous latent feature map is corrected based on the mean matrix and the standard deviation matrix to obtain the corrected continuous latent feature map. Based on the corrected continuous latent feature map, a synonym synthesis transformation satisfying semantic equivalence constraints is performed to obtain a reconstructed image corresponding to the signal sequence.

9. A training method for a synonymous variational image transport model, characterized in that, include: Obtain image samples; The image samples are subjected to synonym parsing transformation based on semantic equivalence constraints to obtain continuous latent feature maps; The probability distribution of each element in the continuous latent feature map is estimated based on the variational entropy model, and the probability distribution characteristics of each element in the continuous latent feature map are obtained. Based on the preset sequence numbers of multiple synonym levels, the continuous latent feature map is divided into multiple sets of synonym representation parts and detail representation parts corresponding to each synonym level. The synonym level is identified by the sequence number, and the sequence number of each synonym level corresponds to the preset synonym level division rule. The synonym level division rule is used to characterize the division method of the synonym representation part and detail representation part of the latent feature map in the current synonym level. The sequence numbers of the multiple synonym levels are ordered according to the semantic importance and / or semantic granularity of the latent feature map. Based on the probability distribution characteristics of each element of the continuous hidden feature map, the encoding rate of the synonym representation part of each synonym level is estimated. Determine the channel overhead corresponding to each feature vector in the continuous hidden feature map, and generate a channel overhead sequence; The signal-to-noise ratio threshold corresponding to each synonymous level is estimated based on the coding rate, the reference signal-to-noise ratio, and the channel overhead sequence. The continuous latent feature map is used as the input to the synonym encoder in the synonym variational image transmission model to determine the synonym signal sequence corresponding to the synonym representation part of the target synonym level selected in this training. The target signal-to-noise ratio threshold corresponding to the target synonym level is used as the signal-to-noise ratio of the channel in this training, and the synonym signal sequence is transmitted to the channel; The synonym decoder in the synonym variational image transmission model is input from the synonym signal sequence received from the channel, so that the synonym decoder determines the target synonym level and the continuous latent feature map corresponding to the target synonym level based on the target signal-to-noise ratio threshold. The reconstructed image is obtained by performing a synonym synthesis transformation on the continuous latent feature map that satisfies the semantic equivalence constraint. Calculate the loss value based on the reconstructed image and the image samples; The synonymous variational image transmission model is updated and trained based on the loss value to obtain the trained synonymous variational image transmission model. The loss value includes the total loss obtained by weighted summation of loss values ​​calculated based on at least two different loss functions.

10. A computer device comprising one or more processors, a memory; and one or more programs, wherein the one or more programs are stored in the memory and executed by the one or more processors, the one or more programs comprising instructions for performing the method of any one of claims 1 to 9.