Image bimodal transmission method, device and equipment for discrete semantic representation

CN122824899APending Publication Date: 2026-09-25BEIJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610809167.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-05
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0003]有鉴于此,本申请的目的在于提出一种面向离散语义表示的图像双模态传输方法、装置及设备,解决了语义通信中多模态语义信息难以统一、鲁棒性不足以及语义特征易失真的问题

Benefits of technology

[0015]相对于现有技术,本发明的技术效果在于:将不同模态的语义信息统一表示为离散语义索引,并结合二维星座点学习机制,实现不同模态语义信息在统一通信框架下的表示与传输;设计基于软判决的离散语义索引恢复机制,提高离散语义表示在复杂信道环境下的容错能力;设计面向信道估计误差的自适应语义变换机制,对非完美信道状态条件下的语义特征偏移进行补偿与校正,提升系统在复杂无线环境下的传输稳定性与鲁棒性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122824899A_ABST
    Figure CN122824899A_ABST
Patent Text Reader

Abstract

The application provides an image double-mode transmission method, device and equipment for discrete semantic representation, which comprises the following steps: constructing a transmission system, the transmission system comprising an image semantic layer and a text semantic layer; acquiring image information, extracting shallow basic semantic information and deep semantic auxiliary information, and converting the information into a discrete semantic index sequence; using a semantic-channel encoder provided with an adaptive semantic transformation mechanism to generate a transmission signal from the discrete semantic index sequence; in an image semantic-channel decoder at the receiving end, using a soft decision-based discrete semantic index recovery mechanism to recover the transmission signal of the image semantic layer into a discrete semantic index sequence and continuous semantic features; recovering the transmission signal into discrete image structure semantics and text semantic description, fusing the image structure semantics and the text semantic description, and recovering the image information. The problems of non-uniformity of multi-modal semantic information, insufficient robustness and semantic distortion in semantic communication are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of network communication technology, and in particular to a method, apparatus and device for image bimodal transmission oriented towards discrete semantic representation. Background Technology

[0002] Semantic communication is a novel communication technology oriented towards task-based and semantic information delivery. Unlike the indiscriminate transmission of bit-level symbol sequences in classical communication systems, the core idea of ​​semantic communication lies in extracting, encoding, and transmitting task-related semantic content from information. Multimodal information representation technologies mainly fall into two development directions: multimodal alignment and cross-modal generation. Multimodal alignment primarily constructs a mapping relationship between different modal data and a unified semantic space by comparing basic models such as Contrastive Language-Image Pre-training (CLIP). Existing multimodal semantic communication systems still face challenges in practical wireless transmission environments, including difficulties in uniformly representing multimodal semantic information, insufficient robustness of discrete semantic index transmission, and the susceptibility of semantic features to distortion under imperfect channel state information conditions. Summary of the Invention

[0003] In view of this, the purpose of this application is to propose an image bimodal transmission method, apparatus and device for discrete semantic representation, which solves the problems of difficulty in unifying multimodal semantic information, insufficient robustness and easy distortion of semantic features in semantic communication.

[0004] To achieve one of the aforementioned objectives, this application provides a method for image bimodal transmission oriented towards discrete semantic representation, the method comprising:

[0005] Construct a transmission system; wherein the transmission system includes an image semantic layer and a text semantic layer; Obtain image information, extract shallow basic semantic information of the image information in the image semantic layer, extract deep semantic auxiliary information of the image information in the text semantic layer, and convert both the shallow basic semantic information and the deep semantic auxiliary information into discrete semantic index sequences; Channel state information is obtained, and the discrete semantic index sequence is transformed into an adaptive constellation point set using a semantic-channel encoder with an adaptive semantic transformation mechanism. The discrete semantic index sequence is semantically encoded and modulated by the adaptive constellation points to generate a transmission signal, which is then sent from the transmitting end to the receiving end. In the image semantic-channel decoder at the receiving end, a discrete semantic index recovery mechanism based on soft decision is used to recover the transmitted signal of the image semantic layer into the discrete semantic index sequence and continuous semantic features; In the text semantic-channel decoder at the receiving end, the transmission signal of the text semantic layer is restored to the discrete semantic index sequence; The continuous semantics are converted into image structural semantics using an image semantic synthesizer, and the discrete semantic index sequence of the text semantic layer is converted into a text semantic description using a text semantic synthesizer. The image structural semantics and the text semantic description are then fused to recover the image information.

[0006] As a further improvement to one embodiment of this application, the step of extracting shallow basic semantic information of the image information in the image semantic layer and converting the shallow basic semantic information into a discrete semantic index sequence includes: High-dimensional semantic features of the image information are extracted using an image semantic extractor, and a vector quantization codebook of fixed codebook size is generated based on an exponential moving average algorithm. For each of the high-dimensional semantic features, the codeword with the smallest Euclidean distance is searched in the vector quantization codebook, and the codeword indices corresponding to all the high-dimensional semantic features are set into the discrete semantic index sequence. The step of extracting deep semantic auxiliary information from the image information in the text semantic layer, and converting all of the deep semantic auxiliary information into discrete semantic index sequences, includes: By using a multimodal large model to perform semantic understanding of image information, structured text descriptions of images are generated. The image structured text description is converted into the discrete semantic index sequence using a text semantic extractor.

[0007] As a further improvement to one embodiment of this application, the step of generating a correction component of the channel state information using a semantic-channel encoder equipped with the adaptive semantic transformation mechanism, and converting the discrete semantic index sequence into an adaptive constellation point set based on the correction component, includes: The discrete semantic index sequence is passed through the semantic-channel encoder to extract the latent image representation; wherein the latent image representation includes the discrete semantic index sequence, the learnable parameters of the semantic-channel encoder, and the channel state information; In the semantic-channel encoder, the signal-to-noise ratio is divided into intervals, such that each interval of the signal-to-noise ratio corresponds to a basic constellation set; For each of the aforementioned basic constellation sets, a correction factor is learned; The adaptive constellation point set is generated by superimposing the basic components of the basic constellation set and the correction amount.

[0008] As a further improvement to one embodiment of this application, the semantic-channel encoder equipped with the adaptive semantic transformation mechanism includes: The semantic-channel encoder is configured with alternating layers of adaptive feature scaling and general feature processing. The discrete semantic index sequence, the learnable parameters of the semantic-channel encoder, and the channel state information are input as semantic features to the first layer of the adaptive feature scaling layer. After the semantic features are adaptively scaled by the adaptive feature scaling layer, the scaled semantic features are input to the first layer of the general feature processing layer for processing, and then input to the next layer of the adaptive feature scaling layer. The semantic features are processed in layers to generate the final semantic features.

[0009] As a further improvement to one embodiment of this application, the semantic features are adaptively scaled using the adaptive feature scaling layer, including: The input semantic features are subjected to global average pooling to extract channel-level statistical information; By integrating the aforementioned channel-level statistical information, real-time signal-to-noise ratio, and channel estimation error variance, a joint conditional vector is generated through a multilayer perceptron embedding layer. Based on the joint conditional vector, complex-domain adaptive gating coefficients are generated, and amplitude and phase joint compensation corrections are performed on the real and imaginary parts of the semantic features, respectively. By combining the real-time signal-to-noise ratio to generate channel-level attention weights, the corrected semantic features are adaptively scaled according to the channel dimension to complete a semantic feature adaptive transformation.

[0010] As a further improvement to one embodiment of this application, the soft-decision-based discrete semantic index recovery mechanism recovers the transmission signal of the image semantic layer into the discrete semantic index sequence and continuous semantic features, including: The transmission signal of the image semantic layer is obtained by the receiving end, and the transmission signal is processed by channel estimation and equalization to obtain the recovered symbol. The distance metric between the recovered symbol and each codeword in the vector quantization codebook is calculated. By introducing a Gumbel random perturbation term and a temperature parameter, the distance metric is converted into a posterior probability distribution for each candidate codeword. The vector quantized codebook is weighted and combined using the probability distribution to obtain the recovered continuous semantic features.

[0011] As a further improvement to one embodiment of this application, the soft-decision-based discrete semantic index recovery mechanism, before recovering the transmission signal of the image semantic layer into the discrete semantic index sequence and continuous semantic features, includes: Under noise-free conditions, the image semantic extractor and the image semantic synthesizer are pre-trained; wherein, the forward propagation during the training process uses hard decision vector quantization to generate discrete indices, and the back propagation during the training process uses a gradient pass-through strategy to update parameters. The parameters of the image semantic extractor and the image semantic synthesizer, which have been pre-trained, are frozen. A joint source-channel coding and transmission module and a soft-decision recovery mechanism are introduced. Through joint constraints of index decision loss, semantic embedding alignment loss, and channel transmission loss, the learnable parameters of the image semantic-channel codec and the semantic modulator parameters are optimized.

[0012] As a further improvement to one embodiment of this application, the step of fusing the image structural semantics and the text semantic description to recover the image information includes: The image structural semantics and the text semantic description are used as joint conditional inputs to the cross-modal generation model. Through cross-modal feature alignment and semantic consistency constraints, the latent spatial semantic information is fused and completed to generate the image information.

[0013] Based on the same inventive concept, this application also provides an image dual-modal transmission device for discrete semantic representation, comprising: A construction module is used to construct a transmission system; wherein the transmission system includes an image semantic layer and a text semantic layer; The acquisition module is used to acquire image information, extract shallow basic semantic information of the image information in the image semantic layer, extract deep semantic auxiliary information of the image information in the text semantic layer, and convert both the shallow basic semantic information and the deep semantic auxiliary information into discrete semantic index sequences. The conversion module is used to acquire channel state information and, using a semantic-channel encoder equipped with an adaptive semantic transformation mechanism, convert the discrete semantic index sequence into an adaptive constellation point set. The transmitting module is used to generate a transmission signal by semantically encoding the discrete semantic index sequence and modulating it with the adaptive constellation points, and then transmit the transmission signal from the transmitting end to the receiving end. The image restoration module is used in the image semantic-channel decoder at the receiving end to restore the transmitted signal of the image semantic layer into the discrete semantic index sequence and continuous semantic features based on the soft-decision discrete semantic index restoration mechanism. The text recovery module is used to recover the transmitted signal of the text semantic layer into the discrete semantic index sequence in the text semantic-channel decoder at the receiving end; The fusion module is used to convert the continuous semantics into image structural semantics using an image semantic synthesizer, convert the discrete semantic index sequence of the text semantic layer into a text semantic description using a text semantic synthesizer, and fuse the image structural semantics and the text semantic description to restore the image information.

[0014] Based on the same inventive concept, this application also provides an electronic device, including: a processor and a memory; The memory stores a computer program, which, when executed by the processor, causes the processor to perform the steps of the image bimodal transmission method for discrete semantic representation described above.

[0015] Compared with existing technologies, the technical advantages of this invention are as follows: it unifies the semantic information of different modalities into a discrete semantic index, and combines it with a two-dimensional constellation point learning mechanism to realize the representation and transmission of semantic information of different modalities under a unified communication framework; it designs a discrete semantic index recovery mechanism based on soft decision to improve the fault tolerance capability of discrete semantic representation in complex channel environments; and it designs an adaptive semantic transformation mechanism oriented towards channel estimation errors to compensate and correct semantic feature offsets under imperfect channel conditions, thereby improving the transmission stability and robustness of the system in complex wireless environments. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in this application or related technologies, the drawings used in the description of the implementation methods or related technologies will be briefly introduced below. Obviously, the drawings described below are only the implementation methods of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 A flowchart illustrating an image dual-modal transmission method oriented towards discrete semantic representation, provided in one embodiment of this application; Figure 2 A framework diagram of an image bimodal transmission method oriented towards discrete semantic representation provided in one embodiment of this application; Figure 3 A framework diagram of the adaptive semantic transformation mechanism provided in one embodiment of this application; Figure 4 A flowchart of a training-based soft-decision discrete semantic index transmission mechanism provided in one embodiment of this application; Figure 5 A schematic diagram of an image dual-modal transmission device for discrete semantic representation provided in another embodiment of this application; Figure 6 This is a schematic diagram of the hardware structure of an electronic device provided for another embodiment of this application. Detailed Implementation

[0018] The present invention will now be described in detail with reference to the specific embodiments shown in the accompanying drawings. However, these embodiments do not limit the present invention, and any structural, methodological, or functional modifications made by those skilled in the art based on these embodiments are included within the scope of protection of the present invention.

[0019] It should be noted that, unless otherwise defined, the technical or scientific terms used in the embodiments of this application should have the ordinary meaning understood by those skilled in the art to which this application pertains. The terms "first," "second," and similar terms used in the embodiments of this application do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects.

[0020] Existing technologies include schemes focusing on image compression and multimodal semantic information processing in ultra-low bitrate scenarios. These schemes utilize large multimodal models to automatically generate image description text, achieving efficient text semantic compression through the large model. By combining image coding networks and diffusion generation models, collaborative compression and reconstruction of image and text semantics are achieved. A loss function that integrates perception and semantic constraints is designed to maintain image visual integrity and semantic consistency while reducing bitrate, thereby improving image reconstruction quality. However, this approach does not adequately consider robust transmission mechanisms in complex wireless environments, failing to fully account for the impact of bit errors and noise disturbances on semantic features during channel transmission. It also lacks robust coding strategies for semantic representation sequences and differentiated transmission mechanisms for key semantic information.

[0021] Existing technologies have proposed a joint source-channel coding image wireless transmission method based on an adaptive autoencoder. This method combines image feature learning with varying signal-to-noise ratios (SNR), generates a channel adjustment factor based on the SNR value, and scales the features of each layer of the encoder to control the feature representation under different channel conditions, thereby enhancing the model's robustness to channel noise of varying intensities. However, in real-world wireless transmission environments, due to factors such as channel time-varying characteristics, fading features, and feedback link errors, the channel state information fed back by the receiver may be biased. This can lead to a mismatch between the adaptive adjustment strategy and the actual channel conditions, thus affecting system performance.

[0022] To address the aforementioned problems, this application provides an image dual-modal transmission method oriented towards discrete semantic representation, such as... Figure 1 and Figure 2 As shown, the method includes the following steps: Step S100: Construct a transmission system; wherein the transmission system includes an image semantic layer and a text semantic layer.

[0023] The transmission system of this method is divided into an image semantic layer and a text semantic layer, which respectively carry the basic semantic information of the image and the extracted deep text semantic information of the image.

[0024] Step S200: Obtain image information, extract shallow basic semantic information of image information in image semantic layer, extract deep semantic auxiliary information of image information in text semantic layer, and convert both shallow basic semantic information and deep semantic auxiliary information into discrete semantic index sequences.

[0025] Specifically, the image semantic layer uses miniaturized neural networks to extract shallow, basic semantic information from images. It is responsible for extracting and transmitting basic image structure and contour information, including key steps such as image semantic extraction, vector quantization, image semantic-channel coding, semantic modulation, channel estimation and equalization, image semantic-channel decoding, indexed soft decision, and image semantic synthesis. Building upon the image semantic layer, the text semantic layer further leverages a multimodal large model for semantic understanding, extracting more concise and intuitive deep semantic auxiliary information. This serves as an auxiliary path to provide conditional enhancement information for cross-modal generation at the receiver, including text semantic extraction, text semantic-channel coding, semantic modulation, channel estimation and equalization, text semantic-channel decoding, and text semantic synthesis.

[0026] Step S300: Obtain channel state information and use a semantic-channel encoder with an adaptive semantic transformation mechanism to transform the discrete semantic index sequence into an adaptive constellation point set.

[0027] In step S400, the discrete semantic index sequence is semantically encoded and adaptive constellation point modulated to generate a transmission signal, which is then sent from the transmitting end to the receiving end.

[0028] In step S500, in the image semantic-channel decoder at the receiving end, the transmitted signal of the image semantic layer is restored into a discrete semantic index sequence and continuous semantic features based on the soft-decision discrete semantic index recovery mechanism.

[0029] Step S600: In the text semantic-channel decoder at the receiving end, the transmitted signal of the text semantic layer is restored to a discrete semantic index sequence.

[0030] Step S700: The continuous semantics are converted into image structural semantics using an image semantic synthesizer, and the discrete semantic index sequence of the text semantic layer is converted into a text semantic description using a text semantic synthesizer. The image structural semantics and the text semantic description are then fused to restore the image information.

[0031] In one possible implementation of this application, step S200 in the image semantic layer specifically includes: Step S210: Extract high-dimensional semantic features of image information using an image semantic extractor, and generate a vector quantization codebook of fixed codebook size based on an exponential moving average algorithm.

[0032] Specifically, in the image semantic information transmission stage, image information First, it passes through an image semantic extractor. Extracting high-dimensional feature representations with semantic expressive power ,in These represent the spatial resolution and channel dimension of the feature map, respectively. These are the learnable parameters for the image semantic extractor. The vector quantization codebook is obtained based on the exponential moving average (EMA). , where K represents the codebook size.

[0033] Step S220: For each high-dimensional semantic feature, search for the codeword with the smallest Euclidean distance in the vector quantization codebook, and set the codeword indices corresponding to all high-dimensional semantic features into a discrete semantic index sequence.

[0034] Specifically, for each feature vector The codeword with the smallest distance in the codebook is indexed as follows: ; This yields the discrete semantic index sequence of the image semantic layer: .

[0035] Step S200 in the text semantic layer specifically includes: Step S230: Semantic understanding of image information is performed using a multimodal large model to generate a structured text description of the image.

[0036] Step S240: Use a text semantic extractor to convert the image structured text description into a discrete semantic index sequence.

[0037] Specifically, during the transmission of text semantics, a cross-modal large model is first used for semantic understanding to generate the corresponding discrete descriptive text T. Similarly, a text semantic extractor... Map T to a semantic embedding vector , Further input to the text semantic-channel encoder By combining the current channel state and a channel-adaptive semantic transformation mechanism, a text-robust coding sequence is generated. Semantic modulator Will Mapped to discrete sequence .

[0038] In one possible implementation of this application, step S300 includes: Step S310: The discrete semantic index sequence is passed through a semantic-channel encoder to extract the latent image representation; wherein, the latent image representation includes the discrete semantic index sequence, the learnable parameters of the semantic-channel encoder, and channel state information.

[0039] Specifically, taking the semantic-channel encoder as an example, the image semantic-channel encoder is as follows: Image latent representation ,in This indicates the length of the semantic sequence to be sent from the image semantic layer. These are the learnable parameters of the image semantic-channel encoder. This indicates the current channel state and is used for robust coding of the neural network-guided model.

[0040] Step S320: In the semantic-channel encoder, the signal-to-noise ratio is divided into intervals, so that each interval of signal-to-noise ratio corresponds to a basic constellation set.

[0041] Step S330: For the confidence level of each basic constellation set, learn the correction amount.

[0042] Step S340: Superimpose the basic components and correction values ​​of the basic constellation set to generate an adaptive constellation point set.

[0043] Semantic modulator Based on the confidence level of the signal-to-noise ratio and channel state information, Dynamically mapped to discrete complex constellation point sequences Specifically, the model first divides the possible signal-to-noise ratio (SNR) range into intervals, with each SNR interval corresponding to a set of basic constellations. This is used to determine the basic spatial arrangement of the signal. For different confidence levels, corresponding correction values ​​are learned. This is to compensate for geometric offsets and decision mismatches caused by channel uncertainties under imperfect channel conditions. The final adaptive constellation point set... It is formed by the linear superposition of the basic component and the modified component, and the formula is expressed as: ; Where M is the number of elements in the constellation point set, i.e., the modulation order.

[0044] It should be noted that the semantic-channel encoder includes a text semantic-channel encoder and an image semantic-channel encoder, and steps S310-S340 need to be executed in both the text semantic layer and the image semantic layer. When executed in the text semantic layer, the semantic-channel encoder is a text semantic-channel encoder, and the discrete semantic index sequence is a discrete semantic index sequence of deep semantic auxiliary information; when executed in the image semantic layer, the semantic-channel encoder is an image semantic-channel encoder, and the discrete semantic index sequence is a discrete semantic index sequence of shallow basic semantic information.

[0045] In one possible implementation of this application, to achieve robust transmission of semantic information under non-ideal channel conditions, such as learning imperfect channel state information, an adaptive semantic transformation mechanism oriented towards non-ideal channel state information is designed. This enables the image semantic-channel encoder to adaptively adjust the semantic representation distribution according to different channel conditions, thereby improving the noise resistance and stability of transmission. Specifically, as follows... Figure 3 As shown, in step S300, a semantic-channel encoder equipped with an adaptive semantic transformation mechanism is used, including: Step S301: Set up an alternating layered adaptive feature scaling layer and a general feature processing layer in the semantic-channel encoder.

[0046] In step S302, the discrete semantic index sequence, the learnable parameters of the semantic-channel encoder, and the channel state information are used as semantic features and input to the first layer of adaptive feature scaling layer. After the semantic features are adaptively scaled by the adaptive feature scaling layer, the scaled semantic features are input to the first layer of general feature processing layer for processing, and then input to the next layer of adaptive feature scaling layer. The semantic features are processed layer by layer in sequence to generate the final semantic features.

[0047] Specifically, the semantic-channel codec consists of an overlapping adaptive feature scaling layer and a general feature processing layer. The adaptive feature scaling layer employs an adaptive encoding / decoding strategy, considering the dynamic changes and imperfect channel state feedback under complex real-world communication environments. It introduces uncertainty information, represented by channel estimation error, to achieve image and text semantic encoding / decoding driven by both channel state and estimation error, providing adaptive semantic representation. Compared to classical methods that rely solely on ideal channel state information, this invention can robustly recover semantic features even with incomplete channel information, effectively mitigating the semantic representation shift problem caused by channel disturbances.

[0048] In one possible implementation of this application, step S302, which uses an adaptive feature scaling layer to adaptively scale the semantic features, includes: Step S3021: Perform global average pooling on the input semantic features to extract channel-level statistical information.

[0049] Step S3022: In this step, channel-level statistical information, real-time signal-to-noise ratio, and channel estimation error variance are fused to generate a joint conditional vector through a multilayer perceptron embedding layer.

[0050] Step S3023: Generate complex-domain adaptive gating coefficients based on joint conditional vectors, and perform amplitude and phase joint compensation correction on the real and imaginary parts of the semantic features respectively.

[0051] Step S3024: Combine the real-time signal-to-noise ratio to generate channel-level attention weights, and perform adaptive scaling of the channel dimension on the corrected semantic features to complete an adaptive transformation of semantic features.

[0052] Specifically, the semantic feature representation of one layer is set as follows: ,in These represent the number of channels and the resolution of the semantic features, respectively. First, global average pooling is used to extract the channel-level statistics of the input semantic features, represented as: ; in, This indicates global average pooling, used to generate a global description of the response for each channel.

[0053] Subsequently, the semantic statistics g, signal-to-noise ratio (SNR), and channel estimation error variance are used. To perform fusion, construct joint conditions: ; in It is an embedding layer composed of multilayer perceptrons (MLPs).

[0054] Considering the dual impact of channel coefficients on the amplitude and phase of complex signals, the joint condition vector c is further processed by a lightweight nonlinear mapping function to pre-correct the features based on the degree of estimation error, generating adaptive gating coefficients. Based on this, dynamic phase and amplitude modulation operators in complex space are calculated, and joint compensation is performed on the real and imaginary parts of the features to correct the semantic features. The real and imaginary parts of each element in the expression are represented as follows: ; ; Among them, symbols Represents element-wise multiplication. and These are the input semantic features. The real and imaginary parts; and The real and imaginary parts of the corresponding generated gating coefficients.

[0055] Subsequently, after completing the complex domain compensation, channel-level attention weights are generated based on the signal-to-noise ratio, and feature scaling is applied to the real and imaginary parts of the features respectively: ; in, This is used to generate attention coefficients driven by the signal-to-noise ratio, which strengthen high-confidence channels and suppress the responses of channels susceptible to noise, thereby enhancing the model's adaptability to different channel quality conditions.

[0056] In one possible implementation of this application, in step S400, based on the adaptive constellation set, the semantic feature mapping process follows the nearest distance principle, that is, for the input semantic features... The constellation point with the smallest Euclidean distance from the constellation set is selected as the output symbol. The hard-decision mapping process can be represented as follows: ; Modulated complex sequence The received signal transmitted through the physical channel is represented as: ; Where n represents the channel noise, i.e. Channel coefficient Follows a complex Gaussian distribution .

[0057] In one possible implementation of this application, step S500 includes: Step S510: The receiving end acquires the transmission signal of the image semantic layer, and the transmission signal is processed by channel estimation and equalization to obtain the recovered symbol. The distance metric between the recovered symbol and each codeword in the vector quantization codebook is calculated.

[0058] Step S520 introduces the Gumbel random perturbation term and temperature parameter to convert the distance metric into the posterior probability distribution of each candidate codeword.

[0059] Step S530: The vector quantized codebook is weighted and combined using the probability distribution to obtain the recovered continuous semantic features.

[0060] Specifically, the receiver corrects the signal using a channel estimation and equalization module to obtain the recovered symbol. In the image reconstruction stage, the image semantic-channel decoder First of all Perform nonlinear inverse mapping to recover the sequence of semantic information of the image. ,in These are learnable parameters for the image semantic-channel decoder. In the semantic feature recovery process based on vector quantization codebooks, to adapt to bit error transmission in noisy channels and avoid semantic feature jumps caused by index error perturbations, a soft-decision mechanism for discrete semantic indexes is designed and used. This soft-decision mechanism restores the discrete index sequence into continuous semantic features. .

[0061] Specifically, let the image vector quantization codebook be... For each received feature Calculate the distance metric between it and each codeword: ; Subsequently, a soft-decision mechanism is employed to transform the distance metric into a probability distribution. To make the training process more closely resemble discrete sampling, a Gumbel random perturbation is introduced, where the probability of the i-th position belonging to the k-th codeword is... Represented as: ; in, This represents the discrimination score of the k-th codeword corresponding to the i-th position in the output of the receiving network. This represents a random perturbation term that follows a Gumbel distribution.

[0062] This implementation can better approximate the discrete index selection process while maintaining the differentiability of the probability distribution. Temperature parameter Used to adjust the smoothness of the distribution. When When the probability distribution is smaller, the probability distribution is sharper, and the recovery result is closer to a hard decision; when When the probability distribution is larger, the probability distribution is smoother, and multiple candidate codewords can obtain non-zero weights, which is beneficial for preserving more potential semantic information when there is strong noise or uncertainty in the decision. Further utilizing this probability distribution to perform weighted combination of the codebook, the... The restored semantic features corresponding to each position are represented as follows: The resulting semantic sequence The image content is then restored by inputting it into the subsequent image semantic synthesizer.

[0063] Specifically, the soft-decision method described above is used not only during training to ensure the differentiability of the training process, but also during inference. This ensures that the receiver output is no longer a discrete and abrupt single-point index, but a continuous semantic representation containing uncertain information. When channel conditions are poor or index decisions are ambiguous, soft decision-making can preserve the relative confidence of multiple candidate semantic bases, allowing the semantic recovery result to exhibit a smooth transition in the latent space, rather than a complete erroneous jump. Therefore, even if some symbols are disturbed by noise, it only manifests as a slight shift in semantic distribution, without directly causing serious structural distortion, thereby improving the system's tolerance to bit errors.

[0064] It should be noted that in the transmission system, both image and text semantic information are transmitted in the form of discrete semantic indices, thereby reducing transmission overhead. However, while the discrete semantic index sequence formed at the transmitting end is beneficial for forming structured, low-redundancy, and interpretable semantic units, its decision mapping process is non-differentiable, making it difficult to integrate into an end-to-end training framework based on gradient propagation. Furthermore, in actual wireless propagation, there are non-ideal factors such as noise, fading, and interference. A disturbance in a single transmitted symbol can easily cause discrete jumps in the index output at the receiving end, ultimately resulting in significant image distortion. Therefore, this method designs a soft-decision recovery mechanism based on probabilistic modeling at the receiving end. This mechanism no longer directly maps received features to a single discrete index. Instead, it constructs the posterior probability distribution of each candidate index based on the received signal and recovers the continuous semantic representation in a probabilistically weighted manner. Through this approach, the output at the receiving end is no longer a single-point discrete decision result, but a continuous semantic estimation result containing uncertainty information and contributions from multiple candidate semantics, which can better adapt to bit error disturbances in noisy channels.

[0065] In one possible implementation of this application, such as Figure 4 As shown, before using the soft-decision discrete semantic index recovery mechanism in step S500, it is necessary to train the soft-decision discrete semantic index recovery mechanism, which specifically includes: Step S501: Under noise-free conditions, the image semantic extractor and the image semantic synthesizer are pre-trained; wherein, the forward propagation during the training process uses hard decision vector quantization to generate discrete indices, and the backpropagation during the training process uses a gradient pass-through strategy to update parameters.

[0066] Step S502: Freeze the pre-trained image semantic extractor and image semantic synthesizer parameters, introduce the joint source-channel coding transmission module and soft decision recovery mechanism, and optimize the learnable parameters of the image semantic-channel codec and the semantic modulator parameters through joint constraints of index decision loss, semantic embedding alignment loss and channel transmission loss.

[0067] Specifically, in order to further ensure the transmission performance of vector quantization and discrete index, this method adopts a phased training strategy to gradually optimize the semantic representation at the sending end, the channel transmission module, and the soft decision recovery module at the receiving end, thereby alleviating the training instability problems caused by discrete representation, channel disturbances, and joint optimization.

[0068] Step S501 is the first training phase, mainly involving semantic extraction and codebook generation. Firstly, under conditions of no channel disturbance, the image semantic extractor... Image semantic synthesizer Pre-training is performed. In this stage, the inverse mapping process of vector quantization first uses a hard decision and straight-through gradient (STE) strategy to achieve stable learning of discrete semantic representations.

[0069] During the forward propagation process, for each consecutive semantic feature The corresponding codeword index is selected based on the principle of minimum Euclidean distance. And obtain quantitative results. .because Since the operation is not differentiable, this invention introduces a STE strategy during backpropagation to approximate this discrete mapping process. In the backpropagation phase, the quantization operation is approximated as an identity mapping: .

[0070] This allows the gradient to be directly passed from the decoder to the semantic extractor. Correspondingly, the quantization process can be represented in the computational graph as follows: ; in, This indicates that gradient operations are stopped, and the values ​​are only retained in the forward propagation, not in the gradient calculation in the back propagation.

[0071] Through the above construction, hard decision quantization is still performed during forward propagation, thereby ensuring the discreteness and stability of the semantic index; while during backward propagation, the gradient pass-through mechanism bypasses the non-differentiable operation, enabling the encoder parameters to obtain effective gradient updates.

[0072] At this stage, reconstruction loss and codebook update loss can be used. And to impose constraints on the system by submitting losses: ; in This represents the image reconstruction loss, specifically including pixel-level loss. and cross-modal semantic alignment loss , and These represent the visual encoder and text encoder based on the CLIP model, respectively, which map visual semantics and text semantics to a unified cross-modal semantic space.

[0073] Step S502 is the second training phase, mainly involving the transmission of discrete index files. To further improve the robustness of discrete semantic index transmission in noisy channels, this invention introduces a JSCC transmission module and a soft-decision recovery mechanism at the receiving end in the second phase, specifically optimizing the channel transmission process of the discrete semantic index. During this process, the image semantic extractor... Image semantic synthesizer The parameters are frozen, and only the semantic-channel codec and semantic modulator are trained and updated to ensure the robust transmission of discrete indices in noisy channels. The receiver uses the soft decision mechanism proposed in this paper to recover continuous semantic features.

[0074] The total loss in the second phase is expressed as follows: ; in and These are the weighting coefficients for each loss function.

[0075] Specifically, The decision loss for the discrete index is used to constrain the probability distribution of the receiver's output, making it as concentrated as possible in the codebook dimension around the codewords corresponding to the true index, thereby improving index recovery accuracy. A supervision signal is constructed using the true discrete index from the transmitter, and its corresponding one-hot representation is denoted as... , Defined as: ; in, Represents the semantic embedding alignment loss, used to constrain the consistency between the continuous semantic representation recovered by the receiver and the corresponding codeword embedding at the sender: ; Finally, transmission loss Used to constrain intermediate representations during channel coding, modulation, and reception to improve the stability of channel transmission.

[0076] In one possible implementation of this application, in step S600, at the receiving end of the text semantic layer, the signal is equalized and then processed by the text semantic-channel decoder. By combining semantic priors, the text index is reconstructed. Finally, it passes through the text semantic synthesizer. Descriptive information that is semantically consistent with the original text .in, , , and These are the learnable parameters of the relevant neural network modules.

[0077] In one possible implementation of this application, step S700 includes: inputting image structural semantics and text semantic description as joint conditions into a cross-modal generation model, and completing the fusion and completion of latent spatial semantic information through cross-modal feature alignment and semantic consistency constraints to generate image information.

[0078] Specifically, continuous semantic features of the image semantic layer Input to training parameters An image semantic synthesizer is used to obtain a single-modal reconstructed image. .

[0079] Finally, cross-modal generative models utilize image structural semantics. With text semantic description As a joint conditional input, semantic information is fused and completed in the latent space through cross-modal feature alignment and semantic consistency constraints, thereby generating a reconstructed image with high fidelity in structure, semantics, and visual details. .

[0080] Compared with existing technologies, this method has the following advantages: By unifying the semantic information of different modalities into discrete semantic indices and combining this with a two-dimensional constellation point learning mechanism, the representation and transmission of semantic information of different modalities within a unified communication framework are achieved. This method does not rely on a complex cross-modal semantic space alignment process, which can reduce the complexity of system design and enhance the transmission adaptability and scalability of multimodal semantic information in complex wireless channel environments.

[0081] By setting up a discrete semantic index transmission mechanism based on soft decision, continuous semantic features are mapped to discrete semantic indices to improve representation efficiency. At the receiving end, the received index is inversely mapped and recovered based on probability distribution, which enables the semantic decoding and image generation process to have a certain semantic fault tolerance capability. This effectively reduces the discrete index recovery error caused by channel disturbances or bit errors, and improves the robustness of discrete semantic representation in the joint source channel transmission process.

[0082] By establishing an adaptive semantic transformation mechanism for imperfect channel state information, a pre-channel awareness module is introduced during transmission. Corresponding correction factors are generated based on the degree of channel estimation error to compensate for and restore semantic features, thereby reducing the impact of semantic feature shift on system performance when channel estimation errors exist. This method provides a concrete solution for robust semantic transmission under imperfect channel state information.

[0083] Another embodiment of this application discloses an image dual-modal transmission device oriented towards discrete semantic representation, such as... Figure 5 As shown, it includes: The building blocks are used to construct the transmission system; the transmission system includes an image semantic layer and a text semantic layer. The acquisition module is used to acquire image information, extract shallow basic semantic information of image information in the image semantic layer, extract deep semantic auxiliary information of image information in the text semantic layer, and convert both shallow basic semantic information and deep semantic auxiliary information into discrete semantic index sequences. The transformation module is used to acquire channel state information and use a semantic-channel encoder with an adaptive semantic transformation mechanism to transform the discrete semantic index sequence into an adaptive constellation point set. The transmitting module is used to generate a transmission signal by semantically encoding and adaptive constellation point modulation of the discrete semantic index sequence, and then transmit the transmission signal from the transmitting end to the receiving end. The image restoration module is used in the image semantic-channel decoder at the receiving end to restore the transmitted signal of the image semantic layer into a discrete semantic index sequence and continuous semantic features based on a soft-decision discrete semantic index restoration mechanism. The text recovery module is used to recover the transmitted signal of the text semantic layer into a discrete semantic index sequence in the text semantic-channel decoder at the receiving end; The fusion module is used to convert continuous semantics into image structural semantics using an image semantic synthesizer, convert discrete semantic index sequences of the text semantic layer into text semantic descriptions using a text semantic synthesizer, and fuse image structural semantics and text semantic descriptions to restore image information.

[0084] Figure 6 This diagram illustrates a more specific hardware structure of an electronic device provided in this embodiment. The device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, memory 1020, input / output interface 1030, and communication interface 1040 are interconnected internally via the bus 1050.

[0085] The processor 1010 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.

[0086] The memory 1020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 1020 can store the operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program code is stored in the memory 1020 and is called and executed by the processor 1010.

[0087] The input / output interface 1030 is used to connect input / output modules to realize information input and output. The input / output modules can be configured as components in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Input devices may include keyboards, mice, touch screens, microphones, various sensors, etc., and output devices may include displays, speakers, vibrators, indicator lights, etc.

[0088] The communication interface 1040 is used to connect the communication module (not shown in the figure) to enable communication between this device and other devices. The communication module can communicate via wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, radio (shortwave / ultra-shortwave) communication, satellite communication, data link communication, etc.).

[0089] Bus 1050 includes pathways for transmitting information between various components of the device, such as processor 1010, memory 1020, input / output interface 1030, and communication interface 1040.

[0090] It should be noted that although the above-described device only shows the processor 1010, memory 1020, input / output interface 1030, communication interface 1040, and bus 1050, in specific implementations, the device may also include other components necessary for normal operation. Furthermore, those skilled in the art will understand that the above-described device may only include the components necessary for implementing the embodiments described in this specification, and need not include all the components shown in the figures.

[0091] The electronic device described above is used to implement the corresponding image bimodal transmission method oriented towards discrete semantic representation in any of the foregoing embodiments, and has the beneficial effects of the corresponding method implementation, which will not be elaborated here.

[0092] Based on the same inventive concept, corresponding to any of the above-described embodiments, this application also provides a non-transitory computer-readable storage medium storing computer instructions for causing the computer to execute the image bimodal transmission method oriented towards discrete semantic representation as described in any of the above embodiments.

[0093] The computer-readable medium in this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.

[0094] The computer instructions stored in the storage medium of the above embodiments are used to cause the computer to execute the image bimodal transmission method oriented towards discrete semantic representation as described in any of the above embodiments, and have the beneficial effects of the corresponding method implementations, which will not be repeated here.

[0095] Those skilled in the art should understand that the discussion of any of the above embodiments is merely exemplary and is not intended to imply that the scope of this application (including the claims) is limited to these examples; this manner of description is merely for clarity, and those skilled in the art should consider the specification as a whole. Within the framework of this application, the above embodiments or the technical features of different embodiments can also be appropriately combined, the steps can be implemented in any order, and there are many other variations of different aspects of the embodiments of this application as described above, which are not provided in the details for the sake of brevity.

[0096] Additionally, to simplify the description and discussion, and to avoid obscuring the embodiments of this application, the well-known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided drawings. Furthermore, the apparatus may be shown in block diagram form to avoid obscuring the embodiments of this application, and this also takes into account the fact that the details of the implementation of these block diagram apparatuses are highly dependent on the platform on which the embodiments of this application will be implemented (i.e., these details should be entirely within the understanding of those skilled in the art). While specific details (e.g., circuits) are set forth to describe exemplary embodiments of this application, it will be apparent to those skilled in the art that the embodiments of this application can be implemented without these specific details or with variations thereof. Therefore, these descriptions should be considered illustrative rather than restrictive.

[0097] Although this application has been described in conjunction with specific embodiments thereof, many substitutions, modifications, and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may use the embodiments discussed.

[0098] The embodiments described herein are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc., made without departing from the spirit and principles of the embodiments described herein should be included within the protection scope of this application.

Claims

1. A method for image dual-modal transmission oriented towards discrete semantic representation, characterized in that, The method includes: Construct a transmission system; wherein the transmission system includes an image semantic layer and a text semantic layer; Obtain image information, extract shallow basic semantic information of the image information in the image semantic layer, extract deep semantic auxiliary information of the image information in the text semantic layer, and convert both the shallow basic semantic information and the deep semantic auxiliary information into discrete semantic index sequences; Channel state information is obtained, and the discrete semantic index sequence is transformed into an adaptive constellation point set using a semantic-channel encoder with an adaptive semantic transformation mechanism. The discrete semantic index sequence is semantically encoded and modulated by the adaptive constellation points to generate a transmission signal, which is then sent from the transmitting end to the receiving end. In the image semantic-channel decoder at the receiving end, a discrete semantic index recovery mechanism based on soft decision is used to recover the transmitted signal of the image semantic layer into the discrete semantic index sequence and continuous semantic features; In the text semantic-channel decoder at the receiving end, the transmission signal of the text semantic layer is restored to the discrete semantic index sequence; The continuous semantics are converted into image structural semantics using an image semantic synthesizer, and the discrete semantic index sequence of the text semantic layer is converted into a text semantic description using a text semantic synthesizer. The image structural semantics and the text semantic description are then fused to recover the image information.

2. The image dual-modal transmission method for discrete semantic representation according to claim 1, characterized in that, The step of extracting shallow basic semantic information of the image information in the image semantic layer and converting the shallow basic semantic information into a discrete semantic index sequence includes: High-dimensional semantic features of the image information are extracted using an image semantic extractor, and a vector quantization codebook of fixed codebook size is generated based on an exponential moving average algorithm. For each of the high-dimensional semantic features, the codeword with the smallest Euclidean distance is searched in the vector quantization codebook, and the codeword indexes corresponding to all the high-dimensional semantic features are set into the discrete semantic index sequence. The step of extracting deep semantic auxiliary information from the image information in the text semantic layer, and converting all of the deep semantic auxiliary information into discrete semantic index sequences, includes: By using a multimodal large model to perform semantic understanding of image information, structured text descriptions of images are generated. The image structured text description is converted into the discrete semantic index sequence using a text semantic extractor.

3. The image dual-modal transmission method for discrete semantic representation according to claim 1, characterized in that, The step of using a semantic-channel encoder equipped with the adaptive semantic transformation mechanism to transform the discrete semantic index sequence into an adaptive constellation point set includes: The discrete semantic index sequence is passed through the semantic-channel encoder to extract the latent image representation; wherein the latent image representation includes the discrete semantic index sequence, the learnable parameters of the semantic-channel encoder, and the channel state information; In the semantic-channel encoder, the signal-to-noise ratio is divided into intervals, such that each interval of the signal-to-noise ratio corresponds to a basic constellation set; For each of the aforementioned basic constellation sets, a correction factor is learned; The adaptive constellation point set is generated by superimposing the basic components of the basic constellation set and the correction amount.

4. The image dual-modal transmission method for discrete semantic representation according to claim 3, characterized in that, The semantic-channel encoder equipped with the adaptive semantic transformation mechanism includes: The semantic-channel encoder is configured with alternating layers of adaptive feature scaling and general feature processing. The discrete semantic index sequence, the learnable parameters of the semantic-channel encoder, and the channel state information are input as semantic features to the first layer of the adaptive feature scaling layer. After the semantic features are adaptively scaled by the adaptive feature scaling layer, the scaled semantic features are input to the first layer of the general feature processing layer for processing, and then input to the next layer of the adaptive feature scaling layer. The semantic features are processed in layers to generate the final semantic features.

5. The image dual-modal transmission method for discrete semantic representation according to claim 4, characterized in that, The semantic features are adaptively scaled using the adaptive feature scaling layer, including: The input semantic features are subjected to global average pooling to extract channel-level statistical information; By integrating the aforementioned channel-level statistical information, real-time signal-to-noise ratio, and channel estimation error variance, a joint conditional vector is generated through a multilayer perceptron embedding layer. Based on the joint conditional vector, complex-domain adaptive gating coefficients are generated, and amplitude and phase joint compensation corrections are performed on the real and imaginary parts of the semantic features, respectively. By combining the real-time signal-to-noise ratio to generate channel-level attention weights, the corrected semantic features are adaptively scaled according to the channel dimension to complete a semantic feature adaptive transformation.

6. The image dual-modal transmission method for discrete semantic representation according to claim 1, characterized in that, The soft-decision-based discrete semantic index recovery mechanism recovers the transmitted signal of the image semantic layer into the discrete semantic index sequence and continuous semantic features, including: The transmission signal of the image semantic layer is obtained by the receiving end, and the transmission signal is processed by channel estimation and equalization to obtain the recovered symbol. The distance metric between the recovered symbol and each codeword in the vector quantization codebook is calculated. By introducing a Gumbel random perturbation term and a temperature parameter, the distance metric is converted into a posterior probability distribution for each candidate codeword. The vector quantized codebook is weighted and combined using the probability distribution to obtain the recovered continuous semantic features.

7. The image dual-modal transmission method for discrete semantic representation according to claim 6, characterized in that, The soft-decision-based discrete semantic index recovery mechanism, before recovering the transmitted signal of the image semantic layer into the discrete semantic index sequence and continuous semantic features, includes: Under noise-free conditions, the image semantic extractor and the image semantic synthesizer are pre-trained; wherein, the forward propagation during the training process uses hard decision vector quantization to generate discrete indices, and the back propagation during the training process uses a gradient pass-through strategy to update parameters. The parameters of the image semantic extractor and the image semantic synthesizer, which have been pre-trained, are frozen. A joint source-channel coding and transmission module and a soft-decision recovery mechanism are introduced. Through joint constraints of index decision loss, semantic embedding alignment loss, and channel transmission loss, the learnable parameters of the image semantic-channel codec and the semantic modulator parameters are optimized.

8. The image dual-modal transmission method for discrete semantic representation according to claim 1, characterized in that, The process of fusing the image structural semantics and the text semantic description to recover the image information includes: The image structural semantics and the text semantic description are used as joint conditional inputs to the cross-modal generation model. Through cross-modal feature alignment and semantic consistency constraints, the latent spatial semantic information is fused and completed to generate the image information.

9. An image dual-modal transmission device for discrete semantic representation, characterized in that, The system includes: A construction module is used to construct a transmission system; wherein the transmission system includes an image semantic layer and a text semantic layer; The acquisition module is used to acquire image information, extract shallow basic semantic information of the image information in the image semantic layer, extract deep semantic auxiliary information of the image information in the text semantic layer, and convert both the shallow basic semantic information and the deep semantic auxiliary information into discrete semantic index sequences. The conversion module is used to acquire channel state information and, using a semantic-channel encoder equipped with an adaptive semantic transformation mechanism, convert the discrete semantic index sequence into an adaptive constellation point set. The transmitting module is used to generate a transmission signal by semantically encoding the discrete semantic index sequence and modulating it with the adaptive constellation points, and then transmit the transmission signal from the transmitting end to the receiving end. The image restoration module is used in the image semantic-channel decoder at the receiving end to restore the transmitted signal of the image semantic layer into the discrete semantic index sequence and continuous semantic features based on the soft-decision discrete semantic index restoration mechanism; The text recovery module is used to recover the transmitted signal of the text semantic layer into the discrete semantic index sequence in the text semantic-channel decoder at the receiving end; The fusion module is used to convert the continuous semantics into image structural semantics using an image semantic synthesizer, convert the discrete semantic index sequence of the text semantic layer into a text semantic description using a text semantic synthesizer, and fuse the image structural semantics and the text semantic description to restore the image information.

10. An electronic device, characterized in that, include: Processor and memory; The memory stores a computer program that, when executed by the processor, causes the processor to perform the steps of the image bimodal transmission method for discrete semantic representation as described in any one of claims 1 to 8.