Image transmission method and system based on source-channel joint coding and non-orthogonal multiple access

By extracting semantic features from images using a deep neural network model and allocating differentiated power, the problem of multi-user channel modeling and semantic awareness in the integration of JSCC and NOMA is solved, realizing efficient and robust semantic communication for multi-user image transmission.

CN121908018APending Publication Date: 2026-04-21TONGJI UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TONGJI UNIV
Filing Date
2025-12-30
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing research on the integration of JSCC and NOMA suffers from limitations in multi-user channel modeling capabilities, lack of semantic awareness, and defects in power allocation strategies, making it difficult to achieve optimal end-to-end coding performance and semantic communication protection in multi-user non-orthogonal environments.

Method used

The semantic feature sequence of the image is extracted by a semantic encoder of a deep neural network model and semantic importance weights are assigned. Different transmit powers are allocated in the NOMA transmission domain using a learnable power mapping function. Image transmission is optimized by combining semantic-signal-noise ratio metrics, and an end-to-end trainable deep learning framework is constructed.

Benefits of technology

It achieves joint optimization of semantic fidelity and spectral efficiency in image transmission under multi-user interference conditions, thereby improving perception quality and system bandwidth efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121908018A_ABST
    Figure CN121908018A_ABST
Patent Text Reader

Abstract

The invention provides an image transmission method and system based on source-channel joint coding and non-orthogonal multiple access, and the method comprises the steps: carrying out the source-channel joint coding of input images of a plurality of users through a semantic encoder of a deep neural network model, and generating a semantic feature sequence corresponding to the input image for each user, generating a semantic importance weight corresponding to each semantic feature in the semantic feature sequence; according to the semantic importance weight, distributing different transmitting powers used in a non-orthogonal multiple access transmission domain for each semantic feature to obtain semantic feature signals of different transmitting powers; superposing the semantic feature signals with different transmitting powers in a non-orthogonal multiple access power domain to form a composite signal; and reconstructing a reconstructed image of the corresponding user from the composite signal through a special decoder configured for each user, wherein the reconstructed image is a full-resolution image. According to the method, global joint optimization of semantic fidelity and spectral efficiency in multi-user image transmission is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of deep learning-driven information encoding and transmission technology, applicable to various information input formats such as text, voice, images, and video, and particularly to an image transmission method and system based on source-channel joint coding and non-orthogonal multiple access. Background Technology

[0002] Non-orthogonal multiple access (NOMA), a key technology in next-generation mobile communication systems, overcomes the resource utilization limitations of traditional orthogonal multiple access methods. Its core idea is to allow multiple users to transmit signals simultaneously on the same time-frequency resources, distinguishing users through power or code domain, and achieving signal separation at the receiving end using a serial interference cancellation (SIC) algorithm. This mechanism significantly improves the system's spectral efficiency and connectivity, and is widely considered one of the crucial supporting technologies for 5G and future 6G communication systems.

[0003] Source-channel joint coding (JSCC) is a communication paradigm that breaks through the traditional Shannon separation principle. Its core lies in jointly optimizing the source coding and channel coding processes, thereby achieving higher robustness and transmission efficiency under complex channel conditions. In recent years, with the development of deep learning, end-to-end JSCC models based on deep neural networks (such as DeepJSCC and DeepWiVe) can achieve performance superior to traditional separate coding schemes in low signal-to-noise ratio environments, providing new ideas and methods for the design of intelligent communication systems.

[0004] Despite the individual successes of JSCC and NOMA, and their integration widely considered a potential path to building next-generation efficient semantic communication systems, current integration research remains in its early stages and suffers from fundamental limitations. Some studies propose a heterogeneous semantic and bit-based communication framework, where an access point simultaneously sends semantic and bit streams to a semantically interested user and a bit-interested user, respectively; this proposed semi-NOMA outperforms both NOMA and OMA. Other studies introduce joint image compression and transmission schemes, where devices transmit their compressed image representations in a non-orthogonal manner, achieving significant improvements in image reconstruction quality compared to orthogonal transmission. These combinations aim to leverage the semantic coding capabilities of DeepJSCC with the efficient multiplexing framework of NOMA, potentially enabling more resilient and efficient communication systems.

[0005] However, the existing methods described above still face the following three main limitations: First, there are limitations in multi-user channel modeling capabilities. Existing JSCC frameworks are typically based on single-user or orthogonal transmission assumptions, making it difficult to effectively characterize the multi-user signal superposition and interference characteristics introduced by power domain multiplexing in NOMA scenarios. This simplified channel modeling limits the model's adaptability to inter-user power allocation and interference cancellation mechanisms, thus making it difficult to achieve optimal end-to-end coding performance in multi-user non-orthogonal environments.

[0006] Second, NOMA transmission lacks semantic awareness. Existing NOMA frameworks primarily focus on the superposition of physical layer signals and power allocation optimization during design, without fully considering the differences in the semantic importance of transmitted content. Because all data packets are treated equally, the system cannot provide differentiated protection for information with high semantic value or high perceptual sensitivity, thus limiting its potential applications in semantic communication and perception-driven networks.

[0007] Third, the semantic awareness deficiency of power allocation strategies. Traditional NOMA power allocation strategies mainly allocate resources based on user channel state or differences in received power, ignoring the semantic importance of different content features. This optimization approach, which only focuses on the physical layer signal-to-noise ratio (SNR), cannot effectively reflect the perceived value of information at the semantic layer. Summary of the Invention

[0008] This invention provides an image transmission method and system based on source-channel joint coding and nonorthogonal multiple access to solve the above-mentioned problems.

[0009] According to a first aspect of the present invention, an image transmission method based on source-channel joint coding and non-orthogonal multiple access is provided, comprising the following steps: performing source-channel joint coding on input images of multiple users using a semantic encoder of a deep neural network model, generating a semantic feature sequence corresponding to the input image of each user, and generating a semantic importance weight corresponding to each semantic feature in the semantic feature sequence; assigning different transmit powers for each semantic feature in the transmission domain of non-orthogonal multiple access according to the semantic importance weight, obtaining semantic feature signals with different transmit powers; and superimposing the semantic feature signals with different transmit powers in the power domain of non-orthogonal multiple access to form a composite signal for image reconstruction.

[0010] Optionally, the method further includes reconstructing a reconstructed image of the corresponding user from the composite signal using a dedicated decoder configured for each user, wherein the reconstructed image is a full-resolution image.

[0011] Optionally, the semantic encoder is a visual encoder based on the Transformer architecture, used to assign semantic importance weights to each semantic feature by calculating the attention score in the final self-attention block. The visual encoder is any one of the ViT, Swin Transformer, DeiT, PVT or MiT models.

[0012] Optionally, the semantic importance weights are mapped to transmission power through a semantic-to-power mapping module, so as to assign different transmission powers to each semantic feature:

[0013]

[0014] in, For a learnable power mapping function, This represents a semantic importance weight score. This represents a predefined attention allocation method, where L represents a sequence of L non-overlapping patches. Indicates the transmission power.

[0015] Optionally, the learnable power mapping function Specifically:

[0016] in, Different expressions exist in different contexts: Using a soft-mapped layer of a two-layer MLP with Sigmoid activation:

[0017] Mapping to M levels of discrete steps: .

[0018] Optionally, the method further includes using semantic signal-to-noise ratio (S-SNR) as a performance evaluation metric to quantify the energy concentration of semantically key components for users. k The semantic signal-to-noise ratio is defined as S-SNR. k :

[0019] in, and These represent the set of high-importance tokens and the set of low-importance tokens, respectively, based on semantic importance. The tokens correspond to the output of the semantic encoder. When performing semantic importance classification, firstly for each token... Calculate its semantic importance score Subsequently, a semantic importance score was assigned. Token Categorized to or ; Among them, semantic importance score The weights are determined based on the attention weight criterion, specifically as follows:

[0020] in, Indicates the first Layer The CLS token in the attention head points to the first Attention weights for each token, This refers to the selected set of layers and headers.

[0021] Optionally, the dedicated decoder is a decoder based on a convolutional neural network or a Transformer architecture.

[0022] Optionally, the deep neural network model is obtained through end-to-end training, and the joint loss function used during training includes a semantic distortion term and a power constraint term. The semantic distortion term is used to measure the difference between the reconstructed image and the input image.

[0023] According to a second aspect of the present invention, an image transmission system based on source-channel joint coding and non-orthogonal multiple access is provided, comprising: a semantic coding module, configured to perform source-channel joint coding on input images of multiple users through a semantic encoder of a deep neural network model, generate a semantic feature sequence corresponding to the input image of each user, and generate semantic importance weights corresponding to each semantic feature in the semantic feature sequence; a semantic-to-power mapping module, configured to assign different transmit powers used in the transmission domain of non-orthogonal multiple access to each semantic feature according to the semantic importance weights, thereby obtaining semantic feature signals with different transmit powers; and a superposition module, configured to superimpose the semantic feature signals with different transmit powers in the power domain of non-orthogonal multiple access to form a composite signal for image reconstruction.

[0024] Optionally, the system further includes a decoding module for reconstructing a reconstructed image of the corresponding user from the composite signal using a dedicated decoder configured for each user, the reconstructed image being a full-resolution image.

[0025] According to a third aspect of the present invention, an electronic device is provided, including a processor and a memory storing a program. The program includes instructions that, when executed by the processor, cause the processor to perform the steps performed by the method of the first aspect described above.

[0026] According to a fourth aspect of the present invention, a computer storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the method of the first aspect described above.

[0027] Compared to existing technologies, this invention proposes the first semantically aware DeepJSCC framework integrated with NOMA. By utilizing a Transformer-based semantic encoder, it effectively extracts hierarchical semantic representations from input images and assigns different power levels to them in the NOMA transmission domain according to their importance, achieving intelligent mapping from semantic importance to power. This invention also introduces a novel semantic signal-to-noise ratio metric to measure the perceptual resilience of transmitted semantic content. By developing an end-to-end trainable architecture that simultaneously optimizes semantic distortion and power constraints, perceptual quality and system bandwidth efficiency are significantly improved in image transmission. Attached Figure Description

[0028] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings.

[0029] Figure 1 This is a flowchart of the steps of the method of the present invention.

[0030] Figure 2 This is a schematic diagram of the system architecture of the present invention. Detailed Implementation

[0031] To provide a clearer understanding of the technical features, objectives, and effects of the embodiments of the present invention, specific implementation methods of the embodiments of the present invention will now be described with reference to the accompanying drawings.

[0032] In this document, “exemplary” means “serving as an example, illustration or description”, and any illustrations or implementations described herein as “exemplary” should not be construed as a more preferred or advantageous technical solution.

[0033] To enable those skilled in the art to better understand the technical solutions in the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art should fall within the protection scope of the present invention.

[0034] The specific implementation of the embodiments of the present invention will be further described below with reference to the accompanying drawings.

[0035] See Figure 1 The present invention provides an image transmission method based on source-channel joint coding and nonorthogonal multiple access, comprising the following steps: Step S1: Perform source-channel joint coding on the input images of multiple users through the semantic encoder of the deep neural network model, generate a semantic feature sequence corresponding to the input image of each user, and generate the semantic importance weight of each semantic feature in the semantic feature sequence. Step S2: Based on the semantic importance weights, assign different transmit powers to each semantic feature in the transmission domain of non-orthogonal multiple access to obtain semantic feature signals with different transmit powers; Step S3: Superimpose the semantic feature signals with different transmission powers in the power domain of non-orthogonal multiple access to form a composite signal for image reconstruction.

[0036] Optionally, the method further includes reconstructing a reconstructed image of the corresponding user from the composite signal using a dedicated decoder configured for each user, wherein the reconstructed image is a full-resolution image.

[0037] Optionally, the semantic encoder is a visual encoder based on the Transformer architecture, used to assign semantic importance weights to each semantic feature by calculating the attention score in the final self-attention block. The visual encoder is any one of the ViT, Swin Transformer, DeiT, PVT or MiT models.

[0038] Optionally, the semantic importance weights are mapped to transmission power through a semantic-to-power mapping module, so as to assign different transmission powers to each semantic feature:

[0039]

[0040] in, For a learnable power mapping function, This represents a semantic importance weight score. This represents a predefined attention allocation method, where L represents a sequence of L non-overlapping patches. Indicates the transmission power.

[0041] Optionally, the learnable power mapping function Specifically:

[0042] in, Different expressions exist in different contexts: Using a soft-mapped layer of a two-layer MLP with Sigmoid activation:

[0043] Mapping to M levels of discrete steps: .

[0044] Optionally, the method further includes using semantic signal-to-noise ratio (S-SNR) as a performance evaluation metric to quantify the energy concentration of semantically key components for users. k The semantic signal-to-noise ratio is defined as S-SNR. k :

[0045] in, and These represent the set of high-importance tokens and the set of low-importance tokens, respectively, based on semantic importance. The tokens correspond to the output of the semantic encoder. When performing semantic importance classification, firstly for each token... Calculate its semantic importance score Subsequently, a semantic importance score was assigned. Token Categorized to or ; Among them, semantic importance score The weights are determined based on the attention weight criterion, specifically as follows:

[0046] in, Indicates the first Layer The CLS token in the attention head points to the first Attention weights for each token, This refers to the selected set of layers and headers.

[0047] Optionally, the dedicated decoder is a decoder based on a convolutional neural network or a Transformer architecture.

[0048] Optionally, the deep neural network model is obtained through end-to-end training, and the joint loss function used during training includes a semantic distortion term and a power constraint term. The semantic distortion term is used to measure the difference between the reconstructed image and the input image.

[0049] As another example, the present invention also provides an image transmission system based on source-channel joint coding and nonorthogonal multiple access, comprising: The semantic encoding module is used to perform source-channel joint encoding on the input images of multiple users through the semantic encoder of the deep neural network model, generate a semantic feature sequence corresponding to the input image of each user, and generate the semantic importance weight of each semantic feature in the semantic feature sequence. The semantic-to-power mapping module is used to assign different transmit powers used in the transmission domain of non-orthogonal multiple access to each semantic feature according to the semantic importance weight, so as to obtain semantic feature signals with different transmit powers. The superposition module is used to superimpose the semantic feature signals with different transmission powers in the power domain of non-orthogonal multiple access to form a composite signal for image reconstruction.

[0050] Optionally, the system further includes a decoding module for reconstructing a reconstructed image of the corresponding user from the composite signal using a dedicated decoder configured for each user, the reconstructed image being a full-resolution image.

[0051] The system in this embodiment is used to implement the corresponding methods in the foregoing multiple method embodiments and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0052] Specifically, the solution of the present invention is further described with reference to the following examples: This invention provides an image transmission method and system based on source-channel joint coding and nonorthogonal multiple access, using the semantically conscious DEEPJSCC framework of Noma ensemble, and introduces a novel metric: semantic-signal-noise ratio. The specific implementation includes the following technical solutions: A. Semantic signal-to-noise ratio This invention defines a new metric, semantic signal-to-noise ratio (S-SNR), as a performance evaluation indicator to quantify the energy concentration of semantically key components:

[0053] in, and These represent the sets of high-importance tokens and low-importance tokens, respectively, based on semantic importance. Each token corresponds to the output of the semantic encoder. This metric reflects the semantically meaningful content of the system under noisy and multi-user conditions.

[0054] For example, the semantic features output by the semantic encoder correspond to a set of tokens. For each token Calculate its semantic importance score The tokens are then divided into a high-importance set based on the semantic importance score. With low importance set .

[0055] Among them, semantic importance score The attention weight criterion can be used to determine this: Based on ViT's self-attention, let

[0056] in Indicates the first Layer The CLS token in the attention head points to the first Attention weights for each token, This refers to the selected set of layers and heads. Intuitively, CLS is more "concerned" about the token semantics that are more critical.

[0057] B. Semantic relevance In semantic relevance, an attention mechanism is introduced, which is applied to each token. Semantic importance weights were assigned to scores. This indicates its semantic relevance:

[0058] in, This represents a pre-defined attention allocation method. These semantic features are then divided into M priority levels (e.g., high, medium, low), and each level is mapped to a different NOMA power level. The mapping power of the token is:

[0059] in, It is a learnable or rule-based power mapping function.

[0060] C. System Framework The overall system architecture is as follows Figure 2 As shown, the framework includes a transformer encoder, semantic importance mapping, NOMA overlay, channel model, and per-user decoder. The entire framework minimizes semantic distortion between encoding and decoding while satisfying NOMA decoding constraints and the total power budget.

[0061] D. Deep Learning Strategies This invention proposes a fully distinguishable, end-to-end trainable deep neural architecture to solve the cross-layer optimization problem, mainly including: (1) Transformer-based semantic encoder: A visual encoder (ViT) is used. The encoder computes the context embedding and assigns it to each token based on the attention score in the final self-attention block. Assign semantic importance weights to scores These attention weights reflect the relative semantic contribution of each token and are used in the power mapping phase.

[0062] (2) Semantic-to-Power Mapping Module: This module maps semantic importance to power levels in the NOMA domain, controlling the transmission method of each semantic tag. This invention proposes a learnable power mapping function. a. A soft-mapped layer of a two-layer MLP with Sigmoid activation:

[0063] b. Discrete steps up to M levels

[0064] (3) Decoder architecture and semantic reconstruction: Each user k is assigned a decoder It is implemented as a multi-layer CNN or Transformer decoder, which reconstructs the received semantic representation into a full-resolution image. The decoder is trained to minimize .

[0065] It should be understood that the visual encoder of the present invention is any one of ViT (Vision Transformer), Swin Transformer, DeiT, PVT or MiT model. ViT is preferred for illustrative purposes here, but is not intended to limit the embodiments of the present invention.

[0066] It should also be understood that the input information format of the present invention is not limited to images, but can also be text, voice, video and other types of information input formats. Here, the present invention is preferably illustrated by images, but is not intended to limit the embodiments of the present invention.

[0067] Example: This invention is applied to wireless image transmission, specifically multi-user image transmission over a NOMA channel. It aims to address the fundamental disconnect between semantic communication and multiple access systems in terms of optimization objectives. For the first time, it constructs an end-to-end deep learning framework capable of simultaneously "understanding content importance" and "intelligently allocating resources accordingly," thereby achieving global joint optimization of semantic fidelity and spectral efficiency in multi-user image transmission. Specific details are as follows: Step 1. Input image encoding A transformer-based semantic encoder is used to extract high-level semantic representations and attention-based importance scores from the input image. A visual transformer (ViT) is employed as the encoder. Input image A sequence labeled as L non-overlapping patches is linearly embedded into the feature vector. The encoder computes the context embedding and assigns it to each tag based on the attention score in the final self-attention block. Assign importance weights .

[0068] Step 2. Transmit by mapping semantic importance to power levels. This invention employs a semantic-to-power mapping module to map semantic importance to power levels in the NOMA domain, controlling the transmission method of each semantic tag. A learnable power mapping function is proposed in this invention.

[0069] in Different expressions exist in different contexts: a. Using a soft-mapped layer of a two-layer MLP with Sigmoid activation:

[0070] b. Discrete steps mapped to M levels

[0071] Step 3. Image Reconstruction An image is reconstructed from the output of the noisy channel using a set of user-specific decoders. Each user k is assigned one decoder. It is implemented as a multi-layer CNN or Transformer decoder, which reconstructs the received semantic representation into a full-resolution image. .

[0072] To verify the effectiveness of this invention, the CIFAR-10 dataset was used, with images resized to 64×64 and normalized to [0, 1]. The ransformer encoder employed 6 layers and 8 attention heads, with each user having an independent CNN-based head. The entire system was trained using the Adam optimizer with a batch size of 128 for 2000 epochs.

[0073] This invention simulates K = 2 Five users, each experiencing an independent AWGN channel with different average SNR levels, were tested. The channel remained constant during each transmission. The deep neural network model of this invention was compared with DeepJSCC-NOMA, DeepJSCC-OMA, and BPG + LDPC. Simulation results show that compared with existing related schemes, the sensing quality and bandwidth efficiency of this invention are significantly improved.

[0074] In particular, in a two-user scenario, the semantically important adaptive power allocation strategy provides stronger protection for key content in the image compared to a fixed power allocation scheme, thus bringing additional performance gains.

[0075] Specifically, the present invention achieves its beneficial effects through the following technical solutions: First, the framework proposed in this invention (corresponding to steps S1 to S4) breaks through the traditional paradigm of separate design of JSCC and NOMA. Step S1 uses a semantic encoder of a deep neural network model to perform joint source-channel coding, which not only extracts the semantic feature sequence of the input image for each user, but also generates weights representing the perceived importance of each feature. Step S2 then assigns differentiated transmit power to each semantic feature in the NOMA transmission domain based on these semantic importance weights, realizing intelligent mapping from semantic importance to power. Steps S3 and S4 complete the signal superposition in the NOMA domain and the targeted reconstruction for each user. This complete process, consisting of a series of steps, enables semantic features to be hierarchically mapped to different power levels according to their importance, thereby achieving joint optimization of semantic fidelity and multi-user access efficiency at the system level for the first time.

[0076] Secondly, this invention introduces a novel semantic signal-to-noise ratio metric, which can accurately quantify the degree of energy concentration and protection strength of key semantic components under noise and multi-user interference conditions, providing an objective and effective standard for evaluating the "perceptual resilience" of semantic transmission.

[0077] Ultimately, the technology of this invention is implemented and optimized through an end-to-end trainable deep neural network model architecture. This architecture is jointly trained with the goal of simultaneously minimizing semantic distortion and power constraints.

[0078] In summary, this invention solves three core problems—encoding mismatch under multi-user interference, semantic blindness in resource allocation, and fragmented optimization objectives—by creating a semantically aware DeepJSCC-NOMA integration framework, providing a core technical solution for realizing the next generation of intelligent, efficient, and robust multi-user image transmission systems.

[0079] As another example, embodiments of the present invention also provide an electronic device, which is an example of a hardware device that can be applied to various aspects of the present invention. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0080] The electronic device may include a processor, a communications interface, memory, and a communications bus.

[0081] The processor, communication interface, and memory communicate with each other via a communication bus. The communication interface is used to communicate with other electronic devices or servers.

[0082] The processor is used to execute programs, specifically the relevant steps in the above method embodiments.

[0083] Specifically, the program may include program code, which includes computer operation instructions.

[0084] The processor may be a CPU, an Application Specific Integrated Circuit (ASIC), or one or more integrated circuits configured to implement embodiments of the present invention. The one or more processors included in a smart device may be of the same type, such as one or more CPUs; or they may be of different types, such as one or more CPUs and one or more ASICs.

[0085] Memory is used to store programs. Memory may include high-speed RAM, and may also include non-volatile memory, such as at least one disk drive.

[0086] The program, when executed by a processor, is used to cause an electronic device to perform the method of the present invention.

[0087] Furthermore, the specific implementation of each step in the program can be found in the corresponding descriptions of the steps and units in the above method embodiments, and will not be repeated here. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the devices and modules described above can be referred to the corresponding process descriptions in the foregoing method embodiments, and will not be repeated here.

[0088] An exemplary embodiment of the present invention also provides a computer storage medium storing a computer program, wherein when the computer program is executed by a processor, it implements the methods of the various embodiments of the present invention. The corresponding process descriptions in the foregoing method embodiments can be referred to, and will not be repeated here.

[0089] The methods described above according to embodiments of the present invention can be implemented in hardware, firmware, or as software or computer code that can be stored in a recording medium (such as a CD-ROM, RAM, floppy disk, hard disk, or magneto-optical disk), or as computer code originally stored on a remote recording medium or a non-transitory machine-readable medium and subsequently stored on a local recording medium, downloaded via a network. Thus, the methods described herein can be processed by software stored on a recording medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as an ASIC or FPGA). It is understood that the computer, processor, microprocessor controller, or programmable hardware includes storage components (e.g., RAM, ROM, flash memory, etc.) capable of storing or receiving software or computer code, which, when accessed and executed by the computer, processor, or hardware, implements the methods described herein. Furthermore, when a general-purpose computer accesses code used to implement the methods shown herein, the execution of the code transforms the general-purpose computer into a dedicated computer for executing the methods shown herein.

[0090] Specific embodiments of the invention have now been described. Other embodiments are within the scope of the appended claims. In some cases, the actions described in the claims can be performed in a different order and still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing can be advantageous.

[0091] It should be understood that although this specification is described according to various embodiments, not every embodiment contains only one independent technical solution. This way of describing the specification is only for clarity. Those skilled in the art should regard the specification as a whole. The technical solutions in each embodiment can also be appropriately combined to form other implementation methods that can be understood by those skilled in the art.

[0092] Finally, it should be noted that the above embodiments are only used to illustrate the embodiments of the present invention, and are not intended to limit the embodiments of the present invention. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the embodiments of the present invention. Therefore, all equivalent technical solutions also fall within the scope of the embodiments of the present invention, and the patent protection scope of the embodiments of the present invention should be defined by the claims.

Claims

1. An image transmission method based on source-channel joint coding and non-orthogonal multiple access, characterized in that, Includes the following steps: The semantic encoder of the deep neural network model performs source-channel joint coding on the input images of multiple users, generates a semantic feature sequence corresponding to the input image of each user, and generates the semantic importance weight of each semantic feature in the semantic feature sequence. Based on the semantic importance weights, different transmit powers are assigned to each semantic feature in the transmission domain of non-orthogonal multiple access, resulting in semantic feature signals with different transmit powers. The semantic feature signals with different transmission powers are superimposed in the power domain of non-orthogonal multiple access to form a composite signal for image reconstruction.

2. The method according to claim 1, characterized in that, The method further includes: A reconstructed image of the corresponding user is reconstructed from the composite signal using a dedicated decoder configured for each user; the reconstructed image is a full-resolution image.

3. The method according to claim 1, characterized in that, The semantic encoder is a visual encoder based on the Transformer architecture, which is used to assign semantic importance weights to each semantic feature by calculating the attention score in the final self-attention block. The visual encoder can be any one of the ViT, Swin Transformer, DeiT, PVT or MiT models.

4. The method according to claim 1, characterized in that, The semantic-to-power mapping module maps semantic importance weights to transmission power, thereby enabling the allocation of different transmission powers to each semantic feature: in, For a learnable power mapping function, This represents a semantic importance weight score. This represents a predefined attention allocation method, where L represents a sequence of L non-overlapping patches. Indicates the transmission power.

5. The method according to claim 4, characterized in that, The learnable power mapping function Specifically: in, Different expressions exist in different contexts: Using a soft-mapped layer of a two-layer MLP with Sigmoid activation: Mapping to M levels of discrete steps: 。 6. The method according to claim 1, characterized in that, The method also includes using semantic signal-to-noise ratio (S-SNR) as a performance evaluation metric to quantify the energy concentration of semantically key components for users. k The semantic signal-to-noise ratio is defined as S-SNR. k : in, and These represent the set of high-importance tokens and the set of low-importance tokens, respectively, based on semantic importance. The tokens correspond to the output of the semantic encoder. When performing semantic importance classification, firstly for each token... Calculate its semantic importance score Subsequently, a semantic importance score was assigned. Token Categorized to or ; Among them, semantic importance score The weights are determined based on the attention weight criterion, specifically as follows: in, Indicates the first Layer The CLS token in the attention head points to the first Attention weights for each token, This refers to the selected set of layers and headers.

7. The method according to claim 2, characterized in that, The dedicated decoder is a decoder based on a convolutional neural network or Transformer architecture.

8. The method according to claim 2, characterized in that, The deep neural network model is obtained through end-to-end training. The joint loss function used during training includes a semantic distortion term and a power constraint term. The semantic distortion term is used to measure the difference between the reconstructed image and the input image.

9. An image transmission system based on source-channel joint coding and non-orthogonal multiple access, characterized in that, include: The semantic encoding module is used to perform source-channel joint encoding on the input images of multiple users through the semantic encoder of the deep neural network model, generate a semantic feature sequence corresponding to the input image of each user, and generate the semantic importance weight of each semantic feature in the semantic feature sequence. The semantic-to-power mapping module is used to assign different transmit powers used in the transmission domain of non-orthogonal multiple access to each semantic feature according to the semantic importance weight, so as to obtain semantic feature signals with different transmit powers. The superposition module is used to superimpose the semantic feature signals with different transmission powers in the power domain of non-orthogonal multiple access to form a composite signal for image reconstruction.

10. The system according to claim 9, characterized in that, The system also includes: The decoding module is used to reconstruct the reconstructed image of the corresponding user from the composite signal using a dedicated decoder configured for each user. The reconstructed image is a full-resolution image.