Semantic communication image processing method, system, device, and storage medium

By extracting and quantizing the semantic feature vectors of the source image, the semantic feature signals of multiple users are encoded and transmitted in separate subspaces, which solves the problems of limited user capacity and high base station complexity in orthogonal multiple access technology and improves the processing efficiency of the communication system.

CN118869402BActive Publication Date: 2026-02-03PENG CHENG LAB
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410855286.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-27
Publication Date
2026-02-03
Estimated Expiration
2044-06-27

AI Technical Summary

Technical Problem

Existing orthogonal multiple access technologies have user capacity approaching the Shannon limit, high base station receiver complexity, and severe inter-user interference, resulting in limited communication system capacity and low efficiency.

Method used

The semantic feature vector of the source image is extracted by the feature extraction unit, quantized into discrete semantic feature vectors and modulated. Channel equalization and semantic reasoning are performed by the semantic reasoning model of the base station, so that the semantic feature signals of multiple users can be encoded and transmitted in separate semantic feature subspaces, reducing the processing difficulty of the base station.

Benefits of technology

This reduces the processing difficulty of the total received signals for the base station, improves the processing efficiency and communication transmission efficiency of the base station, and solves the problem of limited user capacity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118869402B_ABST
    Figure CN118869402B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a semantic communication image processing method, system, device and storage medium, relating to the technical field of communication. The method extracts a semantic feature vector of a source image, quantizes and modulates the semantic feature vector to obtain a semantic feature signal, obtains a transmission signal according to a transmission power, channel state information and the semantic feature signal, and transmits the transmission signal to a base station, so that the base station receives a total received signal formed by a plurality of transmission signals, performs channel equalization on the total received signal, obtains an estimated signal of each transmission end, and inputs the estimated signal into a semantic inference model corresponding to the transmission end to perform semantic inference, and obtains target image processing data corresponding to the source image. By separating the feature domain into a plurality of semantic feature subspaces, the semantic feature signals of a plurality of users are encoded and transmitted in the separated semantic feature subspaces, are approximately orthogonal to each other, the processing difficulty of the base station for the total received signal is reduced, the processing efficiency of the base station is improved, and the communication transmission efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of communication technology, and in particular to semantic communication image processing methods, systems, devices and storage media. Background Technology

[0002] Orthogonal Multiple Access (OMA) is a multi-user access method in wireless communication systems. It considers the orthogonal allocation of physical resources in the time, frequency, spatial, and power domains to distinguish multiple users, ensuring the orthogonality of signals between different users and effectively avoiding interference between multiple users. However, the user capacity of each orthogonal channel in current OMA technology is approaching the Shannon limit, limiting the number of users that can access the system.

[0003] In related technologies, complex algorithms are set up in the base station to distinguish the signals of different users. In this case, the base station receiver is highly complex, and the increase in the number of users will exponentially increase the signal interference between users, further increasing the processing difficulty of the receiver. Summary of the Invention

[0004] The main objective of this application is to propose a semantic communication image processing method, system, device, and storage medium, which can reduce the processing difficulty of base stations, improve the processing efficiency of base stations, and improve communication transmission efficiency.

[0005] To achieve the above objectives, a first aspect of this application proposes a semantic communication image processing method applied at a transmitting end, the method comprising:

[0006] The semantic feature vector of the source image is extracted using a feature extraction unit;

[0007] The semantic feature vector is quantized to obtain a discrete semantic feature vector, and the discrete semantic feature vector is modulated using a digital modulation unit to obtain a semantic feature signal;

[0008] The transmitted signal is obtained based on the transmit power, channel state information, and semantic feature signal corresponding to the transmitter.

[0009] The transmitted signal is transmitted to the base station so that the base station receives a total received signal composed of the transmitted signals from multiple different transmitters. Channel equalization is performed on the total received signal to obtain an estimated signal corresponding to each transmitter. The estimated signal is then input into a pre-trained semantic reasoning model corresponding to the transmitter for semantic reasoning to obtain the target image processing data corresponding to the source image.

[0010] In some embodiments, quantizing the semantic feature vector to obtain a discrete semantic feature vector includes:

[0011] The semantic feature vector is input into the linear layer to obtain the linear vector corresponding to the number of quantization bits;

[0012] Values ​​greater than or equal to zero in the linear vector are quantized to one, and otherwise quantized to zero, to obtain the discrete semantic feature vector.

[0013] In some embodiments, obtaining the transmitted signal based on the transmit power corresponding to the transmitting end, channel state information, and the semantic feature signal includes:

[0014] Obtain the conjugate transpose matrix of the channel state information;

[0015] The transmission parameters corresponding to each transmitter are obtained based on the transmission power and the conjugate transpose matrix.

[0016] The transmitted signal is obtained by multiplying the transmitted parameters and the semantic feature signal, and the total received signal is obtained by accumulating the transmitted signal and the noise vector.

[0017] In some embodiments, the source image includes an image to be classified or an image to be reconstructed;

[0018] When the source image is an image to be classified, the semantic reasoning model is a neural network model, and the target image processing data is the classification result of the image to be classified.

[0019] When the source image is an image to be reconstructed, the feature extraction unit is an encoder unit, the semantic reasoning model is a decoder unit, and the target image processing data is the reconstruction result corresponding to the image to be reconstructed.

[0020] In some embodiments, when the feature extraction unit is an encoder unit, the step of extracting the semantic feature vector of the source image using the feature extraction unit includes:

[0021] After performing a first downsampling and self-attention processing on the source image using the first coding block, a first feature map is obtained, wherein the width and height of the first feature map are both half of the source image.

[0022] After performing a second downsampling and self-attention processing on the first feature map using the second coding block, a second feature map is obtained, wherein the width and height of the second feature map are both half of the first feature map.

[0023] The semantic feature vector is obtained by performing a third downsampling and self-attention processing on the second feature map using a third coding block.

[0024] To achieve the above objectives, a second aspect of this application proposes a semantic communication image processing method applied to a base station, the method comprising:

[0025] A total received signal is obtained based on a transmission signal from at least one of the transmitting ends, wherein the transmission signal is obtained according to the semantic communication image processing method according to any one of the first aspects;

[0026] Channel equalization is performed on the total received signal to obtain the estimated signal corresponding to each transmitter;

[0027] The estimated signal is input into a pre-trained semantic reasoning model corresponding to the transmitter to perform semantic reasoning, thereby obtaining the target image processing data corresponding to the source image.

[0028] In some embodiments, performing channel equalization on the total received signal to obtain an estimated signal corresponding to each transmitter includes:

[0029] The estimated coefficients are obtained based on the channel state information and the conjugate transpose matrix;

[0030] The estimated signal is obtained by multiplying the estimated coefficients and the total received signal.

[0031] In some embodiments, when the semantic reasoning model is a neural network model, before inputting the estimated signal into a pre-trained semantic reasoning model corresponding to the transmitting end for semantic reasoning, the method further includes:

[0032] A training dataset is obtained, which includes a total sample signal and image processing labels. The total sample signal is obtained by the transmitted sample signal generated by each of the transmitting ends based on the sample image.

[0033] The estimated sample signal corresponding to each of the transmitting ends is obtained based on the total sample signal;

[0034] The semantic reasoning model is jointly trained using the estimated sample signals to obtain image prediction results;

[0035] A first loss value is calculated based on the image processing label and the image prediction result; a second loss value is calculated based on the transmitted sample signal and the total sample signal; and a third loss value is calculated based on the transmitted sample signal and the sample image.

[0036] The total loss value is obtained based on the first loss value, the second loss value, and the third loss value. The semantic reasoning model is then weighted based on the total loss value to obtain the trained semantic reasoning model corresponding to each of the transmitting ends.

[0037] In some embodiments, when the semantic reasoning model is a decoder unit, before inputting the estimated signal into a pre-trained semantic reasoning model corresponding to the transmitter for semantic reasoning, the method further includes:

[0038] A training dataset is obtained, which includes a total sample signal and image processing labels. The total sample signal is obtained by the transmitted sample signal generated by each of the transmitting ends based on the sample image.

[0039] The estimated sample signal corresponding to each of the transmitting ends is obtained based on the total sample signal;

[0040] The semantic reasoning model is jointly trained using the estimated sample signals to obtain image prediction results;

[0041] The total loss value is calculated based on the image processing label and the image prediction result;

[0042] The semantic reasoning model is weighted based on the total loss value to obtain a trained semantic reasoning model corresponding to each of the transmitting ends.

[0043] To achieve the above objectives, a third aspect of this application proposes a semantic communication image processing system, comprising:

[0044] Multiple transmitters, wherein the transmitters are used to obtain corresponding transmission signals according to the semantic communication image processing method according to any one of the first aspects;

[0045] A base station, wherein the base station is configured to acquire the total received signal obtained based on the transmitted signal, and to obtain target image processing data corresponding to each of the transmitting ends according to the semantic communication image processing method described in any of the second aspects.

[0046] To achieve the above objectives, a fourth aspect of the present application provides an electronic device, the electronic device including a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method described in the first or second aspect.

[0047] To achieve the above objectives, a fifth aspect of the present application provides a storage medium that stores a computer program, which, when executed by a processor, implements the method described in the first or second aspect.

[0048] The semantic communication image processing method, system, device, and storage medium proposed in this application utilize a feature extraction unit to extract semantic feature vectors from a source image, quantize these semantic feature vectors to obtain discrete semantic feature vectors, and modulate these discrete semantic feature vectors using a digital modulation unit to obtain semantic feature signals. A transmission signal is obtained based on the transmission power, channel state information, and semantic feature signals of the transmitting end, and is transmitted to a base station. This allows the base station to receive a total received signal composed of transmission signals from multiple different transmitting ends. Channel equalization is performed on the total received signal to obtain an estimated signal corresponding to each transmitting end. The estimated signal is then input into a pre-trained semantic inference model corresponding to the transmitting end for semantic inference, resulting in target image processing data corresponding to the source image. This application's embodiments extract semantic feature signals from the source image through feature extraction units corresponding to different transmitting ends, separating the feature domain into multiple semantic feature subspaces. This enables the encoding and transmission of semantic feature signals from multiple users within these separated semantic feature subspaces. The discrete semantic feature vectors of different users constitute the total received signal, but they are approximately orthogonal to each other, thereby reducing the processing difficulty of the total received signal for the base station, improving the base station's processing efficiency, and enhancing communication transmission efficiency. Attached Figure Description

[0049] Figure 1 This is a schematic diagram of the semantic communication image processing system provided in the embodiments of this application.

[0050] Figure 2 This is another schematic diagram of the semantic communication image processing system provided in the embodiments of this application.

[0051] Figure 3 This is an optional flowchart of the semantic communication image processing method provided in the embodiments of this application.

[0052] Figure 4 This is a flowchart illustrating the extraction of semantic feature vectors from a source image using a feature extraction unit, as provided in an embodiment of this application.

[0053] Figure 5 This is a schematic diagram of the encoder unit in an embodiment of this application.

[0054] Figure 6 This is a flowchart provided in an embodiment of the present application for obtaining a transmitted signal based on the transmit power, channel state information and semantic feature signals corresponding to the transmitter.

[0055] Figure 7 This is an optional flowchart of the semantic communication image processing method provided in the embodiments of this application.

[0056] Figure 8 This is a schematic diagram illustrating the training process of the semantic reasoning model provided in the embodiments of this application.

[0057] Figure 9 This is a schematic diagram of the decoder unit provided in the embodiments of this application.

[0058] Figure 10 This is another schematic diagram of the semantic communication image processing system provided in the embodiments of this application.

[0059] Figure 11 This is another schematic diagram illustrating the training process of the semantic reasoning model provided in the embodiments of this application.

[0060] Figure 12 This is a schematic diagram illustrating the classification performance of the image to be classified provided in an embodiment of this application.

[0061] Figure 13 This is another schematic diagram illustrating the classification performance of the image to be classified provided in the embodiments of this application.

[0062] Figure 14 This is a schematic diagram illustrating the reconstruction performance of the image to be reconstructed provided in an embodiment of this application.

[0063] Figure 15 This is another schematic diagram illustrating the reconstruction performance of the image to be reconstructed provided in an embodiment of this application.

[0064] Figure 16 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0065] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0066] It should be noted that although functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart.

[0067] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0068] First, let's analyze some of the terms used in this application:

[0069] Artificial intelligence (AI) is a new branch of computer science that studies, develops, and applies theories, methods, technologies, and systems to simulate, extend, and expand human intelligence. It aims to understand the essence of intelligence and produce intelligent machines that can react in a way similar to human intelligence. Research in this field includes robotics, speech recognition, image recognition, natural language processing, and expert systems. AI can simulate the information processes of human consciousness and thought. Furthermore, AI utilizes digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceiving the environment, acquiring knowledge, and using that knowledge to achieve optimal results.

[0070] With the rapid development of large-scale wireless communication and the increasing demand for intelligent processing, various emerging intelligent services based on wireless communication technologies are emerging one after another, such as augmented reality, virtual reality, holographic communication, and intelligent connected vehicles. This has led to a significant increase in the number of connected devices worldwide, which has brought new challenges to communication technology. From 1G to 5G, mobile communication systems, guided by Shannon's information theory, have seen a significant increase in transmission rates, but they also face problems such as system capacity gradually approaching the Shannon limit, significant energy consumption, and a scarcity of high-quality spectrum resources.

[0071] Furthermore, Orthogonal Multiple Access (OMA) technology is a multi-user access method in wireless communication systems. It considers the orthogonal allocation of physical resources in the time, frequency, spatial, and power domains to differentiate between users, ensuring signal orthogonality between different users and effectively avoiding interference between them. However, the user capacity of each orthogonal channel (time, frequency, subspace, sequence, etc.) in current OMA technology is approaching the Shannon limit, limiting the number of users that can access the network and making it difficult to meet the challenges facing communication technology. Therefore, 6G communication research needs to explore new theories of information transmission and suitable multiple access technologies to enable massive numbers of users to efficiently and flexibly connect to the network and transmit data on given wireless resources, thereby meeting the ever-increasing communication demands.

[0072] In related technologies, complex algorithms are used in base stations to distinguish signals from different users. In this case, the base station receiver is highly complex, and the increase in the number of users will exponentially increase the signal interference between users, further increasing the processing difficulty of the receiver. As a result, current orthogonal multiple access communication faces problems such as a large amount of data redundancy, high system complexity, limited capacity, high energy consumption, and low information transmission efficiency.

[0073] Based on this, embodiments of this application provide a semantic communication image processing method, system, device, and storage medium. The semantic feature signals of the source image are extracted by feature extraction units corresponding to different transmitters, and the feature domain is separated into multiple semantic feature subspaces. This enables the semantic feature signals of multiple users to be encoded and transmitted in the separated semantic feature subspaces. The discrete semantic feature vectors of different users constitute the total received signal, but they are approximately orthogonal to each other, thereby reducing the processing difficulty of the base station for the total received signal, improving the processing efficiency of the base station, and improving the communication transmission efficiency.

[0074] This application provides a semantic communication image processing method, system, device, and storage medium, which are specifically described through the following embodiments. First, the semantic communication image processing method in this application embodiment is described.

[0075] This application's embodiments can acquire and process relevant data based on artificial intelligence (AI) technology. AI is the theory, methods, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new type of intelligent machine that can react in a way similar to human intelligence. AI also studies the design principles and implementation methods of various intelligent machines, enabling them to possess perception, reasoning, and decision-making capabilities.

[0076] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0077] The semantic communication image processing method provided in this application relates to the field of communication technology. This method can be applied to a terminal, a server, or a computer program running on either the terminal or the server. For example, the computer program can be a native program or software module in an operating system; it can be a native application (APP), i.e., a program that needs to be installed in the operating system to run, such as a client supporting semantic communication image processing; it can also be a small program, i.e., a program that only needs to be downloaded to a browser environment to run; or it can be a small program that can be embedded in any APP. In short, the above-mentioned computer program can be any form of application, module, or plugin. The terminal communicates with the server via a network. The semantic communication image processing method can be executed by the terminal or the server, or by the terminal and the server working together.

[0078] In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, or smartwatch, etc. The server can be a standalone server, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms; it can also be a service node in a blockchain system, where the service nodes form a peer-to-peer (P2P) network. The P2P protocol is an application layer protocol running on top of the Transmission Control Protocol (TCP). The terminal and server can connect via Bluetooth, Universal Serial Bus (USB), or a network, etc., and this embodiment does not impose any limitations.

[0079] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0080] It should be noted that in all specific embodiments of this application, when processing data related to user identity or characteristics, such as user information, user behavior data, user historical data, and user location information, user permission or consent is obtained first. Furthermore, the collection, use, and processing of this data comply with relevant laws, regulations, and standards of the relevant countries and regions. In addition, when embodiments of this application require access to sensitive personal information of users, separate permission or consent from the user is obtained through pop-ups or redirects to confirmation pages. Only after obtaining the user's separate permission or consent is the necessary user-related data for the proper functioning of the embodiments of this application obtained.

[0081] First, the semantic communication image processing system provided in the embodiments of this application is described.

[0082] Reference Figure 1 , Figure 1 This is a schematic diagram of the semantic communication image processing system provided in the embodiments of this application. Figure 1 The semantic communication image processing system is based on a multi-user semantic multiple access network using single-carrier frequency-division multiple access (SFDMA). This includes multiple transmitters and base stations, each of which can be considered a user, for example... Figure 1 The following example uses N users to illustrate the point.

[0083] Each user corresponds to a semantic encoder. Taking the i-th transmitter as an example, the parameters of the semantic encoder are represented as φ. i It can be based on the user's source image u i Obtain semantic feature signal x i , is represented as: Then, the channel fading coefficient g between the i-th transmitter and the base station i As channel state information, semantic feature signal x is obtained based on the channel state information. i The corresponding transmitted signal. Since each user's semantic encoder is different, each semantic feature signal corresponds to a semantic feature subspace.

[0084] Then, the base station's receiver simultaneously receives transmitted signals from multiple users. At this point, the transmitted signals are superimposed, forming the total received signal y. The base station includes semantic decoders corresponding one-to-one with the transmitters. These semantic decoders correspond to semantic encoders, and the base station uses the semantic decoders to obtain the source image u from the total received signal y. i Corresponding target image processing data

[0085] In one embodiment, reference is made to Figure 2 , Figure 2 This is another schematic diagram of the semantic communication image processing system provided in the embodiments of this application. Combined with... Figure 1 , Figure 2 The semantic encoder at the transmitter includes a feature extraction unit, a quantizer, and a digital modulation unit, while the semantic decoder at the base station includes an equalizer and a semantic reasoning model.

[0086] Taking the i-th transmitter as an example, the transmitter uses a feature extraction unit to extract information from the user's source image u. i The semantic feature vector z is obtained from i Then the quantizer Q processes the semantic feature vector z i Quantization is performed to obtain discrete semantic feature vector b. i Then, the discrete semantic feature vector b is processed by the digital modulation unit. i Modulation is performed to obtain semantic feature signal x i .

[0087] In addition, correspondingly, after the base station obtains the total received signal y, it uses an equalizer (EQ) to perform channel equalization on the total received signal y to obtain the estimated signal for each transmitter. For example, the estimated signal for the i-th transmitter is... Then, the semantic reasoning model corresponding to the i-th transmitter is used. Semantic reasoning is performed on the estimated signal to obtain target image processing data. Where, θ i Representation of semantic reasoning model Network parameters.

[0088] The following is combined with Figure 1 and Figure 2 This application describes a semantic communication image processing method applied to the transmitting end in an embodiment of the present application.

[0089] Figure 3 This is an optional flowchart of the semantic communication image processing method provided in the embodiments of this application. Figure 3 The method may include, but is not limited to, steps 110 to 140. It is also understood that this embodiment... Figure 3 The order of steps 110 to 140 is not specifically limited. The order of steps can be adjusted or some steps can be reduced or added according to actual needs.

[0090] Step 110: Extract the semantic feature vector of the source image using the feature extraction unit.

[0091] In one embodiment, due to different image processing tasks, the source image includes an image to be classified or an image to be reconstructed. The image to be classified is used for image classification, and the image to be reconstructed is used for image reconstruction. When the source image is an image to be classified, the semantic reasoning model is a neural network model, and the target image processing data is the classification result of the image to be classified. When the source image is an image to be reconstructed, the feature extraction unit is an encoder unit, the semantic reasoning model is a decoder unit, and the target image processing data is the reconstruction result corresponding to the image to be reconstructed.

[0092] In one embodiment, when the source image is an image to be classified, for the i-th transmitter, the feature extraction unit It can be a neural network model, capable of analyzing source images u i Features are extracted and encoded to obtain the semantic feature vector z. i This feature extraction process can be represented as:

[0093]

[0094] In one embodiment, when the source image is the image to be reconstructed, the feature extraction unit is an encoder unit, referring to... Figure 4 , Figure 4 This is a flowchart illustrating the extraction of semantic feature vectors from a source image using a feature extraction unit, provided in an embodiment of this application. Specifically, it includes:

[0095] Step 410: After performing the first downsampling and self-attention processing on the source image using the first coding block, the first feature map is obtained.

[0096] In one embodiment, when the source image is the image to be reconstructed, the source image is represented as: Where H and W represent the height and width of the source image, and 3 represents the number of channels of the source image.

[0097] At this time, refer to Figure 5 , Figure 5This is a schematic diagram of the encoder unit in an embodiment of this application. Figure 5 The encoder unit comprises three coding blocks: the first coding block, the second coding block, and the third coding block. All three coding blocks are Swing Transformer structures, including a coding patch block and a Swing Transformer coding block. The coding patch block is used to divide the input image into multiple non-overlapping patches, which are then fed into the Swing Transformer coding block. Self-attention calculation is performed within each patch, and information is integrated through shift operations to achieve the encoding process of the input image. The Swing Transformer structure utilizes the coding patch block and the Swing Transformer coding block to work together to extract and transform semantic information in the image.

[0098] In one embodiment, the first coding block first processes the source image u i The first downsampling is performed to divide the source image into... Non-overlapping patches are labeled in order from top left to bottom right to obtain a concatenated sequence, which is then embedded into the embedding. Self-attention processing is then performed using a Swing Transformer block to obtain the first feature map, the size of which is [value missing]. C equals 3, representing the number of channels. The width and height of the first feature map are both half that of the source image, and the number of channels is the same as that of the source image.

[0099] Step 420: After performing a second downsampling and self-attention processing on the first feature map using the second coding block, a second feature map is obtained.

[0100] In one embodiment, in the second coding block, the first feature map is first downsampled a second time using a coding patch block to reduce its size and the number of patches, thereby reducing the image resolution and adjusting the number of channels, which helps reduce subsequent computation. After patching, the height and width of the feature map are reduced to half of their original values, while the number of channels is doubled. Subsequently, the adjusted patches are processed again by the Swing Transformer coding block to output the second feature map, which has a size of [missing information]. The width and height of the second feature map are both half that of the first feature map, and the number of channels is twice that of the first feature map.

[0101] Step 430: After performing a third downsampling and self-attention processing on the second feature map using the third coding block, the semantic feature vector is obtained through encoding.

[0102] In one embodiment, the second feature map is further downsampled a third time by the coding patch block of the third coding block, and then processed by the Swing Transformer coding block to obtain the final feature map, which is represented as follows: Its width and height are both half that of the second feature map, and its number of channels is twice that of the second feature map. Finally, the feature map is encoded to output a semantic feature vector. This semantic feature vector contains key semantic information from the source image, and its size is significantly smaller than that of the source image.

[0103] Overall, the feature extraction unit In other words, the encoder unit's process of extracting the semantic feature vector of the source image can be represented as:

[0104]

[0105] Step 120: Quantize the semantic feature vector to obtain a discrete semantic feature vector, and use a digital modulation unit to modulate the discrete semantic feature vector to obtain a semantic feature signal.

[0106] In one embodiment, in order to improve the transmission efficiency of the communication system and simplify the base station processing, the semantic feature vector z in this application embodiment... i Binary quantization is performed, using a quantizer to quantize the semantic feature vector to obtain a discrete semantic feature vector. The specific process includes: inputting the semantic feature vector into a linear layer to obtain a linear vector corresponding to the number of quantization bits; quantizing values ​​greater than or equal to zero in the linear vector to one, and otherwise quantizing to zero, thus obtaining the discrete semantic feature vector.

[0107] In one embodiment, the number of quantization bits d is related to the image processing result; for example, d is 16 or 64. Taking d as an example of 64, the assumed semantic feature vector z... i With a dimension of 1024, inputting the semantic feature vector into a linear layer yields a linear vector corresponding to the number of quantization bits. This means that the linear layer is used to quantize the 1024-dimensional semantic feature vector z. i Transform it into a 64-dimensional linear vector.

[0108] Next, the linear vector is binarized. Values ​​greater than or equal to zero in the linear vector are quantized to one, and otherwise quantized to zero, resulting in a discrete semantic feature vector. This discrete semantic feature vector is a discrete vector composed of 0s and 1s, represented as:

[0109]

[0110] Among them, b i Let T represent the discrete semantic feature vector corresponding to the i-th transmitter, and Q represent the transpose. i() indicates the quantization process using 0 or 1, z i,j This represents the numerical value in a linear vector.

[0111] In one embodiment, a discrete semantic feature vector b is obtained. i Then, a digital modulator is further utilized. For discrete semantic feature vector b i Modulation is performed to convert the bitstream into a signal form suitable for transmission over the channel, such as BPSK modulation. The modulated signal is then normalized to obtain the normalized semantic feature signal, represented as:

[0112] x i =[x i,1 ,x i,2 ,…,x i,d ] H =Norm(Θ m (b i ),i∈{1,..,N}

[0113] Where H represents the conjugate transpose, x i Let represent the semantic feature signal of the i-th transmitter. Norm(·) represents the normalization operation. The purpose of normalization is to ensure that the semantic feature signal can adapt to the characteristics of the channel during transmission and to ensure that the signal has a uniform power level, which helps the base station to accurately recover the transmitted semantic feature signal.

[0114] Step 130: Obtain the transmitted signal based on the transmit power, channel state information, and semantic feature signals corresponding to the transmitter.

[0115] In one embodiment, N transmitters simultaneously transmit corresponding semantic feature signals to the base station in the time-frequency domain, and transmit them via a multiple access communication network. For the i-th transmitter, g... i p represents the channel fading coefficient from the base station to the channel state information. i Given its transmission power, the process for obtaining its corresponding transmission signal is as follows: Figure 6 , Figure 6 This is a flowchart provided in this application embodiment for obtaining a transmitted signal based on the transmit power, channel state information, and semantic feature signals corresponding to the transmitter, specifically including:

[0116] Step 610: Obtain the conjugate transpose matrix of the channel state information.

[0117] Step 620: Obtain the transmission parameters corresponding to each transmitter based on the transmission power and the conjugate transpose matrix.

[0118] Step 630: Obtain the transmitted signal by multiplying the transmission parameters and the semantic feature signal.

[0119] The total received signal is obtained by summing the transmitted signal and the noise vector.

[0120] In one embodiment, the conjugate transpose matrix of the channel state information is g. i H The launch parameters are expressed as follows:

[0121]

[0122] The i-th transmitted signal is represented as:

[0123]

[0124] Step 140: Transmit the transmitted signal to the base station.

[0125] In one embodiment, N transmitters transmit signals simultaneously in the time-frequency domain. Therefore, the total received signal y received by the base station is obtained by accumulating the transmitted signal and the noise vector. The total received signal y is expressed as:

[0126]

[0127] Where n is a number with a mean of 0 and a variance of δ. 2 A complex noise vector, satisfying n:

[0128] When a base station receives a total received signal consisting of transmitted signals from multiple different transmitters, it performs channel equalization on the total received signal to obtain an estimated signal corresponding to each transmitter. The estimated signal is then input into a pre-trained semantic reasoning model corresponding to the transmitter for semantic reasoning to obtain the target image processing data corresponding to the source image.

[0129] In this application embodiment, the semantic feature signals of the source image are extracted by the feature extraction units corresponding to different transmitters. The feature domain is separated into multiple semantic feature subspaces, so that the semantic feature signals of multiple users are encoded and transmitted in the separated semantic feature subspaces. The discrete semantic feature vectors of different users constitute the total received signal, but they are approximately orthogonal to each other, thereby reducing the processing difficulty of the base station on the total received signal, improving the processing efficiency of the base station, and improving the communication transmission efficiency.

[0130] The following is combined with Figure 1 and Figure 2 This application describes a semantic communication image processing method applied to a base station in an embodiment of the present application.

[0131] Figure 7 This is an optional flowchart of the semantic communication image processing method provided in the embodiments of this application. Figure 7The method may include, but is not limited to, steps 710 to 740. It is also understood that this embodiment... Figure 7 The order of steps 710 to 740 is not specifically limited, and the order of steps can be adjusted or some steps can be reduced or added according to actual needs.

[0132] Step 710: Obtain the total received signal.

[0133] In one embodiment, the total received signal is obtained based on the transmitted signal of at least one transmitter, which is obtained by the semantic communication image processing method described in any of the above embodiments.

[0134] Step 720: Perform channel equalization on the total received signal to obtain the estimated signal corresponding to each transmitter.

[0135] In one embodiment, channel equalization is performed under ideal conditions, meaning that the channel state information between each transmitter and the base station is assumed to be known. It is understood that if the channel state information is unknown, channel estimation is required. Common channel estimation methods can be used, such as LS channel estimation, MMSE, LMMSE, etc., and this embodiment does not limit this approach.

[0136] The process of performing channel equalization on the total received signal to obtain the estimated signal corresponding to each transmitter specifically includes: obtaining the estimation coefficients based on the channel state information and the conjugate transpose matrix, and obtaining the estimated signal by multiplying the estimation coefficients and the total received signal.

[0137] The i-th estimated coefficient is expressed as:

[0138]

[0139] The i-th estimated signal is represented as:

[0140]

[0141] By transforming and calculating the estimated signal, another representation of the estimated signal corresponding to the i-th transmitted signal can be obtained:

[0142]

[0143] In the above formula, This represents the estimated signal corresponding to the i-th transmitter. This represents noise in a Rayleigh fading channel. This represents the interference signal between the transmitted signals of other transmitters and the transmitted signal of the i-th transmitter. This is due to the channel state information g. i Therefore, in the channel equalization process, embodiments of this application can divide the total received signal by g. iThis transforms the channel effect from multiplication to addition, thereby reducing the learning burden on the subsequent semantic reasoning model.

[0144] Step 730: Input the estimated signal into the pre-trained semantic reasoning model corresponding to the transmitter to perform semantic reasoning and obtain the target image processing data corresponding to the source image.

[0145] In one embodiment, when the source image is an image to be classified, the semantic reasoning model is a neural network model, and the target image processing data is the classification result of the image to be classified. When the source image is an image to be reconstructed, the feature extraction unit is an encoder unit, the semantic reasoning model is a decoder unit, and the target image processing data is the reconstruction result corresponding to the image to be reconstructed.

[0146] In one embodiment, after signal equalization processing, the i-th estimated signal is... Feed into semantic reasoning network The semantic reasoning network extracts semantic features from the data, performs classification or image reconstruction on them, and outputs the final result, obtaining the target image processing data corresponding to the i-th transmitter. The reasoning process is represented as follows:

[0147]

[0148] Where, θ i This represents the parameters of the i-th semantic reasoning network.

[0149] The semantic reasoning models for different types of source images and their corresponding training processes are described below.

[0150] In one embodiment, when the source image is an image to be classified, the semantic reasoning model is a neural network model. (See also...) Figure 8 , Figure 8 The training process of the semantic reasoning model provided in this application embodiment is illustrated in the following diagram:

[0151] Step 810: Obtain the training dataset.

[0152] First, a corresponding transmitted sample signal is generated based on the sample image from each transmitter. Then, the transmitted sample signals are accumulated, and noise is superimposed to obtain the total sample signal. The image processing label corresponding to each sample image is then obtained, and the total sample signal and the image processing label are associated. Multiple sets of total sample signals are generated in this manner.

[0153] In one embodiment, the training dataset contains L samples, corresponding to L total sample signals. The semantic encoder at each transmitter samples the semantic feature signals M times. That is, for the semantic feature signal of the transmitted sample corresponding to each total sample signal, the semantic encoder divides the semantic feature signal into M discrete data points. Correspondingly, the transmitted sample signal also corresponds to M discrete transmitted data points, and the image processing labels also correspond to M discrete label data points. The number of samplings here affects the complexity and accuracy of base station data processing. This embodiment sets the number of samplings according to actual needs.

[0154] Step 820: Obtain the estimated sample signal corresponding to each transmitter based on the total sample signal.

[0155] In one embodiment, the estimated sample signal for each transmitter is obtained based on the total sample signal, according to the method described above for calculating the estimated signal.

[0156] Step 830: Jointly train the semantic reasoning model using the estimated sample signals to obtain the image prediction results.

[0157] In one embodiment, data from multiple transmitters are jointly trained to perform multi-user joint source-channel-destination coding, achieving the goal of encoding and transmitting semantically relevant information from multiple users in separate semantic feature domains. In other words, the estimated sample signals are input into the semantic inference model corresponding to each transmitter to obtain the corresponding image prediction results.

[0158] Step 840: Calculate the first loss value based on the image processing label and image prediction result, calculate the second loss value based on the transmitted sample signal and the total sample signal, and calculate the third loss value based on the transmitted sample signal and the sample image.

[0159] In one embodiment, the first loss value is expressed as:

[0160]

[0161] Among them, y (l,m) This represents the discrete label data of the m-th image processing label corresponding to the transmitted sample of the i-th transmitter in the l-th sample, and represents the image prediction result of the i-th transmitter in the l-th sample. This indicates that the semantic reasoning model has parameters θ. i Under the premise that y (l,m) get The probability distribution is shown. Understandably, the first loss value here is used to measure the difference between the image processing label and the image prediction result, so the smaller the difference, the better.

[0162] The second loss value is expressed as:

[0163]

[0164] in, Y represents the j-th dimension of the m-th discrete transmitted data corresponding to the transmitted sample signal of the i-th transmitter in the l-th sample. j Let the j-th dimension of the total sample signal be denoted as . express and Y j The information entropy between them. It can be understood that the optimization objective of the second loss value here is to maximize the semantic transmission rate, that is, to maximize the corresponding information entropy.

[0165] The third loss value is expressed as:

[0166]

[0167] in, X represents the sample image of the i-th transmitter in the l-th sample. i,j This represents the j-th dimension of the transmitted sample signal from the i-th transmitter after quantization and modulation. X represents i,j and The information entropy between them. It's understandable that the optimization objective of the third loss value here is to measure the quality of semantic compression; the higher the compression level, the lower the corresponding information entropy.

[0168] Step 850: Obtain the total loss value based on the first loss value, the second loss value, and the third loss value. Adjust the weights of the semantic reasoning model based on the total loss value to obtain the trained semantic reasoning model corresponding to each transmitter.

[0169] In one embodiment, the total loss value is expressed as:

[0170]

[0171] in, β represents the total loss value corresponding to the image to be classified. i ≥0 represents a weighted parameter between the inference performance and model robustness of the i-th transmitter, which can be set according to actual needs.

[0172] The design of the total loss value in this embodiment achieves a trade-off between the amount of transmitted semantic information, inter-user interference, and inference accuracy. Therefore, the semantic inference model is weighted according to the total loss value until a preset training termination condition is reached, thus obtaining a trained semantic inference model corresponding to each transmitter. It is understood that different semantic inference models have different weights and can adapt to different channel states of transmitters.

[0173] In one embodiment, when the source image is the image to be reconstructed, the semantic reasoning model is a decoder unit. (Refer to...) Figure 9 , Figure 9 This is a schematic diagram of the decoder unit provided in an embodiment of this application. (Combined with...) Figure 5 ,from Figure 9 As can be seen from the diagram, the decoder unit and encoder unit have a symmetrical structure, including three decoding blocks: a first decoding block, a second decoding block, and a third decoding block. The first, second, and third decoding blocks are all Swing Transformer structures, each including a Swing Transformer decoding block and a decoding patch block. The Swing Transformer decoding block corresponds to the Swing Transformer encoding block and is used for decoding and reconstructing the signal. The decoding patch block is used to upsample the input image. Therefore, in this embodiment, the estimated signal... First, the code enters the first decoding block. After reconstruction and upsampling, it is then... Become Next, the second decoding block is entered. After reconstruction and upsampling, the code is then... Become Finally, entering the third decoding block, after reconstruction and upsampling, by The result becomes H×W×3. At this point, the target image processing data is the reconstruction result corresponding to the image to be reconstructed. The entire process is represented as follows:

[0174]

[0175] in, This represents a semantic reasoning network.

[0176] In one embodiment, corresponding to Figure 2 When the source image is the image to be reconstructed, refer to Figure 10 , Figure 10 This is another schematic diagram of the semantic communication image processing system provided in the embodiments of this application. The semantic encoder at the transmitting end includes a feature extraction unit, a quantizer, and a digital modulation unit, and the semantic decoder at the base station includes an equalizer and a semantic reasoning model. The feature extraction unit is an encoder unit, and the semantic reasoning model is a decoder unit.

[0177] The training process of the decoder unit is described below. In one embodiment, referring to... Figure 11 , Figure 11 Another schematic diagram illustrating the training process of the semantic reasoning model provided in this application embodiment, specifically including:

[0178] Step 1110: Obtain the training dataset.

[0179] The training dataset is generated in a similar manner to the training dataset in step 810. The training dataset includes the total sample signal and image processing labels. The total sample signal is obtained from the transmitted sample signal of each transmitter, which is generated based on the sample images.

[0180] Step 1120: Obtain the estimated sample signal corresponding to each transmitter based on the total sample signal.

[0181] Step 1130: Jointly train the semantic reasoning model using the estimated sample signals to obtain the image prediction results.

[0182] Step 1140: Calculate the total loss value based on the image processing labels and image prediction results.

[0183] The total loss value is expressed as:

[0184]

[0185] in, Let E represent the total loss value corresponding to the image to be reconstructed, and let E represent the expectation. u represents the joint probability distribution of image processing labels and image prediction results. i ' represents an image processing label. This represents the image prediction result.

[0186] Step 1150: Adjust the weights of the semantic reasoning model based on the total loss value to obtain the trained semantic reasoning model corresponding to each transmitter.

[0187] In one embodiment, the semantic inference model is weighted according to the total loss value until a preset training termination condition is met, thus obtaining a trained semantic inference model corresponding to each transmitter. The training termination condition can be that the trained semantic inference model meets the evaluation metric MS-SSIM. MS-SSIM stands for Multi-Scale Structural Similarity, a metric used to measure the quality of image reconstruction. It evaluates the similarity between images by considering the degree of structural similarity at different scales. The calculation process of MS-SSIM includes dividing the image processing labels and image prediction results into sub-images of different scales, calculating the SSIM index for each scale sub-image (the SSIM index is a measure of structural similarity), and then weighting these indices to obtain the final MS-SSIM value. The MS-SSIM value ranges from 0 to 1; the closer the value is to 1, the higher the similarity between the reconstructed image and the original image, and the better the image quality. It is understood that different semantic inference models have different weights and can adapt to different channel conditions of the transmitters.

[0188] In one embodiment, reference is made to Figure 12 , Figure 12This is a schematic diagram illustrating the classification performance of the image to be classified provided in an embodiment of this application. Figure 12 Image classification is performed on two users, U1 and U2, respectively, with quantization bits of d=16 and d=64, and a training signal-to-noise ratio of 5dB. In the figure, the vertical axis represents classification accuracy (percentage), the horizontal axis represents signal-to-noise ratio (dB), Upperbound represents the performance upper bound, DeepJSCC represents the deep joint source-channel coding scheme in related technologies, and SFDMA represents the semantic communication image processing method provided in this embodiment. Figure 12 As can be seen from the above, the classification accuracy of the SFDMA scheme provided in this application embodiment is always better than that of the Deep JSCC scheme in related technologies, and it approaches the upper bound of performance for different users, which proves the effectiveness of the semantic communication image processing method in image classification in this application embodiment.

[0189] In one embodiment, reference is made to Figure 13 , Figure 13 This is another schematic diagram illustrating the classification performance of the image to be classified provided in the embodiments of this application. The figure shows that, with a quantization bit depth of d128 and a training signal-to-noise ratio of 5dB, it is necessary to recognize the handwritten digits of different users. According to the recognition results, the semantic feature signals of users U1, U2, and U3 are separated in the semantic feature subspace, and the classification accuracy of each user is above 94%.

[0190] In one embodiment, reference is made to Figure 14 , Figure 14 This is a schematic diagram illustrating the reconstruction performance of the image to be reconstructed provided in an embodiment of this application. Figure 14 For visualization of the semantic feature subspace, from Figure 14 As can be seen, the images of the two users U1 and U2 can be separated in the semantic feature subspace, and the reconstruction process will not be affected by other users.

[0191] In one embodiment, reference is made to Figure 15 , Figure 15 This is another schematic diagram illustrating the reconstruction performance of the image to be reconstructed provided in an embodiment of this application. Figure 15 The paper presents a comparison of reconstruction performance when Peak Signal-to-Noise Ratio (PSNR) and Multi-Scale Structural Similarity (MS-SSIM) are used as evaluation metrics. Figure 15 As can be seen, whether it is the peak signal-to-noise ratio method (PSNR) or the multi-scale structural similarity method (MS-SSIM), the reconstruction performance of the SFDMA scheme provided in this application embodiment is always better than the DeepJSCC scheme in related technologies and approaches the upper bound of performance, which proves the effectiveness of the semantic communication image processing method in image reconstruction in this application embodiment.

[0192] This application's embodiments extract semantic features from source images and combine quantization and encoding to perform multi-user joint source-channel-destination coding. This enables the encoded transmission of multi-user semantic feature signals in separate semantic feature subspaces, effectively resolving inter-user interference and increasing the number of users. Furthermore, the semantic feature signals of different users are approximately orthogonal, allowing for encoded transmission of semantic information from each user within separate feature subspaces. This enhances the robustness of the semantic inference model and ensures the quality of image classification or reconstruction.

[0193] The technical solution provided in this application embodiment extracts semantic feature vectors from the source image using a feature extraction unit, quantizes the semantic feature vectors to obtain discrete semantic feature vectors, and modulates the discrete semantic feature vectors using a digital modulation unit to obtain semantic feature signals. A transmission signal is obtained based on the transmission power, channel state information, and semantic feature signals of the transmitting end, and is transmitted to the base station so that the base station receives a total received signal composed of transmission signals from multiple different transmitting ends. Channel equalization is performed on the total received signal to obtain an estimated signal corresponding to each transmitting end. The estimated signal is then input into a pre-trained semantic inference model corresponding to the transmitting end for semantic inference to obtain target image processing data corresponding to the source image. This application embodiment extracts semantic feature signals from the source image through feature extraction units corresponding to different transmitting ends, separating the feature domain into multiple semantic feature subspaces. This enables the encoding and transmission of semantic feature signals from multiple users in the separated semantic feature subspaces. The discrete semantic feature vectors of different users constitute the total received signal, but they are approximately orthogonal to each other, thereby reducing the processing difficulty of the total received signal for the base station, improving the processing efficiency of the base station, and improving communication transmission efficiency.

[0194] This application also provides an electronic device, including:

[0195] At least one memory;

[0196] At least one processor;

[0197] At least one program;

[0198] The program is stored in a memory, and the processor executes the at least one program to implement the semantic communication image processing method described above. The electronic device can be any smart terminal, including mobile phones, tablets, personal digital assistants (PDAs), and in-vehicle computers.

[0199] Please see Figure 16 , Figure 16 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes:

[0200] The processor 1601 can be implemented using a general-purpose central processing unit (CPU), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application.

[0201] The memory 1602 can be implemented as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 1602 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1602 and is called and executed by the processor 1601 using the semantic communication image processing method of the embodiments of this application.

[0202] The input / output interface 1603 is used to implement information input and output;

[0203] The communication interface 1604 is used to enable communication and interaction between this device and other devices. Communication can be achieved via wired means (e.g., USB, Ethernet cable) or wireless means (e.g., mobile network, Wi-Fi, Bluetooth).

[0204] Bus 1605 transmits information between various components of the device (e.g., processor 1601, memory 1602, input / output interface 1603, and communication interface 1604);

[0205] The processor 1601, memory 1602, input / output interface 1603 and communication interface 1604 are connected to each other within the device via bus 1605.

[0206] This application embodiment also provides a storage medium that stores a computer program, which, when executed by a processor, implements the above-described semantic communication image processing method.

[0207] Memory, as a non-transitory storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0208] The semantic communication image processing method, system, device, and storage medium proposed in this application extract semantic feature vectors from the source image using a feature extraction unit. These semantic feature vectors are quantized to obtain discrete semantic feature vectors, and then modulated using a digital modulation unit to obtain semantic feature signals. A transmission signal is obtained based on the transmission power, channel state information, and semantic feature signals of the transmitting end. This transmission signal is then transmitted to the base station, enabling the base station to receive a total received signal composed of transmission signals from multiple different transmitting ends. Channel equalization is performed on the total received signal to obtain an estimated signal corresponding to each transmitting end. This estimated signal is then input into a pre-trained semantic inference model corresponding to the transmitting end for semantic inference, resulting in target image processing data corresponding to the source image. This application's embodiments extract semantic feature signals from the source image using feature extraction units corresponding to different transmitting ends, separating the feature domain into multiple semantic feature subspaces. This allows the semantic feature signals of multiple users to be encoded and transmitted within these separated semantic feature subspaces. The discrete semantic feature vectors of different users constitute the total received signal, but they are approximately orthogonal to each other, thereby reducing the processing difficulty of the total received signal for the base station, improving the base station's processing efficiency, and enhancing communication transmission efficiency.

[0209] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.

[0210] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.

[0211] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0212] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.

[0213] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0214] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0215] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0216] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0217] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0218] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0219] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.

Claims

1. A semantic communication image processing method, characterized in that, Applied to the transmitting end, the method includes: The semantic feature vector of the source image is extracted using a feature extraction unit, wherein the source image includes an image to be classified or an image to be reconstructed; The semantic feature vector is quantized to obtain a discrete semantic feature vector, and the discrete semantic feature vector is modulated using a digital modulation unit to obtain a semantic feature signal; The process involves obtaining the transmit power, channel state information, and semantic feature signal corresponding to the transmitting end; obtaining the conjugate transpose matrix of the channel state information; obtaining the transmit parameters corresponding to each transmitting end based on the transmit power and the conjugate transpose matrix; and obtaining the transmit signal by multiplying the transmit parameters and the semantic feature signal. The transmitted signal is transmitted to a base station so that the base station receives a total received signal composed of the transmitted signals from multiple different transmitters. The total received signal is obtained by accumulating the transmitted signals and noise vectors. Channel equalization is performed on the total received signal to obtain an estimated signal corresponding to each transmitter. The estimated signal is then input into a pre-trained semantic reasoning model corresponding to the transmitter for semantic reasoning to obtain target image processing data corresponding to the source image. When the source image is an image to be classified, the semantic reasoning model is a neural network model, and the target image processing data is the classification result of the image to be classified. When the source image is an image to be reconstructed, the feature extraction unit is an encoder unit, the semantic reasoning model is a decoder unit, and the target image processing data is the reconstruction result corresponding to the image to be reconstructed. When the feature extraction unit is an encoder unit, the step of extracting the semantic feature vector of the source image using the feature extraction unit includes: After performing a first downsampling and self-attention processing on the source image using a first coding block, a first feature map is obtained, the width and height of which are both half of the source image. After performing a second downsampling and self-attention processing on the first feature map using a second coding block, a second feature map is obtained, the width and height of which are both half of the first feature map. After performing a third downsampling and self-attention processing on the second feature map using a third coding block, the semantic feature vector is obtained through encoding.

2. The semantic communication image processing method according to claim 1, characterized in that, The process of quantizing the semantic feature vector to obtain a discrete semantic feature vector includes: The semantic feature vector is input into the linear layer to obtain the linear vector corresponding to the number of quantization bits; Values ​​greater than or equal to zero in the linear vector are quantized to one, and otherwise quantized to zero, to obtain the discrete semantic feature vector.

3. A semantic communication image processing method, characterized in that, Applied to a base station, the method includes: A total received signal is obtained based on a transmission signal from at least one of the transmitting ends, wherein the transmission signal is obtained by the semantic communication image processing method according to any one of claims 1 to 2; Channel equalization is performed on the total received signal to obtain the estimated signal corresponding to each transmitter; The estimated signal is input into a pre-trained semantic reasoning model corresponding to the transmitter to perform semantic reasoning, thereby obtaining the target image processing data corresponding to the source image.

4. The semantic communication image processing method according to claim 3, characterized in that, The step of performing channel equalization on the total received signal to obtain the estimated signal corresponding to each transmitter includes: The estimated coefficients are obtained based on the channel state information and the conjugate transpose matrix; The estimated signal is obtained by multiplying the estimated coefficients and the total received signal.

5. The semantic communication image processing method according to claim 3, characterized in that, When the semantic reasoning model is a neural network model, before inputting the estimated signal into the pre-trained semantic reasoning model corresponding to the transmitting end for semantic reasoning, the method further includes: A training dataset is obtained, which includes a total sample signal and image processing labels. The total sample signal is obtained by the transmitted sample signal generated by each of the transmitters based on the sample image. The estimated sample signal corresponding to each transmitter is obtained based on the total sample signal; The semantic reasoning model is jointly trained using the estimated sample signals to obtain image prediction results; A first loss value is calculated based on the image processing label and the image prediction result; a second loss value is calculated based on the transmitted sample signal and the total sample signal; and a third loss value is calculated based on the transmitted sample signal and the sample image. The total loss value is obtained based on the first loss value, the second loss value, and the third loss value. The semantic reasoning model is then weighted based on the total loss value to obtain the trained semantic reasoning model corresponding to each of the transmitting ends.

6. The semantic communication image processing method according to claim 5, characterized in that, When the semantic reasoning model is a decoder unit, before inputting the estimated signal into the pre-trained semantic reasoning model corresponding to the transmitter for semantic reasoning, the method further includes: A training dataset is obtained, which includes a total sample signal and image processing labels. The total sample signal is obtained by the transmitted sample signal generated by each of the transmitters based on the sample image. The estimated sample signal corresponding to each transmitter is obtained based on the total sample signal; The semantic reasoning model is jointly trained using the estimated sample signals to obtain image prediction results; The total loss value is calculated based on the image processing label and the image prediction result; The semantic reasoning model is weighted based on the total loss value to obtain a trained semantic reasoning model corresponding to each of the transmitting ends.

7. A semantic communication image processing system, characterized in that, include: Multiple transmitting ends, wherein the transmitting ends are used to obtain corresponding transmission signals according to the semantic communication image processing method according to any one of claims 1 to 2; A base station, wherein the base station is configured to acquire the total received signal obtained based on the transmitted signal, and to obtain target image processing data corresponding to each of the transmitting ends according to the semantic communication image processing method according to any one of claims 3 to 6.

8. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the semantic communication image processing method according to any one of claims 1 to 6.

9. A storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the semantic communication image processing method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Uplink communication multi-user signal detection method and device based on generalized spatial modulation

    CN108259073A

  • Communication image transmission method, electronic equipment and readable storage medium

    CN116847089A