A lightweight image semantic communication method, system and readable storage medium
By segmenting and parallel processing of image semantic communication tasks on edge devices, combining deep learning and traditional vision algorithms, the delay problem caused by excessive computing power in the prior art is solved, and low-latency image semantic communication is achieved.
Patent Information
- Application Number
- CN202510074433.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-01-17
AI Technical Summary
The existing semantic communication methods rely on deep learning models to cause excessive computing power and are difficult to deploy efficiently on edge devices, especially in delay-sensitive communication systems that cannot meet real-time requirements.
The deep learning model is used to extract image texture semantics and combine traditional visual algorithms to extract color semantics. The segmentation computing task is processed in parallel on multiple auxiliary devices to reduce the computing load of a single device.
Low-latency image semantic communication is realized on edge devices, reducing the demand for device computing power, and improving the deployment flexibility and efficiency of the system.
Smart Images

Figure CN119495100B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of semantic communication, and particularly to a lightweight image semantic communication method, system and readable storage medium. Background Art
[0002] In the field of image semantic communication, existing research has shown that deep learning models can replace most of the original codecs and achieve better compression and accuracy performance than traditional methods. Therefore, the research on joint source-channel coding based on deep learning has greatly promoted the research on semantic communication systems, and almost all existing semantic communication schemes use the form of joint source-channel coding.
[0003] However, even on GPU servers with relatively high computing power, almost all of these better semantic communication algorithms based on deep learning joint source-channel coding require at least one hundred to thousands of milliseconds for model initialization and model inference. This level of latency is unacceptable when developing real-time communication systems that are sensitive to latency. Especially considering that a large part of communication occurs on edge devices, and the computing power of edge devices is usually lower and their memory is often limited in addition to latency issues. Therefore, when applying deep learning-based semantic coding algorithms on such edge devices, one of the biggest challenges is that they may not be able to provide sufficient computing power and complete the heavy calculations of model inference in a short time. In practical applications, the time consumed to complete these large-scale deep learning-based semantic encoding and decoding may even be longer than the time required to complete the entire communication process using traditional communication algorithms, which is unreasonable in most application scenarios. Summary of the Invention
[0004] The purpose of the present invention is to provide a lightweight image semantic communication method, system and readable storage medium, aiming to solve the problem that the existing semantic communication methods require too much computing power and are difficult to use under existing hardware conditions.
[0005] The technical solution adopted by the present invention to solve the technical problems is as follows:
[0006] The present invention provides a lightweight image semantic communication method, which includes:
[0007] The sending end obtains the source image, extracts the texture semantics of the source image using a deep learning model, and extracts the color semantics of the source image using a traditional vision algorithm;
[0008] The sending end sends the color semantics and the texture semantics to the receiving end;
[0009] The receiving end fuses the color semantics and the texture semantics to reconstruct the source image.
[0010] Furthermore, the sender sending the color semantics and the texture semantics to the receiver specifically includes:
[0011] The sender quantizes the color semantics into quantized color semantics and quantizes the texture semantics into quantized texture semantics;
[0012] The sender transmits the quantized color semantics and the quantized texture semantics to the receiver;
[0013] The receiver uses a traditional vision algorithm to restore the quantized color semantics to obtain the color semantics, and uses a texture restoration model to process the quantized texture semantics to restore the texture semantics.
[0014] Furthermore, both the sender and the receiver include multiple auxiliary devices, and the deep learning model includes a texture extraction model and a shallow neural network. Extracting the texture semantics of the source image using the deep learning model specifically includes:
[0015] Inputting the source image into the shallow neural network to extract the shallow features of the source image;
[0016] Dividing the shallow features into several feature blocks;
[0017] Allocating the several feature blocks to multiple auxiliary devices, and each auxiliary device uses the texture extraction model to process the feature blocks respectively to obtain the texture semantics.
[0018] Furthermore, processing the quantized texture semantics using the texture restoration model to restore the texture semantics specifically includes:
[0019] Dividing the quantized texture semantics into several semantic blocks;
[0020] Allocating the several semantic blocks to multiple auxiliary devices, and each auxiliary device uses the texture restoration model to process the semantic blocks respectively to restore the texture semantics.
[0021] Furthermore, dividing the shallow features into several feature blocks specifically includes:
[0022] Obtaining a segmentation consumption model, and according to the segmentation consumption model, calculating the segmentation size of each auxiliary device with the goal of minimizing the total time consumption:
[0023] ;
[0024] ;
[0025] ;
[0026] ;
[0027] Among them, is the total time consumption, represents a constraint, represents the th memory consumption of the auxiliary device under the corresponding segmentation size, represents the th upper limit of the available memory of the auxiliary device, is the maximum number of available devices, is the number of auxiliary devices;
[0028] According to each of the said segmentation sizes, the shallow features are segmented into several feature blocks;
[0029] The quantization texture semantics is segmented into several semantic blocks, specifically including:
[0030] Obtain a segmentation consumption model, and according to the segmentation consumption model, calculate the segmentation size of each auxiliary device with the lowest total time consumption as the goal:
[0031] ;
[0032] ;
[0033] ;
[0034] ;
[0035] According to each of the said segmentation sizes, the quantization texture semantics is segmented into several semantic blocks.
[0036] Furthermore, the obtaining of the segmentation consumption model specifically includes:
[0037] Establish a parallel time consumption model under the real protocol of the homogeneous system:
[0038] ;
[0039] Among them, is the actual protocol parallel time consumption of the homogeneous system, is the number of CPU cycles required to complete the inference calculation of unit input data, is the data transmission rate between the receiving end and the auxiliary device, is the data rate of the auxiliary device, represents the number of CPU cycles that the auxiliary device can complete in unit time, is the size of the data that needs to be transmitted additionally, is the quantization texture semantics size, is the number of available auxiliary devices;
[0040] Build a parallel time-consuming model under the actual protocol of the heterogeneous system:
[0041] ;
[0042] Wherein, represents the actual protocol parallel time consumption of the heterogeneous system, represents the auxiliary device the size of the data block processed, represents the main device and the auxiliary device the transmission rate between, represents the auxiliary device data rate, represents the auxiliary device the number of CPU cycles that can be completed per unit time;
[0043] Build a memory consumption model of the auxiliary device:
[0044] ;
[0045] Wherein, is the memory required to occupy for the inference calculation of the unit input data, is the memory required to occupy for loading the model parameter weights;
[0046] Take the set of the parallel time-consuming model under the actual protocol of the homogeneous system, the memory consumption model, and the parallel time-consuming model under the actual protocol of the heterogeneous system as the consumption model.
[0047] Furthermore, the receiving end fuses the color semantics and the texture semantics to reconstruct the source image, which specifically includes:
[0048] Input the color semantics and the texture semantics into the fusion model;
[0049] The fusion model reconstructs and outputs the source image according to the color semantics and the texture semantics.
[0050] Furthermore, the fusion model, the texture extraction model, and the texture reduction model are obtained through training, and the loss function of the training is:
[0051] ;
[0052] Wherein, represents the loss value, represents the source image, represents the reconstructed image, and are training hyperparameters, represents a texture extraction operator, represents the texture distribution of the source image, represents the texture distribution of the reconstructed image.
[0053] In addition, to achieve the above object, the present invention further provides a lightweight semantic communication system, the lightweight semantic communication system includes: a sending end and a receiving end, and a communication connection is established between the sending end and the receiving end;
[0054] The sending end is used to obtain a source image, extract the texture semantics of the source image by using a neural network, extract the color semantics of the source image by using a traditional vision algorithm, and send the color semantics and the texture semantics to the receiving end;
[0055] The receiving end is used to reconstruct the source image by fusing the color semantics and the texture semantics.
[0056] In addition, to achieve the above object, the present invention further provides a readable storage medium, the readable storage medium stores a lightweight image semantic communication program, and when the lightweight image semantic communication program is executed by a processor, the steps of the above-mentioned lightweight image semantic communication method are realized.
[0057] The present invention adopts the above technical solutions and has the following effects:
[0058] By splitting the semantic communication task and introducing a simple traditional computer vision algorithm to replace some deep learning models used in the semantic communication process in the prior art, the present invention simplifies the goal that deep learning needs to achieve, avoids the significant high latency that may be caused by overusing deep models, and solves the problem that the semantic communication method in the prior art completely depends on the neural network model, resulting in excessive computing power required and being difficult to be deployed and used on a real machine in the full scenario under the existing hardware conditions. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Figure 1 is a flowchart of the steps of a lightweight image semantic communication method in a preferred embodiment of the present invention;
[0060] Figure 2 is a full flowchart of a lightweight image semantic communication method in a preferred embodiment of the present invention;
[0061] Figure 3 is a schematic diagram of the relationship between the number of devices and the time-consuming in a preferred embodiment of the present invention;
[0062] Figure 4 is a schematic diagram of the calculation of texture information entropy and texture relative entropy in a preferred embodiment of the present invention;
[0063] Figure 5 Schematic diagram of the operating environment of a preferred embodiment of the system of the present invention. Detailed implementation manners
[0064] To make the objectives, technical solutions and advantages of the present invention clearer and more definite, the following further describes the present invention in detail with reference to the accompanying drawings and by way of examples. It should be understood that the specific examples described herein are only used to explain the present invention and are not used to limit the present invention.
[0065] Embodiment 1
[0066] The present invention uses to represent the input source image. The sending end uses the encoding function to map the source image to some finite complex symbols , , represents a complex number, represents a set of complex numbers the first complex symbol in the set, represents a set of complex numbers the second complex symbol in the set, represents a set of complex numbers the th complex symbol in the set, represents the number of elements in the set of complex numbers , and this process can be expressed as . Then, the encoded symbols are transmitted through the channel, and this process is described as . Among them, is the channel gain, is Gaussian noise. Finally, the receiving end applies the decoding function to reconstruct the source image and obtain the reconstructed image , and this process is described as .
[0067] A common strategy for constructing a deep learning-based JSCC (Joint source-channel coding) communication system is to train an autoencoder and use its encoder and decoder as the encoding and decoding functions of the communication system, which greatly simplifies the construction process of the communication system because only the autoencoder needs to be trained end-to-end. However, greedily using deep models may lead to significant high latency. Therefore, the objective of the present invention is to replace some deep models with some efficient traditional computer vision algorithms to solve the high-latency problem.
[0068] Specifically, please refer to Figure 1 and Figure 2 , Embodiment 1 of the present application is a lightweight image semantic communication method, which includes the steps of:
[0069] S1. The sender obtains the source image, extracts the texture semantics of the source image using a deep learning model, and extracts the color semantics of the source image using a traditional vision algorithm.
[0070] A key challenge in the actual development of a deep semantic communication system is the extremely limited edge resources. Therefore, the complexity of the method must be minimized. In computer vision, some efficient traditional algorithms can effectively complete simple tasks. Therefore, in the present invention, some traditional algorithms are used to partially replace the deep learning model.
[0071] Specifically, in this embodiment, two different processing methods are adopted for the source image to extract and restore its color and texture semantics respectively. Finally, the source image is reconstructed by fusing the color and texture semantics.
[0072] In this embodiment, the color semantics are specifically extracted using a Gaussian filter. Gaussian downsampling can smooth the image and remove high-frequency information. Since the color semantics belong to low-frequency information, the filter will retain this information. Therefore, for the color semantics, the extraction and reconstruction steps can consist of Gaussian pyramid downsampling and upsampling. By using simple traditional computer vision algorithms, the color semantics can be extracted with extremely low resource consumption. In step S1, only the extraction step is involved, and its process can be described as:
[0073] ;
[0074] where, is the extracted color semantics, is the Gaussian downsampling.
[0075] For the texture semantics, in this embodiment, an autoencoder with learnable parameters of is trained as a texture extraction model. The texture extraction model uses a dense block to process the features at each stage:
[0076] ;
[0077] where, is the texture extraction model, is the extracted texture semantics.
[0078] S2. The sender sends the color semantics and the texture semantics to the receiver.
[0079] After that, the transmitter sends the texture semantics and the color semantics to the receiver. Specifically, first, the texture semantics and the color semantics are respectively quantized and encoded:
[0080] ;
[0081] ;
[0082] Among them, is the quantization operation, is the quantized texture semantics, is the quantized color semantics. The quantized texture semantics and quantized color semantics after quantization are transmitted to the receiving end, and the receiving end restores the color semantics and texture semantics according to the quantized texture semantics and quantized color semantics. The whole process can be described as:
[0083] ;
[0084] ;
[0085] Among them, is the signal-to-noise ratio (SNR), is the communication channel, is Gaussian upsampling, is the texture restoration model, which is a decoding function.
[0086] In an optional embodiment, in order to deploy and accelerate the lightweight image semantic communication method of this embodiment on a single-edge device or edge cluster with extremely limited computing resources, the present invention adds a feature segmentation step before the first-stage texture processing, divides the computing tasks and distributes them to other limited devices in the local network / cluster for parallel computing, or performs serial computing on the current device after segmentation. Experiments show that the method of the present invention can directly complete the segmented computing tasks on the edge device without reducing the transmission quality, while other previous methods cannot effectively complete them on the edge device.
[0087] Specifically, those skilled in the art know that the first problem usually encountered when deploying a deep model on an edge device is the high inference latency caused by the low computing power of the device. Solving this weakness is crucial for designing a powerful communication system and cannot always be solved by reducing the size of the deep network, because this is likely to lead to a very significant decrease in communication accuracy. In addition to low computing power, memory limitation is also a typical problem of edge devices. The memory consumed by processing visual data using a deep model often exceeds the maximum available memory of the edge device, which greatly limits the deployment scenarios of existing deep learning-based semantic communication methods. To alleviate this problem, designing an edge deployment framework with joint perception-communication-computation functions may be a relatively good choice to accelerate the real-world deep semantic communication system.
[0088] Since the cost of upgrading / replacing the entire device is much higher and the flexibility is much lower, considering the cost issue, such a solution may not always be acceptable. In fact, in the context of edge computing technology, if the inference calculation can be divided into several parts and executed in parallel on multiple devices in a cluster or collaborative network, this problem can be better solved. Because the devices in the edge network are not always busy, they can share some computing resources with other devices in the same sub-network. In addition, a weak edge device can be connected to several weak computing nodes for simple and low-cost computing power upgrade instead of replacing the entire device.
[0089] Therefore, in this embodiment, before texture extraction, the source image is first input into a shallow neural network to extract the shallow features of the source image, and then the shallow features are segmented into several feature blocks, and the several feature blocks are assigned to multiple auxiliary devices. Each auxiliary device processes the feature blocks respectively using a texture extraction model to obtain texture semantics. Before texture restoration, the quantized texture semantics are segmented into several semantic blocks, and the several semantic blocks are assigned to multiple auxiliary devices. Each auxiliary device processes the semantic blocks respectively using a texture restoration model to restore the texture semantics. To achieve the segmented processing of the inference calculation for texture extraction and texture restoration.
[0090] Then, what needs to be considered is how to divide the inference calculation into several blocks without reducing the accuracy and then redistribute them to different devices to complete the calculation. The above problem has been considered in the previously proposed solutions and appropriate solutions have been provided. The semantic block segmentation step in the proposed feature processing can divide the inference calculation into several blocks to achieve redistribution and distributed computing on heterogeneous devices, or complete small-scale calculations multiple times on a single device to greatly reduce the consumption of memory resources.
[0091] In this embodiment, the inference latency of the semantic communication codec based on the convolutional neural network is:
[0092] ;
[0093] Among them, is the size of the quantized texture semantics, represents the inference latency, represents each convolutional layer, represents the total number of convolutional layers, and respectively represent the width and height of the input feature map of the th layer, represents the convolutional kernel size of the th layer, and respectively represent the The number of input and output channels of the layer represents the number of CPU cycles that the device performing the inference calculation can complete per unit time. Since the convolutional neural network model is fixed during the inference phase, the number of CPU cycles required for each unit of input is also fixed. If the number of CPU cycles required for each unit of input is set to , it can be seen that the inference latency has a linear relationship with the quantized texture semantic size .
[0094] First, consider the homogeneous scenario and use to represent the number of auxiliary devices. Assume that semantic blocks are transmitted to each device in a linear rather than parallel manner and an ideal protocol that can establish a connection without additional data transmission (such as headers) is used. The time taken is:
[0095] ;
[0096] where is the ideal protocol time for the homogeneous system, is the data transmission rate between the master device and the auxiliary device, is the data rate of the auxiliary device. It can be found that the time complexity of this process is , as Figure 3 shows, the time taken is inversely proportional to the number of auxiliary devices . As the number of auxiliary devices approaches infinity, the time taken approaches a constant. However, actual protocols are usually not like ideal protocols. They usually need to waste some communication resources to transmit additional data (such as headers). In this case, the time consumption is:
[0097] ;
[0098] where is the actual protocol time for the homogeneous system, and the time complexity is , is the size of the additional data. At the beginning, the time taken is approximately inversely proportional to the number of auxiliary devices because the size of the header is relatively small compared to the size of the semantic block data and can be ignored. Finally, when approaches infinity, the time complexity approaches .
[0099] If semantic block parallel transmission is used instead of semantic block linear transmission, then the time taken using the above ideal protocol or actual protocol will become:
[0100] ;
[0101] ;
[0102] Among them, is the ideal protocol parallel time consumption of the homogeneous system, is the actual protocol parallel time consumption of the homogeneous system, and their time complexities are both , and the latency decreases at an inversely proportional rate. In fact, as long as the number of devices is limited, parallel transmission is a commonly used strategy in practice.
[0103] Based on this, in this scenario, it is possible to try to extend the formula to the heterogeneous scenario, where the total latency is equal to the maximum time consumption of all devices, and its calculation formula is:
[0104] ;
[0105] Among them, represents the actual protocol parallel time consumption of the heterogeneous system, represents the th available auxiliary device, is the total number of auxiliary devices, represents the data size processed by the auxiliary device , represents the transmission rate between the main device and the auxiliary device , represents the data rate of the auxiliary device , represents the number of CPU cycles that the auxiliary device can complete per unit time.
[0106] In this embodiment, another important factor to be considered is the memory consumption of device inference. The memory consumption of the semantic communication codec based on the convolutional neural network is:
[0107] ;
[0108] When the convolutional neural network model is fixed, the memory occupied by the parameter weights and the memory required for the inference operation of each unit of input are both fixed. Set the memory occupied by the parameter weights as , and set the memory required for the inference calculation of the unit input data as , then this formula can be further simplified to a linear relationship. In the heterogeneous case, the memory consumption of the auxiliary device is:
[0109] ;
[0110] When designing a semantic communication system, it is always desired that the total transmission delay be as low as possible, and there is also an upper bound supported by each device on the memory consumption of the system on that device. Therefore, the objective function of the entire acceleration process can be described as:
[0111] ;
[0112] ;
[0113] ;
[0114] ;
[0115] where, is the total time consumption. Using this objective function, the data size can be found to minimize the total time consumption while the memory consumption of each device does not exceed its available memory upper limit . is the maximum number of available devices in the local network or cluster.
[0116] In a homogeneous system, . In a heterogeneous system, . In both cases, to solve the optimization problem, it is necessary to know the amount of computation required for each unit of input, the maximum available computing power and memory of each edge device in the cluster, and the of each device and the transmission rate to the master device. Finally, using the corresponding and available resource information to solve the objective function of the above acceleration process, thereby determining the strategy of block computing.
[0117] S3. The receiver fuses the color semantics and the texture semantics to reconstruct the source image.
[0118] The receiver uses a deep network block with a learnable parameter of as a fusion model to fuse the color and texture semantics to restore the source image:
[0119] ;
[0120] where, is the reconstructed image, is the fusion function for fusing the semantics.
[0121] The present invention trains a texture restoration model, a texture processing model, and a deep network block using a loss function. Specifically, in an alternative embodiment, during the training of the texture processing model, the texture restoration model, and the deep network block, the loss function used in this embodiment is a combined loss function of mean squared error and SSIM (Mean Structural Similarity Index Measure), which is as follows:
[0122] ;
[0123] where, represents the loss value, are two hyperparameters, represents a real number, represents the structural similarity function.
[0124] During the training process, this embodiment uses the Adam optimizer to determine the , and values that minimize the distortion:
[0125] ;
[0126] where, is the joint probability distribution of the source image and the reconstructed image, represents the expected value of under the joint probability distribution, represents parameters, represents parameters, represents parameters.
[0127] Finally, the present invention evaluates the method of the present invention in terms of semantic communication. In practice, the present invention finds that images with similar quality judged by classical metrics such as mean squared error (MSE) and peak signal-to-noise ratio (PSNR) do not always match the quality perceived subjectively by humans.
[0128] This is because these metrics typically focus on the low-frequency information (such as color) of the image and ignore the high-frequency information (such as details and textures). Therefore, such metrics may not accurately display the distortion of all types of information. In addition, since the human visual system is not sensitive to subtle color differences, these metrics do not always correctly approximate human perception, especially when the image quality is relatively high. To alleviate this problem, prior art research has also proposed some other metrics, such as perceptual loss and adversarial loss. Perceptual loss uses a pre-trained deep neural network (such as VGG-16), assuming its ability to perceive semantic information, extracts important features, and then measures the distortion of the image from the perspective of the extracted features. Adversarial loss uses a discriminator based on a deep neural network to check whether the image has a similar probability distribution. Unfortunately, adversarial loss is often unstable during training, and combined with the problem of low interpretability of deep neural networks in the above two losses, they are not good methods for measuring the distortion of high-resolution images. Therefore, a better method is needed to measure quality and explain the differences from the perspectives of evaluation and human perception.
[0129] Specifically, the metric proposed by the present invention needs to consider what helps humans obtain information from images. In daily life, it can be found that humans are good at obtaining a large amount of information from images composed of lines such as sketches. These lines describe the basic structure of the image and provide the main information of the image. Based on this phenomenon, some lines can be used to describe the information of the image. More precisely, the texture of the image can provide information for humans about the image. Secondly, a method for measuring information is needed. Since the microscopic displacement of texture lines usually does not affect the information, a distortion metric method strictly targeting pixels, such as mean absolute error (MAE) or mean squared error (MSE), should not be used.
[0130] In fact, previous information theory research has provided some valuable tools for measuring information. Information entropy is a powerful tool for measuring the amount of information in a given source. As shown in (a) of Figure 4 , the information entropy of the image texture can be calculated to approximately judge the amount of information that can be obtained from the human perspective. Specifically, taking the original image as the standard, it is checked whether the amount of information provided by the reconstructed image is close. However, although the texture information entropy (TIE) can accurately describe the amount of information, it does not always show the quality of the information. If one image contains a lot of incorrect and chaotic information, while another image contains limited but accurate information, then the first image may have a higher TIE value, but the information quality is worse. This weakness will severely limit the use of this measurement method.
[0131] Since the evaluation of information entropy only focuses on a single objective, it is difficult to give the relationship between images. In contrast, relative entropy (also known as Kullback-Leibler (KL) divergence or information divergence) can well describe the divergence of information distribution. Please refer to Figure 4 in (b) of which, which shows the process of determining the texture relative entropy (TRE). First, the texture is extracted from the image using an edge detection algorithm. Specifically, the present invention uses the Canny operator as the edge detection algorithm. Next, a sliding window is used to intercept small texture blocks. The size of each small block intercepted by the sliding window is determined by a hyperparameter which determines the allowed pixel displacement of the texture. ranges from 1 to the size of the image. To obtain more representative evaluation results, in this embodiment, the value that maximizes the standard deviation of the evaluation results of various reconstruction methods is selected.
[0132] Finally, the average relative entropy of the corresponding texture blocks is calculated to determine the divergence degree of the local information distribution. This process can be expressed as:
[0133] .
[0134] where, and are the height and width of the original image / reconstructed image, and are the height and width of the sliding window, represents the local information distribution of the texture block of the original image cut out by the sliding window, represents the local information distribution of the texture block of the reconstructed image cut out by the sliding window, represents the KL divergence, represents the texture relative entropy.
[0135] Among these parameters, , , and can all be regarded as constants. Since and represent the texture distributions of the original image and the reconstructed image respectively, if and are close, the original image and the reconstructed image are relatively similar. To minimize the information distortion, it is only necessary to minimize the gap between and . It is not difficult to see that when and When they are the same, the maximum lower limit of TRE is 0 and there is no upper limit. In addition, the above TRE calculation is asymmetric because the KL divergence is asymmetric. TRE can be easily made symmetric by replacing the KL divergence with the Jensen-Shannon (JS) divergence.
[0136] So far, the present invention has proposed a new standard for checking the texture quality of reconstructed images. Different from most of the previous standards, the method of the present invention can use relative entropy to measure the divergence between image texture information in an interpretable manner. The introduction of a sliding window can flexibly handle the micro-pixel displacement of the texture, thus better approaching human perception.
[0137] In addition, based on the above evaluation of the texture quality of the reconstructed image, in an alternative embodiment, in order to strengthen model training, the present invention hopes to use TRE as a loss function, but according to the research of the present invention, this cannot be directly achieved. Different from using TRE as an evaluation metric, using TRE as one of the loss functions to supervise the training of a deep autoencoder has more limitations. Among them, the main difficulty is that the loss function itself must be differentiable. However, since the Canny edge detection algorithm adopted in the process of calculating TRE is non-differentiable, TRE cannot be directly used as one of the loss functions.
[0138] For this reason, the present invention provides a simple and effective alternative method. The present invention uses another differentiable edge detection operator (such as the Sobel operator) to replace the Canny operator. In order to simplify the calculation and speed up the training process, the present invention omits the use of a sliding window during the training process and uses the MSE loss to supervise the error between the results after applying the Sobel operator to the original image and the reconstructed image. This is because during the training process, the present invention hopes that the higher the accuracy of the texture, the better, so any pixel offset is not acceptable during the training process, and a loss function that focuses on pixel differences can be directly adopted. Based on the above, the new loss function used in this embodiment is:
[0139] ;
[0140] Among them, are two hyperparameters, is the algorithm used by the present invention to extract texture, that is, the Sobel operator, represents the texture distribution of the source image, represents the texture distribution of the reconstructed image. Experiments prove that using such a loss function can significantly improve the texture quality.
[0141] Embodiment 2
[0142] Please refer to Figure 5, Based on the above method, the present invention also provides a lightweight semantic communication system, which includes: a sending end and a receiving end, and a communication connection is established between the sending end and the receiving end;
[0143] The sending end is used to obtain a source image, extract the texture semantics of the source image using a neural network, extract the color semantics of the source image using a traditional vision algorithm, and send the color semantics and the texture semantics to the receiving end;
[0144] The receiving end is used to fuse the color semantics and the texture semantics to reconstruct the source image.
[0145] Embodiment III
[0146] This embodiment provides a storage medium, and the readable storage medium stores a lightweight image semantic communication program. When the lightweight image semantic communication program is executed by a processor, the steps of the lightweight image semantic communication method described above are implemented.
[0147] In summary, by using a simple traditional computer vision algorithm instead of the deep learning color semantic extraction model used in the prior art during semantic communication, the present invention avoids the significant high latency that may be caused by using a deep model, and solves the problem that the semantic communication method of the prior art completely adopts a neural network model, resulting in excessive computing power required and being difficult to use under the existing hardware conditions.
[0148] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such a process, method, article or system. Without more limitations, the element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or system including the element.
[0149] Of course, those of ordinary skill in the art can understand that all or part of the processes of implementing the above method embodiments can be completed by instructing relevant hardware (such as a processor, a controller, etc.) through a computer program. The program can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes of the above method embodiments. The storage medium can be a memory, a magnetic disk, an optical disc, etc.
[0150] It should be understood that the application of the present invention is not limited to the above examples. For those of ordinary skill in the art, improvements or transformations can be made according to the above description, and all such improvements and transformations should fall within the protection scope of the appended claims of the present invention.
Claims
1. A lightweight image semantic communication method, characterized in that, The lightweight image semantic communication method includes: The sender obtains the source image, extracts the texture semantics of the source image using a deep learning model, and extracts the color semantics of the source image using a traditional vision algorithm; The sender sends the color semantics and the texture semantics to the receiver; The receiver fuses the color semantics and the texture semantics to reconstruct the source image; The sender sending the color semantics and the texture semantics to the receiver specifically includes: The sender quantizes the color semantics into quantized color semantics and quantizes the texture semantics into quantized texture semantics; The sender transmits the quantized color semantics and the quantized texture semantics to the receiver; The receiver uses a traditional vision algorithm to restore the quantized color semantics to obtain the color semantics, and uses a texture restoration model to process the quantized texture semantics to restore the texture semantics; The traditional vision algorithm is specifically the Gaussian filter algorithm; Both the sender and the receiver include multiple auxiliary devices. The deep learning model includes a texture extraction model and a shallow neural network. Extracting the texture semantics of the source image using the deep learning model specifically includes: Inputting the source image into the shallow neural network to extract the shallow features of the source image; Dividing the shallow features into several feature blocks; Assigning several of the feature blocks to multiple auxiliary devices, and each auxiliary device uses the texture extraction model to process the feature blocks respectively to obtain the texture semantics; Processing the quantized texture semantics using the texture restoration model to restore the texture semantics specifically includes: Dividing the quantized texture semantics into several semantic blocks; Assigning several of the semantic blocks to multiple auxiliary devices, and each auxiliary device uses the texture restoration model to process the semantic blocks respectively to restore the texture semantics; Dividing the shallow features into several feature blocks specifically includes: Obtaining a segmentation consumption model, and according to the segmentation consumption model, calculating the segmentation size of each auxiliary device with the goal of minimizing the total time consumption: ; ; ; ; in, is the total time consumed, Indicates constraints, Indicates The memory consumption of each auxiliary device at the corresponding partition size, Indicates The maximum amount of memory available to auxiliary devices. is the maximum number of available devices, is the number of auxiliary devices; Dividing the shallow features into several feature blocks according to the respective segmentation sizes; Dividing the quantized texture semantics into several semantic blocks specifically includes: Obtaining a segmentation consumption model, and according to the segmentation consumption model, calculating the segmentation size of each auxiliary device with the goal of minimizing the total time consumption: Dividing the quantized texture semantics into several semantic blocks according to the respective segmentation sizes; Obtaining the segmentation consumption model specifically includes: Establishing a parallel time consumption model under the real protocol of a homogeneous system: ; Among them, is the actual protocol parallel time consumption of the isomorphic system, is the number of CPU cycles required to complete the inference calculation of the unit input data, is the data transfer rate between the receiving end and the auxiliary device, is the data of the auxiliary device rate, represents the number of CPU cycles that the auxiliary device can complete per unit time, is the size of the additional data to be transferred, is the quantization texture semantic size, is the number of available auxiliary devices; Establishing a parallel time consumption model under the real protocol of a heterogeneous system: ; Among them, represents the actual protocol parallel time consumption of the heterogeneous system, represents the auxiliary device which is the data block size processed, represents the transmission rate between the main device and the auxiliary device ; represents the data rate of the auxiliary device ; represents the number of CPU cycles that the auxiliary device can complete per unit time; Establishing a memory consumption model of the auxiliary device: ; Among them, is the memory required to occupy for the inference calculation of the input data of the completion unit, is the memory required to occupy for loading the model parameter weights; Taking the set of the parallel time consumption model under the real protocol of the homogeneous system, the memory consumption model, and the parallel time consumption model under the real protocol of the heterogeneous system as the consumption model.
2. The lightweight image semantic communication method according to claim 1, wherein The receiver fusing the color semantics and the texture semantics to reconstruct the source image specifically includes: Inputting the color semantics and the texture semantics into a fusion model; The fusion model reconstructs and outputs the source image according to the color semantics and the texture semantics.
3. The lightweight image semantic communication method according to claim 2, characterized in that, The fusion model, the texture extraction model, and the texture restoration model are obtained through training, and the loss function for the training is as follows: ; Among them, represents the loss value, represents the source image, represents the reconstructed image, and are training hyperparameters, represents the texture extraction operator, represents the texture distribution of the source image, represents the texture distribution of the reconstructed image.
4. A lightweight semantic communication system, characterized in that, The lightweight semantic communication system is applied to the lightweight image semantic communication method according to any one of claims 1-3. The lightweight semantic communication system includes a sending end and a receiving end, and a communication connection is established between the sending end and the receiving end; The sending end is configured to obtain a source image, extract the texture semantics of the source image by using a neural network, extract the color semantics of the source image by using a traditional vision algorithm, and send the color semantics and the texture semantics to the receiving end; The receiving end is configured to reconstruct the source image by fusing the color semantics and the texture semantics.
5. A readable storage medium, characterized in that, The readable storage medium stores a lightweight image semantic communication program, and when the lightweight image semantic communication program is executed by a processor, the steps of the lightweight image semantic communication method according to any one of claims 1-3 are implemented.
Citation Information
Patent Citations
Channel diversity method and device for semantic communication
CN115149986A
Lightweight image style migration method based on color and texture dual channels
CN119295296A
Cited By
A generative ai-based satellite semantic communication system
CN122512974A