Semantic information compression method, semantic information transmission method, and video image restoration method

By extracting the potential feature map of video frames and merging related channels, efficient compression and transmission of video images are achieved, the problem of index errors in the prior art is solved, and the accuracy of image recovery and communication stability are improved.

WO2025112238A1PCT designated stage expired Publication Date: 2025-06-05BEIJING UNIV OF POSTS & TELECOMM
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/082764
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-28
Filing Date
2024-03-20
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

In the process of compressing and transmitting semantic information of video images, the index of feature vector elements is required to transmit, resulting in errors prone to unreliable channels and failing to restore images.

Method used

The latent feature map of the video frame is extracted through the latent feature extractor, the correlation between the latent feature map channels of each adjacent two frames of images is calculated, and channels with high correlation are merged to form a shared channel to realize the compression and transmission of semantic information.

Benefits of technology

It improves the compression efficiency of video images, reduces the amount of data transmitted, avoids the problem of index errors, and improves the accuracy of image recovery and communication stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024082764_05062025_PF_FP_ABST
    Figure CN2024082764_05062025_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the present disclosure are a semantic information compression method, a semantic information transmission method, and a video image restoration method. The semantic information compression method comprises: by means of a potential feature extractor, extracting potential features of each image frame of a video to be compressed, generating a plurality of potential feature maps correspondingly, and taking the potential feature maps as semantic information, each potential feature map comprising a plurality of channels, and the plurality of channels of different potential feature maps being in one-to-one correspondence; on the basis of the semantic information of every two adjacent image frames, calculating the correlation between a plurality of corresponding channels of potential feature maps of every two adjacent image frames, so as to obtain a plurality of correlation values; and, ranking the plurality of correlation values in a descending order, and on the basis of a preset compression ratio, a channel merging module selecting the plurality of corresponding channels of the potential feature maps of every two adjacent image frames corresponding to the top-ranking correlations, and merging same, so as to achieve compression of the semantic information in said video. The present disclosure can improve the compression efficiency of video semantic information.
Need to check novelty before this filing date? Find Prior Art

Description

Semantic information compression and transmission method and video image restoration method

[0001] Related applications

[0002] This disclosure claims priority to the Chinese patent application filed with the China Patent Office on November 28, 2023, with application number 202311606727.X and invention name “Semantic Information Compression and Transmission Method and Video Image Restoration Method”, the entire contents of which are incorporated by reference into this application. Technical Field

[0003] The present disclosure relates to the field of communication and data processing technology, and in particular to a semantic information compression and transmission method and a video image restoration method. Background Art

[0004] The development of science and technology has greatly enriched people's lives. The evolution of media communication has brought user experiences ever closer to the real world. Communication methods are becoming increasingly popular and modular. Video is playing an increasingly prominent role in production and daily life, and people are spending more time streaming media. Compared to text, video data is many times larger, resulting in video accounting for nearly 70% of all data on the internet.

[0005] Semantic communication uses computing power at both ends of the transmitter and receiver at the expense of computing power. Generally, the feature vector of the video stream is extracted through a neural network, the feature vector is transmitted through the channel, and then the neural network is used to restore the transmitted image.

[0006] Currently, some researchers are working from the perspective of information entropy, calculating the entropy of feature vectors and using this to determine the importance of vector elements. During communication, they suppress information with low entropy and compress the transmitted data. However, this approach requires transmitting the index of the transmitted vector elements, and over unreliable channels, the transmitted index can be erroneous, leading to image recovery failure.

[0007] Summary of the Invention

[0008] In view of this, the embodiments of the present disclosure provide a semantic information compression and transmission method and a video image restoration method to eliminate or improve one or more defects in the prior art.

[0009] One aspect of the present disclosure provides a method for compressing semantic information, the method comprising:

[0010] A latent feature extractor extracts latent features from each frame of the compressed video, and generates a plurality of latent feature maps accordingly. The latent feature maps are used as semantic information. Each latent feature map includes a plurality of channels, and the channels of different latent feature maps correspond to each other one by one.

[0011] Calculating correlations between multiple corresponding channels of potential feature maps of each of the two adjacent frames of images based on semantic information of each of the two adjacent frames of images to obtain multiple correlation values;

[0012] The values ​​of multiple correlations are sorted from large to small, and a channel merging module selects multiple corresponding channels of the potential feature map of each two adjacent frames of images corresponding to the correlations sorted first based on a preset compression ratio to merge to form a shared channel. The semantic information corresponding to the shared channel and the respective non-shared channels constitute the semantic information of each two adjacent frames of images after compression, thereby realizing the compression of the semantic information in the video to be compressed.

[0013] In some embodiments of the present disclosure, the step of calculating the correlation between multiple corresponding channels of the potential feature map of each two adjacent frames of images based on the semantic information of each two adjacent frames of images includes:

[0014] The latent feature maps of multiple corresponding channels in the latent feature maps of each of the two adjacent frames of images are converted into multiple pairs of semantic vectors, where the two semantic vectors of each corresponding channel form a pair of semantic vectors;

[0015] The Pearson correlation coefficient between each pair of semantic vectors is calculated to obtain the correlation between multiple corresponding channels of the potential feature map of each two adjacent frames.

[0016] In some embodiments of the present disclosure, the feature extractor includes multiple convolutional layers and multiple GDN activation functions, wherein a GDN activation function is connected between every two convolutional layers, the first convolutional layer serves as the input of the feature extractor, and the last convolutional layer serves as the output of the feature extractor. The multiple convolutional layers are used to downsample each frame image of the compressed video, generate multiple potential feature maps accordingly, increase the number of channels of the image, and reduce the size of the potential feature map of each channel.

[0017] Another aspect of the present disclosure provides a method for transmitting semantic information, the method comprising:

[0018] The sending end compresses the semantic information of each frame image in the video to be transmitted using the aforementioned semantic information compression method;

[0019] The source-channel joint encoder encodes the semantic information of each frame image in the compressed video to be transmitted to obtain a semantic code;

[0020] The channel transmits the semantic coding of each frame image in the compressed video to be transmitted.

[0021] In some embodiments of the present disclosure, the source-channel joint encoder includes multiple convolutional layers and multiple GDN activation functions, wherein a GDN activation function is connected between every two convolutional layers, the first convolutional layer serves as the input of the source-channel joint encoder, and the last convolutional layer serves as the output of the source-channel joint encoder, and the multiple convolutional layers are used to encode semantic information.

[0022] In some embodiments of the present disclosure, the step of transmitting the semantic coding of each frame image in the compressed video to be transmitted through the channel also includes: transmitting the index of the shared channel of the potential feature map of each frame image in the compressed video to be transmitted.

[0023] Another aspect of the present disclosure provides a video image restoration method, the method comprising:

[0024] The receiving end receives semantic codes of each frame image in the compressed video from the transmitting end, and decodes the semantic codes by a source-channel joint decoder to obtain semantic information of each frame image in the compressed video, wherein the semantic information includes a compressed shared channel and respective non-shared channels of two adjacent frames before compression corresponding to each frame image, where the respective non-shared channels are respectively a first non-shared channel and a second non-shared channel;

[0025] A channel recovery module extracts the shared channel, the first non-shared channel, and the second non-shared channel from semantic information of each frame of the compressed video based on a preset compression ratio;

[0026] Combining the semantic information of the shared channel with the semantic information corresponding to the first non-shared channel and the second non-shared channel to restore the semantic information of each two adjacent frames before compression, thereby obtaining the semantic information of each frame in the decompressed video;

[0027] The latent feature restorer restores each frame of the video based on the semantic information of each frame in the decompressed video, thereby obtaining the transmitted video.

[0028] In some embodiments of the present disclosure, the source-channel joint decoder includes multiple convolutional layers and multiple GDN activation functions, wherein a GDN activation function is connected between every two convolutional layers, the first convolutional layer serves as the input of the source-channel joint decoder, and the last convolutional layer serves as the output of the source-channel joint decoder. The multiple convolutional layers of the source-channel joint decoder are used to decode semantic coding.

[0029] In some embodiments of the present disclosure, the latent feature restorer includes multiple transposed convolutional layers, multiple GDN activation functions, convolutional layers and Tanh activation functions, wherein the number of the transposed convolutional layers and the GDN activation functions is the same, a GDN activation function is connected between every two transposed convolutional layers, and the output of the last GDN activation function is sequentially connected to the convolutional layer and the Tanh activation function. The first transposed convolutional layer serves as the input of the latent feature restorer, and the Tanh activation function serves as the output of the latent feature restorer. The multiple transposed convolutional layers are used to upsample the semantic information of each frame image in the decompressed video, reduce the number of image channels, and increase the size of the image of each channel.

[0030] Another aspect of the present disclosure provides an electronic device, which includes: a computer device, the computer device including a processor and a memory, the memory storing computer instructions, the processor being used to execute the computer instructions stored in the memory, and when the computer instructions are executed by the processor, the electronic device implements the steps of the aforementioned semantic information compression method, semantic information transmission method, and video image restoration method.

[0031] Another aspect of the present disclosure provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the aforementioned semantic information compression method, semantic information transmission method, and video image restoration method.

[0032] The semantic information compression and transmission method and video image restoration method disclosed in the present invention extract shared potential features based on the correlation of feature maps or feature matrices, merge them from the dimension of the channels of the feature maps or feature matrices, and then realize the compression of semantic information, which can improve the compression efficiency and have good compression performance.

[0033] Additional advantages, objects, and features of the present disclosure will be described in part in the following description and will become apparent to those skilled in the art upon study of the following or may be learned from practice of the present disclosure. The objects and other advantages of the present disclosure may be achieved and obtained by the structures specifically pointed out in the description and drawings.

[0034] Those skilled in the art will understand that the purposes and advantages that can be achieved by the present disclosure are not limited to the above specific descriptions, and the above and other purposes that can be achieved by the present disclosure will be more clearly understood from the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] The drawings described herein are used to provide a further understanding of the present disclosure, constitute a part of this application, and do not constitute a limitation of the present disclosure.

[0036] FIG1 is a flow chart of an embodiment of a semantic information compression method disclosed herein;

[0037] FIG2 is a schematic diagram of a network structure of a potential feature extractor according to an embodiment of the present invention;

[0038] FIG3 is a flow chart of an embodiment of step S120 in the semantic information compression method of the present disclosure;

[0039] FIG4 is a schematic diagram of an example of channel merging according to the present disclosure;

[0040] FIG5 is a flow chart of an embodiment of the semantic information transmission method disclosed herein;

[0041] FIG6 is a schematic diagram of a network structure of an embodiment of a source-channel joint encoder and a source-channel joint decoder disclosed herein;

[0042] FIG7 is a flow chart of an embodiment of a video image restoration method disclosed herein;

[0043] FIG8 is a schematic diagram of an example of channel recovery according to the present disclosure;

[0044] FIG9 is a schematic diagram of a network structure of an embodiment of a potential feature restorer disclosed herein;

[0045] FIG10 is a schematic diagram of an example of the semantic information compression, transmission and video image restoration process of the present disclosure;

[0046] FIG11 is a schematic diagram comparing the effects of the video image restoration method disclosed herein and the restoration of a transmitted image in the prior art;

[0047] FIG12 is a schematic diagram showing a comparison between the semantic information transmission method disclosed herein and the index amount transmitted in the prior art. DETAILED DESCRIPTION

[0048] In order to make the purpose, technical solutions and advantages of the present disclosure more clearly understood, the present disclosure is further described in detail below in conjunction with the embodiments and drawings. Here, the illustrative embodiments of the present disclosure and their descriptions are used to explain the present disclosure, but are not intended to limit the present disclosure.

[0049] It should also be noted here that in order to avoid obscuring the present disclosure due to unnecessary details, only structures and / or processing steps closely related to the solutions according to the present disclosure are shown in the accompanying drawings, while other details that are not closely related to the present disclosure are omitted.

[0050] It should be emphasized that the term "include / comprises" when used herein refers to the existence of features, elements, steps or components, but does not exclude the existence or addition of one or more other features, elements, steps or components.

[0051] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. In the accompanying drawings, the same reference numerals represent the same or similar components, or the same or similar steps.

[0052] In response to the problem of video image compression during semantic communication, and how to restore the compressed video images and improve the accuracy of video image restoration, and even restore the transmitted original video images and original video, the embodiments of the present disclosure disclose a semantic information compression method, a semantic information transmission method and a video image restoration method.

[0053] FIG1 is a flow chart of a semantic information compression method according to an embodiment of the present disclosure. Referring to FIG1 , the semantic information compression method according to an embodiment of the present disclosure includes the following steps:

[0054] In step S110, a latent feature extractor extracts latent features of each frame image from the compressed video, and generates multiple latent feature maps accordingly. The latent feature maps are used as semantic information. Each latent feature map includes multiple channels, and the multiple channels of different latent feature maps correspond to each other one by one.

[0055] In an embodiment of the present disclosure, the latent features of each frame image extracted from the video to be compressed in step S110 constitute their respective latent feature maps, which can be represented by a latent feature matrix, that is, the latent feature matrix of each frame image can also serve as their respective semantic information. Since each frame image in the input video to be compressed comes from the same background knowledge base, the semantic information obtained by the latent feature extractor comes from the same information representation space, so the number of channels of the latent feature map of each frame image in the video to be compressed is the same and corresponds one to one. In one embodiment of the present disclosure, the channel of the latent feature map can be each row of the latent feature matrix.

[0056] In an optional embodiment of the present disclosure, the potential feature extractor in step S110 includes multiple convolutional layers and multiple GDN activation functions, wherein a GDN activation function is connected between every two convolutional layers, the first convolutional layer serves as the input of the potential feature extractor, and the last convolutional layer serves as the output of the potential feature extractor. The multiple convolutional layers are used to downsample each frame image in the compressed video, generate multiple potential feature maps accordingly, increase the number of channels of the image, and reduce the size of the potential feature map of each channel.

[0057] In the embodiment of the present disclosure, the potential feature map contained in each channel is the channel feature of the respective channel.

[0058] For example, see Figure 2. The latent feature extractor includes four convolutional layers and three GDN activation functions. A GDN activation function is connected between every two convolutional layers. The first convolutional layer is the input of the latent feature extractor, and the last convolutional layer is the output of the latent feature extractor. Two adjacent frames of the video to be compressed with a size of 3*512*1024 are input to the first convolutional layer of the latent feature extractor. The four convolutional layers all downsample the input images of each layer. The size of the two latent feature maps output by the fourth convolutional layer is 64*64*128, so that the latent feature map output by the last convolutional layer has an increased number of channels (from 3 to 64) compared to the image input to the first convolutional layer, and the size of the feature map on each channel is reduced (from 512*1024 to 64*128). Similarly, the other two adjacent frames of the video to be compressed are processed in the same way, which will not be repeated here.

[0059] Step S120 , calculating the correlation between multiple corresponding channels of the potential feature map of each two adjacent frames of images based on the semantic information of each two adjacent frames of images, and obtaining multiple correlation values.

[0060] In the disclosed embodiment, since each frame image in the input video to be compressed comes from the same background knowledge base, the semantic information obtained by the latent feature extractor comes from the same information representation space and has certain similarities. Therefore, the latent features or semantic information of the video image can be compressed based on the similarity between each two adjacent frames of images. By calculating the correlation between multiple corresponding channels (i.e., each row of the latent feature matrix) in the latent feature map of each two adjacent frames of images, the dimension of the information representation space can be effectively reduced.

[0061] In an optional embodiment of the present disclosure, referring to FIG. 3 , the specific step of calculating the correlation between multiple corresponding channels of the potential feature map of each of the two adjacent frames of image based on the semantic information of each of the two adjacent frames of image in step S120 further includes:

[0062] Step S121, converting the latent feature maps of multiple corresponding channels in the latent feature maps of each of two adjacent frames of images into multiple pairs of semantic vectors, wherein the two semantic vectors of each corresponding channel form a pair of semantic vectors;

[0063] Step S122 : calculating the Pearson correlation coefficient between each pair of semantic vectors to obtain the correlation between multiple corresponding channels of the latent feature map of each two adjacent frames of images.

[0064] In the embodiment of the present disclosure, step S121 is to expand each row of the two potential feature matrices of each two adjacent frames of images and convert them into multiple pairs of one-dimensional semantic vectors. In step S122, the Pearson correlation coefficient between each pair of one-dimensional semantic vectors is calculated according to the following formula:

[0065] Among them, r represents the Pearson correlation coefficient, X i and Y i Represent the elements of the two semantic vectors in each pair of one-dimensional semantic vectors, and represents the average value of the elements in the two semantic vectors in each pair of one-dimensional semantic vectors, i = 1, ..., n, where i represents the index of the semantic vector and n represents the length of the semantic vector. The larger the calculated Pearson correlation coefficient, the greater the correlation.

[0066] For example, the correlation between two latent feature maps of size 64*64*128 output by the latent feature extractor is calculated. Since the two latent feature maps have 64 channels, 64 correlations need to be calculated to obtain 64 Pearson correlation coefficients. Similarly, 64 correlations need to be calculated for each of the other two adjacent frames in the compressed video.

[0067] In step S130, the values ​​of multiple correlations are sorted from large to small, and the channel merging module selects multiple corresponding channels of the potential feature maps of each two adjacent frames of images corresponding to the correlations ranked first based on a preset compression ratio to merge to form a shared channel. The semantic information corresponding to the shared channel and the respective non-shared channels constitute the semantic information of each two adjacent frames of images after compression, thereby achieving compression of the semantic information in the video to be compressed.

[0068] In the disclosed embodiment, the channel features contained in each channel in the shared channel (Channel) - the latent feature map (semantic information) are shared latent features. Non-shared channels are channels that are not merged in the latent feature map of each of two adjacent frames of images.

[0069] Exemplarily, the preset compression ratio is 50%, and the channel merging module is used to sort the values ​​of the 64 Pearson correlation coefficients of the two potential feature maps of size 64*64*128 from large to small, and the corresponding channels of the two potential feature maps corresponding to the 32 Pearson correlation coefficients ranked first (certain rows corresponding to the two potential feature matrices) are selected for merging, and the potential feature maps of these channels are merged as shared potential features, and the size of the merged feature map is output as 96*64*128, so as to realize the compression of two adjacent frames of the video to be compressed with a size of 3*512*1024. Please refer to Figure 4, the feature map 1 (top) and the feature map 2 (bottom) on the left are schematic diagrams of two potential feature maps of size 64*64*128 respectively, and the figure on the right after the feature map 1 and the feature map 2 are merged is a schematic diagram of a feature map of size 96*64*128, and each row in these three figures represents a channel.

[0070] As can be seen, the semantic information compression method of the disclosed embodiment extracts shared latent features based on the correlation of feature maps or feature matrices, merges them along the channel dimension of the feature maps or feature matrices, and thereby achieves semantic information compression, thereby improving compression efficiency and achieving excellent compression performance. By transmitting the semantic information of video images compressed by this semantic information compression method, the receiving end can restore the image and video without considering the index of each transmitted channel, thereby improving communication stability.

[0071] FIG5 is a schematic diagram of a semantic information transmission method provided by an embodiment of the present disclosure. Referring to FIG5 , the steps of a semantic information transmission method according to an embodiment of the present disclosure include:

[0072] In step S210 , the transmitting end compresses the semantic information of each frame image in the video to be transmitted using the aforementioned semantic information compression method.

[0073] Exemplarily, the semantic information compression method shown in Figure 1 is used to compress each frame image of a size of 3*512*1024 in the video to be transmitted, and each two adjacent frames of images are compressed into multiple feature maps of a size of 96*64*128 through the potential feature extractor and the channel merging module.

[0074] In step S220 , the source-channel joint encoder encodes the semantic information of each frame of the compressed video to be transmitted to obtain a semantic code.

[0075] In an optional embodiment of the present disclosure, the source-channel joint encoder in step S220 includes multiple convolutional layers and multiple GDN activation functions, wherein a GDN activation function is connected between every two convolutional layers, the first convolutional layer serves as the input of the source-channel joint encoder, and the last convolutional layer serves as the output of the source-channel joint encoder, and the multiple convolutional layers are used to encode semantic information.

[0076] In the disclosed embodiment, the source-channel joint encoder can improve the noise resistance of semantic information to cope with interference or noise in the transmission channel, while the size of the encoded feature map is the same as that before encoding (without changing the number and size of image channels).

[0077] For example, see Figure 6. The joint source-channel encoder includes two convolutional layers and a GDN activation function. The GDN activation function is connected to each of the two convolutional layers, one of which serves as the input and the other as the output of the joint source-channel encoder. A feature map of size 96*64*128 is input to the joint source-channel encoder for encoding, and the output encoded feature map still has a size of 96*64*128. The joint source-channel encoder does not resize the feature map, which improves noise immunity.

[0078] In step S230 , the channel transmits the compressed semantic code of each frame of the video to be transmitted.

[0079] In an optional embodiment of the present disclosure, the step of transmitting the semantic encoding of each frame of the compressed video to be transmitted via the channel in step S230 further includes transmitting the index of the shared channel of the latent feature map of each frame of the compressed video to be transmitted. That is, the two latent feature maps of each two adjacent frames of the video to be transmitted are merged into a single image via the shared channel, thereby achieving compression of the video semantic information. The index of this shared channel can also be transmitted via the channel.

[0080] FIG7 is a schematic diagram of a video image restoration method provided by an embodiment of the present disclosure. Referring to FIG7 , the steps of a video image restoration method according to an embodiment of the present disclosure include:

[0081] In step S310, the receiving end receives the semantic coding of each frame image in the compressed video from the sending end, and the source channel joint decoder decodes the semantic coding to obtain semantic information of each frame image in the compressed video, wherein the semantic information includes the compressed shared channel and the respective non-shared channels of the two adjacent frame images before compression corresponding to each frame image, and the respective non-shared channels are respectively a first non-shared channel and a second non-shared channel.

[0082] In an embodiment of the present disclosure, a receiving end receives the semantic coding of each frame image in a compressed video transmitted by a sending end through the aforementioned semantic information transmission method, wherein the compressed video is obtained through the aforementioned semantic information compression method, and each frame image in the compressed video is obtained by channel merging of the corresponding latent feature maps of two adjacent frames before compression, the first non-shared channel is the channel of the latent feature map of one of the two adjacent frames that has not been merged or compressed, and the second non-shared channel is the channel of the latent feature map of the other of the two adjacent frames that has not been merged or compressed.

[0083] In an optional embodiment of the present disclosure, the source-channel joint decoder in step S310 includes multiple convolutional layers and multiple GDN activation functions, wherein a GDN activation function is connected between every two convolutional layers, the first convolutional layer serves as the input of the source-channel joint decoder, and the last convolutional layer serves as the output of the source-channel joint decoder, and the multiple convolutional layers are used to decode the semantic code.

[0084] In the disclosed embodiment, the source-channel joint decoder can further improve the noise resistance of semantic information, and the size of the decoded feature map is the same as the size of the encoded feature map (without changing the number and size of image channels).

[0085] For example, as shown in Figure 6, the source-channel joint decoder includes two convolutional layers and a GDN activation function. The GDN activation function is connected to each of the two convolutional layers, one of which serves as the input and the other as the output of the source-channel joint decoder. The transmitted encoded feature map of size 96*64*128 is input into the source-channel joint decoder for decoding, and the decoded output feature map still has a size of 96*64*128. The source-channel joint decoder does not resize the feature map, which improves noise immunity.

[0086] Step S320 : A channel recovery module extracts the shared channel, the first non-shared channel, and the second non-shared channel from the semantic information of each frame of the compressed video based on a preset compression ratio.

[0087] In step S330 , the semantic information of the shared channel is combined with the semantic information corresponding to the first non-shared channel and the second non-shared channel to restore the semantic information of each two adjacent frames before compression, thereby obtaining the semantic information of each frame in the decompressed video.

[0088] Exemplarily, the decoded feature map of size 96*64*128 is input into the channel recovery module, and the preset compression ratio is 50%. The channel recovery module is used to extract the 32 shared channels of the two corresponding potential feature maps before compression and the remaining 32 channels of each of the two potential feature maps excluding the shared channels from the feature map. The remaining 32 channels excluding the shared channels and the potential feature map parts corresponding to the 32 shared channels respectively form two potential feature maps of size 64*64*128, thereby achieving decompression. For the schematic diagram of the two potential feature maps of the two adjacent frames of the decompressed image of size 3*512*1024 recovered here, please refer to the feature map 1 (top) and feature map 2 (bottom) on the right side of Figure 8. Compared with the original potential feature maps output by the potential feature extractor (see the schematic diagram of feature map 1 (top) and feature map 2 (bottom) on the left side of Figure 4), the order of the channels of these two potential feature maps has changed.

[0089] In step S340 , the latent feature restorer restores each frame of the video based on the semantic information of each frame of the decompressed video, thereby obtaining the transmitted video.

[0090] In an optional embodiment of the present disclosure, the potential feature restorer in step S340 includes multiple transposed convolutional layers, multiple GDN activation functions, convolutional layers and Tanh activation functions, wherein the number of transposed convolutional layers and GDN activation functions is the same, a GDN activation function is connected between every two transposed convolutional layers, and the output of the last GDN activation function is connected to the convolutional layer and the Tanh activation function in sequence. The first transposed convolutional layer serves as the input of the potential feature restorer, and the Tanh activation function serves as the output of the potential feature restorer. The multiple transposed convolutional layers are used to upsample the semantic information of each frame image in the decompressed video, reduce the number of image channels, and increase the size of the image of each channel.

[0091] For example, see Figure 9. The latent feature restorer includes three transposed convolutional layers, three GDN activation functions, a convolutional layer, and a Tanh activation function. A GDN activation function is connected between every two transposed convolutional layers. The output of the third GDN activation function is sequentially connected to the convolutional layer and the Tanh activation function. The first transposed convolutional layer serves as the input of the latent feature restorer, and the Tanh activation function serves as the output of the latent feature restorer. The two latent feature maps of size 64*64*128 with the channel order changed after decompression are input into the first transposed convolutional layer of the latent feature restorer as input. After upsampling through the three transposed convolutional layers, two frames of image size 3*512*1024 are finally restored and output. The difference between the two restored images of size 3*512*1024 and the two original images input to the latent feature extractor is only that the image quality is different.

[0092] It can be seen that the semantic information transmission method of the embodiment of the present disclosure does not need to transmit the index of each channel in the latent feature map of the image (that is, the index of each semantic vector in the latent feature matrix), so that the video image restoration method of the embodiment of the present disclosure can restore the semantic information of each frame image of the transmitted video. The semantic information of each frame image of the restored video is the same as the original image of each frame of the original video, which can ensure the accuracy, precision and stability of the semantic information of each frame image of the restored video. At the same time, it can avoid the problem of incorrect transmitted index due to the need to transmit the index of each semantic vector in the latent feature matrix, and the security, reliability and stability of information transmission are high.

[0093] The prior art calculates the entropy of the feature vector from the perspective of information entropy, and judges the importance of the vector elements based on the size of the entropy. Information with low entropy is not transmitted during the communication transmission process, and the transmitted data is compressed by transmitting the index of the vector elements. Figure 11 is a schematic diagram comparing the effects of the video image restoration method of the embodiment of the present disclosure and the restoration of the transmitted image in the prior art. Please refer to Figure 11. The two images in the first row are the original images, the two images in the second row are the transmission images restored by the video image restoration method of the embodiment of the present disclosure, and the two images in the third row are the transmission images restored by the prior art. It can be seen that due to the index error in the transmission process of the prior art, the aliasing of the transmitted image is caused by the aliasing of the transmitted image, while the semantic information transmission method of the embodiment of the present disclosure does not need to transmit the index of each channel, which can improve the stability of the communication transmission and improve the quality of the video image restoration method for video images. Figure 12 is a schematic diagram comparing the semantic information transmission method of the embodiment of the present disclosure and the amount of index transmitted in the prior art. Please refer to Figure 12, where PC-HEM represents the index amount transmitted by the semantic information transmission method of the embodiment of the present disclosure, and ED-HEM represents the index amount transmitted by the prior art for images of different sizes. It can be seen that compared with the prior art, the index transmission amount of the semantic information transmission method of the embodiment of the present disclosure is 0, which saves resources, greatly reduces the transmission amount, and improves the transmission efficiency.

[0094] In an optional embodiment of the present disclosure, when the above-mentioned semantic information transmission method transmits the index of the shared channel through the channel, the index is received by the receiving end. The channel recovery module in the video image restoration method of the embodiment of the present disclosure can also recover the semantic information of each frame image before merging the shared channel using the above-mentioned semantic information compression method according to the index of the shared channel, that is, the original potential feature map obtained by the potential feature extractor.

[0095] It can be seen that the semantic information transmission method of the embodiment of the present disclosure transmits the index of the shared channel of each adjacent two frames of images and the semantic coding of each frame of images in the compressed video to be transmitted to the receiving end, so that the video image restoration method of the embodiment of the present disclosure uses the channel recovery module to perform channel recovery on the semantic information of each frame of images in the compressed video according to the index of the shared channel to obtain the potential feature map of each frame of images in the transmitted original video, and further enables the potential feature restorer to restore the original image of each frame of the original video according to the original potential feature map restored by the channel recovery module, thereby restoring the original video.

[0096] Corresponding to the above method, an electronic device of an embodiment of the present disclosure includes a computer device, the computer device includes a processor and a memory, the memory stores computer instructions, the processor is used to execute the computer instructions stored in the memory, and when the computer instructions are executed by the processor, the electronic device implements the steps of the above-mentioned semantic information compression method, semantic information transmission method and video image restoration method.

[0097] In the disclosed embodiment, referring to FIG10 , the above-mentioned semantic information compression, transmission, and video image restoration process is implemented by utilizing a latent feature extractor, a correlation calculation module (not shown in the figure), a channel merging module, a source-channel joint encoder, a channel, a source-channel joint decoder, a channel recovery module, and a latent feature restorer. The latent feature extractor may be an autoencoder or a transformer encoder, and the latent feature restorer may be an autodecoder or a transformer decoder.

[0098] The present disclosure also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the aforementioned semantic information compression method, semantic information transmission method, and video image restoration method. The computer-readable storage medium can be a tangible storage medium, such as a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a floppy disk, a hard disk, a removable storage disk, a CD-ROM, or any other form of storage medium known in the art.

[0099] It should be understood by those skilled in the art that the various exemplary components, systems and methods described in conjunction with the embodiments disclosed herein can be implemented in hardware, software or a combination of the two. Whether it is specifically performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this disclosure. When implemented in hardware, it can be, for example, an electronic circuit, an application specific integrated circuit (ASIC), appropriate firmware, a plug-in, a function card, etc. When implemented in software, the elements of this disclosure are programs or code segments that are used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted on a transmission medium or a communication link via a data signal carried in a carrier.

[0100] It should be understood that the present disclosure is not limited to the specific configurations and processes described above and illustrated in the figures. For the sake of brevity, a detailed description of known methods is omitted here. In the above embodiments, several specific steps are described and illustrated as examples. However, the method process of the present disclosure is not limited to the specific steps described and illustrated. Those skilled in the art can make various changes, modifications, and additions, or change the order of the steps after understanding the spirit of the present disclosure.

[0101] In the present disclosure, features described and / or illustrated for one embodiment may be used in the same or similar manner in one or more other embodiments, and / or combined with or replace features of other embodiments.

[0102] The foregoing description is merely a preferred embodiment of the present disclosure and is not intended to limit the present disclosure. Those skilled in the art will readily appreciate that various modifications and variations of the present disclosure are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present disclosure shall be included within the scope of protection of the present disclosure.

Claims

1. A semantic information compression method, characterized in that: The method comprises: A latent feature extractor extracts latent features of each frame image from the compressed video, and generates a plurality of latent feature maps accordingly. The latent feature maps are used as semantic information. Each latent feature map includes a plurality of channels, and the plurality of channels of different latent feature maps correspond to each other one by one. Calculating correlations between multiple corresponding channels of potential feature maps of each of the two adjacent frames of images based on semantic information of each of the two adjacent frames of images, and obtaining multiple correlation values; The values ​​of multiple correlations are sorted from large to small, and a channel merging module selects multiple corresponding channels of the potential feature map of each two adjacent frames of images corresponding to the correlations sorted first based on a preset compression ratio to merge to form a shared channel. The semantic information corresponding to the shared channel and the respective non-shared channels constitutes the semantic information of each two adjacent frames of images after compression, thereby achieving compression of the semantic information in the video to be compressed.

2. The method according to claim 1, characterized in that The step of calculating the correlation between multiple corresponding channels of the potential feature map of each two adjacent frames of images based on the semantic information of each two adjacent frames of images includes: The latent feature maps of multiple corresponding channels in the latent feature map of each of two adjacent frames of images are converted into multiple pairs of semantic vectors, wherein the two semantic vectors of each corresponding channel form a pair of semantic vectors; The Pearson correlation coefficient between each pair of semantic vectors is calculated to obtain the correlation between multiple corresponding channels of the potential feature map of each two adjacent frames of images.

3. The method according to claim 1, characterized in that The feature extractor includes multiple convolutional layers and multiple GDN activation functions, wherein a GDN activation function is connected between every two convolutional layers, the first convolutional layer serves as the input of the feature extractor, and the last convolutional layer serves as the output of the feature extractor. The multiple convolutional layers are used to downsample each frame image of the video to be compressed, and correspondingly generate multiple potential feature maps, increase the number of channels of the image, and reduce the size of the potential feature map of each channel.

4. A semantic information transmission method, characterized in that: The method comprises: The sending end compresses the semantic information of each frame image in the video to be transmitted using the method as described in any one of claims 1 to 3; The source-channel joint encoder encodes the semantic information of each frame image in the compressed video to be transmitted to obtain a semantic code; The channel transmits the semantic coding of each frame image in the compressed video to be transmitted.

5. The method according to claim 4, characterized in that The source-channel joint encoder includes multiple convolutional layers and multiple GDN activation functions, wherein a GDN activation function is connected between every two convolutional layers, the first convolutional layer serves as the input of the source-channel joint encoder, and the last convolutional layer serves as the output of the source-channel joint encoder, and the multiple convolutional layers are used to encode semantic information.

6. The method according to claim 4, characterized in that The step of transmitting the semantic coding of each frame image in the compressed video to be transmitted through the channel also includes: transmitting the index of the shared channel of the potential feature map of each frame image in the compressed video to be transmitted.

7. A video image restoration method, characterized in that: The method comprises: The receiving end receives the semantic coding of each frame image in the compressed video from the sending end, and the source channel joint decoder decodes the semantic coding to obtain semantic information of each frame image in the compressed video, wherein the semantic information includes the compressed shared channel and respective non-shared channels of two adjacent frame images before compression corresponding to each frame image, and the respective non-shared channels are respectively a first non-shared channel and a second non-shared channel; The channel recovery module extracts the shared channel, the first non-shared channel and the second non-shared channel from the semantic information of each frame of the compressed video based on a preset compression ratio; The semantic information of the shared channel is combined with the semantic information corresponding to the first non-shared channel and the second non-shared channel respectively to restore the semantic information of each two adjacent frames of images before compression, thereby obtaining the semantic information of each frame of the image in the decompressed video; The latent feature restorer restores each frame of the video based on the semantic information of each frame in the decompressed video, thereby obtaining the transmitted video.

8. The method according to claim 7, characterized in that The source-channel joint decoder comprises a plurality of convolutional layers and a plurality of GDN activation functions, wherein a GDN activation function is connected between every two convolutional layers, the first convolutional layer is used as the input of the source-channel joint decoder, and the last convolutional layer is used as the output of the source-channel joint decoder, and the plurality of convolutional layers of the source-channel joint decoder are used to decode the semantic coding; The potential feature restorer includes multiple transposed convolutional layers, multiple GDN activation functions, convolutional layers and Tanh activation functions, wherein the number of the transposed convolutional layers and the GDN activation functions is the same, a GDN activation function is connected between every two transposed convolutional layers, the output of the last GDN activation function is sequentially connected to the convolutional layer and the Tanh activation function, the first transposed convolutional layer serves as the input of the potential feature restorer, and the Tanh activation function serves as the output of the potential feature restorer, and the multiple transposed convolutional layers are used to upsample the semantic information of each frame image in the decompressed video, reduce the number of channels of the image, and increase the size of the image of each channel.

9. An electronic device comprising a processor and a memory, characterized in that: The memory stores computer instructions, and the processor is used to execute the computer instructions stored in the memory. When the computer instructions are executed by the processor, the electronic device implements the steps of the method as claimed in any one of claims 1 to 8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Image coding method, image decoding method and image compression method based on context recombination modeling

    CN113747163A

  • Coding method and decoding method of feature data, equipment and storage medium

    CN116868570A

  • Code rate adaptive video semantic communication method and related device

    CN116896651A

  • Image encoder and image decoder

    JP1998341433A