Semantic video transmission method and device, electronic equipment and storage medium

By dividing semantic video into multiple segments and dynamically adjusting the compression ratio according to network conditions, and generating incremental data, the problem of high-quality video transmission in the semantic video communication system when the conditions of wireless network fluctuate, achieving efficient video reconstruction effect.

CN120281906APending Publication Date: 2025-07-08BEIJING UNIV OF POSTS & TELECOMM +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311830762.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-28
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing semantic video communication systems have shortcomings in taking into account high semantic accuracy and bandwidth utilization, especially when the conditions of wireless networks fluctuate, it is difficult to meet the needs of high-quality video transmission.

Method used

The semantic video to be transmitted is divided into multiple video clips, and encoded in a first compression ratio and transmitted through the wireless channel. At the same time, the second compression ratio is determined based on the current network condition of the wireless channel, and incremental data of the video clip is generated to realize dynamic adjustment of the video code rate and high-quality reconstruction.

Benefits of technology

By dynamically adjusting the video compression ratio, high-quality video content can be restored at the receiving end, meeting users' video transmission needs, adapting to network fluctuations, and improving the stability and quality of video transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120281906A_ABST
    Figure CN120281906A_ABST
Patent Text Reader

Abstract

The invention provides a semantic video transmission method and device, electronic equipment and a storage medium, and relates to the technical field of communication. According to the specific implementation scheme, a semantic video to be transmitted is divided, and at least one first video clip and at least one second video clip are obtained; performing coding processing on the first video clip and the second video clip according to a first compression ratio, and transmitting a coding processing result to a receiving end through a wireless channel to obtain a first coded video clip and a second coded video clip; determining a second compression ratio according to the current network condition of the wireless channel, processing the second video clip according to the second compression ratio, obtaining incremental data corresponding to the second video clip, and transmitting the incremental data to the receiving end; and the receiving end performs video reconstruction operation according to the first coded video clip, the second coded video clip and the incremental data corresponding to the second video clip to obtain a reconstructed video. According to the invention, the video with high quality can be recovered.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of communication technologies, specifically to the field of semantic video communication technologies, and particularly to a semantic video transmission method, apparatus, electronic device, and storage medium. Background Art

[0002] Semantic video communication technology is a video communication technology based on semantic understanding and analysis. It aims to extract and transmit semantic information in video content to achieve a more intelligent, efficient, and interactive video communication experience. Existing semantic video communication has a wide range of applications in multiple fields, including video conferencing, distance education, video surveillance, augmented reality, and virtual reality, etc.

[0003] Semantic video communication technology can effectively overcome problems faced by traditional video communication to a certain extent, such as bandwidth limitations, high latency, and low-quality transmission. However, without combining with the adaptive bitrate (ABR) algorithm, the video semantic communication system cannot well balance high semantic accuracy and bandwidth utilization. In addition, due to the higher requirements of video semantic communication for video quality and latency control, it is necessary to design an effective adaptive bitrate algorithm to adapt to different network conditions, meet the requirements of the communication system for high semantic accuracy and low latency, and also meet the requirements of users for video service quality. Summary of the Invention

[0004] The present disclosure provides a semantic video transmission method, apparatus, electronic device, and storage medium.

[0005] According to one aspect of the present disclosure, there is provided a semantic video transmission method, including:

[0006] Dividing a semantic video to be transmitted to obtain a plurality of video segments, where the plurality of video segments include at least one first video segment and at least one second video segment;

[0007] Encoding the first video segment and the second video segment both at a first compression ratio, and transmitting the encoding result through a wireless channel to a receiving end to obtain a first encoded video segment and a second encoded video segment;

[0008] Determining a second compression ratio according to the current network condition of the wireless channel, and processing the second video segment according to the second compression ratio to obtain incremental data corresponding to the second video segment and transmitting it to the receiving end, where the first compression ratio is greater than the second compression ratio;

[0009] The receiving end performs a video reconstruction operation according to the first encoded video segment, the second encoded video segment, and the incremental data corresponding to the second video segment to obtain a reconstructed semantic video.

[0010] According to one aspect of the present disclosure, there is provided a semantic video transmission method, including:

[0011] Receiving a first encoded video segment and a second encoded video segment transmitted from a sending end through a wireless channel, where the sending end obtains a semantic video to be transmitted, divides the semantic video to be transmitted to obtain at least one first video segment and at least one second video segment, and encodes both the first video segment and the second video segment at a first compression ratio to obtain the first encoded video segment and the second encoded video segment;

[0012] Receiving incremental data corresponding to the second video segment sent from the sending end, where the sending end determines a second compression ratio according to the current network condition of the wireless channel, and processes the second video segment according to the second compression ratio to obtain incremental data corresponding to the second video segment, where the first compression ratio is greater than the second compression ratio;

[0013] Performing a video reconstruction operation according to the first encoded video segment, the second encoded video segment, and the incremental data corresponding to the second video segment to obtain a reconstructed semantic video.

[0014] According to another aspect of the present disclosure, there is provided a semantic video transmission device, including:

[0015] A dividing module, configured to divide a semantic video to be transmitted to obtain a plurality of video segments, where the plurality of video segments include at least one first video segment and at least one second video segment;

[0016] An encoding module, configured to encode both the first video segment and the second video segment at a first compression ratio, and transmit the encoding result through a wireless channel to a receiving end to obtain a first encoded video segment and a second encoded video segment;

[0017] An incremental module, configured to determine a second compression ratio according to the current network condition of the wireless channel, and process the second video segment according to the second compression ratio to obtain incremental data corresponding to the second video segment and transmit it to the receiving end, where the first compression ratio is greater than the second compression ratio;

[0018] The receiving end performs a video reconstruction operation according to the first encoded video segment, the second encoded video segment, and the incremental data corresponding to the second video segment to obtain a reconstructed semantic video.

[0019] According to another aspect of the present disclosure, there is provided a semantic video transmission device, including:

[0020] A first receiving module, configured to receive a first encoded video segment and a second encoded video segment transmitted from a sending end via a wireless channel. The sending end obtains a semantic video to be transmitted, divides the semantic video to be transmitted to obtain at least one first video segment and at least one second video segment, and encodes both the first video segment and the second video segment at a first compression ratio to obtain the first encoded video segment and the second encoded video segment;

[0021] A second receiving module, configured to receive incremental data corresponding to the second video segment sent from the sending end. The sending end determines a second compression ratio according to the current network condition of the wireless channel, and processes the second video segment according to the second compression ratio to obtain incremental data corresponding to the second video segment, where the first compression ratio is greater than the second compression ratio;

[0022] A reconstruction module, configured to perform a video reconstruction operation according to the first encoded video segment, the second encoded video segment, and the incremental data corresponding to the second video segment to obtain a reconstructed semantic video.

[0023] According to a fifth aspect of the present disclosure, there is provided an electronic device, including:

[0024] At least one processor; and

[0025] A memory communicatively connected to the at least one processor; wherein,

[0026] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method described in any one of the above technical solutions.

[0027] According to a sixth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, where the computer instructions are used to cause the computer to execute the method described in any one of the above technical solutions.

[0028] According to a seventh aspect of the present disclosure, there is provided a computer program product, including a computer program, where the computer program, when executed by a processor, implements the method described in any one of the above technical solutions.

[0029] The present disclosure provides a semantic video transmission method, apparatus, device, and storage medium. For other video segments except the first video segment, after encoding and processing at the first compression ratio, a suitable second compression ratio is selected according to the current network condition of the wireless channel, thereby realizing dynamic adjustment of the video bit rate, and further an incremental data corresponding to the video segment can be obtained, which is further beneficial to restoring a video segment with higher quality at the receiving end. Among them, taking the second video segment as an example, after encoding and processing at the first compression ratio, the second compression ratio is determined according to the current network condition of the wireless channel, and the second video segment is processed according to the second compression ratio to obtain the corresponding incremental data, so that as much information of the second video segment as possible can be obtained, and a second video segment with higher quality can be restored when restoring the second video segment. Therefore, the solution of the present disclosure selects a suitable video compression ratio for transmission according to the real-time network fluctuation condition, realizes dynamic adjustment of the video transmission bit rate, can restore high-quality video content at the receiving end, and meets the user's video transmission requirements.

[0030] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0032] Figure 1 is a schematic diagram of the steps of the semantic video transmission method in an embodiment of the present disclosure;

[0033] Figure 2 is a schematic diagram of the steps of the semantic video transmission method in an embodiment of the present disclosure;

[0034] Figure 3 is a schematic diagram of the system corresponding to the semantic video transmission method in an embodiment of the present disclosure;

[0035] Figure 4 is a schematic diagram of the processing process of the basic video data and the incremental video data in an embodiment of the present disclosure;

[0036] Figure 5 The principle block diagram of the semantic video transmission apparatus in an embodiment of the present disclosure;

[0037] Figure 6 The principle block diagram of the semantic video transmission apparatus in an embodiment of the present disclosure;

[0038] Figure 7 is a block diagram of an electronic device for implementing the semantic video transmission method in an embodiment of the present disclosure. Detailed implementation manners

[0039] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to assist understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and structures are omitted below for clarity and conciseness.

[0040] In order to achieve high-quality video content distribution under the current network conditions of users, dynamic adaptive technologies with adaptive bitrate algorithms as the core are widely used by content providers. The core idea of the adaptive bitrate algorithm is to evaluate the current network state in real time according to network conditions (such as network throughput, buffer size, etc.), and then select the most suitable bitrate version from multiple preset bitrates for video stream encoding and transmission. If the network conditions improve, a higher bitrate is selected for transmission. If the network conditions deteriorate, a lower bitrate video is transmitted. Currently, the widely adopted adaptive bitrate algorithm models mainly include four categories: based on network throughput, based on buffer, hybrid model, and based on deep learning. These existing adaptive bitrate algorithms have achieved high-quality video recovery to a certain extent under network fluctuations and have obtained relatively high user experience values.

[0041] For semantic video communication technology, there is currently no video semantic communication system combined with an adaptive bitrate (ABR) algorithm to balance high semantic accuracy and bandwidth utilization. Most traditional adaptive bitrate algorithms make decisions based on future network predictions. However, due to the sharp increase in the number of terminal devices in wireless scenarios, there are large fluctuations in network throughput, which poses a huge challenge to improving the quality of stable and high-quality Internet services.

[0042] To solve the above technical problems, the present disclosure provides a semantic video transmission method, which is applied to a sending end. Refer to Figure 1 as shown, including:

[0043] Step S101: Divide the semantic video to be transmitted to obtain multiple video segments, where the multiple video segments include at least one first video segment and at least one second video segment.

[0044] Specifically, after determining the semantic video to be transmitted, the semantic video to be transmitted is divided, so that multiple video segments can be obtained. The multiple video segments include, for example, a first video segment, a second video segment, a third video segment... and an Nth video segment. When dividing the semantic video to be transmitted, it can be divided into equal lengths or unequal lengths, which is not limited here. Those skilled in the art can divide it according to actual needs. For example, if the semantic video to be transmitted is V, the semantic video to be transmitted V can be divided into multiple video segments of equal length {v1, v2, v3,..., v N}.

[0045] Optionally, in the embodiments of the present disclosure, at least one first video segment and at least one second video segment are included in the multiple video segments, that is, there can be multiple first video segments and multiple second video segments in the multiple video segments. This solution can be approximately considered as dividing the semantic video to be transmitted into multiple video segment groups, and each video segment group contains a first video segment and a second video segment. In this way, in order to improve the data processing efficiency, the data of different video segment groups can be processed synchronously, that is, parallel processing can be achieved for multiple different first video segments and second video segments, which is conducive to improving the efficiency of video reconstruction.

[0046] Step S102: Encode both the first video segment and the second video segment at a first compression ratio, and transmit the encoding result to the receiving end through a wireless channel to obtain a first encoded video segment and a second encoded video segment.

[0047] Specifically, after obtaining multiple video segments, the first video segment and the second video segment are selected from the multiple video segments. The first video segment can be considered as the first video segment, and the second video segment is the next video segment after the first segment. Taking the above example, v1 is used as the first video segment, and v2 is used as the second video segment. After obtaining the first video segment and the second video segment, both the first video segment and the second video segment are encoded at a first compression ratio to obtain a first encoding result and a second encoding result. Then, the first encoding result and the second encoding result are transmitted to the receiving end through a wireless channel to obtain a first encoded video segment and a second encoded video segment.

[0048] Step S103: Determine a second compression ratio according to the current network condition of the wireless channel, and process the second video segment according to the second compression ratio to obtain incremental data corresponding to the second video segment and transmit it to the receiving end, where the first compression ratio is greater than the second compression ratio.

[0049] Specifically, after transmitting the encoded processing result to the receiving end through the wireless channel to obtain the first encoded video segment and the second encoded video segment, the current network condition of the wireless channel is acquired. The second compression ratio is determined according to the current network condition of the wireless channel, and then the second video segment is processed according to the second compression ratio to obtain the incremental data corresponding to the second video segment. In this way, determining the second compression ratio according to the current network condition of the wireless channel and processing the second video segment according to the second compression ratio to obtain the corresponding incremental data is to obtain as much information of the second video segment as possible, so as to facilitate the subsequent restoration to obtain a second video segment with relatively high quality. It should be noted that, for example, for the third video segment, the fourth video segment until the Nth video segment, it is necessary to calculate the incremental data corresponding to each video segment. Since the calculation method of the incremental data corresponding to each video segment is similar to the method for determining the incremental data of the second video segment, it will not be elaborated here.

[0050] Step S104, the receiving end performs video reconstruction operations based on the first encoded video segment, the second encoded video segment, and the incremental data corresponding to the second video segment to obtain the reconstructed semantic video.

[0051] Specifically, after obtaining the first encoded video segment, the second encoded video segment, and the incremental data corresponding to the second video segment, the incremental data corresponding to other video segments is also calculated for other video segments. Since the calculation method is similar to the method for determining the incremental data of the second video segment, it will not be repeated here. When the receiving end performs video reconstruction, after each video segment is reconstructed, the reconstructed semantic video is finally output. For example, in the order of reception, the first encoded video segment is reconstructed, the second encoded video segment and the incremental data corresponding to the second video segment are merged in the buffer and then reconstructed, the third encoded video segment and the incremental data corresponding to the third video segment are merged in the buffer and then reconstructed, and so on, until the Nth encoded video segment and the incremental data corresponding to the Nth video segment are merged in the buffer and then reconstructed. Finally, the reconstructed semantic video can be obtained according to the reconstructed video segments.

[0052] The present disclosure provides a semantic video transmission method, apparatus, device, and storage medium. For other video segments except the first video segment, after encoding and processing at the first compression ratio, a suitable second compression ratio is selected according to the current network condition of the wireless channel, thereby realizing dynamic adjustment of the video bitrate, and further, incremental data corresponding to the video segment can be obtained, which is further beneficial to restoring a video segment with higher quality at the receiving end. Among them, taking the second video segment as an example, after encoding and processing at the first compression ratio, the second compression ratio is determined according to the current network condition of the wireless channel, and the second video segment is processed according to the second compression ratio to obtain corresponding incremental data, so that as much information of the second video segment as possible can be obtained, and a second video segment with higher quality can be restored when restoring the second video segment. Therefore, the solution of the present disclosure selects a suitable video compression ratio for transmission according to the real-time network fluctuation condition, realizes dynamic adjustment of the video transmission bitrate, can restore high-quality video content at the receiving end, and meets the user's video transmission requirements.

[0053] In some alternative embodiments, at least one third video segment is further included in the plurality of video segments;

[0054] After obtaining the incremental data corresponding to the second video segment, the method further includes:

[0055] Encoding and processing the third video segment at the first compression ratio to obtain a third encoded video segment;

[0056] Transmitting the third encoded video segment and the incremental data corresponding to the second video segment to the receiving end through the wireless channel;

[0057] Determining a third compression ratio according to the current network condition of the wireless channel, and processing the third video segment according to the third compression ratio to obtain incremental data corresponding to the third video segment and transmitting it to the receiving end, wherein the first compression ratio is greater than the third compression ratio;

[0058] The receiving end performs a video reconstruction operation according to the third encoded video segment and the incremental data corresponding to the third video segment to obtain a reconstructed semantic video segment.

[0059] Specifically, for the third video segment, after encoding and processing it at the first compression ratio to obtain the third encoded video segment, determine the third compression ratio according to the current network condition of the wireless channel, and then process the third video segment according to the third compression ratio to obtain the incremental data corresponding to the third video segment. In this way, determining the third compression ratio according to the current network condition of the wireless channel and processing the third video segment according to the third compression ratio to obtain the corresponding incremental data is to obtain as much information of the third video segment as possible, so as to facilitate the subsequent restoration to obtain a third video segment with relatively high quality.

[0060] In some alternative embodiments, encoding and processing both the first video segment and the second video segment at the first compression ratio includes:

[0061] Input the first video segment and the second video segment into the encoder module in sequence and perform encoding and processing at the maximum compression ratio to obtain the first encoding result and the second encoding result.

[0062] Specifically, when encoding and processing both the first video segment and the second video segment at the first compression ratio, input the first video segment and the second video segment into the encoder module in sequence and perform encoding and processing at the maximum compression ratio, so as to obtain the first encoding result and the second encoding result. Among them, regarding the maximum compression ratio, it should be noted that the maximum compression ratio refers to sorting the original video semantic information from largest to smallest in terms of importance, and discarding the information with low importance at the maximum ratio on the premise that the quality of the reconstructed video does not drop sharply or even cannot be reconstructed. Generally, this maximum compression ratio can be, for example, 70%.

[0063] In this way, when encoding and processing the first video segment and the second video segment, the maximum compression ratio is used for processing. On the one hand, it can ensure that the video is restored at the lowest quality, and on the other hand, using the maximum compression ratio for processing is beneficial to ensuring the data transmission efficiency.

[0064] In some alternative embodiments, determining the second compression ratio according to the current network condition of the wireless channel includes:

[0065] Obtain the current network throughput of the wireless channel;

[0066] Input the current network throughput into the bitrate adaptation model, and obtain the second compression ratio through the bitrate adaptation model.

[0067] Specifically, to obtain the current network condition of the wireless channel, it can be determined by obtaining the current network throughput of the wireless channel. After obtaining the current network throughput of the wireless channel, the current network throughput is input into the bitrate adaptation model, and the second compression ratio is determined through the bitrate adaptation model. The bitrate adaptation model can be based on the D3QN model (Dueling Double Deep Q - learning), with real - time network throughput, video semantic information importance, and video buffer - related information as model inputs, and the output is the video compression ratio selected for the next - moment transmission. Metrics such as the peak signal - to - noise ratio of the video frame, multi - scale structural similarity, and video picture quality are used as rewards to optimize the behavior of compression ratio selection.

[0068] Optionally, the algorithm steps corresponding to the bitrate adaptation model include:

[0069] (1) Initialize the current Q - network parameters θ and the target - network Q' parameters θ'.

[0070] (2) Assign the Q - network parameters to the Q' network, θ'←θ.

[0071] (3) Initialize the experience replay pool.

[0072] (3.1) Initialize the simulation player environment and obtain the current state s.

[0073] (3.2) Input the state s into the current Q - network, calculate the Q - values corresponding to each action, and use the ε - greedy algorithm to select the action a corresponding to the current state s.

[0074] (3.3) The player executes the compression - ratio selection of action a to obtain a new state s' and a reward r.

[0075] (3.4) Store {s, a, r, s'} in the experience replay pool D.

[0076] (3.5) Randomly sample m samples {s j , a, r, s' j} from D, where j = 1, 2, …, m.

[0077] (3.6) Calculate the target value of the current Q - network:

[0078] y j = r j + γQ(s′ j , argmax a Q(s′ j , a|θ′)|θ), where j = 1, 2, …, m.

[0079] (3.7) Use the mean - square error Calculate the error value and update the parameters by backpropagation;

[0080] (3.8) At each time interval T, update the network parameter θ'←θ.

[0081] In this way, according to the real-time network fluctuation situation, select the appropriate video compression ratio for transmission, realize the dynamic adjustment of the video transmission bit rate, and be able to recover high-quality video content at the receiving end, thus realizing video transmission that meets user requirements.

[0082] In some alternative embodiments, divide the semantic video to be transmitted to obtain multiple video segments, including:

[0083] Divide the semantic video to be transmitted into equal lengths to obtain multiple video segments.

[0084] Specifically, when dividing the semantic video to be transmitted, divide it into equal lengths, which is beneficial to obtaining video segments with uniform lengths, and thus beneficial to the processing of multiple video segments.

[0085] In some alternative embodiments, after obtaining the first encoded video segment and the second encoded video segment, the method further includes:

[0086] Store the second encoded video segment in a buffer.

[0087] After obtaining the first encoded video segment and the second encoded video segment, store the second encoded video segment in a buffer, which is beneficial to the subsequent processing of the second encoded video segment.

[0088] The present disclosure provides a semantic video transmission method applied to a receiving end. Refer to Figure 2 as shown, including:

[0089] Step S201: Receive the first encoded video segment and the second encoded video segment transmitted from the sending end through a wireless channel. The sending end obtains the semantic video to be transmitted, divides the semantic video to be transmitted to obtain at least one first video segment and at least one second video segment, and encodes the first video segment and the second video segment at a first compression ratio to obtain the first encoded video segment and the second encoded video segment.

[0090] Specifically, after the sending end determines the semantic video to be transmitted, divide the semantic video to be transmitted, so that multiple video segments can be obtained. When dividing the semantic video to be transmitted, it can be divided into equal lengths or unequal lengths, which is not limited herein, and those skilled in the art can divide it according to actual needs. For example, if the semantic video to be transmitted is V, then the semantic video to be transmitted V can be divided into multiple video segments {v1, v2, v3,..., vN}.

[0091] Step S202: Receive the incremental data corresponding to the second video segment sent from the sending end. The sending end determines the second compression ratio according to the current network condition of the wireless channel, and processes the second video segment according to the second compression ratio to obtain the incremental data corresponding to the second video segment, where the first compression ratio is greater than the second compression ratio.

[0092] Specifically, after the sending end obtains multiple video segments, select the first video segment and the second video segment from the multiple video segments. The first video segment can be considered as the first video segment, and the second video segment is the next video segment after the first segment. Taking the above example, v1 is used as the first video segment, and v2 is used as the second video segment. After obtaining the first video segment and the second video segment, both the first video segment and the second video segment are encoded and processed at the first compression ratio to obtain the first encoded processing result and the second encoded processing result. Then, the first encoded processing result and the second encoded processing result are transmitted to the receiving end through the wireless channel to obtain the first encoded video segment and the second encoded video segment.

[0093] After transmitting the encoded processing result to the receiving end through the wireless channel to obtain the first encoded video segment and the second encoded video segment, obtain the current network condition of the wireless channel. The sending end determines the second compression ratio according to the current network condition of the wireless channel, and then processes the second video segment according to the second compression ratio to obtain the incremental data corresponding to the second video segment. In this way, determining the second compression ratio according to the current network condition of the wireless channel and processing the second video segment according to the second compression ratio to obtain the corresponding incremental data is to obtain as much information of the second video segment as possible, so as to facilitate the subsequent restoration to obtain a second video segment with relatively high quality. It should be noted that the same processing is also performed on other video segments, which will not be elaborated here.

[0094] Step S203: Perform video reconstruction operations according to the first encoded video segment, the second encoded video segment, and the incremental data corresponding to the second video segment to obtain the reconstructed semantic video.

[0095] Specifically, after obtaining the first encoded video segment, the second encoded video segment, and the incremental data corresponding to the second video segment, the incremental data corresponding to other video segments is calculated in the same way for the other video segments. Since the calculation method is similar to that for determining the incremental data of the second video segment, it will not be repeated here. When video reconstruction is performed at the receiving end, each video segment is reconstructed and finally the reconstructed semantic video is output. For example, in the order of reception, the first encoded video segment is reconstructed, the second encoded video segment and the incremental data corresponding to the second video segment are merged in the buffer and then reconstructed, the third encoded video segment and the incremental data corresponding to the third video segment are merged in the buffer and then reconstructed, and so on until the Nth encoded video segment and the incremental data corresponding to the Nth video segment are merged in the buffer and then reconstructed. Finally, the reconstructed semantic video can be obtained from the reconstructed video segments.

[0096] The present disclosure provides a semantic video transmission method, apparatus, device, and storage medium. For video segments other than the first video segment, after encoding at the first compression ratio, a suitable second compression ratio is selected according to the current network condition of the wireless channel, thereby realizing dynamic adjustment of the video bit rate. Furthermore, the incremental data corresponding to the video segment can be obtained, which is further beneficial to restoring a higher-quality video segment at the receiving end. Taking the second video segment as an example, after encoding at the first compression ratio, the second compression ratio is determined according to the current network condition of the wireless channel, and the second video segment is processed according to the second compression ratio to obtain the corresponding incremental data. In this way, as much information of the second video segment as possible can be obtained, and a higher-quality second video segment can be restored when restoring the second video segment. Therefore, the solution of the present disclosure selects a suitable video compression ratio for transmission according to real-time network fluctuations, realizes dynamic adjustment of the video transmission bit rate, can restore high-quality video content at the receiving end, and meets the user's video transmission requirements.

[0097] In some alternative embodiments, performing a video reconstruction operation based on the first encoded video segment, the second encoded video segment, and the incremental data corresponding to the second video segment to obtain the reconstructed semantic video includes:

[0098] Input the first encoded video into the decoder module, and obtain the reconstructed first video segment through the decoder module;

[0099] Perform video segment reconstruction based on the second encoded video segment and the incremental data corresponding to the second video segment to obtain the reconstructed second video segment;

[0100] Obtain the reconstructed semantic video based on the reconstructed first video segment and the reconstructed second video segment.

[0101] Specifically, when reconstructing the semantic video, after obtaining the first encoded video segment, the first encoded video is input into the decoder module, and the reconstructed first video segment is obtained through the decoder module. For the second video segment, however, video segment reconstruction needs to be performed based on the second encoded video segment and the incremental data corresponding to the second video segment to obtain the reconstructed second video segment. Similarly, for other video segments, such as the third video segment, the third compression ratio needs to be determined according to the current network condition of the wireless channel, and the third video segment is processed according to the third compression ratio to obtain the incremental data corresponding to the third video segment. Video segment reconstruction is performed based on the third encoded video segment and the incremental data corresponding to the third video segment to obtain the reconstructed third video segment. Since other video segments are processed in the same way, they will not be elaborated here. After reconstructing each video segment, all the segments are integrated to obtain the reconstructed semantic video.

[0102] In this way, corresponding video segments are restored at the receiving end according to each video segment and the incremental data corresponding to each video segment, which is beneficial to restoring high-quality semantic video segments.

[0103] In some optional embodiments, performing video segment reconstruction based on the second encoded video segment and the incremental data corresponding to the second video segment to obtain the reconstructed second video segment includes:

[0104] Merging the second encoded video segment and the incremental data corresponding to the second video segment to obtain the merged second encoded video;

[0105] Inputting the merged second encoded video into the decoder module to obtain the reconstructed second video segment through the decoder module.

[0106] Specifically, after obtaining the second encoded video segment and the incremental data corresponding to the second video segment, the second encoded video segment and the incremental data corresponding to the second video segment are merged to obtain the merged second encoded video; the merged second encoded video is input into the decoder module to obtain the reconstructed second video segment through the decoder module. For other video segments, similar processing is also performed. For example, after obtaining the third encoded video segment and the incremental data corresponding to the third video segment, the third encoded video segment and the incremental data corresponding to the third video segment are merged to obtain the merged third encoded video; the merged third encoded video is input into the decoder module to obtain the reconstructed third video segment through the decoder module. Since other video segments are processed in the same way, they will not be elaborated here. After reconstructing each video segment, all the segments are integrated to obtain the reconstructed semantic video.

[0107] In this way, by restoring the corresponding video segments and the incremental data corresponding to each video segment at the receiving end, it is beneficial to restore a semantic video with relatively high quality.

[0108] To facilitate an overall understanding of the technical solution of the present disclosure, refer to Figure 3 , Figure 3 which is a schematic diagram of the system corresponding to the semantic video transmission method in an embodiment of the present disclosure. The system corresponding to this method includes a semantic encoding module and a semantic decoding module. The semantic encoding module includes modules such as implicit transformation, source-channel joint encoding, and variable-length encoding. The role of variable-length encoding is to discard some data according to the calculated discard threshold to form semantic vectors of different lengths. Among them, the lower the discard threshold, the more data is discarded, and the lower the video quality restored at the receiving end; the higher the discard threshold, the less data is discarded, and the higher the video quality restored at the receiving end. The semantic decoding module includes source-channel joint decoding and implicit inverse transformation modules. There is a wireless channel and a video buffer between the semantic encoding module and the semantic decoding module. The system also includes an entropy model. The role of the entropy model is to obtain the information entropy of each feature data according to the input semantic features, and sort them from largest to smallest according to the entropy value, so as to calculate the importance of different semantic information.

[0109] The semantic transmission method process corresponding to this system is as follows: (1) Divide the original video V into 10 video segments {v1, v2, v3,..., v 10} of equal length. The first two video segments v1 and v2 are sequentially passed into modules such as a semantic encoder, an information entropy model, and a variable-length encoder for processing at a maximum compression ratio of 70% and transmitted to the receiving end; (2) The receiving end reconstructs the video segment v1. At the same time, the third video segment v3 is operated on in the manner of (1) at the sending end, and the optimal video compression ratio under the current network throughput is obtained in real time using a bit rate adaptation module. According to the different importance of semantic information, the video segment v2 is processed at the semantic level to obtain the incremental information v'2 of v2 and sent to the receiving end; (3) The receiving end merges v2 and v'2 and then reconstructs and plays them. At the same time, the fourth video segment v4 and the incremental information v'3 of v3 are operated on in the manner of (2) at the sending end; (4) The remaining video segments are sequentially operated on in the manner of (3) until the video transmission is completed.

[0110] Specifically, refer to Figure 4 , Figure 4It is a schematic diagram of the processing process of the basic video data and the incremental video data in an embodiment of the present disclosure. For video segment 1, video reconstruction is directly performed using basic data 1. For video segment 2, in addition to performing video reconstruction using basic data 2, video reconstruction is also performed using incremental data 2 obtained by determining a second compression ratio according to the network fluctuation situation. The processing of other video segments is similar to that of video segment 2 and will not be elaborated here.

[0111] In this way, the solution of the present disclosure can select a video compression ratio suitable for transmission according to the real-time network fluctuation situation, realize the dynamic adjustment of the video transmission bit rate, be able to restore high-quality video content at the receiving end, and realize video transmission that meets user requirements.

[0112] The following introduces the device embodiments of the present application, which can be used to execute the semantic video transmission method in the above embodiments of the present application. For details not disclosed in the device embodiments of the present application, please refer to the embodiments of the semantic video transmission method above of the present application.

[0113] The present disclosure also provides a semantic video transmission device 500, as Figure 5 shown, including:

[0114] A partitioning module 501, configured to partition the semantic video to be transmitted to obtain a plurality of video segments, where the plurality of video segments include a first video segment and a second video segment;

[0115] An encoding module 502, configured to perform encoding processing on both the first video segment and the second video segment at a first compression ratio, and transmit the encoding processing result to the receiving end through a wireless channel to obtain a first encoded video segment and a second encoded video segment;

[0116] An incremental module 503, configured to determine a second compression ratio according to the current network condition of the wireless channel, and process the second video segment according to the second compression ratio to obtain incremental data corresponding to the second video segment and transmit it to the receiving end, where the first compression ratio is greater than the second compression ratio;

[0117] The receiving end performs video reconstruction operations according to the first encoded video segment, the second encoded video segment, and the incremental data corresponding to the second video segment to obtain the reconstructed semantic video.

[0118] In some alternative embodiments, the encoding module 502 performing encoding processing on both the first video segment and the second video segment at a first compression ratio includes:

[0119] Sequentially inputting the first video segment and the second video segment into an encoder module, and performing encoding processing at the maximum compression ratio to obtain a first encoding processing result and a second encoding processing result.

[0120] In some alternative embodiments, the increment module 503 determines a second compression ratio according to the current network condition of the wireless channel, including:

[0121] Obtain the current network throughput of the wireless channel;

[0122] Input the current network throughput into the bit rate adaptation model, and obtain the second compression ratio through the bit rate adaptation model.

[0123] In some alternative embodiments, the partitioning module 501 partitions the semantic video to be transmitted to obtain a plurality of video segments, including:

[0124] Partition the semantic video to be transmitted at equal lengths to obtain a plurality of video segments.

[0125] In some alternative embodiments, the increment module 503 is further configured to store the second encoded video segment in a buffer.

[0126] In some alternative embodiments, the encoding module 502 is configured to perform encoding processing on the third video segment at a first compression ratio to obtain a third encoded video segment; transmit the third encoded video segment and the incremental data corresponding to the second video segment to the receiving end through the wireless channel;

[0127] The increment module 503 is configured to determine a third compression ratio according to the current network condition of the wireless channel, and process the third video segment according to the third compression ratio to obtain incremental data corresponding to the third video segment and transmit it to the receiving end, where the first compression ratio is greater than the third compression ratio;

[0128] The receiving end performs video reconstruction operations according to the third encoded video segment and the incremental data corresponding to the third video segment to obtain the reconstructed semantic video segment.

[0129] The present disclosure also provides a semantic video transmission device 600, as Figure 6 shown, including:

[0130] A first receiving module 601, configured to receive the first encoded video segment and the second encoded video segment transmitted from the sending end through the wireless channel. The sending end obtains the semantic video to be transmitted, partitions the semantic video to be transmitted to obtain at least one first video segment and at least one second video segment, and performs encoding processing on both the first video segment and the second video segment at a first compression ratio to obtain the first encoded video segment and the second encoded video segment;

[0131] A second receiving module 602, configured to receive the incremental data corresponding to the second video segment sent from the sending end. The sending end determines a second compression ratio according to the current network condition of the wireless channel, and processes the second video segment according to the second compression ratio to obtain the incremental data corresponding to the second video segment, where the first compression ratio is greater than the second compression ratio;

[0132] A reconstruction module 603, configured to perform a video reconstruction operation according to the first encoded video segment, the second encoded video segment, and the incremental data corresponding to the second video segment to obtain a reconstructed semantic video.

[0133] In some alternative embodiments, the reconstruction module 603 performs a video reconstruction operation according to the first encoded video segment, the second encoded video segment, and the incremental data corresponding to the second video segment to obtain a reconstructed semantic video, including:

[0134] Input the first encoded video into a decoder module, and obtain a reconstructed first video segment through the decoder module;

[0135] Perform video segment reconstruction according to the second encoded video segment and the incremental data corresponding to the second video segment to obtain a reconstructed second video segment;

[0136] Obtain a reconstructed semantic video according to the reconstructed first video segment and the reconstructed second video segment.

[0137] In some alternative embodiments, the reconstruction module 603 performs video segment reconstruction according to the second encoded video segment and the incremental data corresponding to the second video segment to obtain a reconstructed second video segment, including:

[0138] Merge the second encoded video segment and the incremental data corresponding to the second video segment to obtain a merged second encoded video;

[0139] Input the merged second encoded video into a decoder module, and obtain a reconstructed second video segment through the decoder module.

[0140] In the technical solution of the present disclosure, the acquisition, storage, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0141] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0142] Figure 7FIG. shows a schematic block diagram of an exemplary electronic device 700 that can be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, personal digital processors, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0143] As Figure 7 shown, the electronic device 700 includes a computing unit 701 that can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the device 700 can also be stored. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0144] A plurality of components in the device 700 are connected to the I / O interface 705, including: an input unit 706, such as a keyboard, a mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a magnetic disk, an optical disk, etc.; and a communication unit 709, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 709 allows the device 700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0145] The computing unit 701 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 701 executes the various methods and processes described above, such as the semantic video transmission method. For example, in some embodiments, the semantic video transmission method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, one or more steps of the small program distribution described above can be executed. Alternatively, in other embodiments, the computing unit 701 can be configured to execute the semantic video transmission method by any other suitable means (e.g., by means of firmware).

[0146] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor, receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0147] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0148] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0149] For purposes of providing an interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic, speech, or tactile input).

[0150] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0151] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, a server of a distributed system, or a server incorporating a blockchain.

[0152] It should be understood that the various forms of processes shown above can be used, with steps reordered, added or deleted. For example, the steps described in the present disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution disclosed in the present disclosure can be achieved, and no limitations are imposed herein.

[0153] The above specific embodiments do not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub - combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present disclosure shall be included within the protection scope of the present disclosure.

Claims

1. A semantic video transmission method, comprising: Dividing the semantic video to be transmitted to obtain multiple video segments, where the multiple video segments include at least one first video segment and at least one second video segment; Encoding the first video segment and the second video segment at a first compression ratio, and transmitting the encoding result through a wireless channel to a receiving end to obtain a first encoded video segment and a second encoded video segment; Determining a second compression ratio according to the current network condition of the wireless channel, and processing the second video segment according to the second compression ratio to obtain incremental data corresponding to the second video segment and transmitting it to the receiving end, where the first compression ratio is greater than the second compression ratio; The receiving end performs a video reconstruction operation according to the first encoded video segment, the second encoded video segment, and the incremental data corresponding to the second video segment to obtain the reconstructed semantic video.

2. The method according to claim 1, wherein, The encoding the first video segment and the second video segment at a first compression ratio includes: Sequentially inputting the first video segment and the second video segment into an encoder module, and encoding them at the maximum compression ratio to obtain a first encoding result and a second encoding result.

3. The method according to claim 1 or 2, wherein The determining a second compression ratio according to the current network condition of the wireless channel includes: Obtaining the current network throughput of the wireless channel; Inputting the current network throughput into a bitrate adaptation model to obtain the second compression ratio through the bitrate adaptation model.

4. The method according to claim 1 or 2, wherein The dividing the semantic video to be transmitted to obtain multiple video segments includes: Dividing the semantic video to be transmitted at equal length to obtain multiple video segments.

5. The method according to claim 1 or 2, wherein After obtaining the first encoded video segment and the second encoded video segment, the method further includes: Storing the second encoded video segment in a buffer.

6. The method according to claim 1, wherein, The multiple video segments further include at least one third video segment; After obtaining the incremental data corresponding to the second video segment, the method further includes: Encoding the third video segment at the first compression ratio to obtain a third encoded video segment; Transmitting the third encoded video segment and the incremental data corresponding to the second video segment through a wireless channel to the receiving end; Determining a third compression ratio according to the current network condition of the wireless channel, and processing the third video segment according to the third compression ratio to obtain incremental data corresponding to the third video segment and transmitting it to the receiving end, where the first compression ratio is greater than the third compression ratio; The receiving end performs a video reconstruction operation according to the third encoded video segment and the incremental data corresponding to the third video segment to obtain the reconstructed semantic video segment.

7. A semantic video transmission method, comprising: Receive a first encoded video segment and a second encoded video segment transmitted from a sender via a wireless channel. The sender obtains a semantic video to be transmitted, divides the semantic video to be transmitted to obtain at least one first video segment and at least one second video segment, and encodes both the first video segment and the second video segment at a first compression ratio to obtain the first encoded video segment and the second encoded video segment; Receive incremental data corresponding to the second video segment sent from the sender. The sender determines a second compression ratio according to the current network condition of the wireless channel, and processes the second video segment according to the second compression ratio to obtain incremental data corresponding to the second video segment, where the first compression ratio is greater than the second compression ratio; Perform a video reconstruction operation according to the first encoded video segment, the second encoded video segment, and the incremental data corresponding to the second video segment to obtain a reconstructed semantic video.

8. The method according to claim 7, wherein The performing a video reconstruction operation according to the first encoded video segment, the second encoded video segment, and the incremental data corresponding to the second video segment to obtain a reconstructed semantic video includes: Input the first encoded video into a decoder module, and obtain a reconstructed first video segment through the decoder module; Perform video segment reconstruction according to the second encoded video segment and the incremental data corresponding to the second video segment to obtain a reconstructed second video segment; Obtain a reconstructed semantic video according to the reconstructed first video segment and the reconstructed second video segment.

9. The method according to claim 8, wherein, The performing video segment reconstruction according to the second encoded video segment and the incremental data corresponding to the second video segment to obtain a reconstructed second video segment includes: Merge the second encoded video segment and the incremental data corresponding to the second video segment to obtain a merged second encoded video; Input the merged second encoded video into a decoder module, and obtain a reconstructed second video segment through the decoder module.

10. A semantic video transmission device, comprising: A division module, configured to divide a semantic video to be transmitted to obtain multiple video segments, where the multiple video segments include at least one first video segment and at least one second video segment; An encoding module, configured to encode both the first video segment and the second video segment at a first compression ratio, and transmit the encoding result to a receiving end via a wireless channel to obtain a first encoded video segment and a second encoded video segment; An incremental module, configured to determine a second compression ratio according to the current network condition of the wireless channel, and process the second video segment according to the second compression ratio to obtain incremental data corresponding to the second video segment and transmit it to the receiving end, where the first compression ratio is greater than the second compression ratio; The receiving end performs a video reconstruction operation according to the first encoded video segment, the second encoded video segment, and the incremental data corresponding to the second video segment to obtain a reconstructed semantic video.

11. A semantic video transmission device, comprising: A first receiving module, configured to receive a first encoded video segment and a second encoded video segment transmitted from a sending end via a wireless channel, where the sending end obtains a semantic video to be transmitted, divides the semantic video to be transmitted to obtain at least one first video segment and at least one second video segment, and encodes both the first video segment and the second video segment at a first compression ratio to obtain the first encoded video segment and the second encoded video segment; A second receiving module, configured to receive incremental data corresponding to the second video segment sent from the sending end, where the sending end determines a second compression ratio according to the current network condition of the wireless channel, and processes the second video segment according to the second compression ratio to obtain incremental data corresponding to the second video segment, where the first compression ratio is greater than the second compression ratio; A reconstruction module, configured to perform a video reconstruction operation according to the first encoded video segment, the second encoded video segment, and the incremental data corresponding to the second video segment to obtain the reconstructed semantic video.

12. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1-8.

13. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-9.

14. A computer program product, comprising a computer program, where the computer program, when executed by a processor, implements the method according to any one of claims 1-9.