Video transmission adaptive optimization method, device, equipment and storage medium

By adaptively optimizing the ratio of common and individual features of the target video frame group, the problem of network environment changes in semantic communication video transmission is solved, a more efficient and stable video transmission effect is achieved, and the user experience is improved.

CN119892808BActive Publication Date: 2025-10-10BEIJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510030706.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-10-10
Estimated Expiration
2045-01-08

AI Technical Summary

Technical Problem

Existing video transmission methods based on semantic communication cannot adapt to real-time changes in the network environment, and cannot balance video transmission efficiency and post-transmission video effects, resulting in a decrease in transmission efficiency or post-transmission video effects, and cannot meet user needs.

Method used

By obtaining the video quality assessment result data of historical video frame groups, adaptively optimizing the ratio of common and individual features of the target video frame group, and performing semantic encoding and encapsulation based on the current ratio, the ratio of common and individual features is dynamically adjusted to adapt to changes in the network environment, thereby improving the stability and quality of video transmission.

Benefits of technology

Effectively reduce data redundancy, lower the amount of data to be transmitted, improve the stability and quality of video transmission, enhance video transmission efficiency and user experience, and adapt to complex and changing network environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119892808B_ABST
    Figure CN119892808B_ABST
Patent Text Reader

Abstract

The application provides a video transmission adaptive optimization method and device, equipment and a storage medium. The method comprises the following steps: if video quality evaluation result data of a historical video frame group is acquired, then the commonness and individuality feature ratio of a target video frame group is adaptively optimized; based on the current commonness and individuality feature ratio, semantic coding is performed on the target video frame group to obtain commonness feature data and individuality feature data of the target video frame group respectively, and the commonness feature data and the individuality feature data are encapsulated and then transmitted to a receiving end, so that the receiving end generates video quality evaluation result data. According to the change of network channel quality and video quality, the commonness and individuality feature ratio of video communication can be dynamically adjusted, so that the stability and quality of video transmission are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of video data processing technology, and in particular to a method, apparatus, device and storage medium for adaptive optimization of video transmission. Background Art

[0002] With the continuous development of the digital age, the proportion of video traffic in the network has gradually increased. Traditional communication methods can no longer meet the growing demand for video transmission. To address this problem, video transmission based on semantic communication has gradually become a new trend in recent years. Semantic communication focuses on the semantic understanding and transmission of video content to improve transmission efficiency and user experience.

[0003] In a video transmission system based on semantic communication, the video semantic features extracted by the system can be divided into common features and individual features. Common features refer to the shared feature information between each video frame in a group of video frames, and individual features refer to the unique feature information of each video frame in a group of video frames. The ratio of common features to individual features will affect the effect of video transmission. Existing video transmission methods based on semantic communication usually extract features from video data based on a fixed ratio of common features to individual features. However, this method cannot adapt to real-time changes in the network environment, and cannot balance the video transmission efficiency and the video effect after transmission, which leads to a decrease in transmission efficiency or the video effect after transmission, and cannot meet user needs. Summary of the Invention

[0004] In view of this, embodiments of the present application provide a method, apparatus, device, and storage medium for adaptive optimization of video transmission to eliminate or improve one or more defects in the prior art.

[0005] One aspect of the present application provides a method for adaptive optimization of video transmission, comprising:

[0006] If the video quality assessment result data of the historical video frame group is obtained, the ratio of the commonality and individuality features corresponding to the target video frame group is adaptively optimized; the historical video frame group is the video frame group corresponding to the video data that has been transmitted to the receiving end; the target video frame group is the video frame group corresponding to the video data that is currently to be transmitted;

[0007] Based on the current ratio of the common and individual features, the target video frame group is semantically encoded to obtain the common feature data and individual feature data of the target video frame group respectively, and the common feature data and individual feature data are encapsulated and transmitted to the receiving end, so that the receiving end obtains the corresponding video frame group based on the received common feature data and individual feature data, and the receiving end uses the video frame group as a new historical video frame group for video quality assessment, and then determines whether to generate video quality assessment result data of the new historical video frame group.

[0008] In some embodiments of the present application, the video quality assessment result data is used to represent the first quality feedback information or the second quality feedback information;

[0009] The first quality feedback information is used to indicate that the video quality evaluation values ​​corresponding to the plurality of consecutive video frames in the historical video frame group are all higher than a preset upper threshold;

[0010] The second quality feedback information is used to indicate that the video quality assessment values ​​corresponding to a plurality of consecutive video frames in the historical video frame group are all lower than a preset lower threshold; wherein the lower threshold is smaller than the upper threshold.

[0011] In some embodiments of the present application, if the video quality assessment result data of the historical video frame group is obtained, the ratio of commonality to individuality features corresponding to the target video frame group is adaptively optimized, including:

[0012] If video quality assessment result data of a first historical video frame group is received and the video quality assessment result data currently represents the first quality feedback information, reducing the commonality-to-individuality feature ratio corresponding to the current target video frame group;

[0013] The first historical video frame group is a historical video frame group located before the current target video frame group in the video data, and the first historical video frame group and the target video frame group are adjacent to or spaced apart by a preset number of video frame groups.

[0014] In some embodiments of the present application, if the video quality assessment result data of the historical video frame group is obtained, the ratio of commonality to individuality features corresponding to the target video frame group is adaptively optimized, including:

[0015] If video quality assessment result data of a second historical video frame group is received and the video quality assessment result data currently represents the second quality feedback information, increasing the commonality-to-individuality feature ratio corresponding to the current target video frame group;

[0016] The second historical video frame group is a historical video frame group located before the current target video frame group in the video data, and the second historical video frame group and the target video frame group are adjacent to or spaced apart by a preset number of video frame groups.

[0017] In some embodiments of the present application, before semantically encoding the target video frame group based on the current commonality-to-individuality feature ratio to obtain the commonality feature data and individuality feature data of the target video frame group, the method further includes:

[0018] If, within a preset time period after the common feature data and individual feature data corresponding to the third historical video frame group are encapsulated and transmitted to the receiving end, no video quality assessment result data for the third historical video frame group is received, then the common feature to individual feature ratio corresponding to a historical video frame group preceding the third historical video frame group in the video data is used as the common feature to individual feature ratio corresponding to the target video frame group to be transmitted currently;

[0019] The third historical video frame group is a historical video frame group located before the current target video frame group in the video data, and the third historical video frame group and the target video frame group are adjacent to or spaced apart by a preset number of video frame groups.

[0020] A second aspect of the present application provides a method for adaptive optimization of video transmission, comprising:

[0021] Based on the common feature data and individual feature data currently received from the sending end, obtaining a video frame group belonging to the video data corresponding to the common feature data and the individual feature data, and using the video frame group as a historical video frame group for video quality assessment to determine whether to generate video quality assessment result data for the historical video frame group;

[0022] If video quality assessment result data of the historical video frame group is generated, the video quality assessment result data of the historical video frame group is sent to the sending end, so that the sending end adaptively optimizes the commonality and individuality feature ratio corresponding to the current target video frame group according to the video quality assessment result data, and enables the sending end to semantically encode the target video frame group based on the current commonality and individuality feature ratio to obtain the commonality feature data and individuality feature data of the target video frame group respectively, so as to encapsulate the commonality feature data and individuality feature data and transmit them, and the target video frame group is the video frame group to be transmitted corresponding to the video data.

[0023] In some embodiments of the present application, the step of obtaining, based on the common feature data and individual feature data currently received from the transmitting end, a video frame group corresponding to the common feature data and the individual feature data and belonging to the video data includes:

[0024] receiving a data packet encapsulating common feature data and individual feature data from a transmitting end, and extracting the common feature data and individual feature data from the data packet;

[0025] It is determined whether the currently extracted common feature data and individual feature data belong to the same target video frame group in the video data. If so, semantic decoding is performed on the common feature data and individual feature data to obtain a corresponding video frame group.

[0026] In some embodiments of the present application, the performing of video quality assessment on the video frame group as a historical video frame group to determine whether to generate video quality assessment result data for the historical video frame group includes:

[0027] The video frame group obtained after semantic decoding is used as a historical video frame group, and video quality assessment is performed on the historical video frame group to obtain video quality assessment values ​​corresponding to each of the multiple consecutive video frames in the historical video frame group, and it is determined whether the video quality assessment values ​​corresponding to each of the multiple consecutive video frames in the historical video frame group are all higher than a preset upper threshold or are all lower than a preset lower threshold, wherein the lower threshold is less than the upper threshold;

[0028] If the video quality assessment values ​​corresponding to the plurality of consecutive video frames in the historical video frame group are all higher than a preset upper threshold, generating first quality feedback information and generating video quality assessment result data corresponding to the historical video frame group for representing the first quality feedback information;

[0029] If the video quality assessment values ​​corresponding to multiple consecutive video frames in the historical video frame group are all lower than the preset lower threshold, second quality feedback information is generated and video quality assessment result data corresponding to the historical video frame group is generated for representing the second quality feedback information.

[0030] A third aspect of the present application provides a first video transmission adaptive optimization device, which is provided in a transmitting end and includes:

[0031] a commonality-to-individuality feature ratio selection module configured to adaptively optimize the commonality-to-individuality feature ratio corresponding to a target video frame group upon obtaining video quality assessment result data of a historical video frame group; the historical video frame group being a video frame group corresponding to the video data that has been transmitted to the receiving end; and the target video frame group being a video frame group corresponding to the video data that is currently to be transmitted;

[0032] The semantic coding and channel transmission module is used to semantically encode the target video frame group based on the current common-to-individual feature ratio to obtain the common feature data and individual feature data of the target video frame group respectively, and to encapsulate the common feature data and individual feature data and transmit them to the receiving end, so that the receiving end obtains the corresponding video frame group based on the received common feature data and individual feature data, and enables the receiving end to use the video frame group as a new historical video frame group for video quality assessment, and then determine whether to generate video quality assessment result data of the new historical video frame group.

[0033] A fourth aspect of the present application provides a second video transmission adaptive optimization device, which is provided in a receiving end and includes:

[0034] A video quality assessment module is configured to obtain, based on the common feature data and individual feature data currently received from the transmitting end, a video frame group belonging to the video data corresponding to the common feature data and the individual feature data, and use the video frame group as a historical video frame group for video quality assessment to determine whether to generate video quality assessment result data for the historical video frame group;

[0035] An evaluation result feedback module is used to send the video quality evaluation result data of the historical video frame group to the sending end if the video quality evaluation result data of the historical video frame group is generated, so that the sending end adaptively optimizes the commonality and individuality feature ratio corresponding to the current target video frame group according to the video quality evaluation result data, and enables the sending end to semantically encode the target video frame group based on the current commonality and individuality feature ratio to obtain the commonality feature data and individuality feature data of the target video frame group respectively, so as to encapsulate the commonality feature data and individuality feature data and transmit them, and the target video frame group is the video frame group to be transmitted corresponding to the video data.

[0036] The fifth aspect of the present application provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the computer program, it implements the video transmission adaptive optimization method described in the first aspect, and / or implements the video transmission adaptive optimization method described in the second aspect.

[0037] The sixth aspect of the present application provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, it implements the video transmission adaptive optimization method described in the first aspect, and / or implements the video transmission adaptive optimization method described in the second aspect.

[0038] A seventh aspect of the present application provides a computer program product comprising a computer program which, when executed by a processor, implements the video transmission adaptive optimization method of the first aspect described above, and / or implements the video transmission adaptive optimization method of the second aspect described above.

[0039] The video transmission adaptive optimization method provided by the present application can effectively reduce data redundancy, reduce the amount of data to be transmitted, thereby supporting real-time video communication; can dynamically adjust the common and individual feature ratio according to the change of network channel quality and video quality, improve the stability and quality of video transmission; and can more flexibly adapt to complex and changeable network environment, effectively improve the efficiency and user experience of video transmission. Therefore, the present application has important application prospect and technical value, and brings new ideas and solutions for the development of video communication field.

[0040] Additional advantages, objects, and features of the application will be set forth in part by the description that follows, and will become apparent to those skilled in the art upon examination of the following detailed description and drawings in which

[0041] Those skilled in the art will understand that the objects and advantages of the present application are not limited to the above specifically described, and the above and other objects that can be achieved by the present application will be more clearly understood according to the following detailed description. BRIEF DESCRIPTION OF DRAWINGS

[0042] The drawings described herein are intended to provide a further understanding of the present application, constitute a part of the present application, and do not constitute a limitation of the present application. The components in the drawings are not drawn to scale, but are only for the purpose of illustrating the principles of the present application. In order to facilitate the illustration and description of some parts of the present application, the corresponding parts in the drawings may be enlarged, that is, they may become larger than other components in the exemplary device actually manufactured according to the present application. In the drawings:

[0043] Figure 1 Schematic diagram of the overall process of semantic communication.

[0044] Figure 2 Schematic diagram of the execution logic of video transmission based on semantic communication.

[0045] Figure 3 This is a schematic diagram of a first flow chart of a video transmission adaptive optimization method performed at a transmitting end in one embodiment of the present application.

[0046] Figure 4 2 is a schematic diagram of a second flow chart of a video transmission adaptive optimization method executed at a transmitting end in one embodiment of the present application.

[0047] Figure 5 1 is a schematic diagram of a first flow chart of a method for adaptive optimization of video transmission performed at a receiving end in one embodiment of the present application.

[0048] Figure 6 2 is a schematic diagram of a second flow chart of a method for adaptive optimization of video transmission performed at a receiving end in one embodiment of the present application.

[0049] Figure 7 2 is a structural diagram of a first video transmission adaptive optimization device in one embodiment of the present application.

[0050] Figure 8 2 is a structural diagram of a second video transmission adaptive optimization device in one embodiment of the present application.

[0051] Figure 9 This is a schematic diagram of the execution logic of the adaptive commonality-individuality feature ratio adjustment method based on video semantic communication in the application example of this application.

[0052] Figure 10 This is a schematic diagram of the execution flow of the receiving end quality assessment feedback module in the application example of this application.

[0053] Figure 11 This is a schematic diagram of the execution flow of the common and individual feature ratio module of the sending end in the application example of this application. DETAILED DESCRIPTION

[0054] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail in conjunction with the embodiments and drawings. Here, the illustrative embodiments of this application and their descriptions are used to explain this application, but are not intended to limit this application.

[0055] It should also be noted here that in order to avoid obscuring the present application due to unnecessary details, the accompanying drawings only show structures and / or processing steps that are closely related to the scheme according to the present application, while other details that are not closely related to the present application are omitted.

[0056] It should be emphasized that the term "include / comprises" when used herein refers to the existence of features, elements, steps or components, but does not exclude the existence or addition of one or more other features, elements, steps or components.

[0057] It should also be noted that, unless otherwise specified, the term "connection" herein may refer not only to a direct connection but also to an indirect connection involving an intermediate.

[0058] Hereinafter, embodiments of the present application will be described with reference to the accompanying drawings. In the accompanying drawings, the same reference numerals represent the same or similar components, or the same or similar steps.

[0059] In traditional communications, even a small bit loss or error in the bit stream of a message can compromise the integrity of the entire data packet, making it impossible for the receiving end to correctly parse and restore the original information. Therefore, traditional communication systems are designed and implemented with a focus on data transmission stability and reliability, employing various error correction and fault tolerance technologies to minimize potential errors during transmission and ensure accurate communication and reception of information.

[0060] Transmission based on semantic communication extracts the video information to be sent at the source, rather than transmitting the message bitstream. This allows for the transmission of more meaningful information within the same bandwidth, effectively compressing data redundancy, improving information transmission efficiency, alleviating network transmission pressure, and reducing processing latency for intelligent tasks. Compared to traditional networks, semantic communication can better meet the needs of situations with poor channel quality and high packet loss rates.

[0061] The process of semantic communication is as follows Figure 1 As shown in the figure, the transmitter first extracts features from the information to be transmitted and then performs semantic and channel coding on the extracted semantic information. The data after semantic and channel coding reaches the receiver at the physical layer in the form of a bit stream via the network channel. Upon receiving the bit stream, the receiver performs semantic and channel decoding on it to recover the information from the bit stream. The process of extracting semantic information on the transmitter and recovering it on the receiver relies primarily on a semantic knowledge base shared between the transmitter and receiver.

[0062] Since video traffic accounts for a large proportion of all mobile data traffic, the increase in video traffic places higher demands on the network environment. Therefore, semantic communication has broad development prospects when facing video transmission tasks.

[0063] Video transmission systems based on semantic communication such as Figure 2 As shown in the figure, it includes five steps: video data splitting, implicit conversion, source-channel joint coding, semantic feature extraction, and variable length coding. Video data splitting divides the video into groups of video frames (Group of Frames, GoF); implicit conversion aims to reduce the amount of data to be transmitted by downsampling, convolution and residual network processing of GOF, thereby supporting real-time video communication; source-channel joint coding is divided into two steps: source coding and channel coding, where the purpose of source coding is to remove redundant information within the source and improve effectiveness; channel coding requires adding check bits to the original bit sequence to realize error detection and correction functions, thereby increasing the reliability of bit sequence transmission over noisy channels; semantic feature extraction obtains the common features of GOF and personality traits In video semantic communication, data containing key video features is typically defined as "common features," which refer to the shared features across all frames in a group of video frames. Each group of frames (GOFs) has only one common feature. Data containing non-key features is defined as "individual features," which refer to the unique features of each frame in a group of video frames. Each GOF has n individual features, where n is the number of video frames in the GOF. The variable-length coding process generates an entropy model structure shared by both the sender and receiver.

[0064] The receiving end combines the received common features and individual features through source-channel joint decoding, and then converts the received data into implicit expression data through deconvolution and activation function; then the receiving end converts the implicit expression data into a video image observable to the human eye through inverse implicit conversion using deconvolution and residual blocks.

[0065] In the above process, the common characteristics and personality traits The ratio of will affect the transmission quality of the video. Here, the Common Individual Feature Ratio (CIFR) is defined as:

[0066]

[0067] Different CIFRs have different effects on the video quality of semantic communication. The setting of CIFR should not be too large or too small. It should be set to ensure effective transmission of M c While ensuring Mi There is a reasonable amount of data.

[0068] Therefore, when video semantic communication faces a complex and changing network environment, the fixed CIFR method cannot fully adapt to the needs of network channel transmission. This is because under different network channel conditions, and The degree of packet loss and interference varies, and a fixed CIFR may prevent the receiver from improving video quality in some cases. Therefore, CIFR should be adjusted to accommodate different network channel environments with different packet loss rates, improving video quality and the robustness of the video semantic communication system.

[0069] In one or more embodiments of the present application, the common personality trait ratio refers to the ratio of the dimension of the personality trait to the dimension of the common trait.

[0070] In order to improve the quality of video transmission based on semantic communication, thereby improving the stability and quality of video transmission, the embodiments of the present application respectively provide a video transmission adaptive optimization method performed at the sending end, a first video transmission adaptive optimization device for performing the video transmission adaptive optimization method, a video transmission adaptive optimization method performed at the receiving end, a second video transmission adaptive optimization device for performing the video transmission adaptive optimization method, an electronic device, a computer-readable storage medium and a computer program product, which can dynamically adjust the common and individual feature ratio of video semantic communication according to feedback information, thereby effectively improving the stability and quality of video transmission.

[0071] The details are described in detail through the following examples.

[0072] Based on this, the embodiment of the present application provides a video transmission adaptive optimization method that can be performed by a first video transmission adaptive optimization device set at a transmitting end, see Figure 3 The video transmission adaptive optimization method performed by the first video transmission adaptive optimization device provided at the transmitting end specifically includes the following contents:

[0073] Step 100: If the video quality assessment result data of the historical video frame group is obtained, the ratio of common and individual features corresponding to the target video frame group is adaptively optimized; the historical video frame group is the video frame group corresponding to the video data that has been transmitted to the receiving end; the target video frame group is the video frame group corresponding to the video data that is currently to be transmitted.

[0074] In step 100, if the first video transmission adaptive optimization device set at the sending end receives the video quality assessment result data of the historical video frame group sent by the receiving end, the ratio of common characteristics to individual characteristics corresponding to the target video frame group is reduced or increased according to the video quality assessment result data to achieve adaptive adjustment of the ratio of common characteristics to individual characteristics corresponding to the target video frame group.

[0075] In one or more embodiments of the present application, the historical video frame group and the target video frame group each refer to one of the video frame groups corresponding to the same video data, wherein the historical video frame group is the video frame group corresponding to the video data that has been transmitted to the receiving end; and the target video frame group is the video frame group corresponding to the video data that is currently to be transmitted. It is understood that the target video frame group, the historical video frame group, and the subsequent references to the first historical video frame group, the second historical video frame group, and the first historical video frame group all refer to video frame groups with the same data structure, and the distinction is only for convenience of expression.

[0076] Before step 100, the transmitter may split the video data to be transmitted into groups of frames (Group of Frames, GoF), and each of the groups of frames contains video frames sorted in time sequence; and then perform downsampling, convolution, and residual network processing on each group of frames through implicit conversion.

[0077] Step 200: Based on the current ratio of the common and individual features, the target video frame group is semantically encoded to obtain the common feature data and individual feature data of the target video frame group respectively, and the common feature data and individual feature data are encapsulated and transmitted to the receiving end, so that the receiving end obtains the corresponding video frame group based on the received common feature data and individual feature data, and the receiving end uses the video frame group as a new historical video frame group for video quality assessment, and then determines whether to generate video quality assessment result data of the new historical video frame group.

[0078] In step 200, the first video transmission adaptive optimization device provided at the transmitting end extracts semantic information from the target video frame group by performing source-channel joint coding (also known as source and channel joint coding) on ​​each implicitly converted video frame group to obtain semantic information corresponding to the target video frame group. Source-channel joint coding is divided into two steps: source coding and channel coding. The purpose of source coding is to remove redundant information within the source and improve efficiency. Channel coding requires adding check bits to the original bit sequence to implement error detection and correction functions, thereby increasing the reliability of bit sequence transmission over noisy channels. Based on the current commonality-to-individuality feature ratio, semantic feature extraction is performed on the semantic information corresponding to the target video frame group obtained by the source-channel joint coding to obtain commonality feature data and individuality feature data of the target video frame group. The commonality feature data and individuality feature data are then encapsulated (channel coded) to obtain a data packet corresponding to the target video frame group, and the data packet is transmitted to the receiving end via a network channel.

[0079] After receiving the data packet, the receiving end will first extract the common feature data and individual feature data from the data packet, and then based on the consistency between the corresponding identifiers of the currently received common feature data and individual feature data, it can determine whether the currently extracted common feature data and individual feature data belong to the same target video frame group in the video data. If so, the common feature data and individual feature data are semantically decoded (that is, the inverse process of the semantic encoding performed by the sending end) to obtain the corresponding video frame group. And the video frame group is used as a historical video frame group for video quality assessment to determine whether to generate video quality assessment result data for the historical video frame group; and then if the receiving end generates video quality assessment result data for the historical video frame group, the video quality assessment result data for the historical video frame group is sent to the sending end, so that the sending end adaptively optimizes the commonality and individuality feature ratio corresponding to the current target video frame group according to the video quality assessment result data, and enables the sending end to semantically encode the target video frame group based on the current commonality and individuality feature ratio to obtain the commonality feature data and individuality feature data of the target video frame group respectively, so as to encapsulate the commonality feature data and individuality feature data and transmit them, and the target video frame group is the video frame group to be transmitted corresponding to the video data.

[0080] In one or more embodiments of the present application, the sending end and the receiving end refer to the two ends of video transmission, and the sending end and the receiving end can both be electronic devices, which can be servers or client devices. It can be understood that the same device can also be a sending end and a receiving end at the same time. For example, a client device can semantically encode and transmit video data A based on the video transmission adaptive optimization method executed by the first video transmission adaptive optimization device, and can also receive and semantically decode video data B based on the video transmission adaptive optimization method executed by the second video transmission adaptive optimization device mentioned in the following embodiments.

[0081] From the above description, it can be seen that the video transmission adaptive optimization method provided by the embodiment of the present application, which is performed by the first video transmission adaptive optimization device set at the transmitting end, can effectively reduce data redundancy and reduce the amount of data to be transmitted, thereby supporting real-time video communication; it can dynamically adjust the ratio of common and individual characteristics according to changes in network channel quality and video quality, thereby improving the stability and quality of video transmission; and it can more flexibly adapt to complex and changing network environments, effectively improving the efficiency of video transmission and user experience. Therefore, this application has important application prospects and technical value, and brings new ideas and solutions to the development of the field of video communication.

[0082] In order to further avoid transmission loss of invalid information, simplify the adaptive optimization process at the receiving end, and further improve video transmission efficiency, in a video transmission adaptive optimization method performed by a first video transmission adaptive optimization device provided at a transmitting end, provided in an embodiment of the present application, the video quality evaluation result data is used to represent the first quality feedback information or the second quality feedback information;

[0083] The first quality feedback information is used to indicate that the video quality evaluation values ​​corresponding to the plurality of consecutive video frames in the historical video frame group are all higher than a preset upper threshold;

[0084] The second quality feedback information is used to indicate that the video quality assessment values ​​corresponding to a plurality of consecutive video frames in the historical video frame group are all lower than a preset lower threshold; wherein the lower threshold is smaller than the upper threshold.

[0085] That is to say, the video quality assessment result data is only used to indicate that the video quality assessment values ​​corresponding to multiple consecutive video frames in the historical video frame group are all higher than the preset upper threshold or the video quality assessment values ​​corresponding to multiple consecutive video frames in the historical video frame group are all lower than the preset lower threshold, but does not indicate the information that the video quality assessment value is between the lower threshold and the upper threshold. This setting is because the video quality assessment value between the lower threshold and the upper threshold indicates that the video quality of the corresponding historical video frame group is in a state that can balance the transmission efficiency and video time quality. Therefore, there is no need to optimize the commonality and individuality feature ratio of the target video frame group after the historical video frame group in this state, and thus there is no need to generate video quality assessment result data containing information corresponding to this state, and there is no need to send data to the sending end when the commonality and individuality feature ratio is not optimized. This can avoid the transmission loss of invalid information, simplify the adaptive optimization process of the receiving end, and further improve the video transmission efficiency.

[0086] In order to further improve the effectiveness, flexibility and intelligence of the ratio of commonality to individuality features corresponding to the adaptive optimization target video frame group, in the video transmission adaptive optimization method provided by the first video transmission adaptive optimization device provided at the transmitting end in the embodiment of the present application, see Figure 4 Step 100 of the video transmission adaptive optimization method executed by the first video transmission adaptive optimization device provided at the transmitting end specifically includes the following contents:

[0087] Step 110: if video quality assessment result data of the first historical video frame group is received and the video quality assessment result data currently represents the first quality feedback information, then reducing the commonality-to-individuality feature ratio corresponding to the current target video frame group;

[0088] The first historical video frame group is a historical video frame group located before the current target video frame group in the video data, and the first historical video frame group and the target video frame group are adjacent to or spaced apart by a preset number of video frame groups.

[0089] In order to further improve the effectiveness, flexibility and intelligence of the ratio of commonality to individuality features corresponding to the adaptive optimization target video frame group, in a video transmission adaptive optimization method provided in an embodiment of the present application, see Figure 4 Step 100 in the video transmission adaptive optimization method further specifically includes the following contents:

[0090] Step 120: If video quality assessment result data of a second historical video frame group is received and the video quality assessment result data currently represents the second quality feedback information, then increase the commonality-to-individuality feature ratio corresponding to the current target video frame group;

[0091] The second historical video frame group is a historical video frame group located before the current target video frame group in the video data, and the second historical video frame group and the target video frame group are adjacent to or spaced apart by a preset number of video frame groups.

[0092] In order to further avoid the transmission loss of invalid information, simplify the adaptive optimization process at the receiving end, and further improve the video transmission efficiency, in the video transmission adaptive optimization method performed by the first video transmission adaptive optimization device provided at the transmitting end in the embodiment of the present application, see Figure 4 The video transmission adaptive optimization method executed by the first video transmission adaptive optimization device provided at the transmitting end further specifically includes the following contents before step 200:

[0093] Step 010: If, within a preset time period after the common feature data and individual feature data corresponding to the third historical video frame group are encapsulated and transmitted to the receiving end, no video quality assessment result data for the third historical video frame group is received, then the common feature to individual feature ratio corresponding to a historical video frame group preceding the third historical video frame group in the video data is used as the common feature to individual feature ratio corresponding to the target video frame group to be transmitted;

[0094] The third historical video frame group is a historical video frame group located before the current target video frame group in the video data, and the third historical video frame group and the target video frame group are adjacent to or spaced apart by a preset number of video frame groups.

[0095] It can be understood that in one or more embodiments of the present application, the first historical video frame group, the second historical video frame group and the third historical video frame group can adopt the same or different historical video frame groups, which can be specifically set according to the actual application scenario.

[0096] For example, in executing steps 110 and 120, in order to avoid the receiving end's evaluation of video quality affecting the sending end's efficiency in video transmission, the first historical video frame group and the second historical video frame group can be selected as historical video frame groups that are separated from the target video frame group by a preset number of video frame groups. For example, the historical video frame group that is three video frame groups apart from the current target video frame group is selected as the first historical video frame group and the second historical video frame, rather than the adjacent video frame group. In this way, the transmission time of three video frame groups apart can be used to enable the receiving end to evaluate the video quality and send the video quality evaluation result data to the sending end, thereby effectively avoiding the receiving end's evaluation of video quality affecting the sending end's efficiency in video transmission.

[0097] For another example, in executing step 010, in order to avoid the transmitter from waiting for a long time for the video quality assessment result data and affecting the transmission efficiency of the target video frame group, a historical video frame group three video frame groups before the current target video frame group can also be selected as the third historical video frame group.

[0098] At the same time, since the video quality assessment result data only indicates information that the video quality is too good or too bad, in order to further avoid waiting time, if the video quality assessment result data for the third historical video frame group is not received within a preset time period after the common feature data and individual feature data corresponding to the third historical video frame group are encapsulated and transmitted to the receiving end, the commonality and individuality feature ratio corresponding to a historical video frame group located before the third historical video frame group in the video data will be used as the commonality and individuality feature ratio corresponding to the target video frame group to be transmitted.

[0099] Among them, the preset time period can be set according to the actual situation. For example, based on the common feature data and individual feature data extracted from the data packet received by the receiving end, it is judged whether the common feature data and individual feature data currently extracted belong to the same target video frame group in the video data. If so, the common feature data and individual feature data are semantically decoded to obtain the corresponding video frame group, and the video frame group is used as a historical video frame group for video quality assessment to determine whether video quality assessment result data of the historical video frame group is generated; and then if the receiving end generates video quality assessment result data of the historical video frame group, the total time required to send the video quality assessment result data to the sending end is set.

[0100] Based on this, the embodiment of the present application provides a video transmission adaptive optimization method that can be implemented by a second video transmission adaptive optimization device provided at the receiving end, see Figure 5 The video transmission adaptive optimization method provided by the second video transmission adaptive optimization device at the receiving end specifically includes the following contents:

[0101] Step 300: Based on the common feature data and individual feature data currently received from the sending end, obtain the video frame group belonging to the video data corresponding to the common feature data and the individual feature data, and use the video frame group as a historical video frame group to perform video quality assessment to determine whether to generate video quality assessment result data for the historical video frame group.

[0102] Step 400: If the video quality assessment result data of the historical video frame group is generated, the video quality assessment result data of the historical video frame group is sent to the sending end, so that the sending end adaptively optimizes the common and individual feature ratios corresponding to the current target video frame group according to the video quality assessment result data, and enables the sending end to semantically encode the target video frame group based on the current common and individual feature ratios to obtain the common feature data and individual feature data of the target video frame group respectively, so as to encapsulate the common feature data and individual feature data and transmit them. The target video frame group is the video frame group to be transmitted corresponding to the video data.

[0103] The technical implementation details involved in the embodiment of the video transmission adaptive optimization method performed by the second video transmission adaptive optimization device set at the receiving end provided in the present application can be specifically implemented by adopting the technical implementation details recorded in the video transmission adaptive optimization method performed by the first video transmission adaptive optimization device set at the sending end in the above embodiment. Its functions will not be repeated here, and you can refer to the detailed description of the embodiment of the video transmission adaptive optimization method performed by the first video transmission adaptive optimization device set at the sending end.

[0104] As can be seen from the above description, the video transmission adaptive optimization method provided by the embodiment of the present application, which is performed by the second video transmission adaptive optimization device set at the receiving end, can effectively reduce data redundancy and reduce the amount of data to be transmitted, thereby supporting real-time video communication; it can dynamically adjust the ratio of common and individual characteristics to improve the stability and quality of video transmission; and it can more flexibly adapt to complex and changing network environments, effectively improving the efficiency of video transmission and user experience. Therefore, this application has important application prospects and technical value, and brings new ideas and solutions to the development of the field of video communication.

[0105] In order to further improve the effectiveness and reliability of obtaining the video frame group corresponding to the common feature data and the individual feature data belonging to the video data, in the video transmission adaptive optimization method provided by the second video transmission adaptive optimization device provided at the receiving end in the embodiment of the present application, see Figure 6 Step 300 of the video transmission adaptive optimization method by the second video transmission adaptive optimization device provided at the receiving end specifically includes the following contents:

[0106] Step 310: Receive a data packet encapsulating common feature data and individual feature data from a transmitting end, and extract the common feature data and individual feature data from the data packet.

[0107] Step 320: Determine whether the currently extracted common feature data and individual feature data belong to the same target video frame group in the video data. If so, perform semantic decoding on the common feature data and individual feature data to obtain a corresponding video frame group.

[0108] In order to further improve the effectiveness and reliability of using the video frame group as a historical video frame group for video quality evaluation, in the video transmission adaptive optimization method provided by the second video transmission adaptive optimization device provided at the receiving end in the embodiment of the present application, see Figure 6 Step 300 of the video transmission adaptive optimization method by the second video transmission adaptive optimization device provided at the receiving end further specifically includes the following contents:

[0109] Step 330: The video frame group obtained after semantic decoding is used as a historical video frame group, and a video quality assessment is performed on the historical video frame group to obtain the video quality assessment values ​​corresponding to the multiple consecutive video frames in the historical video frame group, and determine whether the video quality assessment values ​​corresponding to the multiple consecutive video frames in the historical video frame group are all higher than a preset upper limit threshold or whether they are all lower than a preset lower limit threshold, wherein the lower limit threshold is less than the upper limit threshold.

[0110] Step 340: If the video quality assessment values ​​corresponding to multiple consecutive video frames in the historical video frame group are all higher than the preset upper limit threshold, generate first quality feedback information and generate video quality assessment result data corresponding to the historical video frame group for representing the first quality feedback information.

[0111] Step 350: If the video quality assessment values ​​corresponding to multiple consecutive video frames in the historical video frame group are all lower than the preset lower threshold, generate second quality feedback information and generate video quality assessment result data corresponding to the historical video frame group for representing the second quality feedback information.

[0112] From the software level, this application also provides a method for executing Figure 3 or Figure 4 The first video transmission adaptive optimization device in all or part of the video transmission adaptive optimization method shown is set in the transmitting end, see Figure 7 The first video transmission adaptive optimization device specifically includes the following contents:

[0113] The commonality-to-individuality feature ratio selection module 10 is configured to adaptively optimize the commonality-to-individuality feature ratio corresponding to a target video frame group upon obtaining video quality assessment result data of a historical video frame group; the historical video frame group is a video frame group corresponding to the video data that has been transmitted to the receiving end; and the target video frame group is a video frame group corresponding to the video data that is currently to be transmitted.

[0114] The semantic coding and channel transmission module 20 is used to semantically encode the target video frame group based on the current common and individual feature ratio to obtain the common feature data and individual feature data of the target video frame group respectively, and encapsulate the common feature data and individual feature data and transmit them to the receiving end, so that the receiving end obtains the corresponding video frame group based on the received common feature data and individual feature data, and enables the receiving end to use the video frame group as a new historical video frame group for video quality assessment, and then determine whether to generate video quality assessment result data of the new historical video frame group.

[0115] The embodiment of the first video transmission adaptive optimization device provided in this application can be specifically used to perform the following Figure 3 or Figure 4 The processing flow of the embodiment of the video transmission adaptive optimization method in the above embodiment is shown in FIG. Figure 3 or Figure 4 Detailed description of the embodiment of the above-mentioned video transmission adaptive optimization method is shown.

[0116] The first video transmission adaptive optimization device performs the following Figure 3 or Figure 4 The portion of the video transmission adaptive optimization shown can be completed in the client device. The specific selection can be based on the processing capability of the client device and the limitations of the user's usage scenario. This application does not limit this. If all operations are completed in the client device, the client device may also include a processor for Figure 3 or Figure 4 The specific process of adaptive optimization of video transmission is shown.

[0117] The client device may include a communication module (i.e., a communication unit) that can establish a communication connection with a remote server to implement data transmission with the server. The server may include a server on the task scheduling center side, and in other implementation scenarios, may also include a server on an intermediate platform, such as a server on a third-party server platform that has a communication link with the task scheduling center server. The server may include a single computer device, a server cluster consisting of multiple servers, or a server structure of a distributed device.

[0118] The server and the client device may communicate using any suitable network protocol, including network protocols that have not yet been developed as of the filing date of this application. Examples of such network protocols include TCP / IP, UDP / IP, HTTP, and HTTPS. Furthermore, examples of such network protocols include RPC (Remote Procedure Call Protocol) and REST (Representational State Transfer) protocols, which are used on top of the aforementioned protocols.

[0119] As can be seen from the above description, the first video transmission adaptive optimization device provided in the embodiment of this application can effectively reduce data redundancy and the amount of data to be transmitted, thereby supporting real-time video communication; it can dynamically adjust the ratio of common and individual characteristics to improve the stability and quality of video transmission; and it can more flexibly adapt to complex and changing network environments, effectively improving the efficiency of video transmission and user experience. Therefore, this application has important application prospects and technical value, and brings new ideas and solutions to the development of the video communication field.

[0120] From the software level, this application also provides a method for executing Figure 5 or Figure 6 The second video transmission adaptive optimization device in all or part of the video transmission adaptive optimization method shown in the figure, the first video transmission adaptive optimization device is set in the receiving end, see Figure 8 The second video transmission adaptive optimization device specifically includes the following contents:

[0121] The video quality assessment module 30 is configured to obtain, based on the common feature data and individual feature data currently received from the transmitting end, a video frame group corresponding to the common feature data and the individual feature data belonging to the video data, and use the video frame group as a historical video frame group for video quality assessment to determine whether to generate video quality assessment result data for the historical video frame group;

[0122] The evaluation result feedback module 40 is used to send the video quality evaluation result data of the historical video frame group to the sending end if the video quality evaluation result data of the historical video frame group is generated, so that the sending end adaptively optimizes the commonality and individuality feature ratio corresponding to the current target video frame group according to the video quality evaluation result data, and enables the sending end to semantically encode the target video frame group based on the current commonality and individuality feature ratio to obtain the commonality feature data and individuality feature data of the target video frame group respectively, so as to encapsulate the commonality feature data and individuality feature data and transmit them. The target video frame group is the video frame group to be transmitted corresponding to the video data.

[0123] The embodiment of the second video transmission adaptive optimization device provided in this application can be specifically used to perform the following Figure 5 or Figure 6 The processing flow of the embodiment of the video transmission adaptive optimization method in the above embodiment is shown in FIG. Figure 5 or Figure 6 Detailed description of the embodiment of the above-mentioned video transmission adaptive optimization method is shown.

[0124] As can be seen from the above description, the second video transmission adaptive optimization device provided in the embodiment of this application can effectively reduce data redundancy and the amount of data to be transmitted, thereby supporting real-time video communication; it can dynamically adjust the ratio of common and individual characteristics to improve the stability and quality of video transmission; and it can more flexibly adapt to complex and changing network environments, effectively improving the efficiency of video transmission and user experience. Therefore, this application has important application prospects and technical value, and brings new ideas and solutions to the development of the video communication field.

[0125] To further illustrate the above embodiments, the present application also provides an adaptive common individual feature ratio adjustment method based on video semantic communication, which is interactively executed by the sending end and the receiving end. In actual video transmission scenarios, if video semantic communication adopts an unchanging CIFR in the face of a complex and changeable network channel environment, it is not conducive to maintaining the high quality of the video. Therefore, the CIFR extracted by video semantics can be adjusted to adapt to the changeable network channel environment. The application example of this application designs and implements an adaptive common individual feature ratio adjustment method based on video semantic communication based on the above ideas. As video traffic increases, traditional communications cannot meet the demand, and video transmission based on semantic communication is gradually emerging. Semantic communication focuses on the semantic understanding and transmission of video content to improve transmission efficiency and user experience. The application example of this application proposes an adaptive common individual feature ratio adjustment method, which is dynamically adjusted according to the network channel quality and video quality information to improve the stability and quality of video transmission. See. Figure 9 , the specific steps are as follows:

[0126] Step 1. The sender uses the current CIFR to semantically encode the video to be transmitted and generate common features of the video. and personality traits

[0127] Step 2. Sending end After encapsulation, it is transmitted to the receiving end;

[0128] Step 3. The receiving end analyzes the common features and personality traits Perform quality assessment. If the video quality changes, the receiver will generate feedback information with the adjusted common and individual features and send it to the transmitter. At the same time, the receiver performs semantic decoding to restore the semantic features into video.

[0129] Step 4. The sender adjusts the CIFR accordingly based on the received feedback information;

[0130] Step 5. The above process continues to iterate periodically until the video transmission is completed.

[0131] The specific implementation of this application mainly relies on two modules: a receiving-end video quality assessment feedback module and a sending-end commonality and individuality feature ratio selection module. Detailed descriptions are given below.

[0132] 1. Receiver video quality assessment feedback module

[0133] The main function of the receiving end video quality assessment feedback module is to evaluate the network channel parameters and the quality of the transmitted video, and then generate different feedback information to send to the common individual feature ratio selection module.

[0134] This module evaluates the quality of the network channel parameters and the transmitted video. If the evaluation value is continuously lower than the set threshold V l , then generate quality feedback information M l If the evaluation value is continuously higher than the set threshold V h , then generate quality feedback information M h and feed it back to the sender to reduce CIFR and save network bandwidth. l With V h If the quality feedback information is between , no quality feedback information is generated and the current CIFR is maintained. The specific algorithm of the receiving end quality assessment feedback module is shown in Table 1. The process is as follows Figure 10 shown.

[0135] Table 1 Specific algorithm description of the receiving end video quality assessment feedback module

[0136]

[0137] 2. Common and individual characteristics ratio selection module of the sending end

[0138] The transmitter first parses the received quality feedback and then passes the parsed information to the Common-Individual Feature Ratio Selection module. This module actively adjusts the semantically encoded CIFR based on the different feedback signals sent by the receiver, selecting the CIFR that best suits the current network channel environment to maintain high video quality.

[0139] When the parsed feedback information is M lWhen the feedback information is M, it indicates that the quality of the current video semantic communication is poor, so it is necessary to increase CIFR to improve the video quality; h When , the quality of the current video semantic communication is high, and the packet loss rate of the network channel has little impact on the video quality, which will not affect the user's intuitive experience. In this case, the sender's commonality and individuality feature ratio selection module can reduce CIFR to reduce network bandwidth consumption and improve transmission efficiency and video quality.

[0140] Subsequently, the video is semantically encoded based on the CIFR provided by the common and individual feature ratio selection module at the sending end to generate the common and individual features of the video. After encapsulating the header information, the video data is transmitted. The specific algorithm of the common and individual feature ratio selection module is shown in Table 2. The process is as follows: Figure 11 shown.

[0141] Table 2 Specific algorithm description of the common individual feature ratio selection module at the sending end

[0142]

[0143] The adaptive commonality-individuality feature ratio adjustment method proposed in the application example of this application brings significant advantages to the video transmission system based on semantic communication. First, by extracting the common characteristics and individual characteristics of the video, it is possible to more effectively reduce data redundancy and reduce the amount of data to be transmitted, thereby supporting real-time video communication. Secondly, by adopting an adaptive commonality-individuality feature ratio adjustment strategy, the commonality-individuality feature ratio can be dynamically adjusted according to the real-time network channel quality and video quality information, thereby improving the stability and quality of video transmission. Compared with the traditional fixed feature ratio method, this application can more flexibly adapt to complex and changeable network environments, effectively improving the efficiency of video transmission and user experience. Therefore, this application has important application prospects and technical value, and brings new ideas and solutions to the development of the field of video communication.

[0144] The embodiment of the present application also provides an electronic device, which may include a processor, a memory, a receiver and a transmitter, wherein the processor is configured to execute the following Figure 3 or Figure 4 The video transmission adaptive optimization method shown and / or Figure 5 or Figure 6 In the video transmission adaptive optimization method shown, the processor and the memory may be connected via a bus or other means, with bus connection being taken as an example. The receiver may be connected to the processor and the memory via a wired or wireless means.

[0145] The processor may be a central processing unit (CPU). The processor may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, or a combination of the above chips.

[0146] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs, non-transitory computer executable programs and modules, such as the embodiment of the present application. Figure 3 or Figure 4 The video transmission adaptive optimization method shown and / or Figure 5 or Figure 6 The program instructions / modules corresponding to the video transmission adaptive optimization method shown. The processor executes various functional applications and data processing of the processor by running the non-transient software programs, instructions and modules stored in the memory, that is, realizing the above method embodiment as shown in FIG. Figure 3 or Figure 4 The video transmission adaptive optimization method shown and / or Figure 5 or Figure 6 The video transmission adaptive optimization method shown.

[0147] The memory may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created by the processor, etc. In addition, the memory may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory may optionally include a memory remotely located relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0148] The one or more modules are stored in the memory and, when executed by the processor, perform the following steps in the embodiment: Figure 3 or Figure 4 The video transmission adaptive optimization method shown and / or Figure 5 or Figure 6 The video transmission adaptive optimization method shown.

[0149] In some embodiments of the present application, the user equipment may include a processor, a memory and a transceiver unit, and the transceiver unit may include a receiver and a transmitter. The processor, memory, receiver and transmitter may be connected through a bus system. The memory is used to store computer instructions, and the processor is used to execute the computer instructions stored in the memory to control the transceiver unit to send and receive signals.

[0150] As an implementation method, the functions of the receiver and transmitter in this application can be considered to be implemented through a transceiver circuit or a dedicated transceiver chip, and the processor can be considered to be implemented through a dedicated processing chip, a processing circuit or a general-purpose chip.

[0151] As another implementation method, it is possible to use a general-purpose computer to implement the server provided in the embodiments of the present application. That is, the program code for implementing the functions of the processor, receiver, and transmitter is stored in a memory, and the general-purpose processor implements the functions of the processor, receiver, and transmitter by executing the code in the memory.

[0152] The embodiment of the present application also provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the computer program can realize the following Figure 3 or Figure 4 The video transmission adaptive optimization method shown and / or Figure 5 or Figure 6 The steps of the video transmission adaptive optimization method shown in the figure. The computer-readable storage medium can be a tangible storage medium, such as a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a floppy disk, a hard disk, a removable storage disk, a CD-ROM, or any other form of storage medium known in the art.

[0153] The present application also provides a computer program product, including a computer program, which, when executed by a processor, implements the following Figure 3 or Figure 4 The video transmission adaptive optimization method shown and / or Figure 5 or Figure 6 The video transmission adaptive optimization method shown.

[0154] It should be understood by those skilled in the art that the various exemplary components, systems and methods described in conjunction with the embodiments disclosed herein can be implemented in hardware, software or a combination of the two. Whether it is specifically performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application. When implemented in hardware, it can be, for example, an electronic circuit, an application-specific integrated circuit (ASIC), appropriate firmware, a plug-in, a function card, etc. When implemented in software, the elements of this application are programs or code segments that are used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted on a transmission medium or a communication link via a data signal carried in a carrier.

[0155] It should be understood that the present application is not limited to the specific configurations and processes described above and illustrated in the figures. For the sake of brevity, a detailed description of known methods is omitted here. In the above embodiments, several specific steps are described and illustrated as examples. However, the method process of the present application is not limited to the specific steps described and illustrated. Those skilled in the art can make various changes, modifications, and additions, or change the order of the steps after understanding the spirit of the present application.

[0156] In this application, features described and / or illustrated for one embodiment may be used in the same or similar manner in one or more other embodiments, and / or combined with or replace features of other embodiments.

[0157] The above description is merely a preferred embodiment of the present application and is not intended to limit the present application. Those skilled in the art will appreciate that various modifications and variations of the present embodiment are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present application shall be included within the scope of protection of the present application.

Claims

1. A video transmission adaptive optimization method, characterized in that: include: If the video quality assessment result data of the historical video frame group is obtained, the ratio of the commonality and individuality features corresponding to the target video frame group is adaptively optimized; The historical video frame group is a video frame group corresponding to the video data that has been transmitted to the receiving end; the target video frame group is a video frame group corresponding to the video data that is currently to be transmitted; Based on the current ratio of the common and individual features, the target video frame group is semantically encoded to obtain the common feature data and individual feature data of the target video frame group respectively, and the common feature data and individual feature data are encapsulated and transmitted to the receiving end, so that the receiving end obtains the corresponding video frame group based on the received common feature data and individual feature data, and the receiving end uses the video frame group as a new historical video frame group for video quality assessment, and then determines whether to generate video quality assessment result data of the new historical video frame group.

2. The video transmission adaptive optimization method according to claim 1, characterized in that: The video quality assessment result data is used to represent the first quality feedback information or the second quality feedback information; The first quality feedback information is used to indicate that the video quality evaluation values ​​corresponding to the plurality of consecutive video frames in the historical video frame group are all higher than a preset upper threshold; The second quality feedback information is used to indicate that the video quality assessment values ​​corresponding to a plurality of consecutive video frames in the historical video frame group are all lower than a preset lower threshold; wherein the lower threshold is smaller than the upper threshold.

3. The video transmission adaptive optimization method according to claim 2, characterized in that: If the video quality assessment result data of the historical video frame group is obtained, the ratio of the commonality and individuality features corresponding to the target video frame group is adaptively optimized, including: If video quality assessment result data of a first historical video frame group is received and the video quality assessment result data currently represents the first quality feedback information, reducing the commonality-to-individuality feature ratio corresponding to the current target video frame group; The first historical video frame group is a historical video frame group located before the current target video frame group in the video data, and the first historical video frame group and the target video frame group are adjacent to or spaced apart by a preset number of video frame groups.

4. The video transmission adaptive optimization method according to claim 2, characterized in that: If the video quality assessment result data of the historical video frame group is obtained, the ratio of the commonality and individuality features corresponding to the target video frame group is adaptively optimized, including: If video quality assessment result data of a second historical video frame group is received and the video quality assessment result data currently represents the second quality feedback information, increasing the commonality-to-individuality feature ratio corresponding to the current target video frame group; The second historical video frame group is a historical video frame group located before the current target video frame group in the video data, and the second historical video frame group and the target video frame group are adjacent to or spaced apart by a preset number of video frame groups.

5. The video transmission adaptive optimization method according to claim 1, characterized in that: Before semantically encoding the target video frame group based on the current commonality-to-individuality feature ratio to obtain the commonality feature data and individuality feature data of the target video frame group, the method further includes: If, within a preset time period after the common feature data and individual feature data corresponding to the third historical video frame group are encapsulated and transmitted to the receiving end, no video quality assessment result data for the third historical video frame group is received, then the common feature to individual feature ratio corresponding to a historical video frame group preceding the third historical video frame group in the video data is used as the common feature to individual feature ratio corresponding to the target video frame group to be transmitted currently; The third historical video frame group is a historical video frame group located before the current target video frame group in the video data, and the third historical video frame group and the target video frame group are adjacent to or spaced apart by a preset number of video frame groups.

6. A video transmission adaptive optimization method, characterized in that: include: Based on the common feature data and individual feature data currently received from the sending end, obtaining a video frame group belonging to the video data corresponding to the common feature data and the individual feature data, and using the video frame group as a historical video frame group for video quality assessment to determine whether to generate video quality assessment result data for the historical video frame group; If video quality assessment result data of the historical video frame group is generated, the video quality assessment result data of the historical video frame group is sent to the sending end, so that the sending end adaptively optimizes the commonality and individuality feature ratio corresponding to the current target video frame group according to the video quality assessment result data, and enables the sending end to semantically encode the target video frame group based on the current commonality and individuality feature ratio to obtain the commonality feature data and individuality feature data of the target video frame group respectively, so as to encapsulate the commonality feature data and individuality feature data and transmit them, and the target video frame group is the video frame group to be transmitted corresponding to the video data.

7. The video transmission adaptive optimization method according to claim 6, characterized in that: The step of obtaining a video frame group belonging to the video data corresponding to the common feature data and the individual feature data based on the common feature data and the individual feature data currently received from the transmitting end includes: receiving a data packet encapsulating common feature data and individual feature data from a transmitting end, and extracting the common feature data and individual feature data from the data packet; It is determined whether the currently extracted common feature data and individual feature data belong to the same target video frame group in the video data. If so, semantic decoding is performed on the common feature data and individual feature data to obtain a corresponding video frame group.

8. The video transmission adaptive optimization method according to claim 6, characterized in that: The performing video quality assessment on the video frame group as a historical video frame group to determine whether to generate video quality assessment result data of the historical video frame group includes: The video frame group obtained after semantic decoding is used as a historical video frame group, and video quality assessment is performed on the historical video frame group to obtain video quality assessment values ​​corresponding to each of the multiple consecutive video frames in the historical video frame group, and it is determined whether the video quality assessment values ​​corresponding to each of the multiple consecutive video frames in the historical video frame group are all higher than a preset upper threshold or are all lower than a preset lower threshold, wherein the lower threshold is less than the upper threshold; If the video quality assessment values ​​corresponding to the plurality of consecutive video frames in the historical video frame group are all higher than a preset upper threshold, generating first quality feedback information and generating video quality assessment result data corresponding to the historical video frame group for representing the first quality feedback information; If the video quality assessment values ​​corresponding to multiple consecutive video frames in the historical video frame group are all lower than the preset lower threshold, second quality feedback information is generated and video quality assessment result data corresponding to the historical video frame group is generated for representing the second quality feedback information.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, it implements the video transmission adaptive optimization method according to any one of claims 1 to 5, and / or implements the video transmission adaptive optimization method according to any one of claims 6 to 8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the video transmission adaptive optimization method according to any one of claims 1 to 5, and / or implements the video transmission adaptive optimization method according to any one of claims 6 to 8.

Citation Information

Patent Citations

  • Code rate adaptive video semantic communication method and related device

    CN116896651A

  • Full-reference video quality evaluation method based on semantic communication and network channel parameters

    CN118158388A