Video transmission

By generating redundant packets for each frame and adjusting the redundancy using single-frame redundancy and reception status information, the stuttering problem caused by packet loss in video transmission is solved, achieving real-time recovery and smooth video transmission.

WO2025253214A1PCT designated stage Publication Date: 2025-12-11CLOUD INTELLIGENCE ASSETS HOLDING (SINGAPORE) PTE LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/IB2025/055212
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-03
Filing Date
2025-05-20
Publication Date
2025-12-11

AI Technical Summary

Technical Problem

In poor network conditions, existing video transmission technologies suffer from packet loss, leading to video stuttering. Existing forward error correction technologies require waiting for multiple frames of data before recovery, causing latency and stuttering issues.

Method used

By employing a method of determining redundancy in a single frame and adjusting the initial redundancy based on the received state information, redundant packets are generated for each frame to ensure that the receiving end can recover video frames in real time and reduce video transmission stuttering.

Benefits of technology

By instantly restoring video frames, video transmission stuttering is effectively reduced, improving the smoothness and reliability of video transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2025055212_11122025_PF_FP_ABST
    Figure IB2025055212_11122025_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a video transmission method, a video transmission system, a computing device, a computer-readable storage medium, and a computer program product. The video transmission method is applied to a sending end, and comprises: in response to a video transmission request for a target video, acquiring transmission information of a target video frame, wherein the target video frame is any video frame of the target video; determining an initial redundancy degree of the target video frame on the basis of the transmission information; adjusting the initial redundancy degree on the basis of receiving state information of a receiving end to obtain a target redundancy degree of the target video frame; and generating a target redundancy packet of the target video frame on the basis of the target redundancy degree, and transmitting the target video frame and the target redundancy packet to the receiving end.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Field of video transmission

[0002]

[0001] Embodiments of the present disclosure relate to the technical field of data transmission, and in particular, to video transmission. BACKGROUND

[0003]

[0002] In a video transmission scenario, packet loss has a very serious impact on the quality of real-time video communication. For example, it can cause video mosaics, stalls, second skipping and other problems, resulting in very poor user experience. In particular, in a poor network environment, packet loss is a common problem.

[0004]

[0003] Currently, redundant packets are generated for a certain number of video frames, and forward error correction techniques are used to alleviate the problems caused by packet loss. However, in the above scheme, the receiving end needs several frames of data before it can start packet loss recovery, resulting in a stall problem in the video transmission process. Therefore, there is an urgent need for a video transmission scheme that can effectively alleviate the stall problem. SUMMARY

[0005]

[0004] In view of this, embodiments of the present disclosure provide a video transmission method. One or more embodiments of the present disclosure also relate to a video transmission system, a video transmission device, a computing device, a computer-readable storage medium, and a computer program product, to solve the technical defects in the prior art.

[0006]

[0005] According to a first aspect of embodiments of the present disclosure, a video transmission method is provided, applied to a sending end, including: in response to a video transmission request for a target video, obtaining transmission information of a target video frame, wherein the target video frame is any video frame of the target video; determining an initial redundancy of the target video frame according to the transmission information; adjusting the initial redundancy according to reception state information of a receiving end to obtain a target redundancy of the target video frame, wherein the reception state information is used to describe a state of the receiving end when receiving the target video; generating a target redundant packet of the target video frame according to the target redundancy, and transmitting the target video frame and the target redundant packet to the receiving end.

[0007]

[0006] According to a second aspect of the embodiments of the present disclosure, a video transmission method is provided, applied to a receiving end, comprising: receiving a target video frame of a target video and a target redundancy package of the target video frame, wherein the target redundancy package is obtained based on a target redundancy degree of the target video frame, the target redundancy degree is obtained by adjusting an initial redundancy degree of the target video frame based on receiving state information of the receiving end, the initial redundancy degree is obtained based on transmission information of the target video frame, and the receiving state information is used to describe a state of the receiving end when receiving the target video; in a case where it is determined that the target video is packet lost, recovering the target video according to the target video frame and the target redundancy package to obtain a recovered target video.

[0008]

[0007] According to a third aspect of the embodiments of the present disclosure, a video transmission system is provided, comprising a cloud-side device and an end-side device; the cloud-side device is configured to, in response to a video transmission request for a target video, acquire transmission information of a target video frame, wherein the target video frame is any video frame of the target video; determine an initial redundancy degree of the target video frame according to the transmission information; adjust the initial redundancy degree according to receiving state information of the end-side device to obtain a target redundancy degree of the target video frame, wherein the receiving state information is used to describe a state of the receiving end when receiving the target video; generate a target redundancy package of the target video frame according to the target redundancy degree, and transmit the target video frame and the target redundancy package to the end-side device; and the end-side device is configured to, in a case where it is determined that the target video is packet lost, recover the target video according to the target video frame and the target redundancy package to obtain a recovered target video.

[0008] According to a fourth aspect of the embodiments of the present disclosure, a video transmission apparatus is provided, applied to a sending end, comprising: an acquisition module configured to, in response to a video transmission request for a target video, acquire transmission information of a target video frame, wherein the target video frame is any video frame of the target video; a determination module configured to determine an initial redundancy degree of the target video frame according to the transmission information; an adjustment module configured to adjust the initial redundancy degree according to receiving state information of the receiving end to obtain a target redundancy degree of the target video frame, wherein the receiving state information is used to describe a state of the receiving end when receiving the target video; and a transmission module configured to generate a target redundancy package of the target video frame according to the target redundancy degree, and transmit the target video frame and the target redundancy package to the receiving end.

[0009]

[0009] According to a fifth aspect of some embodiments of the present disclosure, a video transmission apparatus is provided, which is applied to a receiving end and includes: a receiving module configured to receive a target video frame of a target video and a target redundancy packet of the target video frame, wherein the target redundancy packet is obtained based on a target redundancy degree of the target video frame, the target redundancy degree is obtained by adjusting an initial redundancy degree of the target video frame based on receiving state information of the receiving end, the initial redundancy degree is obtained based on transmission information of the target video frame, and the receiving state information is used to describe a state of the receiving end when receiving the target video; and a recovery module configured to, in a case where it is determined that the target video is packet lost, recover the target video based on the target video frame and the target redundancy packet to obtain a recovered target video.

[0010]

[0010] According to a sixth aspect of some embodiments of the present disclosure, a computing device is provided, which includes a memory and a processor, the memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, and the computer programs / instructions, when executed by the processor, implement the steps of the method provided in the first aspect or the second aspect.

[0011]

[0011] According to a seventh aspect of some embodiments of the present disclosure, a computer readable storage medium is provided, which stores computer programs / instructions, and the computer programs / instructions, when executed by a processor, implement the steps of the method provided in the first aspect or the second aspect.

[0012]

[0012] According to an eighth aspect of some embodiments of the present disclosure, a computer program product is provided, which includes computer programs / instructions, and the computer programs / instructions, when executed by a processor, implement the steps of the method provided in the first aspect or the second aspect.

[0013]

[0013] According to an embodiment of the present disclosure, a video transmission method is provided, which is applied to a sending end and includes: in response to a video transmission request for a target video, obtaining transmission information of a target video frame, wherein the target video frame is any video frame of the target video; determining an initial redundancy degree of the target video frame based on the transmission information; adjusting the initial redundancy degree based on receiving state information of a receiving end to obtain a target redundancy degree of the target video frame, wherein the receiving state information is used to describe a state of the receiving end when receiving the target video; generating a target redundancy packet of the target video frame based on the target redundancy degree, and transmitting the target video frame and the target redundancy packet to the receiving end. The method determines the redundancy degree of each frame and adjusts the initial redundancy degree based on the receiving state information to generate a redundancy packet for each frame, so that the receiving end can start the recovery of the video packet for each frame, without waiting for the data of the next frame or several frames, thereby effectively reducing the video transmission lag. BRIEF DESCRIPTION OF DRAWINGS

[0014]

[0014] Figure 1 is an architecture diagram of a video transmission system according to one embodiment of the present disclosure;

[0015]

[0015] Figure 2 is an architecture diagram of another video transmission system according to one embodiment of the present disclosure;

[0016]

[0016] Figure 3 is a flow chart of a video transmission method according to one embodiment of the present disclosure;

[0017]

[0017] Figure 4 is a flow chart of determining redundancy in a video transmission method according to one embodiment of the present disclosure;

[0018] Figure 5 is a schematic diagram of generating a redundant packet in a video transmission method according to one embodiment of the present disclosure;

[0018]

[0019] Figure 6 is a flow chart of a processing procedure of a video transmission method according to one embodiment of the present disclosure;

[0019]

[0020] Figure 7 is a flow chart of a processing procedure of another video transmission method according to one embodiment of the present disclosure;

[0020]

[0021] Figure 8 is a flow chart of another video transmission method according to one embodiment of the present disclosure;

[0021]

[0022] Figure 9 is a structural schematic diagram of a video transmission apparatus according to one embodiment of the present disclosure;

[0022]

[0023] Figure 10 is a structural schematic diagram of another video transmission apparatus according to one embodiment of the present disclosure;

[0023]

[0024] Figure 11 is a structural block diagram of a computing device according to one embodiment of the present disclosure. DETAILED DESCRIPTION

[0024]

[0025] In the following description, a lot of specific details are set forth in order to fully understand the present disclosure. However, the present disclosure can be implemented in many different ways than described herein, and those skilled in the art can make similar extensions without departing from the spirit of the present disclosure, so the present disclosure is not limited to the specific implementation disclosed below.

[0025]

[0027] It should be understood that, although the terms first, second, etc. can be employed in describing various information in one or more embodiments of the present disclosure, the information should not be limited to such terms. These terms are only used to distinguish one particular information from another particular information. For example, a first can be termed a second, and, similarly, a second can be termed a first, without departing from the scope of one or more embodiments of the present disclosure. The word "if' as used herein can be interpreted as meaning "when" or "in response to determining" depending on the context.

[0026]

[0028] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in one or more embodiments of the present disclosure are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation portal for user to choose authorization or refusal.

[0027]

[0029] First, the nomenclature involved in one or more embodiments of the present disclosure is explained.

[0028]

[0030] Cloud desktop: a desktop computing mode based on cloud computing technology, which can convert a traditional local desktop environment into a virtual desktop environment based on a cloud server. Users can access computing resources of a cloud-side device through the Internet to achieve cross-device and cross-platform office and entertainment experience.

[0029]

[0031] Forward error correction (FEC): a streaming media transmission technology that adds redundant packets during data transmission to recover the original data packets through redundant packets at the receiving end, thereby improving the reliability of data transmission and reducing the packet loss rate.

[0032] Round-trip time (RTT): the round-trip time of data from the sending end to the receiving end, usually expressed in milliseconds (ms).

[0030]

[0033] Web real-time communication (Web RTC): a real-time communication protocol used to implement voice, video and data transmission in Web applications.

[0031]

[0034] Windows Remote Desktop Protocol (Windows RDP): A remote desktop protocol used for remote desktop access on Windows operating systems.

[0032]

[0035] Choppiness: The frequency of choppiness during video playback. Specifically, choppiness can measure the proportion of choppiness events relative to the total playback time or viewing session. Choppiness usually includes video playback interruption, frozen screen, buffer wheel waiting for loading, etc., which seriously affects the viewing experience.

[0033]

[0036] Packet loss rate: The proportion of the number of lost packets to the total number of packets during data transmission.

[0034]

[0037] Resolution: The image clarity displayed during video playback, usually represented by the number of pixels.

[0035]

[0038] Key frame: A frame with important significance during video playback, such as action, expression, etc.

[0036]

[0039] In the video transmission scenario, packet loss has a very serious impact on the quality of real-time video communication. Taking the cloud desktop scenario as an example, the cloud desktop has higher requirements than the general network scenario (such as video conference scenario, etc.). In order to reduce the choppiness rate of the cloud desktop scenario under the packet loss scenario, the FEC technology is generally used to generate redundant packets based on a group of video frames, and the lost video packets are recovered at the receiver through the redundant packets.

[0037]

[0040] Currently, the freezing phenomenon in the packet loss scenario can be generally addressed by the solutions of Windows RDP or Web RTC. Taking Web RTC as an example, Web RTC can determine the number of currently cached video frames at the end of a video frame, and start generating redundant packets according to the redundancy when the configured minimum number of frames is reached. In the above solution, the redundancy is obtained by table lookup, which depends on empirical values and has a small range of applicable scenarios. At the same time, Web RTC starts to generate redundant packets only after a certain number of video frames are reached, which may contain both key frames and ordinary frames. At this time, whether the redundancy of ordinary frames or the redundancy of key frames is used to calculate and generate redundant packets, the effect is poor. Moreover, due to the low frame rate of the cloud desktop and the requirement of low latency, the method of generating redundant packets after receiving a certain number of video frames will cause the receiving end to wait for video frames during the recovery process, resulting in freezing, bandwidth consumption and delay problems. For example, assuming a frame rate of 30 frames per second and a data interval of 30 ms per frame, if a frame of data is lost and it still needs to wait for 2 frames of data to recover, it will take 60 ms to recover, which will cause freezing.

[0038]

[0041] To solve the above problems, the embodiments of the present disclosure propose a video transmission scheme that generates redundant packets using FEC technology, calculates the initial redundancy, adjusts the initial redundancy according to the receiving status information of the receiving end to calculate the redundancy, and at the same time adjusts the redundant packet generation method to increase the single frame (single video frame) redundant packet generation method to achieve the purpose of reducing the redundancy packet bandwidth while also reducing the freezing rate.

[0039]

[0042] Specifically, the video transmission scheme is applied to a sending end and includes: in response to a video transmission request for a target video, obtaining transmission information of a target video frame, wherein the target video frame is any video frame of the target video; determining an initial redundancy of the target video frame according to the transmission information; adjusting the initial redundancy according to receiving status information of a receiving end to obtain a target redundancy of the target video frame; generating a target redundant packet of the target video frame according to the target redundancy, and transmitting the target video frame and the target redundant packet to the receiving end. The single frame redundancy determination and the initial redundancy adjustment based on the receiving status information are used to generate a redundant packet for each frame, so that the receiving end can start video packet recovery at each frame without waiting for the next frame or several frames of data, thereby effectively reducing the video transmission freezing.

[0040]

[0043] In the present disclosure, a video transmission method is provided, and the present disclosure also relates to a video transmission system, a video transmission device, a computing device, a computer-readable storage medium and a computer program product, which are described in detail one by one in the following embodiments.

[0041]

[0044] Referring to FIG. 1, FIG. 1 shows an architecture diagram of a video transmission system provided by an embodiment of the present disclosure, and the video transmission system can include a sending end 100 and a receiving end 200.

[0042]

[0045] The sending end 100 is configured to, in response to a video transmission request for a target video, acquire transmission information of a target video frame, wherein the target video frame is any video frame of the target video; determine an initial redundancy of the target video frame according to the transmission information; adjust the initial redundancy according to reception state information of the receiving end 200 to obtain a target redundancy of the target video frame, wherein the reception state information is used to describe a state of the receiving end 200 when receiving the target video; generate a target redundancy packet of the target video frame according to the target redundancy, and transmit the target video frame and the target redundancy packet to the receiving end 200.

[0043]

[0046] The receiving end 200 is configured to, in a case where it is determined that the target video is packet lost, recover the target video according to the target video frame and the target redundancy packet to obtain a recovered target video.

[0044]

[0047] By using the scheme of the embodiments of the present disclosure, the single-frame redundancy is determined, and the initial redundancy is adjusted based on the reception state information to generate a redundancy packet for each frame, so that the receiving end can start the recovery of the video packet for each frame, without waiting for the data of the next frame or several frames, thereby effectively reducing the video transmission lag.

[0045]

[0048] In the embodiments of the present disclosure, the sending end and the receiving end can both be end-side devices, for example, the sending end can be a first end-side device, and the receiving end can be a second end-side device; the sending end and the receiving end can both be cloud-side devices, for example, the sending end can be a first cloud-side device, and the receiving end can be a second cloud-side device; the sending end and the receiving end can also be an end-side device and a cloud-side device respectively, so as to perform end-cloud interaction, for example, the sending end can be an end-side device, and the receiving end can be a cloud-side device, or the sending end can be a cloud-side device, and the receiving end can be an end-side device.

[0046]

[0049] Referring to FIG. 2, FIG. 2 shows an architecture diagram of another video transmission system provided by an embodiment of the present disclosure, the video transmission system can include a cloud-side device 102 and an end-side device 202.

[0047]

[0050] The cloud-side device 102 is configured to, in response to a video transmission request for a target video, acquire transmission information of a target video frame, where the target video frame is any video frame of the target video; determine an initial redundancy of the target video frame according to the transmission information; adjust the initial redundancy according to reception state information of the end-side device 202 to obtain a target redundancy of the target video frame, where the reception state information is used to describe a state of the end-side device 202 when receiving the target video; and generate a target redundancy packet of the target video frame according to the target redundancy, and transmit the target video frame and the target redundancy packet to the end-side device 202.

[0048]

[0051] The end-side device 202 is configured to, in a case where it is determined that the target video is packet loss, recover the target video according to the target video frame and the target redundancy packet to obtain a recovered target video.

[0049]

[0052] By using the scheme provided by the embodiments of the present disclosure, a single-frame redundancy is determined, and the initial redundancy is adjusted based on the reception state information to generate a redundancy packet for each frame, so that the end-side device can start the recovery of the video packet for each frame, without waiting for the data of the next frame or several frames, thereby effectively reducing the video transmission lag.

[0053] In actual application, the end-side device can be referred to as a client device or an edge device, which refers to a hardware device that directly interacts with a user or is located at a data generation source. The end-side device can be deployed in an electronic device and run in dependence on the device or some APP in the device. The electronic device can have a display screen and support information browsing, for example, and can be a personal mobile terminal such as a mobile phone, a tablet computer, a personal computer, etc. Various other types of applications can also be configured in the electronic device, for example, human-computer dialogue applications, text processing applications, web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.

[0050]

[0054] The cloud-side device refers to servers, storage systems, network devices and other infrastructures located in remote data centers, which form the core of cloud computing. The cloud-side device can provide powerful computing resources, storage space and advanced services (such as database management, big data analysis, machine learning platform) for various application programs and services. In the cloud computing architecture, data is usually uploaded from the end-side device to the cloud side for centralized processing, analysis and long-term storage. The advantage of the cloud-side device is that it has good scalability and efficient resource sharing, and can handle a large number of concurrent requests and complex computation tasks. The cloud-side device can include servers that provide various services, such as servers that provide communication services for multiple end-side devices, servers that provide support for models used on end-side devices for background training, servers that process data sent by end-side devices, etc. It should be noted that the cloud-side device can be implemented as a distributed server cluster composed of multiple servers, or as a single server. The server can also be a server of a distributed system, or a server combined with a blockchain. The server can also be a cloud server of cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), big data and artificial intelligence platforms, and other basic cloud computing services, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology.

[0051]

[0055] Referring to FIG. 3, FIG. 3 shows a flowchart of a video transmission method according to an embodiment of the present disclosure. The video transmission method is applied to a sending end, and specifically includes the following steps.

[0052]

[0056] Step 302: In response to a video transmission request for a target video, transmission information of a target video frame is obtained, wherein the target video frame is any video frame of the target video.

[0053]

[0057] In one or more embodiments of the present disclosure, the sending end starts to transmit the target video in response to a video transmission request for the target video. During the transmission of the target video, the transmission information of the target video frame can be obtained, so as to determine the target redundancy of the target video frame based on the transmission information, and generate the target redundancy package of the target video frame based on the target redundancy, and then implement the transmission of the target video frame and the target redundancy package by using the FEC technology.

[0054]

[0058] Specifically, the target video can be a video in different scenarios, such as a conference video in a conference scenario, an office video in an office scenario, and the like. The target video is a dynamic picture composed of a series of static images, each of which is referred to as a frame. The frame is a basic unit of the video. The transmission information of the target video frame can include attribute information of the target video frame, such as resolution, and can also include transmission process information of the target video frame, such as round-trip delay information, packet loss information, code rate bandwidth information, and the like.

[0055]

[0059] It should be noted that the video transmission request is used to request the sending end to transmit the target video to the receiving end. The video transmission request can be sent by the receiving end, or can be sent by other ends than the receiving end, and the specific implementation is selected according to actual conditions, and the embodiments of the present disclosure do not make any limitation on this. Further, in response to the video transmission request for the target video, before the transmission information of the target video frame is obtained, it can be determined whether the target video frame is ended, and in the case that the target video frame is ended, the transmission information of the target video frame is obtained, so as to realize single-frame calculation of redundancy and generation of a redundant packet.

[0060] In actual application, there are various ways to obtain the transmission information of the target video frame, and the specific implementation is selected according to actual conditions, and the embodiments of the present disclosure do not make any limitation on this. In a possible implementation of the present disclosure, the transmission information of the target video frame sent by the receiving end can be received. In another possible implementation of the present disclosure, the video transmission log of the target video can be obtained, the video transmission log is parsed, and the transmission information of the target video frame is obtained.

[0056]

[0061] Step 304: According to the transmission information, the initial redundancy of the target video frame is determined.

[0057]

[0062] In one or more embodiments of the present disclosure, in response to the video transmission request for the target video, after the transmission information of the target video frame is obtained, further, the initial redundancy of the target video frame can be determined according to the transmission information.

[0058]

[0063] Specifically, redundancy refers to redundancy information calculated in a video transmission process for ensuring reliable transmission of a video. The redundancy is used to describe the number of redundant packets, and based on the redundancy, redundant packets for recovering the target video can be generated. The redundant packet refers to a video packet additionally added in the target video in the video transmission process, and the additionally added video packet can be a media packet. In the video transmission process, the receiving end can recover the lost or damaged video packet through the redundant packet even if it does not receive all the video packets of the target video, thereby ensuring the continuity and quality of the video. The initial redundancy refers to the basic redundancy calculated based on the transmission information of the target video frame.

[0059]

[0064] In actual applications, there are various ways to determine the initial redundancy of the target video frame based on the transmission information, which are specifically selected according to actual conditions, and the present disclosure does not make any limitation on this. In a possible implementation manner of the present disclosure, the initial redundancy of the target video frame can be determined from the reference redundancy carried in at least one transmission reference information according to the transmission information. In another possible implementation manner of the present disclosure, the unit redundancy corresponding to the unit transmission information can be acquired, and the initial redundancy of the target video frame is determined according to the unit transmission information, the transmission information and the unit redundancy. For example, assuming that the unit transmission information is “packet loss rate N”, the unit redundancy corresponding to the unit transmission information is n, the transmission information is “packet loss rate 2N”, and then the initial redundancy is 2n.

[0060]

[0065] In an optional embodiment of the present disclosure, the transmission information includes at least one of round-trip delay information, packet loss information, resolution and code rate bandwidth information; the above-mentioned determining the initial redundancy of the target video frame according to the transmission information can include the following steps: acquiring at least one transmission reference information, wherein the transmission reference information carries reference redundancy; according to the transmission information, the target transmission reference information is screened out from the at least one transmission reference information, and the reference redundancy carried by the target transmission reference information is determined as the initial redundancy of the target video frame, wherein the target transmission reference information is the transmission reference information in the at least one transmission reference information, and the similarity between the transmission information is greater than a preset threshold.

[0061]

[0066] Specifically, the round-trip time information is used to describe the round-trip time of the target video frame from the sending end to the receiving end. The packet loss information is used to describe the packet loss condition in the target video transmission process, such as the packet loss rate. The bit rate refers to the amount of information transmitted per second, that is, the data rate after video encoding. The bandwidth is the maximum video transmission rate that the network can provide. The bit rate bandwidth information includes but is not limited to the bit rate bandwidth ratio. The target transmission reference information refers to at least one transmission reference information with a similarity greater than a preset threshold to the transmission information, and the preset threshold is set according to actual conditions.

[0062]

[0067] It should be noted that there are many ways to obtain at least one transmission reference information, which is selected according to actual conditions, and the present disclosure does not make any limitation on this. In a possible implementation of the present disclosure, a plurality of transmission reference information carrying reference redundancy can be read from other data acquisition devices or databases. In another possible implementation of the present disclosure, a plurality of transmission reference information carrying reference redundancy configured by a user based on prior knowledge can be received.

[0068] In practical applications, if the RTT is high, it means that the network delay is large, and there may be a high risk of packet loss, at this time, the redundancy can be appropriately increased to ensure the correct reception of the video; if the network packet loss rate is high, the method of increasing the redundancy can be used to ensure the integrity of the video; under the condition of fixed bandwidth, high-resolution video may need higher bit rate to maintain good picture quality, at this time, a certain redundancy can be increased in the encoding process to allow a certain loss of picture quality in exchange for better anti-packet loss performance. In the process of video encoding and transmission, introducing redundant packets can enhance the anti-packet loss ability and error recovery ability, but this will increase the bit rate. Therefore, the bit rate bandwidth ratio can reflect the degree of redundancy design of the system in resisting network instability. For example, if the bit rate continuously approaches or exceeds the bandwidth, then without taking additional redundant encoding, the system has low tolerance to packet loss; on the contrary, if the bit rate is much lower than the bandwidth, and appropriate redundant encoding technology (such as the video transmission method proposed in the present disclosure) is adopted, even if there is a certain packet loss, the lost content can be recovered through redundant data, improving the reliability and quality of transmission.

[0063]

[0069] According to the scheme of the embodiment of the present disclosure, at least one transmission reference information is obtained, wherein the transmission reference information carries reference redundancy; target transmission reference information is selected from the at least one transmission reference information according to the transmission information, and the reference redundancy carried by the target transmission reference information is determined as the initial redundancy of the target video frame. By comparing the transmission reference information with the transmission information, the reference redundancy carried by the target transmission reference information similar to the transmission information is determined as the initial redundancy of the target video frame, so that the initial redundancy can be quickly and accurately determined.

[0064]

[0070] In an optional embodiment of the present disclosure, in order to reduce the video transmission stall rate and improve the video recovery rate, the initial redundancy of the key video frame can be increased, that is, after the initial redundancy of the target video frame is determined according to the transmission information, the following steps can be further included: type identification is performed on the target video frame to determine the frame type of the target video frame; the initial redundancy is adjusted according to the receiving state information of the receiving end to obtain the target redundancy of the target video frame, which can include the following steps: the adjustment redundancy of the target video frame is determined according to the frame type and the initial redundancy; the adjustment redundancy is adjusted according to the receiving state information of the receiving end to obtain the target redundancy of the target video frame.

[0065]

[0071] Specifically, frame types can be classified as key and non-key. Based on frame types, video frames can be classified as key video frames (key frames) and non-key video frames (non-key frames). Key video frames, such as I frames (Intra Frame), contain complete image information and can be decoded independently without reference to other frames. In a video sequence, I frames are used as a reference to start a new scene or are inserted after a time interval to facilitate random access and editing. The compression rate of I frames is relatively low because they do not utilize temporal redundancy, but they are essential components of a decoded sequence. Non-key video frames, such as P frames (Predictive Frame) and B frames (Bi-directional Predictive Frame), are encoded by referencing a previous frame (which can be an I frame or a P frame) for P frames, and by referencing both a previous frame and a subsequent frame for B frames, to achieve higher compression rates. P frames utilize temporal redundancy to improve compression rates, but require previous frames for decoding. If the previous frame referenced by a P frame is lost, the P frame cannot be decoded correctly. B frames utilize information from both previous and subsequent frames to remove more spatial and temporal redundancy, resulting in higher compression rates, but require context information from both frames for decoding. The presence of B frames can significantly improve compression rates, but requires a specific decoding order and real-time performance.

[0066]

[0072] In practical applications, there are various ways to identify the type of a target video frame and determine the frame type of the target video frame. The specific implementation is selected according to actual conditions, and the present disclosure does not make any limitation in this regard. In one possible implementation of the present disclosure, a video decoding library or tool can be used to identify the type of a target video frame and determine the frame type of the target video frame. For example, open-source multimedia processing libraries such as FFmpeg and Gstreamer can be used, which provide application programming interfaces (APIs) or command-line tools to parse video streams. The frame type of each frame can be directly output or obtained through a programming interface. In another possible implementation of the present disclosure, the specific data of a video frame can be analyzed in depth to find a flag or code used to identify the frame type. For example, in H.264, the frame type of a target video frame can be determined by reading information in a slice header.

[0067]

[0073] It should be noted that there are various ways to determine the adjusted redundancy of the target video frame according to the frame type and the initial redundancy, which can be selected according to actual conditions, and the embodiments of the present disclosure do not make any limitation in this regard. In a possible implementation of the present disclosure, the initial redundancy of the key video frame can be adjusted to determine the adjusted redundancy of the target video frame. In another possible implementation of the present disclosure, the initial redundancies of the key video frame and the non-key video frame can be adjusted to determine the adjusted redundancy of the target video frame. For example, when the transmission bandwidth is constant, the initial redundancy of the key video frame is increased, and the initial redundancy of the non-key video frame is reduced to obtain the adjusted redundancy of the target video frame.

[0068]

[0074] By applying the scheme of the embodiments of the present disclosure, the frame type of the target video frame is identified to determine the frame type of the target video frame, the adjusted redundancy of the target video frame is determined according to the frame type and the initial redundancy, and the adjusted redundancy is adjusted to obtain the target redundancy of the target video frame according to the receiving status information of the receiving end. By determining the adjusted redundancy of the target video frame according to the frame type and the initial redundancy, the redundancy of the key video frame is increased by using the frame information, so that the recovery rate of the key video frame is high, and the stall rate can also be reduced.

[0069]

[0075] In an optional embodiment of the present disclosure, the above-mentioned determination of the adjusted redundancy of the target video frame according to the frame type and the initial redundancy can include the following steps: in a case where the target video frame is determined to be a non-key video frame based on the frame type, the initial redundancy is determined as the adjusted redundancy of the target video frame; and in a case where the target video frame is determined to be a key video frame based on the frame type, the initial redundancy is adjusted to obtain the adjusted redundancy of the target video frame.

[0070]

[0076] It should be noted that when the initial redundancy of the key video frame is adjusted, the initial redundancy can be increased to obtain the adjusted redundancy of the target video frame in order to ensure the high recovery rate of the key video frame, that is, the adjusted redundancy of the key video frame is greater than the initial redundancy of the key video frame.

[0071]

[0077] In actual applications, there are various ways to increase the initial redundancy, which are specifically selected according to actual conditions, and the embodiments of the present disclosure do not make any limitation in this regard. In a possible implementation manner of the present disclosure, a preset redundancy can be acquired, the preset redundancy is added on the basis of the initial redundancy, and the adjustment redundancy of the target video frame is obtained, where the preset redundancy is specifically selected according to actual conditions, and the embodiments of the present disclosure do not make any limitation in this regard. In another possible implementation manner of the present disclosure, the initial redundancy can be randomly increased to obtain the adjustment redundancy of the target video frame.

[0072]

[0078] According to the scheme of the embodiments of the present disclosure, in a case where it is determined based on the frame type that the target video frame is a non-key video frame, the initial redundancy is determined as the adjustment redundancy of the target video frame; in a case where it is determined based on the frame type that the target video frame is a key video frame, the initial redundancy is adjusted to obtain the adjustment redundancy of the target video frame. The initial redundancy of the non-key video frame is not adjusted, and only the initial redundancy of the key video frame is increased, so that the non-key frame generates less redundant packets, the bandwidth can be reduced, the key frame generates more redundant packets, the recovery rate of the receiving end is higher, and the freezing is reduced.

[0073]

[0079] Step 306: adjusting the initial redundancy according to the receiving state information of the receiving end to obtain the target redundancy of the target video frame, where the receiving state information is used to describe the state of the receiving end when receiving the target video.

[0074]

[0080] In one or more embodiments of the present disclosure, in response to a video transmission request for a target video, transmission information of a target video frame is acquired; after the initial redundancy of the target video frame is determined according to the transmission information, further, the initial redundancy can be adjusted according to the receiving state information of the receiving end to obtain the target redundancy of the target video frame.

[0075]

[0081] Specifically, the receiving state information includes but is not limited to a receiving freezing rate and a receiving bandwidth, which are specifically selected according to actual conditions, and the embodiments of the present disclosure do not make any limitation in this regard.

[0076]

[0082] In actual applications, there are various ways to adjust the initial redundancy according to the receiving status information of the receiving end to obtain the target redundancy of the target video frame, which are selected according to actual conditions, and the embodiments of the present disclosure do not make any limitation on this. In a possible implementation manner of the present disclosure, reference status information can be obtained, the receiving status information is compared with the reference status information to obtain a status comparison result, and the initial redundancy is adjusted according to the status comparison result to obtain the target redundancy of the target video frame. The reference status information is, for example, a reference stall rate, which is set according to actual needs of video transmission, and the embodiments of the present disclosure do not make any limitation on this. When the initial redundancy is adjusted according to the status comparison result, a preset redundancy adjustment value can be adjusted each time until the difference between the receiving status information and the reference status information is within a preset status difference range. The preset redundancy adjustment value and the preset status difference range are set according to actual needs, and the embodiments of the present disclosure do not make any limitation on this. In another possible implementation manner of the present disclosure, a candidate redundancy corresponding to the receiving status information can be determined, and the initial redundancy is adjusted by using the candidate redundancy to obtain the target redundancy of the target video frame. The candidate redundancy can be positive or negative, corresponding to increasing or decreasing the initial redundancy.

[0077]

[0083] In an optional embodiment of the present disclosure, taking the receiving status information including a receiving stall rate as an example, the above-mentioned adjusting the initial redundancy according to the receiving status information of the receiving end to obtain the target redundancy of the target video frame can include the following steps: obtaining a reference stall rate; comparing the receiving stall rate with the reference stall rate to obtain a comparison result; and adjusting the initial redundancy according to the comparison result to obtain the target redundancy of the target video frame.

[0078]

[0084] Specifically, the reference stall rate refers to a stall rate meeting the needs of video transmission, which is determined based on a video transmission environment and a bandwidth. The reference stall rate is, for example, 5%, and the reference stall rate is set according to actual conditions. There are various ways to obtain the reference stall rate, which are selected according to actual conditions, and the embodiments of the present disclosure do not make any limitation on this. In a possible implementation manner of the present disclosure, the reference stall rate can be read from other data acquisition devices or databases. In another possible implementation manner of the present disclosure, the reference stall rate configured by a user based on the needs of video transmission can be received. The comparison result can be a size comparison result of the receiving stall rate and the reference stall rate, for example, the receiving stall rate is greater than the reference stall rate. The comparison result can also be a numerical comparison result of the receiving stall rate and the reference stall rate, for example, the receiving stall rate is twice as large as the reference stall rate.

[0079]

[0085] In actual applications, according to the comparison result, the initial redundancy is adjusted to obtain the target redundancy of the target video frame. There are various methods for adjusting the initial redundancy to obtain the target redundancy of the target video frame, which are selected according to actual conditions, and the embodiments of the present disclosure do not make any limitation on this. In a possible implementation manner of the present disclosure, in the case that the receiving frame freezing rate is greater than the reference frame freezing rate, the initial redundancy is adjusted to obtain the target redundancy. In the case that the receiving frame freezing rate is less than or equal to the reference frame freezing rate, the initial redundancy is determined as the target redundancy. In another possible implementation manner of the present disclosure, since the excessively high or excessively low receiving frame freezing rate may affect the video transmission quality, if it is determined according to the comparison result that the difference between the receiving frame freezing rate and the reference frame freezing rate is outside the preset frame freezing rate range, the initial redundancy is adjusted to obtain the target redundancy of the target video frame.

[0080]

[0086] By applying the scheme of the embodiments of the present disclosure, the reference frame freezing rate is obtained; the receiving frame freezing rate and the reference frame freezing rate are compared to obtain a comparison result; and the initial redundancy is adjusted according to the comparison result to obtain the target redundancy of the target video frame. By comparing the receiving frame freezing rate fed back by the receiving end with the reference frame freezing rate, the initial redundancy is dynamically adjusted, so that the target redundancy is more accurate. Referring to FIG. 4, FIG. 4 shows a flowchart of determination of redundancy in a video transmission method provided by an embodiment of the present disclosure. The determination process of the single-frame redundancy includes the following three stages: In the first stage, the initial redundancy of the target video frame is determined according to the transmission information, wherein the transmission information includes at least one of the round-trip delay information, the packet loss information, the resolution and the code rate bandwidth information; In the second stage, the adjustment redundancy of the target video frame is determined according to the frame type and the initial redundancy; for example, when the adjustment redundancy of the target video frame is determined according to the frame type and the initial redundancy, it can be determined whether the video frame is a key video frame, so as to increase the redundancy of the key video frame; In the third stage, the adjustment redundancy is adjusted according to the receiving state information of the receiving end to obtain the target redundancy of the target video frame.

[0081]

[0088] Step 308: According to the target redundancy, a target redundancy packet of the target video frame is generated, and the target video frame and the target redundancy packet are transmitted to the receiving end.

[0082]

[0089] In one or more embodiments of the present disclosure, in response to a video transmission request for a target video, transmission information of a target video frame is obtained; an initial redundancy of the target video frame is determined according to the transmission information; the initial redundancy is adjusted according to reception status information of a receiving end to obtain a target redundancy of the target video frame. Further, a target redundancy package of the target video frame can be generated according to the target redundancy, and the target video frame and the target redundancy package are transmitted to the receiving end.

[0083]

[0090] In actual applications, when the target redundancy package of the target video frame is generated according to the target redundancy, a coding scheme such as Reed-Solomon coding (RS) or Low Density Parity Check Code (LDPC) can be used to generate the target redundancy package of the target video frame according to the target redundancy package, and the target video frame and the target redundancy package are transmitted to the receiving end.

[0084]

[0091] By using the scheme of the present disclosure, the single-frame redundancy is determined, and the initial redundancy is adjusted based on the reception status information to generate a redundancy package for each frame, so that the receiving end can start video package recovery for each frame, without waiting for data of the next frame or several frames, thereby effectively reducing video transmission lag.

[0085]

[0092] In an optional embodiment of the present disclosure, in order to ensure that the packet loss of the target video can be successfully recovered, the scheme of single-frame redundancy package generation as the main and group-frame redundancy package generation as the auxiliary is proposed in the embodiment of the present disclosure, that is, the above-mentioned video transmission method can further include the following steps: in response to a video transmission request for a target video, a set redundancy of a video frame set is determined, wherein the video frame set includes a plurality of video frames of the target video; a set redundancy package of the video frame set is generated according to the set redundancy, and the video frame set and the set redundancy package are transmitted to the receiving end.

[0086]

[0093] Specifically, the video frame set includes a plurality of video frames, so the video frame set can be understood as a group frame. The set redundancy can be understood as a group frame redundancy, and the set redundancy package can be understood as a group frame redundancy package.

[0087]

[0094] It should be noted that the manner of "generating a set redundancy package of a set of video frames according to a set redundancy, and transmitting the set of video frames and the set redundancy package to a receiving end" is the same as the above-mentioned "generating a target redundancy package of a target video frame according to a target redundancy, and transmitting the target video frame and the target redundancy package to a receiving end", and the embodiments of the present disclosure will not be described again.

[0088]

[0095] In practical applications, there are various ways to determine the set redundancy of the set of video frames, which are specifically selected according to actual conditions, and the embodiments of the present disclosure do not make any limitation. In a possible implementation manner of the present disclosure, the set redundancy of the set of video frames can be queried from a redundancy configuration table. In the redundancy configuration table, a plurality of frame sets and the set redundancy of each frame set are included. In another possible implementation manner of the present disclosure, the set transmission information of the set of video frames can be obtained; and the set redundancy of the set of video frames is determined according to the set transmission information.

[0089]

[0096] By applying the scheme of the embodiments of the present disclosure, in response to a video transmission request for a target video, the set redundancy of a set of video frames is determined, wherein the set of video frames includes a plurality of video frames of the target video; and a set redundancy package of the set of video frames is generated according to the set redundancy, and the set of video frames and the set redundancy package are transmitted to a receiving end. By the manner of taking a single-frame redundancy package as the main and a group-frame redundancy package as the auxiliary, not only can it be ensured that the receiving end can start video package recovery at each frame, but also the group-frame redundancy can be reduced, and in the case that the single-frame redundancy package cannot recover the video package, the group-frame redundancy package is used for continuous recovery, so as to ensure that the video packet loss problem is effectively solved.

[0090]

[0097] In an optional embodiment of the present disclosure, the above-mentioned "determining the set redundancy of the set of video frames in response to the video transmission request for the target video" can include the following steps: in response to the video transmission request for the target video, obtaining set transmission information of the set of video frames, wherein the set transmission information is obtained based on transmission information of a plurality of video frames; determining an initial set redundancy of the set of video frames according to the set transmission information, wherein the initial set redundancy is an initial redundancy of the plurality of video frames; and adjusting the initial set redundancy according to reception state information of the receiving end to obtain the set redundancy of the set of video frames, wherein the set redundancy is a target redundancy of the plurality of video frames.

[0091]

[0098] It should be noted that the implementation manner of "obtaining the set transmission information of the video frame set in response to the video transmission request for the target video; determining the initial set redundancy of the video frame set according to the set transmission information; and adjusting the initial set redundancy according to the reception state information of the receiving end to obtain the set redundancy of the video frame set" is the same as the implementation manner of "obtaining the transmission information of the target video frame in response to the video transmission request for the target video; determining the initial redundancy of the target video frame according to the transmission information; and adjusting the initial redundancy according to the reception state information of the receiving end to obtain the target redundancy of the target video frame", and the present embodiment of the present disclosure will not be described in detail.

[0092]

[0099] According to the scheme of the present embodiment of the present disclosure, the set transmission information of the video frame set is obtained in response to the video transmission request for the target video; the initial set redundancy of the video frame set is determined according to the set transmission information; and the initial set redundancy is adjusted according to the reception state information of the receiving end to obtain the set redundancy of the video frame set. The initial set redundancy is dynamically adjusted according to the reception state information fed back by the receiving end, so that the set redundancy is more accurate.

[0093]

[0100] Referring to FIG. 5, FIG. 5 shows a schematic diagram of generation of redundant packets in a video transmission method according to an embodiment of the present disclosure. As shown in FIG. 5, the target video includes three video frames, which are video frame 1, video frame 2 and video frame 3. The video frame 2 is a key video frame, and the video frame 1 and the video frame 3 are non-key video frames. The video frame 1 includes four video packets, the video frame 2 includes five video packets, and the video frame 3 includes three video packets. According to the single-frame redundant packet generation method, one redundant packet is generated for each video packet in the video frame 1, two redundant packets are generated for each video packet in the video frame 2, and one redundant packet is generated for each video packet in the video frame 3. The redundant packets generated for the three video frames can be referred to as single-frame redundant packets. Then, according to the group-frame redundant packet generation method, one redundant packet is generated for the video frame 1, the video frame 2 and the video frame 3, which can be referred to as a group-frame redundant packet.

[0094]

[0101] Referring to FIG. 6, FIG. 6 shows a processing procedure flow chart of a video transmission method according to an embodiment of the present disclosure. As shown in FIG. 6, the manner of generating a single-frame redundancy packet can include: when a sending end transmits a plurality of video packets of a target video, determining whether a single target video frame is ended, if yes, determining whether the target video frame is a key frame, if yes, determining the redundancy of the key frame and generating a redundancy packet, if no, determining the redundancy of a non-key frame and generating a redundancy packet. Finally, the sending end sends the video packets and the redundancy packets corresponding to the video packets to a receiving end.

[0095]

[0102] It should be noted that the determining the redundancy of the key frame and generating the redundancy packet can include the following steps: obtaining transmission information of the target video frame; determining an initial redundancy of the target video frame according to the transmission information; in a case where it is determined that the target video frame is a key frame based on the frame type, adjusting the initial redundancy to obtain an adjusted redundancy of the target video frame; adjusting the adjusted redundancy according to reception state information of the receiving end to obtain a target redundancy of the target video frame; and generating a target redundancy packet of the target video frame according to the target redundancy.

[0096]

[0103] The determining the redundancy of the non-key frame and generating the redundancy packet can include the following steps: obtaining transmission information of the target video frame; determining an initial redundancy of the target video frame according to the transmission information; adjusting the initial redundancy according to reception state information of the receiving end to obtain a target redundancy of the target video frame; and generating a target redundancy packet of the target video frame according to the target redundancy.

[0097]

[0104] It should be noted that the manner of generating a single-frame redundancy packet does not need to wait for other frames in the video recovery process, and the recovery is fast, which can reduce the stall rate. In addition, the frame type can be used to increase the redundancy of the key frame, so that the recovery rate of the key frame is high, and the stall rate can also be reduced.

[0098]

[0105] Referring to FIG. 7, FIG. 7 shows a process flow chart of another video transmission method provided by an embodiment of the present disclosure. As shown in FIG. 7, the single-frame redundancy packet generation can be combined with the group-frame redundancy packet generation in the following manner. When the sending end transmits the multiple video packets of the target video, it determines whether a single target video frame ends. If yes, it determines whether the target video frame is a key frame. If yes, it determines the redundancy of the key frame and generates a redundancy packet. If no, it determines the redundancy of the non-key frame and generates a redundancy packet. Meanwhile, when the sending end transmits the multiple video packets of the target video, it determines whether a group frame ends. If yes, it determines the group-frame redundancy and generates a group-frame redundancy packet. Finally, the sending end sends the video packets and the redundancy packets corresponding to the video packets to the receiving end.

[0099]

[0106] It should be noted that the determination of the group-frame redundancy and the generation of the group-frame redundancy packet can include the following steps. The group-frame redundancy of the group frame is determined. The group-frame redundancy packet of the group frame is generated according to the group-frame redundancy.

[0100]

[0107] Referring to FIG. 8, FIG. 8 shows a flow chart of another video transmission method provided by an embodiment of the present disclosure. The video transmission method is applied to the receiving end and specifically includes the following steps.

[0101]

[0108] Step 802: receiving a target video frame of a target video and a target redundancy packet of the target video frame, wherein the target redundancy packet is obtained based on a target redundancy of the target video frame, the target redundancy is obtained by adjusting an initial redundancy of the target video frame based on receiving state information of the receiving end, and the initial redundancy is obtained based on transmission information of the target video frame, and the receiving state information is used to describe the state of the receiving end when receiving the target video.

[0102]

[0109] Step 804: in the case where it is determined that the target video packet is lost, the target video is recovered according to the target video frame and the target redundancy packet, and a recovered target video is obtained.

[0103]

[0110] In an optional embodiment of the present disclosure, before determining the target video packet loss and recovering the target video according to the target video frame and the target redundant packet to obtain the recovered target video, it can be determined whether the target video has packet loss. Specifically, network diagnostic tools such as Ping and Traceroute can be used for packet loss diagnosis: the ping command is used to continuously send data packets to key nodes (such as server IP) in the video streaming path, and it is observed whether there is a "Request timed out" prompt, which indicates that there is packet loss. Alternatively, the continuity and integrity of the data packet can be detected by analyzing the timestamp, sequence number and other information in the video stream, and then it can be determined whether there is packet loss.

[0104]

[0111] By applying the scheme of the embodiments of the present disclosure, since the target redundant packet is generated for each frame by determining the single-frame redundancy and adjusting the initial redundancy based on the reception state information, the video packet recovery can be started at each frame, without waiting for the data of the next frame or several frames, thereby effectively reducing the video transmission lag.

[0105]

[0112] In an optional embodiment of the present disclosure, the above-mentioned recovering the target video according to the target video frame and the target redundant packet to obtain the recovered target video in the case of determining the target video packet loss can include the following steps: receiving a video frame set of the target video and a set redundant packet of the video frame set; in the case of determining the target video packet loss and failing to recover the target video according to the target video frame and the target redundant packet, recovering the target video according to the video frame set and the set redundant packet to obtain the recovered target video.

[0106]

[0113] Specifically, the set redundant packet can be obtained based on the set redundancy of the video frame set, the set redundancy can be obtained by adjusting the initial set redundancy based on the reception state information of the receiving end, and the initial set redundancy can be obtained based on the set transmission information of the video frame set.

[0107]

[0114] It should be noted that, in the case of determining the target video packet loss, the target video can be directly recovered according to the target video frame and the target redundant packet, further, in the case of successfully recovering the target video according to the target video frame and the target redundant packet, the recovered target video is obtained; in the case of failing to recover the target video according to the target video frame and the target redundant packet, the target video can be recovered according to the video frame set and the set redundant packet to obtain the recovered target video.

[0108]

[0115] By applying the scheme of the embodiment of the present disclosure, by means of the mode of single-frame redundancy packet as the main and group-frame redundancy packet as the auxiliary, it can not only ensure that the receiving end can start the recovery of the video packet in each frame, but also the group-frame redundancy degree can be low, and in the case that the single-frame redundancy packet cannot recover the video packet, the group-frame redundancy packet is used to continue the recovery, so as to ensure that the video packet loss problem is effectively solved.

[0109]

[0116] Corresponding to the above-mentioned video transmission method embodiment applied to the sending end, the present disclosure also provides a video transmission device embodiment applied to the sending end. FIG. 9 shows a structure schematic diagram of a video transmission device provided by one embodiment of the present disclosure. As shown in FIG. 9, the device applied to the sending end includes: an acquisition module 902 configured to acquire transmission information of a target video frame in response to a video transmission request for a target video, wherein the target video frame is any video frame of the target video; a determination module 904 configured to determine an initial redundancy degree of the target video frame according to the transmission information; an adjustment module 906 configured to adjust the initial redundancy degree according to receiving state information of the receiving end to obtain a target redundancy degree of the target video frame, wherein the receiving state information is used to describe the state of the receiving end when receiving the target video; and a transmission module 908 configured to generate a target redundancy packet of the target video frame according to the target redundancy degree, and transmit the target video frame and the target redundancy packet to the receiving end.

[0110]

[0117] Optionally, the device further includes: an identification module configured to identify the type of the target video frame to determine the frame type of the target video frame; and the adjustment module 906 is further configured to determine an adjustment redundancy degree of the target video frame according to the frame type and the initial redundancy degree, and adjust the adjustment redundancy degree according to the receiving state information of the receiving end to obtain the target redundancy degree of the target video frame.

[0111]

[0118] Optionally, the identification module is further configured to, in a case that the target video frame is determined to be a non-key video frame based on the frame type, determine the initial redundancy degree as the adjustment redundancy degree of the target video frame; and in a case that the target video frame is determined to be a key video frame based on the frame type, adjust the initial redundancy degree to obtain the adjustment redundancy degree of the target video frame.

[0119] Optionally, the receiving state information includes a receiving freezing rate; and the adjustment module 906 is further configured to acquire a reference freezing rate, compare the receiving freezing rate and the reference freezing rate to obtain a comparison result, and adjust the initial redundancy degree according to the comparison result to obtain the target redundancy degree of the target video frame.

[0112]

[0120] Optionally, the transmission information comprises at least one of round-trip delay information, packet loss information, resolution and code rate bandwidth information; the determining module 904 is further configured to obtain at least one transmission reference information, wherein the transmission reference information carries reference redundancy; according to the transmission information, the target transmission reference information is selected from the at least one transmission reference information, and the reference redundancy carried by the target transmission reference information is determined as the initial redundancy of the target video frame, wherein the target transmission reference information is the transmission reference information in the at least one transmission reference information which has a similarity greater than a preset threshold with the transmission information.

[0113]

[0121] Optionally, the apparatus further comprises a generating module configured to, in response to a video transmission request for the target video, determine a set redundancy of a video frame set, wherein the video frame set comprises a plurality of video frames of the target video; generate a set redundancy packet of the video frame set according to the set redundancy, and transmit the video frame set and the set redundancy packet to the receiving end.

[0114]

[0122] Optionally, the generating module is further configured to, in response to a video transmission request for the target video, obtain set transmission information of the video frame set, wherein the set transmission information is obtained based on transmission information of the plurality of video frames; determine an initial set redundancy of the video frame set according to the set transmission information, wherein the initial set redundancy is the initial redundancy of the plurality of video frames; and adjust the initial set redundancy according to the reception state information of the receiving end to obtain the set redundancy of the video frame set, wherein the set redundancy is the target redundancy of the plurality of video frames.

[0115]

[0123] By applying the scheme of the embodiment of the present disclosure, the single-frame redundancy is determined, and the initial redundancy is adjusted based on the reception state information to generate a redundancy packet for each frame, so that the receiving end can start the recovery of the video packet at each frame, without waiting for the data of the next frame or several frames, thereby effectively reducing the video transmission lag.

[0116]

[0124] The present embodiment is a schematic scheme of a video transmission apparatus applied to a sending end. It should be noted that the technical scheme of the video transmission apparatus applied to the sending end belongs to the same concept as the technical scheme of the video transmission method applied to the sending end described above, and the details not described in the technical scheme of the video transmission apparatus applied to the sending end can be referred to the description of the technical scheme of the video transmission method applied to the sending end.

[0117]

[0125] Corresponding to the above-described video transmission method embodiment applied to the receiving end, this disclosure also provides an embodiment of a video transmission device applied to the receiving end. FIG10 shows a schematic diagram of another video transmission device provided in an embodiment of this disclosure. As shown in FIG10, the device is applied to the receiving end and includes: a receiving module 1002, configured to receive a target video frame and a target redundancy packet of the target video frame, wherein the target redundancy packet is obtained based on the target redundancy of the target video frame, the target redundancy is obtained by adjusting the initial redundancy of the target video frame based on the receiving state information of the receiving end, the initial redundancy is obtained based on the transmission information of the target video frame, and the receiving state information is used to describe the state of the receiving end when receiving the target video; a recovery module 1004, configured to recover the target video according to the target video frame and the target redundancy packet when it is determined that the target video has lost packets, and obtain the recovered target video.

[0118]

[0126] Optionally, the recovery module 1004 is further configured to receive a set of video frames of the target video and a set of redundant packets of the video frame set; if it is determined that the target video is lost and the recovery of the target video based on the target video frames and the target redundant packets fails, the recovery of the target video is performed based on the set of video frames and the set of redundant packets to obtain the recovered target video.

[0119]

[0127] By applying the scheme of the present disclosure, since the target redundancy packet is generated for each frame by determining the redundancy of a single frame and adjusting the initial redundancy based on the receiving state information, the recovery of the video packet can start in each frame without waiting for the data of the next frame or several frames, thereby effectively reducing video transmission stuttering.

[0120]

[0128] The above is an illustrative scheme of a video transmission device applied to a receiving end according to this embodiment. It should be noted that the technical solution of the video transmission device applied to the receiving end and the technical solution of the video transmission method applied to the receiving end described above belong to the same concept. For details not described in detail in the technical solution of the video transmission device applied to the receiving end, please refer to the description of the technical solution of the video transmission method applied to the receiving end described above.

[0121]

[0129] FIG11 shows a structural block diagram of a computing device provided in an embodiment of the present disclosure. The components of the computing device 1100 include, but are not limited to, a memory 1110 and a processor 1120. The processor 1120 is connected to the memory 1110 via a bus 1130, and a database 1150 is used to store data.

[0122]

[0130] The computing device 1100 also includes an access device 1140 that enables the computing device 1100 to communicate via one or more networks 1160. Examples of such networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or an intranet, such as a company's internal network, or the Internet. The access device 1140 can include one or more of any type of network interface (for example, a network interface card (NIC)) such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a Global System for Mobile (GSM) wireless interface, a Bluetooth interface, a near field communication (NFC) interface, a cellular network interface, an Ethernet interface, a universal serial bus (USB) interface, a Wi-Fi® interface, a Wi-MAX® interface, or the like.

[0123]

[0131] In one embodiment of the disclosure, the above-described components of the computing device 1100, as well as other components not shown in FIG. 11, can be connected to each other by a bus, for example. It should be understood that the structure block diagram of the computing device shown in FIG. 11 is merely for the purpose of example, and is not a limitation on the scope of the disclosure. Those skilled in the art can add or replace other components as needed.

[0124]

[0132] The computing device 1100 can be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (for example, a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook, and the like), a mobile phone (for example, a smartphone), a wearable computing device (for example, a smartwatch, smartglasses, and the like), or other type of mobile device, or a stationary computing device such as a desktop computer or a personal computer (PC). The computing device 1100 can also be a mobile or stationary server.

[0125]

[0133] The processor 1120 is configured to execute the computer program / instructions, and realize the steps of the video transmission method when the computer program / instructions are executed by the processor.

[0126]

[0134] The above describes a schematic solution of the computing device of the embodiment. It should be noted that the technical solution of the computing device and the technical solution of the video transmission method belong to the same concept, and the details of the technical solution of the computing device not described in detail can be referred to the description of the technical solution of the video transmission method.

[0127]

[0135] The embodiment of the present disclosure further provides a computer readable storage medium, which stores computer program / instructions, and the computer program / instructions realize the steps of the video transmission method when executed by the processor.

[0128]

[0136] The above describes a schematic solution of the computer readable storage medium of the embodiment. It should be noted that the technical solution of the storage medium and the technical solution of the video transmission method belong to the same concept, and the details of the technical solution of the storage medium not described in detail can be referred to the description of the technical solution of the video transmission method.

[0129]

[0137] The embodiment of the present disclosure further provides a computer program product, which includes computer program / instructions, and the computer program / instructions realize the steps of the video transmission method when executed by the processor.

[0130]

[0138] The above describes a schematic solution of the computer program product of the embodiment. It should be noted that the technical solution of the computer program product and the technical solution of the video transmission method belong to the same concept, and the details of the technical solution of the computer program product not described in detail can be referred to the description of the technical solution of the video transmission method.

[0131]

[0139] The above describes specific embodiments of the present disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order and still achieve desirable results. Also, the processes depicted in the figures do not necessarily require the particular order shown or sequential order to achieve desirable results. In certain implementations, multitasking and parallel processing can be advantageous.

[0132]

[0140] The computer readable medium can include any entity or apparatus capable of carrying the computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal, software distribution medium, etc. It should be noted that the computer readable medium can include appropriate additions or subtractions according to the requirements of patent practice. For example, according to the patent practice in some regions, the computer readable medium does not include electric carrier signals and telecommunication signals.

[0133]

[0141] It should be noted that, for the foregoing method embodiments, in order to facilitate description, each is described as a combination of a series of actions, but those skilled in the art should know that the disclosed embodiments are not limited to the order of the actions described, because according to the disclosed embodiments, certain steps can be performed in other orders or at the same time. In addition, those skilled in the art should know that the embodiments described in the specification all belong to preferred embodiments, and the actions and modules involved are not necessarily required by the disclosed embodiments.

[0134]

[0142] In the above embodiments, the description of each embodiment has its own focus, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments.

[0135]

[0143] The preferred embodiments of the present disclosure disclosed above are only used to help explain the present disclosure. The alternative embodiments do not describe all the details and limit the invention to the specific embodiments described. Obviously, according to the content of the disclosed embodiments, many modifications and changes can be made. The present disclosure selects and specifically describes these embodiments in order to better explain the principles and practical applications of the disclosed embodiments, so that those skilled in the art can well understand and use the present disclosure. The present disclosure is limited only by the claims and their full scope and equivalents.

Claims

CLAIM 1. A video transmission method, applied to a sending end, the method comprising: In response to a video transmission request for a target video, transmission information of a target video frame is acquired, wherein the target video frame is any video frame of the target video; an initial redundancy of the target video frame is determined according to the transmission information; the initial redundancy is adjusted according to reception state information of a receiving end to obtain a target redundancy of the target video frame, wherein the reception state information is used to describe a state of the receiving end when the target video is received; a target redundancy packet of the target video frame is generated according to the target redundancy, and the target video frame and the target redundancy packet are transmitted to the receiving end.

2. The method of claim 1, wherein after determining the initial redundancy of the target video frame based on the transmission information, the method further comprises: The target video frame is subjected to type identification to determine a frame type of the target video frame. The initial redundancy is adjusted according to the reception state information of the receiving end to obtain the target redundancy of the target video frame, including: an adjustment redundancy of the target video frame is determined according to the frame type and the initial redundancy. The adjustment redundancy is adjusted according to the reception state information of the receiving end to obtain the target redundancy of the target video frame.

3. The method of claim 2, wherein determining the adjusted redundancy of the target video frame according to the frame type and the initial redundancy comprises: In a case where the target video frame is determined to be a non-key video frame based on the frame type, the initial redundancy is determined as the adjustment redundancy of the target video frame. In a case where the target video frame is determined to be a key video frame based on the frame type, the initial redundancy is adjusted to obtain the adjustment redundancy of the target video frame.

4. The method of claim 1, wherein the receiving the reception state information comprises receiving a stall rate; and wherein the adjusting the initial redundancy level to obtain the target redundancy level of the target video frame based on the reception state information of the receiving end comprises: A reference stall rate is acquired; a comparison result is obtained by comparing the reception stall rate and the reference stall rate; The initial redundancy is adjusted according to the comparison result to obtain the target redundancy of the target video frame.

5. The method of claim 1, wherein the transmission information comprises at least one of round trip time information, packet loss information, resolution and bitrate bandwidth information; and wherein determining the initial redundancy of the target video frame based on the transmission information comprises: At least one transmission reference information is acquired, wherein the transmission reference information carries a reference redundancy; target transmission reference information is screened out from the at least one transmission reference information according to the transmission information, and a reference redundancy carried by the target transmission reference information is determined as the initial redundancy of the target video frame, wherein the target transmission reference information is transmission reference information in the at least one transmission reference information that has a similarity greater than a preset threshold with the transmission information. ​ 6. The method of claim 1, further comprising: In response to a video transmission request for a target video, a set redundancy of a video frame set is determined, wherein the video frame set includes a plurality of video frames of the target video; a set redundancy packet of the video frame set is generated according to the set redundancy, and the video frame set and the set redundancy packet are transmitted to the receiving end.

7. The method of claim 6, wherein the determining the collective redundancy of the set of video frame sets in response to the video transmission request for the target video comprises: In response to a video transmission request for a target video, collection transmission information of a video frame set is obtained, wherein the collection transmission information is obtained based on transmission information of the plurality of video frames; an initial collection redundancy of the video frame set is determined according to the collection transmission information, wherein the initial collection redundancy is an initial redundancy of the plurality of video frames; and the initial collection redundancy is adjusted according to reception state information of the receiving end to obtain a collection redundancy of the video frame set, wherein the collection redundancy is a target redundancy of the plurality of video frames.

8. A video transmission method, applied to a receiving end, the method comprising: The target video frame and a target redundancy packet of the target video frame are received, wherein the target redundancy packet is obtained based on a target redundancy of the target video frame, the target redundancy is obtained by adjusting an initial redundancy of the target video frame based on reception state information of the receiving end, the initial redundancy is obtained based on transmission information of the target video frame, and the reception state information is used to describe a state of the receiving end when the target video is received; and in a case where it is determined that the target video is lost, the target video is recovered based on the target video frame and the target redundancy packet to obtain a recovered target video.

9. The method of claim 8, wherein in the case where it is determined that the target video is lost, the target video is recovered based on the target video frame and the target redundancy packet to obtain a recovered target video, including: receiving a video frame set of the target video and a collection redundancy packet of the video frame set; and in a case where it is determined that the target video is lost and the target video fails to be recovered based on the target video frame and the target redundancy packet, the target video is recovered based on the video frame set and the collection redundancy packet to obtain a recovered target video.

10. A video transmission system, comprising a cloud-side device and an end-side device; the cloud-side device is configured to, in response to a video transmission request for a target video, acquire transmission information of a target video frame, wherein, The target video frame is any video frame of the target video; and the initial redundancy of the target video frame is determined based on the transmission information. The initial redundancy is adjusted based on reception state information of the terminal-side device to obtain a target redundancy of the target video frame, wherein the reception state information is used to describe a state of the terminal-side device when the target video is received; the target redundancy packet of the target video frame is generated based on the target redundancy, and the target video frame and the target redundancy packet are transmitted to the terminal-side device; and the terminal-side device is configured to recover the target video based on the target video frame and the target redundancy packet to obtain a recovered target video in a case where it is determined that the target video is lost.

11. A computing device comprising: A memory and a processor; The memory is configured to store computer programs / instructions, and the processor is configured to execute the computer programs / instructions, so as to implement the steps of the method of any one of claims 1 to 7 or any one of claims 8 to 9. ​ 12. A computer readable storage medium, storing computer programs / instructions, which, when executed by a processor, implement the steps of the method according to any one of claims 1 to 7 or any one of claims 8 to 9.

13. A computer program product, comprising computer programs / instructions, which, when executed by a processor, implement the steps of the method according to any one of claims 1 to 7 or any one of claims 8 to 9.

Citation Information

Patent Citations

  • Data transmission method and system

    CN113810769A

  • Video transmission method and device based on forward error correction and computer storage medium

    CN114900698A

  • Video stream code rate adjustment method and apparatus, computer device, and storage medium

    WO2024051426A1