Video data transmission method and device, electronic equipment, storage medium and program product
By generating video data packets carrying extended information at the push end, and selecting the target video stream at the pull end based on the extended information, the problem of poor playback effect in weak uplink network scenarios caused by the push end's inability to perceive the multi-stream push situation is solved, achieving high smoothness and stability of video communication and improving user experience.
Patent Information
- Application Number
- CN202511172043.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-20
- Publication Date
- 2025-11-21
AI Technical Summary
In real-time communication, the streaming end cannot effectively perceive the multi-stream push situation, resulting in poor playback effect of the pull-down streaming end in the case of weak uplink network, with stuttering, delay and unnecessary stream degradation, and frequent stream switching, which affects the user experience.
The push end generates video data packets carrying extended information. The video data packets include the status attributes of candidate video streams at multiple spatial levels. The pull end selects the target video stream based on the extended information, realizing dynamic optimization and adaptive adjustment.
It improves the smoothness and stability of video communication, reduces stuttering and packet loss rates, enhances user experience, and ensures high-quality video playback in different network environments.
Smart Images

Figure CN121000902A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of video processing technology, and more specifically, to a video data transmission method, apparatus, electronic device, storage medium, and program product. Background Technology
[0002] With the development of video processing technology, real-time communication (RTC) technology has emerged. RTC has been widely used in scenarios such as video conferencing, online education, and telemedicine. To adapt to different network environments and device performance, multi-stream transmission (simulcast) technology can be combined during real-time communication. Simulcast technology can simultaneously generate and transmit multiple video streams with different resolutions and bitrates, i.e., video streams at different spatial levels. The streaming end selects the appropriate video stream for playback based on network conditions and device capabilities.
[0003] In realizing the concept disclosed herein, the inventors discovered at least the following problems in the related technology: since the pull-streaming end is unaware of the specific situation of multi-stream push from the push-streaming end, it is impossible to effectively guarantee the playback effect of the pull-streaming end when the push-streaming end encounters a weak uplink network scenario. Summary of the Invention
[0004] In view of the above, this disclosure provides a video data transmission method, apparatus, electronic device, storage medium, and program product.
[0005] According to one aspect of this disclosure, a video data transmission method is provided, applied to a streaming end, the method comprising: generating a video data packet based on a video stream to be transmitted, wherein the video data packet includes video data of each of a plurality of candidate video streams at a spatial level corresponding to the video stream to be transmitted, the video data carrying extended information for characterizing the state attributes of the candidate video streams at the spatial level; and sending the video data packet to a streaming end, so that the streaming end selects a target video stream from the plurality of candidate video streams at the spatial levels based on the plurality of extended information in the video data packet.
[0006] According to another aspect of this disclosure, a video data transmission method is provided, applied to a streaming end, the method comprising: in response to receiving a video data packet from a streaming end, parsing the video data packet to obtain video data for each of a plurality of candidate video streams at multiple spatial levels, wherein the video data carries extended information for characterizing attributes of the candidate video streams at the spatial levels; and selecting a target video stream from the plurality of candidate video streams at the multiple spatial levels according to a selection strategy determined based on the plurality of extended information.
[0007] According to another aspect of this disclosure, a video data transmission apparatus is provided, applied to a streaming end. The method includes: a generation module, configured to generate a video data packet based on a video stream to be transmitted, wherein the video data packet includes video data of each of a plurality of candidate video streams at multiple spatial levels corresponding to the video stream to be transmitted, and the video data carries extended information for characterizing the state attributes of the candidate video streams at the spatial levels; and a transmission module, configured to send the video data packet to a streaming end, so that the streaming end selects a target video stream from the plurality of candidate video streams at the multiple spatial levels based on the multiple extended information in the video data packet.
[0008] According to another aspect of this disclosure, a video data transmission apparatus is provided, applied at a streaming end, the method comprising: a parsing module, configured to parse the video data packet received from a streaming end to obtain video data for each of a plurality of candidate video streams at multiple spatial levels, wherein the video data carries extended information for characterizing the attributes of the candidate video streams at the spatial levels; and a selection module, configured to select a target video stream from the plurality of candidate video streams at the multiple spatial levels according to a selection strategy determined based on the plurality of extended information.
[0009] According to another aspect of this disclosure, an electronic device is provided, comprising: one or more processors; and a memory for storing one or more instructions, wherein, when executed by the one or more processors, the one or more processors cause the one or more processors to perform the method as described in this disclosure.
[0010] According to another aspect of this disclosure, a computer-readable storage medium is provided having executable instructions stored thereon, which, when executed by a processor, cause the processor to perform the methods described in this disclosure.
[0011] According to another aspect of this disclosure, a computer program product is provided, which includes computer-executable instructions that, when executed, are used to perform the methods described in this disclosure. Attached Figure Description
[0012] The above and other objects, features and advantages of this disclosure will become clearer from the following description of embodiments with reference to the accompanying drawings, in which:
[0013] Figure 1 This illustration schematically shows a system architecture to which a video data transmission method can be applied according to embodiments of the present disclosure;
[0014] Figure 2 A flowchart illustrating a video data transmission method according to an embodiment of the present disclosure is shown schematically.
[0015] Figure 3A An example schematic diagram of a video data packet according to an embodiment of the present disclosure is shown;
[0016] Figure 3B An example schematic diagram illustrating extended information according to embodiments of the present disclosure is shown;
[0017] Figure 4 A flowchart illustrating a video data transmission method according to an embodiment of the present disclosure is shown schematically.
[0018] Figure 5 This illustration schematically shows an example diagram of a target video stream determination process according to an embodiment of the present disclosure;
[0019] Figure 6 A block diagram of a video data transmission apparatus according to an embodiment of the present disclosure is shown schematically;
[0020] Figure 7 A block diagram schematically illustrates a video data transmission apparatus according to embodiments of the present disclosure; and
[0021] Figure 8 A block diagram of an electronic device suitable for implementing a video data transmission method according to an embodiment of the present disclosure is shown schematically. Detailed Implementation
[0022] The embodiments of the present disclosure will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the disclosure. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of the present disclosure for ease of explanation. However, it will be apparent that one or more embodiments may be practiced without these specific details. Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concepts of the present disclosure.
[0023] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this disclosure. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0024] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.
[0025] When using expressions such as "at least one of A, B and C", they should generally be interpreted in accordance with the meaning that is commonly understood by those skilled in the art (e.g., "a system having at least one of A, B and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B and C, etc.).
[0026] In the technical solution of this invention, the user information (including but not limited to user personal information, user image information, user device information, such as location information) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of related data all comply with relevant laws, regulations, and standards, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entry points for users to choose to authorize or refuse.
[0027] In scenarios involving automated decision-making using personal information, the methods, devices, and systems provided in this invention offer users corresponding entry points for choosing to agree to or reject the automated decision-making results. If the user chooses to reject, the process proceeds to the expert decision-making stage. Here, "automated decision-making" refers to the activity of automatically analyzing and evaluating an individual's behavioral habits, interests, or economic, health, and credit status through computer programs, and then making a decision. Here, "expert decision-making" refers to the activity of making decisions by personnel who specialize in a particular field, possess specialized experience, knowledge, and skills, and have reached a certain level of professional expertise.
[0028] In practical applications, multi-stream transmission technology typically generates spatial layer video streams of three resolutions: large, medium, and small (e.g., 1080p, 720p, and 360p). The server dynamically switches between different quality video streams based on network conditions and receiver feedback to ensure smooth video calls and maintain image quality.
[0029] In RTC application scenarios of multi-stream transmission technology, the minimum bitrate, target bitrate and maximum bitrate of the three types of video streams with large, medium and small resolutions are usually configured by profiling. The streaming end will allocate bandwidth according to the results obtained by the bandwidth estimation module. If the overall bandwidth is insufficient to support the push of certain data streams, the encoding and transmission of these data streams will be stopped, and only other data streams with lower resolution and lower bandwidth requirements will be sent.
[0030] For example, when an RTC streaming client using multi-stream technology configures minimum bandwidths of 1200Kbps, 700kbps, and 300kbps for the large, medium, and small streams, respectively, and streaming client A requests to pull the large stream while streaming client B requests to pull the small stream, the streaming client's bandwidth estimation module only estimates 1100Kbps of network bandwidth based on the current network conditions. This is insufficient to support the expected 1500Kbps bandwidth (i.e., 1200Kbps for the large stream + 300Kbps for the small stream) required by streaming clients A and B. To keep the total output bandwidth within the estimated bandwidth, the streaming client will abandon pushing the combination of the large and small streams and instead push the combination of the medium and small streams. In this situation, for client A requesting the large stream, the actual stream pulled is the medium stream. This is a case of video stream degradation caused by a weak uplink network.
[0031] The strategy for video stream degradation caused by weak uplink networks is usually controlled by the RTC media server. However, the RTC media server in the relevant technology is unaware of the specific situation of multi-stream push by the push end. Specifically: (1) The server cannot distinguish the real-time push status of the data stream, and can only sense and judge whether a certain source stream has stopped pushing due to insufficient uplink bandwidth through timeout, even though such timeout may only be caused by sudden network fluctuations rather than source stream stoppage; (2) There is no standard and universally used protocol that allows the server to sense the actual allocated bandwidth of the source stream, the trend of allocated bandwidth changes, and the stability of the source stream. As a result, the push end will have the following problems when encountering weak uplink network scenarios:
[0032] On the one hand, stream degradation causes stuttering and latency on the playback end. This is because the RTC media server cannot distinguish the real-time push status of the data stream, and can only detect and determine whether a source stream has stopped pushing due to insufficient uplink bandwidth through timeout. In other words, stream degradation is triggered by timeout events. During the continuous detection of the timeout event, no data packets are sent to the stream, and the streaming end naturally does not have enough data packets to decode, render, and play, which causes playback stuttering on the streaming end.
[0033] The stream degradation mechanism of multi-stream transmission technology in related technologies not only relies on timeout event driving, but also needs to request key frames from the streaming end to complete the smooth switching of different levels of video streams. In fact, the time required to complete one stream degradation is the timeout perception period + uplink RTT (Round-Trip Time). The cost of such stream degradation often easily leads to playback stuttering at the streaming end.
[0034] On the other hand, the result of stream degradation / upgrade may not be the optimal choice. That is, since the stream degradation mechanism on the RTC media server is triggered by a timeout event, certain sudden network jitters (such as base station switching, 4G signal switching to WIFI signal, etc.) may cause the server to misjudge and cause unnecessary stream degradation, even though the original video stream has not actually stopped being pushed.
[0035] For data streams with different resolutions and different encoding parameters, the final quality result is not necessarily positively correlated with the resolution. For example, the actual viewing quality of a low bitrate medium stream (720p) at a fixed frame rate may not be as high as that of a low bitrate stream (360p) at a high bitrate. Although RTC media servers can calculate the received bitrate of video streams at different resolutions, they do not directly know the actual bandwidth allocation of different video streams during encoding at the streaming end. Furthermore, the received bitrate calculated by the server will still be affected by network fluctuations in weak network scenarios, which reduces the reliability of the bitrate used as a reference for stream downgrading / upgrading. This may lead to situations where, in some weak uplink network scenarios, a scenario that should ideally be downgraded to a high bitrate low stream is still using a low bitrate medium stream.
[0036] On the other hand, in certain weak network scenarios, frequent stream switching and noticeable stuttering may occur. Specifically, in some bandwidth-limited scenarios, the streaming of high-level video streams may become unstable due to fluctuations in the bandwidth estimation module. Because the overall estimated bandwidth intermittently meets the minimum bandwidth boundary required by the high-level video stream, although the high-level video stream can occasionally be streamed, there are intermittent interruptions. Specifically, when the estimated bandwidth just meets the push requirements of a certain level of stream, the streaming client may attempt to push that stream. However, due to unstable network conditions, the push quality of this stream is often low, and it will quickly degrade to a lower-level stream due to bandwidth reduction. This frequent switching will cause stuttering and quality fluctuations at the receiving end.
[0037] For example, when the minimum bandwidth configured for the large, medium, and small streams by the RTC push end using Simulcast technology is 1200Kbps, 700kbps, and 300kbps respectively, and the uplink bandwidth of a certain network link fluctuates between 1000Kbps and 1300kbps, it is very likely to cause intermittent interruptions in the large stream push and switching back and forth between the large and medium streams. Such frequent switching may also cause excessively frequent keyframe requests. In extreme cases, a certain level of video stream on the push end may be interrupted and re-pushed several times within the timeout perception period. In this case, even if the video stream of that level is basically unavailable (very few data packets), the RTC media server will still consider the data stream available instead of degrading to a relatively stable lower-level stream, thus causing playback stuttering on the pull end.
[0038] In summary, the aforementioned problems are particularly evident in weak uplink network scenarios for RTC applications, severely impacting the user experience during real-time communication. Therefore, effectively addressing the issue of ensuring playback quality on the streaming end when encountering weak uplink network scenarios at the streaming end is an urgent problem to be solved.
[0039] Therefore, this disclosure provides a video data transmission method, apparatus, electronic device, storage medium, and program product, which can be applied to the field of video processing technology. The video data transmission method includes: generating a video data packet based on a video stream to be transmitted, wherein the video data packet includes video data of candidate video streams at multiple spatial levels corresponding to the video stream to be transmitted, and the video data carries extended information for characterizing the state attributes of the candidate video streams at the spatial levels; and sending the video data packet to a streaming end, so that the streaming end selects a target video stream from the candidate video streams at multiple spatial levels based on the extended information in the video data packet.
[0040] Figure 1 The illustration schematically depicts a system architecture to which a video data transmission method can be applied according to embodiments of the present disclosure. It should be noted that... Figure 1 The examples shown are merely examples of system architectures that can be applied to the embodiments of this disclosure, in order to help those skilled in the art understand the technical content of this disclosure, but do not mean that the embodiments of this disclosure cannot be used in other devices, systems, environments or scenarios.
[0041] like Figure 1 As shown, the system architecture 100 according to this embodiment may include a push stream end 110, a pull stream end 120, and a network 130. The network 130 is used as a medium to provide communication links between different devices.
[0042] In one embodiment, the push streaming end 110 can generate video data packets based on the video stream to be transmitted. The video data packets include video data from multiple candidate video streams at various spatial levels corresponding to the video stream to be transmitted. The video data carries extended information representing the state attributes of the candidate video streams at the spatial levels. After generating the video data packets, the push streaming end 110 can send the video data packets to the pull streaming end 120 via the network 130.
[0043] Upon receiving video data packets from the push end, the pull end 120 can parse the video data packets to obtain video data for each of the candidate video streams at multiple spatial levels. The video data carries extended information representing the attributes of the candidate video streams at each spatial level. Based on this, the pull end 120 selects the target video stream from the candidate video streams at multiple spatial levels according to a selection strategy determined based on the extended information.
[0044] It should be understood that Figure 1The number of push streamers, pull streamers, and networks shown is merely illustrative. Depending on implementation requirements, any number of push streamers, pull streamers, and networks can be included.
[0045] It should be noted that the sequence numbers of the operations in the following methods are for descriptive purposes only and should not be considered as indicating the execution order of the operations. Unless explicitly stated otherwise, the method does not need to be executed in the exact order shown.
[0046] Figure 2 A flowchart illustrating a video data transmission method according to an embodiment of the present disclosure is shown schematically.
[0047] like Figure 2 As shown, the video data transmission method 200 includes operations S210 to S220.
[0048] In operation S210, a video data packet is generated based on the video stream to be transmitted. The video data packet includes video data of each of the candidate video streams at multiple spatial levels corresponding to the video stream to be transmitted. The video data carries extended information on the state attributes of the candidate video streams for characterizing the spatial levels.
[0049] In operation S220, video data packets are sent to the streaming end so that the streaming end can select the target video stream from candidate video streams at multiple spatial levels based on multiple extended information in the video data packets.
[0050] The video data transmission method 200 disclosed herein is applied to real-time communication scenarios, where real-time communication refers to the process of transmitting video data from a push streaming end to a pull streaming end via a network. The push streaming end refers to a device or system that acquires video data, encodes and packages it, and sends it over the network. The pull streaming end refers to a device or system that receives video data packets from the network, decodes them, and plays them. For example, in a video conference, video data acquired by the camera and microphone of a participant at the push streaming end can be transmitted in real-time over the network to other participants at the pull streaming end.
[0051] A video stream to be transmitted is a continuous data stream composed of video data, which includes both audio and video information. The streaming end processes the video stream to obtain video data packets, enabling it to be sent in a format suitable for network transmission. The processing includes steps such as encoding and encapsulation, generating multiple candidate video streams from the video stream to be transmitted and adding necessary header and extension information. A video data packet refers to a data unit with a specific format formed after the video stream has been segmented and encapsulated, facilitating transmission over the network.
[0052] In multi-stream transmission, candidate video streams refer to multiple video streams at different spatial levels generated by the streaming end, corresponding to the video stream to be transmitted. A spatial level refers to the quality grade of the video stream; each spatial level can correspond to different resolutions and bitrates. It's important to note that in real-time communication scenarios, the temporal layer refers to adjusting the frame rate during video encoding to generate video layers of different quality levels, achieving frame rate-based adaptive bitrate streaming. Conversely, in real-time communication scenarios, the spatial layer refers to adjusting the resolution during video encoding to generate video layers of different quality levels, achieving resolution-based multi-resolution adaptive streaming.
[0053] For each candidate video stream at the spatial level, the extended information can carry state attributes that characterize the candidate video stream at the spatial level. This extended information supplements the video data and can help the streaming end to better perform decoding, playback, and quality control.
[0054] After obtaining the video data packets, the video data packets can be sent to the streaming client. Upon receiving the video data packets, the streaming client can select the most suitable target video stream from multiple candidate video streams based on its own network conditions, device performance, and other factors.
[0055] The specific method of transmission can be configured according to actual business needs and is not limited here. For example, video data packets can be transmitted based on the RTP protocol; alternatively, they can be transmitted based on WebRTC; alternatively, they can be transmitted based on Object Real-Time Communications (ORTC); alternatively, they can be transmitted based on Secure Real-time Transport Protocol (SRTP); alternatively, they can be transmitted based on Datagram Transport Layer Security (DTLS).
[0056] According to embodiments of this disclosure, by generating video data packets carrying extended information at the push end and sending them to the pull end for it to select a suitable target video stream, dynamic optimization and adaptive adjustment of video transmission quality are achieved under different network environments and device performance conditions. This can effectively improve the smoothness and stability of video communication, reduce stuttering and packet loss rates, and at the same time take into account the balance of video quality, meeting the high requirements of users for video experience in real-time communication scenarios.
[0057] The following is for reference. Figure 3A and Figure 3BThe video data transmission method 200 according to an embodiment of the present invention will be further described.
[0058] According to embodiments of this disclosure, generating video data packets based on the video stream to be transmitted may include the following operations: determining multiple configuration combinations and extended information for each configuration combination based on multiple candidate resolutions and multiple candidate bitrates for the streaming end, wherein each configuration combination includes a candidate resolution and a candidate bitrate, and the configuration combination corresponds one-to-one with the spatial hierarchy; and generating video data of the candidate video streams at the spatial hierarchy based on the configuration combinations and extended information.
[0059] Candidate resolutions refer to the different resolution options that the streaming end can generate for the video stream. Common video resolutions include 1080p, 720p, 480p, and 360p, each corresponding to different image clarity and bitrate requirements. Candidate bitrates refer to the different bitrate options that the streaming end can set for the video stream. Bitrate determines the transmission speed and quality of video data; for example, a higher bitrate usually corresponds to higher video quality and greater bandwidth requirements. Configuration combinations refer to combining a candidate resolution and a candidate bitrate to form a specific spatial hierarchy of video stream configuration. Each configuration combination corresponds to different video quality and transmission requirements.
[0060] The method for generating configuration combinations can be configured according to actual business needs and is not limited here. For example, configuration combinations can be dynamically generated based on machine learning, that is, machine learning algorithms are used to analyze historical transmission data and network conditions to dynamically generate multiple configuration combinations, and corresponding extended information is generated for each combination. Alternatively, configuration combinations can be generated based on user-defined configurations, that is, users can customize candidate resolutions and bitrates according to their own needs, and the streaming end generates corresponding configuration combinations and extended information based on user input. For example, users can choose 1080p, 720p, and 480p as candidate resolutions, and 2Mbps, 1Mbps, and 500kbps as candidate bitrates, and the streaming end generates configuration combinations based on these selections.
[0061] Extended information refers to additional information appended to video data packets. It conveys crucial information about the spatial layer of the video stream in multi-stream transmission, aiding the receiver in selection and processing. Extended information is an optional part of real-time transport protocol (RTP) data packets, allowing the addition of extra information for specific applications or scenarios without altering the basic structure of the RTP. This extension mechanism provides flexibility, enabling the protocol to adapt to different needs, such as quality feedback, synchronization information, or custom metadata.
[0062] Extended information can be located after the fixed header of the Real-Time Transport Protocol (RTP), and its payload portion can occupy 1 byte. Extended information contains one or more attribute fields, each with its own specific identifier and length, enabling the streaming end to correctly parse this additional information. This enhances the functionality of the RTP and meets various complex real-time communication needs.
[0063] According to embodiments of this disclosure, by determining multiple configuration combinations based on multiple candidate resolutions and candidate bitrates, and generating corresponding extended information for each configuration combination, the streaming end can generate candidate video streams at different spatial levels. This effectively adapts to different network environments and device performance, enabling dynamic selection and optimized transmission of video streams. Furthermore, by using video data packets carrying extended information, the streaming end can understand the video stream status at each spatial level in real time, thereby selecting the most suitable target video stream for playback, improving the smoothness and stability of video communication, reducing stuttering and packet loss rates, and enhancing the user experience.
[0064] Figure 3A An example schematic diagram of a video data packet according to an embodiment of the present disclosure is shown.
[0065] like Figure 3A As shown in 300A, a video data packet 300 may include multiple sub-video data packets, each corresponding one-to-one with a candidate video stream. Furthermore, each sub-video data packet carries extended information, which characterizes the state attributes of the candidate video stream at the spatial hierarchy.
[0066] For example, sub-video data packet 310 corresponds to candidate video stream 311, and sub-video data packet 310 also carries extended information 312 corresponding to candidate video stream 311; sub-video data packet 320 corresponds to candidate video stream 321, and sub-video data packet 320 also carries extended information 322 corresponding to candidate video stream 321; ...; and so on, sub-video data packet 3N0 corresponds to candidate video stream 3N1, and sub-video data packet 3N0 also carries extended information 3N2 corresponding to candidate video stream 3N1. N is a positive integer.
[0067] Figure 3B An example schematic diagram illustrating extended information according to an embodiment of this disclosure is shown.
[0068] like Figure 3BAs shown in 300B, taking extended information 312 as an example, extended information 312 may include attribute fields carried by the extended header of the real-time transport protocol. The attribute fields include at least one of the following: resolution status field 301, stability status field 302, bitrate status field 303, and bandwidth status field 304. Resolution status field 301 is used to characterize whether the candidate video stream is the maximum resolution video stream that the streaming end can send; stability status field 302 is used to characterize whether the pushing status of the candidate video stream is stable; bitrate status field 303 is used to characterize the encoding bitrate level assigned to the candidate video stream by the video data packet; and bandwidth status field 304 is used to characterize the bandwidth change trend of the candidate video stream.
[0069] The resolution status field 301 can be represented as Urgent (U), occupying 1 bit. For example, if the streaming client can send video streams at 1080p, 720p, and 360p resolutions, then the resolution status field corresponding to 1080p will be marked as the maximum resolution. The stability status field 302 can be represented as Stable (S), occupying 1 bit. For example, if there is no packet loss, stuttering, or bitrate fluctuation during the video stream's delivery, then the stability status field 302 will be marked as stable. The bitrate status field 303 can be represented as Bitrate Status (BS), occupying 2 bits. For example, the bitrate status field 303 can indicate whether the current stream's bitrate is low, medium, high, or full. The bandwidth status field 304 can be represented as Bandwidth Trendline (BWTL), occupying 2 bits. For example, the bandwidth status field 304 can indicate whether the bandwidth is increasing, decreasing, fluctuating, or stable. In addition, extended information 312 may also include alternative field 305, which may be represented as a reserved bit (i.e., R).
[0070] According to embodiments of this disclosure, by carrying attribute fields including resolution status field, stability status field, bitrate status field, and bandwidth status field in the RTP extension header, comprehensive video stream status information can be provided to the streaming end. This helps the receiving end understand the real-time trends of video stream resolution, stability, bitrate, and bandwidth changes, make more intelligent choices, optimize video stream playback quality, improve the adaptability of video communication and user experience, reduce the probability of stuttering and buffering, and ensure smooth playback of video streams under different network conditions.
[0071] According to embodiments of this disclosure, the resolution status value corresponding to the resolution status field 301 is determined in the following manner: when the candidate video stream is not the video stream with the maximum resolution that the streaming end can send, the resolution status value is set to a first preset value; and when the candidate video stream is the video stream with the maximum resolution that the streaming end can send, the resolution status value is set to a second preset value.
[0072] The valid values for the resolution status value can include a first preset value and a second preset value. If the resolution status value is the first preset value, it means that the candidate video stream is not the maximum resolution video stream that the streaming end can send; if the resolution status value is the second preset value, it means that the candidate video stream is the maximum resolution video stream that the streaming end can send.
[0073] In one example, the first preset value can be 0, and the second preset value can be 1. For example, if the streaming client is currently pushing a large stream (1080p), a medium stream (720p), and a small stream (360p), since the large stream (1080p) is the highest resolution data stream currently pushed by the streaming client, the position of the resolution status field 301 of the extended information of the large stream is set to 1, and the position of the resolution status field 301 of the extended information of the medium stream (720p) and the small stream (360p) is set to 0.
[0074] According to embodiments of this disclosure, by explicitly identifying whether the video stream is the maximum resolution stream of the pushing end, the pulling end can quickly perceive changes in the capabilities of the pushing end and prioritize the selection of high-quality video streams. This helps to make timely decisions on stream upgrades, optimizes the video stream selection process, and improves the user experience.
[0075] According to embodiments of this disclosure, the stable state value corresponding to the stable state field 302 is determined in the following manner: when the candidate video stream does not meet the preset stability conditions, the stable state value is set to a first preset value; and when the candidate video stream meets the preset stability conditions, the stable state value is set to a second preset value, wherein the preset stability conditions are determined based on the time difference between the generation time of the video data packet and the time when the most recent sufficient bandwidth event was detected, and the stable reference duration.
[0076] The valid values of the stable state value can include a first preset value and a second preset value. If the stable state value is the first preset value, it means that the candidate video stream does not meet the preset stability condition; if the stable state value is the second preset value, it means that the candidate video stream meets the preset stability condition.
[0077] In one example, the first preset value can be 0, and the second preset value can be 1. For example, if the streaming client is currently pushing candidate video stream 1 and candidate video stream 2, if candidate video stream 1 does not meet the preset stability condition, the position of the stability state field 302 of the extended information of candidate video stream 1 can be set to 0; if candidate video stream 2 meets the preset stability condition, the position of the stability state field 302 of the extended information of candidate video stream 2 can be set to 1.
[0078] The preset stability condition is used to determine whether a video stream is stable. It is based on a comparison between the time difference between the generation time of the video data packet (T2) and the time when sufficient bandwidth was detected and the event began to be pushed (T1), and the stability reference duration (stream_stable_referncen_duration_ms). The stability reference duration is a preset time threshold used to determine the stability of the video stream.
[0079] For example, if the time difference exceeds the stable reference duration, the video stream can be considered stable. Furthermore, the stability of the video stream can also be determined based on at least one of the bitrate status value corresponding to bitrate status field 303 and the bandwidth status value corresponding to bandwidth status field 304. For example, if the bitrate status value is 0 and the bandwidth status value is not 0 or 1, the video stream can be considered stable. Conversely, if the bitrate status value is not 0, the video stream is considered temporarily unstable.
[0080] It should be noted that the stable reference duration can be flexibly configured according to the business scenario and is not limited here. For example, the configuration range of the stable reference duration can be between 1000 ms and 5000 ms. In one example, if the requirements for smoothness and real-time performance are high, the stable reference duration can be configured to be relatively large; if the requirements for clarity are high, the stable reference duration can be configured to be relatively small.
[0081] According to embodiments of this disclosure, by defining a stable state field and its value method, the streaming end can prioritize selecting video streams with high stability, thereby improving playback smoothness and user experience. This effectively reduces stuttering and playback interruptions caused by selecting unstable video streams, reduces fluctuations in user experience, and ensures the reliability and stability of video communication.
[0082] According to embodiments of this disclosure, the bitrate status value corresponding to the bitrate status field 303 is determined as follows: when the allocated coding bitrate is between the minimum bandwidth and the critical bandwidth, the bitrate status value is configured to a first preset value; when the allocated coding bitrate is between the critical bandwidth and the target bandwidth, the bitrate status value is configured to a second preset value; when the allocated coding bitrate is between the target bandwidth and the maximum bandwidth, the bitrate status value is configured to a third preset value; and when the allocated coding bitrate is the maximum bandwidth, the bitrate status value is configured to a fourth preset value.
[0083] The valid values for the bitrate status value can include a first preset value to a fourth preset value. In one example, the first preset value can be 0, the second preset value can be 1, the third preset value can be 2, and the fourth preset value can be 3.
[0084] Minimum bandwidth refers to the lowest bitrate required for a video stream to maintain basic usable quality. Critical bandwidth is the bitrate threshold at which video stream quality begins to significantly degrade, falling between the minimum and target bandwidth. Target bandwidth is the expected bitrate for the video stream under normal network conditions. Maximum bandwidth is the highest bitrate that the streaming end can provide.
[0085] If the bitrate status value is the first preset value, it means that the allocated coding bitrate is between the minimum bandwidth and the critical bandwidth, and is in a low bitrate state; if the bitrate status value is the second preset value, it means that the allocated coding bitrate is between the critical bandwidth and the target bandwidth, and is in a reduced bitrate state; if the bitrate status value is the third preset value, it means that the allocated coding bitrate is between the target bandwidth and the maximum bandwidth, and is in a full bitrate state; if the bitrate status value is the fourth preset value, it means that the allocated coding bitrate is the maximum bandwidth, and is in an over-bitrate state.
[0086] The critical bandwidth is determined based on the minimum bandwidth and the critical bitrate factor. For example, the critical bandwidth can be equal to the minimum bandwidth * (1 + critical bitrate factor). The critical bitrate factor can be determined depending on the business scenario and is not limited here. In one example, the critical bitrate factor can be between 0.05 and 0.2.
[0087] According to embodiments of this disclosure, by defining a bitrate status field and its value method, the bitrate status is divided into four levels: critically low, reduced, full bitrate, and over-bitrate. This allows the streaming end to select the most suitable video stream based on the current network conditions and device performance, thereby optimizing playback quality and user experience. It can effectively improve the adaptability and stability of video communication, ensure smooth playback of video streams in different network environments, and maximize the use of available bandwidth to provide the best video quality.
[0088] According to embodiments of this disclosure, the bandwidth status value corresponding to the bandwidth status field 304 is determined by: smoothing the acquired historical bandwidth allocation data to obtain smoothed bandwidth allocation data; performing linear regression analysis on the smoothed bandwidth allocation data to obtain discrete information characterizing the bandwidth change trend, wherein the discrete information includes at least one of slope, intercept, and coefficient of determination; and determining the bandwidth status value based on the discrete information characterizing the bandwidth change trend.
[0089] Historical bandwidth allocation data refers to the bandwidth allocation of a video stream recorded by the streaming end within a certain time range, which can include timestamps and corresponding bandwidth values. In one example, historical bandwidth allocation data is collected periodically and stored in a circular buffer.
[0090] Smoothing refers to processing the raw historical bandwidth allocation data to reduce the impact of noise and fluctuations, making the data smoother. The specific methods of smoothing can be configured according to actual business needs and are not limited here. For example, smoothing methods may include at least one of the following: exponential moving average, simple moving average, weighted moving average, and Gaussian filtering.
[0091] In one example, the original historical bandwidth allocation data can be smoothed using an exponential moving average, as shown in the following formula (1).
[0092] (1);
[0093] in, The exponential moving average at the current moment, Characterizing smoothness factor, The raw data representing the current moment. It represents the exponential moving average of the previous time step.
[0094] After obtaining the smoothed bandwidth allocation data, linear regression analysis can be performed on it to obtain discrete information characterizing the bandwidth change trend. Linear regression analysis is used to model the linear relationship between independent and dependent variables, and outputs parameters including slope, intercept, and coefficient of determination. The specific method of linear regression analysis can be configured according to actual business needs and is not limited here. For example, the specific method of linear regression analysis can include at least one of the following: least squares method, gradient descent method, ridge regression, and principal component regression.
[0095] Discrete information refers to key features or parameters extracted from continuous data to characterize the trend of bandwidth change. In one example, discrete information includes at least one of slope, intercept, and coefficient of determination. Slope refers to the rate at which bandwidth changes over time; for example, a positive slope indicates an increase in bandwidth, and a negative slope indicates a decrease. Intercept refers to the predicted bandwidth value when time is zero in a linear regression model. Coefficient of determination indicates how well the model fits the data; for example, the closer the coefficient of determination is to 1, the better the model fit.
[0096] In one example, the least squares method can be used to perform linear regression analysis on the smoothed bandwidth allocation data, as shown in formulas (2) to (4) below. During this process, the parameters can be updated by sampling incremental calculation.
[0097] (2);
[0098] (3);
[0099] (4);
[0100] in, Characterizing the slope, Characterizing the number of samples, Represents the sum of the cross products of x and y. The sum of x values represents the total value of x. The sum of y values. Characterized by the sum of squares of x, Characteristic of determination coefficient The sum of squares of y is represented.
[0101] After obtaining discrete information, bandwidth state values can be determined based on this information. For example, discrete information can be used as feature input to a clustering algorithm (such as K-means) to classify different bandwidth change trends into different categories, with each category corresponding to a bandwidth state value. Alternatively, decision tree algorithms can be used to classify discrete information and determine bandwidth state values based on different conditional branches. Another option is to construct a neural network model, using discrete information as input and bandwidth state values as output, and train the model to achieve automatic bandwidth state determination.
[0102] According to embodiments of this disclosure, key discrete information is extracted from historical bandwidth allocation data through smoothing and linear regression analysis, and bandwidth status values are determined accordingly. This comprehensively considers long-term trends and short-term fluctuations, accurately characterizing the bandwidth change trend of the video stream. This allows the streaming end to understand the changes in network bandwidth in advance, thereby adjusting the video stream selection strategy in a timely manner, optimizing playback quality, reducing stuttering and buffering, improving the accuracy and reliability of bandwidth status assessment, and ensuring the stability and smoothness of video communication under different network conditions.
[0103] According to embodiments of this disclosure, determining a bandwidth state value based on discrete values characterizing bandwidth change trends includes: configuring the bandwidth state value to a first preset value when the discrete information characterizes a continuously decreasing bandwidth change trend; configuring the bandwidth state value to a second preset value when the discrete information characterizes a fluctuating decreasing bandwidth change trend; configuring the bandwidth state value to a third preset value when the discrete information characterizes a generally increasing bandwidth change trend; and configuring the bandwidth state value to a fourth preset value when the discrete information characterizes a relatively stable bandwidth change trend.
[0104] The valid values for the bandwidth status value can include a first preset value to a fourth preset value. In one example, the first preset value can be 0, the second preset value can be 1, the third preset value can be 2, and the fourth preset value can be 3.
[0105] If the bandwidth status value is the first preset value, it indicates that the bandwidth change trend is continuously decreasing, that is, the slope is negative and the coefficient of determination is high; if the bandwidth status value is the second preset value, it indicates that the bandwidth change trend is fluctuating downward, that is, the slope is negative but the coefficient of determination is low; if the bandwidth status value is the third preset value, it indicates that the bandwidth change trend is generally increasing, that is, the slope is positive; if the bandwidth status value is the fourth preset value, it indicates that the bandwidth change trend is relatively stable, that is, the slope is close to 0 and the coefficient of determination is high.
[0106] According to embodiments of this disclosure, by defining bandwidth status values and their correspondence with bandwidth change trends, a clear bandwidth status indication is provided to the streaming end, enabling the streaming end to adjust the video stream selection strategy in a timely manner according to different bandwidth status values, optimize playback quality, reduce stuttering and buffering, thereby ensuring the stability and smoothness of video communication under different network conditions.
[0107] Figure 4 A flowchart illustrating a video data transmission method according to an embodiment of the present disclosure is shown schematically.
[0108] like Figure 4 As shown, the video data transmission method 400 includes operations S410~S420.
[0109] In operation S410, in response to receiving a video data packet from the streaming end, the video data packet is parsed to obtain video data for each of the candidate video streams at multiple spatial levels. The video data carries extended information about the attributes of the candidate video streams at the spatial levels.
[0110] In operation S420, the target video stream is selected from candidate video streams at multiple spatial levels according to a selection strategy determined based on multiple extended information.
[0111] The video data transmission method 400 provided in this disclosure is applied to real-time communication scenarios. Real-time communication refers to the process of transmitting video data from a push streaming end to a pull streaming end via a network. The push streaming end refers to a device or system that collects video data, encodes and packages it, and sends it over the network. The pull streaming end refers to a device or system that receives video data packets from the network, decodes them, and plays them. For example, in a video conference, video data collected by the camera and microphone of a participant at the push streaming end can be transmitted in real-time over the network to other participants at the pull streaming end.
[0112] After receiving video data packets from the push stream, the pull stream end can parse the video data packets to obtain video data for candidate video streams at multiple spatial levels. Parsing refers to the process of processing the received video data packets and extracting the valid information from them.
[0113] After obtaining extended information about the attributes of each candidate video stream at each spatial level, which characterizes the spatial level, a selection strategy can be determined based on this extended information. The selection strategy is used to choose the target video stream best suited for the streaming end from among the multiple candidate video streams. The target video stream refers to the video stream ultimately selected by the streaming end that best suits the current network conditions and device performance.
[0114] The method for determining the selection strategy can be configured according to actual business needs and is not limited here. In one example, the selection strategy can be dynamically determined based on network conditions. Specifically, the streaming client can dynamically adjust the selection strategy based on the current network conditions (such as bandwidth, latency, and packet loss rate). For example, if the network bandwidth is sufficient and stable, high-resolution, high-bitrate video streams are prioritized; if the network bandwidth is insufficient or unstable, low-resolution, low-bitrate video streams are prioritized. In another example, the selection strategy can be customized based on user preferences. Specifically, the streaming client can set the selection strategy according to the user's custom preferences. For example, the user can set to prioritize high-resolution video streams even if network bandwidth is insufficient; or set to prioritize low-latency video streams even if the resolution is lower.
[0115] According to embodiments of this disclosure, by parsing the received video data packets to obtain candidate video streams and their extended information at multiple spatial levels, and then selecting the most suitable target video stream according to the selection strategy, the adaptability and user experience of video communication can be effectively improved, ensuring smooth playback of video streams under different network conditions and device performance, and helping to adapt to different application scenarios and user needs.
[0116] The following is for reference. Figure 5 The video data transmission method 400 according to an embodiment of the present invention will be further described.
[0117] According to embodiments of this disclosure, the selection strategy can indicate the selection direction; the video data transmission method 400 may further include the following operations: obtaining historical attribute values corresponding to the target attribute field in historical video data packets; and determining the selection direction based on multiple historical attribute values, the current attribute value corresponding to the target attribute field in the video data packets, and the actual attribute value of the streaming end.
[0118] Historical video data packets refer to previously received video data packets, containing past video data and attribute information. Target attribute fields are specific attribute fields that are of primary focus in the selection strategy and used to evaluate the quality or suitability of the video stream. Historical attribute values are values extracted from historical video data packets corresponding to the target attribute fields, used to analyze past video stream characteristics. Current attribute values are values extracted from currently received video data packets corresponding to the target attribute fields, used to reflect the characteristics of the current video stream. Actual attribute values refer to the attribute values of the streaming endpoint itself.
[0119] In one example, the streaming client can use a caching mechanism to store attribute values from recently received video packets for quick access and analysis. In another example, the streaming client can also store attribute values from historical video packets in a database and retrieve historical attribute values by querying the database.
[0120] After obtaining the historical, current, and actual attribute values corresponding to the target attribute field, the selection direction can be determined based on these values. The selection direction refers to the trend or tendency of video stream selection indicated by the selection strategy, such as selecting towards higher or lower resolution.
[0121] In one example, a machine learning model (such as a decision tree or neural network) can be trained on historical attribute values, current attribute values, and the actual attributes of the streaming source. The model outputs the selection direction. In another example, a rule engine can be used to define a series of rules, and the selection direction is determined based on the rule matching results. For example, if the current network bandwidth is below a certain threshold, the rule engine triggers a rule to select the low-resolution video stream.
[0122] According to embodiments of this disclosure, by defining a selection strategy that comprehensively considers historical attribute values, current attribute values, and actual attributes of the streaming end, the selection direction of the video stream can be intelligently determined, thereby adapting to different network conditions and device performance, and improving the adaptability of video communication and user experience.
[0123] According to embodiments of this disclosure, when the target attribute field is a resolution status field, the historical attribute value is the historical resolution status value, the current attribute value is the current resolution status value, and the actual attribute value includes the required spatial level and the obtained spatial level. When the target attribute field is a stable status field, the current attribute value is the current stable status value; when the target attribute field is a bitrate status field, the actual attribute value is the obtained bitrate status value; and when the target attribute field is a bandwidth status field, the actual attribute value is the obtained bandwidth status value.
[0124] The spatial layer corresponding to a video stream with a historical resolution state value of 1 can be represented as `last_urgent_sp_layer`, and the spatial layer corresponding to a video stream with a current resolution state value of 1 can be represented as `new_urgent_sp_layer`. The requested spatial layer refers to the spatial layer of the video stream currently being pulled by the streaming client, and can be represented as `preferred_sp_layer`. The obtained spatial layer refers to the spatial layer of the video stream actually pulled by the streaming client, and can be represented as `current_sp_layer`.
[0125] According to embodiments of this disclosure, by clearly defining the historical attribute values, current attribute values, and actual attribute values under different target attribute fields, detailed basis is provided for the selection of video streams. This ensures that the streaming end can make the optimal selection by fully considering historical data, current conditions, and actual needs when selecting video streams, so as to adapt to different network conditions and device performance, and improve the stability of video communication and user experience.
[0126] According to embodiments of this disclosure, the selection direction includes an upgrade direction or a downgrade direction. Determining the selection direction based on multiple historical attribute values, the current attribute value corresponding to the target attribute field in the video data packet, and the actual attribute value of the streaming end may include the following operations: if the historical attribute value, the target attribute value, and the actual attribute value meet preset upgrade conditions, the selection direction is determined to be an upgrade direction; and if the historical attribute value, the target attribute value, and the actual attribute value meet preset downgrade conditions, the selection direction is determined to be a downgrade direction.
[0127] Preset upgrade conditions refer to predefined conditions used to determine whether a higher quality video stream should be selected. In one example, preset upgrade conditions include at least one of the following: the spatial level corresponding to the current resolution state value is greater than or equal to the spatial level corresponding to the historical resolution state value; the spatial level corresponding to the current resolution state is equal to the required spatial level; the obtained spatial level is less than the required spatial level; or the current stable state field value is a first target value. The first target value can be, for example, 1.
[0128] For example, if the spatial layer (new_urgent_sp_layer) corresponding to the current resolution state value is greater than the spatial layer corresponding to the historical resolution state value (last_urgent_sp_layer), and the required spatial layer (prefered_sp_layer) is equal to the spatial layer (new_urgent_sp_layer) corresponding to the current resolution state value, it indicates that a stream upgrade event may occur later, but it is still necessary to wait for the stream corresponding to the spatial layer (new_urgent_sp_layer) corresponding to the current resolution state value to stabilize before triggering the stream upgrade event.
[0129] Alternatively, if the spatial layer corresponding to the current resolution state value (new_urgent_sp_layer) is equal to the spatial layer corresponding to the historical resolution state value (last_urgent_sp_layer), and the required spatial layer (prefered_sp_layer) is equal to the spatial layer corresponding to the current resolution state value (new_urgent_sp_layer), and the obtained spatial layer (current_sp_layer) is less than the required spatial layer (prefered_sp_layer), it indicates that the flow is stabilizing and a flow upgrade event can be triggered.
[0130] Preset degradation conditions refer to predefined conditions used to determine whether a lower quality video stream should be selected. In one example, preset degradation conditions include at least one of the following: the current resolution status value is less than the historical resolution status value, the historical resolution status value is equal to the required resolution status value, the bitrate status field value is a second target value, and the current stable status field value is a third target value. The second target value can be, for example, 3, and the third target value can be, for example, 2 or 3.
[0131] For example, if the spatial layer (new_urgent_sp_layer) corresponding to the current resolution state value is less than the spatial layer (last_urgent_sp_layer) corresponding to the historical resolution state value, and the preferred spatial layer (prefered_sp_layer) is equal to the spatial layer (last_urgent_sp_layer) corresponding to the historical resolution state value, it means that the video stream at the spatial layer (last_urgent_sp_layer) corresponding to the historical resolution state value has stopped pushing. The stream degradation event can be triggered directly, and a suitable and stable stream layer can be selected for degradation.
[0132] Alternatively, if the bitrate status field of the spatial layer (current_sp_layer) is the second target value and the current stable status field is the third target value, then the stream degradation event can be triggered in advance depending on the business scenario.
[0133] According to embodiments of this disclosure, a flexible video stream selection mechanism is provided by defining preset upgrade conditions and preset downgrade conditions. This mechanism can comprehensively consider multiple factors such as spatial hierarchy, stream stability, bitrate status, and bandwidth trends to intelligently and comprehensively select the most suitable video stream quality. This not only improves the user experience but also ensures the smoothness and stability of video communication, adapting to diverse practical application scenarios.
[0134] According to embodiments of this disclosure, by defining upgrade and downgrade directions and comprehensively judging the selection direction based on historical attribute values, current attribute values, and actual attribute values, it is possible to intelligently adapt to different network conditions and device performance. This helps to ensure that a higher quality video stream is selected to improve the user experience when the conditions are met, while a lower quality video stream is selected when necessary to ensure smooth playback. This allows for flexible responses to various scenarios and improves the stability and reliability of video communication.
[0135] According to embodiments of this disclosure, the selection strategy further indicates a selection evaluation value, and the extended information includes the attribute values of at least one attribute field; the video data transmission method 400 described above may further include the following operations: determining a weight for each attribute field based on a pre-defined weight range and demand priority for each attribute field; and for each candidate video stream, determining a selection evaluation value for each candidate audio stream based on the attribute values of at least one attribute field and the weights for each attribute field.
[0136] For each attribute field, the weight for that attribute field can be determined based on the pre-defined weight range and requirement priority. The weight range refers to the range of values for the pre-defined weight for each attribute field; the weight indicates the relative importance of that attribute field in the selection and evaluation process. Requirement priority refers to the priority set by the user or system for different attribute fields based on actual needs, used to guide the determination of the weights.
[0137] According to embodiments of this disclosure, when the attribute field is a resolution status field, the demand priority of the resolution status field is determined based on the streaming end's preference for spatial hierarchy; when the attribute field is a bitrate status field, the demand priority is determined based on the degree of influence of the bitrate status value on the playback effect of the streaming end; when the attribute field is a bandwidth status field, the demand priority is determined based on the degree of influence of the bandwidth status value on the playback effect of the streaming end.
[0138] For the resolution status field, the weight range corresponding to the resolution status field can be 0.2 to 0.4, and is not limited here. The priority of the requirements is determined based on the streaming client's preference for spatial levels, and the weight for the resolution status field is selected within this weight range based on the requirement priority. The preference refers to the streaming client's preference for spatial levels, which can be determined based on user settings or device characteristics. For example, if the preference for spatial levels is higher, the requirement priority is higher, and the corresponding weight is higher.
[0139] For the bitrate status field, the weight range corresponding to the bitrate status field can be 0.3 to 0.5, and is not limited here. The priority of requirements is determined based on the degree of influence of the bitrate status value on the playback effect of the streaming end, and the weight for the bitrate status field is selected within this weight range based on the requirement priority. The degree of influence refers to the magnitude of the impact of the bitrate status value on the playback effect of the streaming end, which can be evaluated based on historical data or experimental results.
[0140] For the bandwidth status field, the corresponding weight range can be 0.2 to 0.4, and is not limited here. The priority of requirements is determined based on the degree of influence of the bandwidth status value on the playback effect of the streaming device, and the weight for the bandwidth status field is selected within this weight range based on the requirement priority. The degree of influence refers to the magnitude of the impact of the bandwidth status value on the playback effect of the streaming device, which can be evaluated based on historical data or experimental results.
[0141] According to embodiments of this disclosure, by dynamically adjusting the priority and corresponding weight of different attribute fields, the video stream selection strategy can accurately match the actual needs and network conditions of the streaming end, effectively improving the adaptability and user experience of video communication, ensuring that the most suitable video stream is selected in different scenarios, and optimizing the playback effect.
[0142] In one example, machine learning algorithms can be used to dynamically adjust weights, automatically optimizing weight allocation based on historical selection results and user feedback. In another example, different weight allocations can be preset according to different use cases.
[0143] After obtaining the weight of each attribute field, the selection evaluation value of the candidate audio stream can be determined based on the attribute value and weight of each attribute field. The selection evaluation value is a numerical value used to evaluate the suitability of the candidate video stream; for example, the higher the selection evaluation value, the more the candidate video stream meets the selection criteria.
[0144] The evaluation value can be represented as SourceStreamScore, and the specific calculation method can be adjusted according to the business scenario. In one example, the evaluation value can be determined by comprehensively considering multiple factors such as spatial hierarchy, stream stability, bitrate status and bandwidth trend, as shown in the following formula (5).
[0145] (5)
[0146] in, Characterize the selection of evaluation values, This represents a spatial hierarchy index, with higher spatial levels corresponding to larger index values. Characterizes the weights used for spatial hierarchy. The resolution status field represents the resolution status field. The field representing the bit rate status. Characterizes the weights used in the bitrate status field. A field representing bandwidth status. Weights representing bandwidth trends This represents the smooth selection evaluation value from the previous calculation.
[0147] According to embodiments of this disclosure, by assigning weights to each attribute field and calculating the selection evaluation value based on the attribute value and weight, a flexible and configurable video stream selection mechanism can dynamically adjust the selection strategy according to actual needs and scenarios, ensuring that the most suitable video stream is selected. This allows for general adaptation to different application scenarios and user needs, thereby improving the stability of video communication and user experience.
[0148] Figure 5 The illustration shows an example schematic diagram of a target video stream determination process according to an embodiment of the present disclosure.
[0149] like Figure 5 As shown in diagram 500, after receiving video data packet 510, the streaming end can parse video data packet 510 to obtain video data for each of the candidate video streams at multiple spatial levels. For example, by parsing video data packet 510, sub-video data packets 501, 502, ..., 50N are obtained. N is a positive integer.
[0150] Each sub-video data packet corresponds one-to-one with a candidate video stream. In addition, each sub-video data packet carries extended information to characterize the spatial hierarchy of the candidate video stream's state attributes. For example, sub-video data packet 501 corresponds to candidate video stream 5011, and also carries extended information 5012 corresponding to candidate video stream 5011; sub-video data packet 502 corresponds to candidate video stream 502, and also carries extended information 5022 corresponding to candidate video stream 5021; ...; and so on, sub-video data packet 50N corresponds to candidate video stream 50N1, and also carries extended information 50N2 corresponding to candidate video stream 50N1. N is a positive integer.
[0151] After parsing and obtaining multiple pieces of extended information, a selection direction 520 and a selection evaluation value 530 can be determined based on the extended information, and a selection strategy 540 can be determined based on the selection direction 520 and the selection evaluation value 530. On this basis, a target video stream 550 can be selected from candidate video streams at multiple spatial levels according to the selection strategy 540.
[0152] Based on the above-described video data transmission method 200, the present invention also provides a video data transmission device. The following will be combined with... Figure 6 The device is described in detail.
[0153] Figure 6 A block diagram of a video data transmission apparatus according to an embodiment of the present disclosure is shown schematically.
[0154] like Figure 6As shown, the video data transmission device 600 may include a generation module 610 and a transmission module 620.
[0155] The generation module 610 is used to generate video data packets based on the video stream to be transmitted. The video data packets include video data of candidate video streams at multiple spatial levels corresponding to the video stream to be transmitted. The video data carries extended information for characterizing the state attributes of the candidate video streams at the spatial levels.
[0156] The transmission module 620 is used to send video data packets to the streaming end, so that the streaming end can select the target video stream from candidate video streams at multiple spatial levels based on multiple extended information in the video data packets.
[0157] According to embodiments of this disclosure, the extended information includes attribute fields carried by the extended header of a real-time transport protocol. The attribute fields include at least one of the following: a resolution status field, a stability status field, a bitrate status field, and a bandwidth status field. The resolution status field is used to characterize whether the candidate video stream is the maximum resolution video stream that the streaming end can send. The stability status field is used to characterize whether the pushing status of the candidate video stream is stable. The bitrate status field is used to characterize the encoding bitrate level assigned to the candidate video stream by the video data packet. The bandwidth status field is used to characterize the bandwidth change trend of the candidate video stream.
[0158] According to embodiments of this disclosure, the resolution status value corresponding to the resolution status field is determined in the following manner: when the candidate video stream is not the maximum resolution video stream that the streaming end can send, the resolution status value is set to a first preset value; and when the candidate video stream is the maximum resolution video stream that the streaming end can send, the resolution status value is set to a second preset value.
[0159] According to embodiments of this disclosure, the stable state value corresponding to the stable state field is determined in the following manner: when the candidate video stream does not meet the preset stability conditions, the stable state value is set to a first preset value; and when the candidate video stream meets the preset stability conditions, the stable state value is set to a second preset value, wherein the preset stability conditions are determined based on the time difference between the video data packet generation time and the time when the most recent sufficient bandwidth event was detected and the stable reference duration.
[0160] According to embodiments of this disclosure, the bitrate status value corresponding to the bitrate status field is determined as follows: when the allocated coding bitrate is between the minimum bandwidth and the critical bandwidth, the bitrate status value is configured to a first preset value; when the allocated coding bitrate is between the critical bandwidth and the target bandwidth, the bitrate status value is configured to a second preset value; when the allocated coding bitrate is between the target bandwidth and the maximum bandwidth, the bitrate status value is configured to a third preset value; and when the allocated coding bitrate is the maximum bandwidth, the bitrate status value is configured to a fourth preset value.
[0161] According to embodiments of this disclosure, the bandwidth status value corresponding to the bandwidth status field is determined by: smoothing the acquired historical bandwidth allocation data to obtain smoothed bandwidth allocation data; performing linear regression analysis on the smoothed bandwidth allocation data to obtain discrete information characterizing the bandwidth change trend, wherein the discrete information includes at least one of slope, intercept, and coefficient of determination; and determining the bandwidth status value based on the discrete information characterizing the bandwidth change trend.
[0162] According to embodiments of this disclosure, determining a bandwidth state value based on discrete values characterizing bandwidth change trends includes: configuring the bandwidth state value to a first preset value when the discrete information characterizes a continuously decreasing bandwidth change trend; configuring the bandwidth state value to a second preset value when the discrete information characterizes a fluctuating decreasing bandwidth change trend; configuring the bandwidth state value to a third preset value when the discrete information characterizes a generally increasing bandwidth change trend; and configuring the bandwidth state value to a fourth preset value when the discrete information characterizes a relatively stable bandwidth change trend.
[0163] According to embodiments of this disclosure, the generation module 610 may include a first determining unit and a generation unit.
[0164] The first determining unit is used to determine multiple configuration combinations and extended information for each configuration combination based on multiple candidate resolutions and multiple candidate bitrates for the streaming end. Each configuration combination includes a candidate resolution and a candidate bitrate, and the configuration combination corresponds one-to-one with the spatial hierarchy.
[0165] The generation unit is used to generate video data for candidate video streams at the spatial level based on configuration combinations and extended information.
[0166] Based on the aforementioned video data transmission method 400, the present invention also provides a video data transmission device. The following will be combined with... Figure 7 The device is described in detail.
[0167] Figure 7 A block diagram of a video data transmission apparatus according to an embodiment of the present disclosure is shown schematically.
[0168] like Figure 7 As shown, the video data transmission device 700 may include a parsing module 710 and a selection module 720.
[0169] The parsing module 710 is used to parse the video data packets received from the streaming end in response to obtain video data for each of the candidate video streams at multiple spatial levels. The video data carries extended information about the attributes of the candidate video streams at the spatial levels.
[0170] Selection module 720 is used to select a target video stream from candidate video streams at multiple spatial levels based on a selection strategy determined based on multiple extended information.
[0171] According to embodiments of this disclosure, the selection strategy is determined based on a selection evaluation value; the video data transmission apparatus 700 may further include an acquisition module and a first determination module.
[0172] The acquisition module is used to retrieve the historical attribute values corresponding to the target attribute fields in historical video data packets.
[0173] The first determining module is used to determine the selection direction based on multiple historical attribute values, the current attribute value corresponding to the target attribute field in the video data packet, and the actual attribute value of the streaming end.
[0174] According to embodiments of this disclosure, the selected direction includes an upgrade direction or a downgrade direction; the determining module may include a second determining unit and a third determining unit.
[0175] The second determining unit is used to determine the selected direction as the upgrade direction when the historical attribute value, target attribute value, and actual attribute value meet the preset upgrade conditions.
[0176] The third determining unit is used to determine the selection direction as the degradation direction when the historical attribute value, target attribute value, and actual attribute value meet the preset degradation conditions.
[0177] According to embodiments of this disclosure, when the target attribute field is a resolution status field, the historical attribute value is the historical resolution status value, the current attribute value is the current resolution status value, and the actual attribute value includes the required resolution status value and the obtained resolution status value; when the target attribute field is a stable status field, the current attribute value is the current stable status value; when the target attribute field is a bitrate status field, the actual attribute value is the obtained bitrate status value; and when the target attribute field is a bandwidth status field, the actual attribute value is the obtained bandwidth status value.
[0178] According to embodiments of this disclosure, the preset upgrade conditions include at least one of the following: the current resolution status value is greater than or equal to the historical resolution status value, the current resolution status value is equal to the required resolution status value, the obtained resolution status value is less than the required resolution status value, and the current stable status field value is a first target value; the preset downgrade conditions include at least one of the following: the current resolution status value is less than the historical resolution status value, the historical resolution status value is equal to the required resolution status value, the bitrate status field value is a second target value, and the current stable status field value is a third target value.
[0179] According to embodiments of this disclosure, the selection strategy further indicates the selection of an evaluation value, and the extended information includes attribute values for each of at least one attribute field; the video data transmission device 700 may include a second determination module and a third determination module.
[0180] The second determining module is used to determine the weight for each attribute field based on the pre-defined weight range and priority of each attribute field.
[0181] The third determining module is used to determine the selection evaluation value of each candidate audio stream for each candidate video stream based on the attribute values of at least one attribute field and the weights used for each attribute field.
[0182] According to embodiments of this disclosure, when the attribute field is a resolution status field, the demand priority of the resolution status field is determined based on the streaming end's preference for spatial hierarchy; when the attribute field is a bitrate status field, the demand priority is determined based on the degree of influence of the bitrate status value on the playback effect of the streaming end; when the attribute field is a bandwidth status field, the demand priority is determined based on the degree of influence of the bandwidth status value on the playback effect of the streaming end.
[0183] Any one or more of the modules, submodules, units, and subunits according to embodiments of the present disclosure, or at least part of the functions of any one or more of them, can be implemented in one module. Any one or more of the modules, submodules, units, and subunits according to embodiments of the present disclosure can be implemented by dividing them into multiple modules. Any one or more of the modules, submodules, units, and subunits according to embodiments of the present disclosure can be at least partially implemented as hardware circuitry, such as a Field-Programmable Gate Array (FPGA), a Programmable Logic Array (PLA), a System-on-Chip, a System-on-a-Substrate, a System-on-Package, an Application-Specific Integrated Circuit (ASIC), or implemented in hardware or firmware by any other reasonable means of integrating or packaging circuitry, or implemented in software, hardware, or firmware, or in any suitable combination of any of these three implementation methods. Alternatively, one or more of the modules, submodules, units, and subunits according to embodiments of the present disclosure can be at least partially implemented as computer program modules, which, when run, can perform corresponding functions.
[0184] It should be noted that the video data transmission device part in the embodiments of this disclosure corresponds to the video data transmission method part in the embodiments of this disclosure. For a detailed description of the video data transmission device part, please refer to the video data transmission method part, which will not be repeated here.
[0185] Figure 8 A block diagram of an electronic device suitable for implementing a video data transmission method according to an embodiment of the present disclosure is shown schematically. Figure 8 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0186] like Figure 8 As shown, a computer electronic device 800 according to an embodiment of the present disclosure includes a processor 801, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage portion 809 into a random access memory (RAM) 803. The processor 801 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or an associated chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 801 may also include onboard memory for caching purposes. The processor 801 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.
[0187] RAM 803 stores various programs and data required for the operation of electronic device 800. Processor 801, ROM 802, and RAM 803 are interconnected via bus 804.
[0188] According to embodiments of this disclosure, the electronic device 800 may further include an input / output (I / O) interface 805, which is also connected to a bus 804. The electronic device 800 may also include one or more of the following components connected to the input / output (I / O) interface 805: an input section 806 including a keyboard, mouse, etc.; an output section 807 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 808 including a hard disk, etc.; and a communication section 809 including a network interface card such as a LAN card, modem, etc. The communication section 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to the input / output (I / O) interface 805 as needed. A removable medium 811, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 810 as needed so that computer programs read from it can be installed into the storage section 808 as needed.
[0189] This disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into the device / apparatus / system. The computer-readable storage medium carries one or more programs, which, when executed, implement the video data transmission method according to the embodiments of this disclosure.
[0190] In this disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0191] Embodiments of this disclosure also include a computer program product comprising a computer program containing program code for performing the methods provided in the embodiments of this disclosure. When the computer program product is run on an electronic device, the program code is used to enable the electronic device to implement the video data transmission method provided in the embodiments of this disclosure.
[0192] When the computer program is executed by the processor 801, it performs the functions defined in the system / apparatus of this disclosure embodiments. According to embodiments of this disclosure, the systems, apparatuses, modules, units, etc., described above can be implemented by computer program modules.
[0193] According to embodiments of this disclosure, program code for executing computer programs provided in embodiments of this disclosure can be written in any combination of one or more programming languages.
[0194] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. It should also be noted that in some alternative implementations, the functions indicated in the boxes may occur in a different order than those shown in the drawings.
[0195] The embodiments of this disclosure have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of this disclosure. Although various embodiments have been described above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. The scope of this disclosure is defined by the appended claims and their equivalents. Various substitutions and modifications can be made by those skilled in the art without departing from the scope of this disclosure, and all such substitutions and modifications should fall within the scope of this disclosure.
Claims
1. A video data transmission method, applied at a streaming end, the method comprising: Based on the video stream to be transmitted, a video data packet is generated, wherein the video data packet includes video data of candidate video streams at multiple spatial levels corresponding to the video stream to be transmitted, and the video data carries extended information for characterizing the state attributes of the candidate video streams at the spatial levels; and The video data packet is sent to the streaming end, so that the streaming end selects the target video stream from the candidate video streams of the multiple spatial levels based on the multiple extended information in the video data packet.
2. The method according to claim 1, wherein, The extended information includes attribute fields carried in the extended header of the real-time transport protocol, and the attribute fields include at least one of the following: resolution status field, stability status field, bitrate status field, and bandwidth status field; The resolution status field is used to characterize whether the candidate video stream is the highest resolution video stream that the streaming end can send. The stability status field is used to characterize whether the pushing status of the candidate video stream is stable. The bitrate status field is used to characterize the encoding bitrate level allocated by the video data packet to the candidate video stream. The bandwidth status field is used to characterize the bandwidth change trend of the candidate video stream.
3. The method according to claim 2, wherein, The resolution status value corresponding to the resolution status field is determined in the following way: If the candidate video stream is not the highest resolution video stream that the streaming end can send, the resolution status value is set to a first preset value. as well as If the candidate video stream is the highest resolution video stream that the streaming end can send, the resolution status value is set to a second preset value.
4. The method according to claim 2, wherein, The steady-state value corresponding to the steady-state field is determined in the following way: If the candidate video stream does not meet the preset stability condition, the stable state value is set to the first preset value; as well as If the candidate video stream meets the preset stability condition, the stable state value is set to a second preset value, wherein the preset stability condition is determined based on the time difference between the generation time of the video data packet and the time of the most recent detection of a sufficient bandwidth event and the stable reference duration.
5. The method according to claim 2, wherein, The bitrate status value corresponding to the bitrate status field is determined in the following way: When the allocated coding rate is between the minimum bandwidth and the critical bandwidth, the code rate status value is configured to a first preset value; When the allocated coding rate is between the critical bandwidth and the target bandwidth, the code rate status value is configured to a second preset value; When the allocated coding rate is between the target bandwidth and the maximum bandwidth, the code rate status value is configured to a third preset value; and When the allocated coding rate is the maximum bandwidth, the code rate status value is configured to a fourth preset value.
6. The method according to claim 2, wherein, The bandwidth status value corresponding to the bandwidth status field is determined in the following way: The acquired historical bandwidth allocation data is smoothed to obtain smoothed bandwidth allocation data. Linear regression analysis is performed on the smoothed bandwidth allocation data to obtain discrete information characterizing the bandwidth change trend, wherein the discrete information includes at least one of slope, intercept, and coefficient of determination; and The bandwidth state value is determined based on the discrete information that characterizes the bandwidth change trend.
7. The method according to claim 6, wherein, Determining the bandwidth state value based on the discrete value characterizing the bandwidth change trend includes: When the discrete information characterizes the bandwidth change trend as a continuous decrease, the bandwidth state value is configured to a first preset value; When the discrete information characterizes the bandwidth change trend as fluctuating downward, the bandwidth state value is configured to a second preset value; When the discrete information characterizes the bandwidth change trend as generally increasing, the bandwidth state value is configured to a third preset value; and When the discrete information characterizes the bandwidth change trend as relatively stable, the bandwidth state value is configured to a fourth preset value.
8. The method according to any one of claims 1 to 7, wherein, The process of generating video data packets based on the video stream to be transmitted includes: Based on multiple candidate resolutions and multiple candidate bitrates used for the streaming end, multiple configuration combinations and extended information for each configuration combination are determined, wherein each configuration combination includes one candidate resolution and one candidate bitrate, and the configuration combination corresponds one-to-one with the spatial hierarchy; and Based on the configuration combination and the extended information, video data for the candidate video streams at the spatial level is generated.
9. A video data transmission method, applied at a streaming end, the method comprising: In response to receiving a video data packet from the streaming end, the video data packet is parsed to obtain video data for each of multiple spatial level candidate video streams, wherein the video data carries extended information for characterizing the attributes of the candidate video streams at the spatial level; and The target video stream is selected from the candidate video streams of the multiple spatial levels according to the selection strategy determined based on the multiple extended information.
10. The method according to claim 9, wherein, The selection strategy indicates the selection direction; The method further includes, after parsing the video data packets to obtain video data for each of the candidate video streams at multiple spatial levels: Retrieve the historical attribute values corresponding to the target attribute field from the historical video data packets; as well as The selection direction is determined based on multiple historical attribute values, the current attribute value in the video data packet corresponding to the target attribute field, and the actual attribute value of the streaming end.
11. The method according to claim 10, wherein, The selection direction includes either an upgrade direction or a downgrade direction; Determining the selection direction based on multiple historical attribute values, the current attribute value in the video data packet corresponding to the target attribute field, and the actual attribute value of the streaming end includes: If the historical attribute value, the target attribute value, and the actual attribute value meet the preset upgrade conditions, the selected direction is determined as the upgrade direction; and If the historical attribute value, the target attribute value, and the actual attribute value meet the preset degradation conditions, the selected direction is determined as the degradation direction.
12. The method according to claim 11, wherein, When the target attribute field is the resolution status field, the historical attribute value is the historical resolution status value, the current attribute value is the current resolution status value, and the actual attribute value includes the required spatial level and the obtained spatial level; If the target attribute field is a stable state field, the current attribute value is the current stable state value; When the target attribute field is a bitrate status field, the actual attribute value is the obtained bitrate status value; when the target attribute field is a bandwidth status field, the actual attribute value is the obtained bandwidth status value.
13. The method according to claim 12, wherein, The preset upgrade conditions include at least one of the following: the spatial level corresponding to the current resolution state value is greater than or equal to the spatial level corresponding to the historical resolution state value; the spatial level corresponding to the current resolution state is equal to the required spatial level; the obtained spatial level is less than the required spatial level; and the current stable state field value is a first target value. The preset degradation conditions include at least one of the following: the spatial level corresponding to the current resolution status value is less than the spatial level corresponding to the historical resolution status value; the spatial level corresponding to the historical resolution status value is equal to the required spatial level; the bitrate status field value is a second target value; and the current stable status field value is a third target value.
14. The method of claim 10, wherein, The selection strategy is determined based on the selection evaluation value, and the extended information includes the attribute values of at least one attribute field. The method further includes: Based on the pre-defined weight range and priority of each attribute field, determine the weight for each attribute field; and For each of the candidate video streams, a selection evaluation value for each of the candidate audio streams is determined based on the attribute values of the at least one attribute field and the weights used for each attribute field.
15. The method according to claim 14, wherein, When the attribute field is a resolution status field, the priority of the requirement for the resolution status field is determined based on the degree of preference of the streaming end for the spatial hierarchy; When the attribute field is a bitrate status field, the demand priority is determined based on the degree of influence of the bitrate status value on the playback effect of the streaming end; When the attribute field is a bandwidth status field, the demand priority is determined based on the degree of influence of the bandwidth status value on the playback effect of the streaming end.
16. A video data transmission apparatus, applied at a streaming end, the method comprising: A generation module is used to generate video data packets based on the video stream to be transmitted, wherein the video data packets include video data of candidate video streams at multiple spatial levels corresponding to the video stream to be transmitted, and the video data carries extended information for characterizing the state attributes of the candidate video streams at the spatial levels. A transmission module is used to send the video data packet to the streaming end, so that the streaming end selects the target video stream from multiple candidate video streams at multiple spatial levels based on multiple extended information in the video data packet.
17. A video data transmission device, applied at a streaming end, the method comprising: The parsing module is used to parse the video data packet received from the streaming end in response to obtain video data of each of the candidate video streams at multiple spatial levels, wherein the video data carries extended information for characterizing the attributes of the candidate video streams at the spatial level. The selection module is used to select a target video stream from candidate video streams at multiple spatial levels according to a selection strategy determined based on multiple extended information.
18. An electronic device comprising: One or more processors; Memory, used to store one or more computer programs. The characteristic feature is that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 15.
19. A computer-readable storage medium having a computer program or instructions stored thereon, characterized in that, When the computer program or instructions are executed by a processor, they implement the steps of the method according to any one of claims 1 to 15.
20. A computer program product comprising a computer program or instructions, characterized in that, When the computer program or instructions are executed by a processor, they implement the steps of the method according to any one of claims 1 to 15.