Method and system for reducing transmission bandwidth and improving transmission stability based on WebRTC

Through AI scene detection and ROI detection in area of ​​interest, the funnel algorithm in WebRTC is optimized, and the code rate and frame rate of video frames are dynamically adjusted, which solves the problem that WebRTC is difficult to transmit high-bit rate video data in low bandwidth environments, and achieves more stable and smooth video transmission.

CN120034651APending Publication Date: 2025-05-23SHANGHAI WONDERTEK SOFTWARE CORP LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510137075.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-07
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

WebRTC is difficult to effectively transmit high-bit rate video data in low bandwidth environments, resulting in significant video delays in motion scenarios, frequent picture stuttering, and high packet loss rate, affecting real-time and fluency.

Method used

Through AI scene detection and ROI detection in area of ​​interest, the funnel algorithm in WebRTC is optimized, the code rate and frame rate of video frames are dynamically adjusted, and different encoding parameters are set according to different scenarios and regions to reduce transmission bandwidth and improve stability.

Benefits of technology

In a low bandwidth environment, the bandwidth requirement for video transmission is significantly reduced, the stability and fluency of video transmission is improved, the picture lag and packet loss are reduced, and the user's visual experience is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120034651A_ABST
    Figure CN120034651A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of video acquisition coding and transmission, and provides a WebRTC-based method for reducing transmission bandwidth and improving transmission stability, which comprises the following steps: S1, a video acquisition module acquires a video frame from a camera and sends the video frame to a video coding module; s2, performing AI scene detection on the video frame to obtain a picture scene of the video frame, and performing region-of-interest ROI detection on the video frame to obtain a region-of-interest ROI range of the video frame; s3, optimizing the original funnel algorithm, setting different code rates for different video frames according to different video frame scenes, and setting different code rates for different areas of the video frames according to different ROI (Region of Interest) ranges of the video frames; s4, encoding the video frame, and sending the encoded video frame to a receiving end; s5, the receiving end decodes the received compressed video data into original video frames by using a video decoder; and S6, transmitting the decoded video frame to a rendering module for video rendering and display. The transmission bandwidth is reduced, and the transmission stability is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video acquisition coding and transmission, and in particular to a method and system for reducing transmission bandwidth and improving transmission stability based on WebRTC. Background Art

[0002] WebRTC (Web Real-Time Communication) protocol, as a revolutionary and widely used technical standard in the field of Internet communication technology today, plays a vital role in many key areas such as video calls, video conferencing, and real-time monitoring. With the rapid development of the Internet, the demand for real-time communication is growing. Traditional communication methods have gradually exposed many shortcomings in terms of real-time, convenience, and cross-platform compatibility. The WebRTC protocol came into being, which provides direct real-time communication capabilities between browsers and mobile applications without the need for additional plug-ins or third-party software support. This feature enables developers to easily integrate real-time communication functions into web pages and mobile applications, greatly promoting the popularization and development of real-time communication technology, and bringing unprecedented convenience to people's communication and information interaction. The WebRTC processing flow covers acquisition, encoding, transmission, decoding, and rendering and display.

[0003] When verifying WebRTC, the following issues were found: (1) Problems with low-bandwidth transmission in sports scenarios In today's uneven network environment, low-bandwidth scenarios are still common. In such an environment, WebRTC shows good adaptability and stability for static scenes. When the picture is static, WebRTC can effectively collect, encode and transmit video data to achieve a low-latency transmission effect. At this time, the picture presented by the receiving end is clear and the playback process is very smooth, which can provide users with a good visual experience. This is due to the fact that the picture content changes very little in static scenes, and the encoding process does not need to process a large amount of dynamic information, so that the picture quality can be guaranteed at a lower bit rate.

[0004] However, when the scene switches to a moving scene, the situation becomes complicated. Moving scenes mean that the objects in the picture are constantly moving and changing. In order to ensure the picture quality and allow users to clearly see the trajectory and details of the moving objects, the encoding process requires a higher bit rate. Because a high bit rate can carry more image information to accurately present the dynamic changes of moving objects. But in a low-bandwidth environment, high-bit rate video data transmission faces huge challenges.

[0005] Due to limited bandwidth, high-bitrate data transmission cannot be carried in time, which leads to a significant increase in transmission delay. Users will clearly feel the delay of the picture. For example, in a video call, the other party's actions and sounds cannot be presented synchronously, which seriously affects the real-time and smoothness of communication. At the same time, the picture freezes frequently, the movements of moving objects are no longer coherent, and there are jumps and pauses, which brings a very poor visual experience to users. In addition, the high bit rate will also lead to a significant increase in the packet loss rate. Some video data is lost during the transmission process, causing problems such as screen distortion and blurring of the picture, further reducing the picture quality.

[0006] (2) Image quality control strategy for high bandwidth requirements In order to ensure the image quality, native WebRTC adopts a relatively conservative strategy, that is, to control the quality of the entire image within a certain range. The starting point of this strategy is good, aiming to provide users with a stable and high-quality visual experience. However, it has obvious drawbacks in practical applications.

[0007] In actual video content, not all parts are equally important. For example, in a video conference, facial expressions, body language, and speech content are key information, while background and other aspects are relatively less important. However, native WebRTC does not distinguish between content, and both important and secondary content must maintain high picture quality. This means that even some parts that have little impact on users' understanding of the core content of the video will take up a lot of bandwidth resources to ensure their picture quality.

[0008] This "one-size-fits-all" quality control strategy results in a high bitrate for the entire image. When bandwidth is sufficient, the problem may not be noticeable, but in a network environment with limited bandwidth, it will have serious consequences. High bitrates overload the network and easily cause network congestion, which in turn leads to problems such as video transmission freezes and packet loss, affecting the overall use effect. Moreover, for some mobile devices or users with poor network conditions, excessively high bandwidth requirements may make it impossible to use WebRTC services normally.

[0009] (3) Limitations of the funnel algorithm for judging frame loss Native WebRTC uses the funnel algorithm to control frame loss. This algorithm plays an important role in the WebRTC system, which is to dynamically adjust the frame rate and transmission rate of the video stream to adapt to changes in network conditions and ensure the stability of video transmission. The working principle of the funnel algorithm is to dynamically adjust the sending rate and frame loss rate of video frames based on the processing power of the receiver and the actual situation of the network bandwidth. When the network bandwidth is sufficient and the receiver has strong processing power, the algorithm will allow a higher frame rate and sending rate to ensure the smoothness and clarity of the picture; when the network bandwidth is tight or the receiver's processing power is insufficient, the algorithm will appropriately reduce the frame rate or even discard some video frames to ensure the smoothness and quality of video transmission.

[0010] However, this algorithm has obvious flaws. It only judges whether a frame is lost based on the bit rate of the encoded video frame. This single judgment standard is too simple and crude, and does not take into account the importance of the content. In actual video scenes, the importance of information contained in different video frames is different. For example, in a live broadcast of a sports event, the video frames of key moments such as athletes shooting and scoring contain the most important information, and these frames must not be lost for users; while some insignificant scene switching frames or background picture frames, even if lost, will not have much impact on users' understanding of the game content. However, since the funnel algorithm does not refer to the importance of the content, it may mistakenly delete key frames during the frame loss process, causing users to miss important information.

[0011] In addition, the funnel algorithm does not make frame drop decisions based on specific scenarios. Different scenarios have different requirements for video quality and frame rate. In monitoring scenarios, for some static images that remain still for a long time, appropriately reducing the frame rate or even discarding some frames may not affect the monitoring effect; in video call scenarios, even if the network conditions are poor, it is necessary to ensure the smoothness of facial expressions and language communication as much as possible, and key frames cannot be discarded at will. However, the funnel algorithm cannot be flexibly adjusted according to the characteristics of these specific scenarios, and thus cannot provide the optimal video transmission solution in some scenarios.

[0012] In summary, although WebRTC has important application value in the field of real-time communication, these problems exposed in actual use require further research and improvement to improve its performance and stability in various network environments and application scenarios. Summary of the invention

[0013] In view of the above problems, the purpose of the present invention is to provide a method and system for reducing transmission bandwidth and improving transmission stability based on WebRTC, which involves reducing the bit rate of encoded video frames, reducing bandwidth, and improving the stability of WebRTC transmission. Specifically, based on the WebRTC architecture, different bit rates are set through scene detection, and different encoding bit rates are set for the region of interest and the region of non-interest by ROI (region of interest) to reduce the bit rate of the entire video frame, and the original funnel algorithm in WebRTC is optimized to reduce the transmission bandwidth and improve the stability of transmission.

[0014] The above-mentioned object of the present invention is achieved through the following technical solutions: A method for reducing transmission bandwidth and improving transmission stability based on WebRTC includes the following steps: S1: The camera of the sending end device collects video data, the video acquisition module obtains video frames from the camera, and sends the video frames to the video encoding module; S2: performing AI scene detection on the video frame to obtain the picture scene of the video frame, and performing region of interest ROI detection on the video frame to obtain the region of interest ROI range of the video frame; S3: Optimizing the original funnel algorithm, setting different bit rates for different video frames according to different video frame scenes, and setting different bit rates for different regions of the video frame according to different ranges of the region of interest ROI of the video frame; S4: After encoding the video frame, send it to the receiving end; S5: The receiving end uses a video decoder to decode the received compressed video data into the original video frame; S6: The decoded video frame is transmitted to a rendering module for video rendering and display.

[0015] Furthermore, before step S1, the method further includes: Setting a signaling service, the signaling service being used to establish and manage a connection between the sending end and the receiving end, the signaling service including session control information, codec information and other control signaling; The signaling service negotiates video codecs, network parameters and other session-related settings by exchanging Session Description Protocol (SDP) information.

[0016] Further, in step S2, AI scene detection is performed on the video frame to obtain the picture scene of the video frame, specifically: Using the AI ​​capabilities of the RK3588 chip, the AI ​​model is used to set and judge the scene of the video frame; A preset number of the video frames are set to be identified as the same picture scene, and when the number of the video frames reaches the preset number or a key frame change occurs, the picture scene is re-determined, wherein the picture scene is determined by sending the YUV data of the first key frame among the preset number of the video frames into a scene detection model for inference; Different qualities are set for different picture scenes, and different quantization parameter QP ranges are set in the video coding standard H.264 for the picture scenes including simple scenes, motion scenes, and complex scenes so that different picture scenes have different bit rates, wherein the larger the value of the quantization parameter QP, the lower the bit rate of the picture scene and the worse the picture quality.

[0017] Further, in step S2, a region of interest ROI detection is performed on the video frame to obtain a region of interest ROI range of the video frame, specifically: Obtaining YUV data of the first key frame in a group of continuous video frames of the same picture scene and feeding the data into the ROI detection model, and obtaining the ROI range of the region of interest of the video frame by using AI reasoning; By adopting region of interest ROI encoding, the bit rate of the original scene is maintained for the scene within the region of interest ROI, and the bit rate of the non-interested region is reduced.

[0018] Further, in step S3, the original funnel algorithm is optimized, different bit rates are set for different video frames according to different video frame scenes, and different bit rates are set for different regions of the video frame according to different ranges of the region of interest ROI of the video frame, specifically: According to the judgment result of the AI ​​scene detection, on the basis of the frame loss strategy of the original funnel algorithm of WebRTC, the frame loss frequency is reduced in turn according to the simple scene, the motion scene, and the complex scene, and the frame rate of the motion scene and the complex scene is increased to improve the fluency; According to the difference between the simple scene, the motion scene and the complex scene, different ranges of quantization parameters QP are set in the video coding standard H.264; The bit rate of the original scene is maintained for the scene within the region of interest ROI, and the bit rate of the non-region of interest is reduced to perform encoding control.

[0019] Furthermore, in step S4, encoding the video frame further includes: Use the hardware encoding on the RK3588 chip to encode to increase the speed of video encoding.

[0020] Furthermore, in step S4, after encoding the video frame, the video frame is sent to the receiving end, which also includes: After the video encoding is completed, the encoded data is encapsulated into a real-time transport protocol RTP packet, and the real-time transport protocol RTP packet is sent to the receiving end through the user datagram protocol UDP.

[0021] A system for reducing transmission bandwidth and improving transmission stability based on WebRTC for executing the above-mentioned method for reducing transmission bandwidth and improving transmission stability based on WebRTC comprises: A video data acquisition module, used to provide video data to a camera of a transmitting device, the video acquisition module obtains video frames from the camera, and sends the video frames to a video encoding module; A picture scene detection module is used to perform AI scene detection on the video frame to obtain the picture scene of the video frame, and at the same time perform region of interest ROI detection on the video frame to obtain the region of interest ROI range of the video frame; The original algorithm optimization module is used to optimize the original funnel algorithm, set different bit rates for different video frames according to different video frame scenes, and set different bit rates for different regions of the video frame according to different ranges of the region of interest ROI of the video frame; A video data sending module, used for encoding the video frame and sending it to a receiving end; A video data decoding module, used for providing the receiving end with a video decoder to decode the received compressed video data into the original video frame; The video data rendering module is used to provide the decoded video frame to the rendering module for video rendering and display.

[0022] A computer device comprises a memory and one or more processors, wherein the memory stores computer codes, and when the computer codes are executed by the one or more processors, the one or more processors execute the above method.

[0023] A computer-readable storage medium stores computer codes. When the computer codes are executed, the above method is executed.

[0024] Compared with the prior art, the present invention has at least one of the following beneficial effects: (1) Accurate bitrate setting to suit different scenarios The present invention uses the powerful AI capabilities of the RK3588 chip and uses the AI ​​model to accurately judge the picture scene of the video frame. Different quantization parameter QP ranges are set in the video coding standard H.264 for different picture scenes, such as simple scenes, motion scenes, and complex scenes. The quantization parameter QP is closely related to the bit rate and picture quality. The larger its value, the lower the bit rate and the worse the picture quality. By reasonably adjusting the QP range for different scenes, it is possible to minimize unnecessary bit rate consumption while ensuring that the picture quality meets basic requirements.

[0025] In addition to the scene-based bit rate setting, the present invention also performs region of interest (ROI) detection on the video frame. By feeding the YUV data of the first key frame in a group of continuous video frames of the same picture scene into the ROI detection model, the region of interest (ROI) range of the video frame is obtained by AI reasoning. During the encoding process, the bit rate of the original scene is maintained for the scene within the region of interest (ROI) to ensure the clarity and integrity of key information; while the bit rate is reduced for non-regions of interest. This targeted bit rate control strategy enables limited bandwidth resources to be more concentratedly used to transmit important information, avoiding bandwidth waste caused by uniform encoding of the entire picture, thereby further reducing the demand for transmission bandwidth.

[0026] (2) Optimize the funnel algorithm to improve the frame loss strategy The present invention is optimized based on the frame loss strategy of the original funnel algorithm of WebRTC. According to the judgment result of AI scene detection, the frame loss frequency is reduced in turn for simple scenes, motion scenes and complex scenes. In simple scenes, due to the small changes in the picture, appropriately increasing the frame loss frequency has little effect on the user's viewing experience, and can further reduce the amount of data transmission; while in motion scenes and complex scenes, reducing the frame loss frequency and increasing the frame rate can effectively improve the smoothness and continuity of the picture. For example, in a video conferencing scenario, if the picture containing key information such as the movements and expressions of the participants belongs to a motion scene or a complex scene, reducing frame loss can ensure that other participants can see this information clearly, avoid problems such as screen freezes and information loss, thereby improving the quality and efficiency of communication. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 This is an overall flow chart of the method for reducing transmission bandwidth and improving transmission stability based on WebRTC in the present invention; Figure 2 Flowchart of the native WebRTC implementation; Figure 3 A flow chart of an optimization implementation scheme of the present invention; Figure 4This is an overall structural diagram of the system for reducing transmission bandwidth and improving transmission stability based on WebRTC in the present invention. DETAILED DESCRIPTION

[0028] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0029] Those skilled in the art will appreciate that, unless otherwise stated, the singular forms "a", "an", "said" and "the" used herein may also include plural forms. It should be further understood that the term "comprising" used in the specification of the present invention refers to the presence of the features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.

[0030] First embodiment like Figure 1 As shown, this embodiment provides a method for reducing transmission bandwidth and improving transmission stability based on WebRTC. Figure 2 An improved technical solution is made based on the native WebRTC implementation solution shown in FIG. Figure 3 The optimized implementation scheme of the present invention shown in the figure specifically comprises the following steps: S1: The camera of the transmitting device collects video data, the video collection module obtains video frames from the camera, and sends the video frames to the video encoding module.

[0031] In this embodiment, the collected video data is usually stored in the YUV format. Before step S1, it includes: setting a signaling service, the signaling service is used to establish and manage a connection between the sending end and the receiving end, the signaling service includes session control information, codec information and other control signaling; the signaling service negotiates video codecs, network parameters and other session-related settings by exchanging session description protocol SDP information.

[0032] Detailed explanation of signaling service content: Session control information: This part of information is used to control the process of the entire communication session. For example, when initiating a video call, the session control information determines when to start the connection, when to end the session, and how to handle abnormal situations during the session (such as a short network interruption and then recovery), to ensure the continuity and stability of communication. With this information, the sender and receiver can coordinate and orderly transmit and interact data.

[0033] Codec information: This is the key information that determines how the video data is encoded and decoded. Different codecs have different coding efficiencies, picture quality, and bandwidth requirements. The codec information in the signaling service enables the sender and receiver to reach an agreement on which codec to use. For example, when the network bandwidth is limited, the two parties may negotiate to use a codec with a higher compression rate to reduce bandwidth usage while ensuring basic picture quality; when the picture quality requirements are extremely high and the bandwidth is sufficient, a codec that can provide higher picture quality will be selected.

[0034] Other control signaling: In addition to the above two types of information, the signaling service also includes a variety of other control signaling, such as signaling for adjusting video resolution and frame rate. These signaling can adjust video parameters in real time according to network conditions and device performance to adapt to different usage scenarios. When the network suddenly becomes congested, the sender can reduce the video resolution and frame rate through control signaling, thereby reducing the amount of data transmission and ensuring that the video call is not interrupted.

[0035] Negotiation mechanism based on SDP information: The signaling service negotiates parameters by exchanging Session Description Protocol (SDP) messages. SDP messages are a format for describing the attributes of a multimedia session, detailing video codecs, network parameters, and other session-related settings.

[0036] Video codec negotiation: When the sender and receiver exchange SDP information, they will list the video codecs they support. Then, both parties will select the most suitable codec based on the network conditions, device performance, and video quality requirements. For example, the sender supports codecs such as H.264 and VP9, ​​and the receiver supports H.264 and AV1. Through the interaction of SDP information, both parties found that H.264 has the best compatibility and performance in the current network environment, so they jointly selected H.264 as the video codec for this communication.

[0037] Network parameter negotiation: Network parameters include network bandwidth, delay tolerance, etc. Through SDP information, the sender and receiver can understand each other's network conditions and negotiate appropriate network parameters. If the network bandwidth of the receiver is low, the sender can adjust the video bit rate and resolution based on this information to ensure that the video can be played smoothly at the receiver. At the same time, the negotiation of network delay tolerance can also help both parties adopt corresponding strategies when transmitting data, such as setting an appropriate buffer size to reduce the jamming caused by network delay.

[0038] Negotiation of other session-related settings: This includes settings such as the video frame rate, resolution, and audio parameters. By exchanging SDP information, both parties can determine the most suitable session parameters based on actual needs and device capabilities. For example, in a video conference scenario, in order to ensure that participants can clearly see each other's expressions and movements, the video frame rate may be negotiated to be set to a higher value; while in some monitoring scenarios that do not require high image quality but have extremely high real-time requirements, the resolution may be appropriately reduced to increase the transmission speed.

[0039] S2: Perform AI scene detection on the video frame to obtain the picture scene of the video frame, and perform region of interest ROI detection on the video frame to obtain the region of interest ROI range of the video frame.

[0040] In step S2, AI scene detection is performed on the video frame to obtain the picture scene of the video frame, specifically: In order to set different bit rates according to different scenes, the present invention uses the AI ​​capability of the RK3588 chip and uses the AI ​​model to set and judge the scene of the video frame; In order to ensure low latency of WebRTC, the present invention uses a certain number of frame intervals to detect scenes, reducing the time consumption caused by scene detection. A preset number (e.g., 900 frames) of the video frames are identified as the same picture scene. When the number of the video frames reaches the preset number or a key frame change occurs, the picture scene is re-determined, wherein the picture scene is determined by sending the YUV data of the first key frame of the preset number of video frames into the scene detection model for inference; Different qualities are set for different picture scenes, and different quantization parameter QP ranges are set in the video coding standard H.264 for the picture scenes including simple scenes, motion scenes, and complex scenes so that different picture scenes have different bit rates, wherein the larger the value of the quantization parameter QP, the lower the bit rate of the picture scene and the worse the picture quality. For example, in native WebRTC, H264 sets the QP range to 24~37 (the larger the value, the lower the bit rate and the worse the picture quality). The present invention modifies it to 24~45 for simple scenes, 24~40 for motion scenes, and 24~37 for complex scenes, thereby reducing the video frame rate while ensuring the picture quality as much as possible.

[0041] In step S2, a region of interest ROI detection is performed on the video frame to obtain a region of interest ROI range of the video frame, specifically: The YUV data of the first key frame in a group of continuous video frames of the same picture scene is obtained and sent to the ROI detection model, and the ROI range of the video frame is obtained by AI reasoning; the ROI encoding of the region of interest is adopted to maintain the bit rate of the original scene for the scene within the ROI range of the region of interest, and reduce the bit rate of the non-interested area.

[0042] The video in a GOP (a group of continuous images) is the same scene, and only the scene of the key frame in a GOP needs to be detected. For example, for a 25fps video, the GOP is 2 seconds, so for 50 pictures, only the scene of the first picture needs to be detected. The present invention uses AI reasoning to determine the ROI range. For example, for the human body, the human body AI model returns the human body range, and the non-human body range can reduce the bit rate.

[0043] S3: Optimizing the original funnel algorithm, setting different bit rates for different video frames according to different video frame scenes, and setting different bit rates for different regions of the video frame according to different ranges of the region of interest ROI of the video frame.

[0044] In this embodiment, step S3 is specifically as follows: According to the judgment result of the AI ​​scene detection, on the basis of the frame loss strategy of the original funnel algorithm of WebRTC, the frame loss frequency is reduced in turn according to the simple scene, the motion scene, and the complex scene, and the frame rate of the motion scene and the complex scene is increased to improve the smoothness.

[0045] Optimize the WebRTC funnel algorithm. On the basis of the original frame drop strategy through bit rate control, the frame drop strategy is also combined with scene optimization. That is, in static or less moving scenes, the frame rate can be reduced to reduce the amount of data transmitted, thereby saving bandwidth and resources. In such scenes, the video content does not change frequently, and reducing the frame rate will not significantly affect the viewing experience; in dynamic or frequently changing scenes, a high frame rate should be maintained to ensure that motion details and dynamic changes are captured. This can be achieved by dynamically adjusting the frame drop strategy of the funnel algorithm, giving priority to retaining key frames or frames in dynamic areas.

[0046] According to the difference between the simple scene, the motion scene and the complex scene, different ranges of quantization parameters QP are set in the video coding standard H.264. For example, H264 sets QP with a simple scene range of 24-45, a motion scene range of 24-40, and a complex scene range of 24-37, so as to control different scenes to have different bit rates; The bit rate of the original scene is maintained for the scene within the region of interest ROI, and the bit rate of the non-region of interest is reduced to perform encoding control.

[0047] S4: After encoding the video frame, send it to the receiving end.

[0048] In this embodiment, the hardware encoding on the RK3588 chip is used for encoding to increase the speed of video encoding. After the video encoding is completed, the encoded data is encapsulated into a real-time transport protocol RTP packet, and the real-time transport protocol RTP packet is sent to the receiving end via the user datagram protocol UDP.

[0049] S5: The receiving end uses a video decoder to decode the received compressed video data into the original video frame. In edge devices, in order to increase the decoding speed, the decoding process usually uses hardware acceleration.

[0050] S6: The decoded video frames are transmitted to a rendering module for video rendering and display. The rendering module uses hardware acceleration to display the video frames on the screen to ensure smooth video playback and low latency.

[0051] Second embodiment like Figure 4 As shown, this embodiment provides a system for reducing transmission bandwidth and improving transmission stability based on WebRTC for executing the method for reducing transmission bandwidth and improving transmission stability based on WebRTC in the first embodiment, including: Video data acquisition module 1, used to provide video data to the camera of the sending end device, the video acquisition module obtains video frames from the camera, and sends the video frames to the video encoding module; A picture scene detection module 2 is used to perform AI scene detection on the video frame to obtain the picture scene of the video frame, and at the same time perform region of interest ROI detection on the video frame to obtain the region of interest ROI range of the video frame; The original algorithm optimization module 3 is used to optimize the original funnel algorithm, set different bit rates for different video frames according to different video frame scenes, and set different bit rates for different regions of the video frame according to different ranges of the region of interest ROI of the video frame; The video data sending module 4 is used to encode the video frame and send it to the receiving end; The video data decoding module 5 is used to provide the receiving end with a video decoder to decode the received compressed video data into the original video frame; The video data rendering module 6 is used to provide the decoded video frame to the rendering module for video rendering and display.

[0052] A computer-readable storage medium stores computer code. When the computer code is executed, the above method is executed. A person skilled in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing related hardware through a program. The program can be stored in a computer-readable storage medium. The storage medium can include: a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk, etc.

[0053] The above is only a preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions under the concept of the present invention belong to the protection scope of the present invention. It should be pointed out that for ordinary technicians in this technical field, some improvements and modifications without departing from the principle of the present invention should also be regarded as the protection scope of the present invention.

[0054] The technical features of the above-described embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0055] It should be noted that the above embodiments can be freely combined as needed. The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered as the protection scope of the present invention.

Claims

1. A method for reducing transmission bandwidth and improving transmission stability based on WebRTC, characterized in that: The following steps are involved: S1: The camera of the transmitting device collects video data, the video acquisition module obtains video frames from the camera, and sends the video frames to the video encoding module; S2: performing AI scene detection on the video frame to obtain the picture scene of the video frame, and performing region of interest ROI detection on the video frame to obtain the region of interest ROI range of the video frame; S3: Optimizing the original funnel algorithm, setting different bit rates for different video frames according to different video frame scenes, and setting different bit rates for different regions of the video frame according to different ranges of the region of interest ROI of the video frame; S4: After encoding the video frame, send it to the receiving end; S5: The receiving end uses a video decoder to decode the received compressed video data into the original video frame; S6: The decoded video frame is transmitted to a rendering module for video rendering and display.

2. The method for reducing transmission bandwidth and improving transmission stability based on WebRTC according to claim 1 is characterized in that: Before step S1, the method further includes: Setting a signaling service, the signaling service being used to establish and manage a connection between the sending end and the receiving end, the signaling service including session control information, codec information and other control signaling; The signaling service negotiates video codecs, network parameters and other session-related settings by exchanging Session Description Protocol (SDP) information.

3. The method for reducing transmission bandwidth and improving transmission stability based on WebRTC according to claim 1 is characterized in that: In step S2, AI scene detection is performed on the video frame to obtain the picture scene of the video frame, specifically: Using the AI ​​capabilities of the RK3588 chip, the AI ​​model is used to set and judge the scene of the video frame; A preset number of the video frames are set to be identified as the same picture scene, and when the number of the video frames reaches the preset number or a key frame change occurs, the picture scene is re-determined, wherein the picture scene is determined by sending the YUV data of the first key frame among the preset number of the video frames into a scene detection model for inference; Different qualities are set for different picture scenes, and different quantization parameter QP ranges are set in the video coding standard H.264 for the picture scenes including simple scenes, motion scenes, and complex scenes so that different picture scenes have different bit rates, wherein the larger the value of the quantization parameter QP, the lower the bit rate of the picture scene and the worse the picture quality.

4. The method for reducing transmission bandwidth and improving transmission stability based on WebRTC according to claim 1 is characterized in that: In step S2, a region of interest ROI detection is performed on the video frame to obtain a region of interest ROI range of the video frame, specifically: Obtaining YUV data of the first key frame in a group of continuous video frames of the same picture scene and feeding the data into the ROI detection model, and obtaining the ROI range of the region of interest of the video frame by using AI reasoning; By adopting region of interest ROI encoding, the bit rate of the original scene is maintained for the scene within the region of interest ROI, and the bit rate of the non-interested region is reduced.

5. The method for reducing transmission bandwidth and improving transmission stability based on WebRTC according to claim 3 is characterized in that: In step S3, the original funnel algorithm is optimized, different bit rates are set for different video frames according to different video frame scenes, and different bit rates are set for different regions of the video frame according to different ranges of the region of interest ROI of the video frame, specifically: According to the judgment result of the AI ​​scene detection, on the basis of the frame loss strategy of the original funnel algorithm of WebRTC, the frame loss frequency is reduced in turn according to the simple scene, the motion scene, and the complex scene, and the frame rate of the motion scene and the complex scene is increased to improve the fluency; According to the difference between the simple scene, the motion scene and the complex scene, different ranges of quantization parameters QP are set in the video coding standard H.264; The bit rate of the original scene is maintained for the scene within the region of interest ROI, and the bit rate of the non-region of interest is reduced to perform encoding control.

6. The method for reducing transmission bandwidth and improving transmission stability based on WebRTC according to claim 1, characterized in that: In step S4, encoding the video frame further includes: Use the hardware encoding on the RK3588 chip to encode to increase the speed of video encoding.

7. The method for reducing transmission bandwidth and improving transmission stability based on WebRTC according to claim 1, characterized in that: In step S4, after encoding the video frame, the video frame is sent to the receiving end, which also includes: After the video encoding is completed, the encoded data is encapsulated into a real-time transport protocol RTP packet, and the real-time transport protocol RTP packet is sent to the receiving end through the user datagram protocol UDP.

8. A system for reducing transmission bandwidth and improving transmission stability based on WebRTC for executing the method for reducing transmission bandwidth and improving transmission stability based on WebRTC as described in any one of claims 1 to 7, characterized in that: include: A video data acquisition module, used to provide video data to a camera of a transmitting device, the video acquisition module obtains video frames from the camera, and sends the video frames to a video encoding module; A picture scene detection module is used to perform AI scene detection on the video frame to obtain the picture scene of the video frame, and at the same time perform region of interest ROI detection on the video frame to obtain the region of interest ROI range of the video frame; The original algorithm optimization module is used to optimize the original funnel algorithm, set different bit rates for different video frames according to different video frame scenes, and set different bit rates for different regions of the video frame according to different ranges of the region of interest ROI of the video frame; A video data sending module, used for encoding the video frame and sending it to a receiving end; A video data decoding module, used for providing the receiving end with a video decoder to decode the received compressed video data into the original video frame; The video data rendering module is used to provide the decoded video frame to the rendering module for video rendering and display.

9. A computer device comprising a memory and one or more processors, wherein the memory stores computer codes, and when the computer codes are executed by the one or more processors, the one or more processors execute the method according to any one of claims 1 to 7. 10 . A computer-readable storage medium storing a computer code. When the computer code is executed, the method according to claim 1 is executed.

Citation Information

Cited By

  • Panoramic video playing method and playing device

    CN120897129A

  • Remote consultation system and method

    CN121509606A