A video encoding method and system based on video conferencing communication

By selecting the long-term reference frame with the nearest display sequence as the encoding basis in the video conference, replacing the IDR frame logic, the frequent IDR frame problems caused by network exceptions are solved, and the smoothness and encoding performance of the video conference are improved.

CN115174847BActive Publication Date: 2025-07-25YEALINK (XIAMEN) NETWORK TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210852696.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-20
Publication Date
2025-07-25
Estimated Expiration
2042-07-20

AI Technical Summary

Technical Problem

In real-time video conferencing, due to network abnormalities, IDR frames frequently appear, affecting the user experience and causing network bandwidth fluctuations, and the existing technology is difficult to effectively solve.

Method used

By acquiring video frames in real time and selecting several frames as long-term reference frames according to the long-term reference frame strategy, clearing the long-term reference frames of frame dropout when detecting network abnormalities, and selecting the long-term reference frame with the nearest display sequence as the encoding basis, replacing the low-compression rate IDR frame logic, and using the intra-frame refresh strategy to adjust the quantization parameters to reduce the IDR frame frequency.

Benefits of technology

It reduces the frequency of IDR frames, improves the smoothness and user experience of video conferencing, reduces network fluctuations, and improves encoding performance and rate distortion performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115174847B_ABST
    Figure CN115174847B_ABST
Patent Text Reader

Abstract

The present invention discloses a video encoding method and system based on video conferencing communication, including: obtaining multiple video frames corresponding to a video conference in real time, and selecting several video frames as several long-term reference frames corresponding to the video conference according to a preset long-term reference frame strategy; encoding each video frame corresponding to the video conference with reference to all long-term reference frames; when detecting network anomalies, clearing the long-term reference frames with frame loss phenomena, and determining the video frames with mosaic phenomena caused by frame loss as the video frames to be matched, and then selecting, from all long-term reference frames, the long-term reference frame with the closest display order to the video frames to be matched in the video conference as the reference basis for encoding the video frames to be matched. The present invention reduces the frequency of IDR frames and the overall encoding time by selecting long-term reference frames and, when detecting network anomalies, selecting another long-term reference frame nearby as the new encoding reference basis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video coding, and in particular to a video coding method and system based on video conferencing communication. Background Art

[0002] In the information age, people often use videos to record their lives, communicate with others, etc. Based on people's pursuit of video fidelity, smoothness, etc., the frame rate of videos is increasing day by day, and the volume of videos is also increasing accordingly. In order to make high-frame-rate videos suitable for actual storage and transmission, people usually perform video coding on the videos.

[0003] Currently, in real-time video conferencing communication, the network conditions of some participants often affect all participants. Based on this, the traditional video coding method is to periodically send IDR (Instant Decoding Refresh) frames from the encoding end to the decoding end. Among them, the IDR frame is the first I frame (Intra-Picture) of a GOP (Group Of Pictures). For those users with insufficient downstream network bandwidth, it is already difficult to receive an IDR frame. If the IDR frame experiences an abnormality again due to transmission reasons, then the IDR frame will appear multiple times within a short period of time, resulting in a "vicious cycle". This phenomenon will also affect all participants, resulting in a poor user experience for all participants. And because the compression rate of the IDR frame is small, when the IDR frame appears, the bandwidth will experience a fluctuation. Therefore, the frequent appearance of IDRs in a short period of time is also a great test for network transmission. Summary of the Invention

[0004] The present invention provides a video coding method and system based on video conferencing communication, which reduces the frequency of IDR frame appearance in real-time video conferencing and reduces the overall encoding time consumption to improve the smoothness of video conferencing.

[0005] To solve the above technical problems, an embodiment of the present invention provides a video coding method based on video conferencing communication, including:

[0006] Obtaining multiple video frames corresponding to a video conference in real time, and selecting several of the video frames as several long-term reference frames corresponding to the video conference according to a preset long-term reference frame strategy;

[0007] Encoding each of the video frames corresponding to the video conference with reference to all the long-term reference frames;

[0008] When a network anomaly is detected, clear the long-term reference frames with frame loss, and determine the video frames with a screen freeze phenomenon caused by the frame loss as the video frames to be matched. Then, from all the long-term reference frames, select the long-term reference frame with the closest display order to the video frames to be matched in the video conference as the reference basis for encoding the video frames to be matched.

[0009] Implementing the embodiments of the present invention, when a network anomaly occurs, determine and clear the long-term reference frames (LTR) with frame loss, avoiding other video frames referring to the long-term reference frames with frame loss, which may cause a screen freeze phenomenon and affect subsequent encoding. At the same time, use the video frames that have already had a screen freeze phenomenon as the video frames to be matched, and select a long-term reference frame with the closest display order to the video frames to be matched in the video conference from several selected long-term reference frames as the reference basis for encoding the video frames to be matched, rather than re-establishing long-term reference frames according to the preset long-term reference frame strategy, thereby reducing the overall encoding time and improving the smoothness of the video conference. In addition, when frame loss occurs due to poor network bandwidth, using the long-term reference frame selection logic to replace the IDR frame logic with a low compression rate can reduce the frequency of IDR frames, reduce network fluctuations, and improve the experience of participants.

[0010] As a preferred solution, the video encoding method based on video conference communication further includes:

[0011] When a new user enters the video conference, real-time obtain the first I-frame corresponding to the video conference, and obtain a plurality of first sub-frames corresponding to the first I-frame. Then, place each of the first sub-frames in the corresponding P-frame to form a plurality of first Intra Refresh frames; where the first sub-frames and the first Intra Refresh frames correspond one by one;

[0012] According to the preset quantization parameter adjustment rule, adjust the quantization parameters of the refresh area and the area to be refreshed in each of the first Intra Refresh frames, and refresh the refresh area in each of the first Intra Refresh frames according to the preset refresh direction and the adjusted quantization parameters;

[0013] After completing the refresh of the refresh area in all the first Intra Refresh frames, set the first Intra Refresh frames that currently do not have the area to be refreshed as the long-term reference frames corresponding to the first I-frame as the reference basis for encoding the first I-frame.

[0014] Implementing the preferred solution of the embodiments of the present invention, when a new participant appears in a video conference, several first sub-frames corresponding to the current I-frame are respectively placed in the corresponding P-frames to form several first Intra Refresh frames, and the P-frames with high compression rate are used to replace the IDR frames with low compression rate, so as to further reduce the occurrence frequency of IDR frames, thereby improving the overall coding performance and being beneficial to network transmission. Additionally, by using the Intra Refresh strategy and according to the preset quantization parameter adjustment rule, the quantization parameter allocation of the refreshed area and the area to be refreshed in each first Intra Refresh frame is adjusted, so as to achieve the effect of improving the overall rate-distortion performance of the conference.

[0015] As a preferred solution, the video coding method based on video conference communication further includes:

[0016] When there are frame loss phenomena in all the long-term reference frames, all the long-term reference frames with frame loss phenomena are cleared, and the second I-frame corresponding to the video conference is obtained in real time;

[0017] According to the second I-frame, several second sub-frames corresponding to the second I-frame are obtained, and each of the second sub-frames is respectively placed in the corresponding P-frame to form several second Intra Refresh frames; wherein, the second sub-frames and the second Intra Refresh frames correspond one by one;

[0018] According to the preset quantization parameter adjustment rule, the quantization parameters of the refreshed area and the area to be refreshed in each of the second Intra Refresh frames are adjusted, and according to the preset refresh direction and the adjusted quantization parameters, the refreshed area in each of the second Intra Refresh frames is refreshed;

[0019] After finishing the refreshing of the refreshed area in all the second Intra Refresh frames, all the second Intra Refresh frames that have completed the refreshing of the refreshed area are set as several long-term reference frames corresponding to the video conference.

[0020] Implementing the preferred solution of the embodiments of the present invention, when there are frame loss phenomena in all the marked long-term reference frames, the original long-term reference frames are all cleared. At this time, there is no reference frame available. Therefore, all the second Intra Refresh frames that have completed intra-frame refreshing are used as several new long-term reference frames corresponding to the video conference, avoiding the appearance of IDR frames, thereby reducing network fluctuations and improving the participation experience of participants.

[0021] As a preferred solution, the acquisition of the preset refresh direction is specifically as follows:

[0022] Obtain the first mv information corresponding to each macroblock in all the first sub-frames, and in accordance with a preset refresh direction determination rule, combine all the first mv information to determine the refresh direction corresponding to the first Intra Refresh frame; wherein, the first mv information includes the mv abscissa and mv ordinate of each macroblock in the first sub-frame;

[0023] Alternatively, obtain the second mv information corresponding to each macroblock in all the second sub-frames, and in accordance with a preset refresh direction determination rule, combine all the second mv information to determine the refresh direction corresponding to the second Intra Refresh frame; wherein, the second mv information includes the mv abscissa and mv ordinate of each macroblock in the second sub-frame.

[0024] Implementing the preferred solution of the embodiments of the present invention, determine the in-frame refresh direction according to the mv information of all the first sub-frames or the mv information corresponding to all the second sub-frames, thereby improving the refresh efficiency of the corresponding first Intra Refresh frame or second Intra Refresh frame, so as to further improve the overall rate-distortion performance of the conference.

[0025] As a preferred solution, select a number of the video frames in accordance with a preset long-term reference frame strategy to serve as a number of long-term reference frames corresponding to the video conference, specifically:

[0026] Arrange all the video frames in the display order from the earliest to the latest in the video conference, and obtain the first N video frames placed at the bottom according to the arrangement result and mark them as the first long-term reference frames;

[0027] Starting from the last video frame marked as the first long-term reference frame, every M video frames, mark the current video frame as the second long-term reference frame;

[0028] Use all the video frames marked as the first long-term reference frames and all the video frames marked as the second long-term reference frames as a number of the long-term reference frames corresponding to the video conference.

[0029] Implementing the preferred solution of the embodiments of the present invention, first mark N video frames placed at the bottom layer as the first long-term reference frames, so that under any SVCT (Scalable Video Codec - Temporal) structure, the first long-term reference frames can be referred to, avoiding being restricted by cross-layer reference. Additionally, after marking the first long-term reference frames, starting from the last long-term reference frame, perform periodic marking with a period of M on the subsequent video frames, that is, every M video frames, mark the current video frame as the second reference frame, and use all the first long-term reference frames and all the second long-term reference frames as several long-term reference frames corresponding to the current video conference, increasing the marking frequency of the long-term reference frames, preventing the lack of long-term reference frames for reference in a long-term continuous network abnormal environment and affecting the overall coding quality, and by regularly marking the long-term reference frames, it can be avoided that a certain long-term reference frame stays in the DPB (Decoded Picture Buffer) for too long, thus preventing the problem of poor coding quality of the coding frames referring to this long-term reference frame from occurring.

[0030] To solve the same technical problem, an embodiment of the present invention also provides a video coding system based on video conference communication, including:

[0031] A data acquisition module, configured to acquire multiple video frames corresponding to a video conference in real time, and select several of the video frames as several long-term reference frames corresponding to the video conference according to a preset long-term reference frame strategy;

[0032] A video coding module, configured to encode each of the video frames corresponding to the video conference with reference to all the long-term reference frames;

[0033] A first frame dropping processing module, configured to, when detecting network abnormality, clear the long-term reference frames with frame dropping phenomena, determine the video frames with the mosaic phenomena caused by the frame dropping phenomena as the video frames to be matched, and then select, from all the long-term reference frames, the long-term reference frame with the closest display order to the video frames to be matched in the video conference as the reference basis for encoding the video frames to be matched

[0034] As a preferred solution, the video coding system based on video conference communication further includes:

[0035] An intra-frame refresh module, which is used to, when detecting that a new user enters the video conference, obtain the first I-frame corresponding to the video conference in real time, obtain a plurality of first sub-frames corresponding to the first I-frame, and then place each of the first sub-frames into the corresponding P-frame respectively to form a plurality of first Intra Refresh frames; wherein, the first sub-frames and the first Intra Refresh frames are in one-to-one correspondence; according to a preset quantization parameter adjustment rule, adjust the quantization parameters of the refresh area and the area to be refreshed in each of the first Intra Refresh frames, and refresh the refresh area in each of the first Intra Refresh frames according to a preset refresh direction and the adjusted quantization parameters; after completing the refresh of the refresh area in all the first Intra Refresh frames, set the first Intra Refresh frames that do not have the area to be refreshed currently as the long-term reference frames corresponding to the first I-frame, so as to be used as a reference basis for encoding the first I-frame.

[0036] As a preferred solution, the video encoding system based on video conference communication further includes:

[0037] A second frame dropping processing module, which is used to, when there are frame dropping phenomena in all the long-term reference frames, clear all the long-term reference frames with frame dropping phenomena and obtain the second I-frame corresponding to the video conference in real time; according to the second I-frame, obtain a plurality of second sub-frames corresponding to the second I-frame, and place each of the second sub-frames into the corresponding P-frame respectively to form a plurality of second Intra Refresh frames; wherein, the second sub-frames and the second Intra Refresh frames are in one-to-one correspondence; according to a preset quantization parameter adjustment rule, adjust the quantization parameters of the refresh area and the area to be refreshed in each of the second Intra Refresh frames, and refresh the refresh area in each of the second Intra Refresh frames according to a preset refresh direction and the adjusted quantization parameters; after completing the refresh of the refresh area in all the second Intra Refresh frames, set all the second Intra Refresh frames that have completed the refresh of the refresh area as a plurality of the long-term reference frames corresponding to the video conference.

[0038] As a preferred solution, the video encoding system based on video conference communication further includes:

[0039] A refresh direction determination module is configured to obtain the first mv information corresponding to each macroblock in all the first sub-frames, and determine the refresh direction corresponding to the first Intra Refresh frame in combination with all the first mv information according to a preset refresh direction determination rule; or obtain the second mv information corresponding to each macroblock in all the second sub-frames, and determine the refresh direction corresponding to the second Intra Refresh frame in combination with all the second mv information according to a preset refresh direction determination rule; wherein, the first mv information includes the mv abscissa and mv ordinate of each macroblock in the first sub-frame, and the second mv information includes the mv abscissa and mv ordinate of each macroblock in the second sub-frame.

[0040] As a preferred solution, the data acquisition module specifically includes:

[0041] A data acquisition unit for real-time acquisition of multiple video frames corresponding to a video conference;

[0042] A first marking unit for arranging all the video frames in the display order from the earliest to the latest in the video conference, and obtaining the first N video frames placed at the bottom according to the arrangement result and marking them as the first long-term reference frames;

[0043] A second marking unit for starting from the last video frame marked as the first long-term reference frame, and marking the current video frame as the second long-term reference frame every M video frames;

[0044] A setting unit for using all the video frames marked as the first long-term reference frames and all the video frames marked as the second long-term reference frames as several long-term reference frames corresponding to the video conference. Description of the Drawings

[0045] Figure 1 : A schematic flowchart of a video encoding method based on video conference communication provided in Embodiment 1 of the present invention;

[0046] Figure 2 : A schematic diagram of an Intra Refresh refresh principle provided in Embodiment 1 of the present invention;

[0047] Figure 3 : A schematic structural diagram of a video encoding system based on video conference communication provided in Embodiment 1 of the present invention. Detailed Embodiment

[0048] Next, in combination with the accompanying drawings in the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts belong to the scope of protection of the present invention.

[0049] Embodiment 1:

[0050] Please refer to Figure 1 , a video encoding method based on video conferencing communication provided by an embodiment of the present invention. This method includes steps S1 to S3, and the specific steps are as follows:

[0051] Step S1: Real-time obtain multiple video frames corresponding to the video conference, and select several video frames according to a preset long-term reference frame strategy to serve as several long-term reference frames corresponding to the video conference.

[0052] As a preferred solution, step S1 includes steps S101 to S104, and the specific steps are as follows:

[0053] Step S101: Real-time obtain multiple video frames corresponding to the video conference.

[0054] Step S102: Arrange all video frames in the display order from the earliest to the latest in the video conference, and obtain the first N video frames placed at the bottom according to the arrangement result and mark them as the first long-term reference frames.

[0055] Step S103: Starting from the last video frame marked as a long-term reference frame, every M video frames, mark the current video frame as the second long-term reference frame.

[0056] In this embodiment, the marking of long-term reference frames is divided into automatic marking inside the encoder and periodic marking triggered by the network engine layer. Among them, automatic marking inside the encoder means that the number of the first long-term reference frames to be marked is preset as N, and in the display order from the earliest to the latest in the video conference, each video frame with tid = 0, that is, the video frames located at the bottom layer, is sequentially marked as the first long-term reference frame until the number of the marked first long-term reference frames reaches the preset number N, then the current automatic marking of the encoder ends. Periodic marking triggered by the network engine layer means that the network engine layer uses the end signal of the automatic marking of the encoder as the start signal of the periodic marking, that is, starting from the last video frame marked as the first long-term reference frame, every M video frames, determine a video frame to be marked as the second long-term reference frame, and the network engine layer sends a long-term reference frame instruction to the encoder, and the current video frame is marked as the second long-term reference frame inside the encoder.

[0057] As an example, N = 4, M = 8, and in multiple video frames corresponding to a video conference, the frame numbers of the video frames at the bottom layer are successively 1, 5, 9, 13, …… Therefore, the video frames with frame numbers 1, 5, 9, and 13 are marked as the first long-term reference frames, and the internal automatic marking of the encoder ends. Then, starting from the video frame with frame number 13, every 8th video frame is marked as a second long-term reference frame, that is, the video frames with frame numbers 21, 29, 37, …… are marked as the second long-term reference frames.

[0058] It should be noted that if an abnormal situation occurs during the marking process, for example, during the process of marking the video frame with frame number 37 as the second long-term reference frame, the internal encoder does not take effect. Then, it is necessary to abandon marking the current video frame and mark the next frame. Specifically, as follows: during the periodic marking process triggered at the network engine layer, the network engine layer determines that the video frame with frame number 37 is the video frame to be marked. Then, the network engine layer sends a long-term reference frame instruction to the encoder. The encoder internally marks the video frame and feeds back the corresponding interface return value to the network engine layer. The network engine layer determines whether the marking of the video frame with frame number 37 is successful based on this interface return value. If the marking is successful, the periodic marking continues. If the marking fails, the periodic marking continues with the next frame (i.e., the video frame with frame number 38). At this time, if the video frame with frame number 38 is marked successfully, the marking objects for subsequent periodic marking are the video frames with frame numbers 38, 46, 54, ……

[0059] Step S104, use all the first long-term reference frames and all the second long-term reference frames as several long-term reference frames corresponding to the video conference.

[0060] Step S2, encode each video frame corresponding to the video conference with reference to all the long-term reference frames.

[0061] Step S3, when a network anomaly is detected, clear the long-term reference frames with frame loss phenomena, and determine the video frames with the mosaic phenomenon caused by the frame loss phenomenon as the video frames to be matched. Then, select the long-term reference frame that is closest to the display order of the video frames to be matched in the video conference from all the long-term reference frames as the reference basis for encoding the video frames to be matched.

[0062] In this embodiment, if the network engine layer determines that there is a frame loss in the current network (i.e., network anomaly), it sends a corresponding frame loss handling instruction to the encoder through the network engine layer to inform the encoder of the information of "the long-term reference frame number with frame loss phenomenon and the long-term reference frame number that the video frame (i.e., the video frame to be matched) causing the screen freeze phenomenon due to the frame loss phenomenon needs to refer to". After receiving the frame loss handling instruction, the encoder determines and clears the long-term reference frame with frame loss phenomenon according to the above information, and selects the long-term reference frame with the closest display order to the video frame to be matched in the video conference as the reference basis for encoding the video frame to be matched.

[0063] It should be noted that the determination process of whether there is a frame loss in the current network by the network engine layer is specifically as follows: The network engine layer checks the number of frames and the frame numbers of all long-term reference frames at the sending end and the receiving end respectively. If the number of frames or the frame numbers are inconsistent, it is determined that there is a frame loss in the current network.

[0064] As a preferred solution, a video encoding method based on video conference communication further includes steps S4 to S6, and the specific steps are as follows:

[0065] Step S4, when it is detected that a new user enters the video conference, the first I-frame corresponding to the video conference is obtained in real time, and a plurality of first sub-frames corresponding to the first I-frame are obtained, and then each first sub-frame is respectively placed in the corresponding P-frame to form a plurality of first Intra Refresh frames.

[0066] As an example, when it is detected that a new user enters the video conference, the network engine layer sends a signal that the current video frame needs Intra Refresh (intra-frame refresh) to the encoder. When the encoder receives the above signal, it splits the I-frame coordinated and generated by the network engine layer at the decoding end to obtain the corresponding four first sub-frames, and places the above four first sub-frames in the corresponding P-frames respectively to form Figure 2 the four first Intra Refresh frames Frame1, Frame2, Frame3 and Frame4 for reference. Among them, the shaded area is the refresh area in the refresh state, the black area is the to-be-refreshed area in the to-be-refreshed state, and the white area is the refreshed area that has been refreshed.

[0067] Step S5, according to the preset quantization parameter adjustment rule, adjust the quantization parameters of the refresh area and the to-be-refreshed area in each first Intra Refresh frame, and refresh the refresh area in each first Intra Refresh frame according to the preset refresh direction and the adjusted quantization parameters.

[0068] In this embodiment, in order to improve the overall rate-distortion performance, it is necessary to adjust the quantization parameters of the refreshed regions and the regions to be refreshed in each first Intra Refresh frame, so as to increase the quantization parameters of the macroblocks in the refreshed regions and reduce the quantization parameters of the macroblocks in the regions to be refreshed. The adjustment process includes steps S51 to S53, which are specifically as follows:

[0069] Step S51, determine the current quantization parameter QP0 of each macroblock.

[0070] Step S52, referring to equations (1) and (2), calculate the quantization parameter reduction QP1 corresponding to each macroblock, and reduce QP1 for the quantization parameters of each macroblock in all refreshed regions.

[0071] QP1 = α * log2 Var - β + QP offset,refresh (1)

[0073]

[0074] where α and β are the first coefficient and the second coefficient determined based on multiple experiments respectively, Var is the variance of the current macroblock, mb refresh is the total number of macroblocks in the refreshed region, mb is the total number of macroblocks in the current frame, and QP offset,refresh is the quantization parameter offset of the macroblocks in the refreshed region of the current frame.

[0075] Additionally, in this embodiment, α = 1.0397 and β = 14.427.

[0076] Step S53, referring to equations (3) and (4), calculate the quantization parameter increase QP2 corresponding to each macroblock, and increase QP2 for the quantization parameters of each macroblock in all regions to be refreshed.

[0077] QP2 = α * log2 Var - β + QP offset,unrefresh (3)

[0079]

[0080] where α and β are the first coefficient and the second coefficient determined based on multiple experiments respectively, Var is the variance of the current macroblock, mb refresh is the total number of macroblocks in the refreshed region, mb is the total number of macroblocks in the current frame, and QP offset,unrefresh is the quantization parameter offset of the macroblocks in the region to be refreshed of the current frame.

[0081] Additionally, in this embodiment, α = 1.0397 and β = 14.427.

[0082] As a preferred solution, the process for obtaining the refresh direction corresponding to the first Intra Refresh frame is step S54, which is specifically as follows:

[0083] Step S54: Obtain the first mv information corresponding to each macroblock in all the first sub - frames, and determine the refresh direction corresponding to the first Intra Refresh frame in combination with all the first mv information according to a preset refresh direction determination rule; wherein, the first mv information includes the mv abscissa and mv ordinate of each macroblock in the first sub - frame.

[0084] In this embodiment, record the first mv information corresponding to each macroblock in the first sub - frame, where the first mv information includes the mv abscissa and mv ordinate of each macroblock in the first sub - frame. According to all the first mv information, obtain the number N0 of macroblocks whose mv abscissa and mv ordinate are not all 0, the number N1 of macroblocks whose mv abscissa is 0 and mv ordinate is greater than 0, the number N2 of macroblocks whose mv abscissa is 0 and mv ordinate is less than 0, and the number M1 of macroblocks whose mv abscissa is greater than 0. If the ratio of N1 to N0 is greater than the first preset threshold, the refresh direction of the current first sub - frame is vertically upward; if the ratio of N2 to N0 is greater than the first preset threshold, the refresh direction of the current first sub - frame is vertically downward; if both the ratio of N1 to N0 and the ratio of N2 to N0 are less than or equal to the first preset threshold, then judge the horizontal refresh direction of the current first sub - frame according to the size relationship between the ratio of M1 to N0 and the second preset threshold. Among them, if the ratio of M1 to N0 is greater than the second preset threshold, the refresh direction of the current first sub - frame is horizontally to the left; if the ratio of M1 to N0 is less than or equal to the second preset threshold, the refresh direction of the current first sub - frame is horizontally to the right.

[0085] As an example, the first preset threshold is 0.6 and the second preset threshold is 0.5.

[0086] Step S6: After completing the refresh of the refresh regions in all the first Intra Refresh frames, set the first Intra Refresh frame that currently has no region to be refreshed as the long - term reference frame corresponding to the first I - frame to be used as the reference basis for encoding the first I - frame.

[0087] As a preferred solution, a video coding method based on video conferencing communication further includes steps S7 to S10, and the specific steps of each are as follows:

[0088] Step S7, when there are frame loss phenomena in all long-term reference frames, clear all long-term reference frames with frame loss phenomena, and obtain the second I-frame corresponding to the video conference in real time. At this time, since all the original long-term reference frames are cleared and there are no reference frames to select, through steps S8 to S10, several new long-term reference frames are obtained to avoid the appearance of IDR frames, thereby reducing network fluctuations and enhancing the participation experience of participants.

[0089] Step S8, according to the second I-frame, obtain several second sub-frames corresponding to the second I-frame, and place each second sub-frame in the corresponding P-frame respectively to form several second Intra Refresh frames.

[0090] Step S9, according to the preset quantization parameter adjustment rule, adjust the quantization parameters of the refreshed area and the area to be refreshed in each second Intra Refresh frame, and refresh the refreshed area in each second Intra Refresh frame according to the preset refresh direction and the adjusted quantization parameters.

[0091] Step S10, after completing the refresh of the refreshed areas in all second Intra Refresh frames, set all the second Intra Refresh frames that have completed the refresh of the refreshed areas as several long-term reference frames corresponding to the video conference.

[0092] In this embodiment, for the specific process of refreshing the refreshed areas in all second Intra Refresh frames, reference can be made to the process of refreshing the refreshed areas in the first Intra Refresh frame in the foregoing step S5, which will not be elaborated here.

[0093] Please refer to Figure 3 , which is a schematic structural diagram of a video coding system based on video conference communication provided by an embodiment of the present invention. The video coding system based on video conference communication includes a data acquisition module 1, a video coding module 2, and a first frame loss processing module 3. The specific functions of each module are as follows:

[0094] The data acquisition module 1 is configured to obtain multiple video frames corresponding to the video conference in real time, and select several video frames according to a preset long-term reference frame strategy as several long-term reference frames corresponding to the video conference;

[0095] The video coding module 2 is configured to encode each video frame corresponding to the video conference with reference to all long-term reference frames;

[0096] The first frame dropping processing module 3 is used to, when detecting network anomalies, clear the long-term reference frames with frame dropping phenomena, and determine the video frames with mosaic phenomena caused by frame dropping phenomena as the video frames to be matched. Then, from all the long-term reference frames, select the long-term reference frame with the closest display order to the video frames to be matched in the video conference as the reference basis for encoding the video frames to be matched.

[0097] As a preferred solution, please refer to Figure 3 , a video encoding system based on video conference communication, further includes an intra-frame refresh module 4, specifically as follows:

[0098] The intra-frame refresh module 4 is used to, when detecting that a new user enters the video conference, obtain the first I-frame corresponding to the video conference in real time, and obtain a plurality of first sub-frames corresponding to the first I-frame. Then, place each first sub-frame in the corresponding P-frame respectively to form a plurality of first Intra Refresh frames; adjust the quantization parameters of the refreshed area and the area to be refreshed in each first Intra Refresh frame according to the preset quantization parameter adjustment rule, and refresh the refreshed area in each first Intra Refresh frame according to the preset refresh direction and the adjusted quantization parameters; after completing the refresh of the refreshed areas in all the first Intra Refresh frames, set the first Intra Refresh frames that do not have areas to be refreshed currently as the long-term reference frames corresponding to the first I-frame as the reference basis for encoding the first I-frame.

[0099] As a preferred solution, please refer to Figure 3 , a video encoding system based on video conference communication, further includes a second frame dropping processing module 5, specifically as follows:

[0100] The second frame dropping processing module 5 is used to, when all the long-term reference frames have frame dropping phenomena, clear all the long-term reference frames with frame dropping phenomena, and obtain the second I-frame corresponding to the video conference in real time; obtain a plurality of second sub-frames corresponding to the second I-frame according to the second I-frame, and place each second sub-frame in the corresponding P-frame respectively to form a plurality of second Intra Refresh frames; adjust the quantization parameters of the refreshed area and the area to be refreshed in each second Intra Refresh frame according to the preset quantization parameter adjustment rule, and refresh the refreshed area in each second Intra Refresh frame according to the preset refresh direction and the adjusted quantization parameters; after completing the refresh of the refreshed areas in all the second Intra Refresh frames, set all the second Intra Refresh frames that have completed the refresh of the refreshed areas as a plurality of long-term reference frames corresponding to the video conference.

[0101] As a preferred solution, please refer toFigure 3 , a video coding system based on video conferencing communication, further includes a refresh direction determination module 6, specifically as follows:

[0102] The refresh direction determination module 6 is configured to obtain the first mv information corresponding to each macroblock in all the first sub-frames, and determine the refresh direction corresponding to the first Intra Refresh frame by combining all the first mv information according to a preset refresh direction determination rule; or, obtain the second mv information corresponding to each macroblock in all the second sub-frames, and determine the refresh direction corresponding to the second Intra Refresh frame by combining all the second mv information according to a preset refresh direction determination rule; wherein, the first mv information includes the mv abscissa and mv ordinate of each macroblock in the first sub-frame, and the second mv information includes the mv abscissa and mv ordinate of each macroblock in the second sub-frame.

[0103] As a preferred solution, the data acquisition module 1 specifically includes a data acquisition unit 11, a first marking unit 12, a second marking unit 13, and a setting unit 14, and each unit is specifically as follows:

[0104] The data acquisition unit 11 is configured to acquire multiple video frames corresponding to the video conference in real time;

[0105] The first marking unit 12 is configured to arrange all the video frames in the display order from the earliest to the latest in the video conference, and obtain the first N video frames placed at the bottom according to the arrangement result and mark them as the first long-term reference frames;

[0106] The second marking unit 13 is configured to start from the last video frame marked as a long-term reference frame, and mark the current video frame as the second long-term reference frame every M video frames;

[0107] The setting unit 14 is configured to use all the first long-term reference frames and all the second long-term reference frames as several long-term reference frames corresponding to the video conference.

[0108] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working process of the above-described system can refer to the corresponding process in the foregoing method embodiment, and will not be repeated here.

[0109] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:

[0110] The present invention provides a video encoding method and system based on video conferencing communication. When a network anomaly occurs, long-term reference frames with frame loss phenomena are cleared to prevent other video frames from referring to these long-term reference frames with frame loss phenomena, which may cause a mosaic phenomenon and affect subsequent encoding. Meanwhile, from several original long-term reference frames, a long-term reference frame that is closest to the display order of the to-be-matched video frame in the video conference is selected as the reference basis for encoding video frames that have already exhibited a mosaic phenomenon, rather than re-establishing long-term reference frames according to a preset long-term reference frame strategy, thereby reducing the overall encoding time consumption and improving the smoothness of the video conference. In addition, when frame loss occurs due to poor network bandwidth conditions, the long-term reference frame selection logic is used to replace the low compression rate IDR frame logic, which can reduce the frequency of IDR frame occurrences, reduce network fluctuations, and improve the experience of participants.

[0111] Furthermore, when a new participant appears, in-frame refreshing of the current I-frame is performed in regions, and high compression rate P-frames are used to replace low compression rate IDR frames, thereby reducing the frequency of IDR frame occurrences, improving the overall encoding performance, and facilitating network transmission. Additionally, using the in-frame refreshing strategy, according to a preset quantization parameter adjustment rule, the quantization parameter allocation of the refreshed regions and the regions to be refreshed in each first IntraRefresh frame is adjusted, thereby achieving the effect of improving the overall rate-distortion performance of the conference.

[0112] The specific embodiments described above further elaborate on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. It is particularly pointed out that for those skilled in the art, any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A video encoding method based on video conferencing communication, characterized in that, Including: Obtaining multiple video frames corresponding to a video conference in real time, and selecting several of the video frames as several long-term reference frames corresponding to the video conference according to a preset long-term reference frame strategy; Encoding each of the video frames corresponding to the video conference with reference to all the long-term reference frames; When a network anomaly is detected, clearing the long-term reference frames with frame loss phenomena, determining the video frames with screen freeze phenomena caused by the frame loss phenomena as video frames to be matched, and then selecting, from all the long-term reference frames, the long-term reference frame with the closest display order to the video frames to be matched in the video conference as the reference basis for encoding the video frames to be matched; When it is detected that a new user enters the video conference, obtaining a first I-frame corresponding to the video conference in real time, and obtaining several first sub-frames corresponding to the first I-frame, and then respectively placing each of the first sub-frames in a corresponding P-frame to form several first Intra Refresh frames; wherein, the first sub-frames and the first Intra Refresh frames correspond one by one; Adjusting the quantization parameters of the refreshed area and the area to be refreshed in each of the first Intra Refresh frames according to a preset quantization parameter adjustment rule, and refreshing the refreshed area in each of the first Intra Refresh frames according to a preset refresh direction and the adjusted quantization parameters; After completing the refreshing of the refreshed area in all the first Intra Refresh frames, setting the first Intra Refresh frames that currently do not have the area to be refreshed as the long-term reference frames corresponding to the first I-frame as the reference basis for encoding the first I-frame.

2. The video encoding method based on video conferencing communication according to claim 1, wherein, Further including: When all the long-term reference frames have frame loss phenomena, clearing all the long-term reference frames with frame loss phenomena, and obtaining a second I-frame corresponding to the video conference in real time; Obtaining several second sub-frames corresponding to the second I-frame according to the second I-frame, and respectively placing each of the second sub-frames in a corresponding P-frame to form several second Intra Refresh frames; wherein, the second sub-frames and the second Intra Refresh frames correspond one by one; Adjusting the quantization parameters of the refreshed area and the area to be refreshed in each of the second Intra Refresh frames according to a preset quantization parameter adjustment rule, and refreshing the refreshed area in each of the second Intra Refresh frames according to a preset refresh direction and the adjusted quantization parameters; After completing the refreshing of the refreshed area in all the second Intra Refresh frames, setting all the second Intra Refresh frames that have completed the refreshing of the refreshed area as several long-term reference frames corresponding to the video conference.

3. The video encoding method based on video conference communication according to claim 2, wherein, The obtaining of the preset refresh direction is specifically as follows: Obtain the first mv information corresponding to each macroblock in all the first sub-frames, and determine the refresh direction corresponding to the first Intra Refresh frame by combining all the first mv information according to a preset refresh direction determination rule; wherein, the first mv information includes the mv abscissa and mv ordinate of each macroblock in the first sub-frame. Alternatively, obtain the second mv information corresponding to each macroblock in all the second sub-frames, and determine the refresh direction corresponding to the second Intra Refresh frame by combining all the second mv information according to a preset refresh direction determination rule; wherein, the second mv information includes the mv abscissa and mv ordinate of each macroblock in the second sub-frame.

4. A video encoding method based on video conferencing communication according to claim 1, characterized in that, The step of selecting several of the video frames as several long-term reference frames corresponding to the video conference according to a preset long-term reference frame strategy is specifically as follows: Arrange all the video frames in the display order from the earliest to the latest in the video conference, and obtain the first N video frames placed at the bottom according to the arrangement result and mark them as the first long-term reference frames; Starting from the last video frame marked as the first long-term reference frame, every M video frames, mark the current video frame as the second long-term reference frame; Use all the video frames marked as the first long-term reference frames and all the video frames marked as the second long-term reference frames as several long-term reference frames corresponding to the video conference.

5. A video coding system based on video conferencing communication, characterized in that, It includes: A data acquisition module, configured to acquire multiple video frames corresponding to a video conference in real time, and select several of the video frames as several long-term reference frames corresponding to the video conference according to a preset long-term reference frame strategy; A video encoding module, configured to encode each of the video frames corresponding to the video conference with reference to all the long-term reference frames; A first frame dropping processing module, configured to, when detecting a network anomaly, clear the long-term reference frames with frame dropping phenomena, and determine the video frames with a mosaic phenomenon caused by the frame dropping phenomenon as the video frames to be matched, and then select, from all the long-term reference frames, the long-term reference frame with the closest display order to the video frames to be matched in the video conference as the reference basis for encoding the video frames to be matched. An intra refresh module, configured to, when detecting that a new user enters the video conference, obtain the first I-frame corresponding to the video conference in real time, obtain a plurality of first sub-frames corresponding to the first I-frame, and then place each of the first sub-frames into a corresponding P-frame respectively to form a plurality of first Intra Refresh frames; wherein, the first sub-frames and the first Intra Refresh frames are in one-to-one correspondence; adjust the quantization parameters of the refresh region and the region to be refreshed in each of the first Intra Refresh frames according to a preset quantization parameter adjustment rule, and refresh the refresh region in each of the first Intra Refresh frames according to a preset refresh direction and the adjusted quantization parameters; after completing the refresh of the refresh regions in all the first Intra Refresh frames, set the first Intra Refresh frames that do not have the region to be refreshed currently as the long-term reference frames corresponding to the first I-frame to be used as a reference basis for encoding the first I-frame.

6. The video encoding system based on video conferencing communication according to claim 5, characterized in that, Further comprising: A second frame dropping processing module, configured to, when there are frame dropping phenomena in all the long-term reference frames, clear the long-term reference frames with frame dropping phenomena and obtain the second I-frame corresponding to the video conference in real time; according to the second I-frame, obtain a plurality of second sub-frames corresponding to the second I-frame, and place each of the second sub-frames into a corresponding P-frame respectively to form a plurality of second Intra Refresh frames; wherein, the second sub-frames and the second Intra Refresh frames are in one-to-one correspondence; adjust the quantization parameters of the refresh region and the region to be refreshed in each of the second Intra Refresh frames according to a preset quantization parameter adjustment rule, and refresh the refresh region in each of the second Intra Refresh frames according to a preset refresh direction and the adjusted quantization parameters; after completing the refresh of the refresh regions in all the second Intra Refresh frames, set all the second Intra Refresh frames that have completed the refresh of the refresh region as a plurality of the long-term reference frames corresponding to the video conference.

7. A video coding system based on video conferencing communication according to claim 6, characterized in that, Further comprising: A refresh direction determination module, configured to obtain the first mv information corresponding to each macroblock in all the first sub-frames, and determine the refresh direction corresponding to the first Intra Refresh frame by combining all the first mv information according to a preset refresh direction determination rule; or, obtain the second mv information corresponding to each macroblock in all the second sub-frames, and determine the refresh direction corresponding to the second Intra Refresh frame by combining all the second mv information according to a preset refresh direction determination rule; wherein, the first mv information includes the mv abscissa and mv ordinate of each macroblock in the first sub-frame, and the second mv information includes the mv abscissa and mv ordinate of each macroblock in the second sub-frame.

8. A video encoding system based on video conferencing communication according to claim 5, characterized in that, The data acquisition module specifically includes: A data acquisition unit, configured to acquire multiple video frames corresponding to a video conference in real time; A first marking unit, configured to arrange all the video frames in the display order from the front to the back in the video conference, and acquire the first N video frames placed at the bottom according to the arrangement result and mark them as first long-term reference frames; A second marking unit, configured to start from the last video frame marked as the first long-term reference frame, and mark the current video frame as a second long-term reference frame every M video frames; A setting unit, configured to use all the video frames marked as the first long-term reference frames and all the video frames marked as the second long-term reference frames as several long-term reference frames corresponding to the video conference.

Citation Information

Patent Citations

  • Feedback based reference frame selection for video coding

    CN103430538A

  • Video picture frame sending method and device and video picture frame receiving method and device

    CN106713913A