Video processing method and related device
Through MEMC processing and B-frame encoding on the MCU side, intermediate frames are generated and adjacent frames are inserted, the problems of code stream increase and delay in video conferencing are solved, and efficient 60FPS video playback is achieved.
Patent Information
- Application Number
- CN202410029579.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-08
- Publication Date
- 2025-07-08
AI Technical Summary
The prior art is difficult to realize 60FPS video playback in video conferencing, resulting in an increase in video code stream, an increase in bandwidth demand, affecting transmission effect, and the delay of frame insertion processing is large, which cannot meet the delay requirements of video conferencing.
Motion estimation and motion compensation (MEMC) processing are used on the multi-point control unit (MCU) side to generate intermediate frames and encode them in B-frame format, insert them between adjacent frames, reduce the increase in code streams, and calculate the optical flow through optical flow method or deep learning algorithm to improve accuracy.
It effectively reduces the increase in code streams, reduces processing delays, improves transmission efficiency and video conferencing quality, and meets the playback needs of high-frame-rate videos.
Smart Images

Figure CN120281938A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of video processing, and in particular, to a video processing method and related devices. Background Art
[0002] With the development of Internet technology, video conferencing has been more and more widely used. Video conferencing enables people located at two or more locations to have a face-to-face conversation through a network and video conferencing system software. Most of the current video conferencing frame rates are 30 frames per second (FPS). However, on large displays, such as displays over 65 inches, a frame rate of 60 FPS will bring a smoother effect. For example, when playing a video of people walking or waving on a large display, there will be a feeling of stuttering at a frame rate of 30 FPS, while it is very smooth at 60 FPS.
[0003] Currently, few devices can support the acquisition of 60 FPS video images, and the acquisition of 60 FPS video has high requirements for ambient brightness. Therefore, it is not easy to implement a video conference at 60 FPS from the sender side. To achieve the playback of a 60 FPS video conference, the common practice in the industry is to perform frame doubling processing on the video image at the receiver side, interpolating the video image from 30 FPS to 60 FPS. However, after interpolation, the video bitstream will increase significantly, which will in turn lead to an increase in bandwidth requirements and affect the transmission effect. Summary of the Invention
[0004] This application provides a video processing method and related devices, which can reduce the increase in bitstream when performing frame interpolation on a video.
[0005] In a first aspect of this application, a video processing method is provided, which can be applied to a multi-point control unit (MCU). The method includes: obtaining a first video stream; performing motion estimation and motion compensation processing on two adjacent frames of images in the first video stream to obtain an intermediate frame image, where the encoding formats of the two frames of images are respectively an intra picture (I) frame and a predictive picture (P) frame, or both are P frames, and the intermediate frame image is encoded in the format of a bidirectionally predicted picture (B) frame; inserting a target number of intermediate frame images into the two frames of images based on the frame rate required by the receiving device to obtain a second video stream.
[0006] The MCU performs MEMC processing on two adjacent frames in the first video stream to obtain corresponding intermediate frame images. Among them, the encoding formats of the two adjacent frames are an I frame and a P frame respectively, or both are P frames, and the obtained intermediate frame images are encoded in the format of B frames. The inserted intermediate frame images can be 1 frame, or two frames, or three frames. That is to say, the encoding format of the video stream after inserting intermediate frames is IBPBP, IBBPBBP, or IBBBPBBBP. For the convenience of description, in this application, IBPBP is used to represent the encoding format of the video stream, and the number of inserted intermediate frames is not limited.
[0007] The MCU performs frame insertion processing according to the frame rate required by the receiving device, and inserts the corresponding (i.e., the intermediate frame obtained based on these two frames) target number of intermediate frames between every two adjacent frames, so as to obtain a second video stream that meets the requirements of the receiving device and has a higher frame rate. The MCU then sends the second video stream to the receiving device, and the receiving device can decode and play it.
[0008] Since the inserted intermediate frame is a B frame and not a reference frame, it does not affect decoding. Therefore, if the intermediate frame is lost during the transmission to the receiving device, retransmission can be not performed, thereby improving the transmission efficiency. If the adjacent reference frame is lost, retransmission is required to avoid affecting the decoding of the receiving device. Of course, retransmission of the lost intermediate frame can also be performed, and specific details are not limited here.
[0009] In the first aspect of this application, the intermediate frame image is obtained through MEMC processing, and the intermediate frame image is encoded in the format of B frames. Since both MEMC and B frames are calculated based on the difference between the previous and next two frames, and the calculation methods are similar, only a small amount of bitstream will be increased after the frame insertion processing. Compared with the current IPPPP encoding format (i.e., the intermediate frame is a P frame) and the deep learning frame doubling scheme used in video conferencing, the increase in bitstream can be greatly reduced, and the occupation of network bandwidth can be reduced.
[0010] Moreover, when performing MEMC processing, only the adjacent previous and next two frames are referenced, which can greatly reduce the processing delay, thereby meeting the requirements of the video conferencing scenario for delay and ensuring the quality of the video conferencing.
[0011] In a possible implementation manner of the first aspect, the above steps: performing motion estimation and motion compensation processing based on two adjacent frames in the first video stream, include: calculating the optical flow of the two frames; performing motion estimation and motion compensation processing on the two frames based on the optical flow.
[0012] In this possible implementation, MEMC processing is performed on two adjacent frames of images through the optical flow method. Optical flow is the movement of an object between consecutive frames of a sequence, which is caused by the relative movement between the object and the camera. The optical flow method is a method that uses the changes in pixels in the time domain of an image and the correlation between adjacent frames to find the corresponding relationships existing between two adjacent frames, thereby calculating the motion information of the object between adjacent frames. Specifically, first calculate the optical flow of two adjacent frames of images, perform motion estimation based on the optical flow result to obtain a motion vector. Then perform motion compensation processing based on the motion vector to obtain an intermediate frame image. The optical flow-guided MEMC algorithm has a higher computing power requirement than traditional MEMC algorithms, but is far lower than the computing power requirement for frame interpolation through deep learning, and can balance the effects and performance of frame interpolation.
[0013] In a possible implementation of the first aspect, the above step of calculating the optical flow of two frames of images includes: calculating the optical flow through the dense inverse search optical flow method. In this possible implementation, the optical flow of two adjacent frames of images is calculated through the dense inverse search optical flow method. The dense inverse search optical flow method performs a reverse search for correlation at the block level, enabling gradient calculation once and multiple reverse searches for use without re-initializing gradient calculation each time, saving a large amount of computation and thus improving performance.
[0014] In a possible implementation of the first aspect, the above step of calculating the optical flow of two frames of images includes: calculating the optical flow through a deep learning algorithm. In this possible implementation, the optical flow of two adjacent frames of images is calculated through deep learning, for example, through a convolutional neural network. Calculating the optical flow through a deep learning algorithm has a relatively low computing power requirement.
[0015] In a possible implementation of the first aspect, the above step of obtaining the first video stream includes: receiving the third video stream and the fourth video stream; merging the third video stream and the fourth video stream into the first video stream.
[0016] In this possible implementation, there can be multiple sending-end devices, and each sending-end device sends one video stream. The first video stream is the video stream obtained after the MCU merges multiple video streams.
[0017] In addition, the first video stream can also be a single video stream, and the MCU performs frame doubling processing after receiving the single video stream.
[0018] In this possible implementation, there can be multiple sending-end devices, which expands the application scenarios.
[0019] In a possible implementation of the first aspect, before the above step of merging the third video stream and the fourth video stream into the first video stream, the method further includes: adjusting the third video stream and the fourth video stream to the same frame rate.
[0020] The video capture capabilities of the sending devices may vary. For example, most devices have a capture capability of 30 FPS, so videos with a frame rate of 30 FPS are sent. However, some devices can achieve a capture capability of 60 FPS, and the frame rate of the sent videos is 60 FPS. Or, to save the transmission bitstream, the sending device only sends videos with a frame rate of 15 FPS. When the MCU receives multiple video streams with different frame rates, these video streams need to be adjusted to the same frame rate before they can be merged into one video stream. Specifically, the third video stream and the fourth video stream can be interpolated to a same frame rate. Exemplarily, if the frame rate of the third video stream is 15 FPS and the frame rate of the fourth video stream is 30 FPS, the third video stream is double-framed to 30 FPS. Or, both the third video stream and the fourth video stream are double-framed to 60 FPS, which can be specifically determined according to the frame rate required by the receiving device. The method of interpolating the third video stream and the fourth video stream can be the method provided in this application. In this case, it is equivalent to the situation where the first video stream is a single video stream. The method of interpolating the third video stream and the fourth video stream can also be a commonly used interpolation method currently, such as frame sampling or frame blending for interpolation. Frame sampling means directly copying the previous and next frames to obtain the intermediate frame, and frame blending means blending the previous and next two frames to obtain the intermediate frame.
[0021] In addition to interpolating the third video stream and the fourth video stream, it is also possible to reduce the frame rate of the two video streams to the same frequency. For example, if the frame rate of the third video stream is 15 FPS and the frame rate of the fourth video stream is 30 FPS, the fourth video stream is reduced to 15 FPS, so that the third video stream and the fourth video stream can be merged.
[0022] In this possible implementation, the multiple video streams received by the MCU can have different frame rates, which expands the application scenarios.
[0023] In a possible implementation of the first aspect, the receiving device includes a first receiving device and a second receiving device. Among them, the frame rate required by the first receiving device is the first frame rate, and the frame rate required by the second receiving device is the second frame rate. The above step: inserting a target number of intermediate frame images into two frame images based on the frame rate required by the receiving device to obtain a second video stream, includes: inserting a first number of intermediate frame images into two frame images based on the first frame rate to obtain a fifth video stream; inserting a second number of intermediate frame images into two frame images based on the second frame rate to obtain a sixth video stream.
[0024] In this possible implementation, the receiving end includes a first receiving device and a second receiving device, and the required video frame rates of the two receiving devices are different. In this case, the MCU inserts a corresponding first number of intermediate frame images into every two adjacent images, so as to obtain a fifth video stream that meets the frame rate requirement of the first receiving device. Similarly, the MCU inserts a corresponding second number of intermediate frame images into every two adjacent images, so as to obtain a sixth video stream that meets the frame rate requirement of the second receiving device. Exemplarily, if the frame rate of the first video stream is 30 FPS, the required frame rate of the first receiving device is 60 FPS, and the required frame rate of the second receiving device is 90 FPS, then the MCU inserts 1 intermediate frame into every two adjacent frames of the first video stream to obtain a fifth video stream with 60 FPS. And inserts two intermediate frames into every two adjacent frames of the first video stream to obtain a sixth video stream with 90 FPS.
[0025] It can be understood that, in addition to the first receiving device and the second receiving device, there may be more receiving devices, and the required frame rates of the multiple receiving devices may also be the same. Specific details are not limited here.
[0026] In this possible implementation, there may be multiple receiving devices, and the required frame rates of the multiple receiving devices may be different, which greatly expands the application scenarios.
[0027] In a possible implementation of the first aspect, the frame rate of the first video stream is 30 frames per second, and the required frame rate of the receiving device is 60 frames per second.
[0028] In the current video conferencing scenario, the acquisition capability of the sending device is mostly 30 FPS. However, on a large display screen, a frame rate of 60 FPS will bring a smoother display effect. In this possible implementation, it is to perform frame interpolation on the video with a frame rate of 30 FPS to 60 FPS, which has more practical application significance.
[0029] The second aspect of this application provides a multi-point control unit, including an acquisition module, a processing module, and an insertion module. The acquisition module is used to acquire the first video stream; the processing module is used to perform motion estimation and motion compensation processing on two adjacent images in the first video stream to obtain intermediate frame images. The encoding formats of the two images are respectively an I frame and a P frame, or both are P frames, and the intermediate frame images are encoded in the format of a B frame; the insertion module is used to insert a target number of intermediate frame images into the two images based on the frame rate required by the receiving device to obtain a second video stream.
[0030] In a possible implementation of the second aspect, the processing module is specifically used to calculate the optical flow of the two images; and perform motion estimation and motion compensation processing on the two images based on the optical flow.
[0031] In a possible implementation of the second aspect, the processing module is specifically configured to calculate the optical flow by the dense inverse search optical flow method.
[0032] In a possible implementation of the second aspect, the processing module is specifically configured to calculate the optical flow by a deep learning algorithm.
[0033] In a possible implementation of the second aspect, the acquisition module is specifically configured to receive a third video stream and a fourth video stream; and combine the third video stream and the fourth video stream into a first video stream.
[0034] In a possible implementation of the second aspect, the multi-point control unit further includes an adjustment module for adjusting the third video stream and the fourth video stream to the same frame rate.
[0035] In a possible implementation of the second aspect, the receiving-end device includes a first receiving-end device and a second receiving-end device. Among them, the frame rate required by the first receiving-end device is the first frame rate, and the frame rate required by the second receiving-end device is the second frame rate. The insertion module is specifically configured to insert a first number of intermediate frame images into two frames of images based on the first frame rate to obtain a fifth video stream; and insert a second number of intermediate frame images into two frames of images based on the second frame rate to obtain a sixth video stream.
[0036] In a possible implementation of the second aspect, the frame rate of the first video stream is 30 frames per second, and the frame rate required by the receiving-end device is 60 frames per second.
[0037] The multi-point control unit provided in the second aspect of the present application is used to execute the method described in the first aspect or any one of the possible implementations of the first aspect.
[0038] The third aspect of the present application provides a multi-point control unit, including a processor and a memory. The memory is used to store instructions, and the processor is used to obtain the instructions stored in the memory to execute the method described in the first aspect or any one of the possible implementations of the first aspect.
[0039] The fourth aspect of the present application provides a computer-readable storage medium, and the computer-readable storage medium includes instructions. When the instructions run on a computer, the computer is caused to execute the method described in the first aspect or any one of the possible implementations of the first aspect.
[0040] The fifth aspect of the present application provides a computer program product containing instructions. When the computer program product runs on a computer, the computer is caused to execute the method described in the first aspect or any one of the possible implementations of the first aspect.
[0041] The sixth aspect of the present application provides a chip system, which includes at least one processor and a communication interface. The communication interface and the at least one processor are interconnected by a line. The at least one processor is used to run a computer program or instruction to execute the method described in the first aspect or any possible implementation manner of the first aspect. Description of the Drawings
[0042] Figure 1 It is a schematic diagram of an application scenario of the video processing method provided by the embodiment of the present application;
[0043] Figure 2 It is a schematic diagram of an embodiment of the video processing method provided by the embodiment of the present application;
[0044] Figure 3 It is a schematic diagram of the system architecture of the video processing method provided by the embodiment of the present application;
[0045] Figure 4 It is a schematic diagram of another embodiment of the video processing method provided by the embodiment of the present application;
[0046] Figure 5 It is a schematic diagram of the frame doubling process in the embodiment of the present application;
[0047] Figure 6 It is a schematic diagram of another embodiment of the video processing method provided by the embodiment of the present application;
[0048] Figure 7 It is a schematic diagram of another embodiment of the video processing method provided by the embodiment of the present application;
[0049] Figure 8 It is a schematic diagram of the technical effect brought by the video processing method provided by the embodiment of the present application;
[0050] Figure 9 It is a schematic diagram of the structure of the multi-point control unit provided by the embodiment of the present application;
[0051] Figure 10 It is a schematic diagram of another structure of the multi-point control unit provided by the embodiment of the present application. Detailed Embodiments
[0052] The embodiment of the present application provides a video processing method, which can reduce the increase in the bitstream when performing frame interpolation on a video. The embodiment of the present application also provides corresponding devices, computer-readable storage media, computer program products, etc. The following will be described separately.
[0053] The embodiments of the present application will be described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Those of ordinary skill in the art can understand that with the development of technology and the emergence of new scenarios, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.
[0054] In the description and claims of the present application and the above accompanying drawings, terms such as "system" and "network", "video stream" and "video bitstream" can be used interchangeably. Unless otherwise specified, ordinal numbers such as "first", "second", etc. are used to distinguish multiple objects and are not used to limit the order, timing, priority or importance of multiple objects. It should be understood that such terms can be interchanged under appropriate circumstances so that the embodiments described here can be implemented in an order other than that shown or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0055] For the sake of easy understanding, the relevant terms and concepts mainly involved in the embodiments of the present application will be introduced below.
[0056] 1. Multi-point control unit (MCU)
[0057] The MCU is an important part of the video conferencing system. When a video conference is held at more than two venues, the MCU converges, distributes, and switches the signals of multiple venues and controls the conference. The MCU is like a switch, which is the processing point for the convergence and switching of audio, video, data, signaling and other signals of each video conferencing terminal device, and is cascaded with other MCUs.
[0058] After synchronously separating the information flows of each video conferencing terminal, the MCU extracts various information such as audio, video, data, and signaling, and sends the same information of each terminal to the corresponding information processing module to complete the processing of the corresponding information. For example: mixing or switching of audio, mixing or switching of video information, broadcasting and routing of data information, timing and signaling control, etc. Finally, the processed audio, video, data, and signaling are recombined and sent to each corresponding venue.
[0059] 2. H.264
[0060] H.264, also known as Advanced Video Coding (AVC), is a video coding standard and is also Part 10 of the fourth edition of the Moving Pictures Experts Group (MPEG-4). It is a highly compressed digital video codec standard video coding standard proposed by the Joint Video Team (JVT), which is jointly composed of the Video Coding Experts Group (VCEG) and the Moving Pictures Experts Group (MPEG).
[0061] 3. SVC
[0062] Scalable Video Coding (SVC) is an extension of AVC and is Appendix G of the AVC standard. Therefore, SVC is also called H.264 Annex G. SVC technology greatly reduces the computational power requirements for the MCU and significantly improves network adaptability.
[0063] 4. H.265
[0064] H.265, also known as High Efficiency Video Coding (HEVC), is a newly developed video coding standard after H.264. Compared with H.264, H.265 provides approximately twice the data compression ratio at the same video quality level, or can significantly improve video quality at the same bit rate.
[0065] 5. I / P / B Frames
[0066] When compressing images, H.264 divides several frames of images into a Group of Pictures (GOP), which includes Intra Picture frames, Predictive Picture frames, and Bidirectionally Predicted Picture frames.
[0067] The Intra Picture frame is the I frame, also known as the key frame. The I frame is a fully encoded frame, usually the first frame of each GOP. After moderate compression, it serves as a random access reference point and can be decoded without referring to other frame data.
[0068] The predicted image frame is the P frame. The P frame represents the difference between this frame and a previous I frame or P frame. When decoding, the previously cached picture needs to be superimposed with the difference defined in this frame to generate the final picture. That is to say, the P frame is a difference frame and does not have complete picture data, which is used to represent the difference from the previous frame of the picture.
[0069] The bi-directionally predicted image frame is the B frame. The B frame records the differences between this frame and the frames before and after it. In other words, to decode the B frame, not only the previously cached picture needs to be obtained, but also the pictures after it need to be decoded. The final picture can be obtained through the superposition of the pictures before and after and the data of this frame. The B frame has a high compression ratio, but also requires higher computing power when decoding. Generally, the I frame has the lowest compression efficiency, the P frame is higher, and the B frame is the highest.
[0070] 6. Motion Estimation and Motion Compensation
[0071] Motion Estimation and Motion Compensation (MEMC) is a method for describing the differences between adjacent frames. MEMC includes two parts, namely Motion Estimation (ME) and Motion Compensation (MC). Among them, the ME part finds the motion vector to represent the motion information between two frames by comparing the pixel differences between the two frames. The MC part compensates the image according to the motion vector calculated by the ME part.
[0072] 7. Optical Flow
[0073] Optical flow is the instantaneous velocity of the pixels of a moving object in space on the observation imaging plane. It is a method that uses the changes of pixels in the time domain in the image sequence and the correlation between adjacent frames to find the corresponding relationship between the previous frame and the current frame, so as to calculate the motion information of the object between adjacent frames.
[0074] Next, first in combination with Figure 1 The application scenarios of the video processing method provided by the embodiments of the present application will be described.
[0075] As Figure 1 shown, this application scenario is a video conferencing scenario, which includes multiple sending devices, multiple receiving devices, and a multi-point control unit. The general process of a video conference is that multiple sending devices collect videos, encode the videos and send them to the MCU. The MCU receives multiple video streams and decodes them. After decoding, the multiple video streams are merged into a video picture, and then a video stream is re-encoded for each receiving device and sent to the corresponding receiving device. After receiving the video stream, the receiving device decodes and plays the video stream. It can be understood that Figure 1The number of the sending-end device and the receiving-end device is only an example and is not specifically limited.
[0076] Currently, most of the frame rates of video conferences are 30 frames per second (FPS). However, on a large display, a frame rate of 60 FPS will bring a smoother effect. For example, when playing images of people walking or waving on a large display, there will be an obvious sense of stuttering at a frame rate of 30 FPS, while the 60 FPS frame rate is very smooth. However, few devices can support the acquisition of 60 FPS video images. Especially when a 4K resolution is required, almost no device can support it. Moreover, the acquisition of 60 FPS video has high requirements for the ambient brightness, and the video effect is poor under low illuminance. Therefore, it is not very possible to implement a video conference at a frame rate of 60 FPS from the sending-end side.
[0077] To achieve the playback of a 60 FPS video conference, the common practice in the industry is to perform frame doubling processing on the video image on the receiving-end side, interpolating the video image from 30 FPS to 60 FPS. For example, the traditional MEMC scheme is used for frame doubling, or the deep learning scheme is used for interpolation. However, after interpolation, the video bitstream will increase significantly, resulting in an increase in bandwidth requirements and seriously affecting the transmission effect. Moreover, due to the limited computing power at the end side, the latency of the interpolation process is relatively large, and video conferences have very high requirements for latency, so the overall effect of the scheme is poor.
[0078] In view of this, the embodiments of the present application provide a video processing method, which is mainly applied to the video conference scenario. In the embodiments of the present application, the MEMC scheme is used for interpolation processing on the MCU side, and the intermediate frames obtained after MEMC processing are encoded in the format of B frames. After the interpolation processing, the video stream is sent to the receiving-end device, and the receiving-end device decodes and plays it. Since the calculation methods of MEMC and B frames are similar, the bitstream will only increase slightly.
[0079] It should be understood that the video processing method provided by the embodiments of the present application can be applied to AVC, and can also be applied to SVC and H.265. The encoding standard applied here is not specifically limited.
[0080] Next, in combination with Figure 2 A general description of the video processing method provided by the embodiments of the present application will be given.
[0081] As Figure 2As shown in the figure, after the MCU obtains the video stream, it decodes the video stream. After decoding, it caches two adjacent frames of images in the video stream. These two frames of images can both be P frames, or one can be an I frame and the other a P frame (i.e., at the start of the GOP). Then, motion estimation and motion compensation processing are performed based on the two frames of images. Exemplarily, motion estimation can be performed by calculating the optical flow of the two frames to obtain the motion vector. Then, motion compensation processing is performed according to the motion vector to obtain the intermediate frame image. For the intermediate frame image, it is encoded in the encoding format of a B frame, and after encoding, the intermediate frame image is inserted into the two frames of images. Finally, the video stream after frame interpolation processing is sent to the receiving device, and the receiving device can play the high-frame-rate video after decoding. It can be understood that since the inserted intermediate frame is a B frame and not a reference frame, it does not affect decoding. Therefore, if the intermediate frame is lost during the transmission to the receiving device, retransmission can be not performed, thereby improving the transmission efficiency. If the adjacent reference frame is lost, retransmission is required so as not to affect the decoding of the receiving device. Of course, retransmission of the lost intermediate frame can also be performed, and specific details are not limited here.
[0082] Please refer to the following Figure 3 , Figure 3 , which is a schematic diagram of a system architecture applied to the video processing method provided by an embodiment of the present application.
[0083] As Figure 3 shown, the MCU includes a video decoding module, an ME module, an MC module, a video encoding module, and a video transmission module. Among them, the video decoding module is used to decode the received video stream, the ME module is used to perform motion estimation on two adjacent frames of images, the MC module is used to perform motion compensation according to the motion vector obtained by the ME module, the video encoding module is used to encode the intermediate frame obtained by MEMC, and the video transmission module is used to transmit the video stream after frame interpolation processing to the receiving device. In the case of including multiple sending devices, the MCU further includes a picture synthesis module for combining the video streams of multiple venues into a single video picture. Optionally, if MEMC processing is performed by the optical flow method, the MCU further includes an optical flow calculation module for calculating the optical flow information of two adjacent frames of images.
[0084] Please refer to the following Figure 4 , Figure 4 , which is a schematic diagram of an embodiment of the video processing method provided by an embodiment of the present application. As Figure 4 shown, this embodiment includes steps 401 to 404.
[0085] 401. Obtain the first video stream.
[0086] The MCU obtains the first video stream to be processed. In a possible solution, the MCU receives the first video stream sent by the sending device and performs frame interpolation on the first video stream. In another possible solution, the first video stream is a video stream obtained by the MCU after merging multiple video streams sent by multiple sending devices. Specifically, the MCU receives the third video stream and the fourth video stream. After decoding the third video stream and the fourth video stream, the two video streams are merged into one picture to obtain the first video stream. It should be understood that in addition to merging two video streams into the first video stream, the MCU may also merge more video streams into the first video stream, which is not specifically limited here.
[0087] Optionally, the third video stream and the fourth video stream are two video streams with different frame rates. In this case, the MCU needs to first adjust the third video stream and the fourth video stream to the same frame rate before merging the third video stream and the fourth video stream. Specifically, the third video stream and the fourth video stream can be frame-interpolated to the same frame rate. Exemplarily, if the frame rate of the third video stream is 15 FPS and the frame rate of the fourth video stream is 30 FPS, then the third video stream is double-framed to 30 FPS. Or, both the third video stream and the fourth video stream are double-framed to 60 PFS, which can be specifically determined according to the frame rate required by the receiving device. The method of frame interpolation for the third video stream and the fourth video stream can be the method provided in the embodiments of the present application or a conventional frame interpolation method, such as frame sampling or frame blending. Frame sampling is to obtain intermediate frames by directly copying the front and back frames, and frame blending is to obtain intermediate frames by blending the front and back frames. Frame sampling and frame blending can increase the frame rate of the video, but the improvement of the picture effect is very little. When the video processing method provided in the embodiments of the present application is adopted, it is the first possible solution in the above-mentioned MCU to obtain the first video stream, that is, the MCU performs frame interpolation after receiving the third video stream or the fourth video stream sent by the sending device.
[0088] In addition to performing frame interpolation on the third video stream and the fourth video stream, frame dropping can also be performed on them. For example, if the frame rate of the third video stream is 15 FPS and the frame rate of the fourth video stream is 30 FPS, then the fourth video stream is frame-dropped to 15 FPS, so that the third video stream and the fourth video stream can be merged. In this embodiment, the method of adjusting the third video stream and the fourth video stream to the same frame rate is not specifically limited.
[0089] 402. Perform motion estimation and motion compensation processing on two adjacent frames of images in the first video stream to obtain an intermediate frame image, and the intermediate frame image is encoded in the format of a B frame.
[0090] After obtaining the first video stream, two adjacent frames in the first video stream are used as reference frames for frame interpolation processing. Specifically, two adjacent frames of images in the first video stream are cached. The encoding formats of these two frames of images can both be P frames, or one frame is an I frame and the other frame is a P frame (that is, the two frames of images are located at the start position of the GOP).
[0091] Then, MEMC processing is performed on these two frames of images. Optionally, MEMC processing is performed on the two adjacent frames of images by using the optical flow method. Optical flow is the movement of an object between consecutive frame sequences, caused by the relative movement between the object and the camera. When the time interval is very small, such as between two consecutive front and back frames, it is also equivalent to the displacement of the target point. The optical flow method is a method that uses the change of pixels in the image in the time domain and the correlation between adjacent frames to find the corresponding relationship existing between two adjacent frames, so as to calculate the motion information of the object between two adjacent frames. Specifically, the optical flow of two adjacent frames of images is calculated, motion estimation is performed according to the optical flow result to obtain a motion vector. Then, motion compensation processing is performed according to the motion vector to obtain an intermediate frame image.
[0092] In a possible solution, the optical flow of two adjacent frames of images is calculated by using the dense inverse search (DIS) optical flow method. The DIS optical flow method performs a reverse search for correlation at the block level, realizes calculating the gradient once and using it for multiple reverse searches without re-initializing the gradient calculation each time, saving a large amount of calculations, thereby improving the performance.
[0093] In another possible solution, the optical flow of two adjacent frames of images is calculated by using a deep learning algorithm. For example, the optical flow is calculated by using a convolutional neural network (CNN). Calculating the optical flow by using a deep learning algorithm has a lower requirement for computing power.
[0094] In addition, the optical flow of two adjacent frames of images can also be calculated by using methods such as the (kanade-lucas-tomasi, KLT) sparse optical flow method or the Farneback dense optical flow method, and the specific calculation method of the optical flow is not specifically limited here.
[0095] Optionally, MEMC processing can also be performed on the two adjacent frames of images by using methods such as the block matching method, the pixel recursion method, or the energy-based method, and the specific method is not limited here.
[0096] For the intermediate frame images, they are encoded in the format of B frames. When compressing a frame into a B frame, compression is performed based on the differences between the adjacent previous frame, the current frame, and the subsequent frame data, that is, the B frame records the differences between the current frame and the previous and subsequent frames. Since the formats of the two reference frame images are I frame and P frame respectively, or both are P frames, the encoding format of the video stream is IBPBP. It can be understood that for different frame rates required by the receiving device, the number of inserted intermediate frames is also different, and it can be one or two intermediate frame images inserted. That is to say, actually the encoding format can be IBBPBBP or IBBBBPBBBP, etc. For the sake of brevity in writing, in the embodiments of this application, it is represented by IBPBP, and the number of intermediate frames is not limited.
[0097] 403. Insert the target number of intermediate frame images into two adjacent frames based on the frame rate required by the receiving device to obtain a second video stream.
[0098] After obtaining the intermediate frame images through MEMC processing, perform frame insertion processing based on the frame rate required by the receiving device, and insert the target number of intermediate frame images into two adjacent frames, thereby obtaining a second video stream with a higher frame rate. Exemplarily, if the frame rate of the first video stream is 30 FPS and the frame rate required by the receiving device is 60 FPS, then inserting one intermediate frame image between every two adjacent frames of the video stream can change the video stream from 30 FPS to 60 FPS.
[0099] In a possible solution, the receiving device includes a first receiving device and a second receiving device, and the frame rates required by the first receiving device and the second receiving device are different. The frame rate required by the first receiving device is the first frame rate, and the frame rate required by the second receiving device is the second frame rate. Then the MCU inserts the first number of intermediate frame images into two adjacent frames based on the first frame rate to obtain a fifth video stream. Then insert the second number of intermediate frame images into two adjacent frames based on the second frame rate to obtain a sixth video stream. That is to say, the second video stream includes the fifth video stream and the sixth video stream. Exemplarily, if the frame rate of the first video stream is 30 FPS, the frame rate required by the first receiving device is 60 FPS, and the frame rate required by the second receiving device is 90 FPS, then for the first receiving device, insert one intermediate frame image between every two adjacent frames of the first video stream to obtain a fifth video stream with a frame rate of 60 FPS. For the second receiving device, insert two intermediate frame images between every two adjacent frames of the first video stream to obtain a sixth video stream with a frame rate of 90 FPS. It can be understood that in addition to the first receiving device and the second receiving device, there can be more receiving devices, and specific details are not limited here.
[0100] Inserting the intermediate frame images into the video stream can be combined with Figure 5 for understanding. AsFigure 5 As shown, 2999 / 3000 and 3001 / 3002 are adjacent front and rear frames. Based on these two frames of images, MEMC processing is performed to obtain the intermediate frame 3002 / 3001, and then the intermediate frame 3002 / 3001 is inserted between 2999 / 3000 and 3001 / 3002, thus completing the frame interpolation process. Similarly, 2998 / 2997 is the intermediate frame obtained based on 2995 / 2996 and 2997 / 2998, 3000 / 2999 is the intermediate frame obtained based on 2997 / 2998 and 2999 / 3000, and 3004 / 3003 is the intermediate frame obtained based on 3001 / 3002 and 3003 / 3004. Inserting a corresponding intermediate frame image between every two adjacent frames can double the video frame rate.
[0101] Figure 5 In [the figure], the two front and rear frames serving as reference frames are P frames with a size of 3 kilobytes (k). Since the obtained intermediate frame is obtained through MEMC processing and encoded in the format of B frames, the file size is much smaller, about 200 bytes (b).
[0102] 404. Send the second video stream to the receiving end device.
[0103] After performing frame interpolation processing to obtain the second video stream, the MCU sends the second video stream to the receiving end device. Specifically, when there is only one receiving end device, the MCU directly sends the second video stream to the receiving end device. When there are multiple receiving end devices, the MCU encodes a second video stream for each receiving end device and sends a corresponding video stream to each receiving end device.
[0104] Since the inserted intermediate frame is a B frame and not a reference frame, it does not affect decoding. Therefore, the retransmission and frame loss strategies for the video stream can be adjusted accordingly. Exemplarily, if an intermediate frame is lost during transmission to the receiving end device, retransmission may not be performed, thereby improving the transmission efficiency. If an adjacent reference frame is lost, retransmission is required to avoid affecting the decoding of the receiving end device. Of course, the lost intermediate frame can also be retransmitted to ensure the picture quality of the video stream. The retransmission strategy for the intermediate frame is not limited here.
[0105] In this embodiment, on the MCU side, the MEMC processing is performed on two adjacent frames of images to obtain an intermediate frame image, and then the intermediate frame image is inserted between the two adjacent frames to achieve the frame interpolation processing of the video stream. Moreover, the intermediate frame image is encoded in the format of B frame. Since both MEMC and B frame are calculated based on the differences between the two adjacent frames before and after, and the calculation methods are similar, inserting the intermediate frame image with the B frame encoding format obtained by MEMC will only increase a small amount of bitstream. Compared with the IPPPP encoding format (i.e., the intermediate frame is P frame) adopted by the current video conference, in this embodiment, after obtaining the intermediate frame by MEMC and then using B frame encoding (i.e., the IBPBP encoding format), the increase in bitstream can be greatly reduced.
[0106] In this embodiment, the MEMC processing is performed on the MCU side. The MCU has a relatively high computing power and only needs to refer to the two adjacent frames before and after during the MEMC processing. Therefore, the latency of the frame interpolation processing can be greatly reduced, so as to meet the requirements for latency in the video conference scenario and ensure the quality of the video conference.
[0107] Optionally, in this embodiment, the optical flow method is used to guide the MEMC, which can improve the accuracy of calculation and thus ensure the quality of the intermediate frame image.
[0108] Optionally, in this embodiment, the MCU can receive multiple video streams with different frame rates, and can perform frame interpolation processing on the video streams into multiple video streams with different frame rates, greatly expanding the application scenarios.
[0109] Reference Figure 4 According to the content described in the embodiments shown below, in combination with Figure 6 and Figure 7 an exemplary description will be given of the situation when the MCU processes multiple video streams.
[0110] Please refer to Figure 6 , which is the processing flow of the MCU when the frame rates of the video streams sent by multiple sending devices are the same and the frame rates required by multiple receiving devices are also the same. Figure 6 Taking the optical flow method for MEMC as an example. As shown in Figure 6As shown, the MCU receives the video stream 1 sent by the first sending device and the video stream 2 sent by the second sending device. Then the MCU decodes the video stream 1 and the video stream 2, and combines the decoded video streams into one picture. For the combined picture, cache two adjacent frames of images, and calculate the optical flow based on the two adjacent frames of images. Then assist in calculating the ME according to the optical flow result, perform MC according to the result of the ME, generate an intermediate frame image, and encode the intermediate frame image in the format of B frame. Then insert the intermediate frame image into the video stream based on the frame rate required by the receiving device. Specifically, insert the corresponding generated intermediate frame image into each group of adjacent frames (that is, every two adjacent frames), so as to obtain a second video stream with a higher frame rate. Finally, the MCU sends one path of the second video stream to each receiving device. Exemplarily, the frame rate of the video stream sent by the sending device is 30FPS, and the frame rate required by the receiving device is 60FPS. Then insert a corresponding intermediate frame image into each group of adjacent frames, and a second video stream with a frame rate of 60FPS can be obtained. The MCU encodes a video stream with a frame rate of 60FPS for each receiving device and sends the video stream to each receiving device.
[0111] In a video conferencing scenario, the display capabilities of multiple receiving devices may vary, such as the size of the display screen. Some receiving devices require a video with a frame rate of 60FPS, some may require a frame rate of 30FPS or 90FPS, or even 120FPS. Also, the video capture capabilities of the sending devices may also vary. For example, most devices have a capture capability of 30FPS, so they send a video with a frame rate of 30FPS. However, some devices have a capture capability of 60FPS, so the frame rate of the video they send is 60FPS. Or to save the transmitted bitstream, the sending device only sends a video with a frame rate of 15FPS. For these situations, the processing flow of the MCU is described below in combination with Figure 7 to illustrate the processing flow of the MCU. Figure 7 That is, the processing flow of the MCU when the frame rates of the video streams sent by multiple sending devices are different and the frame rates required by multiple receiving devices are also different.
[0112] Such as Figure 7As shown, the MCU receives video stream 1 sent by the first sending device and video stream 2 sent by the second sending device. Among them, the frame rates of video stream 1 and video stream 2 are different. The MCU determines the frame rates of video stream 1 and video stream 2, decodes video stream 1 and video stream 2, and then adjusts the two video streams to the same frame rate. The adjustment method can be to perform frame interpolation on video stream 1 and video stream 2 to a same frame rate, or to perform frame rate reduction on video stream 1 and video stream 2 to a same frame rate. The frame interpolation processing for video stream 1 and video stream 2 can adopt the video processing method provided by the embodiments of the present application, that is, for video stream 1, two adjacent frames of images in video stream 1 are obtained, MEMC processing is performed based on these two frames of images to obtain an intermediate frame image, the intermediate frame image is encoded in the format of a B frame, and then the intermediate frame image is inserted into video stream 1 to complete the frame interpolation processing. The same is true for video stream 2, which will not be elaborated here. Of course, other methods can also be used for frame interpolation processing, such as frame sampling or frame mixing, which are not specifically limited here.
[0113] When adjusting the frame rate of the video stream, the adjustment can be made with reference to the frame rate required by the receiving device. Exemplarily, the frame rate of video stream 1 is 15 FPS, the frame rate of video stream 2 is 30 FPS, and the MCU determines that the frame rate required by the receiving device is 60 FPS. Then, the video processing method provided by the embodiments of the present application can be used to perform frame doubling on video stream 1 and video stream 2 to 60 FPS, and the video can be directly sent after the picture composition. Of course, it can also be to perform frame doubling on video stream 1 to 30 FPS, and then perform frame doubling on the synthesized video stream to 60 FPS after the picture composition.
[0114] After video stream 1 and video stream 2 are adjusted to the same frame rate, video stream 1 and video stream 2 are merged into one picture, two adjacent frames of images are cached from the synthesized video stream, and MEMC processing is performed based on the adjacent frames to obtain an intermediate frame image. The intermediate frame image is encoded in the format of a B frame. Then, a frame doubling strategy is determined based on the frame rate required by the receiving device, and the corresponding number of intermediate frame images are inserted into the video stream to obtain a video stream with a higher frame rate. After frame doubling processing, the video stream is sent to the corresponding receiving device to complete the transmission of the video stream.
[0115] Exemplarily, the frame rate of video stream 1 is 15 FPS, the frame rate of video stream 2 is 30 FPS, the required frame rate of the first receiving device is 60 FPS, and the required frame rate of the second receiving device is 90 FPS. Then the MCU first doubles the frame rate of video stream 1 to 30 FPS, and then combines video stream 1 and video stream 2 into one picture. MEMC processing is performed on every two adjacent frames in the combined video stream to obtain intermediate frame images, and the intermediate frame images are encoded in the format of B frames. Since the required frame rate of the first receiving device is 60 FPS, one intermediate frame is inserted between every two adjacent frames. For the second receiving device with a required frame rate of 90 FPS, two intermediate frame images are inserted between every two adjacent frames.
[0116] In addition, the MCU can also process in combination with the resolution of the video. Specifically, the received video streams with different resolutions are adjusted to the same resolution and combined into one picture. Then, the combined video stream is adjusted based on the required resolution of the receiving device, so as to obtain the resolution required by the receiving device.
[0117] Please refer to Figure 8 , Figure 8 for the bitstream sizes after frame doubling processing by different methods. Among them, the first file is the original video stream with a resolution of 1080P and an IPPP encoding format, and the bitstream size of the original video stream is 71678 KB. The second, third, and fourth files are video stream files obtained by doubling the frame rate of the original video stream from 30 FPS to 60 FPS. The second file is processed for frame doubling by the deep learning method and uses the IBPBP encoding format (i.e., the intermediate frame is a B frame). The file size after frame doubling is 90628 KB, which is about 40% more than the original video stream in terms of bitstream. The fourth file is also processed for frame doubling by the deep learning method and uses the IPPP encoding format (i.e., the intermediate frame is a P frame). The file size after frame doubling is 108056 KB, which is about 50% more than the original video stream in terms of bitstream. The third file is processed for frame doubling by using the video processing method provided in the embodiment of the present application, that is, MEMC processing is performed on two adjacent frames of images through optical flow to obtain intermediate frames, and the IBPBP encoding format is used. The file size after frame doubling is 74847 KB, which is almost the same as the original file and much smaller than the file size obtained by the deep learning frame doubling method. This is because residual encoding is also performed during deep learning frame doubling, resulting in a relatively large increase in bitstream.
[0118] The above describes the embodiments of the present application from the perspective of methods. Next, the related devices in the embodiments of the present application will be introduced from the perspective of specific device implementation.
[0119] Please refer to Figure 9, an embodiment of the present application provides a schematic diagram of a multi-point control unit 900. Among them, the multi-point control unit 900 includes an acquisition module 901, a processing module 902, and an insertion module 903.
[0120] The acquisition module 901 is used to acquire a first video stream.
[0121] The processing module 902 is used to perform motion estimation and motion compensation processing on two adjacent frames of images in the first video stream to obtain an intermediate frame image. The encoding formats of the two frames of images are I-frame and P-frame respectively, or both are P-frame, and the intermediate frame image is encoded in the format of B-frame.
[0122] The insertion module 903 is used to insert a target number of intermediate frame images into the two frames of images based on the frame rate required by the receiving-end device to obtain a second video stream.
[0123] Optionally, the processing module 902 is specifically used to calculate the optical flow of the two frames of images; perform motion estimation and motion compensation processing on the two frames of images based on the optical flow.
[0124] Optionally, the processing module 902 is specifically used to calculate the optical flow by the dense inverse search optical flow method.
[0125] Optionally, the processing module 902 is specifically used to calculate the optical flow by a deep learning algorithm.
[0126] Optionally, the acquisition module 901 is specifically used to receive a third video stream and a fourth video stream; merge the third video stream and the fourth video stream into a first video stream.
[0127] Optionally, the multi-point control unit further includes an adjustment module 904, which is used to adjust the third video stream and the fourth video stream to the same frame rate.
[0128] Optionally, the receiving-end device includes a first receiving-end device and a second receiving-end device. Among them, the frame rate required by the first receiving-end device is the first frame rate, and the frame rate required by the second receiving-end device is the second frame rate. The insertion module 903 is specifically used to insert a first number of intermediate frame images into the two frames of images based on the first frame rate to obtain a fifth video stream; insert a second number of intermediate frame images into the two frames of images based on the second frame rate to obtain a sixth video stream.
[0129] Optionally, the frame rate of the first video stream is 30 frames per second, and the frame rate required by the receiving-end device is 60 frames per second.
[0130] Each module in the multi-point control unit 900 executes the operations of the multi-point control unit in the embodiments as described above Figure 2 、 Figure 4 、 Figure 6 and Figure 7 shown, and the details are not described here again.
[0131] Please refer to the following Figure 10 , which is a possible structural schematic diagram of the multi-point control unit 1000 provided by the embodiment of the present application, including a processor 1001, a communication interface 1002, a memory 1003, and a bus 1004. The processor 1001, the communication interface 1002, and the memory 1003 are interconnected through the bus 1004. In the embodiment of the present application, the processor 1001 is used to control and manage the actions of the multi-point control unit. For example, the processor 1001 is used to execute Figure 4 the steps executed by the multi-point control unit in the method embodiment shown. The communication interface 1002 is used to support the multi-point control unit to communicate. The memory 1003 is used to store the program code and data of the multi-point control unit.
[0132] Among them, the processor 1001 can be a central processing unit, a general-purpose processor, a digital signal processor, an application-specific integrated circuit, a field-programmable gate array, or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logical blocks, modules, and circuits described in combination with the disclosure of the present application. The processor can also be a combination that realizes computing functions, such as a combination of one or more microprocessors, a combination of a digital signal processor and a microprocessor, and so on. The bus 1004 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 10 only a thick line is used to represent it in
[0133] The embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium includes instructions. When the instructions run on a computer, the computer is caused to execute the foregoing Figure 2 , Figure 4 , Figure 6 and Figure 7 the methods in the embodiments shown.
[0134] The embodiment of the present application also provides a computer program product containing instructions. When the computer program product runs on a computer, the computer is caused to execute the foregoing Figure 2 , Figure 4 , Figure 6 and Figure 7 the methods in the embodiments shown.
[0135] The embodiments of the present application also provide a chip system, which includes at least one processor and a communication interface. The communication interface and the at least one processor are interconnected by a line. The at least one processor is used to run a computer program or instruction to execute the methods in the foregoing Figure 2 , Figure 4 , Figure 6 and Figure 7 illustrated embodiments.
[0136] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described system, device, and unit can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
[0137] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described system, device, and unit can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
[0138] In several embodiments provided by the present application, it should be understood that the disclosed system, device, and method can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling, direct coupling, or communication connection can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in electrical, mechanical, or other forms.
[0139] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0140] In addition, in each embodiment of the present application, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0141] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, read-only memory), random access memories (RAM, random access memory), magnetic disks, or optical discs.
Claims
1. A video processing method, characterized in that, The method is applied to a multi-point control unit (MCU), and the method includes: Obtain a first video stream; Perform motion estimation and motion compensation processing on two adjacent frames of images in the first video stream to obtain an intermediate frame image. The coding formats of the two frames of images are an intra-frame image (I-frame) and a predicted image (P-frame) respectively, or both are P-frames. The intermediate frame image is encoded in the format of a bi-directionally predicted image (B-frame); Insert a target number of intermediate frame images into the two frames of images based on the frame rate required by the receiving-end device to obtain a second video stream.
2. The method according to claim 1, wherein The performing motion estimation and motion compensation processing on two adjacent frames of images in the first video stream includes: Calculate the optical flow of the two frames of images; Perform motion estimation and motion compensation processing on the two frames of images based on the optical flow.
3. The method according to claim 2, wherein The calculating the optical flow of the two frames of images includes: Calculate the optical flow by means of a dense inverse search optical flow method.
4. The method according to claim 2, characterized in that, The calculating the optical flow of the two frames of images includes: Calculate the optical flow by means of a deep learning algorithm.
5. The method according to any one of claims 1 to 4, characterized in that, The obtaining a first video stream includes: Receive a third video stream and a fourth video stream; Merge the third video stream and the fourth video stream into a first video stream.
6. The method according to claim 5, wherein Before merging the third video stream and the fourth video stream into a first video stream, the method further includes: Adjust the third video stream and the fourth video stream to the same frame rate.
7. The method according to any one of claims 1 to 6, characterized in that, The receiving-end device includes a first receiving-end device and a second receiving-end device. Among them, the frame rate required by the first receiving-end device is a first frame rate, and the frame rate required by the second receiving-end device is a second frame rate. The inserting a target number of intermediate frame images into the two frames of images based on the frame rate required by the receiving-end device to obtain a second video stream includes: Insert a first number of the intermediate frame images into the two frames of images based on the first frame rate to obtain a fifth video stream; Insert a second number of the intermediate frame images into the two frames of images based on the second frame rate to obtain a sixth video stream.
8. The method according to claim 1, wherein The frame rate of the first video stream is 30 frames per second, and the frame rate required by the receiving-end device is 60 frames per second.
9. A multi-point control unit, characterized in that, It includes: An obtaining module, configured to obtain a first video stream; A processing module, configured to perform motion estimation and motion compensation processing on two adjacent frames of images in the first video stream to obtain an intermediate frame image. The coding formats of the two frames of images are an intra-frame image (I-frame) and a predicted image (P-frame) respectively, or both are P-frames. The intermediate frame image is encoded in the format of a bi-directionally predicted image (B-frame); An inserting module, configured to insert a target number of intermediate frame images into the two frames of images based on the frame rate required by the receiving-end device to obtain a second video stream.
10. The multi-point control unit according to claim 9, characterized in that, The processing module is specifically configured to: Calculate the optical flow of the two frames of images; Perform motion estimation and motion compensation processing on the two frames of images based on the optical flow.
11. The multi-point control unit according to claim 10, characterized in that, The processing module is specifically configured to: Calculate the optical flow by means of a dense inverse search optical flow method.
12. The multi-point control unit according to claim 10, wherein The processing module is specifically configured to: Calculate the optical flow by means of a deep learning algorithm.
13. The multi-point control unit according to any one of claims 9 to 12, characterized in that, The obtaining module is specifically configured to: Receive a third video stream and a fourth video stream; Merge the third video stream and the fourth video stream into a first video stream.
14. The multi-point control unit according to claim 13, wherein, The multi-point control unit further includes: An adjustment module, configured to adjust the third video stream and the fourth video stream to the same frame rate.
15. The multi-point control unit according to any one of claims 9 to 14, characterized in that, The receiving-end device includes a first receiving-end device and a second receiving-end device. Among them, the frame rate required by the first receiving-end device is a first frame rate, and the frame rate required by the second receiving-end device is a second frame rate. The insertion module is specifically configured to: Insert a first number of the intermediate frame images into the two frame images based on the first frame rate to obtain a fifth video stream; Insert a second number of the intermediate frame images into the two frame images based on the second frame rate to obtain a sixth video stream.
16. The multi-point control unit according to claim 9, characterized in that, The frame rate of the first video stream is 30 frames per second, and the frame rate required by the receiving-end device is 60 frames per second.
17. A multi-point control unit, characterized in that, It includes: A processor and a memory; The memory is used to store instructions; The processor is used to execute the instructions stored in the memory to implement the method according to any one of claims 1 to 8.
18. A computer-readable storage medium having a computer program stored thereon, characterized in that, The computer program, when executed by one or more processors, implements the method according to any one of claims 1 to 8.
19. A computer program product containing instructions, characterized in that, When the computer program product runs on a computer, the computer is caused to execute the method according to any one of claims 1 to 8.
Citation Information
Cited By
Picture frame processing method and device
CN120833249A
Picture frame processing method and apparatus
CN120833249B