Video processing method, apparatus, device, and storage medium
Patent Information
- Application Number
- CN202310761192.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-26
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2043-06-26
Smart Images

Figure CN116801040B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of Internet technologies, and in particular, to a video processing method, apparatus, device, and storage medium. Background Art
[0002] In the production-consumption model of video rendering, video capture devices such as cameras produce video frames at their own frame rates, and playback devices consume video frames at their own frame rates, that is, render video frames.
[0003] In video playback scenarios such as live streaming, after obtaining the video frames produced by a video capture device, a playback device usually needs to perform key point detection on the objects included in the video frames. For example, perform facial key point detection and limb key point detection on the people in the video frames, etc., in order to add rendering special effects such as animations to the video frames.
[0004] In practical applications, since the time required for key point detection of different video frames is different, it may not be possible to complete the key point detection of the current video frame to be rendered by the time point when the playback device renders the video frame, resulting in problems such as being unable to render the rendering special effects corresponding to the current video frame to be rendered, or the rendered special effects not matching the rendered video frames, which affects the user experience. Summary of the Invention
[0005] Embodiments of the present invention provide a video processing method, apparatus, device, and storage medium to improve the quality of the rendered video image.
[0006] In a first aspect, an embodiment of the present invention provides a video processing method applied to a playback device. The method includes:
[0007] Obtain a plurality of video frames sequentially produced by a video capture device at a target time interval, and a plurality of historical detection durations corresponding to special effect key points in a plurality of historical video frames, where the target time interval matches the production frame rate of the video capture device;
[0008] Detect the special effect key points corresponding to the plurality of video frames respectively;
[0009] Determine the special effects to be rendered corresponding to the plurality of video frames according to the special effect key points;
[0010] Determine the target rendering time points corresponding to the plurality of video frames according to the plurality of historical detection durations and the rendering frame rate of the playback device, and the target rendering time point corresponding to each video frame is later than the time point when the key point detection of the special effect corresponding to each video frame is completed;
[0011] Render the plurality of video frames and the special effects to be rendered corresponding to the plurality of video frames respectively at the target rendering time points.
[0012] In a second aspect, an embodiment of the present invention provides a video processing apparatus applied to a playback device, and the apparatus includes:
[0013] An acquisition module, configured to acquire a plurality of video frames sequentially produced by a video acquisition device at a target time interval, and the playback device detects a plurality of historical detection durations corresponding to special effect key points in a plurality of historical video frames, where the target time interval matches the production frame rate of the video acquisition device;
[0014] A detection module, configured to detect special effect key points corresponding to the plurality of video frames respectively; and determine special effects to be rendered corresponding to the plurality of video frames according to the special effect key points;
[0015] A processing module, configured to determine target rendering time points corresponding to the plurality of video frames respectively according to the plurality of historical detection durations and the rendering frame rate of the playback device, and the target rendering time point corresponding to each video frame is later than the time point when the detection of the special effect key point corresponding to each video frame is completed;
[0016] A rendering module, configured to render the plurality of video frames and the special effects to be rendered corresponding to the plurality of video frames respectively at the target rendering time points.
[0017] In a third aspect, an embodiment of the present invention provides an electronic device, including: a memory, a processor, and a communication interface; wherein, an executable code is stored on the memory, and when the executable code is executed by the processor, the processor can at least implement the video processing method as described in the first aspect.
[0018] In a fourth aspect, an embodiment of the present invention provides a non-transitory machine-readable storage medium, on which an executable code is stored, and when the executable code is executed by a processor of an electronic device, the processor can at least implement the video processing method as described in the first aspect.
[0019] The solution provided by the embodiments of the present invention is applied to the video special effect rendering and playing scenario. In this scenario, the video acquisition device continuously produces multiple video frames at equal time intervals corresponding to its own production frame rate and transmits them to the playing device for the playing device to perform key point detection and rendering of the special effects corresponding to the key points in the video frames. In the specific implementation process, after the playing device obtains multiple video frames sequentially produced by the video acquisition device at the target time interval, first, it respectively detects the special effect key points corresponding to the multiple video frames, and determines the special effects to be rendered corresponding to the multiple video frames according to the special effect key points. In order to ensure that the detection of the special effect key points corresponding to the currently to-be-rendered video frame is completed during rendering, in this solution, the playing device also obtains multiple historical detection durations corresponding to the detection of the special effect key points in multiple historical video frames, and determines the target rendering time points corresponding to the multiple video frames according to the multiple historical detection durations and the rendering frame rate of the playing device. Among them, the target rendering time point corresponding to each video frame is later than the time point when the detection of the special effect key points corresponding to each video frame is completed. Finally, at the target rendering time points, multiple video frames and the special effects to be rendered corresponding to the multiple video frames are rendered.
[0020] In this solution, according to the multiple historical detection durations corresponding to the detection of the special effect key points in multiple historical video frames by the playing device, the rendering time points for the playing device to render the multiple obtained video frames are re-determined, so as to ensure that while rendering the video frames, the special effects to be rendered corresponding to the special effect key points in the video frames can also be rendered, ensuring the matching between the special effects and the video frames, and improving the quality of video frame rendering. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.
[0022] Figure 1 It is a flowchart of a video processing method provided by the embodiments of the present invention;
[0023] Figure 2 It is a schematic diagram of the scenario during the video playing process provided by the embodiments of the present invention;
[0024] Figure 3 It is another schematic diagram of the scenario during the video playing process provided by the embodiments of the present invention;
[0025] Figure 4 It is yet another schematic diagram of the scenario during the video playing process provided by the embodiments of the present invention;
[0026] Figure 5 Schematic structural diagram of a video processing device provided by an embodiment of the present invention;
[0027] Figure 6 For Figure 5 Schematic structural diagram of an electronic device corresponding to the video processing device provided by the illustrated embodiment. Detailed implementation manners
[0028] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0029] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the embodiments of the present invention are all information and data that have been authorized by the user or fully authorized by all parties. Moreover, the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to select authorization or rejection.
[0030] In addition, the step timings in the following method embodiments are only examples and are not strictly limited.
[0031] Figure 1 Flowchart of a video processing method provided by an embodiment of the present invention, which is applied to a playback device, such as Figure 1 shown, and may include the following steps:
[0032] 101. Obtain a plurality of video frames sequentially produced by a video capture device at a target time interval, and a plurality of historical detection durations corresponding to special effect key points in a plurality of historical video frames, where the target time interval matches the production frame rate of the video capture device.
[0033] 102. Detect special effect key points corresponding to the plurality of video frames respectively.
[0034] 103. Determine special effects to be rendered corresponding to the plurality of video frames respectively according to the special effect key points corresponding to the plurality of video frames respectively.
[0035] 104. Determine target rendering time points corresponding to the plurality of video frames respectively according to the plurality of historical detection durations and the rendering frame rate of the playback device, where the target rendering time point corresponding to each video frame is later than the time point when the detection of the special effect key point corresponding to each video frame is completed.
[0036] 105. Render multiple video frames and the rendering effects to be rendered corresponding to the multiple video frames respectively at the target rendering time point.
[0037] In this embodiment, the playback device includes terminal devices with playback and display functions such as a PC, a laptop, and a smart phone. An application program capable of implementing video playback functions and special effect rendering functions is set in the playback device. For example, in scenarios such as live streaming and short video shooting, the user interface corresponding to the application program can usually be provided with a variety of special effect interaction controls for the user to select to achieve different video rendering effects. For example: in response to the user's selection of the beauty makeup control, by detecting the key points corresponding to the human face in the video frame, the beauty makeup effect is added; and for another example: in response to the user's selection of the action special effect, by detecting the key points corresponding to the limbs of the human in the video frame, the action special effect is added, etc. The video acquisition device includes devices with video shooting functions such as a camera and a video camera. Among them, the playback device is communicatively connected to the video acquisition device. Optionally, the video acquisition device can be integrated into the playback device or externally connected to the playback device.
[0038] Optionally, for rendering the video frames that need to render special effects, the playback device can establish two threads, respectively denoted as the first thread and the second thread. Among them, the first thread is used to detect the special effect key points corresponding to the multiple video frames respectively, and determine the rendering effects to be rendered corresponding to the multiple video frames respectively according to the special effect key points; the second thread is used to render the multiple video frames and the rendering effects to be rendered corresponding to the multiple video frames respectively at the target rendering time point. There is a communication relationship between the first thread and the second thread, and the rendering effects to be rendered determined by the first thread can be transmitted to the second thread.
[0039] It can be understood that in application scenarios that require special effect rendering, the playback device needs to render both video frames and special effects matching the video frames during rendering. If a certain video frame does not render special effects during rendering, or the rendered special effects do not match the video frame, problems such as uneven video playback and incorrect display positions of the picture special effects will occur when the video frame is displayed on the playback device, greatly affecting the user's viewing experience.
[0040] Generally, in scenarios such as live streaming and short video shooting that require real-time playback and display, a video capture device will produce multiple video frames corresponding to a target shooting object at a certain fixed production frame rate, and sequentially transmit the produced multiple video frames to a playback device at the same target time interval. Among them, the target time interval matches the production frame rate of the video capture device. For example, if the production frame rate corresponding to the video capture device is 30 frames per second (fps), then the video capture device transmits a produced video frame to the playback device every 33.3 milliseconds. Here, 33.3 milliseconds is the target time interval matching the frame rate of 30 fps.
[0041] After obtaining multiple video frames sequentially transmitted by the video capture device, the playback device detects the special effect key points of the obtained video frames in real time, so as to determine the special effects to be rendered corresponding to each video frame according to the detected special effect key points. For example, detecting the key points of the facial features of the person in video frame i to determine to add beauty effects to the person in video frame i. Further, when reaching the rendering time point corresponding to the rendering frame rate of the playback device, the obtained video frames and the special effects to be rendered corresponding to the video frames are sequentially rendered. That is, the playback device renders the obtained video frames and the special effects to be rendered corresponding to the video frames at a time interval matching the rendering frame rate. For example, if the rendering frame rate corresponding to the playback device is 30 fps, then the playback device renders a video frame and the special effect to be rendered corresponding to the video frame every 33.3 milliseconds.
[0042] It can be understood that detecting the special effect key points of any video frame requires a certain detection time, and different video frames have different detection times due to different objects to be detected. In practical applications, when reaching the rendering time point corresponding to the playback device, if the special effect key points of the corresponding video frame have not been detected yet, problems such as unable to render the special effect or inaccurate rendering position of the special effect may occur. In some special application scenarios, such as the beauty effect rendering scenario, if the rendering position of the special effect is inaccurate, it will bring a bad user experience.
[0043] To facilitate understanding the impact of the detection time corresponding to the special effect key points on the special effect rendering of video frames, the following is combined with Figure 2 for description.
[0044] Figure 2 FIG. is a schematic diagram of a scenario of a video playback process provided by an embodiment of the present invention, as shown in Figure 2As shown, starting from time point t0, the video capture device sequentially transmits video frame 1, video frame 2, video frame 3, and video frame 4 produced at time points t0, t1, t2, and t3 to the playback device respectively. Among them, the time interval between any two adjacent time points among t0, t1, t2, and t3 is Δt, that is, the target time interval corresponding to the video capture device is Δt. Assuming that the rendering frame rate of the playback device is the same as the production frame rate of the video capture device, the rendering time points corresponding to the playback device are time points t1, t2, t3, and t4 with adjacent time intervals of Δt.
[0045] After the playback device obtains video frame 1 at time point t0, since there is no video frame currently undergoing special effect key point detection, therefore, it can directly perform special effect key point detection on video frame 1 to determine the special effect to be rendered corresponding to video frame 1 based on the detected special effect key points. As Figure 2 shown, the special effect key point detection duration of video frame 1 is Δt1. Since Δt1 is less than Δt and t0 + Δt1 is less than t1, therefore, when reaching the rendering time point t1, video frame 1 and the special effect to be rendered corresponding to video frame 1 can be rendered simultaneously.
[0046] After the playback device obtains video frame 2 at time point t1, since there is no video frame currently undergoing special effect key point detection, therefore, it can directly perform special effect key point detection on video frame 2 to determine the special effect to be rendered corresponding to video frame 2 based on the detected special effect key points. As Figure 2 shown, the special effect key point detection duration of video frame 2 is Δt2. Since Δt2 is greater than Δt and t1 + Δt2 is greater than t2, therefore, when reaching the rendering time point t2, the special effect key point detection of video frame 2 has not been completed, and video frame 2 and the special effect to be rendered corresponding to video frame 2 cannot be rendered simultaneously. Usually, the playback device will render the previous video frame and the corresponding special effect, that is, video frame 1 and the special effect to be rendered corresponding to video frame 1, or render video frame 2 and the special effect corresponding to the previous video frame ( Figure 2 not shown in the figure).
[0047] After the playback device obtains video frame 3 at time point t2, since video frame 2 is still undergoing special effect key point detection currently, therefore, it is impossible to perform special effect key point detection on video frame 3. As Figure 2 shown, since t1 + Δt2 is greater than t2 and less than t3, therefore, when reaching the rendering time point t3, the special effect key point detection of video frame 2 is completed, and video frame 2 and the special effect to be rendered corresponding to video frame 2 can be rendered simultaneously.
[0048] After the playback device obtains video frame 4 at time point t3, since there is currently no video frame undergoing special effect key point detection, therefore, the special effect key point detection can be directly performed on video frame 4 (the special effect key point detection of video frame 3 is no longer performed), so as to determine the special effect to be rendered corresponding to video frame 4 according to the detected special effect key points. As Figure 2 shown, similar to video frame 2, the detection duration Δt2 of the special effect key points of video frame 4 is greater than Δt, and t3 + Δt2 is greater than t4. Therefore, when the rendering time point t4 is reached, the detection of the special effect key points of video frame 4 has not been completed, and it is impossible to render video frame 4 and the special effect to be rendered corresponding to video frame 4 simultaneously.
[0049] Based on the above Figure 2 illustrated situation, it is not difficult to find that when the detection durations of the special effect key points corresponding to video frames are inconsistent and the detection duration is greater than the time interval Δt between the rendering time points of the playback device, rendering according to the rendering time points corresponding to the rendering frame rate of the playback device will cause some video frames and corresponding special effects not to be rendered in time, or the rendered video frames and special effects do not match, such as video frame 2, etc., or are not rendered, such as video frame 3, affecting the smoothness and accuracy of video rendering.
[0050] To solve the above at least one technical problem, the embodiment of the present invention provides a video processing method as Figure 1 shown. This method recombines the multiple historical detection durations corresponding to the special effect key points in multiple historical video frames detected by the playback device and the frame rate of the playback device to re-determine the rendering time points for rendering the multiple obtained video frames that match the rendering frame rate of the playback device, that is, the target rendering time points, so as to ensure that when each target rendering time point is reached, the detection of the current video frame to be detected has been completed, ensuring the uniformity and correctness of the rendering of the obtained video frames and the special effects corresponding to the video frames.
[0051] Optionally, obtaining the multiple historical detection durations corresponding to the special effect key points in multiple historical video frames detected by the playback device includes: selecting multiple historical video frames from the multiple video frames detected by the playback device's history, and using the detection durations of the special effect key points corresponding to the selected multiple historical video frames as the multiple historical detection durations for determining the target rendering time point of the playback device.
[0052] Among them, selecting multiple historical video frames from the multiple video frames detected by the playback device's history includes: selecting multiple historical video frames from the multiple video frames detected by the playback device's history whose detected special effect key point types are the same as the special effect key point type to be currently detected, for example: both are for detecting special effect key points of the limbs, and the detected special effect key points all include joint 1, joint 2,..., joint n, etc.
[0053] Optionally, obtaining multiple historical detection durations corresponding to special effect key points in multiple historical video frames by the playback device further includes: before performing special effect key point detection, the playback device performs special effect key point detection of the same type on multiple pre-stored video frames, and determines that the detection durations corresponding to the multiple video frames respectively are multiple historical detection durations for determining the target rendering time point of the playback device. Among them, before performing key point detection, the multiple historical detection durations obtained by the playback device performing special effect key point detection on multiple pre-stored video frames can more accurately reflect the current special effect key point detection ability of the playback device.
[0054] In the specific implementation process, determining the target rendering time point corresponding to each of the multiple sequentially obtained video frames according to the multiple historical detection durations and the rendering frame rate of the playback device includes:
[0055] Determining a target detection duration from the multiple historical detection durations, where the historical detection duration corresponding to the target detection duration is greater than other historical detection durations among the multiple historical detection durations; and determining the target rendering time point corresponding to each of the multiple video frames according to the target detection duration and the rendering frame rate of the playback device.
[0056] In practical applications, the production frame rate corresponding to the video acquisition device and the rendering frame rate corresponding to the playback device may be the same or different. In this embodiment, an example is given with the production frame rate being the same as the rendering frame rate. It is easy to understand that when the production frame rate is the same as the rendering frame rate, the time interval between the rendering time points matching the rendering frame rate is equal to the target time interval matching the production frame rate.
[0057] In this embodiment, the rendering time point originally determined by the playback device according to the frame rate is called the initial rendering time point, such as Figure 2 the time points t1, t2, etc. in, and the time point obtained by the playback device adjusting the initial rendering time point according to the multiple historical detection durations is called the target rendering time point, and the time interval between adjacent initial rendering time points or target rendering time points is the target time interval.
[0058] When determining the target rendering time point corresponding to each of the multiple video frames according to the target detection duration and the rendering frame rate of the playback device, if the target detection duration is less than or equal to the target time interval, the initial rendering time point determined by the playback device according to the rendering frame rate remains unchanged, that is, rendering is performed according to the original corresponding rendering time point of the playback device, and the initial rendering time point of the playback device is the target rendering time point. If the target detection duration is greater than the target time interval, it is determined that the time point corresponding to the initial rendering time point delayed by the target duration is the target rendering time point, and the target duration matches the target detection duration.
[0059] Optionally, when determining the target duration, the ratio of the target detection duration to the target time interval can be determined; the product of the rounded-up value of the ratio and the target time interval is determined as the target duration. Alternatively, the target detection duration is used as the target duration.
[0060] For ease of understanding, for example, assume that the multiple historical detection durations corresponding to the special effect key points in multiple historical video frames obtained by the playback device are T1, T2, …, Tn respectively, and the maximum value among T1, T2, …, Tn is Tn, then the target detection duration is determined as Tn. If Tn is less than or equal to the target time interval Δt, it indicates that the playback device has a strong ability to detect special effect key points and can usually complete the detection of the special effect key points of the video frame within Δt. In this case, at any initial rendering time point, the playback device can render the current video frame to be rendered and the corresponding special effect to be rendered for the video frame.
[0061] For example, Figure 3 FIG. is a schematic diagram of another scenario in the video playback process provided by an embodiment of the present invention. As Figure 3 shown, the video acquisition device sequentially transmits the produced video frames 1, video frames 2, video frames 3, and video frames 4 to the playback device at time points t0, t1, t2, and t3 respectively. Among them, the time interval between any two adjacent time points among the time points t0, t1, t2, and t3 is Δt, that is, the target time interval corresponding to the video acquisition device is Δt. Assume that the initial rendering time points corresponding to the playback device are time points t1, t2, t3, and t4 with adjacent time intervals of Δt.
[0062] As Figure 3 shown, if the detection durations corresponding to video frames 1, 2, 3, and 4 are all Δt1, since Δt1 does not exceed the target time interval Δt, therefore, when reaching the rendering time point t1, the video frame 1 and the corresponding special effect to be rendered for video frame 1 can be rendered simultaneously; when reaching the rendering time point t2, the video frame 2 and the corresponding special effect to be rendered for video frame 2 can be rendered simultaneously; when reaching the rendering time point t3, the video frame 3 and the corresponding special effect to be rendered for video frame 3 can be rendered simultaneously, and so on. The initial rendering time points being time points t1, t2, t3, and t4 are the target rendering time points.
[0063] When the target detection duration Tn corresponding to the multiple historical detection durations T1, T2, …, Tn is greater than the target time interval Δt, it indicates that the playback device has a weak ability to detect special effect key points and usually cannot complete the detection of the special effect key points of the video frame within Δt. In this case, when reaching the initial rendering time point corresponding to the playback device, there may be a situation where the special effect key points of the current video frame to be rendered have not been detected yet, so that the current video frame to be rendered and the corresponding special effect to be rendered for the video frame cannot be rendered.
[0064] For example, Figure 4 FIG. is another schematic diagram of a scenario during the video playback process provided by an embodiment of the present invention. As Figure 4 shown, the video capture device sequentially transmits the produced video frames 1, video frames 2, video frames 3, and video frames 4 to the playback device at time points t0, t1, t2, and t3, respectively. Among them, the time interval between any two adjacent time points among the time points t0, t1, t2, and t3 is Δt, that is, the target time interval corresponding to the video capture device is Δt. Assume that the initial rendering time points corresponding to the playback device are time points t1, t2, t3, and t4 with adjacent time intervals of Δt.
[0065] As Figure 4 shown, if the detection durations corresponding to the video frames 1, video frames 2, video frames 3, and video frames 4 are all Δt2, where Δt2 is greater than Δt and less than 2Δt, then when rendering according to the initial rendering time points, similar to the Figure 2 embodiment shown, situations where the video frame 1 and its corresponding special effect key points cannot be synchronously rendered will occur at time points such as t1. Assume that Δt2 is 1.5Δt, then the ratio of Δt2 to Δt is 1.5, and after rounding up 1.5, it is 2. Further, determine the time points t3, t4, etc. corresponding to the initial rendering time points t1, t2, t3, and t4 delayed by the target duration 2Δt as the target rendering time points. Specifically, the time point (t1 + 2Δt = t3) corresponding to the initial rendering time point t1 delayed by the target duration 2Δt is the target rendering time point of the initial rendering time point, and so on. Thus, it is ensured that the target rendering time points corresponding to the video frames 1, video frames 2, video frames 3, and video frames 4 are later than the time points when the detection of the special effect key points corresponding to the video frames 1, video frames 2, video frames 3, and video frames 4 is completed.
[0066] In practical applications, since the detection duration of the special effect key points corresponding to the video frames is relatively long, it may not be possible to detect the key points for every video frame, but this has a relatively small impact on the special effect rendering. As Figure 4As shown, although only Video Frame 1 and Video Frame 3 were detected during the special effect key point detection for Video Frame 1, Video Frame 2, Video Frame 3, and Video Frame 4 due to the long detection duration, the special effect key points corresponding to Video Frame 2 and Video Frame 4 can be determined by other means. For example, the special effect key points corresponding to Video Frame 1 and Video Frame 3 are interpolated to determine the special effect key points corresponding to Video Frame 2. Thus, when reaching the target rendering time point delayed from the initial rendering time point, it can be ensured that both the current video frame to be rendered and the special effects to be rendered corresponding to the video frame to be rendered can be rendered. For example, when reaching the rendering time point t3, Video Frame 1 and the special effects to be rendered corresponding to Video Frame 1 can be rendered simultaneously; when reaching the rendering time point t4, Video Frame 2 and the special effects to be rendered corresponding to Video Frame 2 can be rendered simultaneously; when reaching the rendering time point t5, Video Frame 3 and the special effects to be rendered corresponding to Video Frame 3 can be rendered simultaneously, and so on.
[0067] In practical applications, to ensure the sequential rendering of video frames, optionally, the video frames carry timestamps, which are used to indicate the production order of each video frame. Thus, during rendering, each video frame can be rendered in chronological order from earliest to latest according to the timestamps carried by each video frame.
[0068] The capabilities of different playback devices to detect special effect key points in video frames vary. To improve the detection efficiency and accuracy of special effect key points in each video frame of the video frame sequence, optionally, the playback device can set a corresponding detection interval according to its own detection capabilities. For example, perform special effect key point detection every other frame. During the special effect key point detection, at the preset detection interval, detect the special effect key points corresponding to some video frames among multiple video frames; determine the special effect key points corresponding to the other video frames in the video frame sequence that have not been detected based on the special effect key points corresponding to the partial video frames. For example, if the detection interval is 1, for Video Frame 1, Video Frame 2, and Video Frame 3, detect the special effect key points corresponding to Video Frame 1 and Video Frame 3, and determine the special effect key points corresponding to Video Frame 2 by interpolating the special effect key points of Video Frame 1 and Video Frame 3.
[0069] In this solution, based on the multiple historical detection durations corresponding to the special effect key points detected in multiple historical video frames by the playback device, the rendering time points of the multiple video frames obtained by the playback device for rendering are re-determined, ensuring that the target rendering time point corresponding to each video frame is later than the time point when the detection of the special effect key points corresponding to each video frame is completed. Thus, while rendering the video frame at the target rendering time point, the special effects to be rendered corresponding to the special effect key points in the video frame can also be rendered, making the special effects match the video frames and improving the quality of video frame rendering.
[0070] The video processing device of one or more embodiments of the present invention will be described in detail below. Those skilled in the art can understand that these devices can all be configured by using commercially available hardware components through the steps taught by this solution.
[0071] Figure 5 FIG. is a schematic structural diagram of a video processing device provided for an embodiment of the present invention, which is applied to a playback device, such as Figure 5 As shown, the device includes: an acquisition module 11, a detection module 12, a processing module 13, and a rendering module 14.
[0072] The acquisition module 11 is configured to acquire a plurality of video frames sequentially produced by a video acquisition device at a target time interval, and the playback device detects a plurality of historical detection durations corresponding to special effect key points in a plurality of historical video frames, and the target time interval matches the production frame rate of the video acquisition device.
[0073] The detection module 12 is configured to detect special effect key points respectively corresponding to the plurality of video frames; and determine special effects to be rendered respectively corresponding to the plurality of video frames according to the special effect key points;
[0074] The processing module 13 is configured to determine target rendering time points respectively corresponding to the plurality of video frames according to the plurality of historical detection durations and the rendering frame rate of the playback device, and the target rendering time point corresponding to each video frame is later than the time point when the detection of the special effect key point corresponding to each video frame is completed.
[0075] The rendering module 14 is configured to render the plurality of video frames and the special effects to be rendered respectively corresponding to the plurality of video frames at the target rendering time points.
[0076] Optionally, the rendering frame rate is the same as the production frame rate, and the time interval between adjacent target rendering time points is the target time interval. The processing module 13 is specifically configured to determine a target detection duration from the plurality of historical detection durations, and the historical detection duration corresponding to the target detection duration is greater than other historical detection durations in the plurality of historical detection durations; and determine the target rendering time points respectively corresponding to the plurality of video frames according to the target detection duration and the rendering frame rate of the playback device.
[0077] Optionally, the processing module 13 is further configured to keep the initial rendering time point determined by the playback device according to the rendering frame rate unchanged if the target detection duration is less than or equal to the target time interval; and determine the time point corresponding to the initial rendering time point after delaying the target duration as the target rendering time point if the target detection duration is greater than the target time interval, and the target duration matches the target detection duration.
[0078] Optionally, the processing module 13 is further specifically configured to determine a ratio of the target detection duration to the target time interval; and determine a product of a value obtained by rounding up the ratio and the target time interval as the target duration.
[0079] Optionally, the detection module 12 is further configured to detect special effect key points corresponding to some video frames in the multiple video frames at a preset detection interval; and determine special effect key points corresponding to other video frames in the video frame sequence that are not detected according to the special effect key points corresponding to the some video frames.
[0080] Optionally, the processing module 13 is further specifically configured to establish a first thread and a second thread, where there is a communication relationship between the first thread and the second thread; wherein, the first thread is configured to detect special effect key points corresponding to the multiple video frames respectively, and determine special effects to be rendered corresponding to the multiple video frames respectively according to the special effect key points; and the second thread is configured to render the multiple video frames and the special effects to be rendered corresponding to the multiple video frames respectively at the target rendering time point.
[0081] Figure 5 The device shown can execute the steps described in the foregoing embodiments. For the detailed execution process and technical effects, refer to the descriptions in the foregoing embodiments, which will not be repeated here.
[0082] In a possible design, the above Figure 5 structure of the video processing device shown can be implemented as an electronic device, as Figure 6 shown, the electronic device may include: a memory 21, a processor 22, and a communication interface 23. Wherein, executable code is stored on the memory 21, and when the executable code is executed by the processor 22, the processor 22 can at least implement the video processing method provided in the foregoing embodiments.
[0083] In addition, an embodiment of the present invention provides a non-transitory machine-readable storage medium, on which executable code is stored. When the executable code is executed by a processor of an electronic device, the processor can at least implement the video processing method provided in the foregoing embodiments.
[0084] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative labor.
[0085] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of adding a necessary general hardware platform, and of course, it can also be implemented by the combination of hardware and software. Based on such an understanding, the above technical solutions, in essence, or the parts that contribute to the prior art can be embodied in the form of a computer product. The present invention can be implemented in the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) that contain computer-usable program codes.
[0086] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A video processing method, characterized in that, Applied to a playback device, including: Obtaining a plurality of video frames sequentially produced by a video capture device at a target time interval, and the playback device detecting a plurality of historical detection durations corresponding to special effect key points in the plurality of historical video frames, where the target time interval matches the production frame rate of the video capture device; Detecting special effect key points corresponding to the plurality of video frames respectively; Determining special effects to be rendered corresponding to the plurality of video frames respectively according to the special effect key points; Determining target rendering time points corresponding to the plurality of video frames respectively according to the plurality of historical detection durations and the rendering frame rate of the playback device, and the target rendering time point corresponding to each video frame is later than the time point when the detection of the special effect key point corresponding to each video frame is completed; Rendering the plurality of video frames and the special effects to be rendered corresponding to the plurality of video frames respectively at the target rendering time points; Among them, determining the target rendering time points corresponding to the plurality of video frames respectively according to the plurality of historical detection durations and the rendering frame rate of the playback device includes: Determining a target detection duration from the plurality of historical detection durations, and the historical detection duration corresponding to the target detection duration is greater than other historical detection durations in the plurality of historical detection durations; Determining the target rendering time points corresponding to the plurality of video frames respectively according to the target detection duration and the rendering frame rate of the playback device; wherein, if the target detection duration is less than or equal to the target time interval, the initial rendering time point determined by the playback device according to the rendering frame rate remains unchanged; if the target detection duration is greater than the target time interval, determining the time point corresponding to the initial rendering time point after delaying by a target duration as the target rendering time point, and the target duration matches the target detection duration.
2. The method according to claim 1, wherein The rendering frame rate is the same as the production frame rate, and the time interval between adjacent target rendering time points is the target time interval.
3. The method according to claim 1, characterized in that, The target duration is determined by the following method: Determining the ratio of the target detection duration to the target time interval; Determining the product of the value obtained by rounding up the ratio and the target time interval as the target duration.
4. The method according to claim 1, wherein The detecting special effect key points corresponding to the plurality of video frames respectively includes: Detecting special effect key points corresponding to some video frames in the plurality of video frames at a preset detection interval; Determining special effect key points corresponding to other video frames in the video frame sequence that have not been detected according to the special effect key points corresponding to the some video frames.
5. The method according to any one of claims 1 to 4, characterized in that, The method further includes: Establishing a first thread and a second thread, and there is a communication relationship between the first thread and the second thread; Among them, the first thread is used to detect special effect key points corresponding to the plurality of video frames respectively, and determine special effects to be rendered corresponding to the plurality of video frames respectively according to the special effect key points; The second thread is used to render the plurality of video frames and the special effects to be rendered corresponding to the plurality of video frames respectively at the target rendering time points.
6. A video processing device, characterized in that, Applied to a playback device, including: An acquisition module, configured to acquire a plurality of video frames sequentially produced by a video acquisition device at a target time interval, and the detection durations of special effect key points corresponding to a plurality of historical video frames detected by a playback device, where the target time interval matches the production frame rate of the video acquisition device; A detection module, configured to detect special effect key points corresponding to the plurality of video frames respectively; and determine special effects to be rendered corresponding to the plurality of video frames respectively according to the special effect key points; A processing module, configured to determine target rendering time points corresponding to the plurality of video frames respectively according to the plurality of historical detection durations and the rendering frame rate of the playback device, where the target rendering time point corresponding to each video frame is later than the time point when the detection of the special effect key point corresponding to each video frame is completed; wherein, determining the target rendering time points corresponding to the plurality of video frames respectively according to the plurality of historical detection durations and the rendering frame rate of the playback device includes: determining a target detection duration from the plurality of historical detection durations, where the historical detection duration corresponding to the target detection duration is greater than other historical detection durations in the plurality of historical detection durations; determining the target rendering time points corresponding to the plurality of video frames respectively according to the target detection duration and the rendering frame rate of the playback device; wherein, if the target detection duration is less than or equal to the target time interval, keep the initial rendering time point determined by the playback device according to the rendering frame rate unchanged; if the target detection duration is greater than the target time interval, determine the time point corresponding to the initial rendering time point after delaying by a target duration as the target rendering time point, where the target duration matches the target detection duration; A rendering module, configured to render the plurality of video frames and the special effects to be rendered corresponding to the plurality of video frames respectively at the target rendering time points.
7. An electronic device, characterized in that, Including: A memory, a processor, and a communication interface; wherein, executable code is stored on the memory, and when the executable code is executed by the processor, the processor executes the video processing method according to any one of claims 1 to 5.
8. A non-transitory machine-readable storage medium, characterized in that, Executable code is stored on the non-transitory machine-readable storage medium, and when the executable code is executed by the processor of an electronic device, the processor executes the video processing method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Video data processing method and device, electronic equipment and storage medium
CN113422980A
Image processing method and device, electronic equipment and storage medium
CN114170632A