Video processing methods, apparatus, systems and computer-readable storage media

By determining the position of the target object at the video acquisition end and performing image stabilization, the problem of image shake caused by video jitter is solved, improving video quality and viewing experience.

CN122138046APending Publication Date: 2026-06-02GREE ELECTRIC APPLIANCE INC OF ZHUHAI

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GREE ELECTRIC APPLIANCE INC OF ZHUHAI
Filing Date
2026-02-25
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

During live streaming, shaking of the video acquisition device can cause the video image to shakiness, making viewers prone to dizziness.

Method used

The system sequentially acquires video images using a photosensitive unit, determines whether the target object is located within the core area of ​​the image, obtains the initial position of the target object in consecutive video frames, performs image stabilization based on these positions, and sends the image stabilization parameters to the video display end for further processing.

Benefits of technology

It effectively reduces the impact of video acquisition device jitter on video image quality, improves video image quality, and enhances the viewing experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122138046A_ABST
    Figure CN122138046A_ABST
Patent Text Reader

Abstract

This application relates to a video processing method, apparatus, system, and computer-readable storage medium. The method includes: sequentially acquiring video images frame by frame using a photosensitive unit, and sequentially determining whether a target object in n consecutive video images is located within the image core region, where n is any integer greater than 1; if the target object is located within the image core region in the n consecutive video images, obtaining n initial positions of the target object relative to the center point of the photosensitive unit in the n consecutive video images; performing image stabilization processing based on the n initial positions to obtain image stabilization parameters corresponding to the Nth video image in the n consecutive video images, and sending the image stabilization parameters to a video display end for image stabilization processing based on the received image stabilization parameters corresponding to each frame of video images, where N is a positive integer less than or equal to n. This can reduce video image jitter caused by video acquisition terminal jitter and improve the viewing experience for video viewers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of video processing technology, and in particular to a video processing method, apparatus, system and computer-readable storage medium. Background Technology

[0002] With the development of internet technology, live streaming has become a popular form of social interaction and economic activity. However, during live streaming, if there is any shaking at the video acquisition end, such as a person walking or a vehicle moving, the video displayed on the screen will be shaky, making viewers prone to dizziness. Therefore, how to reduce the impact of video acquisition shaking on the video image has become an urgent technical problem to be solved. Summary of the Invention

[0003] This application provides a video processing method, apparatus, system, and computer-readable storage medium to solve the problem that video image jitter caused by video acquisition terminal shaking can easily cause dizziness when watching video.

[0004] In a first aspect, embodiments of this application provide a video processing method applied to a video acquisition end, the method comprising: The system uses a photosensitive unit to sequentially acquire each frame of video image and sequentially determines whether the target object in n consecutive video images is located within the image core area. The image core area is the region within a preset distance from the center point of the photosensitive unit, and n is any integer greater than 1. When the target object in the n consecutive video frames is located within the core area of ​​the image, n initial positions of the target object relative to the center point of the photosensitive unit in the n consecutive video frames are obtained, wherein the n initial positions correspond one-to-one with the n consecutive video frames; Based on the n initial positions, image stabilization is performed to obtain the image stabilization parameters corresponding to the Nth video image in the n consecutive video images. The image stabilization parameters are then sent to the video display terminal so that the video display terminal can perform image stabilization based on the image stabilization parameters corresponding to each received video image. N is a positive integer less than or equal to n.

[0005] Optionally, the step of performing image stabilization based on the n initial positions to obtain the image stabilization parameters corresponding to the Nth frame of the consecutive n frames of video images includes: Calculate the center point of the n initial positions, and determine the center point of the n initial positions as the target position of the target object in the Nth frame of the video image; Based on the n initial positions, a jitter coefficient is determined, wherein the jitter coefficient is used to characterize the degree of jitter between the n consecutive video frames; The target position and the jitter coefficient are determined as the anti-shake parameters.

[0006] Optionally, the formula for calculating the jitter coefficient is as follows: ; in, Indicates the jitter coefficient, ( , ) represents the i-th initial position among the n initial positions, where i∈[1,n].

[0007] Optionally, after determining whether the target objects in n consecutive video frames are all located within the image core region, the method further includes: If the target object in the n consecutive video frames is not located within the core area of ​​the image, the latest image stabilization parameter obtained before the n consecutive video frames is used to replace the image stabilization parameter corresponding to the Nth video frame and sent to the video display terminal.

[0008] Secondly, embodiments of this application also provide a video processing method applied to a video display end, the method comprising: The receiving end sends anti-shake parameters corresponding to each frame of video image. Each anti-shake parameter is obtained by the video acquisition end based on n initial positions. The n initial positions are obtained by the video acquisition end when the target object in the n consecutive video images is located within the image core area. The n initial positions are used to characterize the distance of the target object relative to the center point of the photosensitive unit in the n consecutive video images. The image core area is the area within a preset distance range from the center point of the photosensitive unit, and n is an integer greater than 1. Anti-shake processing is performed based on the anti-shake parameters corresponding to each received video frame.

[0009] Optionally, the image stabilization processing based on the stabilization parameters corresponding to each received video frame includes: Based on the anti-shake parameters corresponding to the received m consecutive video images, determine the m target positions and m jitter coefficients of the target object in the m consecutive video images, where m is any integer greater than 1; Determine the m display positions of the m target positions on the video display terminal, and based on the m display positions, determine the geometric center of the target object in the Mth video image of the consecutive m video images, wherein the m target positions correspond one-to-one with the m display positions, and M is a positive integer less than or equal to m; Based on the geometric center and the stabilization parameters corresponding to the Mth frame video image, the Mth frame video image is subjected to stabilization processing.

[0010] Thirdly, embodiments of this application also provide a video processing apparatus applied to a video acquisition end, the apparatus comprising: The judgment module is used to sequentially acquire each frame of video image using the photosensitive unit, and sequentially judge whether the target objects in n consecutive video images are all located within the image core area, wherein the image core area is the area within a preset distance range from the center point of the photosensitive unit, and n is any integer greater than 1. The acquisition module is used to acquire n initial positions of the target object relative to the center point of the photosensitive unit in the n consecutive video images when the target object in the n consecutive video images is located within the core area of ​​the image. The n initial positions correspond one-to-one with the n consecutive video images. The first image stabilization module is used to perform image stabilization processing based on the n initial positions, obtain the image stabilization parameters corresponding to the Nth video image in the n consecutive video images, and send the image stabilization parameters to the video display terminal so that the video display terminal can perform image stabilization processing based on the image stabilization parameters corresponding to each received video image, where N is a positive integer less than or equal to n.

[0011] Fourthly, embodiments of this application also provide a video processing apparatus applied to a video display end, the apparatus comprising: The receiving module is used to receive the anti-shake parameters corresponding to each frame of video image sent by the video acquisition end. Each anti-shake parameter is obtained by the video acquisition end based on n initial positions. The n initial positions are obtained by the video acquisition end when the target object is located within the image core area in n consecutive video images. The n initial positions are used to characterize the distance of the target object relative to the center point of the photosensitive unit in the n consecutive video images. The image core area is the area within a preset distance range from the center point of the photosensitive unit, and n is an integer greater than 1. The second image stabilization module is used to perform image stabilization processing based on the image stabilization parameters corresponding to each received video frame.

[0012] Fifthly, embodiments of this application also provide a video processing system, the system including a video acquisition end and a video display end; The video acquisition terminal is used to execute the video processing method described in the first aspect; The video display terminal is used to execute the video processing method described in the second aspect.

[0013] Sixthly, embodiments of this application also provide a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the video processing method described in the first aspect, or when executed, implements the video processing method described in the second aspect.

[0014] Compared with the prior art, the technical solution provided in this application has the following advantages: The method provided in this application uses a photosensitive unit to sequentially acquire each frame of video images and sequentially determines whether the target objects in n consecutive video images are all located within the image core area, wherein the image core area is a region within a preset distance range from the center point of the photosensitive unit, and n is any integer greater than 1; when the target objects in the n consecutive video images are all located within the image core area, n initial positions of the target objects relative to the center point of the photosensitive unit in the n consecutive video images are obtained, wherein the n initial positions correspond one-to-one with the n consecutive video images; anti-shake processing is performed based on the n initial positions to obtain the anti-shake parameters corresponding to the Nth video image in the n consecutive video images, and the anti-shake parameters are sent to the video display end so that the video display end can perform anti-shake processing based on the anti-shake parameters corresponding to each received frame of video images, where N is a positive integer less than or equal to n. In the above manner, the video acquisition end can perform image stabilization based on the n initial positions of the target object relative to the center point of the photosensitive unit in n consecutive video images, thereby obtaining the image stabilization parameters corresponding to the Nth video image in the n consecutive video images. Following the same method, the image stabilization parameters corresponding to each video image are obtained and sent to the video display end. In this way, the video display end can perform image stabilization based on the image stabilization parameters corresponding to each received video image, thereby effectively reducing the problem of video image jitter caused by the shaking of the video acquisition end, improving the video image quality, and thus improving the viewing experience of the video viewer. Attached Figure Description

[0015] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0016] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] One embodiment or practice is illustrated by way of example with the corresponding pictures in the accompanying drawings. These illustrative descriptions do not constitute a limitation on the embodiments. Elements with the same reference numerals in the drawings are represented as similar elements. Unless otherwise stated, the figures in the drawings are not to be limited by scale.

[0018] Figure 1 A flowchart illustrating a video processing method provided in an embodiment of this application; Figure 2 A schematic diagram of an image core area and an image compensation area provided in an embodiment of this application; Figure 3 A flowchart illustrating yet another video processing method provided in an embodiment of this application; Figure 4 A schematic diagram of three consecutive video frames provided in an embodiment of this application; Figure 5 A schematic diagram of a geometric center provided for an embodiment of this application; Figure 6 This is a schematic diagram of the structure of a video processing apparatus provided in an embodiment of this application; Figure 7 This is a schematic diagram of the structure of another video processing device provided in the embodiments of this application; Figure 8 This is a schematic diagram of the structure of a video processing system provided in an embodiment of this application. Detailed Implementation

[0019] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0020] The following disclosure provides numerous different embodiments or examples for implementing various structures of this application. To simplify the disclosure, specific examples of components and arrangements are described below. These are merely examples and are not intended to limit the scope of this application. Furthermore, reference numerals and / or letters may be repeated in different examples. Such repetition is for simplification and clarity and does not in itself indicate a relationship between the various embodiments and / or arrangements discussed.

[0021] To address the problem of video image jitter caused by video acquisition device judder, which can easily cause dizziness for viewers, this application provides a video processing method, apparatus, system, and computer-readable storage medium that can reduce the impact of video acquisition device judder on video image.

[0022] See Figure 1 , Figure 1 This is a flowchart illustrating a video processing method provided in an embodiment of this application. Figure 1 As shown, this video processing method can be applied to the video acquisition end, and the video processing method may include the following steps: Step S101: Use the photosensitive unit to sequentially acquire each frame of video image, and sequentially determine whether the target objects in the consecutive n frames of video images are all located within the image core area, where the image core area is the area within a preset distance range from the center point of the photosensitive unit, and n is any integer greater than 1.

[0023] Specifically, the aforementioned photosensitive unit can be implemented using a complementary metal-oxide-semiconductor (CMOS) image sensor or a charge-coupled device (CCD) image sensor; this application embodiment does not impose specific limitations. The aforementioned target object can be any object to be photographed, such as a person, product, or building. The aforementioned image core area can be an area within a preset distance range from the center point of the photosensitive unit, wherein this preset distance can be set according to actual needs. For example, as... Figure 2As shown, the image core area 201 can be a circular region with the center point of the photosensitive unit 200 as the center and a preset distance r as the radius. An image compensation area 202 can be defined around the image core area 201, and the image compensation area 202 can be a ring with a diameter of Rr. Of course, as in other embodiments, the image core area 201 and the image compensation area 202 can also have other shapes and sizes; this application embodiment does not impose specific limitations. When the object 203 falls into the image core area 201, it indicates that the shaking is small, and the image stabilization process in this application can be used to solve the image shaking problem. When the object 203 falls into the image compensation area 202, it indicates that the shaking is large. At this time, traditional compensation algorithms such as Digital Image Stabilization (DIS), Electronic Image Stabilization (EIS), and Artificial Intelligence Stabilization (AIS) are needed to solve the image shaking problem. Therefore, this application focuses on the image stabilization process when the object falls into the image core area.

[0024] The value of n can be any positive integer such as 2 or 3, and can be set according to actual needs. As an optional implementation, the value of n can be 2, that is, the video acquisition end sequentially determines whether the target object in two consecutive video frames is located within the image core area. If the target object in two consecutive video frames is located within the image core area, then step S102 is executed. If the target object in two consecutive video frames is not located within the image core area, then no image stabilization processing is performed on these two consecutive video frames. Instead, the most recently acquired image stabilization parameters are used to replace the image stabilization parameters corresponding to the Nth video frame in these two consecutive video frames, and then sent to the video display end.

[0025] Step S102: When the target object in n consecutive video images is located within the core area of ​​the image, obtain n initial positions of the target object relative to the center point of the photosensitive unit in the n consecutive video images, wherein the n initial positions correspond one-to-one with the n consecutive video images.

[0026] Specifically, the initial position here corresponds one-to-one with the video image, that is, in each frame of the video image, there is an initial position of the target object relative to the center point of the photosensitive unit.

[0027] The video acquisition end can establish a coordinate system with the center point of the photosensitive unit as the origin, and determine the coordinate position of the target object in each frame of video image relative to the origin, thereby obtaining n initial positions.

[0028] Step S103: Perform image stabilization processing based on n initial positions to obtain the image stabilization parameters corresponding to the Nth frame of video images in n consecutive frames, and send the image stabilization parameters to the video display end so that the video display end can perform image stabilization processing based on the image stabilization parameters corresponding to each received frame of video images, where N is a positive integer less than or equal to n.

[0029] Specifically, N can be a positive integer less than or equal to n. For example, when n is 3, the Nth frame of video can be the first, second, or third frame of a consecutive three-frame video image. When N is 1, the video acquisition end can acquire the first, second, and third frames of the video as three consecutive frames, perform image stabilization on the three initial positions of these three consecutive frames, and send the obtained image stabilization parameters as the image stabilization parameters of the first frame to the video display end. Then, the second, third, and fourth frames of the video are acquired as new three consecutive frames, and image stabilization is performed on the three initial positions of these new three consecutive frames. The obtained image stabilization parameters are sent as the image stabilization parameters of the second frame to the video display end, and so on, to obtain the image stabilization parameters for each frame of video image, and then send the image stabilization parameters of each frame of video image to the video display end in sequence. When N is 2, the video acquisition end can acquire the first, second, and third frames of the video as three consecutive video frames, perform image stabilization on the three initial positions of these three consecutive video frames, and send the obtained image stabilization parameters as the image stabilization parameters of the second frame (at this time, the image stabilization parameters of the first frame are reused from the image stabilization parameters of the second frame) to the video display end. Then, the second, third, and fourth frames of the video are acquired as three new consecutive video frames, and image stabilization is performed on the three initial positions of these three new consecutive video frames. The obtained image stabilization parameters are sent as the image stabilization parameters of the third frame to the video display end, and so on, to obtain the image stabilization parameters of each frame of the video image, and send the image stabilization parameters of each frame of the video image to the video display end in sequence. When N is 3, the video acquisition end can acquire the first, second, and third frames of the video as three consecutive video frames, perform image stabilization on the three initial positions of these three consecutive video frames, and send the obtained image stabilization parameters as the image stabilization parameters of the third video frame (at this time, the image stabilization parameters of the first and second video frames reuse the image stabilization parameters of the third video frame) to the video display end. Then, the second, third, and fourth frames of the video are acquired as new three consecutive video frames, and image stabilization is performed on the three initial positions of these new three consecutive video frames. The obtained image stabilization parameters are sent as the image stabilization parameters of the fourth video frame to the video display end, and so on, to obtain the image stabilization parameters of each video frame, and send the image stabilization parameters of each video frame to the video display end in sequence.

[0030] For example, when n is 2, the Nth frame of video can be either the first or second frame of a consecutive two-frame video sequence. When N is 1, the video acquisition end can acquire the first and second frames of the video as two consecutive frames, perform image stabilization on the two initial positions of these two frames, and send the resulting image stabilization parameters as the image stabilization parameters for the first frame to the video display end. Then, the second and third frames of the video are acquired as new two consecutive frames, and image stabilization is performed on the two initial positions of these new two consecutive frames. The resulting image stabilization parameters are then sent as the image stabilization parameters for the second frame to the video display end, and so on, to obtain the image stabilization parameters for each frame of video, and then send the image stabilization parameters for each frame of video to the video display end in sequence. When N is 2, the video acquisition end can acquire the first and second frames of the video as two consecutive video frames, perform image stabilization on the two initial positions of these two consecutive video frames, and send the obtained image stabilization parameters as the image stabilization parameters of the second frame (at this time, the image stabilization parameters of the first frame are reused from the image stabilization parameters of the second frame) to the video display end. Then, the second and third frames of the video are acquired as two new consecutive video frames, and image stabilization is performed on the two initial positions of these two new consecutive video frames. The obtained image stabilization parameters are sent as the image stabilization parameters of the third frame to the video display end, and so on, to obtain the image stabilization parameters of each frame of the video image, and send the image stabilization parameters of each frame of the video image to the video display end in sequence.

[0031] In the above manner, the video acquisition end can perform image stabilization based on the n initial positions of the target object relative to the center point of the photosensitive unit in n consecutive video images, thereby obtaining the image stabilization parameters corresponding to the Nth video image in the n consecutive video images. Following the same method, the image stabilization parameters corresponding to each video image are obtained and sent to the video display end. In this way, the video display end can perform image stabilization based on the image stabilization parameters corresponding to each received video image, thereby effectively reducing the problem of video image jitter caused by the shaking of the video acquisition end, improving the video image quality, and thus improving the viewing experience of the video viewer.

[0032] In an optional embodiment, step S103 above, which involves performing image stabilization based on n initial positions to obtain the image stabilization parameters corresponding to the Nth frame of video images in a series of n frames, includes: Calculate the center points of n initial positions, and determine the center points of the n initial positions as the target position of the target object in the Nth frame of the video image; Based on n initial positions, the jitter coefficient is determined, where the jitter coefficient is used to characterize the degree of jitter between consecutive n frames of video images; The target position and the jitter coefficient are determined as the stabilization parameters.

[0033] Specifically, after acquiring n initial positions of the target object relative to the center point of the photosensitive unit in n consecutive video images, the video acquisition end can calculate the center point between each of these n initial positions, and determine the center point of the n initial positions as the target position of the target object in the Nth video image. Based on the n initial positions, the jitter coefficient can be determined, and then the target position and the jitter coefficient can be determined as the anti-shake parameters.

[0034] When calculating the center point of n initial positions, the initial positions of the target object in two consecutive video frames are respectively ( , )and( , When ), the x-coordinates of the center points of these two initial positions can be calculated as ( ) / 2, the vertical axis is ( If ) / 2, then the target position of the target object in the Nth frame of the two consecutive video frames is [( ) / 2, ( For example, when the initial positions of the target object in three consecutive video frames are respectively ( ) / 2]. , ), ( , )and( , When ), the x-coordinates of the center points of these three initial positions can be calculated as ( + ) / 2, the vertical axis is ( + If ) / 2, then the target position of the target object in the Nth frame of these three consecutive video frames is [( + ) / 2, ( + ) / 2).

[0035] Using the above method, the video acquisition end can accurately obtain the stabilization parameters corresponding to each frame of video image, which is convenient for subsequent transmission to the video display end for stabilization processing.

[0036] In one optional embodiment, the jitter coefficient is calculated using the following formula: ; in, Indicates the jitter coefficient, ( , Let represent the i-th initial position among n initial positions, where i∈[1,n].

[0037] Specifically, when the initial positions of the target object in two consecutive video frames are respectively ( , )and( , When ), the jitter coefficient When the initial positions of the target object in three consecutive video frames are respectively ( , ), ( , )and( , When ), the jitter coefficient By analogy, we can obtain the jitter coefficient when n is any positive integer.

[0038] Of course, as other alternative implementation methods, the above formula can also be simply modified, such as... ,or The jitter coefficient can be calculated using a formula (where a is a preset coefficient and b is a preset constant). Other formulas can also be used, such as... Then we calculate the jitter coefficient.

[0039] The jitter coefficient can be accurately calculated using the above method, which facilitates subsequent image stabilization processing based on this coefficient.

[0040] In an optional embodiment, after step 101, which sequentially determines whether the target objects in n consecutive video frames are all located within the image core region, the method further includes: If the target object in n consecutive video frames is not located within the core area of ​​the image, the latest image stabilization parameter obtained before the n consecutive video frames is used to replace the image stabilization parameter corresponding to the Nth video frame and sent to the video display terminal.

[0041] Specifically, when the target object in a series of n consecutive video frames is partially or entirely outside the image core region, the video acquisition end does not need to perform image stabilization on these n consecutive video frames. Instead, it uses the most recently acquired stabilization parameters from the previous n frames as the stabilization parameters for the Nth frame and sends them to the video display end. In other words, when the video acquisition end detects that the target object in a frame of the series of n consecutive video frames is located within the image compensation region, it can use the previously acquired stabilization parameters as the stabilization parameters for the Nth frame and send them to the video display end. Simultaneously, the video display end can use traditional compensation algorithms to perform jitter compensation on the Nth frame.

[0042] In this way, multiple methods can be used to stabilize video images with large shakiness, thereby achieving a better stabilization effect.

[0043] It should be noted that when the video acquisition terminal detects that the target object in the video image for a consecutive preset number of frames is not located in the core area of ​​the image, it indicates that the video acquisition terminal may have switched the shooting angle. Therefore, the video acquisition terminal can select a new object as the target object and repeat the above steps S101 to S104.

[0044] See Figure 3 , Figure 3 This is a flowchart illustrating another video processing method provided in an embodiment of this application. Figure 3 As shown, this video processing method can be applied to a video display end, and the video processing method may include the following steps: Step S301: Receive the anti-shake parameters corresponding to each frame of video image sent by the video acquisition end. Each anti-shake parameter is obtained by the video acquisition end based on n initial positions. The n initial positions are obtained by the video acquisition end when the target object is located within the core area of ​​the image in n consecutive video images. The n initial positions are used to characterize the distance of the target object relative to the center point of the photosensitive unit in the n consecutive video images. The core area of ​​the image is the area within a preset distance range from the center point of the photosensitive unit. n is an integer greater than 1.

[0045] Specifically, the video acquisition terminal mentioned above is the aforementioned... Figure 1 In the illustrated embodiment, the video acquisition end can sequentially acquire each frame of video image using a photosensitive unit, and sequentially determine whether the target object in each of the n consecutive video images is located within the image core area. If the target object in each of the n consecutive video images is located within the image core area, then n initial positions of the target object relative to the center point of the photosensitive unit in the n consecutive video images can be obtained. Based on the n initial positions, image stabilization processing is performed to obtain the image stabilization parameters corresponding to the Nth video image in the n consecutive video images, and then the image stabilization parameters are sent to the video display end. In this way, the video display end can receive the image stabilization parameters corresponding to each frame of video image sent by the video acquisition end. The image stabilization parameters here may include, but are not limited to, parameters such as the target position of the target object in the video image and the jitter coefficient.

[0046] Step S302: Perform image stabilization processing based on the image stabilization parameters corresponding to each received video frame.

[0047] Specifically, after receiving the anti-shake parameters corresponding to each frame of video image sent by the video acquisition end, the video display end can perform anti-shake processing based on the received anti-shake parameters, thereby improving the stability of the video image.

[0048] In the above manner, the video acquisition end can perform image stabilization based on the n initial positions of the target object relative to the center point of the photosensitive unit in n consecutive video images, thereby obtaining the image stabilization parameters corresponding to the Nth video image in the n consecutive video images. Following the same method, the image stabilization parameters corresponding to each video image are obtained and sent to the video display end. In this way, the video display end can perform image stabilization based on the image stabilization parameters corresponding to each received video image, thereby effectively reducing the problem of video image jitter caused by the shaking of the video acquisition end, improving the video image quality, and thus improving the viewing experience of the video viewer.

[0049] In an optional embodiment, step S302, which involves performing image stabilization processing based on the image stabilization parameters corresponding to each received video frame, includes: Based on the anti-shake parameters corresponding to the received m consecutive video images, determine the m target positions and m jitter coefficients of the target object in the m consecutive video images, where m is any integer greater than 1; Determine the m display positions of m target positions on the video display end, and based on the m display positions, determine the geometric center of the target object in the Mth video image in a series of m video images, where the m target positions correspond one-to-one with the m display positions, and M is a positive integer less than or equal to m. Based on the geometric center and the stabilization parameters corresponding to the Mth frame video image, stabilization processing is performed on the Mth frame video image.

[0050] Specifically, the value of m can be any positive integer such as 2 or 3, and can be set according to actual needs. As an optional implementation, the value of m can be 3, that is, the video display end can determine the three target positions and three jitter coefficients of the target object in the three consecutive video frames based on the anti-shake parameters corresponding to the three consecutive video frames received. Then, it can determine the three display positions of these three target positions on the video display end, and based on these three display positions, determine the geometric center of the target object in the Mth video frame in the three consecutive video frames. Then, it can perform anti-shake processing on the Mth video frame based on the geometric center and the anti-shake parameters corresponding to the Mth video frame. As another optional implementation, the value of m can be 2. That is, the video display end can determine two target positions and two jitter coefficients of the target object in the two consecutive video frames based on the anti-shake parameters corresponding to the two received consecutive video frames. Then, it can determine two display positions of these two target positions on the video display end, and based on these two display positions, determine the geometric center of the target object in the Mth video frame in the two consecutive video frames. Then, it can perform anti-shake processing on the Mth video frame based on the geometric center and the anti-shake parameters corresponding to the Mth video frame. Of course, m can also be other positive integers, which will not be elaborated here.

[0051] It should be noted that M can be a positive integer less than or equal to m. For example, when m is 3, the Mth frame of video can be the first, second, or third frame of a consecutive 3-frame video image. When M is 1, the video display end can use the stabilization parameters corresponding to the first, second, and third frames of the received video image as the stabilization parameters for the consecutive 3 frames. It then determines the three target positions and three jitter coefficients of the target object in these three consecutive frames, determines the three display positions of these three target positions on the video display end, obtains the geometric center of the target object in the first frame of the video image, and then performs stabilization processing on the first frame of the video image based on the geometric center and jitter coefficients. Next, the video display end can use the stabilization parameters corresponding to the second, third, and fourth video frames as the stabilization parameters for three consecutive video frames. It then determines the three target positions and three jitter coefficients of the target object in these three consecutive video frames, and determines the three display positions of these three target positions on the video display end. This yields the geometric center of the target object in the second video frame. Based on the geometric center and jitter coefficient of the second video frame, stabilization processing is performed on the second video frame. This process is repeated for each video frame.

[0052] When M is 2, the video display end can use the anti-shake parameters corresponding to the first, second, and third frames of the received video image as the anti-shake parameters for three consecutive frames of video image. It then determines the three target positions and three jitter coefficients of the target object in these three consecutive frames of video image, and determines the three display positions of these three target positions on the video display end. This yields the geometric center of the target object in the second frame of video image (the geometric center of the first frame of video image can be reused from the geometric center of the second frame of video image). Based on the geometric center and jitter coefficients of the first frame of video image, anti-shake processing can be performed on the first frame of video image, and the same applies to the second frame of video image. Next, the video display end can use the stabilization parameters corresponding to the second, third, and fourth video frames as the stabilization parameters for three consecutive video frames. It then determines the three target positions and three jitter coefficients of the target object in these three consecutive video frames, and determines the three display positions of these three target positions on the video display end. This yields the geometric center of the target object in the second video frame. Based on the geometric center and jitter coefficient of the second video frame, stabilization processing is performed on the second video frame. This process is repeated for each video frame.

[0053] When M is 3, the video display end can use the stabilization parameters corresponding to the first, second, and third frames of the received video image as the stabilization parameters for three consecutive frames. It then determines the three target positions and three jitter coefficients of the target object in these three consecutive frames, and determines the three display positions of these three target positions on the video display end. This yields the geometric center of the target object in the third frame (the geometric center corresponding to the first and second frames can be reused from the geometric center corresponding to the third frame). Based on the geometric center and jitter coefficients of the first frame, stabilization processing can be applied to the first frame; similarly, stabilization processing can be applied to the second frame; and finally, stabilization processing can be applied to the third frame. Next, the video display end can use the stabilization parameters corresponding to the second, third, and fourth video frames as the stabilization parameters for three consecutive video frames. It then determines the three target positions and three jitter coefficients of the target object in these three consecutive video frames, and determines the three display positions of these three target positions on the video display end. This yields the geometric center of the target object in the fourth video frame. Based on the geometric center and jitter coefficient of the fourth video frame, stabilization processing is performed on the fourth video frame. This process is repeated for each video frame.

[0054] like Figure 4 and Figure 5 As shown, m is 3 and M is 2, meaning the video display uses the stabilization parameters of three consecutive video frames to perform stabilization on the middle frame. Figure 4 As shown, assume that the display positions of the target object in these three consecutive video frames on the video display are respectively ( , ), ( , )and( , Based on these three display positions, the geometric center of the target object in the second frame of the three consecutive video frames can be determined as [( ) / 3, ( ) / 3 ].

[0055] The above methods can be used to stabilize each frame of video images, reducing image drift caused by shaking and discomfort such as visual dizziness caused by shaking.

[0056] See Figure 6, Figure 6 This is a schematic diagram of the structure of a video processing device provided in an embodiment of this application, as shown below. Figure 6 As shown, the video processing device 600 can be applied to a video acquisition end, and the video processing device 600 includes: The judgment module 601 is used to sequentially acquire each frame of video images using the photosensitive unit, and sequentially judge whether the target objects in the consecutive n frames of video images are all located within the image core area, wherein the image core area is the area within a preset distance range from the center point of the photosensitive unit, and n is any integer greater than 1. The acquisition module 602 is used to acquire n initial positions of the target object relative to the center point of the photosensitive unit in the consecutive n video images when the target object is located in the core area of ​​the image in the consecutive n video images. The n initial positions correspond one-to-one with the consecutive n video images. The first image stabilization module 603 is used to perform image stabilization processing based on n initial positions, obtain the image stabilization parameters corresponding to the Nth frame of video images in n consecutive frames, and send the image stabilization parameters to the video display end so that the video display end can perform image stabilization processing based on the image stabilization parameters corresponding to each received frame of video images, where N is a positive integer less than or equal to n.

[0057] Furthermore, the first image stabilization module 603 includes: The first determination submodule is used to calculate the center points of n initial positions and determine the center points of the n initial positions as the target position of the target object in the Nth frame of the video image; The second determining submodule is used to determine the jitter coefficient based on n initial positions, wherein the jitter coefficient is used to characterize the degree of jitter between consecutive n frames of video images; The third determination submodule is used to determine the target position and jitter coefficient as the stabilization parameters.

[0058] Furthermore, the formula for calculating the jitter coefficient is as follows: ; in, Indicates the jitter coefficient, ( , Let represent the i-th initial position among n initial positions, where i∈[1,n].

[0059] Furthermore, the video processing device 600 also includes: The replacement module, when the target object in n consecutive video frames is not located within the image core area, uses the most recently acquired anti-shake parameters from the previous n consecutive video frames to replace the anti-shake parameters corresponding to the Nth video frame, and sends them to the video display end. It should be noted that the video processing device 600 can achieve the aforementioned... Figure 1 The video processing method provided in the illustrated embodiment can achieve the same technical effect, and will not be described in detail here.

[0060] See Figure 7 , Figure 7 This is a schematic diagram of the structure of another video processing device provided in the embodiments of this application, as shown below. Figure 7 As shown, the video processing device 700 can be applied to a video display end, and the video processing device 700 includes: The receiving module 701 is used to receive the anti-shake parameters corresponding to each frame of video image sent by the video acquisition end. Each anti-shake parameter is obtained by the video acquisition end based on n initial positions. The n initial positions are obtained by the video acquisition end when the target object is located within the image core area in n consecutive video images. The n initial positions are used to characterize the distance of the target object relative to the center point of the photosensitive unit in the n consecutive video images. The image core area is the area within a preset distance range from the center point of the photosensitive unit. n is an integer greater than 1. The second image stabilization module 702 is used to perform image stabilization processing based on the image stabilization parameters corresponding to each received video frame.

[0061] Furthermore, the second image stabilization module 702 includes: The fourth determination submodule is used to determine the m target positions and m jitter coefficients of the target object in the m consecutive video images based on the anti-shake parameters corresponding to the received m consecutive video images, where m is any integer greater than 1; The fifth determination submodule is used to determine the m display positions of m target positions on the video display end, and based on the m display positions, determine the geometric center of the target object in the Mth video image in the consecutive m video images, where the m target positions correspond one-to-one with the m display positions, and M is a positive integer less than or equal to m; The image stabilization submodule is used to perform image stabilization on the Mth frame video image based on the geometric center and the image stabilization parameters corresponding to the Mth frame video image.

[0062] It should be noted that the video processing device 700 can achieve the aforementioned... Figure 3 The video processing method provided in the illustrated embodiment can achieve the same technical effect, and will not be described in detail here.

[0063] See Figure 8 , Figure 8 This is a schematic diagram of the structure of a video processing system provided in an embodiment of this application, as shown below. Figure 8As shown, the video processing system 800 may include a video acquisition terminal 801 and a video display terminal 802; Among them, the video acquisition terminal 801 is used to achieve the aforementioned Figure 1 The video processing method provided in the illustrated embodiment can achieve the same technical effect, and will not be described in detail here; The video display terminal 802 is used to achieve the aforementioned Figure 3 The video processing method provided in the illustrated embodiment can achieve the same technical effect, and will not be described in detail here.

[0064] In addition, embodiments of this application also provide a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the video processing method provided in the foregoing method embodiments.

[0065] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0066] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented using software plus a general-purpose hardware platform, or of course, using hardware. Based on this understanding, the above technical solutions, in essence or the parts that contribute to the related technology, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0067] It should be understood that the terminology used herein is for the purpose of describing particular exemplary embodiments only and is not intended to be limiting. Unless the context clearly indicates otherwise, the singular forms “a,” “an,” and “described” as used herein may also mean including the plural forms. The terms “comprising,” “including,” “containing,” and “having” are inclusive and therefore indicate the presence of the stated features, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, elements, components, and / or combinations thereof. The method steps, processes, and operations described herein are not construed as requiring them to be performed in a particular order described or illustrated unless the order of performance is explicitly indicated. It should also be understood that additional or alternative steps may be used.

[0068] The above description is merely a specific embodiment of this application, enabling those skilled in the art to understand or implement this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.

Claims

1. A video processing method, characterized in that, Applied to a video acquisition terminal, the method includes: The system uses a photosensitive unit to sequentially acquire each frame of video image and sequentially determines whether the target object in n consecutive video images is located within the image core area. The image core area is the region within a preset distance from the center point of the photosensitive unit, and n is any integer greater than 1. When the target object in the n consecutive video frames is located within the core area of ​​the image, n initial positions of the target object relative to the center point of the photosensitive unit in the n consecutive video frames are obtained, wherein the n initial positions correspond one-to-one with the n consecutive video frames; Based on the n initial positions, image stabilization is performed to obtain the image stabilization parameters corresponding to the Nth video image in the n consecutive video images. The image stabilization parameters are then sent to the video display terminal so that the video display terminal can perform image stabilization based on the image stabilization parameters corresponding to each received video image. N is a positive integer less than or equal to n.

2. The method according to claim 1, characterized in that, The stabilization process based on the n initial positions, to obtain the stabilization parameters corresponding to the Nth video frame in the consecutive n video frames, includes: Calculate the center point of the n initial positions, and determine the center point of the n initial positions as the target position of the target object in the Nth frame of the video image; Based on the n initial positions, a jitter coefficient is determined, wherein the jitter coefficient is used to characterize the degree of jitter between the n consecutive video frames; The target position and the jitter coefficient are determined as the anti-shake parameters.

3. The method according to claim 2, characterized in that, The formula for calculating the jitter coefficient is as follows: ; in, Indicates the jitter coefficient, ( , ) represents the i-th initial position among the n initial positions, where i∈[1,n].

4. The method according to claim 2, characterized in that, After determining whether the target objects in n consecutive video frames are all located within the image core region, the method further includes: If the target object in the n consecutive video frames is not located within the core area of ​​the image, the latest image stabilization parameter obtained before the n consecutive video frames is used to replace the image stabilization parameter corresponding to the Nth video frame and sent to the video display terminal.

5. A video processing method, characterized in that, Applied to a video display end, the method includes: The receiving end sends anti-shake parameters corresponding to each frame of video image. Each anti-shake parameter is obtained by the video acquisition end based on n initial positions. The n initial positions are obtained by the video acquisition end when the target object in the n consecutive video images is located within the image core area. The n initial positions are used to characterize the distance of the target object relative to the center point of the photosensitive unit in the n consecutive video images. The image core area is the area within a preset distance range from the center point of the photosensitive unit, and n is an integer greater than 1. Anti-shake processing is performed based on the anti-shake parameters corresponding to each received video frame.

6. The method according to claim 5, characterized in that, The image stabilization process based on the stabilization parameters corresponding to each received video frame includes: Based on the anti-shake parameters corresponding to the received m consecutive video images, determine the m target positions and m jitter coefficients of the target object in the m consecutive video images, where m is any integer greater than 1; Determine the m display positions of the m target positions on the video display terminal, and based on the m display positions, determine the geometric center of the target object in the Mth video image of the consecutive m video images, wherein the m target positions correspond one-to-one with the m display positions, and M is a positive integer less than or equal to m; Based on the geometric center and the stabilization parameters corresponding to the Mth frame video image, the Mth frame video image is subjected to stabilization processing.

7. A video processing apparatus, characterized in that, The device, applied to a video acquisition terminal, includes: The judgment module is used to sequentially acquire each frame of video image using the photosensitive unit, and sequentially judge whether the target objects in n consecutive video images are all located within the image core area, wherein the image core area is the area within a preset distance range from the center point of the photosensitive unit, and n is any integer greater than 1. The acquisition module is used to acquire n initial positions of the target object relative to the center point of the photosensitive unit in the n consecutive video images when the target object in the n consecutive video images is located within the core area of ​​the image. The n initial positions correspond one-to-one with the n consecutive video images. The first image stabilization module is used to perform image stabilization processing based on the n initial positions, obtain the image stabilization parameters corresponding to the Nth video image in the n consecutive video images, and send the image stabilization parameters to the video display terminal so that the video display terminal can perform image stabilization processing based on the image stabilization parameters corresponding to each received video image, where N is a positive integer less than or equal to n.

8. A video processing apparatus, characterized in that, The device, applied to a video display end, includes: The receiving module is used to receive the anti-shake parameters corresponding to each frame of video image sent by the video acquisition end. Each anti-shake parameter is obtained by the video acquisition end based on n initial positions. The n initial positions are obtained by the video acquisition end when the target object is located within the image core area in n consecutive video images. The n initial positions are used to characterize the distance of the target object relative to the center point of the photosensitive unit in the n consecutive video images. The image core area is the area within a preset distance range from the center point of the photosensitive unit, and n is an integer greater than 1. The second image stabilization module is used to perform image stabilization processing based on the image stabilization parameters corresponding to each received video frame.

9. A video processing system, characterized in that, The system includes a video acquisition terminal and a video display terminal; The video acquisition terminal is used to execute the video processing method according to any one of claims 1-4; The video display terminal is used to execute the video processing method according to any one of claims 5-6.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the video processing method of any one of claims 1-4, or when it is executed, it implements the video processing method of any one of claims 5-6.