Video processing method, video processing device, electronic device and computer-readable storage medium

By identifying and adjusting the FOV of target regions in video frames, the method simplifies the generation of sliding zoom effects, addressing complexity and ensuring high robustness and real-time performance.

CN111756996BActive Publication Date: 2025-07-15ARASHI VISION INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010556962.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-06-18
Publication Date
2025-07-15
Estimated Expiration
2040-06-18

AI Technical Summary

Technical Problem

The existing sliding zoom video shooting methods are complex, relying on complex mechanical structures or artificial controls, making it difficult to simulate real perspective effects, especially when the subject moves.

Method used

By identifying the target box of the area of interest in the video frame and adjusting its FOV value to achieve the ideal ratio, combined with rendering, generate a sliding zoom image, simplifying operation and improving robustness.

Benefits of technology

It realizes the sliding zoom effect with simple operation and low computational complexity, with good real-time and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111756996B_ABST
    Figure CN111756996B_ABST
Patent Text Reader

Abstract

The present invention provides a video processing method, which includes: obtaining video frames of a video to be processed; identifying target boxes of the same region of interest in each video frame; adjusting the FOV value of the target boxes in the video frames to adjust the proportion of the target boxes in the video frames to an ideal ratio; and rendering the adjusted video frames to generate an image with a sliding zoom effect. The present invention only needs to determine the target box of the region of interest at the initial moment, then adjusts the FOV value of the target boxes in the video frames to adjust the proportion of the target boxes of the region of interest in the video frames to a predetermined value, and then generates an image with a sliding zoom effect by rendering the adjusted video frames. It has the advantages of simple operation, low computational complexity, good real-time performance and high robustness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the technical field of video processing, and particularly to a video processing method, a video processing device, an electronic device, and a computer-readable storage medium. Background Art

[0002] Dolly Zoom refers to a visual effect in which, during video shooting, zooming and camera movement are carried out simultaneously, but the space occupied by a specified target object in the frame remains stationary.

[0003] Currently, there are several ways to shoot Dolly Zoom videos: First, for cameras with optical zoom lenses, through a slide rail, a complex mechanical structure, or the techniques of a photographer, during the movement, the zoom is controlled according to the movement distance to match the movement process and the zoom process; Second, without using zoom during recording, but during post-editing, by means of cropping, to ensure that the size of the target object remains unchanged. Third, by using a depth camera or 3D reconstruction technology to separate the foreground and background, keeping the foreground unchanged and simulating the perspective change of the background.

[0004] In these methods, by fixing the ratio of the pixel height of the target object to the total number of pixels in the photo to adjust the optical zoom, a relatively complex and large-sized zoom lens is required and it is relatively dependent on the mechanical structure of the zoom slide rail or manual control; based on the method of separating the foreground and background, it is difficult to simulate the real perspective effect caused by the change of object distance, and if the subject is moving, it is difficult to reflect this movement by separating the foreground and background. Therefore, the means to achieve the Dolly Zoom effect are currently very complex. Summary of the Invention

[0005] The purpose of the present invention is to provide a video processing method, a video processing device, an electronic device, and a computer-readable storage medium, aiming to solve the problem that the existing Dolly Zoom processing is too complex.

[0006] In a first aspect, the present invention provides a video processing method, which includes: obtaining video frames of a video to be processed; identifying target frames of the same region of interest in each video frame; adjusting the FOV value of the target frame in the video frame to adjust the proportion of the target frame in the video frame to an ideal ratio; rendering the adjusted video frame to generate a Dolly Zoom image.

[0007] In a second aspect, the present invention provides a video processing device, which includes: an obtaining module for obtaining video frames of a video to be processed; an identifying module for identifying target frames of the region of interest in the video frame; an adjusting module for adjusting the FOV value of the target frame in the video frame to adjust the proportion of the target frame of the region of interest in the video frame to a predetermined value; a rendering module for rendering the adjusted video frame to generate a Dolly Zoom image.

[0008] In a third aspect, the present invention provides an electronic device, including a processor, a memory, and a computer program stored on the memory and executable on the processor. The processor executes the computer program to implement the above-mentioned video processing method.

[0009] In a fourth aspect, a computer-readable storage medium is characterized in that a computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the above-mentioned video processing method is implemented.

[0010] The present invention only needs to determine the target box of interest at the initial moment, then adjusts the FOV value of the target box in the video frame to adjust the proportion of the target box in the region of interest to a predetermined value in the video frame, and then generates a sliding zoom image by rendering the adjusted video frame, which has the advantages of simple operation, low computational complexity, good real-time performance and high robustness. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1 is a preferred embodiment of the video processing method of the present invention.

[0012] Figure 2 is the original image of a certain video frame of the video to be processed.

[0013] Figure 3 is Figure 2 the image processed by the video processing method in the embodiment of the present invention.

[0014] Figure 4 is the module schematic diagram of the video processing device in the second embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0015] In order to make the objectives, technical solutions and beneficial effects of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.

[0016] In order to illustrate the technical solutions described in the present invention, the following will be described through specific embodiments.

[0017] Example 1

[0018] As Figure 1 shown, it is a preferred embodiment of the video processing method of the present invention. The video processing method in this embodiment is executed by a computer device with a processor. The computer device can be a terminal device, such as a camera or a mobile phone with a photographing function. The computer device can also be a server. The method includes the following steps.

[0019] S1: Obtain a video frame of the video to be processed.

[0020] Specifically, the computer device receives a video to be processed from an image sensor. For example, the duration of the video to be processed is 2 seconds and it contains 25 video frames per second, so there are a total of 50 video frames. The acquired video frames can be all 50 video frames, or multiple video frames can be randomly selected from these 50 video frames, or one video frame can be extracted every certain period of time (such as 200 milliseconds), or one video frame can be extracted after every certain number of frames (such as every 20 frames, that is, the 1st frame, the 21st frame, and the 41st frame are extracted).

[0021] S2: Identify the target boxes of the same region of interest in each video frame.

[0022] The initial acquisition method of the target box of the target of interest in the first video frame (such as Figure 2 shown by the red box) can be obtained by a preset method of the target of interest. For example, a target box of the current video frame is selected manually on the display interface, or automatically obtained by a target recognition algorithm, or multiple target boxes of interest are obtained by a target recognition algorithm and displayed on the display interface, and then selected manually; among them, the target of interest can be an object such as a person, an animal, a building, etc., the video frame can be a panoramic plane video frame or an ordinary video frame, and the target box of the target of interest can be a rectangular box, a circular box, etc.; the target boxes of the subsequent video frames are obtained by tracking the input video frames using a target tracking algorithm, that is, the target boxes of the subsequent video frames are obtained by tracking the previous input video frames using a target tracking algorithm. The target tracking algorithm can use classic algorithms such as Online objecttracking, or the target tracking algorithm in the applicant's patent application No. 202010023864.0, patent name "Target tracking method, readable storage medium and computer device for panoramic video".

[0023] S3: Adjust the FOV value of the target box in the video frame to adjust the proportion of the target box in the video frame to an ideal ratio.

[0024] In this example, the FOV value Fov of the target box of the video frame i can be obtained through the calculation formula (1):

[0025] Fov i = arctan(tan(F i )*(R i / R g )) (1)

[0026] Among them, Fi is the FOV value of the i-th video frame of the video to be processed currently, and the R i is the proportion of the target box in the current video frame in the video frame, R gis the ideal ratio of the target box in the video frame. The ideal ratio can be determined according to the aspect ratio of the selected target box and the aspect ratio of the input image.

[0027] Since directly calculating the FOV value of the target box using the results of target tracking may result in unstable situations, making it inaccurate in reflecting the changing trend of FOV, vulnerable to noise interference, and affecting the user experience. In this embodiment, the FOV value Fov of the target box in the video frame i is subjected to a smoothing filtering process. Specifically, the Kalman filtering algorithm can be used. First, an adaptive weight is established for the second FOV value of the previous i - 1 frame, and the FOV value Fov of the target box in the i-th video frame i is filtered to obtain the second FOV value Fov i ’. It should be noted that when i equals 1, that is, when processing the first video frame, since there is no previous video frame, the FOV value Fov1 of the target box in the first frame is equal to its second FOV value Fov1’.

[0028] S4: Render the adjusted video frame to generate a sliding zoom image.

[0029] In this embodiment, according to the second FOV value Fov of each video frame i ’ (in the case where the FOV value Fov of the target box i is not smoothed filtered, then according to the FOV value Fov of the target box in each video frame i ), a new planar image is rendered on the current video frame as a sliding zoom image (as Figure 3 is the sliding zoom image generated after several frames of processing), and the sliding zoom video is further generated by continuing to perform sliding zoom processing on the sliding zoom image using target tracking.

[0030] It can be understood that by rendering multiple adjusted video frames and generating multiple video frame images, and then outputting these video frames in chronological order, a sliding zoom video can be obtained.

[0031] Such as Figure 2 、 Figure 3 shows the comparison of the images of the video frame in this embodiment before and after being processed by the video processing method in this embodiment. It can be seen from the comparison of the images of the two video frames that this processing method significantly improves the image quality of the video frame.

[0032] Example 2:

[0033] Please refer to Figure 4 , the module schematic diagram of the video processing device of the present invention. The video processing device in this embodiment includes:

[0034] An acquisition module, configured to acquire video frames of a video to be processed;

[0035] An identification module, configured to identify target boxes of regions of interest in the video frames;

[0036] An adjustment module, configured to adjust the FOV value of the target boxes in the video frames so as to adjust the proportion of the target boxes of the regions of interest in the video frames to a predetermined value;

[0037] A rendering module, configured to render the adjusted video frames to generate an image with smooth zoom.

[0038] Example 3:

[0039] Embodiment 3 of the present invention provides an electronic device, including a processor, a memory, and a computer program stored on the memory and executable on the processor, where the processor executes the computer program to implement the steps of the video processing method in Embodiment 1. A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the photometric measurement method of the panoramic camera provided in Embodiment 1 of the present invention are implemented.

[0040] Example 4:

[0041] This embodiment provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the video processing method provided in Embodiment 1 of the present invention are implemented.

[0042] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program, and this program can be stored in a computer-readable storage medium. The storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disk, etc.

[0043] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A video processing method, characterized in that, S1: Obtain video frames of the video to be processed, and the video frames of the video to be processed are panoramic planar video frames; S2: Identify the target boxes of the same region of interest in each video frame; S3: Adjust the FOV value of the target box in the video frame to adjust the proportion of the target box in the video frame to the ideal ratio, where the FOV value Fov of the target box in the video frame in step S3 i has the following calculation formula (1): Fov i = arctan(tan(F i )*(R i / R g )) (1) In formula (1), Fi is the FOV value of the i-th video frame of the video to be processed, and R i is the proportion of the target box in the current video frame in the video frame, and R g is the ideal ratio of the target box in the video frame. The step S3 further includes smoothing the FOV value Fov i of the target box in the i-th video frame to obtain a second FOV value Fov i '; S4: Render the adjusted video frames to generate an image with a sliding zoom effect, output the adjusted video frames in chronological order, and obtain a sliding zoom video.

2. The video processing method according to claim 1, characterized in that, The target box of the region of interest determined for the first time in step S2 is a target box manually selected from the video frame or a target box obtained by using a target recognition algorithm.

3. The video processing method according to claim 2, characterized in that, The target boxes of the regions of interest in subsequent video frames are obtained by tracking the target box of the previous frame through a target tracking algorithm.

4. The video processing method according to claim 1, characterized in that, The smoothing filter uses the Kalman filter algorithm. First, an adaptive weight is established for the second FOV value of the (i - 1)-th frame, and then the FOV value Fov of the target box in the i-th video frame is filtered to obtain the second FOV value Fov i '. i ’。 5. A video processing device, characterized in that, It includes: An acquisition module, configured to acquire video frames of the video to be processed, and the video frames of the video to be processed are panoramic planar video frames; An identification module, configured to identify the target boxes of the regions of interest in the video frames; Adjustment module, which is used to adjust the FOV value of the target box in the video frame so as to adjust the proportion of the target box of the region of interest in the video frame to a predetermined value, where the FOV value Fov of the target box in the video frame i has the following calculation formula (1): Fov i = arctan(tan(F i ) * (R i / R g )) (1). In formula (1), Fi is the FOV value of the i-th video frame of the video to be processed currently, and R i is the proportion of the target box in the current video frame in the video frame, and R g is the ideal ratio of the target box in the video frame. The adjustment module is further used to perform smoothing filtering on the FOV value Fov i of the target box in the i-th video frame to obtain a second FOV value Fov i '; A rendering module, configured to render the adjusted video frames to generate an image with a sliding zoom effect, output the adjusted video frames in chronological order, and obtain a sliding zoom video.

6. An electronic device, comprising a processor, a memory, and a computer program stored on the memory and executable on the processor, characterized in that, The processor executes the computer program to implement the video processing method according to any one of claims 1 to 4.

7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, the video processing method according to any one of claims 1 to 4 is implemented.

Citation Information

Patent Citations

  • Target tracking methods, readable storage media, and computer devices for panoramic video

    CN111242977B

  • Slide zooming effect realization method and device, electronic device and computer readable storage medium

    CN109379537A