Video shooting method and system and electronic equipment
By acquiring and cropping image data on a miniature video capture terminal and using information processing equipment for image stabilization, the problem of image shakiness on miniature terminals is solved, achieving efficient, lightweight video stability and real-time performance.
Patent Information
- Application Number
- CN202511758542.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-26
- Publication Date
- 2026-02-10
AI Technical Summary
Existing technologies struggle to achieve efficient and lightweight electronic image stabilization on miniature mobile video capture terminals, resulting in severe image shakiness. Furthermore, existing solutions suffer from high latency, high energy consumption, and limited computing resources.
The initial image and pose data are acquired by the shooting device, cropped to generate an intermediate image sequence, and sent to the information processing device for image stabilization. The image is then stabilized by combining optical flow information and transmission transformation matrix.
While balancing power consumption and data transmission volume, it improves the image stabilization effect and real-time performance of video shooting, thereby enhancing video stability and user experience.
Smart Images

Figure CN121509815A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, and more particularly to a video shooting method and apparatus. Background Technology
[0002] With the development of edge intelligence, many miniature mobile video capture terminals (such as smart glasses, watches, wristbands, vehicle dashcams, and AR / VR headsets) are widely used in first-person perspective recording, remote collaboration, and environmental awareness scenarios. These devices are generally compact in structure, power-sensitive, and prone to image shaking due to their high degree of freedom of movement, thus creating a rigid demand for lightweight and efficient electronic image stabilization technology.
[0003] Existing solutions mainly fall into two categories: one is to transmit the original high-resolution video stream to an external high-performance terminal for image stabilization, resulting in significant latency and bandwidth pressure; the other is to execute the image stabilization algorithm entirely locally, which is limited by computing power and energy efficiency, often sacrificing processing quality or significantly shortening battery life, making it difficult to meet the requirements of both real-time performance and stability. Summary of the Invention
[0004] In view of this, the purpose of this invention is to provide a video shooting method and apparatus that can improve the overall image stabilization effect, real-time performance, and stability of video shooting while taking into account the power consumption and data transmission volume of the shooting terminal.
[0005] In a first aspect, embodiments of the present invention provide a video shooting system, the system comprising a shooting device and an information processing device, wherein: The imaging device is used to acquire initial data, which includes multiple initial images and corresponding pose data. The initial images are cropped according to the pose data to obtain a first intermediate image sequence, which includes multiple first intermediate images. The first intermediate image sequence is sent to the information processing device. The information processing device is used to receive a first intermediate image sequence and perform image stabilization processing on the first intermediate image sequence to generate a target video.
[0006] In some embodiments, the imaging device is used for: A rotation angle sequence is determined based on the initial image and pose data, the rotation angle sequence including the relative rotation angle of each two adjacent initial images in the target direction; The displacement vector sequence is determined based on the rotation angle sequence, and the displacement vector sequence includes the displacement vectors of each two adjacent initial images in the target direction; The desired displacement vector for each initial image is determined based on the displacement vector sequence. The initial images are cropped according to the desired displacement vector to obtain a first intermediate image sequence, which includes a plurality of first intermediate images.
[0007] In some embodiments, the information processing device is used for: Receive a first intermediate image sequence, the first intermediate image sequence including a plurality of first intermediate images, the first intermediate images being cropped images; A subsequence is obtained through a pre-set sliding window, the subsequence comprising a predetermined number of first intermediate images; A matrix sequence is obtained based on each first intermediate image in the sub-sequence, and the matrix sequence includes the transmission transformation matrix of each pair of adjacent first intermediate images in the sub-sequence. The first intermediate image is transformed according to the matrix sequence to obtain the second intermediate image; Generate the target video based on the second intermediate image.
[0008] Secondly, embodiments of the present invention provide a video shooting method, applicable to shooting devices, the method comprising: Acquire initial data, which includes multiple initial images and corresponding pose data; A rotation angle sequence is determined based on the initial image and pose data, the rotation angle sequence including the relative rotation angle of each two adjacent initial images in the target direction; The displacement vector sequence is determined based on the rotation angle sequence, and the displacement vector sequence includes the displacement vectors of each two adjacent initial images in the target direction; The desired displacement vector for each initial image is determined based on the displacement vector sequence. The initial images are cropped according to the desired displacement vector to obtain a first intermediate image sequence, the first intermediate image sequence including a plurality of first intermediate images; The first intermediate image sequence is sent to an information processing device to generate a target video based on the first intermediate image sequence.
[0009] In some embodiments, obtaining the initial data includes: Acquire sensor data, which includes multiple initial images, timestamps corresponding to each initial image, pose data, and timestamps corresponding to each pose data. The initial image and pose data are aligned according to the timestamp corresponding to the initial image and the timestamp corresponding to the pose data to obtain the initial data.
[0010] In some embodiments, the target direction includes the vertical direction and / or the left-right direction.
[0011] In some embodiments, the pose data includes the angular velocity of the capturing device in the target direction at various times; The step of determining the rotation angle sequence based on the initial image and pose data includes: Determine the initial images of the two adjacent frames that need to be processed; The time interval between two adjacent initial images is determined based on the timestamps of the two adjacent initial images; The average angular velocity of two adjacent initial frames in the target direction is determined based on the pose data. The relative rotation angle between the two adjacent initial frames is determined based on the average angular velocity and the time interval.
[0012] In some embodiments, determining the displacement vector sequence based on the rotation angle sequence includes: Acquire intrinsic parameter data from the image sensor; Determine the relative rotation angles to be processed from the rotation angle sequence; The displacement vectors of two adjacent initial images in the target direction are determined based on the intrinsic parameter data and the relative rotation angle.
[0013] In some embodiments, determining the desired displacement vector of each initial image based on the displacement vector sequence includes: The displacement vector sequence is subjected to temporal filtering to obtain the desired displacement vector for each initial image.
[0014] Thirdly, embodiments of the present invention provide a video shooting method, applicable to information processing devices, the method comprising: Receive a first intermediate image sequence, the first intermediate image sequence including a plurality of first intermediate images, the first intermediate images being cropped images; A subsequence is obtained through a pre-set sliding window, the subsequence comprising a predetermined number of first intermediate images; A matrix sequence is obtained based on each first intermediate image in the sub-sequence, and the matrix sequence includes the transmission transformation matrix of each pair of adjacent first intermediate images in the sub-sequence. The first intermediate image is transformed according to the matrix sequence to obtain the second intermediate image; Generate the target video based on the second intermediate image.
[0015] In some embodiments, obtaining the matrix sequence based on each of the first intermediate images in the sub-sequence includes: An optical flow information sequence is obtained based on each first intermediate image in the sub-sequence, and the optical flow information sequence includes the optical flow information of each two adjacent first intermediate images in the sub-sequence. The matrix sequence is determined based on the optical flow information sequence, and the matrix sequence includes the affine transformation matrix of each two adjacent first intermediate images in the subsequence.
[0016] In some embodiments, transforming the first intermediate image according to the matrix sequence to obtain the second intermediate image includes: The desired transformation matrix is obtained based on the matrix sequence corresponding to the sliding window. The first intermediate image is transformed according to the desired transformation matrix to obtain the second intermediate image.
[0017] In some embodiments, generating the target video from the second intermediate image includes: The invalid regions in the second intermediate image are cropped to obtain the third intermediate image; The target video is generated based on the third intermediate image.
[0018] Fourthly, embodiments of the present invention provide an electronic device, including a memory and a processor, wherein the memory is used to store one or more computer program instructions, wherein the one or more computer program instructions are executed by the processor to implement the methods as described in the first and second aspects.
[0019] The technical solution of this invention acquires initial data through a shooting device. The initial data includes multiple initial images and corresponding pose data. The initial images are cropped based on the pose data, and the resulting first intermediate image sequence is sent to an information processing device. The first intermediate image sequence includes multiple first intermediate images. The information processing device performs image stabilization processing on the first intermediate image sequence to generate a target video. Therefore, while considering the power consumption and data transmission volume of the shooting terminal, the overall image stabilization effect, real-time performance, and stability of video shooting can be improved. Attached Figure Description
[0020] The above and other objects, features and advantages of the present invention will become clearer from the following description of embodiments of the invention with reference to the accompanying drawings, in which: Figure 1 This is a schematic diagram of the video shooting system according to an embodiment of the present invention; Figure 2 This is a flowchart of a video shooting method using a shooting device according to an embodiment of the present invention; Figure 3 This is a flowchart of obtaining initial data according to an embodiment of the present invention; Figure 4This is a schematic diagram of the coordinate axes of the pose sensor of the shooting device according to an embodiment of the present invention; Figure 5 This is a flowchart illustrating the process of obtaining a rotation angle sequence according to an embodiment of the present invention; Figure 6 This is a flowchart illustrating the determination of the displacement vector sequence according to an embodiment of the present invention; Figure 7 This is a schematic diagram of image cropping according to an embodiment of the present invention; Figure 8 This is a flowchart of a video shooting method using an information processing device according to an embodiment of the present invention; Figure 9 This is a flowchart illustrating the process of obtaining a matrix sequence according to an embodiment of the present invention; Figure 10 This is a flowchart of obtaining the second intermediate image according to an embodiment of the present invention; Figure 11 This is a schematic diagram of the video shooting device of the shooting equipment according to an embodiment of the present invention; Figure 12 This is a schematic diagram of the video shooting device of the information processing equipment according to an embodiment of the present invention; Figure 13 This is a schematic diagram of an electronic device according to an embodiment of the present invention. Detailed Implementation
[0021] The present application is described below based on embodiments, but it is not limited to these embodiments. In the detailed description of the present application below, certain specific details are described in detail. Those skilled in the art can fully understand the present application without these details. To avoid obscuring the substance of the present application, well-known methods, processes, flows, elements, and circuits are not described in detail.
[0022] Furthermore, those skilled in the art should understand that the accompanying drawings provided herein are for illustrative purposes only and are not necessarily drawn to scale.
[0023] Unless the context explicitly requires it, words such as "including" or "contains" throughout the application should be interpreted as including rather than exclusive or exhaustive; that is, meaning "including but not limited to".
[0024] In the description of this application, it should be understood that the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance. Furthermore, in the description of this application, unless otherwise stated, "a plurality of" means two or more.
[0025] The solutions described in this specification and embodiments, if involving the processing of personal information, will be processed only on the premise of having a legal basis (such as obtaining the consent of the personal information subject, or being necessary for the performance of a contract), and will only be processed within the scope stipulated or agreed upon. A user's refusal to process personal information beyond what is necessary for basic functions will not affect the user's use of basic functions.
[0026] With the deep integration of artificial intelligence and edge computing technologies, edge intelligence is accelerating the empowerment of various miniature mobile video capture terminals, including smart glasses, smartwatches, wristbands, vehicle dashcams, and AR / VR headsets. These terminals, leveraging their first-person view (FPV) capture capabilities, demonstrate high application value in scenarios such as remote collaboration, industrial inspection, emergency response, motion recording, immersive interaction, and environmental perception. However, limited by wearing methods, usage habits, and physical structures, these devices are prone to significant image jitter during actual operation due to high-frequency micro-motions, sudden displacements, and even complex motion trajectories, severely impacting the usability of video content and user experience. Therefore, efficient, lightweight, and low-latency electronic image stabilization technology has become a key technological bottleneck ensuring the practicality of the video functions of these terminals.
[0027] Current mainstream electronic image stabilization solutions mainly fall into two categories: one relies on transmitting the original high-resolution video stream in real time via a wireless link to an external high-performance terminal (such as a smartphone, edge server, or cloud) for post-processing. While this method can leverage powerful external computing power to achieve high-quality stabilization, it inevitably introduces significant end-to-end latency, severely limiting applications with extremely high real-time requirements. Furthermore, the continuous transmission of high-bandwidth video data not only exacerbates the load on the communication link but also significantly increases the overall system's energy consumption, contradicting the design principles of low power consumption and long battery life for edge devices.
[0028] Another approach attempts to deploy all image stabilization algorithms locally on the shooting device to avoid transmission latency issues. However, limited by the extremely limited computing resources, memory capacity, and energy efficiency budget of micro-devices, existing local algorithms often have to compromise between stability, image fidelity, and processing speed. This may involve significantly reducing the input frame resolution and processing frame rate, or simplifying motion estimation and compensation models, resulting in poor image stabilization, excessive image cropping, or even motion blur and artifacts, making it difficult to meet the basic video quality requirements of consumer products.
[0029] Taking smart glasses as an example, their close fit to the body and free head movement result in significantly higher-frequency camera shake during shooting compared to traditional handheld devices. Furthermore, their highly compact design prevents the integration of mechanical structures required for optical image stabilization, forcing the stabilization task to rely entirely on software algorithms. At the same time, smart glasses are extremely sensitive to power consumption; high-load local computing can quickly deplete the battery, severely impacting the daily user experience.
[0030] Therefore, the present invention proposes a video shooting method, system, and electronic device to solve the above-mentioned problems.
[0031] Figure 1 This is a schematic diagram of a video shooting system according to an embodiment of the present invention. Figure 1 As shown, the video shooting system of this embodiment includes a shooting device 1, an information processing device 2, and a network 3. The shooting device 1 and the information processing device 2 transmit data via the network 3.
[0032] Among them, the shooting device 1 is a device with shooting and data processing capabilities. For example, the shooting device can be smart glasses, smartwatches, wristbands, vehicle dashcams, and AR / VR headsets.
[0033] Furthermore, the imaging device 1 includes an image sensor and a pose sensor.
[0034] The image sensor is used to capture a series of video frames at a predetermined frame rate after receiving a video capture command. In this embodiment of the invention, the video frame is referred to as the initial image.
[0035] The pose sensor is used to collect pose data. This pose data includes the angular velocity of the imaging device in the target direction at various times.
[0036] The imaging device 1 is used to acquire initial data, which includes multiple initial images and corresponding pose data. The initial images are cropped according to the pose data to obtain a first intermediate image sequence, which includes multiple first intermediate images. The first intermediate image sequence is then sent to the information processing device.
[0037] Information processing device 2 is a device with data processing capabilities. For example, the information processing device can be a terminal device (such as a mobile phone, laptop, tablet computer, or desktop computer), a server (local server, cloud server), or other dedicated data processing device.
[0038] The information processing device 2 is used to receive a first intermediate image sequence and perform image stabilization processing on the first intermediate image sequence to generate a target video.
[0039] In this embodiment, network 3 is used for the exchange of information and / or data between the imaging device 1 and the information processing device 2. Network 3 can be any type of wired or wireless network. In some embodiments, network 3 may include a wired network, wireless network, fiber optic network, telecommunications network, intranet, Internet, local area network (LAN), wide area network (WAN), wireless local area network (WLAN), metropolitan area network (MAN), public switched telephone network (PSTN), Bluetooth network, ZigBee network, or near field communication (NFC) network, or any combination thereof.
[0040] This invention employs a shooting device to acquire initial data, including multiple initial images and corresponding pose data. The initial images are cropped based on the pose data, and the resulting first intermediate image sequence is sent to an information processing device. This first intermediate image sequence comprises multiple first intermediate images. The information processing device performs image stabilization processing on the first intermediate image sequence to generate the target video. Therefore, while considering the power consumption and data transmission volume of the shooting terminal, the overall image stabilization effect, real-time performance, and stability of video shooting can be improved.
[0041] Figure 2 This is a flowchart of a video shooting method using a shooting device according to an embodiment of the present invention. Figure 2 As shown, the video shooting method of the shooting device in this embodiment of the invention includes the following steps: Step S110: Obtain initial data.
[0042] In this embodiment, the initial data includes multiple initial images and corresponding pose data.
[0043] Figure 3 This is a flowchart illustrating the acquisition of initial data according to an embodiment of the present invention. For example... Figure 3 As shown, obtaining the initial data includes the following steps: Step S111: Acquire sensor data.
[0044] In this embodiment, the sensor data includes multiple initial images, timestamps corresponding to each initial image, pose data, and timestamps corresponding to each pose data.
[0045] Specifically, multiple consecutive initial images are acquired using an image sensor, and the timestamp corresponding to each initial image is recorded. Pose data is acquired using a pose sensor, and the timestamp corresponding to each pose data point is recorded.
[0046] The pose sensor can be implemented using a gyroscope. A gyroscope is a device used to measure angular velocity (i.e., the speed and direction of an object's rotation). A gyroscope can measure the angular velocity of an object about three mutually perpendicular axes (X, Y, and Z axes), with units of radians per second (rad / s) or degrees per second (° / s).
[0047] Figure 4 This is a schematic diagram of the coordinate axes of the pose sensor of the imaging device according to an embodiment of the present invention. Figure 4 As shown, the X-axis points to the right of the imaging device, the Y-axis points above the imaging device, and the Z-axis points in front of the imaging device. The pose data acquired at time t can be denoted as (x... t y t , z t ).
[0048] In some embodiments, the frequency at which the pose sensor acquires pose data is greater than or equal to the frequency at which the image sensor acquires the initial image. Further, when the frequency at which the pose sensor acquires pose data is greater than the frequency at which the image sensor acquires the initial image, the frequency at which the pose sensor acquires pose data is an integer multiple of the frequency at which the image sensor acquires the initial image.
[0049] Step S112: Align the initial image and pose data according to the timestamp corresponding to the initial image and the timestamp corresponding to the pose data to obtain the initial data.
[0050] In this embodiment, the initial image and pose data are aligned according to the timestamp corresponding to the initial image and the timestamp corresponding to the pose data to obtain the initial data. That is, for each initial image, pose data with the same timestamp as the initial image is selected as the pose data corresponding to that initial image.
[0051] It should be noted that if the timestamp corresponding to the initial image and the timestamp corresponding to the pose data cannot be perfectly aligned, the pose data with the timestamp closest to the initial image can be selected as the pose data corresponding to the initial image.
[0052] Step S120: Determine the rotation angle sequence based on the initial image and pose data.
[0053] In this embodiment, a rotation angle sequence is determined based on the initial image and pose data. The rotation angle sequence includes the relative rotation angles of two adjacent initial frames in the target direction. The pose data includes the angular velocity of the capturing device in the target direction at each moment. The target direction includes the vertical direction and / or the left-right direction. For the vertical direction, the corresponding pose data is yt; for the left-right direction, the corresponding pose data is xt.
[0054] Figure 5 This is a flowchart illustrating the acquisition of a rotation angle sequence according to an embodiment of the present invention. For example... Figure 5 As shown, determining the rotation angle sequence based on the initial image and pose data includes the following steps: Step S121: Determine the initial images of the two adjacent frames that need to be processed.
[0055] In this embodiment, it is necessary to process each pair of adjacent initial images to obtain the relative rotation angles corresponding to the two adjacent initial images. Therefore, the two adjacent initial images to be processed are selected.
[0056] Step S122: Determine the time interval between two adjacent initial images based on the timestamps of the two adjacent initial images.
[0057] In this embodiment, it is assumed that the number of initial images acquired is N, and the i-th initial image is denoted as F. i Its corresponding timestamp is t i , i = 1, 2, 3, ..., N. Where N is a positive integer greater than 1.
[0058] For the two adjacent initial images F selected in this processing step i and F i+1 Their corresponding timestamps are t i and t i+1 , where t i Let t be the timestamp of the i-th initial image. i+1 This is the timestamp of the (i+1)th initial image. The time interval between two adjacent initial images is calculated using the following formula: in, Let be the time interval between the i-th initial image and the (i+1)-th initial image. This is the timestamp corresponding to the (i+1)th initial image. This is the timestamp corresponding to the i-th initial image.
[0059] Step S123: Determine the average angular velocity of the two adjacent initial images in the target direction based on the pose data.
[0060] In this embodiment, the pose data corresponding to the i-th initial image is denoted as (x i y i , z i This embodiment of the invention uses the target direction, which includes both vertical and horizontal directions, as an example for illustration. For the two adjacent initial images F selected in this processing step... i and F i+1 In the vertical direction, its average angular velocity can be y i With y i+1 The average value; in the left-right direction, its average angular velocity can be x. i With x i+1 The average value. That is, determining the average angular velocity in the target direction of two adjacent initial frames includes: in, Let be the average angular velocity in the left-right direction of the i-th initial image and the (i+1)-th initial image. Let be the average angular velocity in the vertical direction of the i-th initial image and the (i+1)-th initial image.
[0061] Step S124: Determine the relative rotation angle of the two adjacent initial images based on the average angular velocity and the time interval.
[0062] In this embodiment, the product of the average angular velocity in the target direction and the time interval is used as the relative rotation angle between two adjacent initial frames in the target direction. The specific calculation formula is as follows: in, Let be the relative rotation angle between the i-th initial image and the (i+1)-th initial image in the left-right direction. Let be the relative rotation angle between the i-th initial image and the (i+1)-th initial image in the vertical direction.
[0063] Repeating the above steps yields a rotation angle sequence, which includes ( ), ( ), ..., ( , ), ( , ).
[0064] The above method for calculating the relative rotation angle is only an example provided by the embodiments of the present invention. The embodiments of the present invention do not limit the specific calculation method. For example, the relative rotation angle can also be obtained by integrating all angular velocities between the i-th initial image and the (i+1)-th initial image.
[0065] Step S130: Determine the displacement vector sequence based on the rotation angle sequence.
[0066] In this embodiment, a displacement vector sequence is determined based on the rotation angle sequence, and the displacement vector sequence includes the displacement vectors of each two adjacent initial frames in the target direction.
[0067] Figure 6 This is a flowchart illustrating the determination of a displacement vector sequence according to an embodiment of the present invention. For example... Figure 6 As shown, determining the displacement vector sequence based on the rotation angle sequence includes the following steps: Step S131: Obtain the intrinsic parameter data of the image sensor.
[0068] In this embodiment, the intrinsic parameter data is used to describe how the image sensor projects 3D world points onto a 2D image. This includes information such as focal length and principal point; the focal length is the pixel representation of the camera lens focal length, and the principal point is the pixel coordinates of the image center (optical center).
[0069] Step S132: Determine the relative rotation angle to be processed in the rotation angle sequence.
[0070] In this embodiment, each relative rotation angle in the rotation angle sequence needs to be processed one by one, and a relative rotation angle to be processed is selected.
[0071] Step S133: Determine the displacement vector of the two adjacent initial images in the target direction based on the intrinsic parameter data and the relative rotation angle.
[0072] In this embodiment, camera rotation (including X-axis rotation and Y-axis rotation) causes image point movement. The projection change of a 3D point in the camera coordinate system can be expressed by the following formula: in, Let be the second coordinate of the pixel in the i-th initial image. Let be the second coordinate of the pixel in the (i+1)th initial image. A matrix representation of intrinsic parameter data. Let K be the inverse matrix, and R be the matrix representation of the relative rotation angle between the i-th initial image and the (i+1)-th initial image.
[0073] R includes the relative rotation angles of two adjacent initial images in each target direction.
[0074] Assuming the selected relative rotation angle to be processed is the relative rotation angle between the i-th initial image and the (i+1)-th initial image, a reference point is randomly selected in the i-th initial image (e.g., the center point can be selected), and the theoretical position of the point in the (i+1)-th initial image is calculated using the above formula.
[0075] Suppose the coordinates of the reference point randomly selected in the i-th initial image are (C Xi C Yi ), that is The above formula is used to calculate the result. ,but: in, Let be the displacement vectors of the i-th initial image and the (i+1)-th initial image in the left-right direction. Let be the vertical displacement vector of the i-th initial image and the (i+1)-th initial image.
[0076] By repeating the above steps, a displacement vector sequence can be obtained.
[0077] Step S140: Determine the expected displacement vector of each initial image based on the displacement vector sequence.
[0078] In this embodiment, the displacement vector sequence is subjected to temporal filtering to obtain the desired displacement vector for each initial image. Temporal filtering can be performed using inertial filtering, Kalman filtering, Wiener filtering, or similar methods.
[0079] Taking Kalman filtering as an example, after obtaining the displacement vector sequence, a Kalman filter is created to observe the changes in the displacement. The Kalman filter is updated frame by frame, and the following operations are repeated for each frame: based on the state and motion trend of the previous frame, the predicted value of the current frame is predicted; the actual value is extracted from the affine transformation of the current frame; the Kalman filter automatically weighs the predicted value and the actual value, and outputs a more reliable and smoother estimate to obtain the desired displacement vector.
[0080] Step S150: Crop each of the initial images according to the desired displacement vector to obtain a first intermediate image sequence.
[0081] In this embodiment, each of the initial images is cropped according to the desired displacement vector to obtain a corresponding first intermediate image, and a first intermediate image sequence is generated, the first intermediate image sequence including multiple first intermediate images.
[0082] Specifically, the expected displacement vectors of the i-th initial image and the (i+1)-th initial image in the target direction are denoted as ( ), then according to ( Crop the (i+1)th initial image.
[0083] Figure 7 This is a schematic diagram of image cropping according to an embodiment of the present invention. Figure 7 As shown, for the (i+1)th initial image F i+1 Obtain its center point CF i+1 coordinates (X) F,i+1 ,Y F,i+1 Then, based on the (i+1)th initial image F i+1 center point CF i+1 The coordinates of the i-th initial image and the expected displacement vectors of the i+1-th initial image in the target direction. Calculate a new center point, which becomes the center point of the corresponding first intermediate image. The new center point is CG. i+1 coordinates (X) G,i+1 ,Y G,i+1 )for: Next, the new center point CG i+1 coordinates (X) G,i+1 ,Y G,i+1 Using the (i+1)th initial image as the center point, the first intermediate image G is obtained by cropping the initial image. i+1 .
[0084] Specifically, the image can be cropped according to a predetermined cropping ratio, resulting in a first intermediate image whose length / width is proportional to the length / width of the initial image. The length and width ratios can be the same or different. For example, both the length and width ratios can be 90%.
[0085] It should be noted that for the first frame of the initial image, the center point position can be kept unchanged, and then cropped according to the predetermined cropping ratio.
[0086] Step S160: Send the first intermediate image sequence to the information processing device.
[0087] In this embodiment, the obtained first intermediate image sequence is processed by an ISP (Image Signal Processor), and the ISP-processed first intermediate image sequence is sent to an information processing device so that the information processing device can generate a target video based on the first intermediate image sequence. The ISP is a dedicated hardware module, typically integrated into a SoC (System-on-a-Chip), which mainly performs de-mosaicing, white balance, color correction, noise reduction, automatic exposure / focus control, and image enhancement on the first intermediate image sequence.
[0088] Therefore, the shooting device first uses pose data to perform simple cropping to remove the large displacement caused by the rotation of the x and y axes, and only crops the image without performing other deformation compensation to improve efficiency. It is equivalent to only having a memory copy process, and the cropping can be performed on the RAW image before ISP to reduce the load on the entire ISP. It can effectively remove the jitter caused by the rotation of the glasses along the x and y axes, and the resolution of the cropped image is reduced, resulting in a reduction in the amount of data transmitted.
[0089] Figure 8 This is a flowchart of a video capture method using an information processing device according to an embodiment of the present invention. Figure 8 As shown, the video shooting method of the information processing device in this embodiment of the invention includes the following steps: Step S210: Receive the first intermediate image sequence.
[0090] In this embodiment, the information processing device receives a first intermediate image sequence sent by the shooting device. The first intermediate image sequence includes a plurality of first intermediate images, and the first intermediate images are cropped images.
[0091] Step S220: Obtain the subsequence through a pre-set sliding window.
[0092] In this embodiment, the subsequence includes a predetermined number of first intermediate images. That is, a sliding window is pre-set, and a predetermined number of first intermediate images are obtained as a subsequence through the sliding window.
[0093] The number of items can be set according to the actual scenario, such as 10, 15, 20, etc.
[0094] Step S230: Obtain a matrix sequence based on each of the first intermediate images in the sub-sequence.
[0095] In this embodiment, the subsequence within the sliding window is processed pairwise between consecutive frames to obtain a matrix sequence. The matrix sequence includes the transmission transformation matrix of the first intermediate image of each two adjacent frames in the subsequence.
[0096] Specifically, Figure 9 This is a flowchart illustrating the process of obtaining a matrix sequence according to an embodiment of the present invention. For example... Figure 9 As shown, obtaining the matrix sequence based on each first intermediate image in the sub-sequence includes the following steps: Step S231: Obtain optical flow information sequence based on each first intermediate image in the sub-sequence.
[0097] In this embodiment, an optical flow information sequence is obtained based on each first intermediate image in the sub-sequence. The optical flow information sequence includes the optical flow information of each pair of adjacent first intermediate images in the sub-sequence.
[0098] Specifically, the two adjacent first intermediate images that need to be processed in the subsequence are determined and denoted as G. i and G i+1 The optical flow information of the first intermediate image between two adjacent frames is calculated using a predetermined optical flow information calculation method. The optical flow information is the pixel motion vector field estimated from the two consecutive images.
[0099] The predetermined optical flow information calculation method can be any existing method, such as the Lucas-Kanade method (sparse optical flow), the Horn-Schunck method (dense optical flow), and deep learning-based methods.
[0100] The Lucas-Kanade method calculates optical flow information by analyzing feature points (such as corner points) in an image. This method assumes that all pixels share the same motion vector within a small image window. Based on this, it establishes a system of equations for all pixels within the window using the assumption of constant brightness, and solves these equations using the least squares method to obtain the motion vector corresponding to each feature point, i.e., the optical flow information of that point. To handle large displacement motions, it can be combined with an image pyramid, calculating and refining layer by layer from low resolution to high resolution, ultimately outputting a sparse optical flow vector field representing the motion of key points.
[0101] The Horn-Schunck method obtains optical flow information by solving a global energy function. This energy function contains two core constraints: a data term ensures that the obtained optical flow conforms to the constant brightness assumption, and a smoothing term penalizes drastic changes in optical flow vectors between adjacent pixels, thereby forcing the entire optical flow field to change smoothly. By minimizing this energy function through an iterative optimization algorithm, this method ultimately assigns a motion vector to each pixel in the image, resulting in a globally smooth and complete dense optical flow field as the final optical flow information.
[0102] The deep learning-based approach utilizes a pre-trained end-to-end neural network to directly infer optical flow information. This network takes two consecutive image frames as input and automatically learns the complex motion features and matching relationships between the images through internal convolutional and pooling layers. Instead of relying on manually set assumptions, the network directly learns how to associate the two images from the data and generates a two-dimensional vector field at the network's output layer. This vector field represents the calculated optical flow information, providing a precise motion vector for each pixel in the input image pair, thus achieving high-precision dense optical flow prediction.
[0103] Optical flow information is a two-dimensional vector field, which can be represented as a two-dimensional motion vector (u, v). u represents the pixel's displacement in the horizontal direction (x-axis), with positive values representing rightward movement and negative values representing leftward movement. v represents the pixel's displacement in the vertical direction (y-axis), with positive values representing downward movement and negative values representing upward movement. For example, if in two adjacent frames, a pixel in the first frame moves to a position 5 pixels to its right and 3 pixels below it in the second frame, then the optical flow vector for that pixel is (u=5, v=3).
[0104] By repeating the above steps, the optical flow information sequence can be obtained.
[0105] Step S231: Determine the matrix sequence based on the optical flow information sequence.
[0106] In this embodiment, the matrix sequence includes the affine transformation matrix of each two adjacent first intermediate images in the subsequence.
[0107] The formula for the two-dimensional affine transformation is: Where (x, y) is a point in the first frame of two consecutive first intermediate images, and (x', y') is the corresponding point in the second frame of two consecutive first intermediate images, where x' = x + u, y' = y + v, and (u, v) is the optical flow information calculated above.
[0108] Solving for the affine transformation matrix is equivalent to solving for the six unknown parameters in the above formula. .
[0109] First, sample M points and obtain the coordinates (x, y) of each point in the first frame, as well as the optical flow information (u, v) of that point calculated above. For the i-th point (x...y ... i y i Its position in the second frame (x') i y' i ) is: x' i =x i+u,y' i =y i +v.
[0110] The above formula for two-dimensional affine transformation can be expanded into a linear equation as follows: For the M sampled points, the following formula can be used: make: The above formula can then be expressed as: .
[0111] Solve using the least squares method: Among them, S T Let S be the transpose of S.
[0112] After obtaining W, rearrange it to obtain the affine transformation matrix, which is: Repeating the above steps will yield the matrix sequence corresponding to the sliding window.
[0113] The above a 11 This represents a portion of the scaling and rotation components in the x-direction; a 12 This indicates the effect of the shearing transformation in the x-direction, and it is also related to rotation; a 21 This indicates the effect of the shearing transformation in the y-direction, which also participates in the rotation effect; a 22 This represents a portion of the scaling and rotation components in the y-direction; t x Indicates displacement in the horizontal direction; t y It represents the displacement in the vertical direction.
[0114] Step S240: Transform the first intermediate image according to the matrix sequence to obtain the second intermediate image.
[0115] In this embodiment, the first intermediate image is transformed according to the matrix sequence to obtain the second intermediate image.
[0116] Specifically, Figure 10 This is a flowchart illustrating the acquisition of a second intermediate image according to an embodiment of the present invention. Figure 10 As shown, transforming the first intermediate image according to the matrix sequence to obtain the second intermediate image includes the following steps: Step S241: Obtain the desired transformation matrix based on the matrix sequence corresponding to the sliding window.
[0117] In this embodiment, temporal filtering is performed on multiple radiative transformation matrices in the matrix sequence corresponding to the sliding window to obtain the desired transformation matrix of each first intermediate image. Temporal filtering can be performed using inertial filtering, Kalman filtering, Wiener filtering, or similar methods.
[0118] Taking Kalman filtering as an example, after obtaining the matrix sequence, each affine transformation matrix is decomposed into intuitive motion parameters, such as horizontal translation (t). x ), vertical translation (t) y The parameters include rotation angle, scaling ratio, etc. A Kalman filter is created to observe the changes in these four parameters. The Kalman filter is updated frame by frame, repeating the following operations for each frame: based on the state and motion trend of the previous frame, predict the value of the current frame; extract the actual value from the affine transformation of the current frame; the Kalman filter automatically weighs the predicted and actual values, outputting a more reliable and smoother estimate. The filtered parameters are then recombine into the affine transformation matrix, i.e., the expected transformation matrix.
[0119] Step S242: Transform the first intermediate image according to the desired transformation matrix to obtain the second intermediate image.
[0120] In this embodiment, the second frame in the first intermediate image between two adjacent frames is inversely transformed according to the expected transformation matrix to align it with the first frame, so as to obtain the second intermediate image.
[0121] The above steps can complete the processing of a sliding window. Performing the above steps on each sliding window will yield a complete second intermediate image.
[0122] Step S250: Generate a target video based on the second intermediate image.
[0123] In this embodiment, consecutive second intermediate images are encoded into the target video.
[0124] Therefore, by using information processing equipment to perform optical flow calculations on the first intermediate image after large displacement has been compensated, and then performing temporal filtering of optical flow vectors in different regions to obtain the target positions of each region of the image, the image is transformed to accurately compensate for residual displacement, rotation and deformation, ultimately achieving better image stabilization. Since the shooting device has already compensated for the large displacement, the complexity of optical flow calculation can be significantly reduced, and the deployment difficulty of the shooting device is also significantly reduced compared to traditional image content-based image stabilization algorithms.
[0125] This invention employs a shooting device to acquire initial data, including multiple initial images and corresponding pose data. The initial images are cropped based on the pose data, and the resulting first intermediate image sequence is sent to an information processing device. This first intermediate image sequence comprises multiple first intermediate images. The information processing device performs image stabilization processing on the first intermediate image sequence to generate the target video. Therefore, while considering the power consumption and data transmission volume of the shooting terminal, the overall image stabilization effect, real-time performance, and stability of video shooting can be improved.
[0126] Figure 11 This is a schematic diagram of the video shooting device of an embodiment of the present invention. Figure 11 As shown, the video shooting device of the shooting equipment in this embodiment of the invention includes an initial data acquisition unit 111, a rotation angle sequence determination unit 112, a displacement vector sequence determination unit 113, a desired displacement vector determination unit 114, a first intermediate image sequence acquisition unit 115, and a first intermediate image sequence transmission unit 116. The initial data acquisition unit 111 is used to acquire initial data, which includes multiple initial images and corresponding pose data. The rotation angle sequence determination unit 112 is used to determine a rotation angle sequence based on the initial images and pose data, whereby the rotation angle sequence includes the relative rotation angle between each pair of adjacent initial images in the target direction. The displacement vector sequence determination unit 113 is used to determine a displacement vector sequence based on the rotation angle sequence, whereby the displacement vector sequence includes the displacement vector between each pair of adjacent initial images in the target direction. The desired displacement vector determination unit 114 is used to determine the desired displacement vector for each initial image based on the displacement vector sequence. The first intermediate image sequence acquisition unit 115 is used to crop each of the initial images according to the desired displacement vector to obtain a first intermediate image sequence, whereby the first intermediate image sequence includes multiple first intermediate images. The first intermediate image sequence sending unit 116 is used to send the first intermediate image sequence to the information processing device so that the information processing device can generate a target video based on the first intermediate image sequence.
[0127] This invention employs a shooting device to acquire initial data, including multiple initial images and corresponding pose data. The initial images are cropped based on the pose data, and the resulting first intermediate image sequence is sent to an information processing device. This first intermediate image sequence comprises multiple first intermediate images. The information processing device performs image stabilization processing on the first intermediate image sequence to generate the target video. Therefore, while considering the power consumption and data transmission volume of the shooting terminal, the overall image stabilization effect, real-time performance, and stability of video shooting can be improved.
[0128] Figure 12 This is a schematic diagram of the video capturing device of the information processing apparatus according to an embodiment of the present invention. Figure 12 As shown, the video shooting device of the information processing equipment of this embodiment includes a first intermediate image sequence receiving unit 121, a sub-sequence acquisition unit 122, a matrix sequence acquisition unit 123, a second intermediate image acquisition unit 124, and a target video generation unit 125. The first intermediate image sequence receiving unit 121 receives a first intermediate image sequence, which includes multiple first intermediate images, each of which is a cropped image. The sub-sequence acquisition unit 122 acquires a sub-sequence through a pre-set sliding window, the sub-sequence including a predetermined number of first intermediate images. The matrix sequence acquisition unit 123 acquires a matrix sequence based on each first intermediate image in the sub-sequence, the matrix sequence including the transmission transformation matrix of each two adjacent frames of first intermediate images in the sub-sequence. The second intermediate image acquisition unit 124 transforms the first intermediate images according to the matrix sequence to obtain a second intermediate image. The target video generation unit 125 generates a target video based on the second intermediate image.
[0129] This invention employs a shooting device to acquire initial data, including multiple initial images and corresponding pose data. The initial images are cropped based on the pose data, and the resulting first intermediate image sequence is sent to an information processing device. This first intermediate image sequence comprises multiple first intermediate images. The information processing device performs image stabilization processing on the first intermediate image sequence to generate the target video. Therefore, while considering the power consumption and data transmission volume of the shooting terminal, the overall image stabilization effect, real-time performance, and stability of video shooting can be improved.
[0130] Figure 13 This is a schematic diagram of an electronic device according to an embodiment of the present invention. In this embodiment, the electronic device 13 includes a server, a terminal, etc. Figure 13As shown, the electronic device 13 includes at least one processor 131; a memory 132 communicatively connected to at least one processor 131; and a communication component 133 communicatively connected to a scanning device, wherein the communication component 133 receives and transmits data under the control of the processor 131; wherein the memory 132 stores instructions executable by at least one processor 131, the instructions being executed by at least one processor 131 to implement the above-described video capture method.
[0131] Specifically, the electronic device includes: one or more processors 131 and a memory 132. Figure 13 Taking a processor 131 as an example, the processor 131 and the memory 132 can be connected via a bus or other means. Figure 13 Taking a bus connection as an example, memory 132, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. Processor 131 executes various functional applications and data processing of the device by running the non-volatile software programs, instructions, and modules stored in memory 132, thereby realizing the above-mentioned video shooting method.
[0132] Memory 132 may include a program storage area and a data storage area, wherein the program storage area may store the operating system and applications required for at least one function; the data storage area may store an option list, etc. Furthermore, memory 132 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some embodiments, memory 132 may optionally include memory remotely located relative to processor 131, and these remote memories may be connected to external devices via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0133] One or more modules are stored in memory 132, and when executed by one or more processors 131, they perform the video capture method in any of the above method embodiments.
[0134] The above-mentioned products can perform the methods provided in the embodiments of this application, and have the corresponding functional modules and beneficial effects of performing the methods. For technical details not described in detail in this embodiment, please refer to the methods provided in the embodiments of this application.
[0135] This invention employs a shooting device to acquire initial data, including multiple initial images and corresponding pose data. The initial images are cropped based on the pose data, and the resulting first intermediate image sequence is sent to an information processing device. This first intermediate image sequence comprises multiple first intermediate images. The information processing device performs image stabilization processing on the first intermediate image sequence to generate the target video. Therefore, while considering the power consumption and data transmission volume of the shooting terminal, the overall image stabilization effect, real-time performance, and stability of video shooting can be improved.
[0136] Another embodiment of the present invention relates to a non-volatile storage medium for storing a computer-readable program for use by a computer to execute some or all of the above-described method embodiments.
[0137] That is, those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by a program instructing related hardware. This program is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods described in the embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0138] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A video shooting system, characterized in that, The system includes a camera and an information processing device, wherein: The imaging device is used to acquire initial data, which includes multiple initial images and corresponding pose data. The initial images are cropped according to the pose data to obtain a first intermediate image sequence, which includes multiple first intermediate images. The first intermediate image sequence is sent to the information processing device. The information processing device is used to receive a first intermediate image sequence and perform image stabilization processing on the first intermediate image sequence to generate a target video.
2. The system according to claim 1, characterized in that, The imaging device is used for: A rotation angle sequence is determined based on the initial image and pose data, the rotation angle sequence including the relative rotation angle of each two adjacent initial images in the target direction; The displacement vector sequence is determined based on the rotation angle sequence, and the displacement vector sequence includes the displacement vectors of each two adjacent initial images in the target direction; The desired displacement vector for each initial image is determined based on the displacement vector sequence. The initial images are cropped according to the desired displacement vector to obtain a first intermediate image sequence, which includes a plurality of first intermediate images.
3. The system according to claim 1, characterized in that, The information processing device is used for: Receive a first intermediate image sequence, the first intermediate image sequence including a plurality of first intermediate images, the first intermediate images being cropped images; A subsequence is obtained through a pre-set sliding window, the subsequence comprising a predetermined number of first intermediate images; A matrix sequence is obtained based on each first intermediate image in the sub-sequence, and the matrix sequence includes the transmission transformation matrix of each pair of adjacent first intermediate images in the sub-sequence. The first intermediate image is transformed according to the matrix sequence to obtain the second intermediate image; Generate the target video based on the second intermediate image.
4. A video shooting method, applicable to shooting equipment, characterized in that, The method includes: Acquire initial data, which includes multiple initial images and corresponding pose data; A rotation angle sequence is determined based on the initial image and pose data, the rotation angle sequence including the relative rotation angle of each two adjacent initial images in the target direction; The displacement vector sequence is determined based on the rotation angle sequence, and the displacement vector sequence includes the displacement vectors of each two adjacent initial images in the target direction; The desired displacement vector for each initial image is determined based on the displacement vector sequence. The initial images are cropped according to the desired displacement vector to obtain a first intermediate image sequence, the first intermediate image sequence including a plurality of first intermediate images; The first intermediate image sequence is sent to an information processing device to generate a target video based on the first intermediate image sequence.
5. The method according to claim 4, characterized in that, The acquisition of initial data includes: Acquire sensor data, which includes multiple initial images, timestamps corresponding to each initial image, pose data, and timestamps corresponding to each pose data. The initial image and pose data are aligned according to the timestamp corresponding to the initial image and the timestamp corresponding to the pose data to obtain the initial data.
6. The method according to claim 4, characterized in that, The target direction includes the vertical direction and / or the left and right direction.
7. The method according to claim 5, characterized in that, The pose data includes the angular velocity of the imaging device in the target direction at various times; The step of determining the rotation angle sequence based on the initial image and pose data includes: Determine the initial images of the two adjacent frames that need to be processed; The time interval between two adjacent initial images is determined based on the timestamps of the two adjacent initial images; The average angular velocity of two adjacent initial frames in the target direction is determined based on the pose data. The relative rotation angle between the two adjacent initial frames is determined based on the average angular velocity and the time interval.
8. The method according to claim 4, characterized in that, Determining the displacement vector sequence based on the rotation angle sequence includes: Acquire intrinsic parameter data from the image sensor; Determine the relative rotation angles to be processed from the rotation angle sequence; The displacement vectors of two adjacent initial images in the target direction are determined based on the intrinsic parameter data and the relative rotation angle.
9. The method according to claim 4, characterized in that, The step of determining the expected displacement vector of each initial image based on the displacement vector sequence includes: The displacement vector sequence is subjected to temporal filtering to obtain the desired displacement vector for each initial image.
10. A video shooting method, applicable to information processing equipment, characterized in that, The method includes: Receive a first intermediate image sequence, the first intermediate image sequence including a plurality of first intermediate images, the first intermediate images being cropped images; A subsequence is obtained through a pre-set sliding window, the subsequence comprising a predetermined number of first intermediate images; A matrix sequence is obtained based on each first intermediate image in the sub-sequence, and the matrix sequence includes the transmission transformation matrix of each pair of adjacent first intermediate images in the sub-sequence. The first intermediate image is transformed according to the matrix sequence to obtain the second intermediate image; Generate the target video based on the second intermediate image.
11. The method according to claim 10, characterized in that, The step of obtaining the matrix sequence based on each of the first intermediate images in the sub-sequence includes: An optical flow information sequence is obtained based on each first intermediate image in the sub-sequence, and the optical flow information sequence includes the optical flow information of each two adjacent first intermediate images in the sub-sequence. The matrix sequence is determined based on the optical flow information sequence, and the matrix sequence includes the affine transformation matrix of each two adjacent first intermediate images in the subsequence.
12. The method according to claim 10, characterized in that, The step of transforming the first intermediate image according to the matrix sequence to obtain the second intermediate image includes: The desired transformation matrix is obtained based on the matrix sequence corresponding to the sliding window. The first intermediate image is transformed according to the desired transformation matrix to obtain the second intermediate image.
13. The method according to claim 10, characterized in that, The step of generating the target video based on the second intermediate image includes: The invalid regions in the second intermediate image are cropped to obtain the third intermediate image; The target video is generated based on the third intermediate image.
14. An electronic device comprising a memory and a processor, characterized in that, The memory is used to store one or more computer program instructions, wherein the one or more computer program instructions are executed by the processor to implement the method as described in any one of claims 4-13.