Video jitter removal method, electronic device, system, and storage medium

Through neural radiation field technology and smooth camera path rendering method, the video jitter problem is solved, and efficient jitter removal is achieved. The generated video picture is coherent and jitter-free.

WO2025060587A9PCT designated stage expired Publication Date: 2025-06-05ARCSOFT CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/103322
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-09-20
Filing Date
2024-07-03
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

During video shooting, especially during movement, jitter problems often occur, resulting in the increasing demand for jitter removal technology in the video processing field.

Method used

The neural radiation field technology is used to render based on the generated smooth camera path to realize the processing of the video to be processed, thereby removing jitter. The specific steps include obtaining the video to be processed and the camera posture information, training the initial neural radiation field, generating a smooth camera path, and rendering the video using the trained neural radiation field.

Benefits of technology

Effectively removes video jitter, avoiding the possible frame loss caused by traditional methods, and the generated new video images are coherent and smooth.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024103322_05062025_PF_FP_ABST
    Figure CN2024103322_05062025_PF_FP_ABST
Patent Text Reader

Abstract

Provided are a video jitter removal method, an electronic device, a system, and a storage medium. The method comprises: obtaining a video to be processed and camera attitude information corresponding to the video to be processed (110); on the basis of the video to be processed and the camera attitude information, training an initial neural radiance field, and obtaining a trained neural radiance field (120); on the basis of the video to be processed and the camera attitude information, generating a smooth camera path (130); and on the basis of the smooth camera path, rendering a video scene of the video to be processed by using the trained neural radiance field, and generating a new video after jitter removal (140). According to the solution provided by the present application, neural radiance field technology is utilized and rendering is carried out on the basis of the generated smooth camera path, achieving the effect of jitter removal by means of processing of the video to be processed.
Need to check novelty before this filing date? Find Prior Art

Description

Video jitter removal method, electronic device, system and storage medium

[0001] This application claims priority to the Chinese patent application filed on September 20, 2023, with application number 202311221184X and invention name “A method, electronic device, system and storage medium for removing video jitter”, the content of which should be understood as incorporated into this application by reference. Technical Field

[0002] This article relates to but is not limited to the field of video processing technology. Background Art

[0003] Shooting videos has become a common daily activity, a way for people to record information and share their lives. However, jitter often occurs when ordinary users shoot videos with handheld devices, especially when moving around. Consequently, the demand for jitter removal technology in video processing is growing, and the requirements are becoming increasingly demanding.

[0004] Summary of the Invention

[0005] The following is a summary of the subject matter described in detail herein. This summary is not intended to limit the scope of the claims.

[0006] The embodiments of the present application provide a video jitter removal method, electronic device, system and storage medium, which utilize neural radiation field technology to perform rendering based on a generated smooth camera path to achieve processing of the video to be processed to achieve the effect of removing jitter.

[0007] The present invention provides a method for removing video jitter, including:

[0008] Obtaining a video to be processed and camera posture information corresponding to the video to be processed;

[0009] Training the initial neural radiation field according to the video to be processed and the camera posture information to obtain a trained neural radiation field;

[0010] generating a smooth camera path according to the video to be processed and the camera posture information;

[0011] According to the smoothed camera path, the video scene of the video to be processed is rendered using the trained neural radiance field to generate a new video with the jitter removed.

[0012] The embodiment of the present application further provides an electronic device, comprising one or more processors; a storage device for storing one or more programs;

[0013] When the one or more programs are executed by the one or more processors, the one or more processors implement the video jitter removal method as described in any embodiment of the present application.

[0014] The embodiment of the present application also provides a video jitter removal system, including a client and a server;

[0015] The client is configured to obtain a video to be processed and camera posture information corresponding to the video to be processed, and send the video to be processed and the camera posture information to the server;

[0016] The server is configured to train the initial neural radiation field according to the video to be processed and the camera posture information to obtain a trained neural radiation field;

[0017] The server is further configured to generate a smooth camera path according to the video to be processed and the camera posture information;

[0018] The server is further configured to render the video scene of the video to be processed using the trained neural radiation field according to the smooth camera path to generate a new video with the jitter removed.

[0019] An embodiment of the present application further provides a computer storage medium, wherein the storage medium stores a computer program, wherein the computer program is configured to execute the video jitter removal method described in any embodiment of the present application when running.

[0020] The video jitter removal system architecture proposed in the embodiment of the present application adopts a server-remote supported system architecture. The server is responsible for computing tasks that require high computing resources, such as neural radiation field training, camera posture correction, and rendering to generate new videos. This significantly alleviates the computing pressure on the client and ensures the practicality of the video jitter removal solution.

[0021] Other features and advantages of the present application will be described in the following description, and in part will become apparent from the description, or will be understood by practicing the present application. Other advantages of the present application can be realized and obtained by the solutions described in the description and the drawings.

[0022] Still other aspects will become apparent upon reading and understanding the accompanying drawings and detailed description.

[0023] Summary of the Figures

[0024] The accompanying drawings are used to provide an understanding of the technical solution of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the technical solution of the present application and do not constitute a limitation on the technical solution of the present application.

[0025] FIG1 is a flow chart of a method for removing video jitter provided by an embodiment of the present application;

[0026] FIG2 is a schematic structural diagram of a video jitter removal system provided in an embodiment of the present application;

[0027] FIG3 is a schematic structural diagram of another video jitter removal system provided in an embodiment of the present application;

[0028] FIG4 is a flow chart of another method for removing video jitter provided by an embodiment of the present application;

[0029] FIG5 is a flowchart of another method for removing video jitter provided in an embodiment of the present application.

[0030] Details

[0031] To make the purpose, technical solutions and advantages of this application more clear, the embodiments of this application will be described in detail below with reference to the accompanying drawings. It should be noted that, unless there is a conflict, the embodiments and features in the embodiments of this application can be combined with each other in any way.

[0032] To facilitate understanding of the present application, the present application will be described more fully below with reference to the accompanying drawings. The accompanying drawings provide embodiments of the present application. However, the present application may be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided to make the disclosure of the present application more thorough and comprehensive.

[0033] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application pertains. The terms used herein in the specification of this application are for the purpose of describing specific embodiments only and are not intended to limit this application.

[0034] It is understood that the terms "first" and "second" used in this application are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include at least one of such features. In the description of this application, "plurality" means at least two, for example, two, three, etc., unless otherwise specifically defined.

[0035] As used herein, the singular forms "a," "an," and "the" may also include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the terms "include," "comprising," "having," and the like specify the presence of stated features, integers, steps, operations, components, parts, or combinations thereof, but do not preclude the presence or addition of one or more other features, integers, steps, operations, components, parts, or combinations thereof. Furthermore, the term "and / or" as used in this specification includes any and all combinations of the relevant listed items.

[0036] This embodiment of the application provides a feasible solution that uses IMU (Inertial Measurement Unit) data transmitted by the camera to calculate the translation and rotation, and then warps the video frame by frame based on the translation and rotation to remove video jitter. The resulting video produced by this method will lose a certain amount of image quality.

[0037] The present application provides a method for removing video jitter, as shown in FIG1 , including:

[0038] Step 110: Obtain a video to be processed and camera posture information corresponding to the video to be processed;

[0039] Step 120: training the initial neural radiation field according to the video to be processed and the camera posture information to obtain a trained neural radiation field;

[0040] Step 130: generating a smooth camera path according to the video to be processed and the camera posture information;

[0041] Step 140: Render the video scene of the video to be processed using the trained neural radiance field according to the smoothed camera path to generate a new video with the jitter removed.

[0042] Among them, Neural Radiance Field (NERF) is an emerging new viewpoint synthesis technology that implicitly models the input video through a multi-layer perceptron, enabling realistic rendering of images from new viewpoints. This solution uses a fully connected network to represent 3D scenes, taking as input a 3D spatial position and viewpoint information, and outputting volume density and viewpoint-dependent color information at that spatial position. Furthermore, combined with volume rendering technology, this outputtable color information and volume density are rendered onto a 2D image, achieving new viewpoint synthesis and generating images from new viewpoints.

[0043] It should be noted that a camera path refers to the position and motion trajectory of a camera in a three-dimensional scene used to generate each frame of a given video sequence. In computer graphics, a video can be viewed as a sequence of consecutive image frames. To generate these frames, the camera's position, orientation, and movement within the scene must be determined. The camera path corresponding to a video describes the temporal evolution of the camera—that is, how the camera moves from one position and posture to the next, capturing each image frame. The camera path corresponding to a video can be obtained through various methods, including manually setting camera parameters, recording real-world camera motion using a motion capture system, or creating a virtual camera path through mathematical models and interpolation. By defining and controlling the camera path corresponding to a video, it is possible to achieve perspective transformations, object tracking, and the simulation of photographic effects, ultimately generating a video sequence with coherent motion and visual consistency.

[0044] In some exemplary embodiments, the camera path can be represented by a rotation that describes the position and orientation of the camera in three-dimensional space. Rotations are typically represented using Euler angles (such as pitch, yaw, and roll) or quaternions, and translations or displacements are represented using three-dimensional vectors.

[0045] In some exemplary embodiments, a camera path can be described by a series of discrete pathpoints, each of which contains information about the camera's position and pose. These pathpoints can be manually set or generated by other methods, such as real camera motion data recorded by a motion capture system or virtual pathpoints calculated using an interpolation algorithm.

[0046] In some exemplary embodiments, the camera path can be described using a parameterized curve. Common curve types include Bezier curves and spline curves. The parameterized curve describes the temporal changes of the camera and determines the camera's position and orientation at different points in time based on the curve shape and control points.

[0047] It can be understood that in the embodiment of the present application, the new video generated by rendering based on the smooth camera path is a video with more coherent pictures and smoother changes compared to the initial video obtained from the shooting device. Therefore, it is also called a video with removed jitter.

[0048] In some exemplary embodiments, obtaining a video to be processed and camera pose information corresponding to the video to be processed includes:

[0049] The camera posture information corresponding to each frame image in the video to be processed is obtained through a simultaneous localization and mapping (SLAM) algorithm.

[0050] The frame image is also called an image frame, or simply a frame.

[0051] In some exemplary embodiments, obtaining a video to be processed and camera pose information corresponding to the video to be processed includes:

[0052] Obtaining the initial camera posture information corresponding to each frame of the video to be processed through the SLAM algorithm;

[0053] The initial camera pose information is optimized by using a Structure from Motion (SFM) algorithm to obtain camera pose information corresponding to each frame of the video to be processed.

[0054] It can be understood that the camera pose information obtained after SFM algorithm optimization is more accurate than the initial camera pose information.

[0055] In some exemplary embodiments, generating a smooth camera path according to the video to be processed and the camera pose information includes:

[0056] Acquire multiple key frames in the video to be processed;

[0057] Interpolation processing is performed on the camera posture information corresponding to the multiple key frames to obtain a smooth camera path corresponding to the video to be processed.

[0058] In some exemplary embodiments, obtaining a plurality of key frames in the video to be processed includes:

[0059] The plurality of key frames are determined from all frame images included in the video to be processed according to a preset frame interval.

[0060] For example, if the preset frame interval is 5, a key frame is determined every 5 frames from the video to be processed. This setting can be flexibly determined based on the system's processing performance requirements and is not limited to a specific aspect. Alternatively, the frame interval can be dynamically determined based on the user's motion information. For example, during intense exercise, the frame interval is shorter, while during gentle movement, the frame interval is longer. More examples are not listed here.

[0061] In some exemplary embodiments, obtaining a plurality of key frames in the video to be processed includes: selecting, in the video to be processed, a frame image that satisfies a first condition as a key frame, wherein the first condition is that a change in attribute information of the frame image on the original camera path compared to attribute information of a previous frame image on the original camera path is greater than a first threshold;

[0062] The original camera path is a camera path determined according to the camera posture information corresponding to the video to be processed.

[0063] In some exemplary embodiments, obtaining a plurality of key frames in the video to be processed includes: obtaining a plurality of key frames selected by a user from the video to be processed.

[0064] In some exemplary embodiments, a user interaction instruction is received through a user interaction module to determine the selected multiple key frames.

[0065] In some exemplary embodiments, the attribute information of the frame image on the camera path includes one or more of the following:

[0066] Location information and orientation information.

[0067] In some exemplary embodiments, the orientation information is simply referred to as orientation, also called rotation information, or direction information.

[0068] In some exemplary embodiments, when the attribute information includes position information and orientation information, the attribute information is also referred to as camera posture information, or simply posture information.

[0069] In some exemplary embodiments, the position information is coordinates; in some exemplary embodiments, the coordinates are three-dimensional coordinates.

[0070] A corresponding first threshold is set based on the attribute information. For example, if the attribute information includes position information, the first threshold is a distance threshold. If the change in the position information of the frame image on the original camera path compared to the position information of the previous frame image on the original camera path is greater than the distance threshold, the frame image is determined to be a key frame; that is, if the distance between the position of the frame image on the original camera path and the position of the previous frame image on the original camera path is greater than the distance threshold.

[0071] For another example, when the attribute information includes orientation information, the first threshold is the direction angle difference threshold. When the change amplitude of the orientation information of the frame image on the original camera path compared to the orientation information of the previous frame image on the original camera path is greater than the direction angle threshold, the frame image is determined to be a key frame; that is, the angle difference between the orientation of the frame image on the original camera path and the orientation of the previous frame image on the original camera path is greater than the distance direction angle difference threshold.

[0072] For another example, when the attribute information includes position information and orientation information, the first threshold also correspondingly includes a distance threshold and a direction angle difference threshold. If the magnitude of the change in the orientation information of the frame image on the original camera path compared to the orientation information of the previous frame image on the original camera path is greater than the direction angle threshold, or if the magnitude of the change in the position information of the frame image on the original camera path compared to the position information of the previous frame image on the original camera path is greater than the distance threshold, the frame image is determined to be a key frame.

[0073] Alternatively, if the magnitude of the change in the orientation information of the frame image on the original camera path compared to the orientation information of the previous frame image on the original camera path is greater than the direction angle threshold, and the magnitude of the change in the position information of the frame image on the original camera path compared to the position information of the previous frame image on the original camera path is greater than the distance threshold, the frame image is determined to be a key frame. More examples are not listed here one by one.

[0074] It is understood that due to jitter during filming, the original camera path corresponding to the video being processed may not be smooth, exhibiting large jumps in spatial position, orientation, or posture. In some exemplary implementations, frame images at path points where attribute information changes significantly are selected to form the multiple keyframes. Interpolation processing is then performed based on these selected keyframes to obtain a smooth camera path corresponding to the video being processed.

[0075] In some exemplary embodiments, the duration corresponding to the smoothed camera path obtained after interpolation processing based on multiple key frames is the same as the duration of the video to be processed.

[0076] In some exemplary embodiments, the duration of the smoothed camera path obtained by interpolating the multiple key frames is the same as the duration of the video to be processed. It is understood that the new video rendered based on the smoothed camera path of the same duration has the same duration as the video to be processed obtained in step 110.

[0077] In some exemplary embodiments, linear interpolation is used to interpolate the position information, and interpolation processing is performed based on the position information corresponding to multiple key frames to obtain the position information of the insertion point.

[0078] In some exemplary embodiments, spherical linear interpolation is used to interpolate the orientation information, and interpolation processing is performed based on the orientation information corresponding to multiple key frames to obtain the orientation information of the insertion point.

[0079] In some exemplary embodiments, the duration of the smoothed camera path obtained by interpolating the multiple key frames is shorter than the duration of the video to be processed. It is understood that the duration of the new video rendered based on the smoothed camera with a shorter duration is shorter than the duration of the video to be processed obtained in step 110.

[0080] In some exemplary embodiments, generating a smooth camera path according to the video to be processed and the camera pose information includes:

[0081] For each frame image in the video to be processed, correcting the camera posture information corresponding to the frame image using multiple frames of images adjacent to the frame image according to a preset number of iterations;

[0082] A smooth camera path corresponding to the video to be processed is generated according to the corrected camera posture information corresponding to each frame image in the video to be processed.

[0083] In some exemplary embodiments, the camera posture information includes: coordinates and orientation.

[0084] In some exemplary embodiments, the coordinates include three-dimensional coordinates.

[0085] In some exemplary embodiments, the orientation comprises a quaternion.

[0086] In some exemplary embodiments, correcting the posture information includes correcting coordinates and correcting orientation.

[0087] In some exemplary embodiments, correcting the camera pose information corresponding to each frame image in the video to be processed by using multiple frames of images adjacent to the frame image according to a preset number of iterations includes:

[0088] For each frame of the video to be processed, the camera pose information corresponding to the frame is used as the initial camera pose information to be corrected. According to the preset number of iterations, in each round of iteration, the following steps are performed respectively:

[0089] Acquiring image parameters of the adjacent multi-frame images, and correcting the to-be-corrected camera posture information of the frame image according to the image parameters of the adjacent multi-frame images, wherein the image parameters include coordinates, orientation, and weight;

[0090] The corrected camera pose information is used as the camera pose information to be corrected in the next iteration.

[0091] That is, multiple iterative corrections are performed on the camera pose information of each frame image in the video to be processed, and the total number of repetitions is a preset number of iterations N. The camera pose information corresponding to the frame image in the video to be processed is used as the initial camera pose information to be corrected, which is corrected in the first iteration, and the corrected camera pose information is used as the camera pose information to be corrected in the next iteration, until N corrections are completed to obtain the final corrected camera pose information.

[0092] In some exemplary embodiments, the preset number of iterations N=3, that is, each initial camera pose information to be corrected is iterated 3 times to obtain the final correction result. The specific value of the number of iterations can be flexibly set as needed and is not limited to the aspects of the embodiments of the present application.

[0093] In some exemplary embodiments, correcting the camera pose information corresponding to each frame image in the video to be processed by using multiple frames of images adjacent to the frame image according to a preset number of iterations includes:

[0094] Repeat the following method for calibration according to the preset number of iterations:

[0095] according to Determine the corrected coordinate information;

[0096] according to Determine the corrected orientation information;

[0097] Where p is the three-dimensional coordinate to be corrected, d is the orientation quaternion to be corrected; n is the number of multiple frames adjacent to the frame image to be corrected; p i is the three-dimensional coordinate of the i-th frame image, {p i |i=1…n},d i is the orientation quaternion of the i-th frame image, {d i |i=1…n};w i is the weight of the i-th frame image, {w i |i=1…n};λ is the set correction coefficient,λ∈[0,1]; is the Frobenius norm, A(d) is the orthogonal attitude matrix of quaternion d; p * is the corrected three-dimensional coordinate, d * is the corrected orientation quaternion.

[0098] Wherein, in the first iteration, the three-dimensional coordinates to be corrected are the three-dimensional coordinates corresponding to the frame image to be corrected, and the orientation quaternion to be corrected is the quaternion corresponding to the frame image to be corrected;

[0099] In a non-first iteration, the three-dimensional coordinates to be corrected are the corrected three-dimensional coordinates obtained in the previous iteration, and the orientation quaternion to be corrected is the corrected quaternion obtained in the previous iteration.

[0100] In some exemplary embodiments, the weight w i Determined according to the following method:

[0101] It should be noted that, in the above correction method, the multiple frames of images adjacent to the frame image are n frames of images adjacent to the frame image, where n is an integer greater than or equal to 1 and is flexibly set according to needs and is not limited to a specific aspect.

[0102] In some exemplary embodiments, the training of the initial neural radiation field based on the video to be processed and the camera pose information includes:

[0103] Acquire a plurality of training samples according to the video to be processed and the camera posture information, wherein each training sample is composed of light emitted by a pixel point of a frame image in the video to be processed and a corresponding color;

[0104] The initial neural radiation field is trained according to the multiple training samples.

[0105] In some exemplary embodiments, the light emitted by a pixel in a frame image of the video to be processed is determined based on the camera pose information corresponding to the frame image in which the pixel is located and the position of the pixel in the frame image to which the pixel belongs. In other words, the light emitted by the pixel is determined based on the camera pose information corresponding to the frame image in which the pixel is located and the position of the pixel in the image.

[0106] In some exemplary embodiments, the step of obtaining a plurality of training samples according to the video to be processed and the camera pose information includes:

[0107] Determine a plurality of data groups according to the video to be processed and the camera posture information, each data group including a frame image and corresponding camera posture information;

[0108] For each data set containing frame images and camera pose information, perform the following steps to obtain multiple training samples:

[0109] Analyze the frame image and camera posture information into the light emitted by each pixel;

[0110] The light emitted by each pixel and the color of the pixel form a sample.

[0111] In some exemplary embodiments, the color of each pixel p in a frame image is c. Combining the camera pose information and pixel position corresponding to the frame image, the light emitted by the pixel can be obtained and recorded as l(p, d), where p = (x, y, z) is the position coordinate of the pixel in a three-dimensional Cartesian coordinate system, and d = (θ, φ) is the solid angle parameter of the light direction in a spherical coordinate system. θ represents the polar angle or latitude, which is the angle between the vector from a reference axis (usually the positive z-axis) to the point and the reference axis, and the value of θ typically ranges from 0 to π. φ represents the azimuth or longitude, which is the angle between the projection of a reference direction on a reference plane (usually the positive x-axis) to the point and the reference direction, and the value of φ typically ranges from 0 to 2π.

[0112] It can be understood that a frame image in the video to be processed includes multiple pixel points, corresponding to multiple training samples composed of light and color emitted by the pixel points, that is, one pixel point corresponds to one training sample, and the multiple image frames in the video to be processed obtain more training samples.

[0113] In some exemplary embodiments, training the initial neural radiation field according to the plurality of training samples includes:

[0114] All training samples are randomly shuffled and the initial neural radiation field is trained.

[0115] For each input sample (l, c), l is light and c is color. The coordinate p is spatially deformed in the divided space and its position is queried through hash coding. The multi-layer perceptron of the node is used to encode it. The encoded features f and d are encoded together using a global multi-layer perceptron to output color c. pred and dissimilarity disp.

[0116] Among them, the loss function

[0117] In some exemplary embodiments, the spatial division in the neural radiation field may be a uniform grid-based spatial division or an octree-based spatial division.

[0118] In some exemplary embodiments, the spatial deformation in the neural radiation field may be a spatial deformation based on normalized device coordinates, or a spatial deformation based on a perspective projection coordinate system.

[0119] It can be understood that the video de-shaking solution provided in the embodiment of the present application obtains training samples from the video to be processed to train the neural radiation field, utilizes the new viewpoint synthesis technology based on the neural radiation field, and then renders and generates a new video with de-shaking corresponding to the video to be processed by smoothing the camera path obtained according to the video to be processed. It can effectively remove video jitter and avoid the frame loss caused by some de-shaking solutions.

[0120] An embodiment of the present application further provides an electronic device, comprising:

[0121] one or more processors;

[0122] a storage device for storing one or more programs,

[0123] When the one or more programs are executed by the one or more processors, the one or more processors implement the video jitter removal method as described in any embodiment of the present application.

[0124] The embodiment of the present application further provides a video jitter removal system, as shown in FIG2 , comprising:

[0125] Client 210 and server 220;

[0126] The client 210 is configured to obtain a video to be processed and camera posture information corresponding to the video to be processed, and send the video to be processed and the camera posture information to the server 220;

[0127] The server 220 is configured to train the initial neural radiation field according to the video to be processed and the camera posture information to obtain a trained neural radiation field;

[0128] The server 220 is further configured to generate a smooth camera path according to the video to be processed and the camera posture information;

[0129] The server 220 is further configured to render the video scene of the video to be processed using the trained neural radiance field according to the smoothed camera path to generate a new video with the jitter removed.

[0130] In some exemplary embodiments, the server 220 is further configured to send the new video after de-shaking to the client 210;

[0131] Correspondingly, the client 210 is further configured to display the new video after the jitter is removed.

[0132] In some exemplary embodiments, the client 210 is further configured to obtain a plurality of key frames in the video to be processed and send the plurality of key frames to the server 220 .

[0133] In some exemplary embodiments, the client 210 includes: a user interaction module configured to receive a user operation instruction and determine the selected multiple key frames from the video to be processed.

[0134] In some exemplary embodiments, the server 220 is further configured to perform interpolation processing on camera pose information corresponding to multiple key frames in the video to be processed, so as to obtain a smooth camera path corresponding to the video to be processed.

[0135] In some exemplary embodiments, the server 220 or the client 210 selects the multiple key frames from the video to be processed, and the frame images that meet the first condition are used as key frames;

[0136] Among them, the first condition is that the change in the attribute information of the frame image on the original camera path compared to the attribute information of the previous frame image on the original camera path is greater than a first threshold; the original camera path is a camera path determined according to the camera posture information corresponding to the video to be processed.

[0137] As can be seen, the key frames used to determine a smooth camera path can be determined by the server 220 based on the video to be processed, or by the client 210 based on the video to be processed, or by the client 210 based on user operation instructions. For example, during video capture or video playback, the user uses the user interaction module to mark and determine multiple key frames.

[0138] The embodiment of the present application further provides a video jitter removal system, as shown in FIG3 , comprising:

[0139] Client 210 and server 220;

[0140] The client 210 includes: an image acquisition module 2110, a posture information acquisition module 2120, and a user interaction module 2130;

[0141] The server 220 includes: a neural radiation field training module 2210 , a smooth path determination module 2220 , and a neural radiation field rendering module 2230 .

[0142] The image acquisition module 2110 is configured to acquire a video to be processed; it can be understood that the video to be processed includes multiple frame images.

[0143] The posture information acquisition module 2120 is configured to acquire the camera posture information corresponding to the video to be processed; accordingly, the corresponding camera posture information includes the camera posture information corresponding to each frame of image.

[0144] The user interaction module 2130 is configured to receive a user operation instruction and determine the multiple key frames selected from the video to be processed.

[0145] The neural radiation field training module 2210 is configured to train the initial neural radiation field according to the video to be processed and the camera posture information to obtain a trained neural radiation field.

[0146] The smooth path determination module 2220 is configured to generate a smooth camera path according to the video to be processed and the camera posture information.

[0147] The neural radiation field rendering module 2230 is configured to render the video to be processed using the trained neural radiation field according to the smoothed camera path to obtain a rendering result.

[0148] In some exemplary embodiments, the posture information acquisition module 2120 is configured to acquire the camera posture information corresponding to each frame image in the video to be processed through a SLAM algorithm.

[0149] In some exemplary embodiments, the posture information acquisition module 2120 is configured to obtain the initial camera posture information corresponding to each frame image in the video to be processed through the SLAM algorithm; and optimize the initial camera posture information through the SFM algorithm to obtain the camera posture information corresponding to each frame image in the video to be processed.

[0150] In some exemplary embodiments, the client 210 further includes: a first data transceiver module configured to send the video to be processed and camera posture information corresponding to the video to be processed to the server 220 .

[0151] In some exemplary embodiments, the first data transceiver module is further configured to send the multiple key frames selected by the user to the server 220 .

[0152] In some exemplary embodiments, the first data transceiver module is further configured to receive a rendering result from the server 220 .

[0153] In some exemplary embodiments, the user interaction module 2130 is configured to generate a new video after de-jittering according to the rendering result.

[0154] In some exemplary embodiments, the server 220 further includes: a second data transceiver module configured to receive the video to be processed and camera posture information corresponding to the video to be processed from the client 210 .

[0155] In some exemplary embodiments, the second data transceiver module is further configured to receive the selected key frame from the client 210 .

[0156] In some exemplary embodiments, the second data transceiver module is further configured to send the rendering result to the client 210 .

[0157] In some exemplary embodiments, the smooth path determination module 2220 is configured to obtain multiple key frames in the video to be processed; and interpolate the camera posture information corresponding to the multiple key frames to obtain a smooth camera path corresponding to the video to be processed.

[0158] In some exemplary embodiments, the smooth path determination module 2220 is configured to correct the camera posture information corresponding to each frame image in the video to be processed using multiple frame images adjacent to the frame image according to a preset number of iterations; and generate a smooth camera path corresponding to the video to be processed based on the corrected camera posture information corresponding to each frame image in the video to be processed.

[0159] The present application also provides a method for removing video jitter, as shown in FIG4 , including:

[0160] Step 410: The image acquisition module acquires the video to be processed;

[0161] Step 420: The posture information acquisition module acquires the camera posture information corresponding to the video to be processed through the SLAM algorithm;

[0162] Step 430: The first data transceiver module sends the video to be processed and the camera posture information corresponding to the video to be processed to the server;

[0163] Step 440: The neural radiation field training module trains the initial neural radiation field according to the video to be processed and the camera posture information to obtain a trained neural radiation field.

[0164] Step 450: The user interaction module obtains multiple key frames and sends them to the server via the first data transceiver module;

[0165] Step 460: The smooth path determination module performs interpolation processing based on the multiple key frames to obtain a smooth camera path;

[0166] Step 470: The neural radiance field rendering module renders the video scene of the video to be processed using the trained neural radiance field according to the smoothed camera path to generate a new video with the jitter removed.

[0167] Step 480: Send the new video to the client via the second transceiver module;

[0168] Step 490: The user interaction module displays the new video.

[0169] In some exemplary embodiments, steps 450-460 are replaced by steps 451-461, as shown in FIG5 :

[0170] Step 451: The smooth path determination module corrects the camera pose information corresponding to each frame image in the video to be processed using multiple frames of images adjacent to the frame image according to a preset number of iterations.

[0171] Step 461 : The smooth path determination module generates a smooth camera path corresponding to the video to be processed based on the corrected camera posture information corresponding to each frame of the video to be processed.

[0172] It can be seen that the system architecture proposed in the disclosed embodiment is that the client performs the basic steps of video acquisition and camera posture information acquisition, and the server performs neural radiation field training, camera posture information correction, and rendering to generate new videos, which are steps that require high computing resources. Taking full account of the implementation characteristics of the neural radiation field solution, a distributed computing solution combining server and client is adopted to improve the feasibility and practicality of the solution. The powerful computing power of the server is utilized to avoid the insufficient computing power that may be faced by relying solely on the local execution of the solution by the video shooting equipment, which affects the final jitter removal effect.

[0173] An embodiment of the present application further provides a computer storage medium, wherein the storage medium stores a computer program, wherein the computer program is configured to execute the video jitter removal method as described in any embodiment of the present application when running.

[0174] The video jitter removal solution provided in the embodiments of the present application adopts a new neural radiation field solution, and implicitly models the input video through a multi-layer perceptron, so that it can render a realistic picture from a new viewpoint, overcome the frame loss caused by some jitter removal solutions, and achieve a good jitter removal effect. In some exemplary embodiments, the solution that supports users to select key frames and then determine a smooth camera path can fully meet the user's viewpoint setting needs. In some exemplary embodiments, an automatic correction method is used to correct the camera posture of the video to be processed, and then a smooth camera path is obtained, which significantly improves the jitter removal effect. In some exemplary embodiments, a server-remotely supported system architecture is adopted, and the server undertakes computing tasks that require high computing resources, such as neural radiation field training, camera posture correction, and rendering to generate new videos, which significantly alleviates the computing pressure on the client and ensures the practicality of the solution of the present application.

[0175] It will be appreciated by those skilled in the art that all or some of the steps, systems, and functional modules / units in the methods disclosed above may be implemented as software, firmware, hardware, and appropriate combinations thereof. In hardware implementations, the division between the functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed by several physical components in cooperation. Some or all components may be implemented as software executed by a processor, such as a digital signal processor or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on a computer-readable medium, which may include a computer storage medium (or non-transitory medium) and a communication medium (or temporary medium). As is well known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable, and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, it is well known to those skilled in the art that communication media generally embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.

Claims

1. A video jitter removal method, comprising: Obtaining a video to be processed and camera posture information corresponding to the video to be processed; Training the initial neural radiation field according to the video to be processed and the camera posture information to obtain a trained neural radiation field; Generate a smooth camera path according to the video to be processed and the camera posture information; According to the smoothed camera path, the video scene of the video to be processed is rendered using the trained neural radiance field to generate a new video with the jitter removed.

2. The video jitter removal method according to claim 1, wherein: The obtaining of the video to be processed and the camera posture information corresponding to the video to be processed includes: The camera posture information corresponding to each frame image in the video to be processed is obtained by synchronously positioning and mapping the SLAM algorithm.

3. The video jitter removal method according to claim 1, wherein: The obtaining of the video to be processed and the camera posture information corresponding to the video to be processed includes: Acquire the initial camera posture information corresponding to each frame image in the video to be processed by synchronous positioning and mapping SLAM algorithm; The initial camera posture information is optimized by using the structure from motion (SFM) algorithm to obtain the camera posture information corresponding to each frame of the video to be processed.

4. The video jitter removal method according to claim 1, wherein: The step of generating a smooth camera path according to the video to be processed and the camera posture information includes: Acquire multiple key frames in the video to be processed; Interpolation processing is performed on the camera posture information corresponding to the multiple key frames to obtain a smooth camera path corresponding to the video to be processed.

5. The video jitter removal method according to claim 4, wherein: The step of obtaining a plurality of key frames in the video to be processed includes: From the video to be processed, a frame image that meets a first condition is selected as a key frame, wherein the first condition is that a change in attribute information of the frame image on the original camera path compared to attribute information of a previous frame image of the frame image on the original camera path is greater than a first threshold; The original camera path is a camera path determined according to the camera posture information corresponding to the video to be processed.

6. The video jitter removal method according to claim 1, wherein: The step of generating a smooth camera path according to the video to be processed and the camera posture information includes: For each frame image in the video to be processed, according to a preset number of iterations, using multiple frames of images adjacent to the frame image, correct the camera posture information corresponding to the frame image; A smooth camera path corresponding to the video to be processed is generated according to the corrected camera posture information corresponding to each frame image in the video to be processed.

7. The video jitter removal method according to claim 6, wherein: The method of correcting the camera posture information corresponding to each frame image in the video to be processed by using multiple frames of images adjacent to the frame image according to a preset number of iterations includes: For each frame of the video to be processed, the camera posture information corresponding to the frame is used as the initial camera posture information to be corrected. According to the preset number of iterations, in each round of iteration, the following steps are performed respectively: Acquire image parameters of the adjacent multi-frame images, and correct the to-be-corrected camera posture information of the frame image according to the image parameters of the adjacent multi-frame images, wherein the image parameters include coordinates, orientations, and weights; The corrected camera posture information is used as the camera posture information to be corrected for the next round of iteration.

8. The video jitter removal method according to any one of claims 1 to 7, wherein: The training of the initial neural radiation field according to the video to be processed and the camera posture information includes: Acquire a plurality of training samples according to the video to be processed and the camera posture information, wherein each training sample is composed of light emitted by a pixel point of a frame image in the video to be processed and a corresponding color; The initial neural radiation field is trained according to the multiple training samples.

9. The video jitter removal method according to claim 8, wherein: The light emitted by the pixel point of the frame image in the video to be processed is determined according to the camera posture information corresponding to the frame image where the pixel point is located and the position of the pixel point in the belonging frame image.

10. An electronic device comprising: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the video jitter removal method according to any one of claims 1 to 9.

11. A video jitter removal system, comprising: Client and server; in, The client is configured to obtain a video to be processed and camera posture information corresponding to the video to be processed, and send the video to be processed and the camera posture information to the server; The server is configured to train the initial neural radiation field according to the video to be processed and the camera posture information to obtain a trained neural radiation field; The server is further configured to generate a smooth camera path according to the video to be processed and the camera posture information; The server is also configured to render the video scene of the video to be processed using the trained neural radiation field according to the smooth camera path to generate a new video with the jitter removed.

12. The video jitter removal system according to claim 11, wherein: The server is further configured to perform interpolation processing on camera posture information corresponding to multiple key frames in the video to be processed, so as to obtain a smooth camera path corresponding to the video to be processed.

13. The video jitter removal system according to claim 11, wherein: The client is also configured to obtain multiple key frames in the video to be processed and send the multiple key frames to the server.

14. The video jitter removal system according to claim 13, wherein: The client comprises: a user interaction module, which is configured to receive user operation instructions and determine the selected multiple key frames from the video to be processed.

15. A computer storage medium, wherein a computer program is stored in the storage medium, wherein: The computer program is configured to execute the video jitter removal method according to any one of claims 1 to 9 when running.