Video processing method and device, electronic equipment and storage medium

By using prediction and adjustment techniques of multi-frame noise potential variables in video depth estimation, the depth map consistency problem in long videos is solved, and higher video depth estimation accuracy is achieved.

CN119967201APending Publication Date: 2025-05-09TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510122122.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

Existing video depth estimation technology is difficult to maintain the temporal consistency and geometric consistency between generated depth maps in long video application scenarios, which affects the accuracy of video depth estimation.

Method used

By obtaining multi-frame noise latent variables with the same number of frames as the target video, perform noise prediction, remove prediction noise to obtain original depth latent variables, divide overlapping subsequences to determine geometric constraint loss, and adjust predicted noise to improve geometric consistency.

Benefits of technology

While keeping the computing resources unchanged, it supports longer video input, which improves the time consistency and geometric consistency between depth maps, thereby improving the accuracy of video depth estimation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119967201A_ABST
    Figure CN119967201A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a video processing method and device, electronic equipment and a storage medium, and the method comprises the steps: obtaining a noise potential variable of a target video, carrying out the noise prediction based on the noise potential variable, obtaining predicted noise, removing the predicted noise from the noise potential variable, obtaining an original depth potential variable, and carrying out the prediction of the original depth potential variable. Decoding the original depth potential variable to obtain an original depth map corresponding to each video frame in the target video, dividing the prediction noise into a plurality of overlapped sub-sequences, determining geometric constraint loss in each sub-sequence based on the original depth map, adjusting the prediction noise of a non-overlapped part based on the geometric constraint loss, and obtaining the prediction noise of the target video. And respectively removing the adjusted prediction noise from the noise potential variables to obtain a target depth potential variable, and decoding the target depth potential variable to obtain a target depth map corresponding to each video frame, so that the time consistency and geometric consistency between the generated depth maps can be better maintained, and the accuracy of the depth map generation is improved. Therefore, the accuracy of video depth estimation is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to a video processing method, device, electronic device and storage medium. Background Art

[0002] Video depth estimation is a long-standing task in the field of computer vision. In related technologies, when performing video depth estimation, the model used generally only relies on data priors trained in two-dimensional image space. In the application scenario of long videos, it is impossible to maintain the consistency between the generated depth maps well, and the accuracy of video depth estimation needs to be improved. Summary of the invention

[0003] The following is a summary of the subject matter of the detailed description of the present disclosure. This summary is not intended to limit the scope of the claims.

[0004] The embodiments of the present disclosure provide a video processing method, device, electronic device and storage medium, which can better maintain the temporal consistency and geometric consistency between the generated depth maps in the application scenario of long videos, thereby improving the accuracy of video depth estimation.

[0005] On the one hand, an embodiment of the present disclosure provides a video processing method, including:

[0006] Acquire multiple frames of noise latent variables having the same number of frames as the target video, perform noise prediction based on the multiple frames of noise latent variables, and obtain multiple frames of predicted noise;

[0007] Removing the predicted noise from the noise latent variables of each frame respectively to obtain multiple frames of original depth latent variables, decoding the multiple frames of original depth latent variables to obtain original depth maps corresponding to each video frame in the target video;

[0008] Dividing the predicted noise of multiple frames into multiple overlapping subsequences, determining a geometric constraint loss in each of the subsequences based on the original depth map, and adjusting the predicted noise of the non-overlapping part based on the geometric constraint loss, wherein the geometric constraint loss is used to constrain geometric consistency when three-dimensionally projecting the corresponding video frame based on the original depth map;

[0009] The adjusted predicted noise is removed from the noise latent variables of each frame respectively to obtain multi-frame target depth latent variables, and the multi-frame target depth latent variables are decoded to obtain multi-frame target depth maps corresponding to each of the video frames.

[0010] On the other hand, the present disclosure also provides a video processing device, including:

[0011] A noise prediction module is used to obtain multiple frames of noise latent variables with the same number of frames as the target video, and perform noise prediction based on the multiple frames of noise latent variables to obtain multiple frames of predicted noise;

[0012] A first generating module is used to remove the predicted noise from the noise latent variables of each frame respectively to obtain multiple frames of original depth latent variables, and decode the multiple frames of the original depth latent variables to obtain the original depth map corresponding to each video frame in the target video;

[0013] An adjustment module, configured to divide the predicted noise of multiple frames into multiple overlapping subsequences, determine a geometric constraint loss in each of the subsequences based on the original depth map, and adjust the predicted noise of the non-overlapping part based on the geometric constraint loss, wherein the geometric constraint loss is used to constrain geometric consistency when the corresponding video frame is three-dimensionally projected based on the original depth map;

[0014] The second generation module is used to remove the adjusted predicted noise from the noise latent variables of each frame, obtain multi-frame target depth latent variables, decode the multi-frame target depth latent variables, and obtain multi-frame target depth maps corresponding to each of the video frames.

[0015] Furthermore, the adjustment module is also used to:

[0016] Determining a first constraint loss based on the original depth map, wherein the first constraint loss is used to constrain scene consistency when three-dimensionally projecting the corresponding video frame based on the original depth map;

[0017] Determining a second constrained loss based on the original depth map, wherein the second constrained loss is used to constrain detail consistency when three-dimensionally projecting the corresponding video frame based on the original depth map;

[0018] A weighted sum is performed on the first constraint loss and the second constraint loss to obtain a geometric constraint loss.

[0019] Furthermore, the adjustment module is also used to:

[0020] Determine a three-dimensional projection relationship between the video frame of the i-th frame and the video frame of the j-th frame based on the original depth map of the i-th frame, perform three-dimensional projection on the video frame of the j-th frame relative to the video frame of the i-th frame based on the three-dimensional projection relationship to obtain a first projection frame, and determine a reprojection loss according to a difference between the first projection frame and the video frame of the i-th frame, wherein i and j are positive integers;

[0021] Determine a projection depth map when three-dimensionally projecting the j-th video frame, and determine a depth loss according to a difference between the projection depth map and the original depth map;

[0022] The reprojection loss and the depth loss are weightedly summed to obtain a first constraint loss.

[0023] Furthermore, the adjustment module is also used to:

[0024] Taking the predicted noise of the i-th frame and the predicted noise of the j-th frame as a predicted noise pair, determining a first mean value of the reprojection loss corresponding to all the predicted noise pairs in the subsequence, and a second mean value of the depth loss corresponding to all the predicted noise pairs in the subsequence;

[0025] A first constraint loss is obtained by performing a weighted summation on the first mean and the second mean.

[0026] Furthermore, the adjustment module is also used to:

[0027] Based on the original depth maps corresponding to the two adjacent video frames, three-dimensionally project the two adjacent video frames to obtain a first projection point set and a second projection point set;

[0028] Converting the coordinate system of the second projection point set to the coordinate system of the first projection point set, and determining the tracking loss according to the difference between the first projection point set and the converted second projection point set;

[0029] The reprojection loss, the depth loss and the tracking loss are weightedly summed to obtain a first constraint loss.

[0030] Furthermore, the adjustment module is also used to:

[0031] Determining a structural similarity between the first projection frame and the i-th video frame;

[0032] The difference between the first projection frame and the i-th video frame and the structural similarity are weightedly summed to obtain a reprojection loss.

[0033] Furthermore, the adjustment module is also used to:

[0034] Calculating a first surface normal based on the original depth map, generating a second surface normal of the corresponding video frame based on a surface normal generation network, and determining a surface normal loss according to a difference between the first surface normal and the second surface normal;

[0035] Determining a depth gradient of the original depth map, and regularizing the depth gradient to obtain a smoothing loss;

[0036] The surface normal loss and the smoothness loss are weightedly summed to obtain a second constraint loss.

[0037] Furthermore, the adjustment module is also used to:

[0038] Determine a third mean of the surface normal loss corresponding to all the prediction noises in the subsequence, and a fourth mean of the smoothing loss corresponding to all the prediction noises in the subsequence;

[0039] A weighted sum is performed on the third mean and the fourth mean to obtain a second constraint loss.

[0040] Furthermore, the adjustment module is also used to:

[0041] Determining the gradient of the geometric constraint loss, and obtaining an adjustment term according to the product of the gradient and the guidance strength;

[0042] The adjusted prediction noise is obtained according to the sum of the adjustment item and the prediction noise of the non-overlapping part.

[0043] Furthermore, the adjustment module is also used to:

[0044] Obtaining a multi-frame sample depth map corresponding to a sample video, extracting a sample depth latent variable of the sample depth map and a multi-frame sample video latent variable of the sample video;

[0045] Adding reference noise to the sample depth latent variables of multiple frames, inputting the sample depth latent variables after adding the reference noise and the sample video latent variables into the diffusion model, performing noise prediction with the sample video latent variables as conditional signals, and obtaining the sample noise of multiple frames;

[0046] A model loss is determined according to a difference between the sample noise and the reference noise, and the diffusion model is trained based on the model loss.

[0047] Furthermore, the adjustment module is also used to:

[0048] Based on the UNet network, the sample depth latent variables after adding the reference noise are mapped to perform noise prediction to obtain multi-frame sample noise;

[0049] In the mapping process of the UNet network, the sample video latent variable is used as a conditional signal and injected into the UNet network through a cross-attention mechanism.

[0050] On the other hand, an embodiment of the present disclosure further provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor implements the above-mentioned video processing method when executing the computer program.

[0051] On the other hand, an embodiment of the present disclosure further provides a computer-readable storage medium, wherein the storage medium stores a computer program, and the computer program is executed by a processor to implement the above-mentioned video processing method.

[0052] On the other hand, the embodiment of the present disclosure further provides a computer program product, which includes a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device executes the above-mentioned video processing method.

[0053] The disclosed embodiments include at least the following beneficial effects: by obtaining multi-frame noise latent variables with the same number of frames as the target video, performing noise prediction based on the multi-frame noise latent variables to obtain multi-frame predicted noise, removing the predicted noise from each frame noise latent variable respectively to obtain multi-frame original depth latent variables, decoding the multi-frame original depth latent variables to obtain the original depth map corresponding to each video frame in the target video, dividing the multi-frame predicted noise into multiple overlapping sub-sequences, being able to support longer video input while keeping the computing resources unchanged, thereby improving the temporal consistency between the depth maps subsequently generated, and then, in each sub-sequence, determining the geometric constraint loss based on the original depth map, based on The geometric constraint loss is used to adjust the prediction noise of the non-overlapping part. Since the geometric constraint loss is used to constrain the geometric consistency when the corresponding video frame is three-dimensionally projected based on the original depth map, it is equivalent to enhancing the geometric consistency between different subsequences and different prediction noises through the adjustment method of the sliding window, thereby improving the geometric consistency between subsequently generated depth maps. Based on this, the adjusted prediction noise is removed from the noise latent variables of each frame respectively, so as to obtain more accurate multi-frame target depth latent variables. When the multi-frame target depth latent variables are decoded to obtain the multi-frame target depth map corresponding to each video frame, the accuracy of video depth estimation can be effectively improved.

[0054] Other features and advantages of the present disclosure will be set forth in the following description, and in part will be apparent from the description, or may be learned by practicing the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] The accompanying drawings are used to provide further understanding of the technical solution of the present disclosure and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the technical solution of the present disclosure and do not constitute a limitation on the technical solution of the present disclosure.

[0056] Figure 1 A schematic diagram of an optional implementation environment provided for an embodiment of the present disclosure;

[0057] Figure 2An optional flowchart of the video processing method provided by the embodiment of the present disclosure;

[0058] Figure 3 An optional schematic diagram of dividing subsequences provided in an embodiment of the present disclosure;

[0059] Figure 4 An optional schematic diagram of adjusting prediction noise provided by an embodiment of the present disclosure;

[0060] Figure 5 An optional schematic diagram of injecting a conditional signal into a UNet network provided in an embodiment of the present disclosure;

[0061] Figure 6 An optional overall flow chart of a video processing method provided in an embodiment of the present disclosure;

[0062] Figure 7 A schematic diagram of the structure of a video processing device provided by an embodiment of the present disclosure;

[0063] Figure 8 A partial structural block diagram of a terminal provided in an embodiment of the present disclosure;

[0064] Fig. 9 A partial structural block diagram of a server provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0065] In order to make the purpose, technical solution and advantages of the present disclosure more clear, the present disclosure is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present disclosure and are not used to limit the present disclosure.

[0066] It should be noted that in various specific embodiments of the present disclosure, when it comes to the need to perform relevant processing based on data related to the characteristics of the target object such as the target object attribute information or attribute information set, the permission or consent of the target object will be obtained first, and the collection, use and processing of these data will comply with relevant laws, regulations and standards. Among them, the target object can be a user. In addition, when the embodiment of the present disclosure needs to obtain the attribute information of the target object, the separate permission or separate consent of the target object will be obtained through a pop-up window or by jumping to a confirmation page. After clearly obtaining the separate permission or separate consent of the target object, the necessary target object-related data used to enable the normal operation of the embodiment of the present disclosure will be obtained.

[0067] In the embodiments of the present disclosure, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.

[0068] To facilitate understanding of the technical solution provided by the embodiments of the present disclosure, some key terms used in the embodiments of the present disclosure are explained here:

[0069] Video depth estimation: It is a computer vision algorithm that can extract and estimate the depth information of each object in the scene from the video sequence by analyzing the temporal and spatial information in the video sequence.

[0070] Diffusion Model: A deep learning generative model based on random processes. It generates new data samples by simulating the diffusion process of data distribution. The idea is to model the data generation process as an inverse Markov chain. The diffusion process is divided into two stages: noise addition and data generation. The noise addition process simulates the process of data distribution gradually changing to Gaussian noise distribution, gradually adding noise to the original data until the original data becomes pure noise that does not contain any original data information. The data generation process is to reversely restore the original data from the noise state. It is mostly used in scenarios such as image generation and image restoration.

[0071] Three-dimensional projection: Through some mathematical transformation, the points in the three-dimensional space are projected onto a two-dimensional plane to obtain the representation of the three-dimensional space points on the two-dimensional plane.

[0072] Sliding window: It is an array or string-based algorithm technology that defines a fixed-size window on the data stream to achieve efficient data processing and data transmission.

[0073] Video depth estimation is a long-standing task in the field of computer vision. In related technologies, depth maps are usually generated using pre-trained deep learning models. These deep learning models rely on strong image space data priors to normalize depth predictions, so they can generate depth maps with rich depth details based on the zero-sample generalization ability of deep learning models. However, most of the deep learning models used rely on extracting data priors from two-dimensional image space. Although the video depth estimation methods in related technologies can maintain the temporal consistency between depth maps to a certain extent, these deep learning models lack the ability to perceive geometric information in three-dimensional structures, resulting in a lack of geometric consistency between the generated depth maps, which affects the accuracy of video depth estimation.

[0074] Based on this, the embodiments of the present disclosure provide a video processing method, device, electronic device and storage medium, which can better maintain the temporal consistency and geometric consistency between the generated depth maps in the application scenario of long videos, thereby improving the accuracy of video depth estimation.

[0075] Reference Figure 1 , Figure 1 A schematic diagram of an optional implementation environment provided for an embodiment of the present disclosure, the implementation environment includes a terminal 101 and a server 102, wherein the terminal 101 and the server 102 are connected via a communication network.

[0076] Exemplarily, a target video is obtained in the terminal 101, and the target video is sent to the server 102. A plurality of noise latent variables having the same number of frames as the target video are obtained in the server 102. Noise prediction is performed based on the multi-frame noise latent variables to obtain multi-frame predicted noise. The predicted noise is removed from the noise latent variables of each frame to obtain multi-frame original depth latent variables. The multi-frame original depth latent variables are decoded to obtain the original depth map corresponding to each video frame in the target video. The multi-frame predicted noise is divided into a plurality of overlapping sub-sequences. In each sub-sequence, a geometric constraint loss is determined based on the original depth map. The predicted noise of the non-overlapping part is adjusted based on the geometric constraint loss. Then, the adjusted predicted noise is removed from the noise latent variables of each frame to obtain multi-frame target depth latent variables. The multi-frame target depth latent variables are decoded to obtain multi-frame target depth maps corresponding to each video frame. The target depth map is sent to the terminal 101 so that the terminal 101 performs three-dimensional reconstruction based on the target depth map.

[0077] It is understandable that the video processing method provided in the embodiment of the present disclosure may also be executed in the terminal 101 alone.

[0078] Server 102 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. In addition, server 102 can also be a node server in a blockchain network.

[0079] The terminal 101 may be a mobile phone, a computer, an intelligent voice interaction device, an intelligent wearable device, an intelligent home appliance, a vehicle-mounted terminal, etc., but is not limited thereto. The terminal 101 and the server 102 may be directly or indirectly connected via wired or wireless communication, which is not limited in the embodiments of the present disclosure.

[0080] Reference Figure 2 , Figure 2 An optional flowchart of a video processing method provided in an embodiment of the present disclosure. The video processing method can be executed by a terminal, or by a server, or by a terminal and a server in cooperation. The video processing method includes but is not limited to the following steps S201 to S204.

[0081] Step S201: obtaining multi-frame noise latent variables with the same number of frames as the target video, and performing noise prediction based on the multi-frame noise latent variables to obtain multi-frame predicted noise.

[0082] Among them, the target video is the video for which a depth map needs to be generated, which is obtained through devices such as cameras or sensors, including multiple time-continuous video frames, and is a sequence of video frames containing depth information; the noise latent variable is a depth latent variable with noise added, and the depth latent variable is a coding feature obtained after encoding the target video into the latent space. The noise latent variable can be obtained by adding random Gaussian noise or by artificial noise addition. When artificial noise addition is performed, it can be randomly added noise or noise added according to the characteristic distribution of the latent variable; the process of noise prediction is based on a diffusion model, and the diffusion model can be trained based on a sample depth map. When the diffusion model performs noise prediction, the target video can be diffused as a conditional signal.

[0083] Specifically, the target video is acquired through the camera, and the feature dimension of the target video can be T is the number of video frames in the target video, H×W is the video frame size, and C is the number of channels, which can be 3 at this time. The target video is input into the encoder for video frame encoding. The encoder can be a VAE encoding network. The encoder encodes each video frame containing depth information into the latent space to obtain the depth latent variable corresponding to each video frame. The depth latent variable is denoised to obtain a noise latent variable, and the noise latent variable is input into the diffusion model for noise prediction at each time step to obtain the predicted noise corresponding to each video frame in the current time step. By encoding the video frames in the target video from the high-dimensional pixel space to the low-dimensional latent space, the video frame data is effectively compressed while retaining the original key information well, so that representative latent variables can be extracted, which helps to improve the accuracy of the generated depth map. In addition, since the dimension of the latent variables in the latent space is usually lower than the dimension of the features in the pixel space, adding noise, noise prediction and other operations to the video frames in the latent space can effectively reduce the amount of calculation and help to improve the speed of video depth estimation. In addition, the diffusion model can capture the complex potential relationships and structures between video frames in the latent space, so that the diffusion model can have better generalization ability when facing videos of different lengths and contents.

[0084] Step S202: respectively remove the predicted noise from the noise latent variables of each frame to obtain multiple frames of original depth latent variables, decode the multiple frames of original depth latent variables, and obtain the original depth map corresponding to each video frame in the target video.

[0085] Among them, the original depth latent variable is a feature obtained by removing the noise added in the noise latent variable by predicting the noise within a preset time step. It can be understood that the depth latent variable obtained after denoising in the last time step is a depth latent variable without noise, and the depth latent variable obtained after denoising in other time steps except the last time step is a depth latent variable containing noise; the original depth map is used to describe the position information of the pixel points in the video frame, so as to provide the three-dimensional geometric information of the video frame. The original depth map contains the distance between each pixel point in the video frame and the camera. Based on this distance, the three-dimensional geometric information of the scene in the video frame or the relative depth information between objects in the video frame can be provided.

[0086] Specifically, in the diffusion model, the noise latent variables of each video frame are denoised based on the predicted noise of the current time step, and the original noise latent variables with the predicted noise of the current time step removed are obtained. The noise latent variables with the predicted noise of the current time step removed are input into the decoder for decoding, and the original depth map corresponding to each video frame in the target video in the current time step is obtained. By decoding the original depth latent variables corresponding to each time step to obtain the original depth map, the spatial depth information of each time step can be understood, and the change of the depth map over time is analyzed based on the spatial depth information of each time step, so that the diffusion model can learn the relationship between the changing posture and spatial structure of the objects in the scene in the video frame, which helps the diffusion model generate more reliable predicted noise.

[0087] Step S203: Divide the predicted noise into a plurality of overlapping subsequences, determine the geometric constraint loss in each subsequence based on the original depth map, and adjust the predicted noise of the non-overlapping part based on the geometric constraint loss.

[0088] Among them, the subsequence is a feature sequence obtained by dividing the prediction noise corresponding to all video frames in the target video according to a fixed batch size. Except for the first subsequence, the remaining subsequences include the prediction noise of the overlapping part and the prediction noise of the non-overlapping part. The prediction noise of the overlapping part and the non-overlapping part is determined according to the previous subsequence, so the first subsequence can be considered to include only the prediction noise of the non-overlapping part; the geometric constraint loss is used to constrain the geometric consistency when the corresponding video frame is three-dimensionally projected based on the original depth map. The geometric consistency includes global consistency and local consistency; the adjustment operation is based on the geometric constraint loss, which is used to update and fine-tune the prediction noise of the non-overlapping part.

[0089] In a possible implementation, in the process of dividing the predicted noise into a plurality of overlapping subsequences, determining the geometric constraint loss in each subsequence based on the original depth map, and adjusting the predicted noise of the non-overlapping part based on the geometric constraint loss, specifically, based on the sliding window strategy, the predicted noise corresponding to all video frames in the target video is divided into a plurality of subsequences according to a fixed sliding window size and a certain sliding step size, for example, referring to Figure 3 , Figure 3 An optional schematic diagram of dividing subsequences provided in an embodiment of the present disclosure, Figure 3The target video in includes 10 video frames, and the corresponding 10 prediction noises are obtained based on these 10 video frames. According to the setting of sliding window size of 4 and sliding step size of 2, subsequence division is performed to obtain subsequence A, subsequence B, subsequence C, and subsequence D. Subsequence A includes prediction noise 1, prediction noise 2, prediction noise 3, and prediction noise 4; subsequence B includes prediction noise 3, prediction noise 4, prediction noise 5, and prediction noise 6; subsequence C includes prediction noise 5, prediction noise 6, prediction noise 7, and prediction noise 8; subsequence D includes prediction noise 7, prediction noise 8, prediction noise 9, and prediction noise 10. For subsequence B, the prediction noise of the overlapping part of subsequence B is prediction noise 3 and prediction noise 4, and the prediction noise of the non-overlapping part is prediction noise 5 and prediction noise 6.

[0090] Next, the geometric constraint loss is determined based on the original depth map in each subsequence in turn, the corresponding gradient is determined based on the geometric constraint loss, the network parameters of the decoder, diffusion model, and encoder are updated according to the gradient, and the prediction noise of the non-overlapping part of the subsequence is adjusted. By dividing the prediction noise into multiple overlapping subsequences, it is possible to support longer video inputs while keeping the computing resources unchanged, which helps to improve the temporal consistency between the generated depth maps. In addition, since the sliding window size and sliding step size can be set by yourself, the method of subsequence division based on the sliding window strategy can adjust the amount of prediction noise in the subsequence and the ratio of the prediction noise in the overlapping part and the non-overlapping part according to the actual computing resources and depth map generation effect, making the division of subsequences more flexible and effectively improving the adjustment effect of the prediction noise.

[0091] In a possible implementation, in the process of determining the geometric constraint based on the original depth map, specifically, a first constraint loss is determined based on the original depth map, a second constraint loss is determined based on the original depth map, and the first constraint loss and the second constraint loss are weighted summed to obtain the geometric constraint loss. The first constraint loss is used to constrain the scene consistency when the corresponding video frame is three-dimensionally projected based on the original depth map, that is, to globally constrain the image data in the video frame (reconstruction constraint), and the second constraint loss is used to constrain the detail consistency when the corresponding video frame is three-dimensionally projected based on the original depth map, that is, to locally constrain the image data in the video frame.

[0092] Specifically, two video frames are selected from the target video, and the two video frames may be adjacent or non-adjacent. Based on the original depth map, the projection relationship between the two video frames, the projection distance of the pixels of the two video frames, the projection depth of the pixels of the two video frames, the surface normal of any video frame, the depth gradient of any video frame and other information are obtained. The first constraint loss is determined based on the projection relationship between the two video frames, the projection distance of the pixels of the two video frames, and the projection depth of the pixels of the two video frames. The second constraint loss is determined based on the surface normal of any video frame and the depth gradient of any video frame. Weights are configured for the first constraint loss and the second constraint loss respectively, and the first constraint loss and the second constraint loss are weighted and summed to obtain the geometric constraint loss. The scene consistency and detail consistency are constrained by the geometric constraint loss, so that the prediction noise can be adjusted from both global and local aspects, and the generated depth map can be prevented from being inconsistent with the overall scene in the local area, such as dislocation, distortion, etc., thereby improving the accuracy of video depth estimation.

[0093] In addition, when the scale of the scene in the video frame is large, the main focus is on the overall structure of the scene. The first constraint loss can be weighted summed to obtain the geometric constraint loss. The geometric constraint loss is used to constrain the scene consistency, which helps to ensure that the generated depth map has the correct position and orientation, thereby improving the accuracy and coordination of depth map generation, and thus improving the accuracy of video depth estimation. When the scale of the scene in the video frame is small, or only a specific area or specific object in the video frame needs to be focused on, the second constraint loss can be weighted summed to obtain the geometric constraint loss. The geometric constraint loss is used to constrain the consistency of details, and the edge contour of the object can be accurately captured based on the surface normal of the object, thereby reducing the influence of factors such as background, improving the accuracy and reliability of depth map generation, and thus improving the accuracy of video depth estimation.

[0094] In one possible implementation, in the process of determining the second constrained loss based on the original depth map, specifically, a first surface normal is calculated based on the original depth map, a second surface normal of the corresponding video frame is generated based on a surface normal generation network, a surface normal loss is determined according to the difference between the first surface normal and the second surface normal, a depth gradient of the original depth map is determined, the depth gradient is regularized to obtain a smoothing loss, and a weighted sum is performed on the surface normal loss and the smoothing loss to obtain the second constrained loss. Among them, the surface normal generation network is used to predict and enhance the surface normal of the video frame to generate the corresponding surface normal, and the surface normal generation network can be a StableNormal network; the first surface normal is calculated based on the original depth map, which can be regarded as the true value of the surface normal, and the second surface normal is predicted based on the surface normal generation network, which can be regarded as the predicted value of the surface normal; the surface normal loss is used to evaluate the degree of alignment between the second surface normal of the video frame and the first surface normal of the original depth map corresponding to the video frame, which can be determined based on the angle between the second surface normal and the first surface normal; the smoothness loss is used to evaluate the continuity and smoothness of the original depth map, which can be determined based on the derivatives of the original depth map in different directions; the regularization operation can be L1 regularization (also known as L1 penalty), which is used to perform local smoothness constraints on the generated depth map.

[0095] Specifically, a first surface normal is calculated based on the original depth map of the i-th video frame, and the i-th video frame is input into the surface normal generation network for surface normal prediction to generate a second surface normal of the i-th video frame. The arc cosine value of the angle between the first surface normal and the second surface normal is calculated, and the difference between the first surface normal and the second surface normal is represented based on the calculated arc cosine value, thereby determining the surface normal loss. Surface normal loss It can be expressed by the following formula:

[0096]

[0097] Among them, n i represents the first surface normal, Represents the second surface normal, and arccos(·) represents the arccosine function. Next, the original depth map of the i-th frame of the video is differentiated in the horizontal and vertical directions to obtain the horizontal depth gradient in the horizontal direction and the vertical depth gradient in the vertical direction. The i-th frame of the video is then differentiated in the horizontal and vertical directions to obtain the horizontal image gradient in the horizontal direction and the vertical image gradient in the vertical direction. Perform natural exponential operations on the horizontal image gradient and the vertical image gradient respectively to obtain the weights corresponding to the horizontal depth gradient and the vertical depth gradient. The horizontal depth gradient and the vertical depth gradient are weightedly summed based on the weights to obtain the smoothing loss. Smoothing loss It can be expressed by the following formula:

[0098]

[0099] Among them, d i represents the original depth map, I i represents the i-th video frame, represents the horizontal depth gradient, represents the vertical depth gradient, represents the horizontal image gradient, represents the vertical image gradient, is the weight of the horizontal depth gradient, is the weight of the vertical depth gradient.

[0100] Finally, the surface normal weight is configured for the surface normal loss, and the smoothing weight is configured for the smoothing loss. The surface normal loss and the smoothing loss are weighted summed based on the surface normal weight and the smoothing weight to obtain the second constraint loss. The second constraint weight L 2 It can be expressed by the following formula:

[0101] L 2 =a n L n +a s L s ,

[0102] Among them, L n represents the surface normal loss, a n represents the surface normal weight, L s represents the smoothing loss, a s Represents the smoothness weight. By using the surface normal loss and the smoothness loss as the second constraint loss, the consistency between the predicted surface normal and the surface normal of the video frame itself can be optimized and the relative position of the object in the video frame can be maintained, thereby improving the accuracy of the local area constrained by the second constraint loss.

[0103] In addition, you can also select only surface normal loss or smoothness loss as the second constraint loss according to the video frame content of the target video and the generation requirements of the depth map. When only surface normal loss is selected as the second constraint loss, you can focus on the alignment between the predicted surface normal and the surface normal of the video frame itself, and perform detail consistency constraints based on this, so that the generated depth map can more accurately display the surface shape and structure of the object in the video frame; when only smoothness loss is selected as the second constraint loss, you can focus on the optimization direction of the original depth map, and optimize the local area based on the depth gradient of the original depth map, which can avoid visual misalignment or errors in the generated depth map, thereby improving the accuracy of video depth estimation.

[0104] Step S204: remove the adjusted prediction noise from each frame noise latent variable respectively to obtain multi-frame target depth latent variables, decode the multi-frame target depth latent variables, and obtain a multi-frame target depth map corresponding to each video frame.

[0105] According to step S201 to step S203, obtain multi-frame noise latent variables with the same number of frames as the target video, perform noise prediction on the multi-frame noise latent variables based on the diffusion model, obtain multi-frame prediction noise at the current time step, remove the prediction noise from each frame noise latent variable, obtain multi-frame original depth latent variables at the current time step, decode the multi-frame original depth latent variables at the current time step, and obtain the original depth map corresponding to each video frame in the target video. Divide the multi-frame prediction noise into multiple overlapping subsequences, determine the geometric constraint loss based on the original depth map in each subsequence in turn, and adjust the prediction noise of the non-overlapping part based on the geometric constraint loss. Then, remove the adjusted prediction noise from each frame noise latent variable time step by time step until the target depth latent variable of the last time step is obtained, decode the target depth latent variable, and obtain the multi-frame target depth map corresponding to the target video. Among them, the target depth latent variable is the depth latent variable without noise obtained by removing the prediction noise by predicting the noise within the preset time step, and the target depth map is the depth map finally output, and tasks such as three-dimensional reconstruction can be performed based on the target depth map.

[0106] Specifically, in the diffusion model, the noise latent variable is denoised time step by time based on the adjusted prediction noise until the target depth latent variable of the last time step is obtained, the target depth latent variable is input into the decoder for decoding, and the target depth latent variable is decoded from the latent space to the image space to obtain the target depth map corresponding to each video frame.

[0107] In a possible implementation, in the process of determining the first loss based on the original depth map, specifically, the three-dimensional projection relationship between the i-th frame video frame and the j-th frame video frame can be determined based on the i-th frame original depth map, and the j-th frame video frame is three-dimensionally projected relative to the i-th frame video frame based on the three-dimensional projection relationship to obtain a first projection frame, and the reprojection loss is determined according to the difference between the first projection frame and the i-th frame video frame, and the projection depth map when the j-th frame video frame is three-dimensionally projected is determined, and the depth loss is determined according to the difference between the projection depth map and the original depth map, and the reprojection loss and the depth loss are weighted summed to obtain the first constrained loss. Among them, i and j are not equal to each other and are positive integers. The three-dimensional projection relationship involves multiple projection parameters for projecting pixels in the i-th video frame to the j-th video frame, including camera intrinsic parameters, posture transformation parameters, projection depth, etc.; the first projection frame is a projection frame obtained by projecting pixels in the i-th video frame to the j-th video frame, and the first projection frame includes pixels of the i-th video frame projected into the j-th video frame; the reprojection loss is used to evaluate the overall difference between the video frames before and after projection, and the depth loss is used to evaluate the difference between the projection depths before and after projection of the video frames.

[0108] Specifically, obtain the i-th video frame, camera internal parameters, and posture transformation parameters, project all pixel points in the i-th video frame to three-dimensional space, and obtain the first projection point corresponding to the i-th video frame and the projection depth generated when projecting to the three-dimensional space, wherein the first projection point is a point in the three-dimensional space. Perform coordinate transformation on the first projection point corresponding to the i-th video frame, that is, transform the coordinate system of the first projection point into the coordinate system corresponding to the j-th video frame, and then project the first projection point after coordinate transformation into the j-th video frame to obtain the second projection point, which is a point on the plane image (j-th video frame). At this time, the j-th video frame contains the original pixel points of the j-th video frame and the second projection point. The three-dimensional projection relationship between the i-th video frame and the j-th video frame (that is, obtaining the second projection point) can be expressed by the following formula,

[0109] p j ~KT i→j d i (p i )K -1 p i ,

[0110] Among them, p i represents the first projection point, p j represents the second projection point, d i (p i ) represents the projection depth generated when the i-th video frame is projected into the three-dimensional space, K represents the camera internal parameter, T i→j Represents the posture transformation parameters.

[0111] Further, based on the three-dimensional projection relationship, the j-th video frame is three-dimensionally projected relative to the i-th video frame to obtain a first projection frame, and the reprojection loss is determined according to the difference between the first projection frame and the i-th video frame. A projection depth map when the j-th video frame is three-dimensionally projected is obtained, and a depth loss is determined according to the difference between the projection depth map and the original depth map. The depth loss can be an absolute error loss (Mean Absolute Error, MAE). According to the above description, the depth loss It can be expressed by the following formula:

[0112]

[0113] Among them, d i is the original depth map, d j→i is the projected depth map.

[0114] Finally, configure the reprojection weight for the reprojection loss and the depth weight for the depth loss. Based on the reprojection weight and the depth weight, the reprojection loss and the depth loss are weighted and summed to obtain the first constraint loss. 1 It can be expressed by the following formula:

[0115] L 1 =a r L r +a d L d

[0116] Among them, L r represents the reprojection loss, a r represents the reprojection weight, L d represents the depth loss, a d Represents the depth weight. By taking the reprojection loss and the depth loss as the first constraint loss, the prediction noise can be constrained for scene consistency from two aspects: projection difference and depth difference. This can locate the noise introduced in the projection process and quantify the noise level in the depth map, providing a clear basis for adjusting the prediction noise.

[0117] In addition, you can also select only reprojection loss or depth loss as the first constraint loss according to the video frame content of the target video and the generation requirements of the depth map. When only reprojection loss is selected as the first constraint loss, you can focus on the inconsistency between the first projection point when projected to the jth video frame and the i-th video frame, and perform geometric constraints based on the inconsistency, thereby improving the accuracy of the projection; when only depth loss is selected as the first constraint loss, you can focus on the difference between the original depth map and the projected depth map, and perform geometric constraints based on the difference, thereby improving the accuracy of the original depth map generated in subsequent time steps.

[0118] Before determining the three-dimensional projection relationship between the i-th video frame and the j-th video frame based on the i-th original depth map, firstly, the i-th video frame and the j-th video frame must be randomly selected from the target video. In order to avoid the i-th video frame and the j-th video frame having too large an inter-frame interval between the i-th video frame and the j-th video frame, resulting in a significant difference in the content between the two frames, which results in a low credibility of the obtained reprojection loss and depth loss, it is necessary to control the inter-frame interval difference between the i-th video frame and the j-th video frame within a certain range. Specifically, an inter-frame interval threshold can be preset, and the i-th video frame and the j-th video frame are randomly selected based on the time interval threshold. When making the selection, only the inter-frame interval between the two frames is considered, and the time order of the two frames does not need to be considered. For example, the inter-frame interval threshold is 2, and the target video includes a first video frame, a second video frame, a third video frame, a fourth video frame, and a fifth video frame, and the above video frames are arranged in the order of shooting time. When the i-th video frame is the first video frame, the j-th video frame may be the second video frame or the third video frame; when the i-th video frame is the third video frame, the j-th video frame may be the first video frame, the second video frame, the fourth video frame or the fifth video frame.

[0119] In one possible implementation, in the process of performing weighted summation on the reprojection loss and the depth loss to obtain the first constrained loss, specifically, the i-th frame prediction noise and the j-th frame prediction noise are taken as prediction noise pairs, and the reprojection losses corresponding to all prediction noise pairs in the subsequence are determined to obtain a first mean, as well as a second mean of the depth loss corresponding to all prediction noise pairs in the subsequence, and the first mean and the second mean are weightedly summed to obtain the first constrained loss.

[0120] Specifically, for any subsequence, a subsequence includes at least two frames of prediction noise, and the at least two frames of prediction noise include the i-th frame prediction noise and the j-th frame prediction noise. Randomly take out the i-th frame prediction noise and the j-th frame prediction noise from the subsequence, take the i-th frame prediction noise and the j-th frame prediction noise as a prediction noise pair, and obtain the reprojection loss and depth loss corresponding to the prediction noise pair based on the prediction noise pair. Then, randomly pair the other prediction noises in the subsequence except the i-th frame prediction noise and the j-th frame prediction noise, and determine the reprojection loss and depth loss corresponding to each prediction noise pair. Calculate the mean of the reprojection loss corresponding to all prediction noise pairs to obtain a first mean, and calculate the mean of the depth loss corresponding to all prediction noise pairs to obtain a second mean. Perform weighted summation on the first mean and the second mean to obtain a first constrained loss, and adjust all prediction noises in the subsequence based on the first constrained loss. For example, the current subsequence includes the first frame prediction noise, the second frame prediction noise, the third frame prediction noise and the fourth frame prediction noise. The first frame prediction noise and the third frame prediction noise are used as the first prediction noise pair. The first reprojection loss r is obtained based on the first prediction noise pair. 1and the first depth loss d 1 ; The second frame prediction noise and the fourth frame prediction noise are used as the second prediction noise pair, and the second reprojection loss r is obtained based on the second prediction noise pair 2 and the second depth loss d 2 . Calculate the first reprojection loss r 1 and the second reprojection loss r 2 The first mean (r 1 +r 2 / 2), calculate the first depth loss d 1 and the second depth loss d 2 The second mean (d 1 +d 2 / 2). Configure the reprojection weight a for the first mean r , configure the depth weight a for the second mean d , based on the reprojection weight a r , depth weight a d The first mean (r 1 +r 2 / 2), the second mean (d 1 +d 2 / 2) weighted summation is performed, and the first constraint loss is obtained as a r (r 1 +r 2 / 2)+a d (d 1 +d 2 / 2). The first constraint loss is determined by averaging the reprojection loss and depth loss of multiple prediction noise pairs in the subsequence, which can reduce the impact of extreme values ​​on the first constraint loss while smoothing multiple loss values, thereby improving the accuracy of adjusting the prediction noise based on the first constraint loss.

[0121] It should be noted that the prediction noises corresponding to the i-th video frame and the j-th video frame mentioned above can be in the same subsequence or in different subsequences. The combination method of the prediction noise pairs is the same as the method of selecting the video frame pairs. The inter-frame interval threshold can be preset to ensure that the inter-frame interval between the prediction noise pairs is kept within a certain range, and the combination of the prediction noise pairs is only related to the inter-frame interval. Taking the prediction noises corresponding to the i-th video frame and the j-th video frame in the same subsequence as an example, the preset inter-frame interval threshold is 3, and the current subsequence includes the first frame prediction noise, the second frame prediction noise, the third frame prediction noise, the fourth frame prediction noise, the fifth frame prediction noise, and the sixth frame prediction noise. When the i-th frame prediction noise is the first frame prediction noise, the j-th frame prediction noise can select the second frame prediction noise, the third frame prediction noise or the fourth frame prediction noise; when the i-th frame prediction noise is the second frame prediction noise, the j-th frame prediction noise can select the first frame prediction noise, the third frame prediction noise, the fourth frame prediction noise or the fifth frame prediction noise. According to the preset inter-frame interval threshold 3, the prediction noise pairs in the current subsequence can be the first frame prediction noise and the third frame prediction noise, the second frame prediction noise and the fifth frame prediction noise, and the third frame prediction noise and the sixth frame prediction noise. By limiting the inter-frame interval between the prediction noise pairs, each pair of prediction noise pairs can maintain a strong and consistent correlation, reducing the uncertainty factors that may be introduced due to a large inter-frame interval, thereby improving the reliability of the first constraint loss.

[0122] In a possible implementation, in the process of performing weighted summation on the reprojection loss and the depth loss to obtain the first constrained loss, the i-th frame prediction noise and the j-th frame prediction noise in the subsequence can also be used as a prediction noise pair, and the reprojection loss and the depth loss corresponding to the prediction noise pair are determined, and the reprojection loss and the depth loss corresponding to the prediction noise pair are weighted summed to obtain the first constrained loss corresponding to the prediction noise pair, and the i-th frame prediction noise and the j-th frame prediction noise are adjusted based on the first constrained loss. Then, the reprojection loss and the depth loss corresponding to the next pair of prediction noise pairs are determined until all prediction noise pairs in the subsequence are adjusted based on their corresponding first constrained losses. By calculating the first constrained loss for the prediction noise pairs respectively, the prediction noise pairs can be adjusted directly based on the corresponding first constrained loss, avoiding the influence of the first constrained losses corresponding to other prediction noise pairs, thereby improving the accuracy of the adjustment process.

[0123] In a possible implementation, in the process of performing weighted summation of reprojection loss and depth loss to obtain the first constrained loss, specifically, based on the original depth maps corresponding to two adjacent video frames, three-dimensional projection is performed on the two adjacent video frames to obtain a first projection point set and a second projection point set, the coordinate system of the second projection point set is converted to the coordinate system of the first projection point set, the tracking loss is determined according to the difference between the first projection point set and the converted second projection point set, and the reprojection loss, depth loss and tracking loss are weighted summed to obtain the first constrained loss. The projection points in the first projection point set and the second projection point set are all points in three-dimensional space; the tracking loss is used to evaluate the similarity between two adjacent video frames.

[0124] Specifically, the original depth maps corresponding to the adjacent first video frame and the second video frame are obtained, and the first video frame and the second video frame are any two adjacent frames in the video frames corresponding to the predicted noise in a subsequence. The first projection depth is obtained from the original depth map of the first video frame, and the second projection depth is obtained from the original depth map of the second video frame. Based on the first projection depth, all pixels in the first video frame are three-dimensionally projected to obtain a first projection point set. Based on the second projection depth, all pixels in the second video frame are three-dimensionally projected to obtain a second projection point set. Based on the camera extrinsic parameters, the coordinate system of the second projection point set is converted to the coordinate system of the first projection point set, and the tracking loss is determined according to the difference between the first projection point set and the converted second projection point set. The tracking loss can be a mean square error loss (MSE). According to the above description, the tracking loss It can be expressed by the following formula:

[0125]

[0126] Among them, P i is the first projection point set, P j→i is the second projection point set after coordinate system transformation.

[0127] Next, configure the reprojection weight for the reprojection loss, configure the depth weight for the depth loss, and configure the tracking weight for the tracking loss. Based on the reprojection weight, depth weight, and tracking weight, the reprojection loss, depth loss, and tracking loss are weighted and summed to obtain the first constraint loss. The first constraint loss L 1 It can be expressed by the following formula:

[0128] L 1 =a r L r +a d L d +a t L t

[0129] Among them, L r represents the reprojection loss, a r represents the reprojection weight, L d represents the depth loss, a d represents the depth weight, L t represents the tracking loss, a t Represents the tracking weight. By introducing the tracking loss into the first constraint loss, the dynamic changes between two adjacent video frames can be evaluated, including position changes, scale changes, etc., so as to reduce the error accumulation in the noise prediction process and further improve the reliability of the first constraint loss.

[0130] In one possible implementation, in the process of determining the reprojection loss based on the difference between the first projection frame and the i-th video frame, the structural similarity between the first projection frame and the i-th video frame can be specifically determined, and the difference and structural similarity between the first projection frame and the i-th video frame are weightedly summed to obtain the reprojection loss.

[0131] Specifically, the structural similarity between the first projection frame and the i-th video frame, and the absolute error between the first projection frame and the i-th video frame are determined. A first weight is configured for the structural similarity, and a second weight is configured for the absolute error. Based on the first weight and the second weight, the structural similarity and the absolute error between the first projection frame and the i-th video frame are weighted and summed to obtain the reprojection loss. According to the above description, the reprojection loss Based on structural similarity loss and absolute error loss,

[0132]

[0133] Among them, I i Represents the i-th video frame, I i→j represents the first projection frame, SSIM(·) represents the structural similarity loss, and the structural similarity loss can be evaluated from the aspects of brightness similarity, contrast similarity, pixel structure similarity, etc. between the i-th video frame and the first projection frame; a is a hyperparameter, and its value can be 0.85. By calculating the structural similarity between the i-th video frame and the first projection frame, the similarity between the i-th video frame and the first projection frame is evaluated from the three aspects of brightness, contrast and structure, which can more comprehensively reflect the overall difference between the i-th video frame and the first projection frame. On this basis, combined with the absolute error, the pixel-level difference between the i-th video frame and the first projection frame can be further captured, thereby improving the reliability of the reprojection loss.

[0134] In one possible implementation, in the process of performing weighted summation on the surface normal loss and the smoothness loss to obtain the second constrained loss, specifically, a third mean of the surface normal loss corresponding to all prediction noises in the subsequence and a fourth mean of the smoothness loss corresponding to all prediction noises in the subsequence can be determined, and the third mean and the fourth mean are weightedly summed to obtain the second constrained loss.

[0135] Specifically, based on any prediction noise, the surface normal loss and smoothing loss corresponding to the prediction noise can be obtained. The surface normal loss and smoothing loss corresponding to all prediction noises in the subsequence are obtained, the mean of the surface normal losses corresponding to all prediction noises is calculated to obtain a third mean, and the mean of the smoothing losses corresponding to all prediction noises is calculated to obtain a fourth mean. The first mean and the second mean are weighted and summed based on the surface normal weight and the smoothing weight to obtain a second constrained loss, and all prediction noises in the subsequence are adjusted based on the second constrained loss. According to the above description, the second constrained loss L 2 It can be expressed by the following formula:

[0136]

[0137] Among them, L n represents the surface normal loss, a n represents the surface normal weight, L s represents the smoothing loss, a s represents the smoothing weight, M represents the number of prediction noises in the subsequence, and M>1. The second constraint loss is determined by averaging the surface normal loss and smoothing loss of the prediction noise in the subsequence, which can intuitively grasp the overall level of the loss value and reduce the impact of extreme values ​​on the second constraint loss, thereby improving the reliability of the second constraint loss.

[0138] In a possible implementation, in the process of performing weighted summation on the surface normal loss and the smoothing loss to obtain the second constrained loss, the corresponding surface normal loss and smoothing loss may also be calculated for all prediction noises in the subsequence in sequence. For any prediction noise, the surface normal loss and smoothing loss corresponding to the prediction noise are determined, and the surface normal loss and smoothing loss corresponding to the prediction noise are weighted summed to obtain the second constrained loss corresponding to the prediction noise, and the prediction noise is adjusted based on the second constrained loss. Then, the surface normal loss and smoothing loss corresponding to the next prediction noise are determined until all prediction noises in the subsequence are adjusted based on their corresponding second constrained losses. By calculating the second constrained loss for the prediction noises respectively, the prediction noise can be adjusted directly based on the corresponding second constrained loss, thereby avoiding the influence of the second constrained losses corresponding to other prediction noises, thereby improving the accuracy of the adjustment process.

[0139] In a possible implementation, the diffusion model performs noise prediction based on multiple time steps. In the process of adjusting the prediction noise of the non-overlapping part based on the geometric constraint loss, the gradient of the geometric constraint loss can be determined, and the adjustment term is obtained according to the product of the gradient and the guidance strength. The adjusted prediction noise is obtained according to the sum of the adjustment term and the prediction noise of the non-overlapping part. The guidance strength is used to control the guidance degree of each time step.

[0140] Specifically, the geometric constraint loss can be composed of reprojection loss, depth loss, tracking loss, surface normal loss, and smoothness loss. The geometric constraint loss L can be expressed by the following formula:

[0141] L = a r L r +a d L d +a t L t +a n L n +a s L s ,

[0142] Determine the gradient corresponding to the current subsequence based on the geometric constraint loss, and get the adjustment term based on the product of the gradient and the guidance strength. Get the prediction noise of the non-overlapping part of the current subsequence in the current sliding window, add the prediction noise of the non-overlapping part to the adjustment term, and get the adjusted prediction noise. Adjusted prediction noise It can be calculated by the following formula:

[0143]

[0144] Where t represents the time step, z t represents the noise latent variable at time step t, ∈ θ (a t ,t) represents the noise latent variable z at time step t t The predicted noise of the non-overlapping part is predicted, s(t) represents the guidance strength at time step t, and c represents the conditional signal; Represents the predicted pure sample, which is obtained by denoising the predicted noise latent variable and is used to guide the adjustment of the predicted noise in each time step; f(·) represents the guidance function, which can be trained based on the pure sample (noise-free sample); L(·) represents the geometric constraint loss, which is used to measure the compatibility between the conditional signal and the predicted pure sample. For the predicted pure sample, the predicted pure sample It can be expressed by the following formula:

[0145]

[0146] Among them, z t represents the noise latent variable at time step t, a c Represents the noise scale. The prediction noise is adjusted through the gradient of the geometric constraint loss and the guidance strength, so that the diffusion model can focus on the areas or pixels that have a significant impact on the geometric characteristics, thereby achieving accurate adjustment of the prediction noise.

[0147] In a possible implementation, in the process of adjusting the prediction noise of the non-overlapping part based on the geometric constraint loss, specifically, for the divided multiple subsequences, all the prediction noise in the first subsequence is adjusted based on the gradient, and the non-overlapping parts in the subsequences other than the first subsequence are adjusted based on the gradient. First, all the prediction noise in the first subsequence is aligned using the first constraint loss to obtain the adjusted first subsequence. Then, referring to Figure 4 , Figure 4 An optional schematic diagram of adjusting the prediction noise provided by an embodiment of the present disclosure, subsequence A is the first subsequence, and subsequence B is the second subsequence. In subsequence B, prediction noise 4, prediction noise 5 and prediction noise 6 overlap with subsequence A, and prediction noise 7, prediction noise 8 and prediction noise 9 do not overlap with subsequence A. Since the prediction noise of the overlapping part (prediction noise 4, prediction noise 5, prediction noise 6) is separated from the prediction noise of the non-overlapping part (prediction noise 7, prediction noise 8, prediction noise 9), the first constraint loss is used for global constraint, and the prediction noise of the non-overlapping part and the prediction noise of the overlapping part in subsequence B are aligned, so that the prediction noise 4, prediction noise 5 and prediction noise 6 of the adjusted subsequence A can be transferred to subsequence B through the sliding window. Then, the prediction noise of the overlapping part is fixed, and the prediction noise of the non-overlapping part is updated by the local constraint using the second constraint loss, so as to obtain the adjusted subsequence B. The remaining subsequences are adjusted in sequence according to the division order until all subsequences are adjusted. By transferring the prediction noise between subsequences through a sliding window, the transferability and continuity of the prediction noise in each subsequence can be guaranteed. Furthermore, combining the sliding window with the geometric constraint loss can enhance the geometric consistency between different subsequences and between different prediction noises, thereby improving the geometric consistency between subsequently generated depth maps.

[0148] In a possible implementation, the predicted noise is obtained by prediction through a diffusion model, and the training process of the diffusion model can be specifically as follows: obtaining a multi-frame sample depth map corresponding to a sample video, extracting a sample depth latent variable of the sample depth map and a multi-frame sample video latent variable of the sample video, adding reference noise to the multi-frame sample depth latent variable, inputting the sample depth latent variable and the sample video latent variable after adding the reference noise into the diffusion model, using the sample video latent variable as a conditional signal to perform noise prediction, obtaining a multi-frame sample noise, determining a model loss based on the difference between the sample noise and the reference noise, and training the diffusion model based on the model loss. The sample video includes a multi-frame sample video frame, each sample video frame contains depth information, that is, the sample video can be regarded as a depth sequence; the model loss can be a geometric constraint loss, including a remapping loss, a depth loss, a tracking loss, a surface normal loss, and a smoothing loss, or an absolute error loss, or a mean square error loss.

[0149] Specifically, a sample video is obtained based on a camera, a sample depth map corresponding to each sample video frame is extracted from the sample video, the sample video is input into an encoder for sample video frame encoding, each sample video frame containing depth information is encoded into a latent space by the encoder, and at the same time, the sample depth map is normalized into a sample depth map with affine invariance, the sample depth map with affine invariance is copied three times, and the three sample depth maps with affine invariance are encoded into a latent space by the encoder to obtain sample video latent variables corresponding to the sample video frame of the sample video and sample depth latent variables of the sample depth map. Then, reference noise is added to each sample depth latent variable, and the sample depth latent variables and sample video latent variables after adding noise are input into a diffusion model for denoising. In the denoising process, the sample video latent variables are used as conditional signals for noise prediction to obtain sample noise corresponding to each video frame. The model loss is determined according to the difference between the sample noise and the reference noise, and the diffusion model is trained based on the model loss. The diffusion model is trained based on sample depth maps and sample video frames, so that the diffusion model can learn accurate three-dimensional representation based on the sample depth maps and learn the spatiotemporal characteristics of the sample videos based on the sample video frames, thereby maintaining the geometric consistency and temporal consistency of the generated depth maps, thereby effectively improving the accuracy of video depth estimation.

[0150] Furthermore, in order to enable the diffusion model to learn the temporal consistency between each sample video frame in the sample video, the sample video can be divided into multiple overlapping sample subsequences. Except for the first sample subsequence, the remaining sample subsequences include video frames in the overlapping part and video frames in the non-overlapping part. For the sample video frames in the overlapping part of the current sample subsequence, the noise added in the prediction based on the previous sample subsequence can be used as the initial noise of the sample video frames in the overlapping part of the current sample subsequence. In this way, the scale information of the sample video frames can be retained in the overlapping area and passed to the subsequent sample subsequences through the diffusion model, ensuring the transferability and continuity of the sample video data.

[0151] In a possible implementation, the diffusion model includes a UNet network, and in the process of using the sample video latent variable as a conditional signal to predict noise and obtain multiple frames of sample noise, the sample depth latent variable after adding reference noise can be mapped based on the UNet network to predict noise and obtain multiple frames of sample noise. The UNet network includes a cross-attention mechanism, and in the mapping process of the UNet network, the sample video latent variable is used as a conditional signal and injected into the UNet network through the cross-attention mechanism.

[0152] Specifically, the sample video latent variables and the sample depth latent variables with reference noise added are input into the UNet network, and the sample video latent variables and the sample depth latent variables with reference noise added are fused through the cross attention mechanism to obtain fused features. UNet performs noise prediction based on the fused features to obtain sample noise. Figure 5 , Figure 5 An optional schematic diagram of injecting a conditional signal into a UNet network provided in an embodiment of the present disclosure, wherein the sample video latent variable z t Input into the self-attention mechanism layer for feature extraction to obtain the self-attention feature z 1 , based on the self-attention feature z 1 Construct the query matrix Q based on the sample depth latent variables Construct the value matrix V and the key matrix K, input the query matrix Q, the value matrix V and the key matrix K into the cross attention mechanism layer for feature fusion, and obtain the fused feature Z O , based on the fusion feature Z ONoise prediction. The cross-attention mechanism is used to fuse the features of the sample video latent variables and the sample depth latent variables with reference noise added to achieve the injection of conditional signals, so that the cross-attention mechanism can capture the correlation and dependency between the sample video latent variables and the sample depth latent variables with reference noise added, allowing the UNet network to more comprehensively understand the complexity between the sample video latent variables and the sample depth latent variables with reference noise added, thereby improving the accuracy of the UNet network in predicting noise.

[0153] The video processing method provided by the embodiment of the present disclosure uses geometric constraint information as a diffusion guide to update the prediction noise, uses global constraints to align the prediction noise of the overlapping part with the non-overlapping prediction noise, and also uses local constraints to update the non-overlapping prediction noise, which can enhance the geometric consistency between different prediction noises. Then, based on the updated prediction noise, the noise latent variable of the next time step is obtained until the target depth latent variable of the last time step is obtained, the target depth latent variable is decoded, and the target video is processed into a target depth map corresponding to the target video, so that the obtained target depth map can maintain temporal consistency and geometric consistency, thereby effectively improving the accuracy of video depth estimation.

[0154] Reference Figure 6 , Figure 6 An optional overall flow chart of the video processing method provided in the embodiment of the present disclosure is provided below. The principle of the video processing method in the embodiment of the present disclosure is fully described in general as follows:

[0155] The depth generation method provided by the embodiment of the present disclosure can be implemented based on a diffusion model and a decoder, wherein the diffusion model includes a UNet network, and the UNet network is used for denoising operations.

[0156] Before generating the target depth map, it is necessary to train the diffusion model, obtain the sample video, extract the sample depth map corresponding to each sample video frame from the sample video, input the sample video into the encoder for sample video frame encoding, and encode each sample video frame containing depth information into the latent space through the encoder. At the same time, the sample depth map is normalized into a sample depth map with affine invariance. After copying the sample depth map with affine invariance three times, the three sample depth maps with affine invariance are encoded into the latent space through the encoder to obtain the sample video latent variables corresponding to the sample video frame of the sample video and the sample depth latent variables of the sample depth map. Then, add reference noise to each sample depth latent variable, input the sample depth latent variables after adding noise and the sample video latent variables into the diffusion model for denoising. In the process of calling the UNet network for denoising, the sample video latent variables are used as conditional signals, and the sample video latent variables are injected into the UNet network based on the feature fusion operation of the cross attention mechanism, so that the UNet network performs noise prediction based on the fusion features to obtain the sample noise corresponding to each video frame. Finally, the model loss is determined according to the difference between the sample noise and the reference noise, and the diffusion model is trained based on the model loss.

[0157] After obtaining the pre-trained diffusion model, freeze the diffusion model and decoder to ensure the stability and consistency of the diffusion model and decoder. Obtain the target video, input the target video into the encoder, encode the video frame of the target video into the latent space, obtain the deep latent variables corresponding to each video frame, add noise to the deep latent variables, and obtain the noise latent variable z t , the noise latent variable z t The deep latent variables corresponding to the target video Input into the diffusion model for noise prediction, and obtain the predicted noise ε corresponding to each video frame in time step t based on the diffusion prior θ (z t ,t). The prediction noise ε based on time step t θ (z t ,t) Denoise the noise latent variables of each video frame to obtain the original noise latent variables without the prediction noise of the current time step. Input the noise latent variables without the prediction noise of the current time step into the decoder for decoding to obtain the original depth map d corresponding to each video frame in the target video at the current time step t .

[0158] Next, based on the sliding window strategy, the prediction noise is divided into multiple overlapping subsequences, and the subsequences include the prediction noise of the overlapping area and the prediction noise of the non-overlapping area. The geometric constraint loss corresponding to each subsequence is determined based on the original depth map. The gradient corresponding to the prediction noise pair or single prediction noise in each subsequence is determined based on the geometric constraint loss. The gradient corresponding to the prediction noise pair or single prediction noise is guided by gradient accumulation to obtain the gradient corresponding to each subsequence, and the prediction noise ε is updated based on the gradient by back propagation. θ (z t ,t). Based on the gradient update prediction noise ε θ (z t ,t), a sliding refinement strategy is adopted. For the current subsequence, global constraints are applied based on the reprojection loss, depth loss, and tracking loss in the geometric constraint loss to align the prediction noise of the overlapping part and the prediction noise of the non-overlapping part in the current subsequence. Then, the prediction noise of the overlapping part is fixed, and the prediction noise of the non-overlapping part is updated based on the surface normal loss and smoothness loss in the geometric constraint loss. The adjusted prediction noise is obtained.

[0159] In the process of determining the geometric constraint loss, two video frames I are randomly selected from the target video based on the preset inter-frame interval threshold. i and I j , the video frame I i The pixel x in i Projected into three-dimensional space, we get a three-dimensional point X i , the three-dimensional point X i The spatial coordinate system is converted to the video frame I j The spatial coordinate system of the converted three-dimensional point X i→j , the three-dimensional point X i→j Projection to video frame I j In the projection point x i→j , based on the projection point x i→j and video frame I j The pixel x in j Determine the remapping loss L r .

[0160] Next, the video frame I j The pixel x in j Projected into three-dimensional space, we get a three-dimensional point X j , based on the three-dimensional point X j and the 3D point X i→j Determine the depth loss L d Then, based on the video frame I i and video frame I j The original depth map d tGet the respective projection depths, and transform the video frames I i The pixel x in i and video frame I j The pixel x in j Project to three-dimensional space and get the three-dimensional projection point set Y i and the three-dimensional projection point set Y j , using video frame I i The camera pose (R i ,t i ) and video frame I j The camera pose (R j ,t j ) The three-dimensional projection point set Y i The spatial coordinate system is converted to the video frame I j The spatial coordinate system of the transformed three-dimensional projection point set Y i→j In the camera pose (R, t), R represents the rotation matrix and t represents the translation vector. Based on the transformed three-dimensional projection point set Y i→j and the three-dimensional projection point set Y j Determine the tracking loss L t .

[0161] Next, get the video frame I i The second surface normal of the objects in the scene and the normal of the second surface of the objects in the scene based on the video frame I i The corresponding original depth map d t The first surface normal is obtained, and the surface normal loss L is determined based on the difference between the first surface normal and the second surface normal. n ; Get video frame I i The original depth map d t , based on the original depth map d t The depth gradient of determines the smoothness loss. The remapping loss, depth loss, tracking loss, surface normal loss, and smoothness loss are weighted summed to obtain the geometric constraint loss.

[0162] Next, from the noise latent variable z t Remove the adjusted forecast noise from Get the deep latent variable for the next denoising step (time step t-1) Until the target depth latent variable of the last time step is obtained The target deep latent variable The input is sent to the decoder for decoding, and the target depth latent variable is decoded from the latent space to the image space to obtain the target depth map corresponding to each video frame.

[0163] It should also be noted that in Figure 7 In

[15] , guided gradient accumulation and sliding refinement strategies are used to perform inference on long video sequences.

[0164] In one possible implementation, the video processing method provided in this embodiment can be applied to three-dimensional reconstruction and visualization tasks, obtaining the noise latent variable corresponding to the target video, inputting the noise latent variable into the diffusion model to obtain the predicted noise of the current time step, determining the corresponding geometric constraint loss based on the predicted noise of the current time step, updating the predicted noise based on the geometric constraint loss, denoising the noise latent variable based on the updated predicted noise, obtaining the noise latent variable of the next time step, until obtaining the target depth latent variable of the last time step, decoding the target latent variable, obtaining the target depth map corresponding to the target video, performing three-dimensional reconstruction based on the target depth map, and generating an accurate three-dimensional model, which can be used for visualization and analysis in the fields of architecture, product design, and urban planning.

[0165] It is to be understood that, although the steps in the above-mentioned various flow charts are shown in sequence according to the indication of the arrows, these steps are not necessarily performed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in the present embodiment, the execution of these steps does not have a strict order restriction, and these steps can be performed in other orders. Moreover, at least a portion of the steps in the above-mentioned flow charts can include a plurality of steps or a plurality of stages, and these steps or stages are not necessarily performed at the same time, but can be performed at different times, and the execution order of these steps or stages is not necessarily performed in sequence, but can be performed in turn or alternately with at least a portion of the steps or stages in other steps or other steps.

[0166] Reference Figure 7 , Figure 7 The video processing device 700 is a schematic diagram of the structure of the video processing device provided in the embodiment of the present disclosure. The video processing device 700 includes:

[0167] The noise prediction module 701 is used to obtain multi-frame noise latent variables with the same number of frames as the target video, and perform noise prediction based on the multi-frame noise latent variables to obtain multi-frame predicted noise;

[0168] The first generation module 702 is used to remove the prediction noise from each frame noise latent variable respectively, obtain multiple frames of original depth latent variables, decode the multiple frames of original depth latent variables, and obtain the original depth map corresponding to each video frame in the target video;

[0169] An adjustment module 703 is used to divide the multi-frame prediction noise into multiple overlapping subsequences, determine the geometric constraint loss based on the original depth map in each subsequence, and adjust the prediction noise of the non-overlapping part based on the geometric constraint loss, wherein the geometric constraint loss is used to constrain geometric consistency when the corresponding video frame is three-dimensionally projected based on the original depth map;

[0170] The second generation module 704 is used to remove the adjusted prediction noise from each frame noise latent variable, obtain multi-frame target depth latent variables, decode the multi-frame target depth latent variables, and obtain a multi-frame target depth map corresponding to each video frame.

[0171] Further, the adjustment module 703 is also used for:

[0172] Determining a first constraint loss based on the original depth map, wherein the first constraint loss is used to constrain scene consistency when three-dimensionally projecting the corresponding video frame based on the original depth map;

[0173] Determining a second constrained loss based on the original depth map, wherein the second constrained loss is used to constrain detail consistency when three-dimensionally projecting the corresponding video frame based on the original depth map;

[0174] The weighted sum of the first constraint loss and the second constraint loss is used to obtain the geometric constraint loss.

[0175] Further, the adjustment module 703 is also used for:

[0176] Determine a three-dimensional projection relationship between an i-th video frame and a j-th video frame based on the i-th original depth map, perform three-dimensional projection on the j-th video frame relative to the i-th video frame based on the three-dimensional projection relationship to obtain a first projection frame, and determine a reprojection loss according to a difference between the first projection frame and the i-th video frame, wherein i and j are positive integers;

[0177] Determine a projection depth map when performing three-dimensional projection on the j-th video frame, and determine a depth loss according to a difference between the projection depth map and the original depth map;

[0178] The first constraint loss is obtained by weighted summing the reprojection loss and the depth loss.

[0179] Further, the adjustment module 703 is also used for:

[0180] Taking the predicted noise of the i-th frame and the predicted noise of the j-th frame as a predicted noise pair, determining a first mean of the reprojection loss corresponding to all the predicted noise pairs in the subsequence, and a second mean of the depth loss corresponding to all the predicted noise pairs in the subsequence;

[0181] The first constraint loss is obtained by weighted summing the first mean and the second mean.

[0182] Further, the adjustment module 703 is also used for:

[0183] Based on the original depth maps corresponding to the two adjacent video frames, three-dimensionally project the two adjacent video frames to obtain a first projection point set and a second projection point set;

[0184] Converting the coordinate system of the second projection point set to the coordinate system of the first projection point set, and determining the tracking loss according to the difference between the first projection point set and the converted second projection point set;

[0185] The first constraint loss is obtained by weighted summing of the reprojection loss, depth loss and tracking loss.

[0186] Further, the adjustment module 703 is also used for:

[0187] Determining the structural similarity between the first projection frame and the i-th video frame;

[0188] The difference between the first projection frame and the i-th video frame and the structural similarity are weighted summed to obtain the reprojection loss.

[0189] Further, the adjustment module 703 is also used for:

[0190] Calculating a first surface normal based on the original depth map, generating a second surface normal of the corresponding video frame based on a surface normal generation network, and determining a surface normal loss according to a difference between the first surface normal and the second surface normal;

[0191] Determine the depth gradient of the original depth map and regularize the depth gradient to obtain a smooth loss;

[0192] The surface normal loss and smoothness loss are weighted summed to obtain the second constraint loss.

[0193] Further, the adjustment module 703 is also used for:

[0194] Determine a third mean of the surface normal loss corresponding to all prediction noises in the subsequence and a fourth mean of the smoothing loss corresponding to all prediction noises in the subsequence;

[0195] The third mean and the fourth mean are weighted summed to obtain the second constraint loss.

[0196] Further, the adjustment module 703 is also used for:

[0197] Determine the gradient of the geometric constraint loss, and get the adjustment term based on the product of the gradient and the guidance strength;

[0198] The adjusted prediction noise is obtained according to the sum of the adjustment item and the prediction noise of the non-overlapping part.

[0199] Further, the adjustment module 703 is also used for:

[0200] Obtain a multi-frame sample depth map corresponding to the sample video, extract a sample depth latent variable of the sample depth map and a multi-frame sample video latent variable of the sample video;

[0201] Add reference noise to the multi-frame sample depth latent variables, input the sample depth latent variables after adding the reference noise and the sample video latent variables into the diffusion model, use the sample video latent variables as conditional signals to perform noise prediction, and obtain the multi-frame sample noise;

[0202] The model loss is determined according to the difference between the sample noise and the reference noise, and the diffusion model is trained based on the model loss.

[0203] Further, the adjustment module 703 is also used for:

[0204] Based on the UNet network, the sample depth latent variables after adding reference noise are mapped to perform noise prediction and obtain multi-frame sample noise;

[0205] Among them, in the mapping process of the UNet network, the sample video latent variables are used as conditional signals and injected into the UNet network through the cross-attention mechanism.

[0206] In summary, the video processing device 700 in the embodiment of the present disclosure obtains multi-frame noise latent variables with the same number of frames as the target video, performs noise prediction based on the multi-frame noise latent variables, obtains multi-frame predicted noise, removes the predicted noise from each frame noise latent variable, obtains multi-frame original depth latent variables, decodes the multi-frame original depth latent variables, obtains the original depth map corresponding to each video frame in the target video, divides the multi-frame predicted noise into multiple overlapping sub-sequences, and can support longer video input on the basis of keeping the computing resources unchanged, thereby improving the temporal consistency between the depth maps generated subsequently, and then, in each sub-sequence, determines the geometric constraint loss based on the original depth map. , the prediction noise of the non-overlapping part is adjusted based on the geometric constraint loss. Since the geometric constraint loss is used to constrain the geometric consistency when the corresponding video frame is three-dimensionally projected based on the original depth map, it is equivalent to enhancing the geometric consistency between different subsequences and different prediction noises through the adjustment method of the sliding window, thereby improving the geometric consistency between subsequently generated depth maps. Based on this, the adjusted prediction noise is removed from the noise latent variables of each frame respectively, and a more accurate multi-frame target depth latent variable can be obtained. When the multi-frame target depth latent variable is decoded and the multi-frame target depth map corresponding to each video frame is obtained, the accuracy of video depth estimation can be effectively improved.

[0207] The electronic device for executing the above-mentioned video processing method provided in the embodiment of the present disclosure may be a terminal. Figure 8 , Figure 8This is a partial structural block diagram of a terminal provided in an embodiment of the present disclosure, and the terminal includes: a camera assembly 810, a first memory 820, an input unit 830, a display unit 840, a sensor 850, an audio circuit 860, a wireless fidelity (Wi-Fi) module 870, a first processor 880, and a first power supply 890. Those skilled in the art can understand that Figure 8 The terminal structure shown in the figure does not constitute a limitation on the terminal, and may include more or less components than shown in the figure, or combine certain components, or arrange the components differently.

[0208] The camera assembly 810 can be used to capture images or videos. Optionally, the camera assembly 810 includes a front camera and a rear camera. Typically, the front camera is disposed on the front panel of the terminal, and the rear camera is disposed on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth of field camera, a wide-angle camera, and a telephoto camera, so as to realize the fusion of the main camera and the depth of field camera to realize the background blur function, the fusion of the main camera and the wide-angle camera to realize the panoramic shooting and VR (Virtual Reality) shooting function or other fusion shooting functions.

[0209] The first memory 820 may be used to store software programs and modules. The first processor 880 executes various functional applications and data processing of the terminal by running the software programs and modules stored in the first memory 820 .

[0210] The input unit 830 may be used to receive input digital or character information and generate key signal input related to the terminal's settings and function control. Specifically, the input unit 830 may include a touch panel 831 and other input devices 832 .

[0211] The display unit 840 may be used to display input information or provided information and various menus of the terminal. The display unit 840 may include a display panel 841.

[0212] The audio circuit 860, the speaker 861, and the microphone 862 may provide an audio interface.

[0213] The first power source 890 may be alternating current, direct current, a disposable battery, or a rechargeable battery.

[0214] The number of sensors 850 may be one or more, and the one or more sensors 850 include but are not limited to: acceleration sensors, gyroscope sensors, pressure sensors, optical sensors, etc. Among them:

[0215] The acceleration sensor can detect the magnitude of acceleration on the three coordinate axes of the coordinate system established by the terminal. For example, the acceleration sensor can be used to detect the components of gravity acceleration on the three coordinate axes. The first processor 880 can control the display unit 840 to display the user interface in a horizontal view or a vertical view according to the gravity acceleration signal collected by the acceleration sensor. The acceleration sensor can also be used for collecting motion data of games or users.

[0216] The gyroscope sensor can detect the body direction and rotation angle of the terminal, and the gyroscope sensor can cooperate with the acceleration sensor to collect the user's 3D actions on the terminal. The first processor 880 can implement the following functions based on the data collected by the gyroscope sensor: motion sensing (such as changing the UI according to the user's tilt operation), image stabilization during shooting, game control, and inertial navigation.

[0217] The pressure sensor can be set in the side frame of the terminal and / or the lower layer of the display unit 840. When the pressure sensor is set in the side frame of the terminal, the user's holding signal of the terminal can be detected, and the first processor 880 performs left and right hand recognition or shortcut operation according to the holding signal collected by the pressure sensor. When the pressure sensor is set in the lower layer of the display unit 840, the first processor 880 controls the operability controls on the UI interface according to the user's pressure operation on the display unit 840. The operability control includes at least one of a button control, a scroll bar control, an icon control, and a menu control.

[0218] The optical sensor is used to collect the ambient light intensity. In one embodiment, the first processor 880 can control the display brightness of the display unit 840 according to the ambient light intensity collected by the optical sensor. Specifically, when the ambient light intensity is high, the display brightness of the display unit 840 is increased; when the ambient light intensity is low, the display brightness of the display unit 840 is reduced. In another embodiment, the first processor 880 can also dynamically adjust the shooting parameters of the camera assembly 810 according to the ambient light intensity collected by the optical sensor.

[0219] In this embodiment, the first processor 880 included in the terminal can execute the video processing method of the previous embodiment.

[0220] The electronic device for executing the above-mentioned video processing method provided in the embodiment of the present disclosure may also be a server. Fig. 9 , Fig. 9This is a partial structural block diagram of a server provided in an embodiment of the present disclosure. The server may have relatively large differences due to different configurations or performances, and may include one or more second processors 910 and a second memory 930, and one or more storage media 940 (e.g., one or more mass storage devices) storing application programs 943 or data 942. Among them, the second memory 930 and the storage medium 940 may be short-term storage or persistent storage. The program stored in the storage medium 940 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Furthermore, the second processor 910 may be configured to communicate with the storage medium 940 to execute a series of instruction operations in the storage medium 940 on the server.

[0221] The server may also include one or more second power supplies 920, one or more wired or wireless network interfaces 950, one or more input and output interfaces 960, and / or one or more operating systems 941, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.

[0222] The second processor 910 in the server may be configured to execute the video processing method.

[0223] The embodiments of the present disclosure further provide a computer-readable storage medium, which is used to store a computer program, and the computer program is used to execute the video processing methods of the aforementioned embodiments.

[0224] The embodiment of the present disclosure also provides a computer program product, which includes a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device executes the above-mentioned video processing method.

[0225] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate to describe the embodiments of the present disclosure, such as being able to be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0226] It should be understood that in the present disclosure, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the objects associated before and after are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0227] It should be understood that in the description of the embodiments of the present disclosure, the meaning of multiple (or multiple items) is more than two, greater than, less than, exceed, etc. are understood to not include the number, and above, below, within, etc. are understood to include the number.

[0228] In the several embodiments provided in the present disclosure, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0229] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0230] In addition, each functional unit in each embodiment of the present disclosure may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0231] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present disclosure is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the various embodiments of the present disclosure. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory, referred to as ROM), random access memory (Random Access Memory, referred to as RAM), disk or optical disk and other media that can store program codes.

[0232] It should also be understood that the various implementations provided in the embodiments of the present disclosure can be combined arbitrarily to achieve different technical effects.

[0233] The above is a specific description of the preferred implementation of the present disclosure, but the present disclosure is not limited to the above-mentioned implementation mode. Technical personnel familiar with the field can also make various equivalent deformations or substitutions under the shared conditions without violating the spirit of the present disclosure. These equivalent deformations or substitutions are all included in the scope defined by the claims of the present disclosure.

Claims

1. A video processing method, characterized in that: include: Acquire multiple frames of noise latent variables having the same number of frames as the target video, perform noise prediction based on the multiple frames of noise latent variables, and obtain multiple frames of predicted noise; Removing the predicted noise from the noise latent variables of each frame respectively to obtain multiple frames of original depth latent variables, decoding the multiple frames of original depth latent variables to obtain original depth maps corresponding to each video frame in the target video; Dividing the predicted noise of multiple frames into multiple overlapping subsequences, determining a geometric constraint loss in each of the subsequences based on the original depth map, and adjusting the predicted noise of the non-overlapping part based on the geometric constraint loss, wherein the geometric constraint loss is used to constrain geometric consistency when three-dimensionally projecting the corresponding video frame based on the original depth map; The adjusted predicted noise is removed from the noise latent variables of each frame respectively to obtain multi-frame target depth latent variables, and the multi-frame target depth latent variables are decoded to obtain multi-frame target depth maps corresponding to each of the video frames.

2. The video processing method according to claim 1, characterized in that: The determining of the geometric constraint loss based on the original depth map comprises: Determining a first constraint loss based on the original depth map, wherein the first constraint loss is used to constrain scene consistency when three-dimensionally projecting the corresponding video frame based on the original depth map; Determining a second constrained loss based on the original depth map, wherein the second constrained loss is used to constrain detail consistency when three-dimensionally projecting the corresponding video frame based on the original depth map; A weighted sum is performed on the first constraint loss and the second constraint loss to obtain a geometric constraint loss.

3. The video processing method according to claim 2, characterized in that: The determining a first constraint loss based on the original depth map comprises: Determine a three-dimensional projection relationship between the video frame of the i-th frame and the video frame of the j-th frame based on the original depth map of the i-th frame, perform three-dimensional projection on the video frame of the j-th frame relative to the video frame of the i-th frame based on the three-dimensional projection relationship to obtain a first projection frame, and determine a reprojection loss according to a difference between the first projection frame and the video frame of the i-th frame, wherein i and j are positive integers; Determine a projection depth map when three-dimensionally projecting the j-th video frame, and determine a depth loss according to a difference between the projection depth map and the original depth map; The reprojection loss and the depth loss are weightedly summed to obtain a first constraint loss.

4. The video processing method according to claim 3, characterized in that: The step of performing a weighted summation on the reprojection loss and the depth loss to obtain a first constraint loss includes: Taking the predicted noise of the i-th frame and the predicted noise of the j-th frame as a predicted noise pair, determining a first mean value of the reprojection loss corresponding to all the predicted noise pairs in the subsequence, and a second mean value of the depth loss corresponding to all the predicted noise pairs in the subsequence; A first constraint loss is obtained by performing a weighted summation on the first mean and the second mean.

5. The video processing method according to claim 3, characterized in that: The step of performing a weighted summation on the reprojection loss and the depth loss to obtain a first constraint loss includes: Based on the original depth maps corresponding to the two adjacent video frames, three-dimensionally project the two adjacent video frames to obtain a first projection point set and a second projection point set; Converting the coordinate system of the second projection point set to the coordinate system of the first projection point set, and determining the tracking loss according to the difference between the first projection point set and the converted second projection point set; The reprojection loss, the depth loss, and the tracking loss are weightedly summed to obtain a first constraint loss.

6. The video processing method according to claim 3, characterized in that: The determining the reprojection loss according to the difference between the first projection frame and the i-th video frame comprises: Determining a structural similarity between the first projection frame and the i-th video frame; The difference between the first projection frame and the i-th video frame and the structural similarity are weightedly summed to obtain a reprojection loss.

7. The video processing method according to claim 2, characterized in that: The determining a second constraint loss based on the original depth map comprises: Calculating a first surface normal based on the original depth map, generating a second surface normal of the corresponding video frame based on a surface normal generation network, and determining a surface normal loss according to a difference between the first surface normal and the second surface normal; Determining a depth gradient of the original depth map, and regularizing the depth gradient to obtain a smoothing loss; The surface normal loss and the smoothness loss are weightedly summed to obtain a second constraint loss.

8. The video processing method according to claim 7, characterized in that: The weighted summing of the surface normal loss and the smoothness loss to obtain a second constraint loss includes: Determine a third mean of the surface normal loss corresponding to all the prediction noises in the subsequence, and a fourth mean of the smoothing loss corresponding to all the prediction noises in the subsequence; A weighted sum is performed on the third mean and the fourth mean to obtain a second constraint loss.

9. The video processing method according to claim 1, characterized in that: The step of adjusting the prediction noise of the non-overlapping portion based on the geometric constraint loss comprises: Determining the gradient of the geometric constraint loss, and obtaining an adjustment term according to the product of the gradient and the guidance strength; The adjusted prediction noise is obtained according to the sum of the adjustment item and the prediction noise of the non-overlapping part.

10. The video processing method according to claim 1, characterized in that: The predicted noise is obtained by prediction through a diffusion model, and the diffusion model is trained through the following steps: Obtaining a multi-frame sample depth map corresponding to a sample video, extracting a sample depth latent variable of the sample depth map and a multi-frame sample video latent variable of the sample video; Adding reference noise to the sample depth latent variables of multiple frames, inputting the sample depth latent variables after adding the reference noise and the sample video latent variables into the diffusion model, performing noise prediction with the sample video latent variables as conditional signals, and obtaining the sample noise of multiple frames; A model loss is determined according to a difference between the sample noise and the reference noise, and the diffusion model is trained based on the model loss.

11. The video processing method according to claim 10, characterized in that: The diffusion model includes a UNet network, and the noise prediction is performed using the sample video latent variable as a conditional signal to obtain multi-frame sample noise, including: Based on the UNet network, the sample depth latent variables after adding the reference noise are mapped to perform noise prediction to obtain multi-frame sample noise; In the mapping process of the UNet network, the sample video latent variable is used as a conditional signal and injected into the UNet network through a cross-attention mechanism.

12. A video processing device, characterized in that: include: A noise prediction module is used to obtain multiple frames of noise latent variables with the same number of frames as the target video, and perform noise prediction based on the multiple frames of noise latent variables to obtain multiple frames of predicted noise; A first generating module is used to remove the predicted noise from the noise latent variables of each frame respectively to obtain multiple frames of original depth latent variables, and decode the multiple frames of the original depth latent variables to obtain the original depth map corresponding to each video frame in the target video; An adjustment module, configured to divide the predicted noise of multiple frames into multiple overlapping subsequences, determine a geometric constraint loss in each of the subsequences based on the original depth map, and adjust the predicted noise of the non-overlapping part based on the geometric constraint loss, wherein the geometric constraint loss is used to constrain geometric consistency when three-dimensionally projecting the corresponding video frame based on the original depth map; The second generation module is used to remove the adjusted predicted noise from the noise latent variables of each frame, obtain multi-frame target depth latent variables, decode the multi-frame target depth latent variables, and obtain multi-frame target depth maps corresponding to each of the video frames.

13. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the video processing method according to any one of claims 1 to 11 is implemented.

14. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the video processing method according to any one of claims 1 to 11 is implemented.

15. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the video processing method according to any one of claims 1 to 11 is implemented.