A virtual camera posture adjustment method, system and computer program product

Through the implicit neural radiation field and multi-feature collaborative optimization method, the problem of camera pose optimization in virtual scenes was solved, efficient camera pose adjustment was achieved, and the shooting effect and creative expression of virtual performances were improved.

CN119251284BActive Publication Date: 2025-09-26COMMUNICATION UNIVERSITY OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411014647.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-26
Publication Date
2025-09-26
Estimated Expiration
2044-07-26

AI Technical Summary

Technical Problem

When shooting performances in virtual scenes, existing technologies lack a fast and effective camera pose optimization method, resulting in time-consuming and labor-intensive reconstruction with poor results. It is difficult to ensure the consistency of 3D composition information, which affects the shooting effect.

Method used

A multi-feature collaborative optimization method combining implicit neural radiation field with optical flow, joint point heat map and depth features is adopted. The implicit expression model is determined by the dynamic neural radiation field, and the camera pose is optimized using optical flow loss, joint point loss and depth loss function. The shooting parameters of the virtual camera are optimized by guiding the selective back propagation gradient of the region.

Benefits of technology

It significantly improves the quality of generated camera pose sequences, reduces the workload of scene reconstruction, and makes the rendered video consistent with the reference video, enhancing the creativity and visual impact of performances.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119251284B_ABST
    Figure CN119251284B_ABST
Patent Text Reader

Abstract

The present application relates to the field of computer vision and discloses a method for optimizing the pose of a virtual camera. The method comprises: determining an implicit expression model corresponding to a preset reference video through a dynamic neural radiation field, and determining a scene rendering video through the implicit neural expression model; inputting the rendered video and the preset reference video into a preset virtual camera pose optimization model to determine the optical flow map, character joint heat map, and depth map of each video frame in the reference video and the rendered video; determining the optical flow loss function, the joint loss function, and the depth loss function and determining the guide area; for any pixel point of the video frame in the rendered video, determining the first return gradient, the second return gradient, and the third return gradient according to the optical flow loss function, the joint loss function, the depth loss function, and the guide area, and returning them to optimize the shooting parameters of the virtual camera. It can reduce memory consumption and provide a mechanism similar to attention to help convergence and focus on information-rich areas.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical fields of computer vision and virtual scene shooting, and in particular to a virtual camera posture adjustment method, system, and computer program product. Background Art

[0002] The purpose of the background technology description provided here is to give an overall background of the present application. The statements in this section merely provide background related to the present application and do not necessarily constitute prior art.

[0003] Under existing technology, the actual video shooting process for a traditional performing arts program generally includes multiple steps, including scene design, performance arrangement, plot design, and camera position setting. Each shot is affected by these factors, and due to objective conditions, video producers may not be able to capture the best possible shot. Furthermore, due to the uncertainty inherent in such programs, directors often choose to conduct pre-production rehearsals, which requires significant manpower and resources and is also affected by various factors such as time and weather. Therefore, high costs, low controllability, and the need for advance rehearsals are the main challenges of traditional performing arts programs. However, if the program could be filmed in a virtual setting, the impact of these factors could be reduced, helping directors quickly and easily find the optimal shooting plan and achieve the best visual effects, improving production efficiency and significantly reducing costs. However, while filming in virtual environments offers many advantages, virtual environments are currently typically constructed manually using virtual engines. Manually modeling performance scenes is complex and time-consuming, requiring the creation of object materials, textures, and lighting. Furthermore, ensuring model fidelity is difficult. Some objects with specialized materials (such as glass and fog) are difficult to reproduce effectively through manual modeling, resulting in time-consuming and labor-intensive virtual scene reconstruction and poor quality. Consequently, there is still a significant lack of methods for quickly expressing virtual environments.

[0004] When filming a performance in a virtual setting, besides the creativity and quality of the performance itself, camera movement is also crucial, as it directly impacts the audience's experience. Appropriate use of camera language can enhance the artistic expression, quality, and creativity of digital performance content, thereby better meeting audience needs and providing an immersive experience. Therefore, determining the camera trajectory within the virtual setting is crucial. To facilitate this, traditional program filming employs "visual homage" techniques. This involves imitating classic shots from films or performances. This helps filmmakers convey specific emotions or themes, enhancing the depth of their work and resonating with the audience. Inspired by these "visual homages," filmmakers can intelligently analyze the camera trajectories and techniques from existing footage and apply them to their virtual performances. This allows them to determine the camera position within the virtual setting without requiring in-depth knowledge of camera language or camera operation techniques. Furthermore, simply by modifying the existing reference video, they can create diverse and flexible virtual scene video clips. This technology, which facilitates precise manipulation and reproduction of visual elements such as camera movement, will help create more creative and visually impactful performing arts works.

[0005] In order to optimize the camera pose and ensure that the reference video and the rendered video are shot using the same technique, it is necessary to consider whether multiple features between the two are consistent. Existing methods often consider features such as character action, composition, camera motion, and aesthetic evaluation. Among them, composition features generally consider the composition of the character, using the character's joint features to represent the character's position and size, etc. However, the position of the character on the screen is only a 2D composition information. In real-life shooting, it is often necessary to fully consider the perspective relationship in the picture, that is, 3D composition information. Therefore, there is still a lack of a feature that can represent the 3D composition information of the character to assist in guiding the optimization of the camera pose.

[0006] Therefore, a new camera pose optimization method is urgently needed to overcome and solve the above-mentioned defects. Summary of the Invention

[0007] In response to the above problems, the present application proposes a virtual camera posture adjustment method, a virtual camera posture adjustment system and a storage medium.

[0008] A first aspect of the present application provides a virtual camera pose adjustment method, comprising:

[0009] Determine an implicit expression model corresponding to a preset reference video through a dynamic neural radiation field, and determine a rendered video under a current camera trajectory pose through the implicit expression model according to any camera trajectory pose in a preset camera trajectory pose set;

[0010] Inputting the rendered video and the preset reference video into a preset virtual camera pose optimization model to respectively determine the optical flow map, character joint point heat map, and depth map of each video frame in the preset reference video and the rendered video;

[0011] Respectively determining an optical flow loss function corresponding to the optical flow map, a joint point loss function corresponding to the character joint point heat map, and a depth loss function corresponding to the depth map;

[0012] Determining a guide area using a preset guide area determination model based on the optical flow map, character joint heat map, and depth map of the preset reference video and the rendered video to highlight key areas whose gradient contribution meets preset conditions;

[0013] For any pixel point of the video frame in the rendered video, a first gradient is determined according to the optical flow loss function, a second gradient is determined according to the joint point loss function, and a third gradient is determined according to the depth loss function; and a first return gradient corresponding to the first gradient, a second return gradient corresponding to the second gradient, and a third return gradient corresponding to the third gradient are respectively determined according to the guide area, and the first return gradient, the second return gradient, and the third return gradient are returned to optimize the shooting parameters of the virtual camera.

[0014] Furthermore, the preset virtual camera pose optimization model includes:

[0015]

[0016] in, is the camera’s extrinsic matrix, For time, is the focal length of the camera, () is an implicit expression model, is the video frame of the reference video;

[0017]

[0018]

[0019] Where i is the video frame number of the reference video, The video frame for the rendered video.

[0020] Furthermore, the preset guidance area determination model includes:

[0021]

[0022] in,( - ) is the second characteristic difference map, To render the heat map of the joint points of the video frame, is the joint point heat map of the reference video frame; ( - ) is the first characteristic difference map, To render the optical flow map of the video frame, is the optical flow map of the reference video frame; ( - ) is the third characteristic difference map, To render the depth map of the video frame, is the depth map of the reference video frame; Indicates normalization processing.

[0023] Furthermore, determining the first gradient according to the optical flow loss function, determining the second gradient according to the joint point loss function, and determining the third gradient according to the depth loss function includes:

[0024] Performing a derivative operation on the optical flow loss function to obtain the first gradient;

[0025] Performing a derivative operation on the joint point loss function to obtain the second gradient;

[0026] Perform a derivative operation on the depth loss function to obtain the third gradient.

[0027] Furthermore, determining a first return gradient corresponding to the first gradient, a second return gradient corresponding to the second gradient, and a third return gradient corresponding to the third gradient according to the guide area includes:

[0028] Taking the dot product of the guide region and the first gradient as the first backpropagation gradient;

[0029] Taking the dot product of the guide region and the second gradient as the second backpropagation gradient;

[0030] A dot product result of the guide region and the third gradient is used as the third return gradient.

[0031] Furthermore, the optical flow loss function includes:

[0032]

[0033] in, and Represent the horizontal optical flow component and vertical optical flow component of the rendered video respectively, and Represent the horizontal optical flow component and vertical optical flow component of the reference video respectively.

[0034] Furthermore, the joint point loss function includes:

[0035]

[0036]

[0037] in, is the Wasserstein distance, and Represents the character joint heat maps of the rendered video frame and the reference video frame respectively, Represents the value of the i-th joint point in the character joint point heat map, and M is the number of joint points.

[0038] Furthermore, the depth loss function includes:

[0039]

[0040] in, is the number of pixels in the video frame, and Represents the rendered video frame and the reference video frame. The depth value of each pixel, H is the height of the reference video frame, and W is the width of the reference video frame.

[0041] A second aspect of the present application provides a virtual camera pose adjustment system, comprising:

[0042] A scene video rendering module is used to determine the implicit expression model corresponding to a preset reference video through a dynamic neural radiation field, and determine the rendered video under the current camera trajectory pose through the implicit expression model according to any camera trajectory pose in a preset camera trajectory pose set;

[0043] A feature estimation module is used to input the rendered video and the preset reference video into a preset virtual camera pose optimization model to respectively determine the optical flow map, character joint point heat map and depth map of each video frame in the preset reference video and the rendered video;

[0044] A loss function determination module is used to respectively determine an optical flow loss function corresponding to the optical flow map, a joint point loss function corresponding to the character joint point heat map, and a depth loss function corresponding to the depth map;

[0045] A guide region determination module is configured to determine a guide region using a preset guide region determination model based on the optical flow map, character joint heat map, and depth map of the preset reference video and the rendered video, so as to highlight a key region whose gradient contribution meets preset conditions;

[0046] A gradient feedback module is used to determine, for any pixel point of a video frame in the rendered video, a first gradient according to the optical flow loss function, a second gradient according to the joint point loss function, and a third gradient according to the depth loss function; and to determine, according to the guide area, a first feedback gradient corresponding to the first gradient, a second feedback gradient corresponding to the second gradient, and a third feedback gradient corresponding to the third gradient, and feedback the first feedback gradient, the second feedback gradient, and the third feedback gradient to optimize the shooting parameters of the virtual camera.

[0047] According to a third aspect of the present application, a computer program product is provided, which includes a computer program or instructions, and when the computer program or instructions are executed by a processor, the steps of the method described above are implemented.

[0048] Compared with the prior art, the advantages or beneficial effects of the technical solution of this application include:

[0049] This application is based on the implicit neural expression of the performance scene and a method of collaborative optimization using three features: optical flow, joint point heat map and depth. It is more suitable for optimizing the camera pose in the performance scene. The effect achieved is relatively consistent with the reference video, which can significantly improve the quality of the generated camera pose sequence and reduce the workload of scene reconstruction.

[0050] This application is based on the problems existing in the existing camera pose optimization schemes, and uses some pre-trained models in existing work as the basis for data preprocessing, to disclose a multi-feature collaborative virtual camera pose optimization model based on implicit neural expression of performing arts scenes. The scene is implicitly neurally expressed using a dynamic neural radiation field to obtain rendered videos under different camera trajectories. At the same time, three differentiable downstream networks are designed to extract optical flow, character joint heat maps and depth features from the rendered video and the reference video, respectively, and calculate the loss between these three features to measure the differences in camera and character motion, character composition and picture perspective between the reference video and the neural radiation field rendered video, and finally backpropagate to optimize the camera pose. Since the overall model is differentiable, the camera pose can be iteratively optimized end-to-end to reduce the cumulative error, and finally a rendered video with a similar shooting effect to the reference video is obtained. At the same time, in order to solve the problem that the back propagation of gradients of all pixels will lead to excessive computational consumption and the background pixels in some losses provide less information, this application proposes to calculate a guidance map by generating a joint difference map between the joint point heat map, optical flow map and depth map to highlight the key areas whose gradient contribution meets the preset conditions, and selectively backpropagate the gradients of these areas to reduce memory consumption and provide a mechanism similar to attention to help convergence and focus on information-rich areas. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present application. For ordinary technicians in the relevant field, other drawings can be obtained based on the provided drawings without any creative work.

[0052] It should also be noted that, for ease of description, only the portions relevant to the present disclosure are shown in the accompanying drawings. The accompanying drawings, which constitute a part of this application, are intended to provide a further understanding of this application. The exemplary embodiments and their descriptions in this application are intended to explain this application and do not constitute an undue limitation of this application. In the accompanying drawings:

[0053] Figure 1 A flowchart of a virtual camera posture adjustment method provided in an embodiment of the present application;

[0054] Figure 2 is a flowchart of another virtual camera pose adjustment method according to an embodiment of the present application;

[0055] Figure 3 A logical diagram of a virtual camera pose optimization generation method provided in an embodiment of the present application;

[0056] Figure 4 A diagram showing the joint point heat map feature extraction results provided in the embodiment of this application;

[0057] Figure 5 A diagram showing the results of deep feature extraction provided in an embodiment of the present application;

[0058] Figure 6 A diagram showing the optical flow feature extraction results provided in an embodiment of the present application;

[0059] Figure 7 A diagram showing the results of the guidance area provided in the embodiment of the present application;

[0060] Figure 8 A visual display of the experimental results provided in the examples of this application. DETAILED DESCRIPTION

[0061] The following will describe the implementation methods of this application in detail with reference to the accompanying drawings and examples, so that the application can fully understand how technical means are used to solve technical problems and achieve corresponding technical effects, and implement them accordingly. The embodiments of this application and the various features therein can be combined with each other without conflict, and the resulting technical solutions are all within the scope of protection of this application.

[0062] It should be clear that the embodiments described below are only some of the embodiments of this application, rather than all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without making any creative work are within the scope of protection of this application.

[0063] Below, some technical terms in the embodiments of the present application and / or the prior art are explained to facilitate those skilled in the art to understand the technical solutions of the present application:

[0064] RAFT (Recurrent All-Pairs Field Transforms, or RAFT) is an optical flow feature extraction network. It consists of three main components: a feature encoder and a semantic feature encoder. The feature encoder extracts pixel information between two consecutive frames, while the semantic feature encoder only extracts pixel information from the first frame of a video. The second component is the correlation calculation module, which calculates a 4D correlation volume by summing the dot products between the feature vectors of two images. The last two dimensions of this volume use a feature pyramid for multi-scale sampling to capture motion at different scales. The third component is the iterative update module, the core module of RAFT. It contains multiple GRU activation units, which process features and information at different times in the feature pyramid and iteratively update the optical flow.

[0065] Monodepth2 is a deep feature extraction network that infers depth information from monocular RGB (red, green, blue) images. The model primarily consists of a depth network and a pose estimation network, both consisting of an encoder and a decoder. The depth network estimates the depth of the target frame, while the pose network uses the target frame and the previous frame as input to predict the camera's angle and position relative to the scene, ultimately outputting the camera's rotation and translation matrices.

[0066] LitePose is a joint point heat map extraction network that uses fused deconvolution layers to eliminate the redundancy of high-resolution branches, allowing single-branch structures to perform scale-aware multi-resolution fusion. At the same time, large convolution kernels are used in the backbone network to improve computational efficiency and reduce model parameters.

[0067] Colmap is open-source computer vision software for reconstructing the structure and camera pose of 3D scenes. It includes functionality for image feature extraction, matching, 3D reconstruction, and dense point cloud generation. With Colmap, users can create an accurate 3D scene model from a set of images.

[0068] NeRF (Neural Radiance Fields) is a neural network model for rendering realistic 3D scenes. It can learn a 3D representation of a scene from a set of 2D images and then use this representation to generate novel viewpoints, perspectives, lighting, and shading effects. Based on the concept of neural radiance fields, NeRF represents the radiance at each 3D point in the scene as a function of the input viewpoint and direction. Through large-scale training and optimization, NeRF is able to produce highly realistic renderings.

[0069] Example 1

[0070] This embodiment provides a virtual camera posture adjustment method, and takes a performance scene as an example to illustrate the virtual camera posture adjustment method disclosed in this embodiment.

[0071] In this embodiment, the virtual camera pose adjustment method is mainly implemented based on implicit neural expression and multi-feature optimization, which can optimize the camera pose sequence in the virtual scene according to a given reference video so that the shooting effect of the rendered video is consistent with that of the reference video.

[0072] First, in this embodiment, the definition of the camera pose adjustment or optimization task (i.e., the preset virtual camera pose optimization model) can be expressed as:

[0073]

[0074] in, , , are the camera parameters to be optimized, is the camera's extrinsic matrix, It's time, Represents the focal length of the camera. are the parameters of the trained NeRF model, The parameter is The NeRF model when is the video frame of the reference video.

[0075] Figure 1 A flowchart of a virtual camera posture adjustment method provided in an embodiment of the present application is shown as follows: Figure 1 As shown, the method disclosed in this embodiment includes the following steps:

[0076] Step 110: Determine an implicit expression model corresponding to a preset reference video through a dynamic neural radiation field, and determine a rendered video under a current camera trajectory pose through the implicit expression model according to any camera trajectory pose in a preset camera trajectory pose set.

[0077] It can be understood that the preset camera trajectory pose set may be an array or a set, and a plurality of preset camera trajectory poses are stored in the preset camera trajectory pose set.

[0078] In some embodiments, the preset virtual camera pose optimization model includes:

[0079]

[0080] in, is the camera’s extrinsic matrix, For time, is the focal length of the camera, are the parameters of the implicit expression model, () is the parameter The corresponding implicit expression model, is the video frame of the reference video;

[0081]

[0082]

[0083] Where i is the video frame number of the reference video (for example, frame 1, frame 2, etc.), The video frame for the rendered video.

[0084] In this step, a dynamic neural radiance field is first used to implicitly represent a performance scene. A video of the scene captured from multiple angles is used as input, and a function model corresponding to the scene is obtained through the implicit neural representation. After obtaining the implicit representation of the scene, the implicit representation model and the randomly initialized camera trajectory pose are used as input. This model is used to render the scene video captured by the camera under this trajectory.

[0085] For a given scene S, we need to first express the scene with a function through the dynamic neural radiation field. Unlike manually building a scene in a virtual engine, we only need to give a video of the scene. A NeRF model can be trained. For a given scene video , we can get each video frame The corresponding camera extrinsic matrix, specifically, use Colmap to obtain the camera's extrinsic matrix :

[0086]

[0087] Since both the scene and the characters may move, a dynamic neural radiation field is used, which takes into account the time variable. .exist moment, from the focal length The camera extrinsic matrix The camera's viewing angle can be rendered and compared with the corresponding video frame Used together to optimize the NeRF model , get the NeRF parameters of scene S :

[0088]

[0089] The dynamic NeRF model of the scene is obtained after training After that, given any observation angle ,time and the focal length of the camera , you can get the rendered image at the corresponding perspective at the corresponding time :

[0090]

[0091] It should be noted that the video frame size of the rendered video and the video frame size of the reference video (for example, the height H and width W of the video frame) may remain the same.

[0092] For example, first, a multi-angle video of a performance scene is used as input, and Colmap is used to estimate the camera's 4×4 pose transformation matrix. Then, a dynamic neural radiation field model of the corresponding scene is trained. After obtaining the dynamic neural radiation field model of the scene, different camera parameters are used as input to obtain scene rendering videos of different angles. Among them, the camera parameters used mainly include the observation angle ,time and the focal length of the camera .

[0093] Step 120: Input the rendered video and the preset reference video into a preset virtual camera pose optimization model to respectively determine the optical flow map, character joint point heat map, and depth map of each video frame in the preset reference video and the rendered video.

[0094] It can be understood that the role of the preset reference video is to allow the model to learn the shooting style from the reference video, and ultimately generate a video output with a similar shooting style to the reference video in the virtual scene.

[0095] In this step, the reference video and the rendered video are used as input to calculate the optical flow information, character joint information, and depth information of the reference video and the rendered video respectively. These three are used to represent the movement of the camera and the character, the character composition, and the perspective relationship of the picture in the reference video and the rendered video respectively.

[0096] Specifically, in order to make the generated neural-rendered virtual video as consistent as possible with the reference video in terms of shooting effect, appropriate features are extracted based on the following three questions:

[0097] 1. Composition issues. Since characters are key in performances, ensuring consistent composition means ensuring that the position and size of the actors in the reference video and the neural rendering video are as consistent as possible.

[0098] 2. Perspective problem: It is necessary to ensure that the distances between the elements of different scenes and the camera in the reference video and the neural rendering video are as consistent as possible. At the same time, the perspective problem also affects the size of the foreground and the middle scene.

[0099] 3. The problem of camera motion and character motion is also a core issue of the research. It is necessary to ensure that the motion of elements of different shot sizes in the reference video and the neural rendering video is as consistent as possible.

[0100] For the first composition problem, we used the joint features of the human body to describe the composition of a single video. To ensure the differentiability of the overall model, we used joint heatmaps to describe the human body's joints. Joint heatmaps represent the confidence that each pixel position is a specific joint. Joint heatmaps also contain richer information than joint coordinates. We used a pre-trained downstream network, LitePose, to predict joint heatmaps:

[0101]

[0102] in, Represent the heatmap of the rendered video frame and the heatmap of the reference video frame respectively, Represents rendered video and reference video frames respectively, Represents the processing of the LitePose model.

[0103] Optionally, use a pre-trained downstream network LitePose to predict joint heatmaps, use a pre-trained downstream depth estimation network Monodepth2 to estimate the depth map of the video, and use a pre-trained downstream optical flow estimation network RAFT to perform optical flow estimation.

[0104] Optionally, the joint heat map feature uses Wasserstein (EMD) distance as the loss function, the optical flow feature uses EPE loss (End-Point Error Loss) as the loss function, and the depth feature uses MSE (Mean Square Error loss) as the loss function.

[0105] like Figure 4 As shown, Figure 4This is a schematic diagram of the results of joint point heat map feature extraction. Pixels that may be joint points are represented by red and yellow. The higher the probability value, the brighter the color.

[0106] For the second perspective problem, we use the depth features of the video to describe the perspective information of a single video. The perspective of the picture is also part of the shooting composition. Perspective is related to the depth and three-dimensionality of the scene being shot. It is used to reflect the distance relationship between people or objects on the screen. It can also divide the foreground, midground, and background areas of the two videos so that they are consistent. Depth represents the distance between objects and the camera in three-dimensional space and is used to determine the distance relationship between objects. The two are relatively consistent. A pre-trained downstream depth estimation network Monodepth2 is used to estimate the depth map of the video:

[0107]

[0108] in, Represent the depth map of the rendered video frame and the depth map of the reference video frame respectively, Represents rendered video and reference video frames respectively, Represents the processing of the Monodepth2 model.

[0109] like Figure 5 As shown, Figure 5 This is a schematic diagram of the results of depth map feature extraction. In the depth map visualization results, the lighter the pixel color, the farther the pixel is from the camera; conversely, the darker the pixel color, the closer the pixel is to the camera.

[0110] For the third problem of camera and character motion, optical flow is used to describe motion information within a single video. Optical flow is a crucial feature, estimating pixel-level displacement between consecutive frames and capturing the motion of characters within the video. In performance scenes, the background generally does not move, so the background's optical flow can also relatively represent the motion of the camera. We use the pre-trained downstream optical flow estimation network, RAFT, for optical flow estimation:

[0111]

[0112] in, and represents the horizontal and vertical optical flow components of the reference video, and and Represent the horizontal optical flow component and vertical optical flow component of the rendered video respectively, and Respectively represent adjacent frames of rendered video and adjacent frames of reference video, Represents the processing of the RAFT model.

[0113] like Figure 6 As shown, Figure 6 This is a schematic diagram of the results of optical flow map feature extraction. In the visualization of the optical flow map, different hues express different directions of optical flow, and saturation represents the size of the optical flow at that point.

[0114] Step 130 : Determine the optical flow loss function corresponding to the optical flow map, the joint loss function corresponding to the character joint heat map, and the depth loss function corresponding to the depth map.

[0115] In some embodiments, the optical flow loss function includes:

[0116]

[0117] in, and Represent the horizontal optical flow component and vertical optical flow component of the rendered video respectively, and Represent the horizontal optical flow component and vertical optical flow component of the reference video respectively.

[0118] In some embodiments, the joint loss function includes:

[0119]

[0120]

[0121] in, is the Wasserstein distance, and Represents the character joint heat maps of the rendered video frame and the reference video frame respectively, Represents the value of the i-th joint point in the character joint point heat map, and M is the number of joint points.

[0122] In some embodiments, the depth loss function includes:

[0123]

[0124] in, is the number of pixels in the video frame, and Represents the rendered video frame and the reference video frame. The depth value of each pixel, H is the height of the reference video frame, and W is the width of the reference video frame.

[0125] Corresponding to the three different features extracted in step 120, three corresponding loss functions are designed in this step. Specifically:

[0126] 1. For the joint point heat map features used, the character joint point loss is introduced , this loss function helps to optimize the problem of inconsistent position and size of characters in the generated neural rendering video and the reference video, and can also constrain the movement of characters in the two videos to a certain extent. In terms of loss function, Wasserstein (EMD) distance is used. Wasserstein distance pays more attention to the spatial layout and position relationship between pixels in the heat map, that is, the similarity between the whole, rather than emphasizing the difference between pixels (single points) like metrics such as mean square error (MSE). That is, Wasserstein distance pays more attention to the measurement of spatial distribution such as the overall size and shape of the character, while MSE pays more attention to the difference of a single local joint point. By minimizing the EMD distance between the heat maps of the joint points of the two characters, the composition information can be transferred from the reference video to the differentiable rendering space of the model. For a size of Image, its joint point heat map ;in, is the number of joints detected by the joint estimation network (for example, by pre-training the downstream network LitePose for size The joint point loss between the final joint point heat map is The calculation formula is as follows:

[0127]

[0128]

[0129] in, represents the Wasserstein distance, and Represent the joint point heat maps of the rendered video frame and the reference video frame respectively, Represents the value of the i-th joint point in the heat map. The summation term in the formula represents the sum of the Wasserstein distances between the same joint points of the two images, and the regularization term in the formula It represents the Wasserstein distance between different joint points on the same heat map. This item is used to ensure the similarity of the joint point shapes in the rendered video generated during the optimization process and the reference video.

[0130] 2. For deep features, depth loss is introduced This loss helps to optimize the perspective relationship between the reference video and the generated neural rendering video to make them as consistent as possible. The loss function uses MSE as the loss function, and its calculation formula is as follows:

[0131]

[0132] in, Represents the number of pixels in the image. and Represents the rendered video frame and the reference video frame respectively. The depth value of each pixel.

[0133] 3. For optical flow features, optical flow loss is introduced In terms of loss function, EPE loss (End-Point Error Loss) is used. EPE is a commonly used evaluation indicator in optical flow estimation tasks. It represents the Euclidean distance between the estimated optical flow and the true optical flow. Its calculation formula is as follows:

[0134]

[0135] in, and represents the horizontal and vertical optical flow components of the rendered video, and and Represent the horizontal optical flow component and vertical optical flow component of the reference video respectively.

[0136] The final overall model loss is composed of the sum of the loss functions of the three downstream networks mentioned above:

[0137]

[0138] in, are the weights corresponding to each loss, Represents the overall model loss.

[0139] Among them, the weights corresponding to each loss are hyperparameters, and it is necessary to conduct multiple experiments and select better parameters from them before making corresponding settings.

[0140] Step 140: Determine a guide area using a preset guide area determination model based on the optical flow map, character joint heat map, and depth map of the preset reference video and the rendered video to highlight key areas whose gradient contribution meets preset conditions.

[0141] In this step, the optical flow map, character joint heat map, and depth map of the reference video and rendered video are used as input to calculate the feature difference map between the two. After obtaining the difference map, it is normalized and summed. Finally, the guidance area is output to distinguish pixels containing large amounts of information and perform gradient backpropagation, thereby accelerating the running speed of the virtual camera pose optimization model.

[0142] Optionally, the reference video and the rendered video optical flow map, joint point heat map and depth map are subtracted to calculate the difference, and the difference map of each feature is normalized using Min-Max normalization. Finally, the normalized results are added together to obtain the final guidance area.

[0143] In some embodiments, determining the guide area using a preset guide area determination model based on the optical flow map, character joint heat map, and depth map of the reference video and the rendered video includes:

[0144] Determine a first feature difference map based on the optical flow maps of the reference video and the rendered video; determine a second feature difference map based on the heat maps of the character joints of the reference video and the rendered video; and determine a third feature difference map based on the depth maps of the reference video and the rendered video;

[0145] The guide area is determined by the preset guide area determination model according to the first feature difference map, the second feature difference map, and the third feature difference map.

[0146] In some embodiments, the preset guidance area determination model includes:

[0147]

[0148] in,( - ) is the second characteristic difference map, To render the heat map of the joint points of the video frame, is the joint point heat map of the reference video frame; ( - ) is the first characteristic difference map, To render the optical flow map of the video frame, is the optical flow map of the reference video frame; ( - ) is the third characteristic difference map, To render the depth map of the video frame, is the depth map of the reference video frame.

[0149] During backpropagation, the camera parameter optimization process requires calculating the gradient of the loss function with respect to all MLP (Multi-Layer Perceptron) parameters. This requires considering all MLP parameters for every pixel, resulting in increased computational complexity and significant memory usage. The downstream network requires a complete image as input, making it impossible to reduce computational complexity by sampling pixels. Furthermore, some losses are unevenly distributed across the image, particularly joint loss, where many background areas provide no valid information. This uneven distribution results in backpropagating the gradients of pixels in uninformative areas, which can either reduce the model's optimization performance or mislead the optimization direction.

[0150] In order to solve the above problems, this embodiment proposes a guidance area calculated based on optical flow, depth and joint point heat maps. In the final gradient backpropagation process, only the gradients of some important pixels will be backpropagated. This method not only avoids destroying the end-to-end training of the model, but also reduces computational and memory overhead, while making the model more focused on pixels that contribute to the loss function without wasting resources on invalid pixels, thereby improving the efficiency and performance of the model. The guidance map G is calculated by the joint difference map between the generated joint point heat map, the optical flow map and the depth map to highlight the key areas that contribute more to the gradient (the preset condition can be set to a larger gradient contribution). Selectively backpropagating the gradients of these pixels can reduce memory consumption and provide a mechanism similar to attention to help convergence and focus on information-rich areas. As shown in the formula:

[0151]

[0152] in, Represent the joint point heat map, optical flow map and depth map respectively. represents the absolute value, and is the standard Min-Max normalization, which is used to convert The value is mapped to [0-1]:

[0153]

[0154] like Figure 7 As shown, Figure 7 This is a schematic diagram of the result of the guide area. In the visualization of the result of the guide area, black pixels represent pixels with less information that are not back-propagated by the gradient, while white pixels represent pixels with more information that are back-propagated by the gradient.

[0155] Step 150: For any pixel point of the video frame in the rendered video, determine a first gradient according to the optical flow loss function, determine a second gradient according to the joint point loss function, and determine a third gradient according to the depth loss function; and determine a first return gradient corresponding to the first gradient, a second return gradient corresponding to the second gradient, and a third return gradient corresponding to the third gradient according to the guide area, and return the first return gradient, the second return gradient, and the third return gradient to optimize the shooting parameters of the virtual camera.

[0156] During the final gradient return, the gradients of all pixels, calculated by derivatizing the loss function, are multiplied by the gradients of the guidance region. The gradients of some pixels are returned, meaning only the gradient values ​​of pixels in the information-rich region are returned to optimize the camera's shooting parameters. Repeat these steps to iteratively optimize the camera's shooting parameters.

[0157] In some embodiments, determining the first gradient according to the optical flow loss function, determining the second gradient according to the joint point loss function, and determining the third gradient according to the depth loss function includes:

[0158] Performing a derivative operation on the optical flow loss function to obtain the first gradient;

[0159] Performing a derivative operation on the joint point loss function to obtain the second gradient;

[0160] Perform a derivative operation on the depth loss function to obtain the third gradient.

[0161] Optionally, the gradient of the loss is multiplied by the guidance region to obtain the final partial pixel gradient value to be returned. This partial pixel gradient value is then used for backpropagation to optimize the camera parameters. After multiple rounds of iterative optimization, the camera parameters and pose parameters that are most consistent with the reference video are finally obtained.

[0162] In some embodiments, determining, according to the guide area, a first return gradient corresponding to the first gradient, a second return gradient corresponding to the second gradient, and a third return gradient corresponding to the third gradient, respectively, includes:

[0163] Taking the dot product of the guide region and the first gradient as the first backpropagation gradient;

[0164] Taking the dot product of the guide region and the second gradient as the second backpropagation gradient;

[0165] A dot product result of the guide region and the third gradient is used as the third return gradient.

[0166] After obtaining the guide region G, G acts as a "mask". The gradient obtained by the loss is multiplied with the guide region G to obtain the final returned pixel gradient value:

[0167]

[0168] in, Represents the gradient of the returned pixel containing important information, represents the gradient of all pixels, G represents the guide area, Indicates a dot product operation. The Adam optimization algorithm is used to optimize the shooting parameters.

[0169] Furthermore, the above steps 120 to 150 are repeated to iteratively optimize the shooting parameters of the camera.

[0170] After multiple rounds of iterative optimization, the camera's shooting parameters gradually converged to their optimal values. In each iteration, the pixel gradients calculated by derivatizing the loss function are multiplied by the guide region, ensuring that only the important pixel gradients are transferred and updated, preserving key details and structure in the image. This refined gradient transfer method makes the optimization process more efficient and better adapts to the shooting requirements of different scenes and complex lighting conditions.

[0171] Further, in order to facilitate understanding of the technical solution of the application, you can also refer to Figure 2 and Figure 3 .in, Figure 2 is a flowchart of another virtual camera pose adjustment method according to an embodiment of the present application. Figure 3 A logical diagram of the virtual camera pose optimization generation method provided in an embodiment of the present application.

[0172] In the virtual camera pose optimization method disclosed in this embodiment, a video of a performance scene is first used as input to train a dynamic neural radiation field, implicitly express the virtual performance scene, and obtain a function expression model under the scene; then a randomly initialized virtual camera pose sequence is used as input, and the neural radiation field model is combined to obtain a rendered video under the camera pose trajectory; the reference video and the rendered video are used as input to extract the joint point heat map features, depth features, and optical flow features of the two respectively; the three shooting features of the reference video and the rendered video are used as input to calculate the loss between each feature of the reference video and the rendered video in turn; the three shooting features of the reference video and the rendered video are used as input to calculate the difference map of each feature of the reference video and the rendered video in turn, and the guide area is obtained after normalization, and the information-rich area is selected for gradient feedback. Finally, the gradient is calculated through the loss function, and the gradient is multiplied by the guide area to obtain the final gradient to be returned, and the gradient is returned to optimize the camera parameters.

[0173] In summary, in the virtual camera pose optimization method disclosed in this embodiment, a video of a performance scene is first given, and the virtual performance scene is implicitly expressed through dynamic neural radiation fields, so that rendered videos of different angles can be rendered according to different camera pose sequences. Then, three differentiable downstream networks are used to extract the optical flow, character joint heat map, and depth features of the reference video and the rendered video respectively, which are used to represent the motion information, composition information, and camera perspective information of the camera and the character. The loss between the features is calculated, and the gradient is back-propagated to optimize the camera pose parameters. This ensures that the final neural-rendered video and the reference video input by the user (i.e., the preset reference video) have similar shooting techniques. In the gradient return, by calculating the guidance area, pixels with more information are selectively returned, reducing computational overhead.

[0174] Example 2

[0175] Based on the above embodiments, this embodiment discloses a specific example.

[0176] In this embodiment, the data set is selected using a video of a performance held offline in a theater and a virtual scene video corresponding to the performance program shot by a multi-angle camera array.

[0177] To intuitively demonstrate the performance of this method on existing datasets, the root mean square error (RMSE-ATE) of the absolute trajectory error and the root mean square error (RMSE-RPE) of the relative pose error of the final camera pose data are used as evaluation metrics for the accuracy of the captured camera pose. The relative pose error represents the error in the camera pose transformation between two frames within a fixed time interval, while the absolute trajectory error measures the overall similarity between two camera trajectories and the accuracy of pose changes between consecutive frames, representing the local accuracy.

[0178] According to the steps described in the above specific implementation manner, the experimental results obtained are shown in Table 1 below.

[0179] Table 1 Comparison of results between this application and existing methods

[0180]

[0181] As shown in Table 1, the present application has achieved higher performance on the existing dataset. It can be seen that the method based on implicit neural expression of the performance scene and the collaborative optimization of three features, optical flow, joint point heat map and depth, is more suitable for the optimization of camera pose in the performance scene. The effect achieved is relatively consistent with the reference video, which can significantly improve the quality of the generated camera pose sequence and reduce the workload of scene reconstruction.

[0182] In order to more clearly demonstrate the performance and usability of the virtual camera pose optimization method of this application, this application uses a virtual scene video as a test scene. Through the steps in the method disclosed in this application, the camera pose trajectory is visualized in the scene. The results can be referred to Figure 8 .

[0183] like Figure 8 As shown in the figure, the reference video is shot in a way that the camera gradually moves backward, while the camera in the virtual rendering video predicted by EIEFM also gradually moves away, which is consistent with the reference video.

[0184] Example 3

[0185] Based on the above embodiments, this embodiment provides a virtual camera posture adjustment system.

[0186] This system embodiment can be used to execute the method embodiment of this application. For details not disclosed in this system embodiment, please refer to the method embodiment of this application. The system disclosed in this embodiment includes:

[0187] A scene video rendering module is used to determine the implicit expression model corresponding to a preset reference video through a dynamic neural radiation field, and determine the rendered video under the current camera trajectory pose through the implicit expression model according to any camera trajectory pose in a preset camera trajectory pose set;

[0188] A feature estimation module is used to input the rendered video and the preset reference video into a preset virtual camera pose optimization model to respectively determine the optical flow map, character joint point heat map and depth map of each video frame in the preset reference video and the rendered video;

[0189] A loss function determination module is used to respectively determine an optical flow loss function corresponding to the optical flow map, a joint point loss function corresponding to the character joint point heat map, and a depth loss function corresponding to the depth map;

[0190] A guide region determination module is configured to determine a guide region using a preset guide region determination model based on the optical flow map, character joint heat map, and depth map of the preset reference video and the rendered video, so as to highlight a key region whose gradient contribution meets preset conditions;

[0191] A gradient feedback module is used to determine, for any pixel point of a video frame in the rendered video, a first gradient according to the optical flow loss function, a second gradient according to the joint point loss function, and a third gradient according to the depth loss function; and to determine, according to the guide area, a first feedback gradient corresponding to the first gradient, a second feedback gradient corresponding to the second gradient, and a third feedback gradient corresponding to the third gradient, and feedback the first feedback gradient, the second feedback gradient, and the third feedback gradient to optimize the shooting parameters of the virtual camera.

[0192] Those skilled in the art will appreciate that the modules or steps of the present application described above can be implemented using a general-purpose computing device, and can be centralized on a single computing device or distributed across a network of multiple computing devices. Alternatively, they can be implemented using program code executable by a computing device, so that they can be stored in a storage device and executed by the computing device. In some cases, the steps shown or described can be performed in a different order than that shown, or they can be fabricated into separate integrated circuit modules, or multiple modules or steps can be fabricated into a single integrated circuit module for implementation.

[0193] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of each module in the virtual camera pose adjustment system can refer to the corresponding process in the aforementioned method embodiment, and this embodiment will not be repeated here.

[0194] Example 4

[0195] Based on the above embodiments, this embodiment provides a computer-readable storage medium storing a computer program that, when executed by a processor, can implement the method steps in the above method embodiments, which will not be repeated in this embodiment.

[0196] Computer-readable storage media may also include computer programs, data files, data structures, and the like, alone or in combination. Computer-readable storage media or computer programs may be specifically designed and understood by those skilled in the computer software field, or they may be generally known and available to those skilled in the computer software field. Examples of computer-readable storage media include: magnetic media such as hard disks, floppy disks, and magnetic tapes; optical media such as CD-ROMs and DVDs; magneto-optical media such as optical disks; and hardware devices specifically configured to store and execute computer programs, such as read-only memory (ROM), random access memory (RAM), and flash memory; or servers and app stores. Examples of computer programs include machine code (e.g., code generated by a compiler) and files containing higher-level code that can be executed by a computer using an interpreter. The described hardware devices can be configured to function as one or more software modules to perform the operations and methods described above, and vice versa. Furthermore, computer-readable storage media can be distributed across networked computer systems, enabling the program code or computer program to be stored and executed in a decentralized manner.

[0197] Example 5

[0198] Based on the above embodiments, this embodiment provides a computer program product. The computer program product includes a computer program or instructions, which, when executed by a processor, implements all or part of the steps of the method in the above method embodiments, and will not be repeated in this embodiment.

[0199] Furthermore, the computer program product may include one or more computer executable components configured to perform the embodiments when the program is run; the computer program product may also include a computer program tangibly embodied on a computer-readable medium, the computer program including program code for performing any method in the embodiments of the present disclosure. In such an embodiment, the computer program may be downloaded and installed from a network via a communication component and / or installed from a removable medium.

[0200] Example 6

[0201] Based on the above embodiments, this embodiment provides an electronic device, which may include: one or more processors, a memory, a multimedia component, an input / output (I / O) interface, and a communication component.

[0202] The one or more processors are used to execute all or part of the steps in the aforementioned method embodiment. The memory is used to store various types of data, which may include instructions of any application or method in the electronic device, as well as application-related data.

[0203] One or more processors can be implemented as an application specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), a controller, a microcontroller, a microprocessor or other electronic components, and are used to execute the methods in the aforementioned method embodiments.

[0204] Memory can be implemented by any type of volatile or non-volatile memory device, or a combination of them, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0205] The multimedia component may include a screen and an audio component, wherein the screen may be a touch screen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone for receiving external audio signals. The received audio signals may be further stored in a memory or transmitted via a communication component. The audio component also includes at least one speaker for outputting audio signals.

[0206] The I / O interface provides an interface between one or more processors and other interface modules, such as a keyboard, a mouse, buttons, etc. These buttons can be virtual buttons or physical buttons.

[0207] The communication component is used for wired or wireless communication between the electronic device and other devices. Wired communication includes communication through network ports, serial ports, etc.; wireless communication includes: Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G, 4G, 5G, or one or a combination of these.

[0208] It should also be understood that the methods or systems disclosed in the embodiments provided in this application may also be implemented in other ways. The method or system embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings illustrate possible architectures, functions, and operations of the methods and systems according to various embodiments of the present application. In this regard, each box in the flowchart or block diagram may represent a module, a computer program segment, or a portion of a computer program, which contains one or more computer programs for implementing specified logical functions. It should also be noted that in some alternative implementations, the functions marked in the boxes may occur in an order different from that marked in the drawings, and may actually be executed substantially in parallel, or they may sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and combinations of boxes in the block diagram and / or flowchart, may be implemented using a dedicated hardware-based system that performs the specified functions or actions, or may be implemented using a combination of dedicated hardware and computer programs.

[0209] In this application, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed or that are inherent to such process, method, article, or apparatus. In the absence of further restrictions, the elements defined by the sentence "including a..." do not exclude the existence of other identical elements in the process, method, device or equipment including the elements; if there is a description of "first", "second", etc., it is only for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features or implicitly indicating the order of the indicated technical features; in the description of this application, unless otherwise specified, the terms "multiple" and "many" mean at least two; if there is a description of a server, it should be noted that the server can be an independent physical server or terminal, or a server cluster composed of multiple physical servers, or a cloud server that can provide basic cloud computing services such as cloud servers, cloud databases, cloud storage and CDN; if there is a description of a smart terminal or mobile device in this application, it should be noted that the smart terminal or mobile device can be a mobile phone, tablet computer, smart watch, netbook, wearable electronic device, personal digital assistant (PDA), augmented reality technology device (AR), virtual reality device (VR), smart TV, smart speaker, personal computer (PC) Computer, referred to as PC), etc., but not limited to this. This application does not specifically limit the specific form of the smart terminal or mobile device.

[0210] Finally, it should be noted that, in the description of this specification, reference to the terms "one embodiment," "some embodiments," "example," "an example," or "some examples" means that the specific features, structures, materials, or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any appropriate manner in any one or more embodiments or examples.

[0211] Although the embodiments of the present application have been shown and described above, it is understood that the above embodiments are exemplary and the contents described are only embodiments adopted to facilitate understanding of the present application and are not intended to limit the present application. Any person skilled in the art of the present application may make any modifications and changes in the form and details of the implementation without departing from the spirit and scope disclosed in the present application, but the scope of protection of the present application shall still be based on the scope defined by the appended claims.

Claims

1. A virtual camera posture adjustment method, characterized in that: The method comprises: Determine an implicit expression model corresponding to a preset reference video through a dynamic neural radiation field, and determine a rendered video under a current camera trajectory pose through the implicit expression model according to any camera trajectory pose in a preset camera trajectory pose set; Inputting the rendered video and the preset reference video into a preset virtual camera pose optimization model to respectively determine the optical flow map, character joint point heat map, and depth map of each video frame in the preset reference video and the rendered video; Respectively determining an optical flow loss function corresponding to the optical flow map, a joint point loss function corresponding to the character joint point heat map, and a depth loss function corresponding to the depth map; Determining a guide area using a preset guide area determination model based on the optical flow map, character joint heat map, and depth map of the preset reference video and the rendered video to highlight key areas whose gradient contribution meets preset conditions; For any pixel point of the video frame in the rendered video, a first gradient is determined according to the optical flow loss function, a second gradient is determined according to the joint point loss function, and a third gradient is determined according to the depth loss function; and a first return gradient corresponding to the first gradient, a second return gradient corresponding to the second gradient, and a third return gradient corresponding to the third gradient are respectively determined according to the guide area, and the first return gradient, the second return gradient, and the third return gradient are returned to optimize the shooting parameters of the virtual camera.

2. The virtual camera posture adjustment method according to claim 1, characterized in that: The preset virtual camera pose optimization model includes: in, is the camera’s extrinsic matrix, For time, is the focal length of the camera, () is an implicit expression model, () is the overall loss of the implicit expression model, Θ is the implicit expression model Model parameters in (), is the video frame of the reference video; Where i is the video frame number of the reference video, The video frame for the rendered video.

3. The virtual camera posture adjustment method according to claim 1, characterized in that: The preset guidance area determination model includes: in,( - ) is the second characteristic difference map, To render the heat map of the joint points of the video frame, is the joint point heat map of the reference video frame; ( - ) is the first characteristic difference map, To render the optical flow map of the video frame, is the optical flow map of the reference video frame; ( - ) is the third characteristic difference map, To render the depth map of the video frame, is the depth map of the reference video frame; Indicates normalization processing.

4. The virtual camera posture adjustment method according to claim 1, characterized in that: The determining of the first gradient according to the optical flow loss function, the determining of the second gradient according to the joint point loss function, and the determining of the third gradient according to the depth loss function includes: Performing a derivative operation on the optical flow loss function to obtain the first gradient; Performing a derivative operation on the joint point loss function to obtain the second gradient; Perform a derivative operation on the depth loss function to obtain the third gradient.

5. The virtual camera posture adjustment method according to claim 1, characterized in that: The determining, according to the guide area, a first return gradient corresponding to the first gradient, a second return gradient corresponding to the second gradient, and a third return gradient corresponding to the third gradient, respectively, includes: Taking the dot product of the guide region and the first gradient as the first backpropagation gradient; Taking the dot product of the guide region and the second gradient as the second backpropagation gradient; A dot product result of the guide region and the third gradient is used as the third return gradient.

6. The virtual camera posture adjustment method according to claim 1, characterized in that: The optical flow loss function includes: in, and Represent the horizontal optical flow component and vertical optical flow component of the rendered video respectively, and Represent the horizontal optical flow component and vertical optical flow component of the reference video respectively.

7. The virtual camera posture adjustment method according to claim 1, characterized in that: The joint point loss function includes: in, is the Wasserstein distance, and Represents the character joint heat maps of the rendered video frame and the reference video frame respectively, Represents the value of the i-th joint point in the character joint point heat map, and M is the number of joint points.

8. The virtual camera posture adjustment method according to claim 1, characterized in that: The depth loss function includes: in, is the number of pixels in the video frame, and Represents the rendered video frame and the reference video frame, respectively. The depth value of each pixel, H is the height of the reference video frame, and W is the width of the reference video frame.

9. A virtual camera posture adjustment system, characterized in that: include: A scene video rendering module is used to determine the implicit expression model corresponding to a preset reference video through a dynamic neural radiation field, and determine the rendered video under the current camera trajectory pose through the implicit expression model according to any camera trajectory pose in a preset camera trajectory pose set; A feature estimation module is used to input the rendered video and the preset reference video into a preset virtual camera pose optimization model to respectively determine the optical flow map, character joint point heat map and depth map of each video frame in the preset reference video and the rendered video; A loss function determination module is used to respectively determine an optical flow loss function corresponding to the optical flow map, a joint point loss function corresponding to the character joint point heat map, and a depth loss function corresponding to the depth map; A guide region determination module is configured to determine a guide region using a preset guide region determination model based on the optical flow map, character joint heat map, and depth map of the preset reference video and the rendered video, so as to highlight a key region whose gradient contribution meets preset conditions; A gradient feedback module is used to determine, for any pixel point of a video frame in the rendered video, a first gradient according to the optical flow loss function, a second gradient according to the joint point loss function, and a third gradient according to the depth loss function; and to determine, according to the guide area, a first feedback gradient corresponding to the first gradient, a second feedback gradient corresponding to the second gradient, and a third feedback gradient corresponding to the third gradient, and feedback the first feedback gradient, the second feedback gradient, and the third feedback gradient to optimize the shooting parameters of the virtual camera.

10. A computer program product comprising a computer program or instructions, characterized in that: When the computer program or instructions are executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Neural radiation field enhancement method based on joint pose optimization

    CN112613609A

  • Three-dimensional human body reconstruction method based on sequential context clues

    CN115330950A