Model training method, device and medium

By using the deviations of virtual streams and RGB optical streams in the NeRF model for training, the problem of model training under unknown acquisition device parameters is solved, and accurate training of the NeRF model and effective estimation of acquisition device parameters are achieved.

CN120107742APending Publication Date: 2025-06-06HISENSE GRP HLDG CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311653792.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-04
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The prior art cannot effectively realize model training when the parameters of the acquisition device are unknown, especially when accurate parameters of the acquisition device are required, the current joint estimation method cannot accurately complete the task.

Method used

By receiving multiple images and inputting each image and its adjacent images into the initial NeRF model, predicted parameters are obtained, and then projecting each pixel in the image into the adjacent image based on these parameters, a virtual stream matrix and an RGB optical flow matrix are generated, and the initial NeRF model is trained by the deviations of these matrices.

Benefits of technology

This method can accurately train the NeRF model when the acquisition device parameters are unknown, reduce the dependence on the acquisition device parameters, improve the quality of new viewpoint synthesis and the accuracy of acquisition device parameter estimation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107742A_ABST
    Figure CN120107742A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a model training method, equipment and a medium, which are used for solving the problem that model training cannot be realized when parameters of acquisition equipment are unknown in the prior art. In the embodiment of the invention, the electronic equipment predicts the parameters of the acquisition equipment and the adjacent acquisition equipment of the acquisition equipment through the initial NeRF model, and in the specific training process, when the parameters of the acquisition equipment and the adjacent acquisition equipment of the acquisition equipment are accurate, the deviation between the RGB optical flow matrix and the virtual flow matrix is relatively small, so that the accuracy of the data processing is improved. According to the method, the initial NeRF model is trained according to the deviation, so that the NeRF model can determine the parameters of the acquisition equipment and the adjacent acquisition equipment while learning the scene information, and the training of the initial NeRF model is accurately realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of model training technology, and in particular to a model training method, device and medium. Background Art

[0002] Neural Radiance Fields (NeRF) does not require an intermediate 3D reconstruction process, and synthesizes images from a new perspective based only on the parameters and images of the acquisition device. This demonstrates the powerful ability of 3D scene representation and high-fidelity new view synthesis, and is widely used in virtual reality, augmented reality, 3D content generation and creation, real-time positioning and mapping technology, etc. In order to achieve the synthesis of images from a new perspective, it is usually necessary to first implement NeRF reconstruction. Specifically, it is first necessary to implement the training of the NeRF model so that the NeRF model can fully understand the scene information. A key priori condition for the training of the NeRF model is the reliable calibration of the acquisition device parameters. However, in related technologies, accurate acquisition device parameters are usually not so easy to obtain.

[0003] In order to reduce the dependence on the parameters of the acquisition device, many methods have been proposed in the relevant technology to jointly estimate the parameters of the acquisition device and the scene reconstruction. Among them, an end-to-end pipeline is proposed for jointly evaluating the parameters of the acquisition device and the NeRF model. Subsequent work further aligns the connection between the NeRF model and the 2D plane image, and proposes a coarse-to-fine position encoding to improve the large error accumulation of the acquisition device parameter estimation in the early stage of joint optimization. By using different activation functions in the multilayer perceptron (MLP) of the NeRF model, the gradient optimization strategy when the photometric loss is derived from the parameters of the acquisition device is changed, further reducing the difficulty of joint optimization. However, these all require the known parameters of the real acquisition device or the known parameters that are close to the parameters of the real acquisition device. When the parameters are not accurate enough, the current joint estimation methods cannot accurately complete the joint estimation task.

[0004] Related technologies also use the idea of ​​generative adversarial networks to reduce this limitation through the idea of ​​adversarial learning, but the current method still requires a known distribution of posture sampling, which is equivalent to still being unable to completely get rid of the parameters of the acquisition device. Recently, the monocular depth map has been integrated into the joint estimation to optimize the geometric prediction of the NeRF model, thereby better constraining the acquisition device parameters. However, these methods have high requirements for the geometric optimization of the scene, and the current methods still cannot restore enough scene geometry to accurately estimate the acquisition device parameters. As a result, it is impossible to obtain a more accurate NeRF model. Summary of the invention

[0005] The embodiments of the present application provide a model training method, device and medium to solve the problem in the prior art that the model cannot be trained when the parameters of the acquisition device are unknown.

[0006] In a first aspect, an embodiment of the present application provides a model training method, the method comprising:

[0007] Receive multiple images collected for the same scene; for each image, input the image and adjacent images collected by adjacent collection devices of the collection device corresponding to the image into an initial neural radiation field NeRF model, and obtain a first prediction parameter of the collection device and a second prediction parameter of the adjacent collection device output by the initial NeRF model;

[0008] For each image, project each pixel point in the image into the adjacent image according to the first prediction parameter and the second prediction parameter, generate a virtual flow matrix according to the deviation between the position of the same pixel point in the adjacent image after projection and the position of the pixel point in the adjacent image; and generate an RGB optical flow matrix according to the position difference between the same pixel point in the image and the adjacent image; determine the deviation between the virtual flow matrix and the RGB optical flow matrix;

[0009] The initial NeRF model is trained according to the deviation.

[0010] In a second aspect, an embodiment of the present application further provides an electronic device, which includes at least a processor and a memory, and the processor is used to implement the steps of model training as described in any of the above items when executing a computer program stored in the memory.

[0011] In a third aspect, an embodiment of the present application further provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of model training as described in any of the above items.

[0012] In an embodiment of the present application, an electronic device receives multiple images captured for the same scene; for each image, the image and adjacent images captured by adjacent acquisition devices of the acquisition device corresponding to the image are respectively input into an initial neural radiation field NeRF model to obtain a first prediction parameter of the acquisition device and a second prediction parameter of the adjacent acquisition device output by the initial NeRF model; for each image, each pixel point in the image is projected into an adjacent image according to the first prediction parameter and the second prediction parameter, and a virtual flow matrix is ​​generated according to the deviation between the position of the same pixel point in the adjacent image after projection and the position of the pixel point in the adjacent image; and an RGB optical flow matrix is ​​generated according to the position difference of the same pixel point in the image and the adjacent image; the deviation between the virtual flow matrix and the RGB optical flow matrix is ​​determined; and the initial NeRF model is trained according to the deviation. In an embodiment of the present application, the electronic device predicts the parameters of the acquisition device and its adjacent acquisition devices through an initial NeRF model, and in a specific training process, when the parameters of the acquisition device and its adjacent acquisition devices are accurate, the deviation between the RGB optical flow matrix and the virtual flow matrix is ​​small. In the present application, the initial NeRF model is trained according to the deviation, so that the NeRF model can determine the parameters of the acquisition device and its adjacent acquisition devices while learning the scene information, and accurately implement the training of the initial NeRF model. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0014] Figure 1 A schematic diagram of a model training process provided in an embodiment of the present application;

[0015] Figure 2 A schematic diagram of a process for determining a deviation provided in an embodiment of the present application;

[0016] Figure 3 A schematic diagram of a posture optimization process provided in an embodiment of the present application;

[0017] Figure 4 A schematic diagram of a sampling area determined in an embodiment of the present application;

[0018] Figure 5 A schematic diagram of the effect of a virtual view provided in an embodiment of the present application;

[0019] Figure 6 A schematic diagram of an adjusted sampling area provided in an embodiment of the present application;

[0020] Figure 7 A detailed process diagram of a model training provided in an embodiment of the present application;

[0021] Figure 8 A schematic diagram of an overall process provided for an embodiment of the present application;

[0022] Fig. 9 A schematic diagram of the structure of a model training device provided in an embodiment of the present application;

[0023] Fig.10 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0024] The present application will be further described in detail below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present application.

[0025] In order to accurately and effectively implement model training when the parameters of the acquisition device are unknown, the embodiments of the present application provide a model training method, device and medium.

[0026] The model training method includes: an electronic device receives a plurality of images collected for the same scene; for each image, the image and adjacent images collected by adjacent collection devices of the collection device corresponding to the image are respectively input into an initial neural radiation field NeRF model, and a first prediction parameter of the collection device output by the initial NeRF model and a second prediction parameter of the adjacent collection device are obtained; for each image, each pixel point in the image is projected into an adjacent image according to the first prediction parameter and the second prediction parameter, and a virtual flow matrix is ​​generated according to the deviation between the position of the same pixel point in the adjacent image after projection and the position of the pixel point in the adjacent image; and an RGB optical flow matrix is ​​generated according to the position difference of the same pixel point in the image and the adjacent image; the deviation between the virtual flow matrix and the RGB optical flow matrix is ​​determined; and the initial NeRF model is trained according to the deviation.

[0027] Figure 1 A schematic diagram of a model training process provided in an embodiment of the present application, the process includes the following steps:

[0028] S101: Receive multiple images captured for the same scene; for each image, input the image and adjacent images captured by adjacent acquisition devices of the acquisition device corresponding to the image into an initial neural radiation field NeRF model, and obtain the first prediction parameter of the acquisition device and the second prediction parameter of the adjacent acquisition device output by the initial NeRF model.

[0029] The model training method provided in the embodiment of the present application is applied to an electronic device, which may be a smart device such as a PC or a server.

[0030] In order to enable the NeRF model to learn scene information, the electronic device may first receive multiple images collected for the same scene. The multiple images may be collected by collection devices set at different positions, or may be collected by a user holding the same collection device around the scene.

[0031] The electronic device can input each received image and adjacent images acquired by adjacent acquisition devices of the acquisition device that acquires the image into an initial NeRF model, and obtain the output of the initial NeRF model. The output of the initial NeRF model includes prediction parameters of the acquisition device corresponding to the image and prediction parameters of adjacent acquisition devices of the acquisition device corresponding to the image. For ease of distinction, the prediction parameters of the acquisition device can be referred to as first preset parameters, and the prediction parameters of the adjacent acquisition device can be referred to as second preset parameters.

[0032] Among them, if the current application scenario is that the user holds the same acquisition device to capture multiple images around a certain scene, then the predicted parameters of the acquisition device corresponding to the image can be considered to be the parameters of the predicted acquisition device when capturing the image, and the predicted parameters of the adjacent acquisition device can be considered to be the parameters of the predicted adjacent acquisition devices when capturing adjacent images.

[0033] In order to accurately perform model training, after receiving multiple images, the electronic device can first initialize the parameters of the NeRF model and the parameters of the acquisition device. It should be noted that the initialization of the parameters of the acquisition device at this time refers to randomly setting the parameters of the acquisition device to a parameter or setting them to a preset parameter. Among them, this step can be implemented by the image and model initialization module of the electronic device.

[0034] S102: For each image, project each pixel point in the image into the adjacent image according to the first prediction parameter and the second prediction parameter, generate a virtual flow matrix according to the deviation between the position of the same pixel point in the adjacent image after projection and the position of the pixel point in the adjacent image; and generate an RGB optical flow matrix according to the position difference between the same pixel point in the image and the adjacent image; determine the deviation between the virtual flow matrix and the RGB optical flow matrix.

[0035] In order to accurately predict the model, after the electronic device obtains the first prediction parameter and the second prediction parameter, the electronic device can project each pixel point in the image to the adjacent image for each received image according to the obtained first prediction parameter and the second prediction parameter. Specifically, when the parameters of the acquisition device of the image and the parameters of the adjacent acquisition device corresponding to the adjacent image are known, how to project each pixel point in the image to the adjacent image is a prior art and will not be described in detail here.

[0036] For each image, after projecting each pixel in the image into the adjacent image, the electronic device can generate a virtual flow matrix based on the position of the same pixel in the adjacent image after projection and the deviation of the position of the pixel in the adjacent image. Specifically, for each pixel, the electronic device can determine the deviation between the position of the pixel projected into the adjacent image and the horizontal coordinate and vertical coordinate of the position of the pixel in the adjacent image, and generate the corresponding virtual flow matrix based on the deviation determined for each pixel; and generate the RGB optical flow matrix based on the position difference of the same pixel in the image and the adjacent image.

[0037] For each image, after obtaining the virtual flow matrix and the RGB optical flow matrix of the image, the electronic device can determine the deviation between the virtual flow matrix and the RGB optical flow matrix. In a possible implementation, the electronic device can determine a difference matrix between the virtual flow matrix and the RGB optical flow matrix, and determine the norm of the difference matrix, and determine the norm as the deviation between the virtual flow matrix and the RGB optical flow.

[0038] S103: Training the initial NeRF model according to the deviation.

[0039] When the first prediction parameter and the second prediction parameter are accurate, the deviation between the virtual flow matrix and the RGB optical flow is small. Therefore, if the deviation between the virtual flow matrix and the RGB optical flow is small, it can be said that the first prediction parameter and the second prediction parameter are relatively accurate. Therefore, the electronic device can train the initial NeRF model according to the deviation between the virtual flow matrix and the RGB optical flow.

[0040] The key challenge of this application lies in the pose geometry ambiguity of the parameter-free NeRF model. The NeRF model tends to overfit the 3D scene for realistic RGB image restoration, but lacks explicit geometric learning, while accurate camera pose estimation relies on rich geometric guidance. During joint optimization, the geometric ambiguity of the NeRF model leads to gradient uncertainty in pose optimization, and inaccurate pose estimation leads to poor NeRF reconstruction based on the NeRF model. This ambiguity will increase if the scene is a scene where the range of camera movement is too large. In order to reduce the dependence of the NeRF model on the quality of geometric learning, this application proposes to train the NeRF model based on the deviation between the virtual flow and the RGB optical flow. In this method, the directional information embedded in the two-dimensional optical flow is processed to provide direct guidance for pose optimization. Specifically, the RGB optical flow is constrained to be consistent with the virtual flow to mine geometric clues across views.

[0041] The embodiment of the present application is a technology for joint estimation of the pose estimation and implicit neural radiation field of the acquisition device based on deep learning. It should be noted that the trained NeRF model can also output relatively accurate parameters of the acquisition device, and then accurately learn the scene information to achieve NeRF reconstruction. The present application can simultaneously reconstruct the scene and estimate the parameters of the acquisition device corresponding to the shooting angle, the parameters including pose information, without the need to use the previous offline acquisition device calibration method, and the present application can be applied to scenes with large motion amplitudes of the acquisition device, wherein the scene with large motion amplitudes of the acquisition device refers to the scene where the user holds the same acquisition device around a scene to collect multiple images, and compared with the related technology method, there is no need for more accurate initialization of acquisition device parameters. The present application gets rid of the current mainstream camera pose-scene expression joint estimation algorithm that relies on a strong pose prior, and can handle situations where the motion amplitude of the acquisition device is large and the scene information is more complex and broad. Since the present application can accurately implement the training of the NeRF model, the quality of new viewpoint synthesis can be improved in the future, and the accuracy of parameter estimation of the acquisition device is also greatly improved.

[0042] In an embodiment of the present application, the electronic device predicts the parameters of the acquisition device and its adjacent acquisition devices through an initial NeRF model, and in a specific training process, when the parameters of the acquisition device and its adjacent acquisition devices are accurate, the deviation between the RGB optical flow matrix and the virtual flow matrix is ​​small. In the present application, the initial NeRF model is trained according to the deviation, so that the NeRF model can determine the parameters of the acquisition device and its adjacent acquisition devices while learning the scene information, and accurately implement the training of the initial NeRF model.

[0043] In order to accurately project the pixel point into the adjacent image, based on the above embodiment, in the embodiment of the present application, according to the first prediction parameter and the second prediction parameter, each pixel point in the image is projected in front of the adjacent image, and the method further includes:

[0044] Obtain a predicted depth map corresponding to the image output by the initial NeRF model;

[0045] The projecting each pixel point in the image onto the adjacent image according to the first prediction parameter and the second prediction parameter comprises:

[0046] For each pixel in the image, the following formula is used to determine the target position of the pixel after projection when the pixel is projected into the adjacent image:

[0047]

[0048] in, is the horizontal coordinate of the target position, is the ordinate of the target position, K i+1 is the internal parameter matrix composed of the internal parameters in the second prediction parameter according to the first preset method, T i+1 is the external parameter matrix composed of the external parameters in the second prediction parameter according to the second preset method, K i The internal parameters in the first prediction parameter are an internal parameter matrix composed according to a first preset method, is a depth matrix composed of the depth values ​​of each pixel in the predicted depth map, u k is the horizontal coordinate of the pixel point in the image, v k is the vertical coordinate of the position of the pixel in the image.

[0049] In order to accurately project the pixel points into the adjacent image, the electronic device may also obtain a predicted depth map corresponding to the image output by the initial NeRF model.

[0050] After obtaining the predicted depth map corresponding to the image, the electronic device can use the following formula to determine the target position of each pixel in the image when the pixel is projected into the adjacent image:

[0051]

[0052] in, is the horizontal coordinate of the target position, is the ordinate of the target position, K i+1 is the internal parameter matrix composed of the internal parameters in the second prediction parameter according to the first preset method, T i+1is the external parameter matrix composed of the external parameters in the second prediction parameter according to the second preset method, K i The internal parameter matrix of the first prediction parameter is composed of the internal parameters in the first preset manner. is the depth matrix composed of the depth values ​​of each pixel in the predicted depth map, u k is the horizontal coordinate of the pixel point in the image, v k is the vertical coordinate of the position of the pixel in the image.

[0053] In an embodiment of the present application, the electronic device is equivalent to directly using the posture and rendering depth of the acquisition device to perform homography transformation in the calculation of the virtual flow, wherein the posture of the acquisition device refers to the external parameters of the acquisition device, and the rendering depth refers to the depth value of the image.

[0054] In order to accurately determine the virtual flow matrix and the RGB optical flow matrix, on the basis of the above embodiments, in the embodiment of the present application, the adjacent acquisition device includes a first adjacent acquisition device adjacent to the acquisition device in the first preset direction and a second adjacent acquisition device adjacent to the acquisition device in the second preset direction;

[0055] The generating of the virtual flow matrix according to the deviation between the position of the same pixel point in the adjacent image after projection and the position of the pixel point in the adjacent image comprises:

[0056] Generate a first virtual flow matrix based on the deviation between the position of the same pixel point in the first adjacent image collected by the first adjacent acquisition device after projection and the position of the pixel point in the first adjacent image; generate a second virtual flow matrix based on the deviation between the position of the same pixel point in the second adjacent image collected by the second adjacent acquisition device after projection and the position of the pixel point in the second adjacent image; and generate a third virtual flow matrix based on the deviation between the position of the same pixel point in the first adjacent image after projection and the position of the pixel point in the second adjacent image;

[0057] Generating an RGB optical flow matrix according to the position difference between the same pixel in the image and the adjacent image comprises:

[0058] A first RGB optical flow matrix is ​​generated based on the position difference between the same pixel point in the image and the first adjacent image; a second RGB optical flow matrix is ​​generated based on the position difference between the same pixel point in the image and the second adjacent image; and a third RGB optical flow matrix is ​​generated based on the position difference between the same pixel point in the first adjacent image and the second adjacent image.

[0059] In an embodiment of the present application, adjacent acquisition devices include adjacent acquisition devices adjacent to the acquisition device in a first preset direction and adjacent acquisition devices adjacent to the acquisition device in a second preset direction. For ease of distinction, the adjacent acquisition devices adjacent to the first preset direction may be referred to as first adjacent acquisition devices, and the adjacent acquisition devices adjacent to the second preset direction may be referred to as second adjacent acquisition devices.

[0060] And when generating a virtual flow matrix, the electronic device can generate a virtual flow matrix based on the deviation between the position of the same pixel point in the first adjacent image collected by the first adjacent acquisition device after projection and the position of the pixel point in the first adjacent image. For the sake of distinction, the virtual flow matrix is ​​called the first virtual flow matrix; the virtual flow matrix is ​​generated based on the deviation between the position of the same pixel point in the second adjacent image collected by the second adjacent acquisition device after projection and the position of the pixel point in the second adjacent image. For the sake of distinction, the virtual flow matrix is ​​called the second virtual flow matrix; and the virtual flow matrix is ​​generated based on the deviation between the position of the same pixel point in the first adjacent image after projection and the position of the pixel point in the second adjacent image. For the sake of distinction, the virtual flow matrix is ​​called the third virtual flow matrix.

[0061] When generating an RGB optical flow matrix, the electronic device may generate an RGB optical flow matrix based on the position difference between the same pixel point in the image and the first adjacent image. For the sake of distinction, the RGB optical flow matrix may be referred to as a first RGB optical flow matrix; generate an RGB optical flow matrix based on the position difference between the same pixel point in the image and the second adjacent image. For the sake of distinction, the RGB optical flow matrix may be referred to as a second RGB optical flow matrix; generate an RGB optical flow matrix based on the position difference between the same pixel point in the first adjacent image and the second adjacent image. For the sake of distinction, the RGB optical flow matrix may be referred to as a third RGB optical flow matrix.

[0062] In order to accurately train the model, based on the above embodiments, in the embodiment of the present application, determining the deviation between the virtual flow matrix and the RGB optical flow matrix includes:

[0063] Determine a first deviation between the first virtual flow matrix and the first RGB optical flow matrix, a second deviation between the second virtual flow matrix and the second RGB optical flow matrix, and a third deviation between the third virtual flow matrix and the third RGB optical flow matrix;

[0064] A corresponding deviation is determined according to the first deviation, the second deviation, the third deviation and their respective corresponding weights.

[0065] In order to accurately train the model, the electronic device can determine the deviation between the first virtual flow matrix and the first RGB optical flow matrix. For the sake of distinction, the deviation can be referred to as the first deviation. Determine the deviation between the second virtual flow matrix and the second RGB optical flow matrix. For the sake of distinction, the deviation can be referred to as the second deviation. Determine the deviation between the third virtual flow matrix and the third RGB optical flow matrix. For the sake of distinction, the deviation can be referred to as the third deviation.

[0066] After determining the first deviation, the second deviation, and the third deviation, the electronic device may determine the corresponding deviation according to the first deviation, the second deviation, the third deviation, and the weights corresponding to them. In a possible implementation, the weights corresponding to the first deviation, the second deviation, and the third deviation may be 0.4, 0.4, and 0.2, respectively.

[0067] The deviation may also be referred to as a single bidirectional optical flow-depth consistency loss. Specifically, calculating the first deviation, the second deviation, and the third deviation may also be understood as calculating the total optical flow depth consistency loss in a local window.

[0068] Electronic devices can calculate the corresponding deviation using the following formula:

[0069]

[0070] in, is the corresponding determined deviation, is the first deviation, w 1 is the weight corresponding to the first deviation, is the second deviation, w 2 is the weight corresponding to the second deviation, is the third deviation, w 3 is the weight corresponding to the third deviation.

[0071] Figure 2 A schematic diagram of a process for determining a deviation provided in an embodiment of the present application.

[0072] Figure 2 middle, I i The image described in the embodiment of the present application, I i-1 is the first adjacent image described in the embodiment of the present application, I i+1 It is the second adjacent image described in the embodiment of the present application. is a first deviation between a first virtual flow matrix determined based on the image and the first adjacent image and a first RGB optical flow matrix, is a second deviation between a second virtual flow matrix determined based on the image and the second adjacent image and a second RGB optical flow matrix, It is a third deviation between a third virtual flow matrix determined based on the first adjacent matrix image and the second adjacent image and a third RGB optical flow matrix.

[0073] In the embodiment of the present application, the disadvantage of deviation through the first adjacent image and the second adjacent image is equivalent to embedding the image frame in the RGB optical flow and the virtual flow. The embedded two-dimensional direction between the image frames can provide a clearer three-dimensional gradient direction for the posture optimization of the acquisition device.

[0074] Specifically, the electronic device can use an optical flow estimation network (Generalized Matching Flow, GMFlow) to estimate bidirectional optical flow between RGB level frames, and the bidirectional optical flow is the RGB optical flow matrix described in the embodiment of the present application. The occluded part is masked by running a forward and backward state check to prevent incorrect reprojection during learning. When the parameters of the acquisition device are accurate, the virtual flow matrix generated in the manner described in the embodiment of the present application is consistent with the RGB optical flow matrix. Based on this, the electronic device implements training of the NeRF model.

[0075] In order to accurately train the model, based on the above embodiments, in the embodiment of the present application, determining the deviation between the virtual flow matrix and the RGB optical flow matrix includes:

[0076] Determine a difference matrix obtained by subtracting the RGB optical flow matrix from the virtual flow matrix; and determine a target matrix obtained by multiplying the difference matrix by a preset matrix;

[0077] The deviation between the virtual flow matrix and the RGB optical flow matrix is ​​determined according to the ratio of the norm of the target matrix to the norm of the preset matrix.

[0078] In an embodiment of the present application, the electronic device uses a virtual flow matrix to subtract an RGB optical flow matrix, and obtains a difference matrix. After obtaining the difference matrix, the electronic device can determine the target matrix obtained by multiplying the difference matrix by a preset matrix. And according to the ratio of the norm of the target matrix to the norm of the preset matrix, the deviation of the virtual flow matrix from the RGB optical flow matrix is ​​determined. In a possible implementation, the electronic device can determine the ratio of the second norm of the target matrix to the first norm of the preset matrix, and determine the ratio as the deviation of the virtual flow matrix from the RGB optical flow matrix.

[0079] The electronic device can determine the deviation between the virtual flow matrix and the RGB optical flow matrix using the following formula:

[0080]

[0081] in, is the determined deviation, is the virtual flow matrix, Fi→i+1 is the RGB optical flow matrix, M i→i+1 is the preset matrix.

[0082] Among them, the deviation between the virtual flow matrix and the RGB optical flow matrix can be determined by the optical flow depth consistency module of the electronic device.

[0083] Figure 3 A schematic diagram of a posture optimization process provided in an embodiment of the present application.

[0084] Depend on Figure 3 It can be seen that there is an ambiguity problem in posture optimization. In this application, the posture optimization process of the NeRF model is guided based on three-dimensional virtual flow and two-dimensional optical flow (ie, RGB optical flow), based on which the NeRF model can identify the optimal posture.

[0085] Among them, the parameters of the acquisition device include posture.

[0086] In order to accurately train the model, based on the above embodiments, in the embodiment of the present application, the method further includes:

[0087] Obtain a predicted depth map corresponding to the image output by the initial NeRF model;

[0088] Identify the standard depth map corresponding to the image through a monocular depth estimation network (Depth Prediction from a Single Image using a Multi-path Refinement Network, DPT);

[0089] The training of the initial NeRF model according to the deviation comprises:

[0090] A depth map deviation is determined according to the predicted depth map and the standard depth map; and the initial NeRF model is trained according to the deviation and the depth map deviation.

[0091] In order to accurately train the model, the electronic device also obtains a predicted depth map corresponding to the image output by the initial NeRF model, and the electronic device identifies a standard depth map corresponding to the image through DPT.

[0092] After obtaining the standard depth map and predicted depth map corresponding to the image, the depth map deviation is determined. In one possible implementation, the electronic device can determine, for each depth value in the standard depth map, a predicted depth value corresponding to the depth value in the predicted depth map, and determine the deviation between the depth value and the corresponding predicted depth value, and determine the sum of the deviations of each determined depth value as the depth map deviation.

[0093] After determining the depth map deviation, the electronic device may train the initial NeRF model according to the depth map deviation and the deviation determined in the above embodiment.

[0094] The electronic device can determine the depth map deviation using the following formula:

[0095]

[0096] in, is the depth map deviation, N is the number of depth values ​​in the standard depth map and the predicted depth map, I i is the i-th depth value in the standard depth map, is the i-th depth value in the predicted depth map.

[0097] In order to accurately train the NeRF model, after obtaining the standard depth map, the electronic device can use adaptive monocular depth estimation to adjust the depth of NeRF rendering to enhance the representation of scene geometry and make it less prone to falling into local minima. Specifically, the electronic device generates a standard depth map and establishes learnable scale and offset parameters to linearly adjust the standard depth map, and then trains the initial NeRF model based on the adjusted standard depth map. Specifically, for each depth value in the standard depth map, the electronic device determines the product of the depth value and the scale, and determines the sum of the product and the offset parameter, and uses the sum to adjust the depth value. Among them, the scale and offset parameters can be pre-set.

[0098] In an embodiment of the present application, the electronic device's prior prediction module may utilize a pre-trained DPT to obtain a standard depth map, and utilize GMFlow to obtain an RGB optical flow to provide better prior information for the subsequent use.

[0099] The NeRF model internally represents the scene as a view-dependent mapping function and is parameterized by the MLP of the NeRF model. The synthetic image can be rendered by synthesizing the radiation color and density along the acquisition device ray r(t)=o+td between the near plane and the far plane. Volume rendering is expressed as: Where W(t) represents the cumulative transmittance, and the depth can be obtained by rendering the cumulative transmittance. Specifically, how to render and obtain the depth inside the model is an existing technology and will not be described in detail here.

[0100] In order to accurately train the model, based on the above embodiments, in the embodiment of the present application, after receiving multiple images collected for the same scene and before training the initial NeRF model according to the deviation, the method further includes:

[0101] For each image, the SuperPoint method is used to process the image, and each feature point with gesture perception characteristics in the image and its position in the image are obtained. For each feature point, a sampling area of ​​a preset size containing the feature point is determined; the position of the sampling point is randomly obtained in each obtained sampling area;

[0102] Inputting the image into an initial neural radiation field NeRF model, and obtaining a first prediction parameter of the acquisition device output by the initial NeRF model comprises:

[0103] Input the image and the position of each sampling point into the initial NeRF model, and obtain the first prediction parameter of the acquisition device and the predicted pixel value of each sampling point in the image output by the initial NeRF model;

[0104] The training of the initial NeRF model according to the deviation comprises:

[0105] A pixel deviation is determined according to the predicted pixel value and the standard pixel value of each sampling point in the image; and an initial NeRF model is trained according to the pixel deviation and the deviation.

[0106] Since the random ray sampling strategy is applied inside the NeRF model, this may introduce less anisotropic supervision in the low-texture area of ​​the image, further increasing the ambiguity of the parameters of the acquisition device, because the probability of sampling in the low-texture area is increased. In order to accurately train the model, the electronic device can obtain sampling points according to each feature point with gesture perception features. Specifically, the electronic device can process each image using the SuperPoint method to obtain each feature point with gesture perception features in the image and the position of each feature point in the image. For each feature point, determine a sampling area of ​​a preset size containing the feature point. Specifically, the electronic device can use the position of the feature point as the center to determine an area of ​​a preset size, which is the sampling area. After each sampling area is obtained, the electronic device can randomly obtain sampling points in the sampling area for each sampling area and determine the position of the sampling point in the image. It should be noted that when the scene geometry and the parameters of the acquisition device are unclear, each feature point obtained provides more effective supervision for parameter estimation.

[0107] Figure 4 A schematic diagram of a sampling area determined in an embodiment of the present application.

[0108] Depend on Figure 4 It can be seen that the electronic device can determine multiple sampling areas.

[0109] Specifically, the electronic device may input the image and the position of each sampling point into the initial NeRF model, and obtain the first prediction parameter of the acquisition device corresponding to the image output by the initial NeRF model and the predicted pixel value of each sampling point in the image.

[0110] The electronic device can also identify the image and obtain the standard pixel value of each sampling point in the image. After obtaining the standard pixel value and the predicted pixel value, the electronic device can determine the pixel deviation based on the predicted pixel value and the standard pixel value of each sampling point in the image. Specifically, the electronic device can determine the deviation between the predicted pixel value and the standard pixel value of each sampling point for each sampling point, and determine the sum of the deviations of each sampling point as the pixel deviation; after determining the pixel deviation, the electronic device can train the initial NeRF model based on the pixel deviation and the deviation determined above.

[0111] Figure 5 A schematic diagram of the effect of a virtual view provided in an embodiment of the present application.

[0112] Figure 5 The first column is the virtual view generated by NeRFmm. Figure 5 The second column is the virtual view generated by the BARF method. Figure 5 The third column is the virtual view generated by the NoPe-NeRF method. Figure 5 The fourth column is a virtual view generated based on the NeRF model of the embodiment of the present application. Figure 5 The fifth column is an image captured by a capture device set at a virtual viewing angle in a scenario for verifying the effect of the embodiment of the present application. Figure 5 As can be seen from the figures in the same row, the virtual view generated based on the NeRF model of the embodiment of the present application is closer to the real virtual view, more accurate, and has higher clarity.

[0113] In order to accurately train the model, based on the above embodiments, in the embodiment of the present application, after determining the sampling area of ​​the preset size containing the feature point, the method further includes:

[0114] For each sampling area, determining an adjustment size of the sampling area according to a ratio of the intensity of the light in the sampling area to the sum of the intensities of the light in each area in the image; and adjusting the size of the sampling area according to the adjustment size;

[0115] After iterating the initial NeRF model for a preset number of times, before randomly acquiring the position of a sampling point in each acquired sampling area, the method further includes:

[0116] The sampling area is replaced with the adjusted sampling area, and for the replaced sampling area, a subsequent step of randomly acquiring the position of the sampling point in each acquired sampling area is performed.

[0117] In order to accurately train the model, after determining each sampling area, the electronic device may determine the adjustment size of the sampling area according to the ratio of the intensity of the light in the sampling area to the sum of the intensities of the light in each area in the image, and adjust the size of the sampling area according to the adjustment size.

[0118] In a possible implementation, the electronic device may determine the resize based on a ratio of the luminance loss of the sampling area to the sum of the luminance losses of each area in the image. Specifically, how to determine the luminance loss of a certain area is a prior art and will not be described in detail here.

[0119] Electronic devices can be resized using the following formula:

[0120]

[0121] Among them, d max is the maximum adjustment range, d max Can be 5, σ can represent the sigmoid shape function, is the intensity of light in the sampling area, The sum of the light intensities for each area.

[0122] Assume that the size of the sampling area is d i ×d i , the electronic device can determine the adjusted sampling area through the following formula:

[0123]

[0124] in, is the adjusted sampling area, is the sampling area, d max is the maximum adjustment range, d max Can be 5, σ can represent the sigmoid shape function, is the intensity of light in the sampling area, The sum of the light intensities for each area.

[0125] After the initial NeRF model is iterated for a preset number of times, the initial NeRF model has a certain recognition ability. At this time, the electronic device can replace the adopted area with the adjusted sampling area, and for the replaced sampling area, perform the subsequent step of randomly obtaining the position of the sampling point in each acquired sampling area. The preset number of times can be 200.

[0126] It should be noted that the light sampled in the sampling area exhibits indistinguishable photometric errors. To solve this problem, this application proposes a new sampling strategy that adaptively adjusts the sampling area from the posture sensing position to the entire image, that is, adjusts the sampling area, thereby improving the learning ability of the model.

[0127] Among them, the electronic device adjusts the sampling area for subsequent random ray selection inside the NeRF model. Specifically, the electronic device expands the sampling area based on the contribution of the photometric loss to the total loss of all generated, because a smaller photometric error indicates that the 3D scene is fully learned along the ray, and the sampling area will be expanded to a larger area. After the expansion, random sampling of rays is performed to increase the complexity of subsequent training and enable NeRF to capture additional scene details. The operation can be performed by the adaptive ray sampling module of the electronic device.

[0128] After each preset number of iterations, the electronic device adjusts the sampling area again until different areas are combined to cover the entire image. Finally, the strategy will return to random sampling to increase light diversity. The method for obtaining the sampling area in this application can be referred to as obtaining the sampling area through the APAS light sampling strategy.

[0129] Figure 6 A schematic diagram of an adjusted sampling area provided in an embodiment of the present application.

[0130] Figure 6 For Figure 4 The schematic diagram of the sampling area after the sampling area adjustment is as follows: Figure 6 It can be seen that the size of each sampling area after adjustment may be different.

[0131] Figure 7 A schematic diagram of the detailed process of model training provided in an embodiment of the present application.

[0132] Depend on Figure 7 It can be seen that the electronic device can input the image into the initial NeRF model, obtain each feature point with posture perception features in the image through SuperPoint, obtain the sampling area through the APAS light sampling strategy, perform 3D point sampling on the image inside the initial NeRF model, extract information through MLP, process the extracted information through volume rendering, and then perform L rendering. The predicted depth map can be obtained through rendering.

[0133] In addition, the predicted depth corresponding to the image is determined through the depth model (i.e., the standard depth map), and the linearly adjusted depth (i.e., the adjusted standard depth map) is obtained. The depth loss (i.e., the depth map deviation) is determined based on the linearly adjusted depth and the depth obtained by rendering. The electronic device also obtains the bidirectional optical flow (i.e., the RGB optical flow matrix) through the optical flow model, and processes the depth obtained by rendering to obtain a virtual flow matrix. The deviation is obtained according to the optical flow depth consistency constraint between the virtual flow matrix and the RGB optical flow matrix, and the NeRF model is trained according to the depth loss and the deviation.

[0134] Figure 8 An overall flow chart of an embodiment of the present application is provided.

[0135] Depend on Figure 8 It can be seen that the technology mainly includes five parts: image and model initialization module, adaptive light sampling module, prior prediction module, optical flow depth consistency module and rendering and reconstruction module. Among them, the image and model initialization module is used to obtain images with multiple perspectives and high overlap when reconstructing the scene, and initialize the pose and NeRF model parameters for subsequent rendering training; the adaptive light sampling module uses the independently developed dynamic light sampling strategy to expand from the area with rich camera pose estimation features to the random area of ​​the entire image for light sampling, so as to improve the subsequent joint estimation; the prior prediction module uses the offline depth estimation and optical flow estimation pre-training models respectively, the depth estimation predicts the monocular depth value of each image, and the optical flow estimation predicts the optical flow matching between adjacent and cross-views; the optical flow depth consistency module constrains the RGB optical flow and virtual flow, and introduces the directional information of the RGB optical flow to provide direct guidance for pose estimation; the rendering and reconstruction module is used to render the image, so as to obtain the implicit three-dimensional representation of the final rendering of the new viewpoint. Specifically, after obtaining the NeRF model, how to render the image is an existing technology and will not be repeated here.

[0136] Fig. 9 A schematic diagram of the structure of a model training device provided in an embodiment of the present application, the device comprising:

[0137] The receiving and acquiring module 901 is used to receive multiple images collected for the same scene; for each image, the image and the adjacent images collected by the adjacent collection device of the collection device corresponding to the image are respectively input into the initial neural radiation field NeRF model, and the first prediction parameter of the collection device and the second prediction parameter of the adjacent collection device output by the initial NeRF model are obtained;

[0138] The processing module 902 is used for, for each image, projecting each pixel point in the image into the adjacent image according to the first prediction parameter and the second prediction parameter, generating a virtual flow matrix according to the deviation between the position of the same pixel point in the adjacent image after projection and the position of the pixel point in the adjacent image; and generating an RGB optical flow matrix according to the position difference between the same pixel point in the image and the adjacent image; and determining the deviation between the virtual flow matrix and the RGB optical flow matrix;

[0139] The training module 903 is used to train the initial NeRF model according to the deviation.

[0140] In a possible implementation, the processing module 902 is further configured to obtain a predicted depth map corresponding to the image output by the initial NeRF model;

[0141] The processing module 902 is specifically configured to determine, for each pixel in the image, a target position of the pixel after projection when the pixel is projected into the adjacent image using the following formula:

[0142]

[0143] in, is the horizontal coordinate of the target position, is the ordinate of the target position, K i+1 is the internal parameter matrix composed of the internal parameters in the second prediction parameter according to the first preset method, T i+1 is the external parameter matrix composed of the external parameters in the second prediction parameter according to the second preset method, K i The internal parameters in the first prediction parameter are an internal parameter matrix composed according to a first preset method, is the depth matrix composed of the depth values ​​of each pixel in the predicted depth map, u k is the horizontal coordinate of the pixel point in the image, v k is the vertical coordinate of the position of the pixel in the image.

[0144] In a possible implementation, the processing module 902 is specifically used for generating a first virtual flow matrix according to the deviation between the position of the same pixel point in the first adjacent image collected by the first adjacent collection device after projection and the position of the pixel point in the first adjacent image if the adjacent collection device includes a first adjacent collection device adjacent to the collection device in a first preset direction and a second adjacent collection device adjacent to the collection device in a second preset direction; generating a second virtual flow matrix according to the deviation between the position of the same pixel point in the second adjacent image collected by the second adjacent collection device after projection and the position of the pixel point in the second adjacent image; and generating a third virtual flow matrix according to the deviation between the position of the same pixel point in the first adjacent image after projection and the position of the pixel point in the second adjacent image; generating a first RGB optical flow matrix according to the position difference between the image and the same pixel point in the first adjacent image; generating a second RGB optical flow matrix according to the position difference between the image and the same pixel point in the second adjacent image; generating a third RGB optical flow matrix according to the position difference between the first adjacent image and the same pixel point in the second adjacent image.

[0145] In a possible implementation, the processing module 902 is specifically used to determine a first deviation between the first virtual flow matrix and the first RGB optical flow matrix, a second deviation between the second virtual flow matrix and the second RGB optical flow matrix, and a third deviation between the third virtual flow matrix and the third RGB optical flow matrix; and determine corresponding deviations according to the first deviation, the second deviation, the third deviation and their respective corresponding weights.

[0146] In one possible implementation, the processing module 902 is specifically used to determine a difference matrix obtained by subtracting the RGB optical flow matrix from the virtual flow matrix; and determine a target matrix obtained by multiplying the difference matrix by a preset matrix; and determine the deviation between the virtual flow matrix and the RGB optical flow matrix based on the ratio of the norm of the target matrix to the norm of the preset matrix.

[0147] In a possible implementation, the processing module 902 is further used to obtain a predicted depth map corresponding to the image output by the initial NeRF model; and identify a standard depth map corresponding to the image through a monocular depth estimation network DPT;

[0148] The training module 903 is specifically used to determine the depth map deviation according to the predicted depth map and the standard depth map; and train the initial NeRF model according to the deviation and the depth map deviation.

[0149] In a possible implementation, the processing module 902 is further configured to process each image using a SuperPoint method, obtain each feature point having a posture sensing feature in the image and its position in the image, and determine a sampling area of ​​a preset size containing the feature point for each feature point; and randomly obtain the position of a sampling point in each obtained sampling area;

[0150] The receiving and acquiring module 901 is specifically used to input the image and the position of each sampling point into the initial NeRF model, and acquire the first prediction parameter of the acquisition device output by the initial NeRF model and the predicted pixel value of each sampling point in the image;

[0151] The training module 903 is specifically used to determine the pixel deviation according to the predicted pixel value and the standard pixel value of each sampling point in the image; and train the initial NeRF model according to the pixel deviation and the deviation.

[0152] In a possible implementation, the processing module 902 is further configured to determine, for each sampling area, an adjustment size of the sampling area according to a ratio of the intensity of light in the sampling area to the sum of the intensities of light in each area in the image; and adjust the size of the sampling area according to the adjustment size;

[0153] The adjusted sampling area is used to replace the sampling area, and for the replaced sampling area, a subsequent step of randomly acquiring the position of the sampling point in each acquired sampling area is performed.

[0154] Fig.10 The present invention provides a schematic diagram of an electronic device structure according to an embodiment of the present invention. Based on the above embodiments, the present invention further provides an electronic device, such as Fig.10 As shown, it includes: a processor 1001, a communication interface 1002, a memory 1003 and a communication bus 1004, wherein the processor 1001, the communication interface 1002, and the memory 1003 communicate with each other through the communication bus 1004;

[0155] The memory 1003 stores a computer program. When the program is executed by the processor 1001, the processor 1001 performs the following steps:

[0156] Receive multiple images collected for the same scene; for each image, input the image and adjacent images collected by adjacent collection devices of the collection device corresponding to the image into an initial neural radiation field NeRF model, and obtain a first prediction parameter of the collection device and a second prediction parameter of the adjacent collection device output by the initial NeRF model;

[0157] For each image, project each pixel point in the image into the adjacent image according to the first prediction parameter and the second prediction parameter, generate a virtual flow matrix according to the deviation between the position of the same pixel point in the adjacent image after projection and the position of the pixel point in the adjacent image; and generate an RGB optical flow matrix according to the position difference between the same pixel point in the image and the adjacent image; determine the deviation between the virtual flow matrix and the RGB optical flow matrix;

[0158] The initial NeRF model is trained according to the deviation.

[0159] Furthermore, the processor 901 is further configured to obtain a predicted depth map corresponding to the image output by the initial NeRF model;

[0160] The processor 901 is specifically configured to determine, for each pixel in the image, a target position of the pixel after projection when the pixel is projected into the adjacent image using the following formula:

[0161]

[0162] in, is the horizontal coordinate of the target position, is the ordinate of the target position, K i+1 is the internal parameter matrix composed of the internal parameters in the second prediction parameter according to the first preset method, T i+1 is the external parameter matrix composed of the external parameters in the second prediction parameter according to the second preset method, K i The internal parameters in the first prediction parameter are an internal parameter matrix composed according to a first preset method, is the depth matrix composed of the depth values ​​of each pixel in the predicted depth map, u k is the horizontal coordinate of the pixel point in the image, v k is the vertical coordinate of the position of the pixel in the image.

[0163] Further, the processor 901 is specifically configured to: if the adjacent acquisition device includes a first adjacent acquisition device adjacent to the acquisition device in a first preset direction and a second adjacent acquisition device adjacent to the acquisition device in a second preset direction;

[0164] Generate a first virtual flow matrix based on the deviation between the position of the same pixel point in the first adjacent image collected by the first adjacent acquisition device after projection and the position of the pixel point in the first adjacent image; generate a second virtual flow matrix based on the deviation between the position of the same pixel point in the second adjacent image collected by the second adjacent acquisition device after projection and the position of the pixel point in the second adjacent image; and generate a third virtual flow matrix based on the deviation between the position of the same pixel point in the first adjacent image after projection and the position of the pixel point in the second adjacent image;

[0165] The processor 901 is specifically used to generate a first RGB optical flow matrix according to the position difference between the same pixel point in the image and the first adjacent image; generate a second RGB optical flow matrix according to the position difference between the same pixel point in the image and the second adjacent image; and generate a third RGB optical flow matrix according to the position difference between the same pixel point in the first adjacent image and the second adjacent image.

[0166] Further, the processor 901 is specifically configured to determine a first deviation between the first virtual flow matrix and the first RGB optical flow matrix, a second deviation between the second virtual flow matrix and the second RGB optical flow matrix, and a third deviation between the third virtual flow matrix and the third RGB optical flow matrix;

[0167] A corresponding deviation is determined according to the first deviation, the second deviation, the third deviation and their respective corresponding weights.

[0168] Further, the processor 901 is specifically configured to determine a difference matrix obtained by subtracting the RGB optical flow matrix from the virtual flow matrix; and determine a target matrix obtained by multiplying the difference matrix by a preset matrix;

[0169] The deviation between the virtual flow matrix and the RGB optical flow matrix is ​​determined according to the ratio of the norm of the target matrix to the norm of the preset matrix.

[0170] Furthermore, the processor 901 is further configured to obtain a predicted depth map corresponding to the image output by the initial NeRF model;

[0171] Identify the standard depth map corresponding to the image through the monocular depth estimation network DPT;

[0172] The processor 901 is specifically configured to determine a depth map deviation according to the predicted depth map and the standard depth map; and train the initial NeRF model according to the deviation and the depth map deviation.

[0173] Furthermore, the processor 901 is further configured to process each image using the SuperPoint method, obtain each feature point having a gesture sensing feature in the image and its position in the image, and determine a sampling area of ​​a preset size containing the feature point for each feature point; and randomly obtain the position of a sampling point in each obtained sampling area;

[0174] The processor 901 is specifically configured to input the image and the position of each sampling point into the initial NeRF model, and obtain the first prediction parameter of the acquisition device output by the initial NeRF model and the predicted pixel value of each sampling point in the image;

[0175] The processor 901 is specifically configured to determine a pixel deviation according to the predicted pixel value and a standard pixel value of each sampling point in the image; and train an initial NeRF model according to the pixel deviation and the deviation.

[0176] Further, the processor 901 is further configured to determine, for each sampling area, an adjustment size of the sampling area according to a ratio of the intensity of the light in the sampling area to the sum of the intensities of the light in each area in the image; and adjust the size of the sampling area according to the adjustment size;

[0177] The processor 901 is further configured to replace the sampling area with the adjusted sampling area, and execute a subsequent step of randomly acquiring positions of sampling points in each acquired sampling area for the replaced sampling area.

[0178] The communication bus mentioned in the above server can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The communication bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.

[0179] The communication interface is used for communication between the above electronic device and other devices.

[0180] The memory may include a random access memory (RAM) or a non-volatile memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located away from the aforementioned processor.

[0181] The above-mentioned processor can be a general-purpose processor, including a central processing unit, a network processor (Network Processor, NP), etc.; it can also be a digital signal processing processor (Digital Signal Processing, DSP), an application-specific integrated circuit, a field programmable gate array or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component, etc.

[0182] On the basis of the above embodiments, an embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program executable by an electronic device, and when the program is run on the electronic device, the electronic device implements the following steps when executing:

[0183] The memory stores a computer program, and when the program is executed by the processor, the processor performs the following steps:

[0184] Receive multiple images collected for the same scene; for each image, input the image and adjacent images collected by adjacent collection devices of the collection device corresponding to the image into an initial neural radiation field NeRF model, and obtain a first prediction parameter of the collection device and a second prediction parameter of the adjacent collection device output by the initial NeRF model;

[0185] For each image, project each pixel point in the image into the adjacent image according to the first prediction parameter and the second prediction parameter, generate a virtual flow matrix according to the deviation between the position of the same pixel point in the adjacent image after projection and the position of the pixel point in the adjacent image; and generate an RGB optical flow matrix according to the position difference between the same pixel point in the image and the adjacent image; determine the deviation between the virtual flow matrix and the RGB optical flow matrix;

[0186] The initial NeRF model is trained according to the deviation.

[0187] In a possible implementation manner, projecting each pixel point in the image in front of the adjacent image according to the first prediction parameter and the second prediction parameter, the method further includes:

[0188] Obtain a predicted depth map corresponding to the image output by the initial NeRF model;

[0189] The projecting each pixel point in the image onto the adjacent image according to the first prediction parameter and the second prediction parameter comprises:

[0190] For each pixel in the image, the following formula is used to determine the target position of the pixel after projection when the pixel is projected into the adjacent image:

[0191]

[0192] in, is the horizontal coordinate of the target position, is the ordinate of the target position, K i+1 is the internal parameter matrix composed of the internal parameters in the second prediction parameter according to the first preset method, T i+1 is the external parameter matrix composed of the external parameters in the second prediction parameter according to the second preset method, K i The internal parameters in the first prediction parameter are an internal parameter matrix composed according to a first preset method, is a depth matrix composed of the depth values ​​of each pixel in the predicted depth map, u k is the horizontal coordinate of the pixel point in the image, v k is the vertical coordinate of the position of the pixel in the image.

[0193] In a possible implementation manner, the adjacent acquisition device includes a first adjacent acquisition device adjacent to the acquisition device in a first preset direction and a second adjacent acquisition device adjacent to the acquisition device in a second preset direction;

[0194] The generating of the virtual flow matrix according to the deviation between the position of the same pixel point in the adjacent image after projection and the position of the pixel point in the adjacent image comprises:

[0195] Generate a first virtual flow matrix based on the deviation between the position of the same pixel point in the first adjacent image collected by the first adjacent acquisition device after projection and the position of the pixel point in the first adjacent image; generate a second virtual flow matrix based on the deviation between the position of the same pixel point in the second adjacent image collected by the second adjacent acquisition device after projection and the position of the pixel point in the second adjacent image; and generate a third virtual flow matrix based on the deviation between the position of the same pixel point in the first adjacent image after projection and the position of the pixel point in the second adjacent image;

[0196] Generating an RGB optical flow matrix according to the position difference between the same pixel in the image and the adjacent image comprises:

[0197] A first RGB optical flow matrix is ​​generated based on the position difference between the same pixel point in the image and the first adjacent image; a second RGB optical flow matrix is ​​generated based on the position difference between the same pixel point in the image and the second adjacent image; and a third RGB optical flow matrix is ​​generated based on the position difference between the same pixel point in the first adjacent image and the second adjacent image.

[0198] In a possible implementation manner, determining the deviation between the virtual flow matrix and the RGB optical flow matrix includes:

[0199] Determine a first deviation between the first virtual flow matrix and the first RGB optical flow matrix, a second deviation between the second virtual flow matrix and the second RGB optical flow matrix, and a third deviation between the third virtual flow matrix and the third RGB optical flow matrix;

[0200] A corresponding deviation is determined according to the first deviation, the second deviation, the third deviation and their respective corresponding weights.

[0201] In a possible implementation manner, determining the deviation between the virtual flow matrix and the RGB optical flow matrix includes:

[0202] Determine a difference matrix obtained by subtracting the RGB optical flow matrix from the virtual flow matrix; and determine a target matrix obtained by multiplying the difference matrix by a preset matrix;

[0203] The deviation between the virtual flow matrix and the RGB optical flow matrix is ​​determined according to the ratio of the norm of the target matrix to the norm of the preset matrix.

[0204] In a possible implementation, the method further includes:

[0205] Obtain a predicted depth map corresponding to the image output by the initial NeRF model;

[0206] Identify the standard depth map corresponding to the image through the monocular depth estimation network DPT;

[0207] The training of the initial NeRF model according to the deviation comprises:

[0208] A depth map deviation is determined according to the predicted depth map and the standard depth map; and the initial NeRF model is trained according to the deviation and the depth map deviation.

[0209] In a possible implementation, after receiving a plurality of images collected for the same scene and before training the initial NeRF model according to the deviation, the method further includes:

[0210] For each image, the SuperPoint method is used to process the image, and each feature point with gesture perception characteristics in the image and its position in the image are obtained. For each feature point, a sampling area of ​​a preset size containing the feature point is determined; the position of the sampling point is randomly obtained in each obtained sampling area;

[0211] Inputting the image into an initial neural radiation field NeRF model, and obtaining a first prediction parameter of the acquisition device output by the initial NeRF model comprises:

[0212] Input the image and the position of each sampling point into the initial NeRF model, and obtain the first prediction parameter of the acquisition device and the predicted pixel value of each sampling point in the image output by the initial NeRF model;

[0213] The training of the initial NeRF model according to the deviation comprises:

[0214] A pixel deviation is determined according to the predicted pixel value and the standard pixel value of each sampling point in the image; and an initial NeRF model is trained according to the pixel deviation and the deviation.

[0215] In a possible implementation manner, after determining the sampling area of ​​the preset size containing the feature point, the method further includes:

[0216] For each sampling area, determining an adjustment size of the sampling area according to a ratio of the intensity of the light in the sampling area to the sum of the intensities of the light in each area in the image; and adjusting the size of the sampling area according to the adjustment size;

[0217] After iterating the initial NeRF model for a preset number of times, before randomly acquiring the position of a sampling point in each acquired sampling area, the method further includes:

[0218] The adjusted sampling area is used to replace the sampling area, and for the replaced sampling area, a subsequent step of randomly acquiring the position of the sampling point in each acquired sampling area is performed.

[0219] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.

[0220] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0221] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0222] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0223] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.

Claims

1. A model training method, It is characterized in that The method comprises: Receive multiple images collected for the same scene; for each image, input the image and adjacent images collected by adjacent collection devices of the collection device corresponding to the image into an initial neural radiation field NeRF model, and obtain a first prediction parameter of the collection device and a second prediction parameter of the adjacent collection device output by the initial NeRF model; For each image, project each pixel point in the image into the adjacent image according to the first prediction parameter and the second prediction parameter, generate a virtual flow matrix according to the deviation between the position of the same pixel point in the adjacent image after projection and the position of the pixel point in the adjacent image; and generate an RGB optical flow matrix according to the position difference between the same pixel point in the image and the adjacent image; determine the deviation between the virtual flow matrix and the RGB optical flow matrix; The initial NeRF model is trained according to the deviation.

2. The method according to claim 1, It is characterized in that The method further comprises: projecting each pixel point in the image in front of the adjacent image according to the first prediction parameter and the second prediction parameter: Obtain a predicted depth map corresponding to the image output by the initial NeRF model; The projecting each pixel point in the image onto the adjacent image according to the first prediction parameter and the second prediction parameter comprises: For each pixel in the image, the following formula is used to determine the target position of the pixel after projection when the pixel is projected into the adjacent image: in, is the horizontal coordinate of the target position, is the ordinate of the target position, K i+1 is the internal parameter matrix composed of the internal parameters in the second prediction parameter according to the first preset method, T i+1 is the external parameter matrix composed of the external parameters in the second prediction parameter according to the second preset method, K i The internal parameters in the first prediction parameter are an internal parameter matrix composed according to a first preset method, is a depth matrix composed of the depth values ​​of each pixel in the predicted depth map, u k is the horizontal coordinate of the pixel point in the image, v k is the vertical coordinate of the position of the pixel in the image.

3. The method according to claim 1, It is characterized in that The adjacent acquisition devices include a first adjacent acquisition device adjacent to the acquisition device in a first preset direction and a second adjacent acquisition device adjacent to the acquisition device in a second preset direction; The generating of the virtual flow matrix according to the deviation between the position of the same pixel point in the adjacent image after projection and the position of the pixel point in the adjacent image comprises: Generate a first virtual flow matrix based on the deviation between the position of the same pixel point in the first adjacent image collected by the first adjacent acquisition device after projection and the position of the pixel point in the first adjacent image; generate a second virtual flow matrix based on the deviation between the position of the same pixel point in the second adjacent image collected by the second adjacent acquisition device after projection and the position of the pixel point in the second adjacent image; and generate a third virtual flow matrix based on the deviation between the position of the same pixel point in the first adjacent image after projection and the position of the pixel point in the second adjacent image; Generating an RGB optical flow matrix according to the position difference between the same pixel in the image and the adjacent image comprises: A first RGB optical flow matrix is ​​generated based on the position difference between the same pixel point in the image and the first adjacent image; a second RGB optical flow matrix is ​​generated based on the position difference between the same pixel point in the image and the second adjacent image; and a third RGB optical flow matrix is ​​generated based on the position difference between the same pixel point in the first adjacent image and the second adjacent image.

4. The method according to claim 3, It is characterized in that Determining the deviation between the virtual flow matrix and the RGB optical flow matrix includes: Determine a first deviation between the first virtual flow matrix and the first RGB optical flow matrix, a second deviation between the second virtual flow matrix and the second RGB optical flow matrix, and a third deviation between the third virtual flow matrix and the third RGB optical flow matrix; A corresponding deviation is determined according to the first deviation, the second deviation, the third deviation and their respective corresponding weights.

5. The method according to claim 1 or 3 or 4, It is characterized in that Determining the deviation between the virtual flow matrix and the RGB optical flow matrix includes: Determine a difference matrix obtained by subtracting the RGB optical flow matrix from the virtual flow matrix; and determine a target matrix obtained by multiplying the difference matrix by a preset matrix; The deviation between the virtual flow matrix and the RGB optical flow matrix is ​​determined according to the ratio of the norm of the target matrix to the norm of the preset matrix.

6. The method according to claim 1, It is characterized in that The method further comprises: Obtain a predicted depth map corresponding to the image output by the initial NeRF model; Identify the standard depth map corresponding to the image through the monocular depth estimation network DPT; The training of the initial NeRF model according to the deviation comprises: A depth map deviation is determined according to the predicted depth map and the standard depth map; and the initial NeRF model is trained according to the deviation and the depth map deviation.

7. The method according to claim 1, It is characterized in that After receiving a plurality of images collected for the same scene and before training the initial NeRF model according to the deviation, the method further includes: For each image, the SuperPoint method is used to process the image, and each feature point with gesture perception characteristics in the image and its position in the image are obtained. For each feature point, a sampling area of ​​a preset size containing the feature point is determined; the position of the sampling point is randomly obtained in each obtained sampling area; Inputting the image into an initial neural radiation field NeRF model, and obtaining a first prediction parameter of the acquisition device output by the initial NeRF model comprises: Input the image and the position of each sampling point into the initial NeRF model, and obtain the first prediction parameter of the acquisition device and the predicted pixel value of each sampling point in the image output by the initial NeRF model; The training of the initial NeRF model according to the deviation comprises: A pixel deviation is determined according to the predicted pixel value and the standard pixel value of each sampling point in the image; and an initial NeRF model is trained according to the pixel deviation and the deviation.

8. The method according to claim 7, It is characterized in that After determining the sampling area of ​​the preset size containing the feature point, the method further includes: For each sampling area, determining an adjustment size of the sampling area according to a ratio of the intensity of the light in the sampling area to the sum of the intensities of the light in each area in the image; and adjusting the size of the sampling area according to the adjustment size; After iterating the initial NeRF model for a preset number of times, before randomly acquiring the position of a sampling point in each acquired sampling area, the method further includes: The adjusted sampling area is used to replace the sampling area, and for the replaced sampling area, a subsequent step of randomly acquiring the position of the sampling point in each acquired sampling area is performed.

9. An electronic device, It is characterized in that The electronic device comprises at least a processor and a memory, and the processor is used to implement the steps of model training as described in any one of claims 1 to 8 when executing a computer program stored in the memory.

10. A computer-readable storage medium, It is characterized in that It stores a computer program, which, when executed by a processor, implements the steps of model training as described in any one of claims 1 to 8.