Model Training Method, Perspective Image Generation Method, Device, Equipment and Medium

The method enhances NeRF-based new view generation by training with sparse source images, improving image quality and clarity without needing many source images, thus reducing computational complexity.

CN115409949BActive Publication Date: 2025-07-15NETEASE LINGDONG (HANGZHOU) TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211124534.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-15
Publication Date
2025-07-15
Estimated Expiration
2042-09-15

AI Technical Summary

Technical Problem

The existing new viewing angle image generation method based on neural radiation fields has high requirements for the number of input source viewing angle images, resulting in the generated target image being blurred when the number of source viewing angle images is small and lacking sharp details.

Method used

By obtaining the sample preset two-dimensional image in the preset target view angle and multiple sample source view angle images in the preset three-dimensional scene, multiple original projected pixel coordinates and original projected image features of each spatial sampling point are obtained, new projected pixel coordinates and new projected image features are generated, and finally model training is carried out to generate a viewing image generation model.

Benefits of technology

It effectively reduces the requirements for the number of source viewing images, improves the generation quality and clarity of the target image, reduces the computational complexity, and improves the model efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115409949B_ABST
    Figure CN115409949B_ABST
Patent Text Reader

Abstract

The present application provides a model training method, a perspective image generation method, an apparatus, a device and a medium, relating to the technical field of image processing. The model training method includes: obtaining a sample preset two-dimensional image at a preset target perspective and multiple sample source perspective images at multiple preset source perspectives in a preset three-dimensional scene, and then, obtaining multiple original projected pixel coordinates and multiple original projected image features of each spatial sampling point at the preset target perspective. Then, according to the multiple original projected pixel coordinates of each spatial sampling point and the multiple sample source perspective images, multiple new projected pixel coordinates and multiple new projected image features of each spatial sampling point are generated, and further, target projected image features of each spatial sampling point at multiple preset source perspectives and a sample target two-dimensional image at the preset target perspective in the preset three-dimensional scene are generated; finally, model training is performed to generate a perspective image generation model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular, to a model training method, a perspective image generation method, a device, a device and a medium. Background Art

[0002] The image rendering method based on neural radiance fields has achieved great success in the new view generation task, significantly improving the quality of the generated images.

[0003] However, the existing new view image generation methods based on neural radiance fields have high requirements and limitations on the number of input source view images, and often require a relatively large number of dense source view images to ensure the quality of the generated target images. When the number of given source view images decreases, the generated target images are easily blurred and lack sharp details. Summary of the Invention

[0004] The purpose of the present invention is to provide a model training method, a perspective image generation method, a device, a device and a medium for effectively processing the situation where the number of source view images is small and sparse, and significantly improving the visual quality and clarity of the generated target images in view of the above-mentioned deficiencies in the prior art.

[0005] To achieve the above object, the technical solutions adopted in the embodiments of the present application are as follows:

[0006] In a first aspect, an embodiment of the present application provides a model training method, including:

[0007] Obtaining a sample preset two-dimensional image at a preset target view and a plurality of sample source view images at a plurality of preset source views in a preset three-dimensional scene;

[0008] Obtaining a plurality of original projected pixel coordinates and a plurality of original projected image features at each spatial sampling point at the preset target view, where the plurality of original projected pixel coordinates respectively correspond to the plurality of preset source views, and the plurality of original projected image features are the image features at the plurality of original projected pixel coordinates in the plurality of sample source view images;

[0009] Generating a plurality of new projected pixel coordinates and a plurality of new projected image features at each spatial sampling point according to the plurality of original projected pixel coordinates at each spatial sampling point and the plurality of sample source view images, where the plurality of new projected pixel coordinates correspond to the plurality of preset source views, and the plurality of new projected image features are the image features at the plurality of new two-dimensional projected pixel coordinates in the plurality of sample source view images;

[0010] Generate target projection image features of each spatial sampling point at the multiple preset source viewpoints according to the multiple original projection image features and the multiple new projection image features;

[0011] Generate a sample target two-dimensional image in the preset three-dimensional scene at the preset target viewpoint according to the target projection image features of the multiple spatial sampling points at the multiple preset source viewpoints;

[0012] Generate the viewpoint image generation model according to the sample target two-dimensional image in the preset three-dimensional scene at the preset target viewpoint and the sample preset two-dimensional image at the preset target viewpoint for model training.

[0013] In a second aspect, an embodiment of the present application further provides a viewpoint image generation method, including:

[0014] Obtain multiple source viewpoint images in a preset three-dimensional scene at multiple source viewpoints;

[0015] Generate a target two-dimensional image in the preset three-dimensional scene at the target viewpoint according to the multiple source viewpoint images and the target viewpoint by using a pre-trained viewpoint image generation model, where the viewpoint image generation model is a model trained by using any one of the model training methods in the first aspect.

[0016] In a third aspect, an embodiment of the present application further provides a training device for a model, including:

[0017] An acquisition module, configured to acquire a sample preset two-dimensional image in a preset three-dimensional scene at a preset target viewpoint and multiple sample source viewpoint images at multiple preset source viewpoints;

[0018] An original image feature generation module, configured to obtain multiple original projection pixel coordinates and multiple original projection image features of each spatial sampling point at the preset target viewpoint, where the multiple original projection pixel coordinates respectively correspond to the multiple preset source viewpoints, and the multiple original projection image features are image features at the multiple original projection pixel coordinates in the multiple sample source viewpoint images;

[0019] A new projection feature generation module, configured to generate multiple new projection pixel coordinates and multiple new projection image features of each spatial sampling point according to the multiple original projection pixel coordinates of each spatial sampling point and the multiple sample source viewpoint images, where the multiple new projection pixel coordinates correspond to the multiple preset source viewpoints, and the multiple new projection image features are image features at the multiple new two-dimensional projection pixel coordinates in the multiple sample source viewpoint images;

[0020] A target projection image feature generation module, configured to generate target projection image features of each spatial sampling point at the multiple preset source viewpoints according to the multiple original projection image features and the multiple new projection image features;

[0021] A target two-dimensional image generation module, configured to generate a sample target two-dimensional image in the preset three-dimensional scene at the preset target viewpoint according to the target projection image features of multiple spatial sampling points at the multiple preset source viewpoints;

[0022] A viewpoint image generation module, configured to perform model training according to the sample target two-dimensional image in the preset three-dimensional scene at the preset target viewpoint and the sample preset two-dimensional image at the preset target viewpoint, and generate the viewpoint image generation model.

[0023] In a fourth aspect, an embodiment of the present application further provides a viewpoint image generation device, including an image acquisition module and an image generation module:

[0024] The image acquisition module is configured to acquire multiple source viewpoint images in a preset three-dimensional scene at multiple source viewpoints;

[0025] The image generation module is configured to generate a target two-dimensional image in the preset three-dimensional scene at the target viewpoint according to the multiple source viewpoint images and the target viewpoint, using a pre-trained viewpoint image generation model, where the viewpoint image generation model is a model trained by using any of the model training methods in the first aspect above.

[0026] In a fifth aspect, an embodiment of the present application further provides an electronic device, including: a processor, a storage medium, and a bus. The storage medium stores program instructions executable by the processor. When the electronic device runs, communication is performed between the processor and the storage medium through the bus. The processor executes the program instructions to perform the steps of any of the model training methods in the first aspect when executed, or perform the steps of the viewpoint image generation method in the second aspect.

[0027] In a sixth aspect, an embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it performs the steps of any of the model training methods in the first aspect, or performs the steps of the viewpoint image generation method in the second aspect.

[0028] The beneficial effects of the present application are as follows: The embodiment of the present application provides a model training method. First, obtain the sample preset two-dimensional image in the preset target perspective and the multiple sample source perspective images in multiple preset source perspectives in the preset three-dimensional scene. Then, obtain the multiple original projection pixel coordinates and multiple original projection image features of each spatial sampling point in the preset target perspective. Next, generate the multiple new projection pixel coordinates and multiple new projection image features of each spatial sampling point according to the multiple original projection pixel coordinates of each spatial sampling point and the multiple sample source perspective images, and further generate the target projection image features of each spatial sampling point in multiple preset source perspectives. Generate the sample target two-dimensional image of the preset three-dimensional scene in the preset target perspective according to the target projection image features of multiple spatial sampling points in multiple preset source perspectives. Finally, perform model training according to the sample target two-dimensional image in the preset target perspective in the preset three-dimensional scene and the sample preset two-dimensional image in the preset target perspective to generate a perspective image generation model. By processing the original projection image features and the generated new projection image features, generate the sample target two-dimensional image to train the perspective image generation model. The perspective image generation model obtained thereby can effectively reduce and lower the quantity requirements and restrictions on the source perspective images in the new perspective generation process, and significantly improve the generation quality and clarity of the target image when the number of source images is small. In addition, the input of fewer source perspective images can also effectively reduce the computational complexity and improve the model efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention, and thus should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0030] Figure 1 It is a flowchart of a model training method provided by an embodiment of the present application;

[0031] Figure 2 It is a flowchart of a model training method provided by another embodiment of the present application;

[0032] Figure 3 It is a flowchart of a model training method provided by another embodiment of the present application;

[0033] Figure 4 It is a flowchart of a model training method provided by still another embodiment of the present application;

[0034] Figure 5 It is a flowchart of a model training method provided by still another second embodiment of the present application;

[0035] Figure 6 A flowchart of a model training method provided in accordance with another embodiment of the present application;

[0036] Figure 7 A flowchart of a model training method provided in yet another embodiment of the present application;

[0037] Figure 8 A flowchart of a model training method provided in another fifth embodiment of the present application;

[0038] Figure 9 A flowchart of a method for generating a viewing angle image provided by an embodiment of the present application;

[0039] Figure 10 A schematic diagram of a model training device provided in one embodiment of the present application;

[0040] Figure 11 A schematic diagram of a viewing angle image generating device provided in one embodiment of the present application;

[0041] Figure 12 A schematic diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0042] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments.

[0043] In this application, unless otherwise clearly specified and limited, the terms "first" and "second" are used only for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Thus, the features defined as "first" and "second" can explicitly or implicitly include at least one feature. In the description of the present invention, the meaning of "multiple" is at least two, such as two or three, unless otherwise clearly and specifically limited. The terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that the process, method, article or equipment including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or equipment. In the absence of more restrictions, the elements defined by the sentence "include one..." do not exclude the presence of other identical elements in the process, method, article or equipment including the elements.

[0044] Given several 2D source perspective images and their corresponding camera poses in a specific 3D scene, the novel view synthesis task aims to generate realistic 2D images corresponding to the scene from a new target perspective. Currently, this technology has been widely applied in many fields such as 3D reconstruction, augmented reality (AR), virtual reality (VR), etc., with broad application prospects and great market value.

[0045] In recent years, image rendering methods based on Neural Radiance Fields (NeRF) have achieved great success in the novel view synthesis task, significantly improving the quality of the generated images. The current training of generalizable novel view generation models based on NeRF mainly includes the following processes:

[0046] Projected image features of multiple 3D sampling points are selected from the source perspective images at several source perspectives in a given or selected specific scene. By aggregating the above projected image features and predicting color density, an image at the target perspective is rendered. Finally, according to the obtained image at the target perspective, the model parameters are adjusted to obtain a trained novel view generation model.

[0047] The novel view generation model trained by the above method can generate results with acceptable visual quality and can solve some application problems.

[0048] However, the novel view generation models obtained by existing methods have high requirements and limitations on the number of input source perspective images, and a large number of dense source perspective images are required to ensure the quality of the generated target images. When the number of given source perspective images decreases, the generated target images are likely to become blurred and lack sharp details.

[0049] To address the problems existing in the current novel view generation models, the embodiments of this application provide various possible implementation manners to effectively handle the situation where the number of source perspective images is small and sparse, and significantly improve the visual quality and clarity of the generated target images. The following is explained through multiple examples in combination with the accompanying drawings. Figure 1 The flowchart of a model training method provided by an embodiment of this application. This method can be implemented by an electronic device running the above model training method. The electronic device can be, for example, a terminal device or a server. As Figure 1 shown, this method includes:

[0050] Step 101: Obtain a sample preset 2D image at a preset target perspective and multiple sample source perspective images at multiple preset source perspectives in a preset three-dimensional scene.

[0051] It should be noted that the present application does not limit the specific perspective settings of the preset target perspective and multiple preset source perspectives. As long as the sample preset two-dimensional image and the multiple sample source perspective images include the same target observation object, the sample target two-dimensional image obtained in step 105 of the present application is a two-dimensional image of the preset target perspective including the target observation object. In addition, in the present application, the multiple sample source perspective images are two-dimensional images of multiple sample source perspectives, and the number of which is determined by the user according to the specific usage situation. The present application does not limit the specific number of the sample source perspective images.

[0052] It should also be noted that the preset three-dimensional scene can be a real three-dimensional scene (such as the scene of an excavator bucket proposed in the following example), or a virtual three-dimensional scene (such as a three-dimensional game scene, a three-dimensional map scene, etc.). The present application does not limit this.

[0053] In a specific implementation manner, obtain N sample source perspective images (two-dimensional images) under the preset three-dimensional scene and the sample preset two-dimensional image under the preset target perspective.

[0054] Step 102: Obtain multiple original projection pixel coordinates and multiple original projection image features of each spatial sampling point under the preset target perspective, where the multiple original projection pixel coordinates respectively correspond to multiple preset source perspectives, and the multiple original projection image features are the image features at the multiple original projection pixel coordinates in the multiple sample source perspective images.

[0055] It should be noted that a sampling space is generated according to the light rays emitted by each pixel point under the preset target perspective, and multiple (for example, P, where P is a positive integer) spatial sampling points are set on the light rays emitted by each pixel point in the sampling space (that is, each optical fiber in the sampling space).

[0056] It should also be noted that the original projection image features in the present application can include, for example, semantic features, perspective information, etc. The present application does not limit this.

[0057] Since the spatial sampling points are all located under the preset target perspective, there are corresponding pixels of the target observation object on each optical fiber where each spatial sampling point is located. According to the correspondence between the spatial sampling points and the pixels of the target observation object, as well as the preset target perspective and the multiple sample source perspective images, the multiple original projection pixel coordinates corresponding to the pixels of the target observation object in the multiple sample source perspective images can be determined, and then the original projection image features at the positions of each original projection pixel coordinate can be obtained.

[0058] The above is only an example. In actual implementation, there can be other ways to obtain the original projection pixel coordinates and the original projection image features. The present application does not limit this.

[0059] Step 103: Generate multiple new projected pixel coordinates and multiple new projected image features for each spatial sampling point based on the multiple original projected pixel coordinates of each spatial sampling point and the multiple sample source perspective images, where the multiple new projected pixel coordinates correspond to multiple preset source perspectives, and the multiple new projected image features are the image features at the multiple new two-dimensional projected pixel coordinates in the multiple sample source perspective images.

[0060] Based on the multiple original projected pixel coordinates of each spatial sampling point and the multiple sample source perspective images, multiple new projected pixel coordinates and multiple new projected image features of each spatial sampling point can be obtained under a new set of source perspectives (a new set of source perspectives different from the multiple preset source perspectives). The acquisition methods of the new projected pixel coordinates and the new projected image features can refer to the acquisition methods of the multiple original projected pixel coordinates and the multiple original projected image features of each spatial sampling point under the preset target perspective in Step 102, which will not be elaborated herein in this application.

[0061] It should be noted that there are new projected pixel coordinates and new projected image features for the spatial sampling point. In this application, the new projected pixel coordinates can be obtained by processing the original projected pixel coordinates or by other offset calculation methods, and this application does not make any limitations in this regard.

[0062] Step 104: Generate the target projected image features of each spatial sampling point under multiple preset source perspectives based on the multiple original projected image features and the multiple new projected image features.

[0063] It should be noted that the multiple original projected image features and the multiple new projected image features are all image features for the target observed object. Therefore, by processing the multiple original projected image features and the multiple new projected image features, the target projected image features of each spatial sampling point under multiple preset source perspectives can be generated, thereby realizing implicit perspective enhancement.

[0064] In a possible implementation manner, the multiple original projected image features and the multiple new projected image features can be aggregated to generate the target projected image features of each spatial sampling point under multiple preset source perspectives.

[0065] The above is only an example illustration. In actual implementation, there can be other ways to generate the target projected image features, and this application does not make any limitations in this regard.

[0066] Step 105: Generate a sample target two-dimensional image of the preset three-dimensional scene under the preset target perspective based on the target projected image features of multiple spatial sampling points under multiple preset source perspectives.

[0067] Process the target projection image features at multiple spatial sampling points under multiple preset source viewpoints to generate a sample target two-dimensional image at a preset target viewpoint in a preset three-dimensional scene. It should be noted that the specific processing method for processing the target projection image features under multiple preset source viewpoints can be, for example, color prediction, density prediction, etc. for each spatial sampling point. This application does not limit this, as long as a sample target two-dimensional image at a preset target viewpoint in a preset three-dimensional scene can be generated.

[0068] Step 106: Perform model training based on the sample target two-dimensional image at a preset target viewpoint in a preset three-dimensional scene and the sample preset two-dimensional image at the preset target viewpoint to generate a viewpoint image generation model.

[0069] In a possible implementation, calculate the mean squared error (MSE) between the generated sample target two-dimensional image and the sample preset two-dimensional image (that is, between the generated sample target two-dimensional image and the preset ground truth image) as the loss function of the network. Use the gradient descent algorithm to iteratively update and optimize the network parameters of the viewpoint image generation model, so that the generated sample target two-dimensional image is as consistent as possible with the preset two-dimensional image (that is, the output of the model is as consistent as possible with the ground truth), thereby finally obtaining a viewpoint image generation model that meets the requirements of developers.

[0070] In a possible implementation, an iteration count threshold can be preset, and the above steps can be repeated to optimize the network parameters of the viewpoint image generation model until the number of loops reaches the preset iteration count threshold. Save the network parameters of the final model to generate a viewpoint image generation model.

[0071] The above is only an example. In actual implementation, there can be other calculation and training methods for the calculation method of the loss function of the viewpoint image generation model, the model training method, etc. This application does not limit this.

[0072] In summary, the embodiment of the present application provides a model training method. First, obtain a sample preset two-dimensional image at a preset target view angle and multiple sample source view angle images at multiple preset source view angles in a preset three-dimensional scene. Then, obtain multiple original projection pixel coordinates and multiple original projection image features at each spatial sampling point at the preset target view angle. Next, generate multiple new projection pixel coordinates and multiple new projection image features at each spatial sampling point according to the multiple original projection pixel coordinates at each spatial sampling point and the multiple sample source view angle images, and further generate target projection image features at each spatial sampling point at multiple preset source view angles. Generate a sample target two-dimensional image at the preset target view angle in the preset three-dimensional scene according to the target projection image features at multiple spatial sampling points at multiple preset source view angles. Finally, perform model training according to the sample target two-dimensional image at the preset target view angle in the preset three-dimensional scene and the sample preset two-dimensional image at the preset target view angle to generate a view image generation model. By processing the original projection image features and the generated new projection image features to generate a sample target two-dimensional image for training the view image generation model, the view image generation model obtained thereby can effectively reduce and lower the quantity requirements and restrictions on the source view angle images during the new view angle generation process, and significantly improve the generation quality and clarity of the target image when the number of source images is small. In addition, the input of fewer source view angle images can effectively reduce the computational complexity and improve the model efficiency.

[0073] Optionally, on the basis of the above Figure 1 , a possible implementation manner of the model training method provided by the present application is Figure 2 a flowchart of a model training method provided by another embodiment of the present application; as Figure 2 shown, step 103 generates multiple new two-dimensional projection pixel coordinates and multiple new projection image features at each spatial sampling point according to the multiple original projection pixel coordinates at each spatial sampling point and the multiple sample source view angle images, including:

[0074] Step 201: Generate multiple new two-dimensional projection pixel coordinates at each spatial sampling point according to the multiple original projection pixel coordinates at each spatial sampling point.

[0075] Generating multiple new two-dimensional projected pixel coordinates and multiple new projected image features for each spatial sampling point requires new two-dimensional projected pixel coordinates. For this, the new two-dimensional projected pixel coordinates can be obtained based on the multiple original projected pixel coordinates of each spatial sampling point. For example, position calculations are performed on the multiple original projected pixel coordinates of each spatial sampling point through a preset means. The preset means can be, for example, position offset (such as offsetting each original projection coordinate by a preset length in at least one direction), position calculation (such as performing calculations on each original projection coordinate according to a pre-designed calculation method. For specific examples, see step 301 and step 302), etc. The present application does not limit this.

[0076] Step 202: Generate multiple new projected image features based on multiple sample source perspective images and multiple new two-dimensional projected pixel coordinates.

[0077] Generating multiple new projected image features based on multiple sample source perspective images and multiple new two-dimensional projected pixel coordinates is implemented in the same way as generating multiple original projected image features based on multiple sample source perspective images and multiple original projected pixel coordinates in step 102. The present application will not elaborate here.

[0078] In a possible implementation manner, multiple new projected image features can also be generated based on the source image features corresponding to multiple sample source perspective images and multiple new two-dimensional projected pixel coordinates. Among them, the source image features corresponding to multiple sample source perspective images can be obtained by performing feature extraction on each sample source perspective image. Thus, when generating multiple new projected image features, a set of new projected image features of multiple spatial sampling points can be sampled from the source image features corresponding to multiple sample source perspective images according to multiple new two-dimensional projected pixel coordinates.

[0079] Optionally, on the basis of the above Figure 2 , the present application also provides a possible implementation manner of a model training method. Figure 3 It is a flowchart of a model training method provided by another embodiment of the present application; as Figure 3 shown, step 201 generates multiple new two-dimensional projected pixel coordinates for each spatial sampling point according to the multiple original projected pixel coordinates of each spatial sampling point, including:

[0080] Step 301: Generate multiple two-dimensional coordinate offsets for each spatial sampling point according to multiple original projected image features of each spatial sampling point. The multiple two-dimensional coordinate offsets correspond to multiple preset source perspectives respectively.

[0081] It should be noted that, based on the features of multiple original projection images at each spatial sampling point, the offset of each spatial sampling point relative to the original projection pixel coordinates at a preset source view angle can be obtained, that is, the two-dimensional coordinate offset. Thus, each two-dimensional coordinate offset corresponds to a preset source view angle. The specific generation method of the two-dimensional coordinate offset in this application is not limited, as long as the two-dimensional coordinate offset is generated based on the features of the original projection images, thereby ensuring the stability of model training, and the two-dimensional coordinate offset can be used normally in subsequent steps. In a possible implementation, the generation method of multiple two-dimensional coordinate offsets can refer to step 401 and step 402, but it should be noted that this method is not the only way to generate two-dimensional coordinate offsets.

[0082] Step 302: Generate multiple new two-dimensional projection pixel coordinates for each spatial sampling point according to multiple original projection pixel coordinates and multiple two-dimensional coordinate offsets.

[0083] In a possible implementation, according to multiple original projection pixel coordinates and multiple two-dimensional coordinate offsets, the method of coordinate accumulation (that is, adding the two-dimensional coordinate offset to the original projection pixel coordinate to obtain a new two-dimensional projection image coordinate) can be used to generate multiple new two-dimensional projection pixel coordinates for each spatial sampling point.

[0084] In a specific implementation, if multiple original projection pixel coordinates are multiple two-dimensional coordinate offsets are multiple new two-dimensional projection pixel coordinates The calculation method is as follows:

[0085]

[0086] The above is only an example. In actual implementation, there can be other calculation methods for multiple new two-dimensional projection pixel coordinates, and this application does not limit this.

[0087] Optionally, on the basis of the above Figure 3 This application also provides a possible implementation of a model training method. Figure 4 It is a flowchart of a model training method provided by another embodiment of this application; as Figure 4 shown, step 301 generates multiple two-dimensional coordinate offsets for each spatial sampling point according to multiple original projection image features of each spatial sampling point, including:

[0088] Step 401: Aggregate the features of multiple original projection images of each spatial sampling point to generate the original aggregated feature of each spatial sampling point;

[0089] In a possible implementation, the original projection image features can be averaged pooled to integrate the original projection image features from different perspectives, thereby generating the original aggregated features for each spatial sampling point.

[0090] The above is only an example. In actual implementation, there may be other ways to generate the original aggregated features, and the present application does not limit this.

[0091] Step 402: Generate multiple two-dimensional coordinate offsets for each spatial sampling point according to the original aggregated features of each spatial sampling point and multiple original projection image features.

[0092] After obtaining the original aggregated features of each spatial sampling point, there may be differences between the original aggregated features and each original projection image feature. Therefore, multiple two-dimensional coordinate offsets for each spatial sampling point can be generated based on the differences between the original aggregated features of each spatial sampling point and multiple original projection image features.

[0093] In a possible implementation, the specific multiple two-dimensional coordinate offsets of each spatial sampling point can refer to Step 501 and Step 502, or can be obtained by other methods, and the present application does not limit this.

[0094] Optionally, on the basis of the above Figure 4 the present application also provides a possible implementation of a model training method. Figure 5 It is a flowchart of a model training method provided by a second embodiment of the present application; as Figure 5 shown, Step 402 generates multiple two-dimensional coordinate offsets for each spatial sampling point according to the original aggregated features of each spatial sampling point and multiple original projection image features, including:

[0095] Step 501: Generate multiple feature differences for each spatial sampling point according to the original aggregated features of each spatial sampling point and multiple original projection image features, and the multiple feature differences correspond to multiple preset source perspectives;

[0096] In a possible implementation, the multiple original projection image features of each spatial sampling point are feature aggregated to generate the original aggregated features of each spatial sampling point. According to the original aggregated features of each spatial sampling point and multiple original projection image features generate multiple feature differences for each spatial sampling point Each feature difference contains the difference information between the original projection features and the original aggregated features under different preset source perspectives at the high-order feature level.

[0097] In a specific implementation manner, multiple feature differences of each spatial sampling point can be calculated in the following way:

[0098]

[0099] The above is only an example. In actual implementation, there can be other ways to calculate feature differences (such as adding weights, influencing factors, etc.), and the present application does not limit this.

[0100] Step 502: Map multiple feature differences to generate multiple two-dimensional coordinate offsets.

[0101] By mapping the above multiple feature differences, multiple two-dimensional coordinate offsets can be generated.

[0102] In a possible implementation manner, a multi-layer perceptron (MLP) can be used to map multiple feature differences to multiple two-dimensional coordinate offsets. This offset value represents the position offset of the spatial sampling point relative to its original projection pixel coordinates under different preset source perspectives. The larger the feature difference, it indicates that the original projection image feature of the spatial sampling point under the current preset source perspective is more different from its initial aggregated feature (which can be understood as the average projection feature under all preset source perspectives). Then, the pixel point corresponding to the original projection pixel coordinates of this point is more likely to be an inaccurate outlier point. For this situation, the mapping network of the above multi-layer perceptron (MLP) can be adjusted to make it output a relatively large coordinate offset (that is, the larger the two-dimensional coordinate offset) to expand the search range and capture useful information; the smaller the feature difference, it indicates that the pixel point corresponding to the original projection pixel coordinates of the spatial sampling point under the current preset source perspective is more likely to be a pixel point with higher accuracy and confidence. For this situation, the mapping network of the above multi-layer perceptron (MLP) can be adjusted to make it output a relatively small coordinate offset (that is, the smaller the two-dimensional coordinate offset) to appropriately reduce the search range and capture useful information within the vicinity of the original projection pixel coordinates.

[0103] Based on the above steps, adding them to the original projection pixel coordinates respectively can obtain a set of new two-dimensional projection pixel coordinates of the spatial sampling point under different preset source perspectives, and then a set of new projection image features of the spatial sampling point can be sampled;

[0104] Optionally, based on the above Figure 1 The present application also provides a possible implementation manner of a model training method. Figure 6 This is a flowchart of a model training method provided by repeated embodiments of the present application; asFigure 6 As shown in Figure 6 , step 104: Generate the target projection image features of each spatial sampling point at multiple preset source viewpoints according to multiple original projection image features and multiple new projection image features, including:

[0105] Step 601: Determine at least one new projection image feature from multiple new projection image features.

[0106] Since there are differences in the similarity between each new projection image feature and the original projection image feature, at least one new projection image feature can be determined from multiple new projection image features.

[0107] In a possible implementation, for multiple new projection image features, the similarity can be calculated by vector dot product (inner product), and the K (K can be a positive integer greater than or equal to 1) new projection image features that are closest to the original projection image feature can be selected from the above new projection image features as the new projection image features (i.e., the additional implicitly view-enhanced projection image features).

[0108] In another possible implementation, since the original aggregated feature can reflect the average projection features at all preset source viewpoints, therefore, for multiple new projection image features, the similarity can be calculated by vector dot product (inner product), and the K (K can be a positive integer greater than or equal to 1) new projection image features that are closest to the original aggregated feature can be selected from the above new projection image features as the new projection image features (i.e., the additional implicitly view-enhanced projection image features).

[0109] In yet another possible implementation, the similarity (the similarity with the original projection image feature, or the similarity with the original aggregated feature) can be determined by calculating the L1, L2, etc. distances between features, and then the new projection image features can be determined according to the magnitude of the similarity; in addition, a preset neural network can also be used to calculate the similarity between features, etc.

[0110] The above is only an example for illustration. In actual implementation, there can be other implementation manners, and this application does not limit this.

[0111] Step 602: Generate the target projection image features of each spatial sampling point at multiple preset source viewpoints according to multiple original projection image features and at least one new projection image feature.

[0112] In a possible implementation, the original projection image features Stitch with at least one new projected image feature (i.e., the K new projected image features determined in step 601), so as to obtain the target projected image features of each spatial sampling point at multiple preset source viewpoints.

[0113] In a specific implementation manner, feature stitching can be performed on the target projected image features, multiple original projected image features, and at least one new projected image feature of each spatial sampling point at multiple preset source viewpoints along the viewpoint direction, which can be understood as an increase in the viewpoints. Specifically, assume there are N original projected image features (each original projected image feature is d-dimensional), and there are K enhanced projected features of new projected image features (each new projected image feature is also d-dimensional). After stitching, N + K projected features are obtained, that is, there are N + K target projected image features, and each target projected image feature is d-dimensional.

[0114] The above is only an example description. In actual implementation, there may be other generation implementation methods for the target projected image features, and the present application does not limit this.

[0115] Optionally, on the basis of the above Figure 1 The present application also provides a possible implementation manner of a model training method. Figure 7 FIG. 13 is a flowchart of a model training method provided in a fourth embodiment of the present application; as Figure 7 shown, step 105: Generate a sample target two-dimensional image in a preset target viewpoint in a preset three-dimensional scene according to the target projected image features of multiple spatial sampling points at multiple preset source viewpoints, including:

[0116] Step 701: Aggregate the target projected image features of each spatial sampling point at multiple preset source viewpoints to generate the target aggregated feature of each spatial sampling point;

[0117] In a possible implementation manner, all the stitched projected image features can be input into a feature aggregation network together (this feature aggregation network can be, for example, a preset neural network or a network capable of performing average pooling operations, and the present application does not limit this), and the target aggregated feature of each spatial sampling point is obtained.

[0118] The above is only an example description. In actual implementation, there may be other ways to generate the target aggregated feature of the target projected image features, and the present application does not limit this.

[0119] Step 702: Generate a sample target two-dimensional image in a preset target viewpoint in a preset three-dimensional scene according to the target aggregated features of multiple spatial sampling points.

[0120] In a possible implementation, based on the target aggregation features of multiple spatial sampling points, the color, density, etc. of each spatial sampling point can be further predicted, and then a sample target two-dimensional image in a preset three-dimensional scene from a preset target perspective can be generated.

[0121] Optionally, on the basis of the above Figure 1 A possible implementation of a model training method is further provided in this application. Figure 8 FIG. is a flowchart of a model training method provided in a fifth embodiment of this application; as Figure 8 shown, step 105 generates a sample target two-dimensional image in a preset three-dimensional scene from a preset target perspective according to the target aggregation features of multiple spatial sampling points, including:

[0122] Step 801: Determine rendering parameters corresponding to multiple spatial sampling points according to the target aggregation features of multiple spatial sampling points;

[0123] In a possible implementation, based on the target aggregation features of multiple spatial sampling points The spatial color and density (i.e., rendering parameters) corresponding to each spatial sampling point can be predicted through multiple fully connected layers.

[0124] It should be noted that in addition to spatial color and density, there may be other rendering parameters, which are not limited in this application.

[0125] Step 802: Perform image rendering according to the rendering parameters corresponding to multiple spatial sampling points to generate a sample target two-dimensional image.

[0126] In a possible implementation, a stereoscopic rendering method can be used to accumulate and integrate the colors, densities, etc. corresponding to the rendering parameters of different spatial sampling points along the light direction, and finally generate a sample target two-dimensional image.

[0127] This application also provides a perspective image generation method. Figure 9 FIG. is a flowchart of a perspective image generation method provided in an embodiment of this application; as Figure 9 shown, this method includes:

[0128] Step 901: Obtain multiple source perspective images in a preset three-dimensional scene from multiple source perspectives.

[0129] It should be noted that the multiple source perspective images are two-dimensional images from multiple source perspectives, and the multiple source perspectives are multiple different perspectives with a perspective deviation within a preset perspective deviation range from the target perspective.

[0130] In a possible implementation, multiple source perspective images of multiple different perspectives with a perspective deviation from the target perspective within a preset perspective deviation range in a preset three-dimensional scene are obtained. It should be noted that the number of source perspective images obtained in this application is not limited, and the user can select according to actual needs. In addition, the number of source perspective images may be the same as or different from the number of sample source perspective images used during the training of the perspective image generation model, and this application does not limit this.

[0131] Step 902: According to the multiple source perspective images and the target perspective, use a pre-trained perspective image generation model to generate a target two-dimensional image in the target perspective in the preset three-dimensional scene, where the perspective image generation model is a model trained according to any embodiment of the above model training method.

[0132] Taking the multiple source perspective images and the target perspective as inputs and inputting them into the perspective image generation model trained by the above model training method, a target two-dimensional image in the target perspective in the preset three-dimensional scene can be obtained.

[0133] It should be noted that the input target perspective can be parameters such as the position and pose of the target perspective. This application does not limit this, as long as the input parameters can uniquely determine the target perspective.

[0134] Optionally, on the basis of the above Figure 1 A possible implementation of a model training method is further provided in this application. The multiple preset source perspectives are: multiple different perspectives with a perspective deviation from the preset target perspective within a preset perspective deviation range.

[0135] In a possible implementation, the perspective deviation between multiple preset source perspectives and a preset target perspective is within a preset perspective deviation range, which can be defined in the following way: the angle between the line connecting the preset source perspective to the target observation object and the line connecting the preset target perspective to the target observation object is within the preset perspective deviation range. For example, the target observation object is the bucket of an excavator, and the target perspective can be, for example, a determined target observation perspective of the bucket. For simplicity of explanation, a line can be determined between this target observation perspective of the bucket and the bucket of the excavator, and this line can be used as a reference line. There may be other multiple observation perspectives for this excavator bucket. These observation perspectives can be, for example, the shooting perspectives of a camera taking two-dimensional images of the target observation object. Lines can also be determined between these observation perspectives and the excavator bucket. Among these lines, the observation perspectives corresponding to the lines whose angle (or the absolute value of the angle) with the reference line is less than the preset perspective deviation (or within the preset perspective deviation range) can be the preset source perspectives. For example, if the preset perspective deviation range is from -10 degrees to +10 degrees, then when the angle between the lines of other multiple observation perspectives and the reference line is within -10 degrees to +10 degrees (or the absolute value of the angle is less than 10 degrees), the observation perspectives corresponding to the lines within the preset perspective deviation range can be the preset source perspectives.

[0136] It can be understood that if the above method is used to determine the preset source perspective, the smaller the preset perspective deviation range, the greater the overlap between the observation range of the preset target perspective on the target observation object and the observation range of the preset source perspective on the target observation object.

[0137] Optionally, on the basis of the above Figure 1 This application also provides a possible implementation of a model training method, which includes obtaining multiple original projection pixel coordinates and multiple original projection image features of each spatial sampling point under the preset target perspective:

[0138] According to the preset target perspective and the multiple sample source perspective images, the multiple original projection pixel coordinates and the multiple original projection image features are obtained.

[0139] In a possible implementation, for each spatial sampling point P on the light ray emitted from each pixel point under the target perspective, first, each spatial sampling point P is projected onto the N sample source perspective images corresponding to the N preset source perspectives respectively according to the preset source perspective (or the camera pose of the preset source perspective) to obtain the original projection pixel coordinates of each sampling point P under N different preset source perspectives Furthermore, the original projection image features at the positions corresponding to the original projection pixel coordinates are obtained For example, for a spatial sampling point with coordinates (X, Y, Z), since each preset source view is known (that is, the position and pose of each preset source view are known), the original projected pixel coordinates of the spatial sampling point can be obtained by projecting the spatial sampling point onto each preset source view.

[0140] In another possible implementation, the original projected pixel coordinates of each 3D sampling point P under N different preset source views are obtained. After that, sampling can be performed at the spatial position corresponding to the sample source view image according to the original projected pixel coordinates to obtain the original projected image features at the position corresponding to the original projected pixel coordinates.

[0141] In yet another possible implementation, multiple new projected image features can also be generated according to the source image features corresponding to multiple sample source view images and multiple new two-dimensional projected pixel coordinates. Among them, the source image features corresponding to multiple sample source view images can be obtained by performing feature extraction on each sample source view image. Thus, according to the preset target view and the source image features corresponding to multiple sample source view images, multiple original projected pixel coordinates and multiple original projected image features of each spatial sampling point under the preset target view are generated.

[0142] In a specific implementation, for N sample source view images (two-dimensional images) in a preset three-dimensional scene An image feature extraction network with weight sharing can be used to separately extract the corresponding source image features from each input sample source view image. The source image features may include semantic features under the preset source view, etc. (It should be noted that generally, after performing source image feature extraction on the sample source view image, the spatial resolution of the extracted source image features decreases and the depth increases. Thus, using the source image features in subsequent embodiments can, on the one hand, filter out the unnecessary image features in the sample source view image and, on the other hand, reduce the amount of computation and speed up the computation speed); then, according to the preset source view (or the camera pose of the preset source view), each spatial sampling point P is respectively projected onto the source image features extracted from the N sample source view images corresponding to the above N preset source views to obtain the original projected pixel coordinates of each sampling point P under N different preset source views. Furthermore, the original projected image features at the positions corresponding to the original projected pixel coordinates are obtained.

[0143] The above is only an example. In actual implementation, there may be other ways to obtain the original projected image features, and the present application does not limit this.

[0144] The perspective image generation method implemented based on the above model training method effectively reduces and relaxes the requirements and limitations on the number of source perspective images during the new perspective generation process, and significantly improves the generation quality and clarity of the target image when the number of source images is small.

[0145] The following describes the training device, perspective image generation device, electronic device, storage medium, etc. for executing the model provided in this application. For the specific implementation process and technical effects, refer to the above, and will not be elaborated below.

[0146] An exemplary implementation of a model training device provided in an embodiment of this application can execute the model training method provided in the above embodiment. Figure 10 It is a schematic diagram of a model training device provided in an embodiment of this application. As Figure 10 shown, the above model training device 100 includes: an acquisition module 11, an original image feature generation module 13, a new projection feature generation module 15, a target projection image feature generation module 17, a target two-dimensional image generation module 18, and a perspective image generation module 19;

[0147] The acquisition module 11 is configured to acquire a sample preset two-dimensional image at a preset target perspective and multiple sample source perspective images at multiple preset source perspectives in a preset three-dimensional scene;

[0148] The original image feature generation module 13 is configured to acquire multiple original projection pixel coordinates and multiple original projection image features at each spatial sampling point at the preset target perspective, where the multiple original projection pixel coordinates correspond to multiple preset source perspectives respectively, and the multiple original projection image features are the image features at the multiple original projection pixel coordinates in the multiple sample source perspective images;

[0149] The new projection feature generation module 15 is configured to generate multiple new projection pixel coordinates and multiple new projection image features at each spatial sampling point according to the multiple original projection pixel coordinates and the multiple sample source perspective images at each spatial sampling point, where the multiple new projection pixel coordinates correspond to multiple preset source perspectives respectively, and the multiple new projection image features are the image features at the multiple new two-dimensional projection pixel coordinates in the multiple sample source perspective images;

[0150] The target projection image feature generation module 17 is configured to generate target projection image features at each spatial sampling point at multiple preset source perspectives according to the multiple original projection image features and the multiple new projection image features;

[0151] The target two-dimensional image generation module 18 is configured to generate a sample target two-dimensional image at the preset target perspective in the preset three-dimensional scene according to the target projection image features at multiple spatial sampling points at multiple preset source perspectives;

[0152] A perspective image generation module 19 is configured to perform model training based on a sample target two-dimensional image at a preset target perspective and a sample preset two-dimensional image at the preset target perspective in a preset three-dimensional scene, so as to generate a perspective image generation model.

[0153] Optionally, a new projection feature generation module 15 is configured to generate multiple new two-dimensional projection pixel coordinates for each spatial sampling point according to multiple original projection pixel coordinates of each spatial sampling point; and generate multiple new projection image features according to multiple sample source perspective images and the multiple new two-dimensional projection pixel coordinates.

[0154] Optionally, the new projection feature generation module 15 is configured to generate multiple two-dimensional coordinate offsets for each spatial sampling point according to multiple original projection image features of each spatial sampling point, and the multiple two-dimensional coordinate offsets respectively correspond to multiple preset source perspectives; and generate multiple new two-dimensional projection pixel coordinates for each spatial sampling point according to the multiple original projection pixel coordinates and the multiple two-dimensional coordinate offsets.

[0155] Optionally, the new projection feature generation module 15 is configured to perform feature aggregation on multiple original projection image features of each spatial sampling point to generate an original aggregation feature of each spatial sampling point; and generate multiple two-dimensional coordinate offsets for each spatial sampling point according to the original aggregation feature and the multiple original projection image features of each spatial sampling point.

[0156] Optionally, the new projection feature generation module 15 is configured to generate multiple feature differences for each spatial sampling point according to the original aggregation feature and the multiple original projection image features of each spatial sampling point, and the multiple feature differences correspond to multiple preset source perspectives; and map the multiple feature differences to generate multiple two-dimensional coordinate offsets.

[0157] Optionally, a target projection image feature generation module 17 is configured to determine at least one new projection image feature from the multiple new projection image features; and generate target projection image features of each spatial sampling point at multiple preset source perspectives according to the multiple original projection image features and the at least one new projection image feature.

[0158] Optionally, a target two-dimensional image generation module 18 is configured to perform feature aggregation on the target projection image features of each spatial sampling point at multiple preset source perspectives to generate a target aggregation feature of each spatial sampling point; and generate a sample target two-dimensional image at a preset target perspective in a preset three-dimensional scene according to the target aggregation features of the multiple spatial sampling points.

[0159] Optionally, the target two-dimensional image generation module 18 is configured to determine rendering parameters corresponding to multiple spatial sampling points according to the target aggregation features of the multiple spatial sampling points; and perform image rendering according to the rendering parameters corresponding to the multiple spatial sampling points to generate a sample target two-dimensional image.

[0160] Optionally, the original image feature generation module 13 is configured to obtain the multiple original projection pixel coordinates and the multiple original projection image features according to the preset target viewing angle and the multiple sample source view images.

[0161] An exemplary implementation of the perspective image generation device provided in an embodiment of the present application can execute the perspective image generation method provided in the foregoing embodiment. Figure 11 The following is a schematic diagram of a perspective image generation device provided in an embodiment of the present application. As Figure 11 shown, the above perspective image generation device 300 includes: an image acquisition module 21 and an image generation module 23:

[0162] The image acquisition module 21 is configured to acquire multiple source view images of a preset three-dimensional scene at multiple source viewing angles;

[0163] The image generation module 23 is configured to generate a target two-dimensional image of the preset three-dimensional scene at a target viewing angle according to the multiple source view images and the target viewing angle by using a pre-trained perspective image generation model, where the perspective image generation model is a model trained by using the model training method of any of the foregoing embodiments.

[0164] The above device is used to execute the method provided in the foregoing embodiment, and its implementation principle and technical effects are similar and will not be described in detail here.

[0165] These modules above can be one or more integrated circuits configured to implement the above methods, for example: one or more application specific integrated circuits (ASICs), or, one or more digital signal processors (DSPs), or, one or more field programmable gate arrays (FPGAs), etc. Again, when a certain module above is implemented in the form of a processing element scheduling program code, the processing element can be a general-purpose processor, such as a central processing unit (CPU) or other processors that can call program code. Again, these modules can be integrated together and implemented in the form of a system-on-a-chip (SOC).

[0166] An exemplary implementation of an electronic device provided by an embodiment of the present application can execute the model training method or the perspective image generation method provided by the above embodiments. Figure 12 FIG. is a schematic diagram of an electronic device provided by an embodiment of the present application. The device can be integrated into a terminal device or a chip of a terminal device, and the terminal can be a computing device with data processing capabilities.

[0167] The electronic device includes: a processor 1201, a storage medium 1202, and a bus. The storage medium stores program instructions executable by the processor. When the electronic device is running, the processor communicates with the storage medium through the bus. The processor executes the program instructions to perform the steps of the above model training method or the steps of the above perspective image generation method when executed. The specific implementation manners and technical effects are similar and will not be elaborated here.

[0168] An exemplary implementation of a computer-readable storage medium provided by an embodiment of the present application can execute the model training method provided by the above embodiments. A computer program is stored on the storage medium, and when the computer program is run by a processor, it executes the steps of the above model training method or the steps of the above perspective image generation method.

[0169] A computer program stored in a storage medium may include several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor (English: processor) to execute some steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (English: Read-Only Memory, abbreviated as: ROM), a random access memory (English: Random Access Memory, abbreviated as: RAM), a magnetic disk, or an optical disc that can store program codes.

[0170] In several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0171] The unit described as a separation component may or may not be physically separated. The component displayed as a unit may or may not be a physical unit, that is, it may be located in one place or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0172] In addition, each functional unit in various embodiments of the present invention can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware, or in the form of a combination of hardware and software functional units.

[0173] The above integrated unit implemented in the form of a software functional unit can be stored in a computer-readable storage medium. The above software functional unit stored in a storage medium includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor (English: processor) to execute some steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media that can store program codes such as USB flash drives, mobile hard disks, read-only memories (English: Read-Only Memory, abbreviated as: ROM), random access memories (English: Random Access Memory, abbreviated as: RAM), magnetic disks, or optical discs.

[0174] The above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed by this application can easily think of changes or substitutions, which should all be covered within the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

Claims

1. A model training method, characterized in that, Including: Obtaining a sample preset two-dimensional image at a preset target perspective and multiple sample source perspective images at multiple preset source perspectives in a preset three-dimensional scene; Obtaining multiple original projection pixel coordinates and multiple original projection image features at each spatial sampling point at the preset target perspective, where the multiple original projection pixel coordinates respectively correspond to the multiple preset source perspectives, and the multiple original projection image features are image features at the multiple original projection pixel coordinates in the multiple sample source perspective images; Generating multiple new projection pixel coordinates and multiple new projection image features at each spatial sampling point according to the multiple original projection pixel coordinates at each spatial sampling point and the multiple sample source perspective images, where the multiple new projection pixel coordinates correspond to the multiple preset source perspectives, and the multiple new projection image features are image features at multiple new two-dimensional projection pixel coordinates in the multiple sample source perspective images; Generating target projection image features at each spatial sampling point at the multiple preset source perspectives according to the multiple original projection image features and the multiple new projection image features; Generating a sample target two-dimensional image at the preset target perspective in the preset three-dimensional scene according to the target projection image features at the multiple preset source perspectives of multiple spatial sampling points; Performing model training according to the sample target two-dimensional image at the preset target perspective in the preset three-dimensional scene and the sample preset two-dimensional image at the preset target perspective to generate a perspective image generation model.

2. The method according to claim 1, characterized in that, The generating multiple new two-dimensional projection pixel coordinates and multiple new projection image features at each spatial sampling point according to the multiple original projection pixel coordinates at each spatial sampling point and the multiple sample source perspective images includes: Generating the multiple new two-dimensional projection pixel coordinates at each spatial sampling point according to the multiple original projection pixel coordinates at each spatial sampling point; Generating the multiple new projection image features according to the multiple sample source perspective images and the multiple new two-dimensional projection pixel coordinates.

3. The method according to claim 2, characterized in that, The generating the multiple new two-dimensional projection pixel coordinates at each spatial sampling point according to the multiple original projection pixel coordinates at each spatial sampling point includes: Generating multiple two-dimensional coordinate offsets at each spatial sampling point according to the multiple original projection image features at each spatial sampling point, where the multiple two-dimensional coordinate offsets respectively correspond to the multiple preset source perspectives; Generating the multiple new two-dimensional projection pixel coordinates at each spatial sampling point according to the multiple original projection pixel coordinates and the multiple two-dimensional coordinate offsets.

4. The method according to claim 3, wherein The generating multiple two-dimensional coordinate offsets at each spatial sampling point according to the multiple original projection image features at each spatial sampling point includes: Performing feature aggregation on the multiple original projection image features at each spatial sampling point to generate an original aggregation feature at each spatial sampling point; Generate a plurality of two-dimensional coordinate offsets for each spatial sampling point based on the original aggregated feature of each spatial sampling point and the plurality of original projection image features.

5. The method according to claim 4, wherein The generating a plurality of two-dimensional coordinate offsets for each spatial sampling point based on the original aggregated feature of each spatial sampling point and the plurality of original projection image features includes: Generate a plurality of feature differences for each spatial sampling point based on the original aggregated feature of each spatial sampling point and the plurality of original projection image features, where the plurality of feature differences correspond to the plurality of preset source viewpoints; Map the plurality of feature differences to generate the plurality of two-dimensional coordinate offsets.

6. The method according to claim 1, characterized in that, The generating target projection image features for each spatial sampling point at the plurality of preset source viewpoints based on the plurality of original projection image features and the plurality of new projection image features includes: Determine at least one new projection image feature from the plurality of new projection image features; Generate target projection image features for each spatial sampling point at the plurality of preset source viewpoints based on the plurality of original projection image features and the at least one new projection image feature.

7. The method according to claim 1, wherein The generating a sample target two-dimensional image in the preset three-dimensional scene at the preset target viewpoint based on the target projection image features of the plurality of spatial sampling points at the plurality of preset source viewpoints includes: Perform feature aggregation on the target projection image features of each spatial sampling point at the plurality of preset source viewpoints to generate a target aggregated feature for each spatial sampling point; Generate a sample target two-dimensional image in the preset three-dimensional scene at the preset target viewpoint based on the target aggregated features of the plurality of spatial sampling points.

8. The method according to claim 7, wherein The generating a sample target two-dimensional image in the preset three-dimensional scene at the preset target viewpoint based on the target aggregated features of the plurality of spatial sampling points includes: Determine rendering parameters corresponding to the plurality of spatial sampling points based on the target aggregated features of the plurality of spatial sampling points; Perform image rendering according to the rendering parameters corresponding to the plurality of spatial sampling points to generate the sample target two-dimensional image.

9. The method according to claim 1, characterized in that, The plurality of preset source viewpoints are: a plurality of different viewpoints with a viewpoint deviation from the preset target viewpoint within a preset viewpoint deviation range.

10. The method according to claim 1, characterized in that, The obtaining a plurality of original projection pixel coordinates and a plurality of original projection image features for each spatial sampling point at the preset target viewpoint includes: Obtain the plurality of original projection pixel coordinates and the plurality of original projection image features according to the preset target viewpoint and the plurality of sample source viewpoint images.

11. A method for generating a perspective image, characterized in that, Includes: Obtain a plurality of source viewpoint images in a plurality of source viewpoints in a preset three-dimensional scene; Generate a target two-dimensional image in the preset three-dimensional scene at the target viewpoint according to the plurality of source viewpoint images and the target viewpoint by using a pre-trained viewpoint image generation model, where the viewpoint image generation model is a model trained by using the model training method described in any one of the above claims 1-10.

12. A model training device, characterized in that, Includes: An acquisition module, configured to acquire a sample preset two-dimensional image at a preset target view angle and multiple sample source view angle images at multiple preset source view angles in a preset three-dimensional scene; An original image feature generation module, configured to acquire multiple original projection pixel coordinates and multiple original projection image features at each spatial sampling point at the preset target view angle, wherein the multiple original projection pixel coordinates respectively correspond to the multiple preset source view angles, and the multiple original projection image features are image features at the multiple original projection pixel coordinates in the multiple sample source view angle images; A new projection feature generation module, configured to generate multiple new projection pixel coordinates and multiple new projection image features at each spatial sampling point according to the multiple original projection pixel coordinates at each spatial sampling point and the multiple sample source view angle images, wherein the multiple new projection pixel coordinates correspond to the multiple preset source view angles, and the multiple new projection image features are image features at multiple new two-dimensional projection pixel coordinates in the multiple sample source view angle images; A target projection image feature generation module, configured to generate target projection image features at each spatial sampling point at the multiple preset source view angles according to the multiple original projection image features and the multiple new projection image features; A target two-dimensional image generation module, configured to generate a sample target two-dimensional image at the preset target view angle in the preset three-dimensional scene according to the target projection image features at multiple spatial sampling points at the multiple preset source view angles; A view angle image generation module, configured to perform model training according to the sample target two-dimensional image at the preset target view angle in the preset three-dimensional scene and the sample preset two-dimensional image at the preset target view angle, and generate a view angle image generation model.

13. An angular image generation device, characterized in that, Comprising: An image acquisition module, an image generation module: The image acquisition module is configured to acquire multiple source view angle images at multiple source view angles in a preset three-dimensional scene; The image generation module is configured to generate a target two-dimensional image at the target view angle in the preset three-dimensional scene according to the multiple source view angle images and the target view angle by using a pre-trained view angle image generation model, wherein the view angle image generation model is a model trained by using the model training method according to any one of claims 1-10 above.

14. An electronic device, characterized in that, Comprising: A processor, a storage medium and a bus, wherein the storage medium stores program instructions executable by the processor. When the electronic device runs, the processor communicates with the storage medium through the bus, and the processor executes the program instructions to execute the steps of the model training method according to any one of claims 1 to 10, or execute the steps of the view angle image generation method according to claim 11.

15. A computer-readable storage medium, characterized in that, A computer program is stored on the storage medium, and when the computer program is run by the processor, it executes the steps of the model training method according to any one of claims 1 to 10, or executes the steps of the view angle image generation method according to claim 11.

Citation Information

Patent Citations

  • Model training method, three-dimensional face image generation method and equipment

    CN113838176A

  • Object rendering method and device, electronic equipment and storage medium

    CN114693853A