Image processing method and corresponding device

By acquiring multiple images and utilizing a generalized neural radiation field and a three-dimensional spatial adaptive normalized convolutional neural network, images from new perspectives are reconstructed, solving the problem of insufficient generalization ability in image processing in existing technologies, and achieving high-quality image reconstruction and image processing with expanded perspectives.

WO2025246253A1PCT designated stage Publication Date: 2025-12-04HUAWEI TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/136345
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-30
Filing Date
2024-12-03
Publication Date
2025-12-04

AI Technical Summary

Technical Problem

Existing image processing techniques based on neural radiation fields and 3D Gaussian splashing have poor generalization ability under spatial scene constraints and are difficult to process images outside of specific scenes.

Method used

By acquiring multiple first images, a second image from a new perspective is reconstructed using a generalized neural radiation field and a three-dimensional spatial adaptive normalized convolutional neural network. By combining multiple decoders to process features of the foreground, background, and sky regions, the generalization ability and quality of image reconstruction are improved.

Benefits of technology

It achieves high-quality image reconstruction from different perspectives, expands the generalization ability of image processing, and improves image clarity and rendering quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024136345_04122025_PF_FP_ABST
    Figure CN2024136345_04122025_PF_FP_ABST
Patent Text Reader

Abstract

The present application provides an image processing method, which is applicable to various scenarios requiring image processing, such as assisted driving, autonomous driving, AR, VR, high-precision mapping, and three-dimensional scene reconstruction. The method comprises: performing generalized scene reconstruction of a target feature on the basis of a plurality of first images, so as to obtain a second image at a second view angle, wherein the target feature is obtained on the basis of the first images, the target feature comprises spatial geometric features of the scene from which the plurality of first images originate, and the second view angle is different from a first view angle; and, outputting the second image. Since the present application, during scene reconstruction, uses spatial geometric features of the scene from which the first images originate, when the second image from a new view angle is reconstructed, the spatial information and image content of the new view angle can be better determined by utilizing the spatial geometric features, thereby widening the range of image reconstruction, improving the generalization capability of image processing, and improving the quality (such as clarity) of the second image.
Need to check novelty before this filing date? Find Prior Art

Description

A method and corresponding apparatus for image processing

[0001] This application claims priority to Chinese Patent Application No. 202410699256.X, filed on May 30, 2024, entitled "A Method and Apparatus for Image Processing", the entire contents of which are incorporated herein by reference. Technical Field

[0002] This application relates to the field of image processing technology, specifically to an image processing method and corresponding apparatus. Background Technology

[0003] With the development of network technology, various types of applications have emerged, many of which involve image processing. Examples include driver assistance systems, autonomous driving, augmented reality (AR), virtual reality (VR), high-precision surveying and mapping, and 3D scene reconstruction.

[0004] Currently, image processing techniques based on neural radiance fields (NeRF) or three-dimensional (3D) Gaussian splatting (3D-GS) can process images quite well. However, both of these techniques are limited by the spatial scene and can only process images in specific scenarios, exhibiting poor generalization ability.

[0005] Therefore, there is an urgent need for an image processing technology with strong generalization ability. Summary of the Invention

[0006] This application provides an image processing method to improve the generalized reconstruction capability during image processing, thereby obtaining images from more perspectives. This application also provides corresponding apparatus, computer-readable storage media, and computer program products.

[0007] The first aspect of this application provides an image processing method, comprising: acquiring multiple first images, the multiple first images corresponding to a first viewpoint; performing generalized scene reconstruction based on target features of the multiple first images to obtain a second image from a second viewpoint; wherein the target features are obtained based on the first images, the target features include spatial geometric features of the scene from which the multiple first images originate, and the second viewpoint is different from the first viewpoint; and outputting the second image.

[0008] In this application, the multiple first images may correspond to the same scene, and the viewpoint corresponding to the first image may be the viewpoint of the camera, radar, or other image acquisition device that captured the first image. The viewpoints of the multiple images may be the same or different. The first viewpoint is usually a viewpoint that is closer to the second viewpoint, or the first viewpoint is the viewpoint that is closest to the second viewpoint among the viewpoints corresponding to the multiple images.

[0009] In this application, the process of generalized scene reconstruction can be based on processing multiple first images using generalized neural radiance fields (NeRF) to obtain a second image with a new perspective (second perspective).

[0010] In this application, spatial geometric features can also be described as spatial features, including the spatial volume features of each sampling point in the space from which the first image originates. That is, features such as the position and angle of the sampling points in space are encoded and mapped to a high dimension. These high-dimensional features can improve the clarity of the reconstructed image more than the position and angle of the sampling points.

[0011] In the first aspect mentioned above, a second image from a new perspective (second perspective) can be reconstructed using multiple first images from a known viewpoint. Because the spatial geometric features of the scene from which the first image originates are used during scene reconstruction, these features can be used to better determine the spatial information and image content of the new perspective when reconstructing the second image from the new perspective. This increases the range of image reconstruction, improves the generalization ability of image processing, and also enhances the quality (e.g., sharpness) of the second image.

[0012] In one possible implementation, generalized scene reconstruction based on multiple first images includes: generalized scene reconstruction based on multiple first images and an image processing model; wherein, the image processing model includes a 3D spatially-adaptive normalization convolutional neural network (3D SPADE CNN), which is used to determine the 3D global volume features of the multiple first images, and the 3D global volume features are used to indicate the spatial geometric features of the scene represented by the multiple first images.

[0013] In one possible implementation, the 3D SPADE CNN takes a target point cloud as input and outputs three-dimensional global features; the target point cloud is determined based on depth maps of multiple first images.

[0014] In this application, the image processing model can be a convolutional neural network model, or a model combining convolutional neural networks and deep neural networks. The 3D SPADE CNN can include multiple 3D CNNs.

[0015] In this application, the target point cloud can be obtained by accumulating the depth maps corresponding to multiple first images. In this way, the target point cloud contains the spatial information of the scene from which the multiple first images originate. By performing feature extraction on the depth point cloud using 3D SPADE CNN, the 3D global feature volume, or spatial geometric features, of the scene from which the multiple first images originate can be obtained.

[0016] In this possible implementation, extracting the three-dimensional global volumetric features of multiple first images using 3D SPADE CNN can improve the generalization ability of subsequent reconstruction of images from new perspectives, as well as improve the quality of second images from new perspectives.

[0017] In one possible implementation, the image processing model further includes a first decoder, the input of which is a three-dimensional global volume feature and a near-field two-dimensional reference feature; wherein, the near-field two-dimensional reference feature is obtained based on a first reference image, the first reference image being obtained by back-projecting light sampling points from a second viewpoint onto the near-field region of the first image from the first viewpoint.

[0018] In one possible implementation, the output of the first decoder is first color information, which is used to render the second image.

[0019] In this application, foreground is a description relative to background. In an image, the image is divided into foreground region and background region according to the distance between the image and the light from the first viewpoint, or other methods that can decouple foreground and background.

[0020] In this possible implementation, the first decoder can fuse and decode the 3D global volume features and the near-field 2D reference features. In this way, the first color information output by the first decoder is associated with both the 3D global volume features and the 2D reference features, which can improve the closeness of the first color information to the real value in the scene, thereby improving the rendering quality of the second image.

[0021] In one possible implementation, the image processing model further includes a second decoder, which takes three-dimensional global features as input and outputs a first volume density, which is used to reconstruct the second image.

[0022] In this possible implementation, the first volume density is obtained based on three-dimensional global features, which contains richer spatial geometric information. Thus, when reconstructing the second image, a more refined image reconstruction can be performed based on the first volume density, making the reconstructed second image closer to the real image corresponding to the second viewpoint.

[0023] In one possible implementation, the image processing model further includes a third decoder, the input of which is a distant two-dimensional reference feature; wherein, the distant two-dimensional reference feature is obtained based on a second reference image, the second reference image being obtained by back-projecting the light sampling points of the second viewpoint onto the distant region of the first image of the first viewpoint.

[0024] In one possible implementation, the output of the third decoder is the second volume density and the second color information; the second volume density is used to reconstruct the second image, and the second color information is used to render the second image.

[0025] In this possible implementation, since the depth of the distant view is often difficult to estimate, the two-dimensional reference features of the distant view can be directly determined using the second reference image, and the second volume density and second color information can be obtained by decoding based on the two-dimensional reference features of the distant view. This can improve the speed of second image reconstruction.

[0026] In one possible implementation, the image processing model further includes a fourth decoder, which takes a two-dimensional sky reference feature as input and outputs sky color information; wherein the two-dimensional sky reference feature is obtained based on the sky region in the second reference image, and the sky color information is used to render the second image.

[0027] In this possible implementation, because street scenes always contain an infinite sky region where light does not collide with any physical objects, the appearance of the sky changes minimally as one moves forward. Therefore, the sky's two-dimensional reference features can be decoded using a fourth decoder to obtain the sky's color information. This can improve the speed of second image reconstruction.

[0028] In one possible implementation, the image processing model is a zero-shot model with zero adjustment and a few-shot model with fine adjustment. The zero-shot model is not adjusted by multiple images of the scene from which the first image originates, while the few-shot model is finely adjusted by multiple images of the scene from which the first image originates.

[0029] A second aspect of this application provides a method for model training, comprising:

[0030] Acquire training samples, which include multiple sample pairs, wherein each sample pair includes an image from a first viewpoint and an image from a second viewpoint, the first viewpoint and the second viewpoint being different;

[0031] The first image processing model is trained by using images from the first perspective in multiple sample pairs as inputs and images from the second perspective as outputs. The second image processing model is obtained by training the first image processing model. The first image processing model is a model based on the generalized NeRF architecture.

[0032] In this application, the first viewpoint can be different in different sample pairs, and the second viewpoint is usually close to the first viewpoint, or the difference between the two viewpoints is within a certain range. There can be multiple images of the first viewpoint in a sample pair.

[0033] In this second aspect, a second image processing model can be obtained by training a first image processing model based on a generalized NeRF architecture. Thus, during inference, the second image processing model can reconstruct images from new perspectives using images from the first perspective, thereby increasing the range of image reconstruction, improving the generalization ability of image processing, and also enhancing the quality of the reconstructed images.

[0034] In one possible implementation, the first image processing model includes a 3D SPADE CNN, which takes a target point cloud as input and outputs three-dimensional global features; wherein the target point cloud is determined based on a depth map from a first-view perspective.

[0035] In one possible implementation, the first image processing model further includes a first decoder, the input of which is a three-dimensional global volume feature and a near-field two-dimensional reference feature; wherein, the near-field two-dimensional reference feature is obtained based on a first reference image, and the first reference image is obtained by back-projecting light sampling points from a second viewpoint onto the near-field region of the first image from the first viewpoint.

[0036] In one possible implementation, the output of the first decoder is first color information, which is used to render the second image.

[0037] In one possible implementation, the first image processing model further includes a second decoder, the second decoder taking three-dimensional global features as input and outputting a first volume density, which is used to reconstruct the image from a second perspective.

[0038] In one possible implementation, the first image processing model further includes a third decoder, the input of which is a distant two-dimensional reference feature; wherein the distant two-dimensional reference feature is obtained based on a second reference image, the second reference image being obtained by back-projecting light sampling points from a second viewpoint onto the distant region of the first image from a first viewpoint.

[0039] In one possible implementation, the output of the third decoder is a second volume density and a second color information; the second volume density is used to reconstruct the second image, and the second color information is used to render the image from the second perspective.

[0040] In one possible implementation, the first image processing model further includes a fourth decoder, the input of which is a two-dimensional sky reference feature, and the output is the sky color information; wherein, the two-dimensional sky reference feature is obtained based on the sky region in the second reference image, and the sky color information is used to render the image from the second perspective.

[0041] A third aspect of this application provides a computer device, comprising:

[0042] The acquisition unit is used to acquire multiple first images, and the viewpoints corresponding to the multiple first images include the first viewpoint;

[0043] The processing unit is used to perform generalized scene reconstruction based on multiple first images to obtain a second image from a second perspective; wherein the second perspective is different from the first perspective.

[0044] The output unit is used to output the second image.

[0045] In one possible implementation, the processing unit is specifically used to perform generalized scene reconstruction based on multiple first images and an image processing model; wherein, the image processing model includes 3D SPADE CNN, which is used to determine the three-dimensional global volume features of the multiple first images, and the three-dimensional global volume features are used to indicate the spatial geometric features of the scene represented by the multiple first images.

[0046] In one possible implementation, the 3D SPADE CNN takes a target point cloud as input and outputs three-dimensional global features; the target point cloud is determined based on depth maps of multiple first images.

[0047] In one possible implementation, the image processing model further includes a first decoder, the input of which is a three-dimensional global volume feature and a near-field two-dimensional reference feature; wherein, the near-field two-dimensional reference feature is obtained based on a first reference image, the first reference image being obtained by back-projecting light sampling points from a second viewpoint onto the near-field region of the first image from the first viewpoint.

[0048] In one possible implementation, the output of the first decoder is first color information, which is used to render the second image.

[0049] In one possible implementation, the image processing model further includes a second decoder, which takes three-dimensional global features as input and outputs a first volume density, which is used to reconstruct the second image.

[0050] In one possible implementation, the image processing model further includes a third decoder, the input of which is a distant two-dimensional reference feature; wherein, the distant two-dimensional reference feature is obtained based on a second reference image, the second reference image being obtained by back-projecting the light sampling points of the second viewpoint onto the distant region of the first image of the first viewpoint.

[0051] In one possible implementation, the output of the third decoder is the second volume density and the second color information; the second volume density is used to reconstruct the second image, and the second color information is used to render the second image.

[0052] In one possible implementation, the image processing model further includes a fourth decoder, which takes a two-dimensional sky reference feature as input and outputs sky color information; wherein the two-dimensional sky reference feature is obtained based on the sky region in the second reference image, and the sky color information is used to render the second image.

[0053] In one possible implementation, the image processing model is a zero-shot model with zero adjustment and a few-shot model with fine adjustment. The zero-shot model is not adjusted by multiple images of the scene from which the first image originates, while the few-shot model is finely adjusted by multiple images of the scene from which the first image originates.

[0054] A fourth aspect of this application provides a computer device, comprising:

[0055] The acquisition unit is used to acquire training samples, which include multiple sample pairs, wherein each sample pair includes an image from a first viewpoint and an image from a second viewpoint, the first viewpoint and the second viewpoint being different;

[0056] The processing unit is used to take the first viewpoint image from multiple sample pairs as the input of the first image processing model and the second viewpoint image as the output of the first image processing model to train the first image processing model to obtain the second image processing model; wherein, the first image processing model is a model based on the generalized NeRF architecture.

[0057] In one possible implementation, the first image processing model includes a 3D SPADE CNN, which takes a target point cloud as input and outputs three-dimensional global features; wherein the target point cloud is determined based on a depth map from a first-view perspective.

[0058] In one possible implementation, the first image processing model further includes a first decoder, the input of which is a three-dimensional global volume feature and a near-field two-dimensional reference feature; wherein, the near-field two-dimensional reference feature is obtained based on a first reference image, and the first reference image is obtained by back-projecting light sampling points from a second viewpoint onto the near-field region of the first image from the first viewpoint.

[0059] In one possible implementation, the output of the first decoder is first color information, which is used to render the second image.

[0060] In one possible implementation, the first image processing model further includes a second decoder, the second decoder taking three-dimensional global features as input and outputting a first volume density, which is used to reconstruct the image from a second perspective.

[0061] In one possible implementation, the first image processing model further includes a third decoder, the input of which is a distant two-dimensional reference feature; wherein the distant two-dimensional reference feature is obtained based on a second reference image, the second reference image being obtained by back-projecting light sampling points from a second viewpoint onto the distant region of the first image from a first viewpoint.

[0062] In one possible implementation, the output of the third decoder is a second volume density and a second color information; the second volume density is used to reconstruct the second image, and the second color information is used to render the image from the second viewpoint.

[0063] In one possible implementation, the first image processing model further includes a fourth decoder, the input of which is a two-dimensional sky reference feature, and the output is the sky color information; wherein, the two-dimensional sky reference feature is obtained based on the sky region in the second reference image, and the sky color information is used to render the image from the second perspective.

[0064] A fifth aspect of this application provides a computer device including a processor and a computer-readable storage medium storing a computer program; the processor is coupled to the computer-readable storage medium, and the computer program, when executed by the processor, implements the method as described in the first aspect or any possible implementation thereof.

[0065] A sixth aspect of this application provides a computer device including a processor and a computer-readable storage medium storing a computer program; the processor is coupled to the computer-readable storage medium, and the computer program, when executed by the processor, implements the method as described in the second aspect above or any possible implementation thereof.

[0066] A seventh aspect of this application provides a computer-readable storage medium storing one or more computer-executable instructions, wherein when the computer-executable instructions are executed by a processor, the processor performs the method as described in the first aspect or any possible implementation thereof.

[0067] The eighth aspect of this application provides a computer-readable storage medium for storing one or more computer-executable instructions, wherein when the computer-executable instructions are executed by a processor, the processor performs the method as described in the second aspect above or any possible implementation thereof.

[0068] The ninth aspect of this application provides a computer program product storing one or more computer-executable instructions, wherein when the computer-executable instructions are executed by the processor, the processor executes the method described in the first aspect or any possible implementation thereof.

[0069] The tenth aspect of this application provides a computer program product storing one or more computer-executable instructions, wherein when the computer-executable instructions are executed by the processor, the processor executes the method described in the second aspect or any possible implementation thereof.

[0070] The eleventh aspect of this application provides a chip system including a processor for supporting a computer device in implementing the functions involved in the first aspect or any possible implementation thereof. In one possible design, the chip system may further include a memory for storing program instructions and data necessary for the computer device. This chip system may be composed of chips or may include chips and other discrete devices.

[0071] The twelfth aspect of this application provides a chip system including a processor for supporting a computer device in implementing the functions involved in the second aspect or any possible implementation thereof. In one possible design, the chip system may further include a memory for storing program instructions and data necessary for the computer device. This chip system may be composed of chips or may include chips and other discrete devices.

[0072] The technical effects of the third aspect and any possible implementation thereof, the fifth aspect, the seventh aspect, the ninth aspect and the eleventh aspect can be found in the first aspect or the technical effects of different possible implementations of the first aspect, and will not be repeated here.

[0073] The technical effects of the fourth aspect and any possible implementation thereof, the sixth aspect, the eighth aspect, the tenth aspect and the twelfth aspect can be found in the technical effects of the second aspect or different possible implementations of the second aspect, and will not be repeated here. Attached Figure Description

[0074] Figure 1A is a schematic diagram of the principle of a nerve radiation field;

[0075] Figure 1B is a schematic diagram of a vehicle scenario;

[0076] Figure 2A is a schematic diagram of the structure of a cloud system provided in an embodiment of this application;

[0077] Figure 2B is another structural schematic diagram of the cloud system provided in the embodiment of this application;

[0078] Figure 2C is a structural schematic diagram of a data center provided in an embodiment of this application;

[0079] Figure 3A is a schematic diagram of an example of image region division provided in an embodiment of this application;

[0080] Figure 3B is a structural schematic diagram of an image processing model provided in an embodiment of this application;

[0081] Figure 3C is a schematic diagram of the structure of the 3D SPADE CNN provided in the embodiment of this application;

[0082] Figure 4 is a schematic diagram of an embodiment of the model training method provided in this application;

[0083] Figure 5 is a schematic diagram of an embodiment of the image processing method provided in this application;

[0084] Figure 6 is a schematic diagram of another embodiment of the image processing method provided in this application;

[0085] Figure 7 is a schematic diagram showing the relationship between pixels and sampling points provided in an embodiment of this application;

[0086] Figure 8 is a schematic diagram of an example of the zero-shot model provided in an embodiment of this application;

[0087] Figure 9 is a schematic diagram of an example few-shot model provided in an embodiment of this application;

[0088] Figure 10 is a schematic diagram of a computer device provided in an embodiment of this application;

[0089] Figure 11 is another structural schematic diagram of the computer device provided in an embodiment of this application;

[0090] Figure 12 is a structural schematic diagram of the computer device provided in this application. Detailed Implementation

[0091] The embodiments of this application are described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. As those skilled in the art will understand, with the development of technology and the emergence of new scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.

[0092] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in a sequence other than that illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0093] This application provides an image processing method to improve the generalized reconstruction capability during image processing, thereby obtaining images from more perspectives. This application also provides corresponding apparatus, computer-readable storage media, and computer program products. These will be described in detail below.

[0094] For ease of understanding, the technical terms involved in the embodiments of this application are briefly introduced below:

[0095] 1. Image Processing Technology: This refers to the technology of processing one or more images captured by cameras or radar using artificial intelligence (AI) models to obtain images that better meet user needs. Currently, image processing technology is widely used in scenarios such as assisted driving, autonomous driving, augmented reality (AR), virtual reality (VR), high-precision surveying and mapping, and 3D scene reconstruction. Currently, advanced image processing technologies include those based on neural radiance fields (NeRF) or three-dimensional (3D) Gaussian splatting (3D-GS).

[0096] 2. Artificial Intelligence (AI): AI is the theory, methods, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a branch of computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to have the functions of perception, reasoning, and decision-making. Research in the field of artificial intelligence includes robotics, natural language processing, computer vision, decision-making and reasoning, human-computer interaction, recommendation and search, and the fundamental theories of AI. The application of artificial intelligence typically involves pre-designing AI models, then training these models with large amounts of data to obtain reasoning models suitable for different scenarios. AI models can include deep neural networks (DNNs), convolutional neural networks (CNNs), and others.

[0097] 3. NeRF: This is an emerging scheme for scene representation and image rendering. It implicitly records scene representations in a deep neural network, using the deep neural network to implicitly learn a static 3D scene, indirectly completing tasks such as 3D scene reconstruction and novel view synthesis. A schematic diagram of the NeRF algorithm can be seen in Figure 1A. As shown in Figure 1A, the NeRF can be a fully-connected neural network, for example, it can have 9 layers and 256 channels. The input of the NeRF can include the spatial location of features in the image, such as 3D coordinates (x, y, z). The input of the NeRF can also include the viewing direction, such as azimuth (θ) and pitch (φ). That is, the input of the NeRF can include 5 parameters: (x, y, z, θ, φ). The output of NeRF can include color information and density; the color information can be the three primary colors (red (r), green (g), and blue (b)). The density can be represented by σ. That is, the input of NeRF can include five parameters: (r, g, b, σ).

[0098] 4. Neural network: It can be composed of neural units, which can refer to units represented by x. sThe arithmetic unit that takes an intercept of 1 as input can output the following:

[0099] Where s = 1, 2, ..., n, n is a natural number greater than 1, W s For x s The weights are denoted by b, where b is the bias of the neural unit. f is the activation function of the neural unit, used to introduce nonlinear characteristics into the neural network to convert the input signal in the neural unit into an output signal. The output signal of this activation function can be used as the input to the next convolutional layer. The activation function can be the sigmoid function. A neural network is a network formed by connecting many of the above-mentioned individual neural units together, that is, the output of one neural unit can be the input of another neural unit. The input of each neural unit can be connected to the local receptive field of the previous layer to extract the features of the local receptive field, which can be a region composed of several neural units.

[0100] 5. Deep Neural Networks (DNNs): These can be understood as neural networks with many hidden layers. There's no specific metric for "many," and multi-layered neural networks and deep neural networks are generally the same in essence. DNNs can be categorized into three types based on their layer positions: input layers, hidden layers, and output layers. Generally, the first layer is the input layer, the last layer is the output layer, and the layers in between are hidden layers. Layers are fully connected, meaning that any neuron in the i-th layer is connected to any neuron in the (i+1)-th layer. Although DNNs seem complex, the operation of each layer is actually not complicated; it can be simply described by the following linear relationship expression: in, It is the input vector. It is the output vector. α is the offset vector, W is the weight matrix (also called coefficients), and α() is the activation function. Each layer is simply an adjustment of the input vector. The output vector is obtained after the operations in the expression. Because DNNs have many layers, the coefficients W and the offset vector... The number of such coefficients is therefore quite large. The definition of coefficient W, taking a three-layer DNN as an example, is as follows: the linear coefficient from the 4th neuron in the second layer to the 2nd neuron in the third layer is defined as... The superscript 3 represents the layer number where coefficient W resides, while the subscript corresponds to the output third layer index 2 and the input second layer index 4. In summary, the coefficients from the k-th neuron in layer L-1 to the j-th neuron in layer L are defined as... Note that the input layer does not have a W parameter. In deep neural networks, more hidden layers allow the network to better represent complex real-world situations. Theoretically, the more parameters a model has, the higher its complexity and "capacity," meaning it can accomplish more complex learning tasks.

[0101] 6. Convolutional Neural Network (CNN): A deep neural network with a convolutional structure. A CNN contains a feature extractor consisting of convolutional layers and subsampling layers. This feature extractor can be viewed as a filter, and the convolution process can be seen as using a trainable filter to convolve with an input image or a convolutional feature map. A convolutional layer is a layer of neurons in a CNN that performs convolution processing on the input signal. In a convolutional layer of a CNN, a neuron can be connected to only some of its neighboring neurons. A convolutional layer typically contains several feature planes, each composed of rectangularly arranged neural units. Neural units on the same feature plane share weights, which are the convolutional kernel. Shared weights can be understood as the way image information is extracted regardless of location. The underlying principle is that the statistical information of one part of the image is the same as that of other parts. This means that image information learned in one part can also be used in another part. Therefore, the same learned image information can be used for all locations in the image. In the same convolutional layer, multiple convolutional kernels can be used to extract different image information. Generally, the more convolutional kernels there are, the richer the image information reflected by the convolution operation.

[0102] Convolutional kernels can be initialized as matrices of random size, and during the training of a convolutional neural network, they can learn appropriate weights. Furthermore, sharing weights directly reduces the number of connections between layers in the convolutional neural network, while also lowering the risk of overfitting.

[0103] 7. Zero-shot model: This refers to a trained model that is not adaptively adjusted for new scenarios before being used for inference.

[0104] 8. Few-shot model: This refers to a trained model that undergoes minor adjustments to adapt to a new scenario before being used for inference.

[0105] The solution provided in this application includes an image processing process for reconstructing an image from a new perspective based on an image from a known perspective, and a process for training an image processing model using training samples. This image processing solution can be applied to various scenarios involving image processing. Taking a driving scenario as an example, as shown in Figure 1B, a vehicle is equipped with multiple image acquisition devices, which may include cameras or radar devices capable of acquiring images. As illustrated in Figure 1B, image acquisition devices 101 to 106 at different positions have specific perspectives and can acquire images from corresponding perspectives. For example, image acquisition device 101 can acquire an image from perspective 1, image acquisition device 102 can acquire an image from perspective 2, image acquisition device 103 can acquire an image from perspective 3, image acquisition device 104 can acquire an image from perspective 4, image acquisition device 105 can acquire an image from perspective 5, and image acquisition device 106 can acquire an image from perspective 6. Considering factors such as cost and aesthetics, it is unlikely that too many image acquisition devices will be installed on a vehicle without blind spots. Therefore, some perspectives cannot be directly acquired by the image acquisition devices. For example, how can the image of perspective 7, which is not covered by any image acquisition device, be obtained? The image processing procedure provided in this application embodiment can reconstruct the image of viewpoint 7 based on the image of the known viewpoint.

[0106] It should be noted that the image processing can be performed by the terminal device or by the cloud system. As long as the image from the known perspective is transmitted to the cloud system, the cloud system can reconstruct the image from the new perspective.

[0107] In addition, in this embodiment of the application, the model training process can be executed in a cloud system. Of course, the model training process can also be executed in a terminal device or a server.

[0108] The structure of the cloud system will be introduced below. Figure 2A is a schematic diagram of the structure of the cloud system provided in an embodiment of this application.

[0109] As shown in Figure 2A, the cloud system includes a scheduling device and multiple resource nodes. The scheduling device can communicate with these resource nodes. It can communicate with tenant clients, receive training samples from tenants, and schedule training tasks to resource nodes. This allows resource nodes to perform different stages of image processing model training, where the image processing model can be based on the NeRF architecture described in Figure 1A. Alternatively, the scheduling device can receive inference requests and images from known viewpoints from tenant clients, then assign the image reconstruction task to resource nodes for execution. Of course, the scheduling device can also perform the inference task itself. Alternatively, the scheduling device can send the inference request and images from known viewpoints to a specific resource node, which then performs the inference task based on the request and the images from the known viewpoints to reconstruct images from a new viewpoint.

[0110] The functions of the scheduling device can be implemented through software or hardware.

[0111] As an example of a software functional unit, a scheduling device may include code running on compute instances. A compute instance may include at least one of a physical host (computing device), a virtual machine, or a container. Further, the aforementioned compute instance may be one or more. For example, the scheduling device may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed within the same region or in different regions. Further, the multiple hosts / virtual machines / containers used to run the code may be distributed within the same availability zone (AZ) or in different AZs, each AZ comprising one or more geographically proximate data centers. Typically, a region may include multiple AZs.

[0112] Similarly, multiple hosts / virtual machines / containers used to run this code can be distributed within the same Virtual Private Cloud (VPC) or across multiple VPCs. Typically, a VPC is set up within a single region. Communication between two VPCs within the same region, as well as between VPCs in different regions, requires a communication gateway to be set up within each VPC to enable interconnection between VPCs.

[0113] As an example of a hardware functional unit, a scheduling device may include at least one computing device, such as a server. Alternatively, the scheduling device may be implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be implemented using a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), generic array logic (GAL), or any combination thereof.

[0114] The scheduling device includes multiple computing devices that can be distributed in the same region or in different regions. These computing devices can also be distributed within the same Availability Zone (AZ) or in different AZs. Similarly, they can be distributed within the same Virtual Private Cloud (VPC) or multiple VPCs. These multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.

[0115] The cloud system provided in this application embodiment can be a cloud service system. As shown in Figure 2B, the cloud service system includes a cloud platform and basic resources. The cloud platform includes a cloud platform manager, and the scheduling device in Figure 2A can be the cloud platform manager in Figure 2B. The basic resources can include multiple servers, where each server can include multiple resource nodes.

[0116] The resource nodes in Figures 2A and 2B can be computing device cards or virtual machines (VMs). The computing device card can be at least one of a central processing unit (CPU), a graphics processing unit (GPU), and a network processing unit (NPU).

[0117] The cloud platform manager can schedule training samples and return model training responses to tenants, such as model training progress, model training completion, or a trained image processing model. Alternatively, for inference, the cloud platform manager can receive multiple images from known viewpoints, then use the trained image processing model and the known viewpoints to reconstruct images from a new perspective and return the new perspective image to the client.

[0118] The cloud system provided in this application embodiment can be a data center. As shown in Figure 2C, the data center includes a data center management platform, an internal data center network, and multiple servers. Each server includes a hardware layer and a software layer. The hardware layer includes memory, network interface cards (NICs), processors, and disks, which are connected via a bus. The hardware layer provides the necessary hardware resources for the virtual machines in the software layer to run. The software layer includes a host operating system and multiple virtual machines. The host operating system may include a data center management platform client, which can interact with the data center management platform.

[0119] Virtualization technology mainly consists of computing virtualization and input / output (I / O) virtualization. It uses virtual machines as the granularity to share a physical server with multiple tenants, enabling tenants to use physical resources conveniently and flexibly under the premise of secure isolation, and greatly improving the utilization rate of physical resources.

[0120] Computational virtualization provides computing resources such as the server's processor and memory to virtual instances. Virtual instances are, for example, virtual machines, and in other scenarios, containers or bare metal servers.

[0121] In Figure 2C, each server uses virtualization technology to create multiple virtual machines, each of which can be understood as a resource node. The scheduling device in Figure 2A can be the data center management platform in Figure 2C.

[0122] Virtual machines can also be called cloud servers (Elastic Compute Service, ECS) or elastic instances (different cloud service providers may use different names).

[0123] The data center management platform can provide access interfaces (such as user interfaces or application programming interfaces, APIs). Tenants can use clients to remotely access these interfaces, register an account and password on the data center management platform, and log in. After successful account and password authentication, the tenant can send training samples to the platform via the client. The platform can then schedule the training samples and return a model training response to the tenant, such as training progress, training completion, or a trained image processing model. Alternatively, for inference, the platform can receive multiple images from known viewpoints. It can then use the trained image processing model and these known viewpoint images to reconstruct an image from a new viewpoint and return this new viewpoint image to the client.

[0124] The solution provided in this application is particularly applicable to outdoor scenes, but it is also applicable to indoor scenes. The difference is that outdoor scene images contain sky information, while indoor scene images typically do not. Taking an outdoor scene image as an example, as shown in Figure 3A, the image can be divided into a foreground region, a background region, and a sky region. The foreground is a description relative to the background. In an image, based on the distance between the image and the light source from the first-viewpoint, or other methods that decouple the foreground and background, the image is divided into a foreground region and a background region. The foreground region is also called the foreground area, and the background region can be called the background area. If the image does not contain sky information, it only needs to be divided into foreground and background regions. If the image is taken from a close perspective and has no background region, no division is needed, and the image only needs to be processed according to the foreground region's processing method.

[0125] The structure of the image processing model provided in this application embodiment can be understood by referring to Figure 3B. As shown in Figure 3B, the image processing model includes: a three-dimensional spatially-adaptive normalization convolutional neural network (3D SPADE CNN) 301, a first encoder 302, a second encoder 303, a third encoder 304, a first decoder 305, a second decoder 306, a third decoder 307, a fourth decoder 308, and a reconstruction rendering module 309.

[0126] The 3D SPADE CNN301 can process target point clouds, which can be obtained by accumulating the depth maps corresponding to multiple first images. In this way, the target point cloud contains spatial information of the scene from which the multiple first images originate. The 3D SPADE CNN301 performs feature extraction on the depth point cloud to obtain the 3D global feature volume, or spatial geometric features, of the scene from which the multiple first images originate.

[0127] The first encoder 302 can process the first reference image to output close-up two-dimensional reference features; wherein the first reference image is obtained by back-projecting the light sampling points of the second viewpoint onto the close-up region of the first image of the first viewpoint.

[0128] The second encoder 303 can process the second reference image to output distant two-dimensional reference features; wherein the second reference image is obtained by back-projecting the light sampling points of the second viewpoint onto the distant region of the first image of the first viewpoint.

[0129] The third encoder 304 can process the sky region in the second reference image to output a two-dimensional sky reference feature.

[0130] The first decoder 305 takes as input the 3D global volume features output by the 3D SPADE CNN 301 and the close-range 2D reference features output by the first encoder 302. After processing the 3D global volume features and the close-range 2D reference features, the first decoder 305 can output the first color information.

[0131] The input to the second decoder 306 is the three-dimensional global volume features output by the 3D SPADE CNN301. After processing the three-dimensional global volume features, the second decoder 306 can output the first volume density.

[0132] The input to the third decoder 307 is the distant two-dimensional reference feature output by the second encoder 303; after processing the distant two-dimensional reference feature, the second decoder 307 can output the second volume density and the second color information.

[0133] The input to the fourth decoder 308 is the two-dimensional sky reference feature output by the third encoder 304; after processing the two-dimensional sky reference feature, the fourth decoder 308 can output the color information of the sky.

[0134] The inputs to the reconstruction rendering module 309 are the first volume density, the first color information, the second volume density, the second color information, and the sky color information. The reconstruction rendering module 309 reconstructs the second image based on the first volume density and the second volume density, and renders the second image based on the first color information, the second color information, and the sky color information, thus obtaining the second image from the second perspective.

[0135] It should be noted that if only close-up images need to be processed, the image processing model may not include the first encoder 302, the second encoder 303, the third encoder 304, the third decoder 307, and the fourth decoder 308. If the image includes both close-up and distant areas but excludes the sky, the image processing model may not include the third encoder 304 and the fourth decoder 308.

[0136] Additionally, it should be noted that this image processing model may also include a depth processing module, a point cloud generation module, and a reference image generation module. Thus, the depth processing module can process multiple first images to obtain depth maps of multiple first images. The point cloud generation module can process the depth maps of multiple first images to obtain a target point cloud. The reference image generation module can process multiple first images and, combined with a second viewpoint, obtain a first reference image and a second reference image.

[0137] The structure of 3D SPADE CNN can be understood by referring to Figure 3C. As shown in Figure 3C, 3D SPADE CNN includes multiple point cloud feature blocks of different sizes, downsampling modules, upsampling modules, SPADE 3D blocks, normalization modules, addition modules, and multiplication modules.

[0138] Different downsampling modules can perform downsampling processing on point cloud feature blocks of the same size at different magnitudes. The SPADE 3D block can process point cloud feature blocks of the first size after 3D convolution (conv), and the upsampling module can upsample the point cloud feature blocks processed by the SPADE 3D block to obtain point cloud feature blocks of the second size. The structure of the SPADE 3D block can include 3D convolution, normalization processing, and multiplication and addition processes on the point cloud feature blocks. Through the 3D SPADE CNN described in Figure 3C, the target point cloud can be converted into three-dimensional global volumetric features.

[0139] The method for model training provided in the embodiments of this application is described below with reference to Figure 4.

[0140] As shown in Figure 4, the model training method includes:

[0141] 401. Obtain training samples.

[0142] The training samples consist of multiple sample pairs, each containing multiple images. These images include a first-view image and a second-view image, with the first and second views differing. The first-view images may differ between sample pairs, while the second-view images are typically similar to the first-view images, or their viewpoint difference is within a certain range. Each sample pair can contain multiple images of the first-view perspective. These multiple images can be either stereo or monocular images.

[0143] 402. The first image processing model is trained by taking the first viewpoint image from multiple sample pairs as the input of the first image processing model and the second viewpoint image as the output of the first image processing model to obtain the second image processing model; wherein, the first image processing model is a model based on the generalized NeRF architecture.

[0144] The model training process described in this embodiment can be understood in conjunction with the structure of the image processing model in Figure 3B. The 3DSPADE CNN 301, first encoder 302, second encoder 303, third encoder 304, first decoder 305, second decoder 306, third decoder 307, and fourth decoder 308 in Figure 3B each include a series of functional parameters. The weights of these parameters are typically set to large values ​​at the beginning of training. Then, during training, multiple images containing a first-viewpoint are processed using batches of training samples. The predicted values ​​of these multiple images and the true values ​​of the second-viewpoint images are then used to determine the loss function. The weight values ​​are gradually adjusted, and iterative training continues until the first image processing model reaches convergence, thus determining the second image processing model.

[0145] The model training scheme provided in this application provides a second image processing model by training a first image processing model based on a generalized NeRF architecture. In this way, during inference, the second image processing model can reconstruct images from new perspectives using images from the first perspective, thereby increasing the image reconstruction range, improving the generalization ability of image processing, and also improving the quality of the reconstructed images.

[0146] Figure 5 is a schematic diagram of an embodiment of the image processing method provided in this application.

[0147] As shown in Figure 5, the image processing method provided in this embodiment includes:

[0148] 501. Acquire multiple first images, the viewpoints corresponding to the multiple first images include the first viewpoint.

[0149] In this application, multiple first images can be images corresponding to the same scene, and the viewpoint corresponding to the first image can be the viewpoint of the camera, radar, or other image acquisition device that captured the first image. The viewpoints of the multiple images can be the same or different. The first image can be a monocular image or a binocular image.

[0150] 502. Based on multiple first images, generalized scene reconstruction of target features is performed to obtain a second image from a second perspective; wherein, the target features are obtained based on the first images, and the target features include the spatial geometric features of the scene from which the multiple first images originate, and the second perspective is different from the first perspective.

[0151] In this application, the first viewpoint is usually a viewpoint that is closer to the second viewpoint, or the first viewpoint is the viewpoint that is closest to the second viewpoint among the viewpoints corresponding to multiple images.

[0152] In this application, spatial geometric features can also be described as spatial features, including the spatial volume features of each sampling point in the space from which the first image originates. That is, features such as the position and angle of the sampling points in space are encoded and mapped to a high dimension. These high-dimensional features can improve the clarity of the reconstructed image more than the position and angle of the sampling points.

[0153] 503. Output the second image.

[0154] The image processing method provided in this application uses multiple first images from a known viewpoint to reconstruct a second image from a new viewpoint (second viewpoint). Because the spatial geometric features of the scene from which the first image originates are used during scene reconstruction, these spatial geometric features can be used to better determine the spatial information and image content of the new viewpoint when reconstructing the second image from the new viewpoint. This increases the image reconstruction range, improves the generalization ability of image processing, and also improves the quality (e.g., sharpness) of the second image.

[0155] Optionally, step 502 above may include: generalizing scene reconstruction based on multiple first images and an image processing model; wherein, the structure of the image processing model can be understood by referring to the description in Figure 3B.

[0156] The process of image processing using an image processing model can be understood by referring to Figure 6. As shown in Figure 6, the image processing process is illustrated using an outdoor scene as an example:

[0157] Terminal devices or cloud systems can acquire multiple first images 601, and the viewpoints corresponding to these multiple first images include the first viewpoint. Of course, the viewpoints corresponding to these multiple first images can also include other viewpoints besides the first viewpoint. For example, the multiple first images illustrated in Figure 6 are images taken on a street, including vehicles, houses, trees, and the sky. Depth estimation of these multiple first images yields depth maps 602. The depth estimation process can be performed using a depth estimation model, and the depth map can be understood as a depth point cloud. Then, depth summation of these multiple depth maps yields a target point cloud 603 representing the scene from which the multiple first images originate. This target point cloud can be a colored point cloud. Because the depth of the distant and sky regions is too deep, their spatial geometric features have little impact on the reconstructed second image. To save computation, the target point cloud 603 can be the point cloud of the near-field region from the depth maps of the multiple first images.

[0158] By processing the target point cloud 603 using a 3D SPADE CNN image processing model, the corresponding 3D global volumetric features 604 can be extracted. The 3D global volumetric features 604 can be used... This indicates, or, further processing of the 3D global volume feature 604 yields...

[0159] Because the second-view image needs to be reconstructed from multiple first images, the terminal device or cloud system can first obtain a reference image for the second view. The reference image for the second view can be obtained by back-projecting the ray sampling points of the second view (target view) onto the first image of the first view. The first view is the view with the smallest angular difference from the second view among the views corresponding to multiple first images. The first view can also be described as the nearest reference view of the second view; therefore, the reference image obtained through the above back-projection can also be described as the nearest reference image. As previously introduced, outdoor scene images can be divided into near-field and far-field regions. Therefore, a first reference image 605 can be obtained for the near-field region of the first image according to the second view, and a second reference image 606 can be obtained for the far-field region of the first image according to the second view.

[0160] For the first reference image 605, a first encoder in the image processing model can be used. Obtain close-up two-dimensional reference features A second encoder in the image processing model can be used for the second reference image 606. Obtain the two-dimensional reference features of the distant view 607. Alternatively, a third encoder in the image processing model can be used. Obtain two-dimensional reference features of the sky 608.

[0161] Will and Input to the first decoder The first color information c can be obtained. fg .Will Input to the second decoder The first volume density σ can be obtained. fg .in, The values ​​of γ(x) and d can be understood by referring to the five parameters (x, y, z, θ, φ) of NeRF in part 1A of Figure 1. γ(x) can represent the three-dimensional coordinates (x, y, z), and d can represent the azimuth (θ) and pitch (φ).

[0162] Will Input to the third decoder The second color information c can be obtained. bg Second volume density σ bg ;Will Input to the fourth decoder You can obtain the color information of the sky. sky .in, The meanings of γ(x) and d can be understood by referring to the previous introduction.

[0163] From the above c fg σ fg σ bg c bg c sky As can be seen from the relationship, compared to the NeRF shown in Figure 1A, the solution provided in this application embodiment, in addition to the three-dimensional coordinates and viewpoint direction in Figure 1A, also includes... or These features greatly improve the generalization ability of NeRF. Input

[0164] c fg σ fg σ bg c bg and c sky Inputting the color information C of the second image into the reconstruction rendering module yields the color information C of the second image. It's important to note that the second image contains multiple pixels. For each pixel, there are sampling points along the ray direction of the corresponding second viewpoint—that is, in the foreground, background, and sky. As shown in Figure 7, taking pixel 701 as an example, this pixel has multiple sampling points along the ray direction of the second viewpoint, such as sampling points 702, 703, and 704. Sampling point 702 can be understood as a foreground sampling point, sampling point 703 as a background sampling point, and sampling point 704 as a sky sampling point. Therefore, the color information of pixel 701 needs to be determined based on the color information C of sampling point 702. fg Color information c at sampling point 703 bg and the color information c of sampling point 704 sky Therefore, the color value of a pixel can be expressed as C = C (fg+bg) +(1-α (bg+fg) C sky ;in, in, c iσ represents the color information of the i-th sampling point. i This represents the volume density at the i-th sampling point. Finally, the reconstruction rendering module can use the color information C to obtain the second image.

[0165] Additionally, it's important to note that images collected from urban scenes often exhibit variable lighting or other environmental changes. To address this, an appearance encoding module can be designed into the network output to account for exposure variations. During model inference, for unseen scenes, the output of the appearance encoding module can be set to the average value of the training scene images. For the training scene, the output of the appearance encoding module is used for interpolation of the test view.

[0166] When training image processing models, multiple training losses can be used to train on different scenarios and optimize testing time on specific scenarios. For example, the most commonly used sensors in autonomous vehicles are cameras and LiDAR. Compared to cameras, LiDAR data acquisition is more expensive and may not be deployable in unseen testing scenarios. Therefore, LiDAR information can be used for training to enhance the geometric understanding of the image processing model. The training loss function and fine-tuning loss function are as follows:

[0167] Where L rgb Let L be the loss function for layer 1 (L1) or layer 2 (L2), where L is the loss function for layer 1 (L1) or layer 2 (L2). rgb It can be used to measure the difference between rendered colors and actual pixel colors. lidar Let L be the radar loss function. sky L is the sky loss function. entropy This is the entropy regularization loss function.

[0168] In this embodiment of the application, in order to separate the sky from the foreground and background areas, a pre-trained segmentation model can be used to provide a sky mask, and binary cross entropy (BCE) loss can be used to supervise the rendering of the sky mask.

[0169] During model training, LiDAR can be used as an additional sensor to further optimize the reconstruction results.

[0170] The boundary width ∈ is initialized to 0.5, and its exponential decay can be set to a minimum value of 0.1, allowing it to gradually decrease as the number of training iterations increases. The second term L... near The purpose is to increase the volume density of the model within a certain range, but without specifying its distribution, while the first and third terms need to keep the space of the remaining regions empty.

[0171] Simultaneously, in order for the model to represent the distant view as semi-transparent, entropy regularization loss is introduced. entropy =-(α) bg lnα fg +(1-α fg )ln(1-α fg ))

[0172] In addition, this application embodiment also provides a zero-shot model and a few-shot model; wherein, the zero-shot model is not adjusted by the images of the scene from which the first images are sourced, and the few-shot model is finely adjusted by the images of the scene from which the first images are sourced.

[0173] For an understanding of the zero-shot model, please refer to Figure 8. As shown in Figure 8, the "snowflake" marking in Figure 8 indicates that the image processing model does not need to be adjusted and can be directly used in the process described in the corresponding embodiment of Figure 6 above.

[0174] For an understanding of the few-shot model, please refer to Figure 9. As shown in Figure 9, the "snowflake" marker indicates that the 3D SPADE CNN, the first encoder, the second encoder, and the third encoder do not require adjustment. The "flame" marker indicates the first decoder. Third decoder and the fourth decoder Adjustments are needed. Therefore, before using the image processing model for inference, a small number of images from the first image source scene can be used to fine-tune the first decoder in the image processing model. Third decoder and the fourth decoder

[0175] The above describes the methods for image processing and model training. The following section, with reference to the accompanying diagrams, introduces the computer device.

[0176] As shown in Figure 10, the computer device 100 provided in this embodiment includes:

[0177] The acquisition unit 1001 is used to acquire multiple first images, and the viewpoints corresponding to the multiple first images include the first viewpoint;

[0178] The processing unit 1002 is used to perform generalized scene reconstruction based on multiple first images to obtain a second image from a second perspective; wherein the second perspective is different from the first perspective.

[0179] Output unit 1003 is used to output a second image.

[0180] Optionally, the processing unit 1002 is specifically used to perform generalized scene reconstruction based on multiple first images and an image processing model; wherein, the image processing model includes 3D SPADE CNN, which is used to determine the three-dimensional global volume features of the multiple first images, and the three-dimensional global volume features are used to indicate the spatial geometric features of the scene represented by the multiple first images.

[0181] Optionally, the input of 3D SPADE CNN is the target point cloud, and the output is three-dimensional global features; wherein, the target point cloud is determined based on the depth maps of multiple first images.

[0182] Optionally, the image processing model further includes a first decoder, the input of which is a three-dimensional global volume feature and a near-field two-dimensional reference feature; wherein, the near-field two-dimensional reference feature is obtained based on a first reference image, and the first reference image is obtained by back-projecting the light sampling points of the second viewpoint onto the near-field region of the first image of the first viewpoint.

[0183] Optionally, the output of the first decoder is first color information, which is used to render the second image.

[0184] Optionally, the image processing model also includes a second decoder, which takes three-dimensional global features as input and outputs a first volume density, which is used to reconstruct the second image.

[0185] Optionally, the image processing model also includes a third decoder, the input of which is a distant two-dimensional reference feature; wherein the distant two-dimensional reference feature is obtained based on a second reference image, which is obtained by back-projecting the light sampling points of the second viewpoint onto the distant region of the first image of the first viewpoint.

[0186] Optionally, the output of the third decoder is the second volume density and the second color information; the second volume density is used to reconstruct the second image, and the second color information is used to render the second image.

[0187] Optionally, the image processing model also includes a fourth decoder, the input of which is a two-dimensional sky reference feature and the output is the sky color information; wherein, the two-dimensional sky reference feature is obtained based on the sky region in the second reference image, and the sky color information is used to render the second image.

[0188] Optionally, the image processing model is a zero-shot model with zero adjustment and a few-shot model with fine adjustment; wherein, the zero-shot model is not adjusted by multiple images of the scene from which the first image originates, and the few-shot model is fine-tuned by multiple images of the scene from which the first image originates.

[0189] The functions of each unit of the computer device 100 described above can be understood by referring to the corresponding descriptions in the preceding method embodiment section, and will not be repeated here.

[0190] As shown in Figure 11, the computer device 110 provided in this embodiment includes:

[0191] The acquisition unit 1101 is used to acquire training samples, which include multiple sample pairs, wherein each sample pair includes an image from a first viewpoint and an image from a second viewpoint, the first viewpoint and the second viewpoint being different;

[0192] The processing unit 1102 is used to take the first viewpoint image from multiple sample pairs as the input of the first image processing model and the second viewpoint image as the output of the first image processing model to train the first image processing model to obtain the second image processing model; wherein, the first image processing model is a model based on the generalized NeRF architecture.

[0193] Optionally, the first image processing model includes a 3D SPADE CNN, the input of which is a target point cloud and the output is a three-dimensional global feature; wherein, the target point cloud is determined based on a depth map from a first-view perspective.

[0194] Optionally, the first image processing model further includes a first decoder, the input of which is a three-dimensional global volume feature and a near-field two-dimensional reference feature; wherein, the near-field two-dimensional reference feature is obtained based on a first reference image, and the first reference image is obtained by back-projecting the light sampling points of the second viewpoint onto the near-field region of the first image of the first viewpoint.

[0195] Optionally, the output of the first decoder is first color information, which is used to render the second image.

[0196] Optionally, the first image processing model further includes a second decoder, the second decoder taking three-dimensional global features as input and outputting a first volume density, which is used to reconstruct the image from a second perspective.

[0197] Optionally, the first image processing model further includes a third decoder, the input of which is a distant two-dimensional reference feature; wherein the distant two-dimensional reference feature is obtained based on a second reference image, the second reference image being obtained by back-projecting light sampling points from a second viewpoint onto the distant region of the first image from a first viewpoint.

[0198] Optionally, the output of the third decoder is a second volume density and a second color information; the second volume density is used to reconstruct the second image, and the second color information is used to render the image from the second viewpoint.

[0199] Optionally, the first image processing model further includes a fourth decoder, the input of which is a two-dimensional sky reference feature, and the output is the sky color information; wherein, the two-dimensional sky reference feature is obtained based on the sky region in the second reference image, and the sky color information is used to render the image from the second perspective.

[0200] The functions of each unit of the computer device 110 described above can be understood by referring to the corresponding descriptions in the preceding method embodiment section, and will not be repeated here.

[0201] Figure 12 is a schematic diagram of a possible logical structure of a computer device provided in an embodiment of this application. The computer device may be a terminal device, a server, a virtual machine, or a container. As shown in Figure 12, the computer device 1200 may include: a central processing unit 1201, a graphics processor 1202, a display device 1203 (which may not be included), and a memory 1204. Optionally, the computer device 1200 may also include at least one communication bus (not shown in Figure 12) for enabling communication between the various components.

[0202] It should be understood that the various components in computer device 1200 can also be coupled to each other via other connectors, which may include various interfaces, transmission lines, or buses. The various components in computer device 1200 can also be connected radially with the central processing unit 1201 at its center. In various embodiments of this application, coupling refers to mutual electrical connection or communication, including direct connection or indirect connection via other devices.

[0203] There are various ways to connect the central processing unit 1201 and the graphics processing unit 1202, and it is not limited to the method shown in Figure 12. The central processing unit 1201 and the graphics processing unit 1202 in the computer device 1200 can be located on the same chip, or they can be separate chips.

[0204] The functions of the central processing unit 1201, graphics processor 1202, display device 1203, and memory 1204 are briefly introduced below.

[0205] Central Processing Unit 1201: Used to run Operating System 1205 and Application Programming Interface 1206. Application Programming Interface 1206 can be a graphics application, such as a video player. Operating System 1205 provides a system graphics library interface. Application Programming Interface 1206 uses this system graphics library interface, along with drivers provided by Operating System 1205, such as user-mode and / or kernel-mode drivers for the graphics library, to generate instruction streams for rendering graphics or image frames, as well as the necessary rendering data. The system graphics library includes, but is not limited to, OpenGL ES (Open Graphics Library for Embedded Systems), the Khronos platform graphics interface, or Vulkan (a cross-platform graphics application programming interface). The instruction stream contains a series of instructions, which are typically calls to the system graphics library interface.

[0206] Optionally, the central processing unit 1201 may include at least one of the following types of processors: application processor, one or more microprocessors, digital signal processor (DSP), microcontroller unit (MCU), or artificial intelligence processor, etc.

[0207] The central processing unit 1201 may further include necessary hardware accelerators, such as application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or integrated circuits for implementing logic operations. The processor 1201 may be coupled to one or more data buses for transmitting data and instructions between various components of the computer device 1200, such as the feature extraction and feature decoding processes described in the above method embodiments.

[0208] Graphics processor 1202: Receives the graphics instruction stream sent by processor 1201, generates rendering targets through the rendering pipeline, and displays the rendering targets on display device 1203 through the operating system's layer compositing display module. The rendering pipeline, also known as the rendering pipeline, pixel pipeline, or pixel pipeline, is a parallel processing unit within the graphics processor 1202 used to process graphics signals. The graphics processor 1202 may include multiple rendering pipelines, which can process graphics signals independently and in parallel. For example, the rendering pipeline can perform a series of operations during the rendering of graphics or image frames. Typical operations may include: vertex processing, primitive processing, rasterization, fragment processing, image reconstruction, and rendering, etc.

[0209] Optionally, the graphics processor 1202 may include a general-purpose graphics processor that executes software, such as a GPU or other types of dedicated graphics processing units.

[0210] Display device 1203: Used to display various images generated by computer device 1200, which may be the graphical user interface (GUI) of the operating system or image data (including still images and video data) processed by graphics processor 1202.

[0211] Optionally, the display device 1203 may include any suitable type of display screen, such as a liquid crystal display (LCD), a plasma display, or an organic light-emitting diode (OLED) display.

[0212] Memory 1204 is a transmission channel between central processing unit 1201 and graphics processor 1202, and can be double data rate synchronous dynamic random access memory (DDR SDRAM) or other types of cache.

[0213] In another embodiment of this application, a computer-readable storage medium is also provided, which stores computer-executable instructions. When the processor of a computer device executes the computer-executable instructions, the computer device performs the steps performed by the computer device in Figures 1A to 9.

[0214] In another embodiment of this application, a computer program product is also provided, which includes computer program code. When the computer program code is executed on a computer, the computer device performs the steps performed by the computer devices in Figures 1A to 9.

[0215] In another embodiment of this application, a chip system is also provided, comprising one or more interface circuits and one or more processors; the interface circuits and processors are interconnected via lines; the interface circuits are used to receive signals from the memory of a computer device and send signals to the processors, the signals including computer instructions stored in the memory; when the processor executes the computer instructions, the computer device performs the steps performed by the computer devices in Figures 1A to 9 above. In one possible design, the chip system may further include a memory for storing program instructions and data necessary for controlling the device. This chip system may be composed of chips or may include chips and other discrete devices.

[0216] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, or indirect coupling or communication connection between apparatuses or units, and may be electrical, mechanical, or other forms.

[0217] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0218] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated units described above can be implemented wholly or partially through software, hardware, firmware, or any combination thereof.

[0219] When the integrated unit is implemented using software, it can be implemented in whole or in part as a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).

Claims

1. A method of image processing, characterized by, The method comprises: acquiring a plurality of first images, the corresponding view angles of the plurality of first images comprising a first view angle; performing generalization scene reconstruction of a target feature based on the plurality of first images to obtain a second image of a second view angle; wherein the target feature is obtained based on the first images, the target feature comprises a spatial geometric feature of a scene from which the plurality of first images are derived, and the second view angle is different from the first view angle; outputting the second image.

2. The method of claim 1, wherein, The generalization scene reconstruction of the target feature based on the plurality of first images comprises: performing generalization scene reconstruction based on the plurality of first images and an image processing model; wherein the image processing model comprises a three-dimensional space adaptive normalization convolutional neural network (3D SPADE CNN), the 3D SPADE CNN is used to determine a three-dimensional global volume feature of the plurality of first images, and the three-dimensional global volume feature is used to indicate the spatial geometric feature of a scene represented by the plurality of first images.

3. The method of claim 2, wherein, The image processing model further comprises a first decoder, an input of the first decoder being the three-dimensional global volume feature and a close-range two-dimensional reference feature; wherein the close-range two-dimensional reference feature is obtained based on a first reference image, and the first reference image is obtained by back-projecting a close-range region of a first image of the first view angle based on a light sampling point of the second view angle.

4. The method according to claim 2 or 3, characterized in that, An input of the 3D SPADE CNN is a target point cloud, and an output of the 3D SPADE CNN is the three-dimensional global feature; wherein the target point cloud is determined based on a depth map of the plurality of first images.

5. The method according to any one of claims 2-4, characterized in that, An output of the first decoder is first color information, and the first color information is used to render the second image.

6. The method according to any one of claims 2-5, characterized in that, The image processing model further comprises a second decoder, an input of the second decoder being the three-dimensional global feature, and an output of the second decoder being a first volume density, the first volume density being used to reconstruct the second image.

7. The method according to any one of claims 2 to 6, characterized in that, The image processing model further comprises a third decoder, an input of the third decoder being a long-range two-dimensional reference feature; wherein the long-range two-dimensional reference feature is obtained based on a second reference image, and the second reference image is obtained by back-projecting a long-range region of a first image of the first view angle based on a light sampling point of the second view angle.

8. The method of claim 7, wherein, An output of the third decoder is a second volume density and second color information; the second volume density is used to reconstruct the second image, and the second color information is used to render the second image.

9. The method according to claim 7 or 8, characterized in that, The image processing model further comprises a fourth decoder, an input of the fourth decoder being a sky two-dimensional reference feature, and an output of the fourth decoder being color information of a sky; wherein the sky two-dimensional reference feature is obtained based on a sky region in the second reference image, and the color information of the sky is used to render the second image.

10. The method according to any one of claims 2-9, characterized in that, The image processing model is a zero-shot model, and is a few-shot model; wherein the zero-shot model is not adjusted based on images of a scene from which the plurality of first images are derived, and the few-shot model is fine-tuned based on images of the scene from which the plurality of first images are derived.

11. A computer apparatus, comprising: The method comprises: An acquisition unit is configured to acquire a plurality of first images, the corresponding view angles of the plurality of first images including a first view angle; A processing unit is configured to perform generalized scene reconstruction based on the plurality of first images to obtain a second image of a second view angle; wherein the second view angle is different from the first view angle; An output unit is configured to output the second image.

12. The computer device of claim 11, wherein, The processing unit is configured to perform generalized scene reconstruction based on the plurality of first images and an image processing model; wherein the image processing model includes a three-dimensional space adaptive normalization convolutional neural network (3D SPADE CNN), the 3D SPADE CNN being configured to determine three-dimensional global volume features of the plurality of first images, the three-dimensional global volume features being configured to indicate the spatial geometric features of a scene represented by the plurality of first images.

13. The computer device of claim 11, wherein, The image processing model further includes a first decoder, an input of the first decoder being the three-dimensional global volume features and near-view two-dimensional reference features; wherein the near-view two-dimensional reference features are obtained based on a first reference image, the first reference image being obtained by back-projecting a near-view region of a first image of the first view angle based on a light sampling point of the second view angle.

14. A computer device, comprising: A computer readable storage medium having a computer program stored therein; The processor is coupled to the computer readable storage medium, and the computer program is executed by the processor to implement the method of any one of claims 1-10.

15. A computer readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the method of any one of claims 1-10.

16. A computer program product, characterised in that, The computer program product includes computer program code, when the computer program code is run on a computer device, causing the computer device to execute the method of any one of claims 1-10.

17. A chip system, characterized by The computer program product includes a processor, the processor being invoked to execute the method of any one of claims 1-10. The computer program product includes a processor, the processor being invoked to execute the method of any one of claims 1-10.

Citation Information

Patent Citations

  • Image rendering method and device based on neural radiation field, and electronic equipment

    CN113592991A

  • New view angle reconstruction method and training method and device of new view angle reconstruction network

    CN116681818A

  • Generalized new view angle synthesis method based on improved NeRF

    CN117496060A

  • New view synthesis method and system based on real-time rendering generalizable neural radiation field

    CN117635801A

  • Generating synthesized digital images utilizing a multi-resolution generator neural network

    US20230053588A1