Image processing method and corresponding device

By using generalized neural radiation fields and 3D SPADE CNN technology, target features are generalized for scene reconstruction using multiple images, which solves the problem of insufficient generalization ability in image processing in existing technologies and achieves image reconstruction with a wider range and higher quality.

WO2025246253A9PCT designated stage Publication Date: 2026-01-29HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/136345
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-30
Filing Date
2024-12-03
Publication Date
2026-01-29

AI Technical Summary

Technical Problem

Existing image processing techniques based on neural radiation fields and 3D Gaussian splashing have poor generalization ability under spatial scene constraints and cannot effectively process images outside of specific scenes.

Method used

By acquiring multiple first images, a generalized scene reconstruction of target features is performed using a generalized neural radiation field and a 3D spatial adaptive normalized convolutional neural network (3D SPADE CNN). The image is reconstructed from a new perspective, and multiple decoders are combined to process features of the foreground, background, and sky regions to improve image quality.

Benefits of technology

It improves the generalization ability of image processing and the quality of reconstructed images, enabling better determination of spatial information from new perspectives and increasing the range and clarity of image reconstruction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024136345_29012026_PF_FP_ABST
    Figure CN2024136345_29012026_PF_FP_ABST
Patent Text Reader

Abstract

The present application provides an image processing method, which is applicable to various scenarios requiring image processing, such as assisted driving, autonomous driving, AR, VR, high-precision mapping, and three-dimensional scene reconstruction. The method comprises: performing generalized scene reconstruction of a target feature on the basis of a plurality of first images, so as to obtain a second image at a second view angle, wherein the target feature is obtained on the basis of the first images, the target feature comprises spatial geometric features of the scene from which the plurality of first images originate, and the second view angle is different from a first view angle; and, outputting the second image. Since the present application, during scene reconstruction, uses spatial geometric features of the scene from which the first images originate, when the second image from a new view angle is reconstructed, the spatial information and image content of the new view angle can be better determined by utilizing the spatial geometric features, thereby widening the range of image reconstruction, improving the generalization capability of image processing, and improving the quality (such as clarity) of the second image.
Need to check novelty before this filing date? Find Prior Art

Description

Method and corresponding apparatus for image processing

[0001] The present application claims priority to the Chinese patent application No. 202410699256.X, filed on May 30, 2024, and entitled "Method and corresponding apparatus for image processing", the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD

[0002] The present application relates to the technical field of image processing, in particular to a method and corresponding apparatus for image processing. BACKGROUND

[0003] With the development of network technology, various types of applications have emerged, many of which involve image processing. For example, scenarios such as assisted driving, autonomous driving, augmented reality (AR), virtual reality (VR), high-precision mapping, and three-dimensional scene reconstruction all involve image processing.

[0004] Currently, image processing techniques based on neural radiance fields (NeRF) or three-dimensional (3D) Gaussian splatting (3D-GS) can handle images well. However, both of these image processing techniques are limited by the spatial scene and can only process images of specific scenes, with poor generalization ability.

[0005] Therefore, there is an urgent need for an image processing technique with strong generalization ability. SUMMARY

[0006] The present application provides a method for image processing to improve the generalization reconstruction ability in the image processing process, thereby obtaining more images of different perspectives. The present application also provides corresponding apparatus, computer-readable storage medium, and computer program product, etc.

[0007] The first aspect of the present application provides a method for image processing, comprising: obtaining a plurality of first images, the perspectives corresponding to the plurality of first images including a first perspective; performing generalization scene reconstruction of a target feature based on the plurality of first images to obtain a second image of a second perspective; wherein the target feature is obtained based on the first image, the target feature includes the spatial geometric feature of the scene from which the plurality of first images are derived, and the second perspective is different from the first perspective; and outputting the second image.

[0008] In the present application, the plurality of first images can be images corresponding to the same scene, and the first image corresponds to the perspective of the camera, radar or other image acquisition device used to capture the first image. The perspectives of the plurality of images can be the same or different. The first perspective is usually a perspective that is close to the second perspective, or the first perspective is the perspective that is closest to the second perspective among the perspectives corresponding to the plurality of images.

[0009] In the present application, the process of generalization scene reconstruction can be to process the plurality of first images based on a generalization neural radiance field (NeRF) to obtain a second image of a new perspective (second perspective).

[0010] In the present application, the spatial geometric features can also be described as spatial features, including the spatial volume features of each sampling point in the space of the scene from which the first image is derived, that is, the features of the positions and angles of the sampling points in the space are mapped to high-dimensional features through encoding. These high-dimensional features can improve the clarity of the reconstructed image compared to the positions and angles of the sampling points.

[0011] In the above first aspect, the plurality of first images with known perspectives can be used to reconstruct a second image of a new perspective (second perspective). Because the spatial geometric features of the scene from which the first image is derived are used in the process of scene reconstruction, the spatial information of the new perspective and the image content of the new perspective can be better determined using the spatial geometric features when reconstructing the second image of the new perspective, thereby increasing the image reconstruction range, improving the generalization ability of image processing, and also improving the quality (such as clarity) of the second image.

[0012] In one possible implementation, the generalization scene reconstruction of target features based on the plurality of first images includes: performing generalization scene reconstruction of target features based on the plurality of first images and an image processing model; wherein the image processing model includes a 3dimension spatially-adaptive normalization convolutional neural network (3D SPADE CNN), and the 3D SPADE CNN is used to determine a three-dimensional global volume feature of the plurality of first images, and the three-dimensional global volume feature is used to indicate the spatial geometric features of the scene represented by the plurality of first images.

[0013] In one possible implementation, the input of the 3D SPADE CNN is a target point cloud, and the output is a three-dimensional global feature; wherein the target point cloud is determined based on a depth map of the plurality of first images.

[0014] In the present application, the image processing model can be a convolutional neural network model, or a model combining a convolutional neural network and a deep neural network. The 3D SPADE CNN can include multiple 3D CNNs.

[0015] In the present application, the target point cloud can be obtained by accumulating multiple depth maps corresponding to the first images respectively. In this way, the target point cloud contains the spatial information of the scene from which the multiple first images are derived. The 3D SPADE CNN extracts features from the depth point cloud, and can obtain the three-dimensional global feature volume (3D global feature volume) of the scene from which the multiple first images are derived, i.e., the spatial geometric features.

[0016] In this possible implementation, the 3D SPADE CNN extracts the three-dimensional global feature volume of the multiple first images, which can improve the generalization ability of the subsequent reconstruction of the image of the new view, and improve the quality of the second image of the new view.

[0017] In one possible implementation, the image processing model further includes a first decoder, and the input of the first decoder is the three-dimensional global feature volume and the near-view two-dimensional reference feature. The near-view two-dimensional reference feature is obtained based on the first reference image, and the first reference image is obtained by back-projecting the light sampling points of the second view to the near-view region of the first image of the first view.

[0018] In one possible implementation, the output of the first decoder is the first color information, and the first color information is used for rendering the second image.

[0019] In the present application, the near view is described with respect to the far view. In an image, the image is divided into a near-view region and a far-view region according to the distance relationship between the image and the light of the first view, or other ways that can decouple the near view and the far view.

[0020] In this possible implementation, the first decoder can fuse and decode the three-dimensional global feature volume and the near-view two-dimensional reference feature. In this way, the first color information output by the first decoder is associated with both the three-dimensional global feature volume and the two-dimensional reference feature, which can improve the closeness of the first color information to the true value in the scene, thereby improving the rendering quality of the second image.

[0021] In one possible implementation, the image processing model further includes a second decoder, and the input of the second decoder is the three-dimensional global feature volume, and the output of the second decoder is the first volume density, and the first volume density is used for reconstructing the second image.

[0022] In the possible implementation, the first volume density is obtained based on the three-dimensional global feature, which contains richer spatial geometric information. In this way, when the second image is reconstructed, more accurate image reconstruction can be performed based on the first volume density, so that the reconstructed second image is closer to the real image corresponding to the second view.

[0023] In a possible implementation, the image processing model further includes a third decoder, an input of the third decoder being a long-range two-dimensional reference feature; wherein the long-range two-dimensional reference feature is obtained based on a second reference image, the second reference image being obtained by back-projecting the light sampling points of the second view to a long-range region of the first image of the first view.

[0024] In a possible implementation, an output of the third decoder is a second volume density and second color information; the second volume density is used for reconstructing the second image, and the second color information is used for rendering the second image.

[0025] In the possible implementation, because the depth of the long-range region is usually difficult to estimate, the long-range two-dimensional reference feature can be directly determined using the second reference image, and the second volume density and the second color information are obtained based on the long-range two-dimensional reference feature. In this way, the speed of the second image reconstruction can be improved.

[0026] In a possible implementation, the image processing model further includes a fourth decoder, an input of the fourth decoder being a sky two-dimensional reference feature, and an output of the fourth decoder being color information of the sky; wherein the sky two-dimensional reference feature is obtained based on a sky region in the second reference image, and the color information of the sky is used for rendering the second image.

[0027] In the possible implementation, because the street scene always contains an infinite sky region in which the light will not collide with any physical object, resulting in the appearance of the sky changing the least when advancing, the sky two-dimensional reference feature can be decoded by the fourth decoder to obtain the color information of the sky. In this way, the speed of the second image reconstruction can be improved.

[0028] In a possible implementation, the image processing model is a zero-shot model, and the few-shot model is a fine-tuned model; wherein the zero-shot model is not adjusted by images of a scene from which the first images are derived, and the fine-tuned model is fine-tuned by the images of the scene from which the first images are derived.

[0029] The second aspect of the present application provides a model training method, including:

[0030] obtaining training samples, the training samples including a plurality of sample pairs, wherein each sample pair includes an image of a first view and an image of a second view, the first view being different from the second view;

[0031] The first image processing model is trained by using images from the first perspective in multiple sample pairs as inputs and images from the second perspective as outputs. The second image processing model is obtained by training the first image processing model. The first image processing model is a model based on the generalized NeRF architecture.

[0032] In this application, the first viewpoint can be different in different sample pairs, and the second viewpoint is usually close to the first viewpoint, or the difference between the two viewpoints is within a certain range. There can be multiple images of the first viewpoint in a sample pair.

[0033] In this second aspect, a second image processing model can be obtained by training a first image processing model based on a generalized NeRF architecture. Thus, during inference, the second image processing model can reconstruct images from new perspectives using images from the first perspective, thereby increasing the range of image reconstruction, improving the generalization ability of image processing, and also enhancing the quality of the reconstructed images.

[0034] In one possible implementation, the first image processing model includes a 3D SPADE CNN, which takes a target point cloud as input and outputs three-dimensional global features; wherein the target point cloud is determined based on a depth map from a first-view perspective.

[0035] In one possible implementation, the first image processing model further includes a first decoder, the input of which is a three-dimensional global volume feature and a near-field two-dimensional reference feature; wherein, the near-field two-dimensional reference feature is obtained based on a first reference image, and the first reference image is obtained by back-projecting light sampling points from a second viewpoint onto the near-field region of the first image from the first viewpoint.

[0036] In one possible implementation, the output of the first decoder is first color information, which is used to render the second image.

[0037] In one possible implementation, the first image processing model further includes a second decoder, the second decoder taking three-dimensional global features as input and outputting a first volume density, which is used to reconstruct the image from a second perspective.

[0038] In one possible implementation, the first image processing model further includes a third decoder, the input of which is a distant two-dimensional reference feature; wherein the distant two-dimensional reference feature is obtained based on a second reference image, the second reference image being obtained by back-projecting light sampling points from a second viewpoint onto the distant region of the first image from a first viewpoint.

[0039] In a possible implementation, the output of the third decoder is second volume density and second color information; the second volume density is used for reconstructing the second image, and the second color information is used for rendering the image of the second view angle.

[0040] In a possible implementation, the first image processing model further includes a fourth decoder, an input of the fourth decoder is a sky two-dimensional reference feature, and an output of the fourth decoder is color information of the sky; the sky two-dimensional reference feature is obtained based on a sky region in the second reference image, and the color information of the sky is used for rendering the image of the second view angle.

[0041] The third aspect of the present application provides a computer device, comprising:

[0042] The acquisition unit is configured to acquire a plurality of first images, and the view angles corresponding to the plurality of first images include a first view angle.

[0043] The processing unit is configured to perform generalization scene reconstruction based on the plurality of first images to obtain a second image of a second view angle; the second view angle is different from the first view angle.

[0044] The output unit is configured to output the second image.

[0045] In a possible implementation, the processing unit is specifically configured to perform generalization scene reconstruction of a target feature based on the plurality of first images and an image processing model; the image processing model includes a 3D SPADE CNN, the 3D SPADE CNN is used to determine a three-dimensional global volume feature of the plurality of first images, and the three-dimensional global volume feature is used to indicate a spatial geometric feature of a scene represented by the plurality of first images.

[0046] In a possible implementation, an input of the 3D SPADE CNN is a target point cloud, and an output of the 3D SPADE CNN is a three-dimensional global feature; the target point cloud is determined based on a depth map of the plurality of first images.

[0047] In a possible implementation, the image processing model further includes a first decoder, an input of the first decoder is the three-dimensional global volume feature and a close-range two-dimensional reference feature; the close-range two-dimensional reference feature is obtained based on a first reference image, and the first reference image is obtained by back-projecting a light sampling point of the second view angle to a close-range region of the first image of the first view angle.

[0048] In a possible implementation, an output of the first decoder is first color information, and the first color information is used for rendering the second image.

[0049] In a possible implementation, the image processing model further includes a second decoder, an input of the second decoder is the three-dimensional global feature, and an output of the second decoder is first volume density; the first volume density is used for reconstructing the second image.

[0050] In a possible implementation, the image processing model further includes a third decoder, an input of the third decoder is the far-view two-dimensional reference feature; and the far-view two-dimensional reference feature is obtained based on a second reference image, and the second reference image is obtained by back-projecting light sampling points of the second view to a far-view region of the first image of the first view.

[0051] In a possible implementation, an output of the third decoder is second volume density and second color information; the second volume density is used for reconstructing the second image, and the second color information is used for rendering the second image.

[0052] In a possible implementation, the image processing model further includes a fourth decoder, an input of the fourth decoder is sky two-dimensional reference feature, and an output of the fourth decoder is color information of the sky; the sky two-dimensional reference feature is obtained based on a sky region in the second reference image, and the color information of the sky is used for rendering the second image.

[0053] In a possible implementation, the image processing model is a zero-shot model, and the few-shot model is a fine-tuned model; the zero-shot model is not adjusted based on images of a scene from which the first images are derived, and the fine-tuned model is fine-tuned based on the images of the scene from which the first images are derived.

[0054] The fourth aspect of the present application provides a computer device, comprising:

[0055] The acquisition unit is configured to acquire training samples, the training samples comprising a plurality of sample pairs, wherein each sample pair comprises an image of a first view and an image of a second view, and the first view is different from the second view.

[0056] The processing unit is configured to take the image of the first view in the plurality of sample pairs as an input of a first image processing model, take the image of the second view as an output of the first image processing model, and train the first image processing model to obtain a second image processing model; and the first image processing model is a model based on a generalized NeRF architecture.

[0057] In a possible implementation, the first image processing model comprises a 3D SPADE CNN, an input of the 3D SPADE CNN is a target point cloud, and an output of the 3D SPADE CNN is a three-dimensional global feature; and the target point cloud is determined based on a depth map of the first view.

[0058] In a possible implementation, the first image processing model further includes a first decoder, an input of the first decoder is the three-dimensional global volume feature and near-view two-dimensional reference feature; and the near-view two-dimensional reference feature is obtained based on a first reference image, and the first reference image is obtained by back-projecting light sampling points of the second view to a near-view region of the first image of the first view.

[0059] In a possible implementation, the output of the first decoder is first color information, and the first color information is used for rendering the second image.

[0060] In a possible implementation, the first image processing model further includes a second decoder, an input of the second decoder is the three-dimensional global feature, and an output of the second decoder is a first volume density, and the first volume density is used for reconstructing an image of the second view angle.

[0061] In a possible implementation, the first image processing model further includes a third decoder, an input of the third decoder is a long-range two-dimensional reference feature, and the long-range two-dimensional reference feature is obtained based on a second reference image, and the second reference image is obtained by back-projecting a light sampling point of the second view angle to a long-range region of the first image of the first view angle.

[0062] In a possible implementation, an output of the third decoder is a second volume density and second color information, and the second volume density is used for reconstructing the second image, and the second color information is used for rendering the image of the second view angle.

[0063] In a possible implementation, the first image processing model further includes a fourth decoder, an input of the fourth decoder is a sky two-dimensional reference feature, and an output of the fourth decoder is color information of the sky, and the sky two-dimensional reference feature is obtained based on a sky region in the second reference image, and the color information of the sky is used for rendering the image of the second view angle.

[0064] The fifth aspect of the present application provides a computer device, which includes a processor and a computer readable storage medium storing a computer program; the processor is coupled with the computer readable storage medium, and the computer program is executed by the processor to implement the method in the first aspect or any possible implementation manner.

[0065] The sixth aspect of the present application provides a computer device, which includes a processor and a computer readable storage medium storing a computer program; the processor is coupled with the computer readable storage medium, and the computer program is executed by the processor to implement the method in the second aspect or any possible implementation manner.

[0066] The seventh aspect of the present application provides a computer readable storage medium storing one or more computer execution instructions, and when the computer execution instructions are executed by a processor, the processor executes the method in the first aspect or any possible implementation manner of the first aspect.

[0067] The eighth aspect of the present application provides a computer readable storage medium storing one or more computer-executable instructions that, when executed by a processor, cause the processor to perform the method according to the second aspect or any possible implementation of the second aspect.

[0068] The ninth aspect of the present application provides a computer program product storing one or more computer-executable instructions that, when executed by a processor, cause the processor to perform the method according to the first aspect or any possible implementation of the first aspect.

[0069] The tenth aspect of the present application provides a computer program product storing one or more computer-executable instructions that, when executed by a processor, cause the processor to perform the method according to the second aspect or any possible implementation of the second aspect.

[0070] The eleventh aspect of the present application provides a chip system, which includes a processor for supporting a computer device to implement the functions involved in the first aspect or any possible implementation of the first aspect. In a possible design, the chip system can further include a memory for storing necessary program instructions and data of the computer device. The chip system can be composed of a chip, or can include the chip and other discrete devices.

[0071] The twelfth aspect of the present application provides a chip system, which includes a processor for supporting a computer device to implement the functions involved in the second aspect or any possible implementation of the second aspect. In a possible design, the chip system can further include a memory for storing necessary program instructions and data of the computer device. The chip system can be composed of a chip, or can include the chip and other discrete devices.

[0072] The third aspect and any possible implementation thereof, the fifth aspect, the seventh aspect, the ninth aspect and the eleventh aspect can bring the technical effects as the first aspect or any possible implementation of the first aspect, which will not be repeated here.

[0073] The fourth aspect and any possible implementation thereof, the sixth aspect, the eighth aspect, the tenth aspect and the twelfth aspect can bring the technical effects as the second aspect or any possible implementation of the second aspect, which will not be repeated here. BRIEF DESCRIPTION OF DRAWINGS

[0074] FIG. 1A is a schematic diagram of a neural radiation field;

[0075] FIG. 1B is a schematic diagram of a vehicle scene;

[0076] FIG. 2A is a structural schematic diagram of a cloud system according to an embodiment of the present application;

[0077] FIG. 2B is another structural schematic diagram of a cloud system according to an embodiment of the present application;

[0078] FIG. 2C is a structural schematic diagram of a data center according to an embodiment of the present application;

[0079] FIG. 3A is a schematic diagram of an example of image region division according to an embodiment of the present application;

[0080] FIG. 3B is a structural schematic diagram of an image processing model according to an embodiment of the present application;

[0081] FIG. 3C is a structural schematic diagram of a 3D SPADE CNN according to an embodiment of the present application;

[0082] FIG. 4 is a schematic diagram of an embodiment of a model training method according to an embodiment of the present application;

[0083] FIG. 5 is a schematic diagram of an embodiment of an image processing method according to an embodiment of the present application;

[0084] FIG. 6 is a schematic diagram of another embodiment of an image processing method according to an embodiment of the present application;

[0085] FIG. 7 is a schematic diagram of a relationship between a pixel point and a sampling point according to an embodiment of the present application;

[0086] FIG. 8 is a schematic diagram of an example of a zero-shot model according to an embodiment of the present application;

[0087] FIG. 9 is a schematic diagram of an example of a few-shot model according to an embodiment of the present application;

[0088] FIG. 10 is a structural schematic diagram of a computer device according to an embodiment of the present application;

[0089] FIG. 11 is another structural schematic diagram of a computer device according to an embodiment of the present application;

[0090] FIG. 12 is a structural schematic diagram of a computer device according to an embodiment of the present application. DETAILED DESCRIPTION

[0091] The embodiments of the present application will be described below in conjunction with the accompanying drawings. It is obvious that the described embodiments are only a part of the embodiments of the present application, but not all the embodiments. It is known to those skilled in the art that, as technology develops and new scenarios appear, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.

[0092] The terms "first", "second", and the like in the description and in the claims of the present application and above-described drawings are used to distinguish similar objects and are not necessarily used to describe a specific sequential or chronological order. It should be understood that the data thus used can be interchanged, where appropriate, so that the embodiments described herein can be carried out in sequences other than those illustrated or described herein. Moreover, the terms "comprise" and "have", and any variations thereof, are intended to cover a non-exclusive inclusion, for example, a process, method, system, product, or apparatus that comprises a list of steps or units is not necessarily limited to those steps or units that are clearly listed, but can include other steps or units that are not clearly listed or inherent to such processes, methods, products, or apparatuses.

[0093] Embodiments of the present application provide a method for image processing, which is used to improve the generalization reconstruction capability in the image processing process, so as to obtain more images of different perspectives. The present application also provides corresponding devices, computer readable storage media and computer program products, etc. The following are described in detail respectively.

[0094] For ease of understanding, the technical terms related to the embodiments of the present application are briefly introduced as follows:

[0095] 1. Image processing technology: It can be a technology of processing one or more images collected by a camera or a radar through an artificial intelligence (AI) model to obtain an image that is more in line with user needs. At present, image processing technology is widely used in assisted driving, autonomous driving, augmented reality (AR), virtual reality (VR), high-precision mapping, three-dimensional scene reconstruction, etc. At present, better image processing technology includes image processing technology based on neural radiance fields (NeRF) or image processing technology based on three-dimensional (3D) Gaussian splatting (3D-GS).

[0096] 2.Artificial intelligence (AI): AI is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, AI is a branch of computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a similar way to human intelligence. AI is the study of the design principles and implementation methods of various intelligent machines, enabling machines to have perception, reasoning and decision-making functions. The research in the field of artificial intelligence includes robots, natural language processing, computer vision, decision and reasoning, human-computer interaction, recommendation and search, AI basic theory, etc. The application of artificial intelligence usually involves pre-designing an AI model, training the model with a large amount of data, and then obtaining a reasoning model suitable for different scenarios. The AI model can include deep neural network (DNN), convolutional neural network (CNN), etc.

[0097] 3.NeRF: is a new scene expression and image rendering scheme, which records the scene expression in a deep neural network through implicit expression, uses a deep neural network to implicitly learn a static three-dimensional scene, and indirectly completes tasks such as three-dimensional reconstruction of the scene, generation of novel view synthesis, etc. The NeRF algorithm schematic diagram can be understood by referring to FIG. 1A, as shown in FIG. 1A, the NeRF can be a fully-connected neural network, for example, the fully-connected neural network can have 9 layers and 256 channels. The input of the NeRF can include the spatial location of the features in the image, such as three-dimensional coordinates (x, y, z), and the input of the NeRF can also include the viewing direction, such as the azimuth angle (θ) and the elevation angle (φ). That is, the input of the NeRF can include 5 parameters, respectively (x, y, z, θ, φ). The output of the NeRF can include color information and density (denstiy); wherein the color information can be three primary colors (red (r), green (g), blue (b)). The density can be represented by σ. That is, the input of the NeRF can include 5 parameters, respectively (r, g, b, σ).

[0098] 4.Neural network: can be composed of neural units, and the neural unit can refer to an x sand intercept 1 is an operation unit for input, and the output of the operation unit can be:

[0099] where s = 1, 2, … n, n is a natural number greater than 1, W s is the weight of x s , and b is the bias of the neural unit. f is the activation function of the neural unit, which is used to introduce non-linear characteristics into the neural network to convert the input signal in the neural unit into an output signal. The output signal of the activation function can be used as the input of the next convolutional layer. The activation function can be a sigmoid function. The neural network is a network formed by connecting many single neural units as described above, that is, the output of one neural unit can be the input of another neural unit. The input of each neural unit can be connected to the local receptive field of the previous layer to extract the features of the local receptive field, and the local receptive field can be a region composed of several neural units.

[0100] 5. Deep neural network: it can be understood as a neural network with many hidden layers, and "many" here has no special measurement standard. Multi-layer neural network and deep neural network are usually the same in essence. From the position of different layers of DNN, the neural network inside DNN can be divided into three categories: input layer, hidden layer and output layer. Generally, the first layer is the input layer, the last layer is the output layer, and the number of layers in between is the hidden layer. The layers are fully connected, that is, any neuron in the i-th layer is connected to any neuron in the i+1-th layer. Although DNN looks very complex, in terms of the work of each layer, it is actually not complex. Simply put, it is as follows: where, is the input vector, is the output vector, is the offset vector, W is the weight matrix (also called coefficient), and α() is the activation function. Each layer only obtains the output vector from the input vector by the operation in the expression. Since DNN has many layers, the number of coefficients W and offset vectors is also large. Among them, the definition of the coefficient W is taken as an example with a three-layer DNN, such as: the linear coefficient of the fourth neuron of the second layer to the second neuron of the third layer is defined as The superscript 3 represents the layer number of the coefficient W, and the subscript corresponds to the third layer index 2 of the output and the second layer index 4 of the input. In summary, the coefficient of the k-th neuron of the L-1-th layer to the j-th neuron of the L-th layer is defined as Note that the input layer has no W parameters. In deep neural networks, more hidden layers allow the network to better capture the complexity of real-world situations. In theory, the more parameters a model has, the higher its complexity, and the greater its "capacity" to perform more complex learning tasks.

[0101] 6. Convolutional neural network: A deep neural network with a convolutional structure. A convolutional neural network contains a feature extractor composed of convolutional layers and subsampling layers. The feature extractor can be viewed as a filter, and the convolution process can be viewed as using a trainable filter to convolve with an input image or a convolutional feature map. A convolutional layer refers to a layer of neurons in a convolutional neural network that performs convolutional processing on an input signal. In a convolutional layer of a convolutional neural network, a neuron can only be connected to a subset of the neurons in the previous layer. A convolutional layer typically contains several feature maps, each of which can be composed of a number of rectangularly arranged neurons. The neurons in the same feature map share weights, where the shared weights are the convolutional kernels. Shared weights can be understood as the way of extracting image information regardless of location. The underlying principle is that the statistical information of a certain part of the image is the same as that of other parts. That is, the image information learned in one part can also be used in another part. Therefore, the same learned image information can be used for all locations on the image. In the same convolutional layer, multiple convolutional kernels can be used to extract different image information. Generally, the more convolutional kernels there are, the more image information the convolution operation reflects.

[0102] Convolutional kernels can be initialized as matrices of random size, and during the training of a convolutional neural network, the convolutional kernels can learn reasonable weights. In addition, the direct benefit of shared weights is to reduce the connections between the layers of a convolutional neural network, while also reducing the risk of overfitting.

[0103] 7. Zero-shot model: Refers to a trained model that does not need to be adjusted for new scenarios before being used for inference.

[0104] 8. Few-shot model: Refers to a trained model that makes minor adjustments in combination with new scenarios before being used for inference.

[0105] The scheme provided in the embodiment of the application includes an image processing process of reconstructing an image of a new view angle based on an image of a known view angle, and a process of training an image processing model using training samples. The image processing scheme can be applied to various scenarios involving image processing. Taking a driving scenario as an example, as shown in FIG. 1B, a plurality of image acquisition devices are configured on a vehicle, which can include cameras or radars and other devices that can acquire images. As shown in FIG. 1B, the image acquisition devices 101 to 106 have certain view angles and can acquire images of the corresponding view angles. For example, the image acquisition device 101 can acquire an image of view angle 1, the image acquisition device 102 can acquire an image of view angle 2, the image acquisition device 103 can acquire an image of view angle 3, the image acquisition device 104 can acquire an image of view angle 4, the image acquisition device 105 can acquire an image of view angle 5, and the image acquisition device 106 can acquire an image of view angle 6. Due to cost or aesthetic considerations, it is not possible to install too many image acquisition devices on the vehicle without dead angles, so some images of view angles cannot be directly acquired by the image acquisition devices. For example, how to obtain an image of view angle 7 that is not covered by the image acquisition devices? The image processing process provided in the embodiment of the application can be used to reconstruct the image of view angle 7 based on the images of known view angles.

[0106] It should be noted that the image processing process can be performed by a terminal device or a cloud system. The cloud system can reconstruct the image of a new view angle based on the images of known view angles.

[0107] In addition, in the embodiment of the application, the model training process can be performed in the cloud system. Of course, the model training process can also be performed in a terminal device or a server.

[0108] The structure of the cloud system will be introduced first. FIG. 2A is a structural schematic diagram of a cloud system provided in an embodiment of the application.

[0109] As shown in FIG. 2A, the cloud system includes a scheduling device and a plurality of resource nodes, and the scheduling device can communicate with the plurality of resource nodes. The scheduling device can communicate with a client of a tenant, receive training samples of the tenant, and schedule a training task to the resource nodes. Thus, the resource nodes can perform different links of the image processing model training, wherein the image processing model can be a model based on the NeRF architecture introduced in FIG. 1A. Alternatively, the scheduling device receives an inference request and an image of a known view angle sent by the client of the tenant, and then assigns an image reconstruction task to the resource nodes for execution. Of course, the scheduling device can also perform the inference task. Alternatively, the scheduling device sends the inference request and the image of the known view angle to a certain resource node, and the resource node performs the inference task according to the inference request and the image of the known view angle to reconstruct an image of a new view angle.

[0110] The functions of the scheduling apparatus can be implemented by software or hardware.

[0111] As an example of a software functional unit, the scheduling apparatus can include code running on a computing instance. The computing instance can include at least one of a physical host (computing device), a virtual machine, and a container. Further, the computing instance can be one or more. For example, the scheduling apparatus can include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers for running the code can be distributed in the same region, or in different regions. Further, the multiple hosts / virtual machines / containers for running the code can be distributed in the same availability zone (AZ), or in different AZs, each of which includes one data center or multiple data centers in close geographical proximity. Generally, one region can include multiple AZs.

[0112] Similarly, the multiple hosts / virtual machines / containers for running the code can be distributed in the same virtual private cloud (VPC), or in multiple VPCs. Generally, one VPC is set up in one region, and communication between two VPCs in the same region, or between VPCs in different regions, requires a communication gateway to be set up in each VPC to achieve interconnection between VPCs.

[0113] As an example of a hardware functional unit, the scheduling apparatus can include at least one computing device, such as a server, etc. Alternatively, the scheduling apparatus can also be a device implemented by an application-specific integrated circuit (ASIC) or a programmable logic device (PLD), etc. The PLD can be implemented by a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0114] The multiple computing devices included in the scheduling apparatus can be distributed in the same region or in different regions. The multiple computing devices included in the scheduling apparatus can be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in the scheduling apparatus can be distributed in the same VPC or in multiple VPCs. The multiple computing devices can be any combination of servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.

[0115] The cloud system provided by the embodiments of the present application can be a cloud service system. As shown in FIG. 2B, the cloud service system includes a cloud platform and basic resources. The cloud platform includes a cloud platform manager. The scheduling apparatus in FIG. 2A can be the cloud platform manager in FIG. 2B. The basic resources can include multiple servers, each of which can include multiple resource nodes.

[0116] The resource nodes in FIGS. 2A and 2B can be computing device cards or virtual machines (VMs). The computing device cards can be at least one of a central processing unit (CPU), a graphic processing unit (GPU), and a network processing unit (NPU).

[0117] The cloud platform manager can schedule training samples and return a model training response, such as a model training progress or a model training completion, or a trained image processing model, to a tenant. Of course, if an inference process is performed, the cloud platform manager can receive multiple images of known perspectives, then the cloud platform manager can use the trained image processing model and the multiple images of known perspectives to reconstruct images of new perspectives, and return the images of new perspectives to the client.

[0118] The cloud system provided by the embodiments of the present application can be a data center. As shown in FIG. 2C, the data center includes a data center management platform, a data center internal network, and multiple servers. Each server includes a hardware layer and a software layer. The hardware layer includes a memory, a network card, a processor, and a disk, which are connected through a bus. The hardware layer provides the virtual machine of the software layer with the necessary hardware resources for running. The software layer includes a host operating system and multiple virtual machines. The host operating system can include a data center management platform client and can interact with the data center management platform.

[0119] Virtualization technology, mainly composed of compute virtualization and input / output (I / O) virtualization, shares a physical server to multiple tenants with virtual machines as granularity, so that tenants can use physical resources conveniently and flexibly under the premise of security isolation, and greatly improve the utilization of physical resources.

[0120] Compute virtualization provides compute resources such as processors and memories of servers to virtual instances, such as virtual machines, and in other scenarios, containers and bare-metal servers.

[0121] Each server in FIG. 2C gets multiple virtual machines through virtualization technology, and each virtual machine can be understood as a resource node. The scheduling device in FIG. 2A can be a data center management platform in FIG. 2C.

[0122] The virtual machine can also be referred to as a cloud server (Elastic Compute Service, ECS) or an elastic instance (different cloud service providers have different names).

[0123] The data center management platform can provide an access interface (such as an interface or an application programming interface (API)), and the tenant can remotely access the access interface through a client to register an account and a password on the data center management platform and log in to the data center management platform. After the data center management platform authenticates the account and the password successfully, the tenant can send a training sample to the data center management platform through the client, and then the data center management platform can schedule the training sample and return a model training response to the tenant, such as a model training progress or a model training completion, or a trained image processing model. Of course, if an inference process is performed, the data center management platform can receive multiple images of known perspectives, and then the data center management platform can reconstruct an image of a new perspective using the trained image processing model and the multiple images of known perspectives, and return the image of the new perspective to the client.

[0124] The scheme provided in the embodiments of the present application is especially suitable for outdoor scenes, and of course, is also suitable for indoor scenes. The difference is that the image of the outdoor scene has sky information, and the image of the indoor scene usually does not have sky information. Taking the image of the outdoor scene as an example, as shown in FIG. 3A, the image can be divided into a close-up region, a long shot region, and a sky region; wherein, close-up is a description relative to long shot, in an image, the image is divided into a close-up region and a long shot region according to the distance relationship between the image and the light ray of the first perspective, or other ways that can decouple the close-up and the long shot. The close-up region is also called the foreground region, and the long shot region can also be called the background region. If there is no sky information in the image, only the close-up region and the long shot region need to be divided, if the image is taken close, there is no long shot region, then no division is needed, and the image only needs to be processed according to the close-up region.

[0125] The structure of the image processing model provided in the embodiments of the present application can be understood with reference to FIG. 3B, as shown in FIG. 3B, the image processing model includes: a 3dimension spatially-adaptive normalization convolutional neural network (3D SPADE CNN) 301, a first encoder 302, a second encoder 303, a third encoder 304, a first decoder 305, a second decoder 306, a third decoder 307, a fourth decoder 308, and a reconstruction rendering module 309.

[0126] The 3D SPADE CNN 301 can process a target point cloud, which can be obtained by accumulating a plurality of depth maps corresponding to the plurality of first images. In this way, the target point cloud contains the spatial information of the scene from which the plurality of first images are derived. The 3D SPADE CNN 301 extracts features from the depth point cloud, and can obtain a 3D global feature volume of the scene from which the plurality of first images are derived, that is, a spatial geometric feature.

[0127] The first encoder 302 can process the first reference image to output a close-up two-dimensional reference feature; wherein, the first reference image is obtained by back-projecting the light sampling points based on the second perspective to the close-up region of the first image of the first perspective.

[0128] The second encoder 303 can process the second reference image to output the long-range two-dimensional reference features; wherein the second reference image is obtained by projecting the light sampling points of the second view to the long-range region of the first image of the first view.

[0129] The third encoder 304 can process the sky region in the second reference image to output the sky two-dimensional reference features.

[0130] The input of the first decoder 305 is the three-dimensional global volume features output by the 3D SPADE CNN 301 and the close-range two-dimensional reference features output by the first encoder 302. After the first decoder 305 processes the three-dimensional global volume features and the close-range two-dimensional reference features, the first color information can be output.

[0131] The input of the second decoder 306 is the three-dimensional global volume features output by the 3D SPADE CNN 301. After the second decoder 306 processes the three-dimensional global volume features, the first volume density can be output.

[0132] The input of the third decoder 307 is the long-range two-dimensional reference features output by the second encoder 303. After the third decoder 307 processes the long-range two-dimensional reference features, the second volume density and the second color information can be output.

[0133] The input of the fourth decoder 308 is the sky two-dimensional reference features output by the third encoder 304. After the fourth decoder 308 processes the sky two-dimensional reference features, the color information of the sky can be output.

[0134] The input of the reconstruction and rendering module 309 is the first volume density, the first color information, the second volume density, the second color information and the color information of the sky. The reconstruction and rendering module 309 reconstructs the second image according to the first volume density and the second volume density, and renders the second image according to the first color information, the second color information and the color information of the sky, so as to obtain the second image of the second view.

[0135] It should be noted that if only the close-range image needs to be processed, the image processing model can not include the first encoder 302, the second encoder 303, the third encoder 304, the third decoder 307 and the fourth decoder 308. If the image includes the close-range region and the long-range region, but does not include the sky, the image processing model can not include the third encoder 304 and the fourth decoder 308.

[0136] In addition, it should be noted that the image processing model can further include a depth processing module, a point cloud generation module, and a reference image generation module. In this way, the depth processing module can process multiple first images to obtain depth maps of the multiple first images. The point cloud generation module can process the depth maps of the multiple first images to obtain a target point cloud. The reference image generation module can process the multiple first images, and in combination with the second perspective, first and second reference images can be obtained.

[0137] The structure of the 3D SPADE CNN can be understood with reference to FIG. 3C. As shown in FIG. 3C, the 3D SPADE CNN includes multiple point cloud feature blocks of different sizes, a downsample module, an upsample module, a SPADE 3D block, a normalization module, an addition module, and a multiplication module.

[0138] Different downsample modules can perform different amplitude downsample processing on point cloud feature blocks of the same size. The SPADE 3D block can process the first-size point cloud feature block after 3D convolution. The upsample module can perform upsample processing on the point cloud feature block processed by the SPADE 3D block, and then obtain a second-size point cloud feature block. The structure of the SPADE 3D block can include 3D convolution, normalization processing, and multiplication and addition processing of the point cloud feature block. The 3D SPADE CNN described above with reference to FIG. 3C can convert the target point cloud into a three-dimensional global volume feature.

[0139] The model training method provided in the embodiments of the present application will be described below with reference to FIG. 4.

[0140] As shown in FIG. 4, the model training method includes the following steps.

[0141] 401. Obtain a training sample.

[0142] The training sample includes multiple sample pairs, wherein each sample pair includes multiple images, and the multiple images include images of a first perspective and images of a second perspective, the first perspective being different from the second perspective. The first perspective can be different in different sample pairs, and the second perspective is usually close to the first perspective, or the perspective difference between the two is within a certain range. In a sample pair, there can be multiple images of the first perspective. The multiple images can be binocular images or monocular images.

[0143] 402. training the first image processing model to obtain a second image processing model, wherein the first image processing model is a model based on a generalized NeRF architecture.

[0144] The model training process provided in the embodiments of the present application can be understood in combination with the structure of the image processing model in FIG. 3B. In FIG. 3B, the 3DSPADE CNN 301, the first encoder 302, the second encoder 303, the third encoder 304, the first decoder 305, the second decoder 306, the third decoder 307, and the fourth decoder 308 all include a series of functional parameters, and the weights of these parameters are usually set to large values at the beginning of training. Then, during the training process, a batch of training samples are processed, a plurality of images containing the first view are processed, and then the predicted values of the plurality of images and the true values of the second view images are used to determine the loss function, and the weight values are gradually adjusted. Through iterative training, the first image processing model reaches the convergence condition to determine the second image processing model.

[0145] The model training scheme provided in the embodiments of the present application can obtain the second image processing model by training the first image processing model based on the generalized NeRF architecture. In this way, the second image processing model can use the image of the first view to reconstruct the image of the new view during the inference process, thereby increasing the image reconstruction range, improving the generalization ability of image processing, and also improving the quality of the reconstructed image.

[0146] FIG. 5 is a schematic diagram of an embodiment of the image processing method provided in the embodiments of the present application.

[0147] As shown in FIG. 5, the image processing method provided in the embodiments of the present application includes:

[0148] 501. obtaining a plurality of first images, wherein the view corresponding to the plurality of first images includes a first view.

[0149] In the present application, the plurality of first images can be images corresponding to the same scene, and the view corresponding to the first image can be the view of a camera, radar or other image acquisition device that captures the first image. The views of the plurality of images can be the same or different. The first image can be a monocular image or a binocular image.

[0150] 502. performing generalization scene reconstruction of target features based on the plurality of first images to obtain a second image of a second view; wherein the target features are obtained based on the first image, the target features include spatial geometric features of the scene from which the plurality of first images are derived, and the second view is different from the first view.

[0151] In this application, the first view is usually a view close to the second view, or the first view is the view closest to the second view among the views corresponding to the plurality of images.

[0152] In this application, the spatial geometric features can also be described as spatial features, including the spatial volume features of each sampling point in the space of the scene from which the first image is derived, that is, the features of the positions and angles of the sampling points in the space after high-dimensional mapping through coding. These high-dimensional features can improve the clarity of the reconstructed image more than the positions and angles of the sampling points.

[0153] 503. Output the second image.

[0154] The method for image processing provided in the embodiments of this application can reconstruct a second image of a new view (second view) using a plurality of first images of known views. Because the spatial geometric features of the scene from which the first image is derived are used in the process of scene reconstruction, the spatial information of the new view and the image content of the new view can be better determined using the spatial geometric features when the second image of the new view is reconstructed, thereby increasing the image reconstruction range, improving the generalization ability of image processing, and also improving the quality (such as clarity) of the second image.

[0155] Optionally, the step 502 can include: performing generalization scene reconstruction of the target features based on the plurality of first images and the image processing model; wherein the structure of the image processing model can be understood with reference to the introduction of FIG. 3B.

[0156] The process of image processing using the image processing model can be understood with reference to FIG. 6. As shown in FIG. 6, the image processing process is introduced taking an outdoor scene as an example:

[0157] The terminal device or the cloud system can obtain a plurality of first images 601, and the views corresponding to the plurality of first images include a first view. Of course, the views corresponding to the plurality of first images can also include other views in addition to the first view. As shown in FIG. 6, the plurality of first images are images taken on a street, including vehicles, houses, trees, and the sky, etc. Depth estimation can be performed on the plurality of first images to obtain a plurality of depth maps 602 corresponding to the plurality of first images. The process of depth estimation can be completed through a depth estimation model, and the depth map can also be understood as a depth point cloud. Then, depth accumulation is performed on the plurality of depth maps to obtain a target point cloud 603 of the scene from which the plurality of first images are derived, and the target point cloud can be a color point cloud. Because the depths of the long-range area and the sky area are too deep, the spatial geometric features of the long-range area and the sky area have little effect on the second image when the second image is reconstructed, and in order to save the amount of calculation, the target point cloud 603 can be the point cloud of the close-range area in the depth map of the plurality of first images.

[0158] The target point cloud 603 can be processed using the 3D SPADE CNN in the image processing model to extract the three-dimensional global volume feature 604 corresponding to the target point cloud 603. The three-dimensional global volume feature 604 can be represented as or, the three-dimensional global volume feature 604 is further processed to obtain

[0159] Because the second image of the second view is to be reconstructed according to multiple first images, the terminal device or the cloud system can also first obtain a reference image of the second view. The reference image of the second view can be obtained by back-projecting the light sampling points based on the second view (target view) onto the first image of the first view, and the first view is the view with the smallest angle difference from the second view among the views corresponding to the multiple first images. The first view can also be described as the nearest reference view of the second view, so the reference image obtained by the above back-projection can also be described as the nearest reference image. It has been introduced in the foregoing that the image of an outdoor scene can be divided into a close-range region and a long-range region, so the first reference image 605 can be obtained for the close-range region of the first image according to the second view, and the second reference image 606 can be obtained for the long-range region of the first image according to the second view.

[0160] The first encoder in the image processing model can be used for the first reference image 605 to obtain the close-range two-dimensional reference feature The second encoder in the image processing model can be used for the second reference image 606 to obtain the long-range two-dimensional reference feature In addition, the third encoder in the image processing model can also be used to obtain the sky two-dimensional reference feature

[0161] and are input into the first decoder to obtain the first color information c fg . The second decoder is input into the second decoder to obtain the first volume density σ fg . Wherein, Wherein, γ(x) and d can be understood with reference to the 5 parameters (x, y, z, θ, φ) of the part NeRF in FIG. 1A, wherein γ(x) can represent the three-dimensional coordinates (x, y, z), and d can represent the azimuth angle (θ) and the elevation angle (φ).​​​​​

[0162] c is input to the third decoder c bg and the second body density bg ; c is input to the fourth decoder c sky . Wherein, Wherein, y(x), d can be understood by referring to the foregoing description.

[0163] From the above relationship of c fg , s fg , s bg , c bg , c sky , it can be known that compared with the NeRF shown in FIG. 1A, the scheme provided in the embodiments of the present application contains or these feature information in addition to the three-dimensional coordinates and the perspective direction in FIG. 1A, which greatly improves the generalization ability of the NeRF. Input

[0164] c fg , s fg , s bg , c bg and c sky are input into the reconstruction rendering module to obtain the color information C of the second image. It should be noted that the second image contains multiple pixel points, and for each pixel point, there are sampling points in the ray direction of the corresponding second perspective, that is, the near view, the far view and the sky. As shown in FIG. 7: taking pixel point 701 as an example, in the ray direction of the second perspective, the pixel point has multiple sampling points, such as sampling point 702, sampling point 703 and sampling point 704, which are all sampling points of the pixel point 701 in the ray direction of the second perspective. Among them, the sampling point 702 can be understood as a near view region sampling point, the sampling point 703 can be understood as a far view region sampling point, and the sampling point 704 can be understood as a sky region sampling point. In this way, the color information of the pixel point 701 needs to be determined according to the color information c fg of the sampling point 702, the color information c bg of the sampling point 703 and the color information c sky of the sampling point 704, so the color value of the pixel point can be expressed as C=C (fg+bg) +(1-a (bg+fg) )C sky ; wherein, Wherein, c iCi represents color information of the i-th sampling point, σ i ρi represents the body density of the i-th sampling point. Finally, the color information C can be used by the reconstruction rendering module to obtain a second image.

[0165] In addition, it should be noted that the images collected from the city scene often present variable lighting or other environmental changes. To cope with such changes, an appearance encoding module can be designed in the network result to consider exposure changes. During model inference, for scenes that have not been seen, the output of the appearance encoding module can be set to the average value of the training scene images. For training scenes, the output of the appearance encoding module is not interpolated for test views.

[0166] When training the image processing model, multiple training losses can be used to train on different scenes and perform test-time optimization on specific scenes. For example: the most commonly used sensors in autonomous vehicles are cameras and lidar. Compared with cameras, lidar data is more expensive to obtain and may not be deployed in unseen test scenes. Therefore, lidar information can be used for training to enhance the geometric understanding of the image processing model. The training loss function and fine-tuning loss function are as follows:

[0167] where L rgb is a layer 1 (L1) or layer 2 (L2) loss function, L rgb can be used to measure the difference between the rendered color and the real pixel color. L lidar is a radar loss function. L sky is a sky loss function. L entropy is an entropy regularization loss function.

[0168] In the embodiments of the present application, in order to distinguish the sky from the near and far scenes, a pre-trained segmentation model can be used to provide a sky mask, and a binary cross entropy (BCE) loss is used to supervise the rendered sky mask.

[0169] During model training, laser radar can be selected as an additional sensor to further optimize the reconstruction result.

[0170] where the boundary width ∈ is set to 0.5 at initialization, and can be exponentially decayed to a minimum value of 0.1, so that it can gradually decrease with the increase of the number of training iterations. The second term L near The purpose of the first term is to increase the volume density within a certain range, but not to specify its distribution, while the first and third terms need to keep the rest of the region empty.

[0171] At the same time, in order for the model to represent the distant view as semi-transparent, an entropy regularization loss is introduced. entropy = -(a bg ln a fg + (1 - a fg ) ln (1 - a fg ))

[0172] In addition, in the embodiments of the present application, a zero-shot model and a few-shot model are also provided; wherein the zero-shot model is not adjusted by the images of the scene from which the first images are derived, and the few-shot model is fine-tuned by the images of the scene from which the first images are derived.

[0173] The zero-shot model can be understood with reference to Figure 8, as shown in Figure 8, the "snowflake" mark in Figure 8 indicates that the image processing model does not need to be adjusted and can be directly used in the process introduced in the embodiment corresponding to Figure 6.

[0174] The few-shot model can be understood with reference to Figure 9, as shown in Figure 9, the "snowflake" mark in Figure 9 indicates that the 3D SPADE CNN, the first encoder, the second encoder, and the third encoder do not need to be adjusted. The "flame" mark indicates that the first decoder The third decoder And the fourth decoder need to be adjusted. Therefore, before using the image processing model for inference, a small number of images of the scene from which the first images are derived can be used to fine-tune the first decoder The third decoder And the fourth decoder

[0175] The above introduces the image processing method and the model training method, and the following introduces the computer device in combination with the drawings.

[0176] As shown in Figure 10, the computer device 100 provided by the embodiments of the present application comprises:

[0177] The acquisition unit 1001 is configured to acquire a plurality of first images, and the viewing angles corresponding to the plurality of first images include a first viewing angle;

[0178] The processing unit 1002 is configured to perform generalization scene reconstruction based on the plurality of first images to obtain a second image of a second viewing angle; wherein the second viewing angle is different from the first viewing angle;

[0179] The output unit 1003 is configured to output the second image.

[0180] Optionally, the processing unit 1002 is specifically configured to perform generalization scene reconstruction of the target feature based on the plurality of first images and an image processing model; wherein the image processing model comprises a 3D SPADE CNN, and the 3D SPADE CNN is configured to determine a three-dimensional global volume feature of the plurality of first images, and the three-dimensional global volume feature is configured to indicate a spatial geometric feature of a scene represented by the plurality of first images.

[0181] Optionally, an input of the 3D SPADE CNN is a target point cloud, and an output is a three-dimensional global feature; wherein the target point cloud is determined based on a depth map of the plurality of first images.

[0182] Optionally, the image processing model further comprises a first decoder, an input of the first decoder is the three-dimensional global volume feature and a close-range two-dimensional reference feature; wherein the close-range two-dimensional reference feature is obtained based on a first reference image, and the first reference image is obtained by back-projecting a light sampling point of the second view to a close-range region of the first image of the first view.

[0183] Optionally, an output of the first decoder is first color information, and the first color information is configured to render the second image.

[0184] Optionally, the image processing model further comprises a second decoder, an input of the second decoder is the three-dimensional global feature, and an output is a first volume density, and the first volume density is configured to reconstruct the second image.

[0185] Optionally, the image processing model further comprises a third decoder, an input of the third decoder is a long-range two-dimensional reference feature; wherein the long-range two-dimensional reference feature is obtained based on a second reference image, and the second reference image is obtained by back-projecting a light sampling point of the second view to a long-range region of the first image of the first view.

[0186] Optionally, an output of the third decoder is a second volume density and second color information; the second volume density is configured to reconstruct the second image, and the second color information is configured to render the second image.

[0187] Optionally, the image processing model further comprises a fourth decoder, an input of the fourth decoder is a sky two-dimensional reference feature, and an output is color information of the sky; wherein the sky two-dimensional reference feature is obtained based on a sky region in the second reference image, and the color information of the sky is configured to render the second image.

[0188] Optionally, the image processing model is a zero-shot model, and the few-shot model is a fine-tuned model; wherein the zero-shot model is not adjusted by images of a scene from which the plurality of first images are derived, and the few-shot model is fine-tuned by the images of the scene from which the plurality of first images are derived.

[0189] The functions of the units of the computer device 100 described above can be understood with reference to the corresponding descriptions in the method embodiment part, which will not be repeated here.

[0190] As shown in FIG. 11, the computer device 110 provided by the embodiment of the present application comprises:

[0191] The acquisition unit 1101 is configured to acquire training samples, the training samples comprising a plurality of sample pairs, wherein each sample pair comprises an image of a first view and an image of a second view, the first view being different from the second view.

[0192] The processing unit 1102 is configured to take the image of the first view in the plurality of sample pairs as an input of a first image processing model, take the image of the second view as an output of the first image processing model, and train the first image processing model to obtain a second image processing model, wherein the first image processing model is a model based on a general NeRF architecture.

[0193] Optionally, the first image processing model comprises a 3D SPADE CNN, an input of the 3D SPADE CNN being a target point cloud, and an output of the 3D SPADE CNN being a three-dimensional global feature; wherein the target point cloud is determined based on a depth map of the first view.

[0194] Optionally, the first image processing model further comprises a first decoder, an input of the first decoder being the three-dimensional global feature and a near-view two-dimensional reference feature; wherein the near-view two-dimensional reference feature is obtained based on a first reference image, the first reference image being obtained by back-projecting a ray sampling point of the second view to a near-view region of the first image of the first view.

[0195] Optionally, an output of the first decoder is first color information, the first color information being used for rendering the second image.

[0196] Optionally, the first image processing model further comprises a second decoder, an input of the second decoder being the three-dimensional global feature, and an output of the second decoder being a first volume density, the first volume density being used for reconstructing the image of the second view.

[0197] Optionally, the first image processing model further comprises a third decoder, an input of the third decoder being a far-view two-dimensional reference feature; wherein the far-view two-dimensional reference feature is obtained based on a second reference image, the second reference image being obtained by back-projecting the ray sampling point of the second view to a far-view region of the first image of the first view.

[0198] Optionally, an output of the third decoder is a second volume density and second color information; the second volume density being used for reconstructing the second image, and the second color information being used for rendering the image of the second view.

[0199] Optionally, the first image processing model further comprises a fourth decoder, an input of the fourth decoder is the sky two-dimensional reference feature, and an output of the fourth decoder is color information of the sky; wherein the sky two-dimensional reference feature is obtained based on a sky region in the second reference image, and the color information of the sky is used for rendering the image of the second view angle.

[0200] The functions of the units of the computer device 110 described above can be understood with reference to the corresponding descriptions in the method embodiment part, which will not be repeated here.

[0201] FIG. 12 is a possible logical structure of a computer device according to an embodiment of the present application. The computer device can be a terminal device, a server, a virtual machine or a container. As shown in FIG. 12, the computer device 1200 can include a central processor 1201, a graphics processor 1202, a display device 1203 (which can not be included) and a memory 1204. Optionally, the computer device 1200 can further include at least one communication bus (not shown in FIG. 12) for realizing connection and communication between the components.

[0202] It should be understood that the components in the computer device 1200 can also be coupled through other connectors, which can include various interfaces, transmission lines or buses, etc. The components in the computer device 1200 can also be in a radial connection mode with the central processor 1201 as the center. In various embodiments of the present application, coupling means through mutual electrical connection or communication, including direct connection or indirect connection through other devices.

[0203] The connection mode of the central processor 1201 and the graphics processor 1202 can also be various, and is not limited to the mode shown in FIG. 12. The central processor 1201 and the graphics processor 1202 in the computer device 1200 can be located on the same chip, or can be independent chips.

[0204] The functions of the central processor 1201, the graphics processor 1202, the display device 1203 and the memory 1204 will be briefly introduced below.

[0205] The central processor 1201 is configured to run an operating system 1205 and an application 1206. The application 1206 can be a graphic application such as a video player and the like. The operating system 1205 provides a system graphic library interface through which the application 1206 generates an instruction stream for rendering a graphic or image frame together with associated rendering data required by the operating system 1205, a graphic library user mode driver and / or a graphic library kernel mode driver. The system graphic library includes, but is not limited to, an open graphic library for embedded system (OpenGL ES), a Khronos platform graphic interface or Vulkan (a cross-platform drawing application interface) and the like. The instruction stream contains a series of instructions, which are usually call instructions to the system graphic library interface.

[0206] Optionally, the central processor 1201 can include at least one of an application processor, one or more microprocessors, a digital signal processor (DSP), a microcontroller unit (MCU), an artificial intelligence processor and the like.

[0207] The central processor 1201 can further include necessary hardware accelerators such as an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or an integrated circuit for implementing logical operations. The processor 1201 can be coupled to one or more data buses for transmitting data and instructions between various components of the computer device 1200, such as the feature extraction and feature decoding processes in the above method embodiments.

[0208] The graphics processor 1202 is configured to receive a graphics instruction stream sent by the processor 1201, generate a rendering target through a rendering pipeline, and display the rendering target to the display device 1203 through a layer composition display module of the operating system. The rendering pipeline can also be referred to as a rendering pipeline, a pixel pipeline, or a pixel pipeline, and is a parallel processing unit inside the graphics processor 1202 for processing graphics signals. Multiple rendering pipelines can be included in the graphics processor 1202, and the multiple rendering pipelines can independently process graphics signals in parallel. For example, the rendering pipeline can perform a series of operations in the process of rendering a graphics or image frame, and typical operations can include vertex processing, primitive processing, rasterization, fragment processing, image reconstruction and rendering, and the like.

[0209] Optionally, the graphics processor 1202 can include a general-purpose graphics processor that executes software, such as a GPU or other types of specialized graphics processing units.

[0210] The display device 1203 is configured to display various images generated by the computer device 1200, which can be a graphical user interface (GUI) of the operating system or image data processed by the graphics processor 1202, including still images and video data.

[0211] Optionally, the display device 1203 can include any suitable type of display screen, such as a liquid crystal display (LCD) or a plasma display or an organic light-emitting diode (OLED) display, and the like.

[0212] The memory 1204 is a transmission channel between the central processor 1201 and the graphics processor 1202, and can be a double data rate synchronous dynamic random access memory (DDR SDRAM) or other types of cache.

[0213] In another embodiment of the present application, a computer readable storage medium is also provided, and the computer readable storage medium stores computer execution instructions. When the processor of the computer device executes the computer execution instructions, the computer device performs the steps performed by the computer device in FIGS. 1A to 9.

[0214] In another embodiment of the present application, a computer program product is also provided, which includes computer program codes, when the computer program codes are executed on a computer, the computer device executes the steps performed by the computer device as described above in FIG. 1A to FIG. 9.

[0215] In another embodiment of the present application, a chip system is also provided, which includes one or more interface circuits and one or more processors; the interface circuits and the processors are interconnected through lines; the interface circuits are configured to receive signals from a memory of the computer device and send signals to the processors, the signals including computer instructions stored in the memory; when the processors execute the computer instructions, the computer device executes the steps performed by the computer device as described above in FIG. 1A to FIG. 9. In a possible design, the chip system can further include the memory, which is configured to store program instructions and data necessary for the control device. The chip system can be composed of chips, or can include chips and other discrete devices.

[0216] In several embodiments provided in the present application, it should be understood that the disclosed system, device and method can be implemented in other manners. For example, the described device embodiments are merely schematic. For example, the division of the units is only a logical function division. There can be another division manner for the actual implementation, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections between different units, can be indirect couplings or communication connections through some interfaces, devices or units, and can be electrical, mechanical or in other forms.

[0217] The units described as separated components can or can not be physically separated, and the components displayed as units can or can not be physical units, i.e., can be located in one place, or can be distributed on a plurality of network units. Some or all of the units can be selected according to actual needs to achieve the purposes of the embodiments.

[0218] In addition, each functional unit in the embodiments of the present application can be integrated in one processing unit, or each unit can exist physically as a separate unit, or two or more units can be integrated in one unit. The above integrated unit can be implemented in whole or in part by software, hardware, firmware or any combination thereof.

[0219] When the units of the integration are implemented by using software, the units can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are wholly or partially generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions can be transmitted from one website site, computer, server or data center to another website site, computer, server or data center through wired (for example, coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (for example, infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server, data center, etc. integrated with one or more available media. The available media can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk (SSD)), etc.

Claims

1. A method of image processing, characterized by, The method comprises: acquiring a plurality of first images, the corresponding view angles of the plurality of first images comprising a first view angle; performing generalization scene reconstruction of a target feature based on the plurality of first images to obtain a second image of a second view angle; wherein the target feature is obtained based on the first images, the target feature comprises a spatial geometric feature of a scene from which the plurality of first images are derived, and the second view angle is different from the first view angle; outputting the second image.

2. The method of claim 1, wherein, The generalization scene reconstruction of the target feature based on the plurality of first images comprises: performing generalization scene reconstruction based on the plurality of first images and an image processing model; wherein the image processing model comprises a three-dimensional space adaptive normalization convolutional neural network (3D SPADE CNN), the 3D SPADE CNN is used to determine a three-dimensional global volume feature of the plurality of first images, and the three-dimensional global volume feature is used to indicate the spatial geometric feature of a scene represented by the plurality of first images.

3. The method of claim 2, wherein, The image processing model further comprises a first decoder, an input of the first decoder being the three-dimensional global volume feature and a close-range two-dimensional reference feature; wherein the close-range two-dimensional reference feature is obtained based on a first reference image, and the first reference image is obtained by back-projecting a close-range region of a first image of the first view angle based on a light sampling point of the second view angle.

4. The method according to claim 2 or 3, characterized in that, An input of the 3D SPADE CNN is a target point cloud, and an output of the 3D SPADE CNN is the three-dimensional global feature; wherein the target point cloud is determined based on a depth map of the plurality of first images.

5. The method according to any one of claims 2-4, characterized in that, An output of the first decoder is first color information, and the first color information is used to render the second image.

6. The method according to any one of claims 2-5, characterized in that, The image processing model further comprises a second decoder, an input of the second decoder being the three-dimensional global feature, and an output of the second decoder being a first volume density, the first volume density being used to reconstruct the second image.

7. The method according to any one of claims 2-6, characterized in that, The image processing model further comprises a third decoder, an input of the third decoder being a long-range two-dimensional reference feature; wherein the long-range two-dimensional reference feature is obtained based on a second reference image, and the second reference image is obtained by back-projecting a long-range region of a first image of the first view angle based on a light sampling point of the second view angle.

8. The method of claim 7, wherein, An output of the third decoder is a second volume density and second color information; the second volume density is used to reconstruct the second image, and the second color information is used to render the second image.

9. The method according to claim 7 or 8, characterized in that, The image processing model further comprises a fourth decoder, an input of the fourth decoder being a sky two-dimensional reference feature, and an output of the fourth decoder being color information of the sky; wherein the sky two-dimensional reference feature is obtained based on a sky region in the second reference image, and the color information of the sky is used to render the second image.

10. The method according to any one of claims 2-9, characterized in that, The image processing model is a zero-shot model, and is a few-shot model; wherein the zero-shot model is not adjusted based on images of a scene from which the plurality of first images are derived, and the few-shot model is fine-tuned based on images of the scene from which the plurality of first images are derived.

11. A computer apparatus, comprising: The method comprises: An acquisition unit is configured to acquire a plurality of first images, the corresponding view angles of the plurality of first images including a first view angle; A processing unit is configured to perform generalized scene reconstruction based on the plurality of first images to obtain a second image of a second view angle; wherein the second view angle is different from the first view angle; An output unit is configured to output the second image.

12. The computer device of claim 11, wherein, The processing unit is configured to perform generalized scene reconstruction based on the plurality of first images and an image processing model; wherein the image processing model includes a three-dimensional space adaptive normalization convolutional neural network (3D SPADE CNN), the 3D SPADE CNN being configured to determine three-dimensional global volume features of the plurality of first images, the three-dimensional global volume features being configured to indicate the spatial geometric features of a scene represented by the plurality of first images.

13. The computer device of claim 11, wherein, The image processing model further includes a first decoder, an input of the first decoder being the three-dimensional global volume features and near-view two-dimensional reference features; wherein the near-view two-dimensional reference features are obtained based on a first reference image, the first reference image being obtained by back-projecting a near-view region of a first image of the first view angle based on a light sampling point of the second view angle.

14. A computer device, comprising: A computer readable storage medium having a computer program stored therein; The processor is coupled to the computer readable storage medium, and the computer program is executed by the processor to implement the method of any one of claims 1-10.

15. A computer readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the method of any one of claims 1-10.

16. A computer program product, characterised in that, The computer program product includes computer program code, when the computer program code is run on a computer device, causing the computer device to execute the method of any one of claims 1-10.

17. A chip system, characterized by The computer program product includes a processor, the processor being invoked to execute the method of any one of claims 1-10. The computer program product includes a processor, the processor being invoked to execute the method of any one of claims 1-10.