Image generation method and device, electronic equipment, medium and vehicle

By obtaining the target viewing angle, ambient light and sampling point position information, the target information generation model is trained, and the image generation problem caused by the constant light assumption in the neural radiation field algorithm is solved, and high-quality image generation is achieved in the changing light environment.

CN120339489APending Publication Date: 2025-07-18XIAOMI EV TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410064836.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-16
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The neural radiation field algorithm assumes constant light, resulting in the generated new viewing angle image being single and unable to adapt to changing ambient light.

Method used

By obtaining the target viewing angle, ambient light and sampling point position information, the target information generation model is trained to generate an image of the target viewing angle under any ambient light.

Benefits of technology

Improves the diversity of generated images, enriches the user experience, and can generate real and accurate target images in changing light environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339489A_ABST
    Figure CN120339489A_ABST
Patent Text Reader

Abstract

The invention relates to an image generation method and device, electronic equipment, a medium and a vehicle, and relates to the technical field of computers.The method comprises the steps that view angle information of a target view angle corresponding to a target scene, ambient light information and position information of multiple sampling points in a target area are obtained; inputting the visual angle information, the ambient light information and the position information of each sampling point into a pre-trained target information generation model to obtain color information and volume density information corresponding to each sampling point output by the target information generation model; and according to the plurality of pieces of color information and the plurality of pieces of volume density information, generating a target image of the target view angle under the ambient light information. Thus, a user can input any ambient light information according to actual requirements, and through the target information generation model obtained after reconstruction, the target image of the target visual angle under any ambient light is generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and in particular, to an image generation method, apparatus, electronic device, medium, and vehicle. Background Art

[0002] Neural Radiance Fields is a 3D scene reconstruction and image synthesis algorithm that combines computer vision and computer graphics. Given a set of images with pose information, Neural Radiance Fields learns the geometric shape and lighting texture information in the scene through a Multiple Layer Perceptor (MLP), can complete dense 3D reconstruction of the scene, and can generate images from new viewpoints.

[0003] However, the Neural Radiance Fields algorithm is based on the assumption of lighting invariance, that is, all input images are taken under the condition of constant scene lighting, and thus the generated images from new viewpoints can only be images under fixed lighting, resulting in a relatively single image generation effect. Summary of the Invention

[0004] To overcome the problems in the related art, the present disclosure provides an image generation method, apparatus, electronic device, medium, and vehicle.

[0005] According to a first aspect of an embodiment of the present disclosure, an image generation method is provided. The method includes:

[0006] Obtain the view information of a target view corresponding to a target scene, the ambient light information, and the position information of a plurality of sampling points in a target area; the target area is the view area of the target scene under the target view;

[0007] Input the view information, the ambient light information, and the position information of each sampling point into a pre-trained target information generation model to obtain the color information and volume density information corresponding to each sampling point output by the target information generation model;

[0008] Generate a target image of the target view under the ambient light information according to the plurality of color information and the plurality of volume density information.

[0009] Optionally, the target information generation model is trained in the following manner:

[0010] Obtain a plurality of scene images, where the plurality of scene images are images of different views collected for the target scene, and the ambient light corresponding to each collection of the scene images is different;

[0011] For each of the scene images, determine the sample view parameters of the acquisition view corresponding to the scene image, the sample position parameters of multiple sample points within the sampling area, and the sample pixel parameters corresponding to each pixel point in the scene image; the sampling area is the view area of the target scene under the acquisition view;

[0012] Train a preset neural radiance field based on multiple sample view parameters, multiple sample position parameters, and multiple sample pixel parameters to obtain the target information generation model.

[0013] Optionally, the training of the preset neural radiance field based on multiple sample view parameters, multiple sample position parameters, and multiple sample pixel parameters to obtain the target information generation model includes:

[0014] For each of the scene images, determine a target loss value according to the sample view parameters, multiple sample position parameters, and the sample pixel parameters;

[0015] Update the model parameters of the preset neural radiance field according to the target loss value until the preset neural radiance field converges to obtain the target information generation model.

[0016] Optionally, the determination of the target loss value according to the sample view parameters, multiple sample position parameters, and the sample pixel parameters includes:

[0017] Input the sample view parameters, multiple sample position parameters, and randomly generated sample environmental light parameters into the preset neural radiance field to obtain the sample color value and sample volume density corresponding to each sample point output by the preset neural radiance field;

[0018] Determine the target loss value according to multiple sample color values, multiple sample volume densities, and the sample pixel parameters.

[0019] Optionally, the determination of the target loss value according to multiple sample color values, multiple sample volume densities, and the sample pixel parameters includes:

[0020] Determine the sample pixel value corresponding to the sample point according to multiple sample color values and multiple sample volume densities;

[0021] Determine the target loss value according to the sample pixel value and the sample pixel parameters, and the target loss value is used to characterize the deviation degree between the sample pixel value and the sample pixel parameters.

[0022] Optionally, the target scene includes multiple sub - scenes, and the target information generation model is trained in the following way:

[0023] For each of the sub-scenarios, obtain a plurality of sub-scenario images corresponding to the sub-scenario. The plurality of sub-scenario images are images from different perspectives collected for the sub-scenario, and the ambient light corresponding to each collection of the sub-scenario images is different;

[0024] Train a preset neural radiance field based on the plurality of sub-scenario images to obtain a pending information generation model corresponding to the sub-scenario;

[0025] Perform perspective fusion on the plurality of obtained pending information generation models corresponding to the sub-scenarios to obtain the target information generation model.

[0026] Optionally, the generating the target image of the target perspective under the ambient light information according to the plurality of color information and the plurality of volume density information includes:

[0027] Determine the target pixel value corresponding to each sampling point according to the plurality of color information and the plurality of volume density information;

[0028] Generate the target image of the target perspective under the ambient light information based on the plurality of target pixel values.

[0029] According to a second aspect of the embodiments of the present disclosure, there is provided an image generation device, the device includes:

[0030] An acquisition module, configured to acquire the perspective information of the target perspective corresponding to the target scene, the ambient light information, and the position information of a plurality of sampling points within the target area; the target area is the perspective area of the target scene under the target perspective;

[0031] A prediction module, configured to input the perspective information, the ambient light information, and the position information of each sampling point into a pre-trained target information generation model to obtain the color information and volume density information corresponding to each sampling point output by the target information generation model;

[0032] A generation module, configured to generate the target image of the target perspective under the ambient light information according to the plurality of color information and the plurality of volume density information.

[0033] Optionally, the target information generation model is trained in the following manner:

[0034] Obtain a plurality of scene images. The plurality of scene images are images from different perspectives collected for the target scene, and the ambient light corresponding to each collection of the scene images is different;

[0035] For each of the scene images, determine the sample view parameters of the acquisition view corresponding to the scene image, the sample position parameters of multiple sample points within the sampling region, and the sample pixel parameters corresponding to each pixel point in the scene image; the sampling region is the view region of the target scene under the acquisition view;

[0036] Train a preset neural radiance field based on multiple sample view parameters, multiple sample position parameters, and multiple sample pixel parameters to obtain the target information generation model.

[0037] Optionally, the training of the preset neural radiance field based on multiple sample view parameters, multiple sample position parameters, and multiple sample pixel parameters to obtain the target information generation model includes:

[0038] For each of the scene images, determine a target loss value according to the sample view parameters, multiple sample position parameters, and the sample pixel parameters;

[0039] Update the model parameters of the preset neural radiance field according to the target loss value until the preset neural radiance field converges to obtain the target information generation model.

[0040] Optionally, the determining of the target loss value according to the sample view parameters, multiple sample position parameters, and the sample pixel parameters includes:

[0041] Input the sample view parameters, multiple sample position parameters, and randomly generated sample environmental light parameters into the preset neural radiance field to obtain the sample color value and sample volume density corresponding to each sample point output by the preset neural radiance field;

[0042] Determine the target loss value according to multiple sample color values, multiple sample volume densities, and the sample pixel parameters.

[0043] Optionally, the determining of the target loss value according to multiple sample color values, multiple sample volume densities, and the sample pixel parameters includes:

[0044] Determine the sample pixel value corresponding to the sample point according to multiple sample color values and multiple sample volume densities;

[0045] Determine the target loss value according to the sample pixel value and the sample pixel parameters, and the target loss value is used to characterize the deviation degree between the sample pixel value and the sample pixel parameters.

[0046] Optionally, the target scene includes multiple sub - scenes, and the target information generation model is trained in the following way:

[0047] For each of the sub-scenes, obtain a plurality of sub-scene images corresponding to the sub-scene. The plurality of sub-scene images are images from different perspectives collected for the sub-scene, and the ambient light corresponding to each collection of the sub-scene images is different;

[0048] Train a preset neural radiance field based on the plurality of sub-scene images to obtain a pending information generation model corresponding to the sub-scene;

[0049] Perform view synthesis on the plurality of obtained pending information generation models corresponding to the sub-scenes to obtain the target information generation model.

[0050] Optionally, the generation module is configured to determine a target pixel value corresponding to each of the sampling points according to the plurality of color information and the plurality of volume density information; and generate a target image of the target view under the ambient light information based on the plurality of target pixel values.

[0051] According to a third aspect of the embodiments of the present disclosure, there is provided an electronic device, including: a processor; a memory for storing processor-executable instructions; wherein, the processor is configured to implement the steps of the image generation method provided in the first aspect of the present disclosure when calling the executable instructions stored on the memory.

[0052] According to a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, on which computer program instructions are stored, and when the program instructions are executed by a processor, the steps of the image generation method provided in the first aspect of the present disclosure are implemented.

[0053] According to a fifth aspect of the embodiments of the present disclosure, there is provided a vehicle, including the electronic device provided in the third aspect of the present disclosure.

[0054] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects: First, obtain the view information, ambient light information of the target view corresponding to the target scene, and the position information of a plurality of sampling points within the target area; the target area is the view area of the target scene under the target view. Then, input the view information, the ambient light information, and the position information of each sampling point into a pre-trained target information generation model to obtain the color information and volume density information corresponding to each sampling point output by the target information generation model. Finally, generate a target image of the target view under the ambient light information according to the plurality of color information and the plurality of volume density information. Through the above method, the user can input arbitrary ambient light information according to actual needs, and generate a target image of the target view under arbitrary ambient light through the reconstructed target information generation model. In this way, the diversity of the generated target images is improved, and at the same time, the user experience is enriched.

[0055] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and should not limit the present disclosure. Description of the Drawings

[0056] The drawings herein are incorporated into and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure.

[0057] Figure 1 is a flowchart of a method for generating an image shown according to an exemplary embodiment.

[0058] Figure 2 is a flowchart of a method for training a model shown according to an exemplary embodiment.

[0059] Figure 3 is a flowchart of another method for training a model shown according to an exemplary embodiment.

[0060] Figure 4 is a block diagram of an apparatus for generating an image shown according to an exemplary embodiment.

[0061] Figure 5 is a block diagram of an electronic device shown according to an exemplary embodiment. Detailed Embodiments

[0062] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of apparatuses and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0063] It should be noted that all actions of obtaining signals, information, or data in the present disclosure are carried out on the premise of complying with the corresponding data protection regulations and policies of the country where it is located and obtaining authorization from the owner of the corresponding device.

[0064] Before introducing a method, apparatus, electronic device, medium, and vehicle for generating an image provided by the present disclosure, the application scenarios involved in each embodiment of the present disclosure will be introduced first. In various application scenarios such as autonomous driving, gaming, virtual reality (VR), and augmented reality (AR), it is often necessary to render images of new perspectives in a specific scene. Currently, neural radiance fields are often used to complete tasks such as three-dimensional reconstruction of three-dimensional scenes and generation of images from new perspectives. A traditional neural radiance field is a 5D function: the neural radiance field function represents a local scene as a function with a 5D vector as the input. The 5D vector includes the 3D coordinate position x = (x, y, z) of a spatial point and the camera viewing direction (θ, φ); the output is the color c = (r, g, b) of the 3D point related to the viewing direction and the volume density σ of the 3D point position.

[0065] The neural radiance field algorithm includes two multi-layer perceptrons (MLPs). The first MLP is used to calculate the volume density σ and the corresponding implicit feature vector h of each 3D point (x, y, z), which can be expressed as σ, h = MLP_1(x, y, z); the second MLP is used to calculate the color of each 3D point (x, y, z) under the viewing angle (θ, φ), which can be expressed as c = r, g, b = MLP_2(h, θ, φ).

[0066] As can be seen from the above analysis, the determination of the volume density σ and the color c has nothing to do with the ambient light at the moment when the camera takes a photo. That is to say, it is assumed in the neural radiance field algorithm that all photos are obtained under constant ambient light. Furthermore, when generating an image from a new perspective, the neural radiance field can only generate an image under a fixed light.

[0067] However, in actual scenarios, such as in the automatic parking scenario, the ambient light in the parking lot is constantly changing, and the brightness and light and shadow of the input images collected at different times are also changing. In this case, it is impossible to model the changing light environment.

[0068] Even the images generated from new perspectives by the neural radiance field reconstructed under the same ambient light can only be images under fixed lighting, resulting in a relatively single image generation effect.

[0069] To solve the above technical problems, the present invention provides a method, apparatus, electronic device, medium, and vehicle for generating an image. Users can input arbitrary ambient light information according to actual needs, and generate a target image of a target perspective under any ambient light through a target information generation model obtained after reconstruction. In this way, the diversity of the generated target images is improved, and at the same time, the user experience is also enriched.

[0070] The following will describe in detail the specific embodiments of the present invention with reference to the accompanying drawings.

[0071] Figure 1 is a flowchart of a method for generating an image shown according to an exemplary embodiment, as Figure 1 shown, the method may include the following steps:

[0072] In step S101, obtain the perspective information of the target perspective corresponding to the target scene, the environmental light information, and the position information of multiple sampling points in the target area.

[0073] Among them, the target area is the perspective area of the target scene under the target perspective, and the target perspective is the new perspective to be generated, which can be set by the user according to actual needs. The target perspective can be understood as the perspective of the camera (simulating the human eye) looking at the target scene. The sampling point is a spatial point obtained by sampling on multiple light rays emitted from the target perspective to the target scene, and the position information is the coordinate position of the sampling point in space. For the same target perspective, multiple sampling points can be sampled in the target scene. The environmental light information can characterize information such as the brightness, light intensity, and transparency of the environmental light in the target scene that the user expects to obtain.

[0074] In step S102, input the perspective information, the environmental light information, and the position information of each sampling point into a pre-trained target information generation model to obtain the color information and volume density information corresponding to each sampling point output by the target information generation model.

[0075] Exemplarily, for each sampling point, the position information, perspective information, and environmental light information corresponding to the sampling point can be input into the target information generation model to obtain the color information and volume density information corresponding to each sampling point output by the target information generation model.

[0076] In this embodiment, the target information generation model can be trained on a preset neural radiance field (NeRF) based on multiple scene images of the target scene collected from multiple different perspectives, and when collecting each scene image, the environmental light of the target scene can be different. In this way, for changing environmental light, the model can learn the environmental light in the image through the scene images, and finally be able to reconstruct a real and accurate target scene. And because the model has learned the environmental light in different scene images, a dynamic environmental light change range is obtained. When generating an image under a new perspective, the environmental light of the target scene can be set according to this environmental light change range to obtain a new image of the target scene under different environmental lights.

[0077] In step S103, generate a target image of the target perspective under the environmental light information according to the multiple color information and the multiple volume density information.

[0078] Among them, the volume density information can be understood as the opacity of the position where the current sampling point is located. The higher the opacity, the greater the proportion of its color. The color information reflects the RGB color representation of the current sampling point.

[0079] Exemplarily, based on multiple pieces of this color information and multiple pieces of this volume density information, the target pixel value corresponding to each sampling point can be determined. Then, based on multiple pieces of this target pixel value, a target image of the target view under this ambient light information is generated.

[0080] As can be analyzed from the foregoing, the sampling points are spatial points sampled on the light ray emitted from the target view to the target scene. There are often multiple sampling points on one light ray, and the final effect reflected on the target image after superposition is the target pixel value corresponding to the pixel point. Therefore, when synthesizing the target image under the target view, it can be completed by traversing all the pixel points of this image. Specifically: emit a light ray from the position where the camera optical center (i.e., the target view) is located to this pixel point. Through the target information generation model, the volume density σ of each sampling point on this light ray, and the color presented by the position of each sampling point under the ray view can be obtained, that is, c = (r, g, b). Among them, the volume density σ is used to calculate the weight, and by performing weighted summation on the colors of each sampling point on the light ray, the target pixel value corresponding to each pixel point can be obtained.

[0081] Since theoretically an infinite number of sampling points can be sampled on one light ray and volume rendering is completed by using the integral method, but this method is not convenient for programming implementation and gradient backpropagation optimization. Therefore, in this embodiment, first, the size of the target scene can be set, the nearest end and the farthest end of the light ray are set as tn and tf respectively, then [tn, tf] is evenly divided into N parts, and then uniform random sampling is performed in each small area to obtain multiple sampling points.

[0082] Then, the target pixel value C(r) of a certain pixel point can be simplified into a summation form:

[0083]

[0084] Among them, σ i is the volume density of the sampling point, c i is the color of the sampling point, δ i is the distance between two adjacent sampling points, and T i is:

[0085]

[0086] In this way, based on the multiple color information and multiple volume density information corresponding to multiple sampling points, the target image of the target perspective under the environmental light information can be obtained through volume rendering.

[0087] Using the above method, users can input arbitrary environmental light information according to actual needs, and generate a target information generation model through the reconstructed target information, and generate a target image of the target perspective under arbitrary environmental light. In this way, the diversity of the generated target image is improved, and at the same time, the user experience is also enriched.

[0088] The training method of the target information generation model will be described in detail below. As Figure 2 shown, the target information generation model can be trained in the following way:

[0089] S1. Obtain multiple scene images.

[0090] Among them, the multiple scene images are images of different perspectives collected for the target scene, and the environmental light corresponding to each scene image collection is different.

[0091] In an actual scene, for example, in an automatic parking environment, it is necessary to model the parking lot environment on site. At this time, a vehicle-mounted image acquisition device can be used to collect multiple scene images at different acquisition perspectives in the parking lot (i.e., the target scene). During the acquisition process, the light in the parking lot may be constantly changing, which results in different environmental lights corresponding to the finally acquired scene images.

[0092] S2. For each scene image, determine the sample view parameter of the acquisition perspective corresponding to the scene image, the sample position parameters of multiple sample points in the sampling area, and the sample pixel parameters corresponding to each pixel point in the scene image.

[0093] Among them, the sampling area is the view area of the target scene at the acquisition perspective, the sample points are the spatial points obtained by sampling on multiple light rays emitted from the acquisition perspective to the target scene, and the sample position parameter is the coordinate position of the sample point in space. The sample pixel parameter is the true pixel value corresponding to each pixel point in each scene image.

[0094] S3. Train a preset neural radiance field according to the multiple sample view parameters, multiple sample position parameters, and multiple sample pixel parameters to obtain the target information generation model.

[0095] Specifically, for each of the scene images, a target loss value can be determined according to the sample view parameter, multiple sample position parameters, and the sample pixel parameter. Then, according to the target loss value, the model parameters of the preset neural radiance field are updated until the preset neural radiance field converges, and the target information generation model is obtained.

[0096] Exemplarily, the target loss value can be determined in the following manner:

[0097] First, the sample view parameter, the multiple sample position parameters, and a randomly generated sample environmental light parameter can be input into the preset neural radiance field to obtain the sample color value and sample volume density corresponding to each sample point output by the preset neural radiance field.

[0098] The traditional neural radiance field is a 5D function. In this embodiment, in order to enable the model to learn the environmental light in the scene image, a sample environmental light parameter is added as an input. That is, the neural radiance field in this embodiment is a 6D function. Since the environmental light parameter in the scene image cannot be directly obtained, a sample environmental light parameter can be randomly generated as the model input. During the iterative learning process of the model, the sample environmental light parameter is also continuously iteratively optimized. Finally, the environmental light parameter obtained when the model converges can reflect the real environmental light in the scene image.

[0099] Then, according to the multiple sample color values, the multiple sample volume densities, and the sample pixel parameter, the target loss value is determined.

[0100] In this step, the sample pixel value corresponding to the sample point can be determined according to the multiple sample color values and the multiple sample volume densities. For example, according to the multiple sample color values and the multiple sample volume densities, the sample pixel value corresponding to the sample point can be obtained by volume rendering. The specific implementation method can refer to the corresponding embodiment part of step S103 above and will not be elaborated here. Then, according to the sample pixel value and the sample pixel parameter, the target loss value is determined, and the target loss value is used to characterize the deviation degree between the sample pixel value and the sample pixel parameter.

[0101] In this embodiment, the neural radiance field still includes two multi-layer perceptrons (MLPs). The first MLP is used to calculate the volume density σ and the corresponding implicit feature vector h of each sample point (x, y, z), which can be expressed as σ, h = MLP_1(x, y, z); that is, the first MLP is consistent with the traditional neural radiance field algorithm. The second MLP is used to calculate the color of each sample point (x, y, z) under the view (θ, φ), which can be expressed as c = r, g, b = MLP_2(h, θ, φ, L_i).

[0102] It is understandable that \(L_i\) here represents the environmental light embedding feature vector when capturing the scene image \(i\). This vector is a learnable vector, and each scene image \(i\) has a corresponding \(L_i\). During the model training process, \(L_i\) is continuously optimized and iterated. When the model converges, the \(L_i\) corresponding to the scene image is obtained.

[0103] Currently, general neural radiance fields can only handle small-scale scene models, such as constructing scene models with a scale of 10 to 100 meters. Due to limitations in factors such as model learning ability, the number of model parameters, and the image imaging quality at greater distances, it is impossible to model large-scale and large-range scenes. Considering the above problems, in order to achieve large-scale scene reconstruction, a large scene can be split into multiple small scenes. That is, in this embodiment, if the target scene includes multiple sub-scenes, such as Figure 3 As shown, the target information generation model can also be trained in the following way:

[0104] Step A: For each such sub-scene, obtain multiple sub-scene images corresponding to the sub-scene.

[0105] Among them, the multiple sub-scene images are images from different perspectives collected for the sub-scene, and the environmental light corresponding to each sub-scene image collection is different.

[0106] Step B: Train the preset neural radiance field according to the multiple sub-scene images to obtain a to-be-determined information generation model corresponding to the sub-scene.

[0107] Among them, training the preset neural radiance field according to the multiple sub-scene images to obtain a to-be-determined information generation model corresponding to the sub-scene can refer to the above embodiment of training the preset neural network preset neural radiance field according to the scene image to obtain the target information generation model corresponding to the target scene, which will not be elaborated here.

[0108] Step C: Perform perspective fusion on the multiple to-be-determined information generation models corresponding to the sub-scenes to obtain the target information generation model.

[0109] That is to say, in this embodiment, when reconstructing a large scene, the scene can be cut and processed into multiple sub-scenes. Taking the target scene as a parking lot as an example, the parking lot can be cut according to road intersections. For the scenes on the same straight road, the same neural radiance field is used for modeling. For the parking lot areas divided by T-junctions, crossroads, right-angle intersections, etc., different neural radiance fields are used for modeling. In this way, different regions can be modeled separately, corresponding to different neural radiance fields (i.e., to-be-determined information generation models). Finally, the to-be-determined information generation models of different regions can be combined together to complete the scene modeling of the entire parking lot.

[0110] With the above method, users can input arbitrary environmental light information according to actual needs, and generate a target information generation model through the reconstructed target information, so as to generate target images of the target perspective under arbitrary environmental light. In this way, the diversity of the generated target images is improved, and at the same time, the user experience is also enriched.

[0111] Figure 4 It is a block diagram of an image generation device shown according to an exemplary embodiment. As Figure 4 shown, the device 200 includes:

[0112] An acquisition module 201, configured to acquire perspective information of a target perspective corresponding to a target scene, environmental light information, and position information of a plurality of sampling points in a target area; the target area is a perspective area of the target scene under the target perspective;

[0113] A prediction module 202, configured to input the perspective information, the environmental light information, and the position information of each sampling point into a pre-trained target information generation model, so as to obtain color information and volume density information corresponding to each sampling point output by the target information generation model;

[0114] A generation module 203, configured to generate a target image of the target perspective under the environmental light information according to the plurality of color information and the plurality of volume density information.

[0115] Optionally, the target information generation model is trained in the following manner:

[0116] Acquire a plurality of scene images, where the plurality of scene images are images of different perspectives collected for the target scene, and the environmental light corresponding to each collection of the scene images is different;

[0117] For each scene image, determine a sample perspective parameter of the collection perspective corresponding to the scene image, sample position parameters of a plurality of sample points in a sampling area, and sample pixel parameters corresponding to each pixel point in the scene image; the sampling area is a perspective area of the target scene under the collection perspective;

[0118] Train a preset neural radiance field according to the plurality of sample perspective parameters, the plurality of sample position parameters, and the plurality of sample pixel parameters to obtain the target information generation model.

[0119] Optionally, the training of the preset neural radiance field according to the plurality of sample perspective parameters, the plurality of sample position parameters, and the plurality of sample pixel parameters to obtain the target information generation model includes:

[0120] For each such scene image, determine a target loss value according to the sample view parameter, multiple sample position parameters, and the sample pixel parameter;

[0121] Update the model parameters of the preset neural radiance field according to the target loss value until the preset neural radiance field converges to obtain the target information generation model.

[0122] Optionally, determining the target loss value according to the sample view parameter, multiple sample position parameters, and the sample pixel parameter includes:

[0123] Input the sample view parameter, multiple sample position parameters, and randomly generated sample ambient light parameters into the preset neural radiance field to obtain the sample color value and sample volume density corresponding to each sample point output by the preset neural radiance field;

[0124] Determine the target loss value according to multiple sample color values, multiple sample volume densities, and the sample pixel parameter.

[0125] Optionally, determining the target loss value according to multiple sample color values, multiple sample volume densities, and the sample pixel parameter includes:

[0126] Determine the sample pixel value corresponding to the sample point according to multiple sample color values and multiple sample volume densities;

[0127] Determine the target loss value according to the sample pixel value and the sample pixel parameter, where the target loss value is used to characterize the deviation degree between the sample pixel value and the sample pixel parameter.

[0128] Optionally, the target scene includes multiple sub - scenes, and the target information generation model is trained in the following way:

[0129] For each such sub - scene, obtain multiple sub - scene images corresponding to the sub - scene, where the multiple sub - scene images are images of different views collected for the sub - scene, and the ambient light corresponding to each collection of the sub - scene image is different;

[0130] Train the preset neural radiance field according to the multiple sub - scene images to obtain a to - be - determined information generation model corresponding to the sub - scene;

[0131] Perform view fusion on the multiple obtained to - be - determined information generation models corresponding to the sub - scenes to obtain the target information generation model.

[0132] Optionally, the generation module 203 is configured to determine the target pixel value corresponding to each sampling point according to multiple color information and multiple volume density information; based on multiple target pixel values, generate a target image of the target view under the ambient light information.

[0133] With the above device, users can input arbitrary ambient light information according to actual needs, and generate a target image of a target perspective under arbitrary ambient light through a model generated from the target information obtained after reconstruction. In this way, the diversity of the generated target images is improved, and at the same time, the user experience is also enriched.

[0134] Regarding the device in the above embodiment, the specific manners in which each module performs operations have been described in detail in the embodiment related to the method, and will not be elaborated here.

[0135] The present disclosure also provides a computer-readable storage medium, on which computer program instructions are stored, and when the program instructions are executed by a processor, the steps of the method for generating an image provided by the present disclosure are implemented.

[0136] Figure 5 FIG. is a block diagram of an electronic device 300 shown according to an exemplary embodiment. For example, the electronic device 300 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0137] Referring to Figure 5 , the electronic device 300 may include one or more of the following components: a processing component 302, a memory 304, a power component 306, a multimedia component 308, an audio component 310, an input / output interface 312, a sensor component 314, and a communication component 316.

[0138] The processing component 302 generally controls the overall operation of the electronic device 300, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The processing component 302 may include one or more processors 320 to execute instructions to complete all or part of the steps of the above method for generating an image. In addition, the processing component 302 may include one or more modules to facilitate the interaction between the processing component 302 and other components. For example, the processing component 302 may include a multimedia module to facilitate the interaction between the multimedia component 308 and the processing component 302.

[0139] The memory 304 is configured to store various types of data to support the operation of the electronic device 300. Examples of such data include instructions for any application or method operating on the electronic device 300, contact data, phone book data, messages, pictures, videos, and the like. The memory 304 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.

[0140] The power supply component 306 provides power to various components of the electronic device 300. The power supply component 306 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device 300.

[0141] The multimedia component 308 includes a screen that provides an output interface between the electronic device 300 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can not only sense the boundaries of touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 308 includes a front camera and / or a rear camera. When the electronic device 300 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capabilities.

[0142] The audio component 310 is configured to output and / or input audio signals. For example, the audio component 310 includes a microphone (MIC) that is configured to receive external audio signals when the electronic device 300 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 304 or transmitted via the communication component 316. In some embodiments, the audio component 310 further includes a speaker for outputting audio signals.

[0143] The input / output interface 312 provides an interface between the processing component 302 and a peripheral interface module, and the peripheral interface module can be a keyboard, a click wheel, buttons, etc. These buttons can include, but are not limited to: a home button, a volume button, a start button, and a lock button.

[0144] The sensor assembly 314 includes one or more sensors for providing status assessments of various aspects of the electronic device 300. For example, the sensor assembly 314 can detect the on / off state of the electronic device 300, the relative positioning of components, such as the display and keypad of the electronic device 300. The sensor assembly 314 can also detect a change in the position of the electronic device 300 or a component of the electronic device 300, the presence or absence of user contact with the electronic device 300, the orientation or acceleration / deceleration of the electronic device 300, and a change in the temperature of the electronic device 300. The sensor assembly 314 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 314 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 314 can also include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0145] The communication component 316 is configured to facilitate communication between the electronic device 300 and other devices in a wired or wireless manner. The electronic device 300 can access a wireless network based on communication standards, such as WiFi, 2G, or 3G, or a combination thereof. In an exemplary embodiment, the communication component 316 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 316 further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0146] In an exemplary embodiment, the electronic device 300 can be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above-described image generation method.

[0147] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions, such as the memory 304 including instructions, is also provided. The above instructions can be executed by the processor 320 of the electronic device 300 to complete the above-described image generation method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.

[0148] In another exemplary embodiment, a computer program product is also provided. The computer program product includes a computer program that can be executed by a programmable device. The computer program has a code portion for executing the above-described method for generating an image when executed by the programmable device.

[0149] In an embodiment of the present disclosure, a vehicle is also provided, including the above-mentioned electronic device 300.

[0150] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the present disclosure. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.

[0151] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.

Claims

1. A method for generating an image, characterized in that, The method includes: Obtaining the perspective information of the target perspective corresponding to the target scene, the environmental light information, and the position information of multiple sampling points within the target area; the target area is the perspective area of the target scene under the target perspective; Inputting the perspective information, the environmental light information, and the position information of each sampling point into a pre-trained target information generation model to obtain the color information and volume density information corresponding to each sampling point output by the target information generation model; Generating a target image of the target perspective under the environmental light information according to the multiple color information and the multiple volume density information.

2. The method according to claim 1, characterized in that, The target information generation model is trained in the following manner: Obtaining a plurality of scene images, which are images of different perspectives collected for the target scene, and the environmental light corresponding to each collection of the scene images is different; For each scene image, determining the sample perspective parameters of the collection perspective corresponding to the scene image, the sample position parameters of multiple sample points within the sampling area, and the sample pixel parameters corresponding to each pixel point in the scene image; The sampling area is the perspective area of the target scene under the collection perspective; Training a preset neural radiance field according to the multiple sample perspective parameters, the multiple sample position parameters, and the multiple sample pixel parameters to obtain the target information generation model.

3. The method according to claim 2, wherein The training of the preset neural radiance field according to the multiple sample perspective parameters, the multiple sample position parameters, and the multiple sample pixel parameters to obtain the target information generation model includes: For each scene image, determining a target loss value according to the sample perspective parameters, the multiple sample position parameters, and the sample pixel parameters; Updating the model parameters of the preset neural radiance field according to the target loss value until the preset neural radiance field converges to obtain the target information generation model.

4. The method according to claim 3, wherein The determining of the target loss value according to the sample perspective parameters, the multiple sample position parameters, and the sample pixel parameters includes: Inputting the sample perspective parameters, the multiple sample position parameters, and randomly generated sample environmental light parameters into the preset neural radiance field to obtain the sample color values and sample volume densities corresponding to each sample point output by the preset neural radiance field; Determining the target loss value according to the multiple sample color values, the multiple sample volume densities, and the sample pixel parameters.

5. The method according to claim 4, characterized in that, The determining of the target loss value according to the multiple sample color values, the multiple sample volume densities, and the sample pixel parameters includes: Determining the sample pixel value corresponding to the sample point according to the multiple sample color values and the multiple sample volume densities; Determining the target loss value according to the sample pixel value and the sample pixel parameters, where the target loss value is used to characterize the deviation degree between the sample pixel value and the sample pixel parameters.

6. The method according to claim 1, wherein The target scene includes multiple sub-scenes, and the target information generation model is trained in the following manner: For each of the sub-scenarios, obtain a plurality of sub-scenario images corresponding to the sub-scenario. The plurality of sub-scenario images are images from different perspectives collected for the sub-scenario, and the ambient light corresponding to each collection of the sub-scenario images is different; Train a preset neural radiance field based on the plurality of sub-scenario images to obtain a pending information generation model corresponding to the sub-scenario; Perform view synthesis on the pending information generation models corresponding to the obtained sub-scenarios to obtain the target information generation model.

7. The method according to any one of claims 1 to 6, characterized in that The generating the target image of the target view under the ambient light information according to the plurality of color information and the plurality of volume density information includes: Determine the target pixel value corresponding to each sampling point according to the plurality of color information and the plurality of volume density information; Generate the target image of the target view under the ambient light information based on the plurality of target pixel values.

8. An image generation device, characterized in that, The apparatus includes: An acquisition module, configured to acquire the view information of the target view corresponding to the target scene, the ambient light information, and the position information of a plurality of sampling points within the target area; the target area is the view area of the target scene under the target view; A prediction module, configured to input the view information, the ambient light information, and the position information of each sampling point into a pre-trained target information generation model to obtain the color information and volume density information corresponding to each sampling point output by the target information generation model; A generation module, configured to generate the target image of the target view under the ambient light information according to the plurality of color information and the plurality of volume density information.

9. An electronic device, characterized in that, Comprising: A processor; A memory for storing processor-executable instructions; Wherein, the processor is configured to implement the steps of the method according to any one of claims 1 to 7 when calling the executable instructions stored on the memory.

10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the program instructions are executed by the processor, the steps of the method according to any one of claims 1 to 7 are implemented.

11. A vehicle, characterized in that, An electronic device including the above-mentioned electronic device according to claim 9.