Color prediction model training method, apparatus, and device

By training a color prediction model using deep learning, and combining it with a pinhole camera model and 3D grid partitioning, the problem of inaccurate color prediction on shopping websites was solved, achieving more accurate image rendering and video synthesis.

CN115908975BActive Publication Date: 2026-04-21BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING BAIDU NETCOM SCI & TECH CO LTD
Filing Date
2022-11-15
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing shopping websites are limited by network transmission speed and cloud server storage space, making it impossible to effectively utilize video information to display items, resulting in inaccurate color prediction.

Method used

A color prediction model is trained using deep learning methods. The model uses a neural network to predict the color of sample points, calculates the loss function and updates the model parameters. By combining a pinhole camera model and three-dimensional grid division, the accuracy of color prediction is improved.

Benefits of technology

It improves the accuracy of color prediction, enabling the generation of more accurate rendered images and increasing the number of images of objects in different camera poses, thus improving the smoothness of video compositing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115908975B_ABST
    Figure CN115908975B_ABST
Patent Text Reader

Abstract

The present disclosure provides a color prediction model training method, device and equipment, relates to the technical field of artificial intelligence, and in particular to the technical field of deep learning and image processing. A specific embodiment of the method comprises: obtaining a training sample, wherein the training sample comprises a sample object image of a sample object and a sample color vector of a sample pixel point of the sample object image; sampling in a sample object space of the sample object to obtain a sample sampling point corresponding to the sample pixel point of the sample object image; performing color prediction on the sample sampling point by using a neural network to obtain a sample color vector of the sample sampling point; calculating a loss function based on the sample color vector of the sample sampling point and the sample color vector of the corresponding sample pixel point; and updating parameters of the neural network based on the loss function to obtain a color prediction model. The color prediction model trained by the embodiment has relatively accurate color prediction capability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence technology, specifically to the fields of deep learning and image processing. Background Technology

[0002] With the rapid development of mobile internet technology, online shopping has become a consumption habit for more and more people. Typically, shopping websites display product information to attract users' attention.

[0003] Currently, shopping websites have shifted from displaying images to showcasing videos. However, due to limitations in network transmission speed and cloud server storage, a shopping website can only upload one set of images taken around an item, which are then used to synthesize a video using backend algorithms. Summary of the Invention

[0004] This disclosure provides a color prediction model training method, apparatus, device, storage medium, and program product.

[0005] In a first aspect, embodiments of this disclosure propose a color prediction model training method, comprising: acquiring training samples, wherein the training samples include sample object images of sample objects and sample color vectors of sample pixels of the sample object images; sampling within the sample object space of the sample objects to obtain sample sampling points corresponding to the sample pixels of the sample object images; performing color prediction on the sample sampling points using a neural network to obtain sample color vectors of the sample sampling points; calculating a loss function based on the sample color vectors of the sample sampling points and the sample color vectors of the corresponding sample pixels; and updating the parameters of the neural network based on the loss function to obtain a color prediction model.

[0006] Secondly, embodiments of this disclosure propose an image rendering method, comprising: acquiring an image of an object to be rendered; sampling within the object space of the object to obtain sampling points corresponding to pixels of the image to be rendered; using a color prediction model to predict the color of the sampling points to obtain color vectors of the sampling points, wherein the color prediction model is trained using the method described in the first aspect; and rendering the pixels corresponding to the sampling points based on the color vectors of the sampling points to obtain a rendered image.

[0007] Thirdly, embodiments of this disclosure propose a color prediction model training device, comprising: an acquisition module configured to acquire training samples, wherein the training samples include a sample object image of a sample object and sample color vectors of sample pixels in the sample object image; a sampling module configured to sample within the sample object space of the sample object to obtain sample sampling points corresponding to the sample pixels in the sample object image; a prediction module configured to perform color prediction on the sample sampling points using a neural network to obtain sample color vectors of the sample sampling points; a calculation module configured to calculate a loss function based on the sample color vectors of the sample sampling points and the sample color vectors of the corresponding sample pixels; and an update module configured to update the parameters of the neural network based on the loss function to obtain a color prediction model.

[0008] Fourthly, embodiments of this disclosure propose an image rendering apparatus, comprising: an acquisition module configured to acquire an image of an object to be rendered; a sampling module configured to sample within the object space of the object to obtain sampling points corresponding to pixels of the image to be rendered; a prediction module configured to perform color prediction on the sampling points using a color prediction model to obtain color vectors of the sampling points, wherein the color prediction model is trained using the apparatus described in the third aspect; and a rendering module configured to render the pixels corresponding to the sampling points based on the color vectors of the sampling points to obtain a rendered image.

[0009] Fifthly, embodiments of this disclosure provide an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform a method as described in any implementation of the first aspect or a method as described in any implementation of the second aspect.

[0010] In a sixth aspect, embodiments of this disclosure provide a non-transitory computer-readable storage medium storing computer instructions for causing a computer to perform a method as described in any implementation of the first aspect or the method as described in any implementation of the second aspect.

[0011] In a seventh aspect, embodiments of this disclosure provide a computer program product, including a computer program that, when executed by a processor, implements the method described in any implementation of the first aspect or the method described in any implementation of the second aspect.

[0012] The color prediction model training method provided in this disclosure is based on deep learning to train the color prediction model, thereby enabling the color prediction model to have a more accurate color prediction capability.

[0013] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0014] Other features, objects, and advantages of this disclosure will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings. The drawings are provided for a better understanding of the invention and are not intended to limit the scope of this disclosure. Wherein:

[0015] Figure 1 This is a flowchart of an embodiment of the color prediction model training method according to the present disclosure;

[0016] Figure 2 This is a flowchart of yet another embodiment of the color prediction model training method according to the present disclosure;

[0017] Figure 3 This is a flowchart of an embodiment of the image rendering method according to the present disclosure;

[0018] Figure 4 This is a flowchart of yet another embodiment of the image rendering method according to the present disclosure;

[0019] Figure 5 This is a schematic diagram of the structure of an embodiment of a color prediction model training apparatus according to the present disclosure;

[0020] Figure 6 This is a schematic diagram of the structure of an embodiment of the image rendering apparatus according to the present disclosure;

[0021] Figure 7 This is a block diagram of an electronic device used to implement the color prediction model training method or image rendering method of the embodiments of this disclosure. Detailed Implementation

[0022] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0023] It should be noted that, unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other. This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.

[0024] Figure 1A flow 100 of an embodiment of a color prediction model training method according to the present disclosure is shown. The color prediction model training method includes the following steps:

[0025] Step 101: Obtain training samples.

[0026] In this embodiment, the entity executing the color prediction model training method can obtain training samples.

[0027] The training samples can include sample object images of the sample object and sample color vectors of sample pixels in the sample object images. Multiple sample object images of a single sample object can be obtained by taking photos around the sample object. The sample color vectors of the sample pixels can be 4-dimensional vectors used to represent the color characteristics of the sample pixels. The first dimension can be opacity, and the last three dimensions can be RGB (Red, Green, Blue) color values.

[0028] Step 102: Sample within the sample object space of the sample object to obtain the sample sampling points corresponding to the sample pixels of the sample object image.

[0029] In this embodiment, the aforementioned execution entity can sample within the sample object space of the sample object to obtain the sample sampling points corresponding to the sample pixels of the sample object image.

[0030] Here, the sample object space can be the space where the sample object resides. Typically, multiple images of the sample object are taken around it, and based on these images, a sparse point cloud of the sample object can be determined. Further, by determining the bounding box of the sparse point cloud, the sample object space can be quickly determined. Here, the pixels of the multiple sample object images are transformed to the same coordinate system to obtain the sparse point cloud. Determining the bounding box of the sparse point cloud yields the sample object space. The bounding box of the sparse point cloud is the smallest cuboid that completely encloses the sparse point cloud.

[0031] Here, the sample sampling point can be a point in the sample object space that corresponds to the sample pixel. Since the sample object image is a two-dimensional image and the sample object space is a three-dimensional space, one sample pixel corresponds to a line segment in the sample object space. By sampling the line segment corresponding to the sample pixel, the sample sampling point corresponding to the sample pixel can be obtained.

[0032] Step 103: Use a neural network to predict the color of the sample sampling points to obtain the sample color vector of the sample sampling points.

[0033] In this embodiment, the aforementioned execution entity can use a neural network to predict the color of the sample sampling points and obtain the sample color vector of the sample sampling points.

[0034] Typically, the aforementioned execution entity can generate a sample input vector based on sample sampling points. Inputting this sample input vector into a neural network allows the neural network to learn the sample color vector of the sample sampling points. For example, the sample feature vector of a sample sampling point can be used as the sample input vector. In this case, if the sample feature vector of a sample sampling point is an F-dimensional vector, then the sample input vector is also an F-dimensional vector. Another example is concatenating the sample feature vector of a sample sampling point with the direction of the line segment containing the sample sampling point to form the sample input vector. In this case, if the sample feature vector of a sample sampling point is an F-dimensional vector and the direction of the line segment containing the sample sampling point is a 3-dimensional vector, then the sample input vector is an F+3-dimensional vector. The sample feature vector of a sample sampling point can be used to characterize the features of the sample sampling point. The sample color vector of a sample sampling point can be a 4-dimensional vector used to characterize the color features of the sample sampling point. The first dimension can be opacity, and the last three dimensions can be RGB color values.

[0035] Step 104: Calculate the loss function based on the sample color vector of the sample sampling point and the sample color vector of the corresponding sample pixel.

[0036] In this embodiment, the execution entity can calculate the loss function based on the sample color vector of the sample sampling point and the sample color vector of the corresponding sample pixel.

[0037] Typically, fusing the sample color vectors of sampled points yields the sample pixel values ​​of those sampled points. Similarly, fusing the sample color vectors of sample pixels yields the sample pixel values ​​of those sample pixels. The loss function is derived by calculating the error between the sample pixel values ​​of the sampled points and the sample pixel values ​​of the sample pixels.

[0038] Step 105: Update the parameters of the neural network based on the loss function to obtain the color prediction model.

[0039] In this embodiment, the aforementioned execution entity can update the parameters of the neural network based on the loss function to obtain a color prediction model.

[0040] Typically, a color prediction model is obtained by backpropagating the neural network based on the loss function and updating the network's parameters until the loss function is small enough.

[0041] The color prediction model training method provided in this disclosure is based on deep learning to train the color prediction model, thereby enabling the color prediction model to have a more accurate color prediction capability.

[0042] Continue to refer to Figure 2 This illustrates a flow 200 of yet another embodiment of the color prediction model training method according to the present disclosure. The color prediction model training method includes the following steps:

[0043] Step 201: Obtain training samples.

[0044] In this embodiment, the specific operation of step 201 has been described. Figure 1 The steps in step 101 of the illustrated embodiment are described in detail and will not be repeated here.

[0045] Step 202: Calculate the sample rays passing through the sample pixels using the pinhole camera model.

[0046] In this embodiment, the entity executing the color prediction model training method can use a pinhole camera model to calculate the sample rays passing through the sample pixels.

[0047] Typically, a pinhole camera model can be used to describe the mathematical relationship between the coordinates of a point in three-dimensional space and its projection onto the image plane of an ideal pinhole camera. Here, a sample ray can be a ray that originates from the camera center point of the pinhole camera model and passes through the sample pixel. The origin of the sample ray is the camera center point, and the direction of the sample ray is from the camera center point towards the sample pixel.

[0048] Step 203: Extract the sample line segment that falls within the sample object space from the sample ray.

[0049] In this embodiment, the aforementioned execution entity can extract sample line segments that fall within the sample object space from the sample ray.

[0050] Here, the sample object space can be the space where the sample object resides. Typically, multiple images of the sample object are taken around it, and based on these multiple images, a sparse point cloud of the sample object can be determined. Further, by determining the bounding box of the sparse point cloud, the sample object space can be obtained. Here, sample rays can pass through the sample object space and intersect it at two points. The line segment with these two points as endpoints is the sample line segment.

[0051] Step 204: Sample the points on the sample line segment to obtain the sample sampling points corresponding to the sample pixels.

[0052] In this embodiment, the execution entity can sample points on the sample line segment to obtain sample sampling points corresponding to the sample pixels. For example, multiple points can be randomly sampled on the sample line segment as sample sampling points.

[0053] Step 205: Divide the sample object space into multiple sample three-dimensional grids.

[0054] In this embodiment, the execution entity can divide the sample object space into multiple sample three-dimensional grids. At this time, the spatial position of each sample three-dimensional grid and the feature vectors of the eight vertices of each sample three-dimensional grid can also be recorded. The feature vectors of the vertices can be F-dimensional vectors, used to characterize the features of the vertices.

[0055] Typically, the sample object space is a cube, and the sample 3D lattice is also a cube. For example, the sample object space can be divided into a G×G×G sample 3D lattice.

[0056] Step 206: Calculate the sample feature vector of the sample sampling point based on the sample vertex feature vector of the sample 3D grid where the sample sampling point is located.

[0057] In this embodiment, the execution entity can calculate the sample feature vector of the sample sampling point based on the sample vertex feature vector of the sample three-dimensional grid where the sample sampling point is located.

[0058] Typically, linear interpolation is performed using the feature vectors of the eight vertices of the sample 3D grid where the sample sampling point is located, and this interpolation is used as the sample feature vector of the sample sampling point. The initial values ​​of the feature vectors of the eight vertices of the sample 3D grid are randomly initialized, and their values ​​are continuously updated as the neural network is trained.

[0059] Step 207: Concatenate the sample feature vector of the sample sampling point and the direction of the sample ray to form the sample input vector.

[0060] In this embodiment, the execution entity can concatenate the sample feature vector of the sample sampling point and the direction of the sample ray to form the sample input vector. In this case, if the sample feature vector of the sample sampling point is an F-dimensional vector and the direction of the sample ray is a 3-dimensional vector, then the sample input vector is an F+3-dimensional vector. The sample color vector of the sample sampling point can be a 4-dimensional vector, used to characterize the color features of the sample sampling point. The first dimension can be opacity, and the last three dimensions can be RGB color values.

[0061] Step 208: Input the sample input vector into the neural network to obtain the sample color vector of the sample sampling point.

[0062] In this embodiment, the aforementioned execution entity can input the sample input vector into the neural network to learn the sample color vector of the sample sampling point. For example, if the sample input vector is an F+3 dimensional vector and the sample color vector of the sample sampling point is a 4 dimensional vector, then the input of the created neural network is an F+3 dimensional vector, and the output is a 4 dimensional vector.

[0063] Step 209: Calculate the loss function based on the sample color vector of the sample sampling point and the sample color vector of the corresponding sample pixel.

[0064] Step 210: Update the parameters of the neural network based on the loss function to obtain the color prediction model.

[0065] In this embodiment, the specific operations of steps 209-210 have been described. Figure 1 Steps 104-105 in the illustrated embodiments are described in detail and will not be repeated here.

[0066] from Figure 2 It can be seen from this that, with Figure 1 Compared to the corresponding embodiments, the color prediction model training method in this embodiment emphasizes the sampling and prediction steps. Therefore, the scheme described in this embodiment determines the sample ray using a pinhole camera model, and sampling along the sample ray accurately captures the sample sampling points corresponding to the sample pixels. By linearly interpolating the feature vectors of the eight vertices of the sample 3D grid, the sample feature vector of the sample sampling point can be quickly calculated. Furthermore, concatenating the directions of the sample ray enriches the input information, thereby improving the accuracy of the predicted color vector.

[0067] Further reference Figure 3 The diagram illustrates a flow 300 of an embodiment of an image rendering method according to the present disclosure. The image rendering method includes the following steps:

[0068] Step 301: Obtain the image of the object to be rendered.

[0069] In this embodiment, the entity executing the image rendering method can obtain the image of the object to be rendered. This image can be an image captured by a camera simulating a certain camera pose, but its colors have not yet been rendered.

[0070] Step 302: Sample within the object space of the object to obtain the sampling points corresponding to the pixels of the image to be rendered.

[0071] In this embodiment, the aforementioned execution entity can sample within the object space of the object to obtain the sampling points corresponding to the pixels of the image to be rendered.

[0072] Here, object space can be the space in which the object resides. Typically, multiple images of the object are taken around it, and a sparse point cloud of the object can be determined based on these images. Further, by determining the bounding box of the sparse point cloud, the object space can be quickly determined. Here, transforming the pixels of multiple object images to the same coordinate system yields the sparse point cloud of the object. The object space is then determined by identifying the smallest cuboid that completely encloses the sparse point cloud.

[0073] Here, the sampling point can be a point in the object space corresponding to a pixel. Since the object image is a two-dimensional image and the object space is a three-dimensional space, one pixel corresponds to a line segment in the object space. By sampling the line segment corresponding to the pixel, the sampling point corresponding to the pixel can be obtained.

[0074] Step 303: Use the color prediction model to predict the color of the sampling points and obtain the color vector of the sampling points.

[0075] In this embodiment, the aforementioned execution entity can use a color prediction model to predict the color of the sampling points and obtain the color vector of the sampling points.

[0076] Typically, the aforementioned execution entity can generate an input vector based on the sampling points. Inputting this input vector into a color prediction model allows the model to learn the color vector of the sampling points. For example, the feature vector of a sampling point can be used as the input vector. In this case, if the feature vector of the sampling point is an F-dimensional vector, then the input vector is also an F-dimensional vector. Another example is concatenating the feature vector of the sampling point with the direction of the line segment containing the sampling point to form the input vector. In this case, if the feature vector of the sampling point is an F-dimensional vector and the direction of the line segment containing the sampling point is a 3-dimensional vector, then the input vector is an F+3-dimensional vector. The feature vector of the sampling point can be used to characterize the features of the sampling point. The color vector of the sampling point can be a 4-dimensional vector used to characterize the color features of the sampling point. The first dimension can be opacity, and the last three dimensions can be RGB color values.

[0077] It should be noted that the color prediction model can be implemented using... Figure 1 or Figure 2 The method shown is used for training, and will not be elaborated further here.

[0078] Step 304: Render the pixels corresponding to the sampling points based on the color vectors of the sampling points to obtain the rendered image.

[0079] In this embodiment, the aforementioned execution entity can render the pixels corresponding to the sampling points based on the color vector of the sampling points to obtain a rendered image.

[0080] Typically, the pixel values ​​of the sampled points are obtained by fusing the color vectors of the sampled points. Using these pixel values ​​as the pixel values ​​of their corresponding pixels allows for the rendering of the image to be rendered, thus generating the rendered image.

[0081] The image rendering method provided in this disclosure improves color prediction accuracy by predicting color vectors using deep learning. Furthermore, rendering images based on the predicted color vectors increases the number of images of objects in different camera poses.

[0082] Further reference Figure 4This illustrates a flow 400 of another embodiment of the image rendering method according to the present disclosure. The image rendering method includes the following steps:

[0083] Step 401: Acquire images taken around the object.

[0084] In this embodiment, the entity executing the image rendering method can acquire images taken around the object. Typically, multiple images are taken around the object at different camera poses.

[0085] Step 402: Calculate the camera pose of the image.

[0086] In this embodiment, the aforementioned execution entity can calculate the camera pose of the image.

[0087] Typically, colmap is used to calculate the camera pose of an image. Colmap is a general-purpose motion structure and multi-view stereo pipeline with both graphical and command-line interfaces. It provides extensive functionality for reconstructing ordered and unordered image sets. Camera pose can include the camera's intrinsic and extrinsic parameters. The camera's intrinsic parameters can include the camera's focal length and pixel size. The camera's extrinsic parameters can include rotation and translation matrices.

[0088] Step 403: Perform uniform interpolation based on the camera pose of the image to obtain the image to be rendered.

[0089] In this embodiment, the aforementioned execution entity can perform uniform interpolation based on the camera pose of the image to obtain the image to be rendered.

[0090] For example, take K images around an object. Then, evenly insert two images to be rendered between any two adjacent images to obtain 3K images with different camera poses.

[0091] Step 404: Determine the sparse point cloud of the object based on multiple images.

[0092] In this embodiment, the aforementioned execution entity can determine the sparse point cloud of an object based on multiple images of the object.

[0093] Typically, by transforming the pixels of multiple images of an object to the same coordinate system, a sparse point cloud of the object can be obtained.

[0094] Step 405: Determine the bounding box of the sparse point cloud as the object space.

[0095] In this embodiment, the aforementioned execution entity can determine the bounding box of the sparse point cloud as the object space. Here, determining the smallest cuboid that completely encloses the sparse point cloud allows for rapid determination of the object space.

[0096] Step 406: Calculate the rays passing through the pixels using the pinhole camera model.

[0097] In this embodiment, the aforementioned execution entity can calculate the rays passing through the pixels using a pinhole camera model.

[0098] Typically, a pinhole camera model can be used to describe the mathematical relationship between the coordinates of a point in three-dimensional space and its projection onto the image plane of an ideal pinhole camera. Here, a ray can be a ray that originates at the camera center point of the pinhole camera model and passes through a pixel. The origin of the ray is the camera center point, and the direction of the ray is from the camera center point towards the pixel.

[0099] Step 407: Extract the line segment from the ray that falls within the object space.

[0100] In this embodiment, the aforementioned execution entity can extract line segments that fall within the object space from the ray.

[0101] Here, a ray can pass through the object space and intersect the object space at two points. The line segment with these two points as endpoints is the line segment intercepted from the ray that falls within the object space.

[0102] Step 408: Sample the points on the line segment to obtain the sampling points corresponding to the pixels.

[0103] In this embodiment, the execution entity can sample points on the line segment to obtain sampling points corresponding to the pixels. For example, multiple points can be randomly sampled on the line segment as sampling points.

[0104] Step 409: Divide the object space into multiple three-dimensional grids.

[0105] In this embodiment, the aforementioned execution entity can divide the object space into multiple three-dimensional grids. At this time, the spatial position of each three-dimensional grid and the feature vectors of the eight vertices of each three-dimensional grid can also be recorded. The feature vectors of the vertices can be F-dimensional vectors, used to characterize the features of the vertices.

[0106] Typically, the space of an object is a cube, and a three-dimensional lattice is also a cube. For example, the space of an object can be divided into a G×G×G three-dimensional lattice.

[0107] Step 410: Calculate the feature vector of the sampling point based on the vertex feature vector of the three-dimensional grid where the sampling point is located.

[0108] In this embodiment, the aforementioned execution entity can calculate the feature vector of the sampling point based on the vertex feature vector of the three-dimensional grid where the sampling point is located.

[0109] Typically, linear interpolation is performed based on the feature vectors of the eight vertices of the three-dimensional grid where the sampling point is located, and this feature vector is used as the feature vector of the sampling point.

[0110] Step 411: Concatenate the feature vectors of the sampling points and the direction of the ray into an input vector.

[0111] In this embodiment, the execution entity can concatenate the feature vector of the sampling point and the direction of the ray to form an input vector. In this case, if the feature vector of the sampling point is an F-dimensional vector and the direction of the ray is a 3-dimensional vector, then the input vector is an F+3-dimensional vector. The color vector of the sampling point can be a 4-dimensional vector, used to characterize the color features of the sampling point. The first dimension can be opacity, and the last three dimensions can be RGB color values.

[0112] Step 412: Input the input vector into the color prediction model to obtain the color vector of the sampling point.

[0113] In this embodiment, the aforementioned execution entity can input the input vector into the color prediction model to learn the color vector of the sampling point. For example, if the input vector is an F+3 dimensional vector and the color vector of the sampling point is a 4 dimensional vector, then the input of the color prediction model is an F+3 dimensional vector, and the output is a 4 dimensional vector.

[0114] It should be noted that the color prediction model can be implemented using... Figure 1 or Figure 2 The method shown is used for training, and will not be elaborated further here.

[0115] Step 413: Render the pixels corresponding to the sampling points based on the color vectors of the sampling points to obtain the rendered image.

[0116] In this embodiment, the specific operation of step 413 has been described. Figure 3 Step 304 in the illustrated embodiment is described in detail and will not be repeated here.

[0117] Step 414: Based on the object's image and rendered image, synthesize the object's surrounding video.

[0118] In this embodiment, the aforementioned execution entity can synthesize an encircling video of the object based on the object's image and rendered image.

[0119] Typically, FFmpeg can be used to sequentially composite an object's image and a rendered image into a loop video. FFmpeg is an open-source computer program that can record, convert, and stream digital audio and video.

[0120] from Figure 4 It can be seen from this that, with Figure 3Compared to the corresponding embodiments, the image rendering method in this embodiment emphasizes the interpolation, sampling, prediction, and video synthesis steps. Therefore, the scheme described in this embodiment can obtain images to be rendered under different camera poses through pose interpolation. By determining rays using a pinhole camera model and sampling along these rays, sampling points corresponding to pixels can be accurately obtained. By performing linear interpolation on the feature vectors of the eight vertices of the 3D grid, the feature vectors of the sampling points can be quickly calculated. Furthermore, stitching together the ray directions enriches the input information, thereby improving the accuracy of the predicted color vectors. Image rendering expands the images of objects under different camera poses. Using the expanded images to synthesize a video of a ring object makes the synthesized video smoother to watch.

[0121] Further reference Figure 5 As an implementation of the methods shown in the above figures, this disclosure provides an embodiment of a color prediction model training device, which is similar to... Figure 1 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.

[0122] like Figure 5 As shown, the color prediction model training device 500 of this embodiment may include: an acquisition module 501, a sampling module 502, a prediction module 503, a calculation module 504, and an update module 505. The acquisition module 501 is configured to acquire training samples, which include sample object images of sample objects and sample color vectors of sample pixels in the sample object images. The sampling module 502 is configured to sample within the sample object space of the sample objects to obtain sample sampling points corresponding to the sample pixels in the sample object images. The prediction module 503 is configured to use a neural network to perform color prediction on the sample sampling points to obtain sample color vectors of the sample sampling points. The calculation module 504 is configured to calculate a loss function based on the sample color vectors of the sample sampling points and the sample color vectors of the corresponding sample pixels. The update module 505 is configured to update the parameters of the neural network based on the loss function to obtain a color prediction model.

[0123] In this embodiment, the specific processing of the acquisition module 501, sampling module 502, prediction module 503, calculation module 504, and update module 505 in the color prediction model training device 500, and the resulting technical effects, can be found in the following references: Figure 1 The relevant descriptions of steps 101-105 in the corresponding embodiments will not be repeated here.

[0124] In some optional implementations of this embodiment, the sampling module 502 is further configured to: calculate the sample ray passing through the sample pixel using the pinhole camera model; extract the sample line segment falling within the sample object space from the sample ray; and sample the points on the sample line segment to obtain the sample sampling point corresponding to the sample pixel.

[0125] In some optional implementations of this embodiment, the prediction module 503 is further configured to: divide the sample object space into multiple sample three-dimensional grids; calculate the sample feature vector of the sample sampling point based on the sample vertex feature vector of the sample three-dimensional grid where the sample sampling point is located; concatenate the sample feature vector of the sample sampling point and the direction of the sample ray to form a sample input vector; and input the sample input vector into the neural network to obtain the sample color vector of the sample sampling point.

[0126] In some optional implementations of this embodiment, the sample object includes multiple sample object images; and the color prediction model training device 500 further includes: a first determining module configured to determine the sample sparse point cloud of the sample object based on the multiple sample object images; and a second determining module configured to determine the bounding box of the sample sparse point cloud as the sample object space.

[0127] Further reference Figure 6 As an implementation of the methods shown in the above figures, this disclosure provides an embodiment of an image rendering apparatus, which is similar to... Figure 3 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.

[0128] like Figure 6 As shown, the image rendering apparatus 600 of this embodiment may include: an acquisition module 601, a sampling module 602, a prediction module 603, and a rendering module 604. The acquisition module 601 is configured to acquire an image of an object to be rendered; the sampling module 602 is configured to sample within the object space of the object to obtain sampling points corresponding to the pixels of the image to be rendered; the prediction module 603 is configured to perform color prediction on the sampling points using a color prediction model to obtain the color vector of the sampling points, wherein the color prediction model employs... Figure 5 The device shown is trained; the rendering module 604 is configured to render the pixels corresponding to the sampling points based on the color vector of the sampling points to obtain the rendered image.

[0129] In this embodiment, the specific processing of the acquisition module 601, sampling module 602, prediction module 603, and rendering module 604 in the image rendering apparatus 600, and the resulting technical effects, can be found in reference to [reference needed]. Figure 3 The relevant descriptions of steps 301-304 in the corresponding embodiments will not be repeated here.

[0130] In some optional implementations of this embodiment, the acquisition module 601 is further configured to: capture an image around the object; calculate the camera pose of the image; and perform uniform interpolation based on the camera pose of the image to obtain the image to be rendered.

[0131] In some optional implementations of this embodiment, the object includes multiple images; and the image rendering apparatus 600 further includes: a first determining module configured to determine a sparse point cloud of the object based on the multiple images; and a second determining module configured to determine a bounding box of the sparse point cloud as the object space.

[0132] In some optional implementations of this embodiment, the sampling module 602 is further configured to: calculate the ray passing through the pixel using the pinhole camera model; extract the line segment falling within the object space from the ray; and sample the points on the line segment to obtain the sampling point corresponding to the pixel.

[0133] In some optional implementations of this embodiment, the prediction module 603 is further configured to: divide the object space into multiple three-dimensional grids; calculate the feature vector of the sampling point based on the vertex feature vector of the three-dimensional grid where the sampling point is located; concatenate the feature vector of the sampling point and the direction of the ray into an input vector; and input the input vector into the color prediction model to obtain the color vector of the sampling point.

[0134] In some optional implementations of this embodiment, the image rendering apparatus 600 further includes a synthesis module configured to synthesize an object-based ambient video based on an object image and a rendered image.

[0135] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0136] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0137] Figure 7 A schematic block diagram of an example electronic device 700 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0138] like Figure 7 As shown, device 700 includes a computing unit 701, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 702 or a computer program loaded from storage unit 708 into random access memory (RAM) 703. RAM 703 may also store various programs and data required for the operation of device 700. The computing unit 701, ROM 702, and RAM 703 are interconnected via bus 704. Input / output (I / O) interface 705 is also connected to bus 704.

[0139] Multiple components in device 700 are connected to I / O interface 705, including: input unit 706, such as keyboard, mouse, etc.; output unit 707, such as various types of monitors, speakers, etc.; storage unit 708, such as disk, optical disk, etc.; and communication unit 709, such as network card, modem, wireless transceiver, etc. Communication unit 709 allows device 700 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0140] The computing unit 701 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 701 performs the various methods and processes described above, such as color prediction model training methods or image rendering methods. For example, in some embodiments, the color prediction model training method or image rendering method can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed on device 700 via ROM 702 and / or communication unit 709. When the computer program is loaded into RAM 703 and executed by the computing unit 701, one or more steps of the color prediction model training method or image rendering method described above can be performed. Alternatively, in other embodiments, the computing unit 701 may be configured by any other suitable means (e.g., by means of firmware) to perform a color prediction model training method or an image rendering method.

[0141] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0142] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0143] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0144] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0145] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0146] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, distributed system servers, or servers incorporating blockchain technology.

[0147] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution provided in this disclosure can be achieved, and this is not limited herein.

[0148] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A method for training a color prediction model, comprising: Obtain training samples, wherein the training samples include sample object images of sample objects and sample color vectors of sample pixels of the sample object images; Samples are taken within the sample object space of the sample object to obtain sample sampling points corresponding to the sample pixels of the sample object image; The sample color vector of the sample sampling point is obtained by using a neural network to predict the color of the sample sampling point; The loss function is calculated based on the sample color vector of the sampled points and the sample color vector of the corresponding sample pixels. The parameters of the neural network are updated based on the loss function to obtain a color prediction model; The step of using a neural network to predict the color of the sample sampling points to obtain the sample color vector of the sample sampling points includes: The space of the sample object is divided into multiple three-dimensional sample grids; Based on the feature vector of the sample vertex of the sample 3D grid where the sample sampling point is located, calculate the sample feature vector of the sample sampling point; The sample feature vector of the sample sampling point and the direction of the sample ray are concatenated to form the sample input vector. The sample ray is the ray that starts from the center point of the camera of the pinhole camera model and passes through the sample pixel. The sample input vector is input into the neural network to obtain the sample color vector of the sample sampling point.

2. The method according to claim 1, wherein, The step of sampling within the sample object space of the sample object to obtain sample sampling points corresponding to sample pixels of the sample object image includes: Using a pinhole camera model, calculate the sample ray passing through the sample pixels; Extract a sample line segment from the sample ray that falls within the space of the sample object; The points on the sample line segment are sampled to obtain the sample sampling points corresponding to the sample pixels.

3. The method according to claim 1 or 2, wherein, The sample objects include multiple sample object images; The method also includes: Based on the multiple sample object images, determine the sample sparse point cloud of the sample object; The bounding box of the sample sparse point cloud is determined as the sample object space.

4. An image rendering method, comprising: Obtain the image of the object to be rendered; Sampling is performed within the object space of the object to obtain the sampling points corresponding to the pixels of the image to be rendered; The color prediction model is used to predict the color of the sampling points to obtain the color vector of the sampling points, wherein the color prediction model is trained using the method described in any one of claims 1-3; The pixel corresponding to the sampling point is rendered based on the color vector of the sampling point to obtain the rendered image.

5. The method according to claim 4, wherein, The process of obtaining the image to be rendered of the object includes: Acquire images taken around the object; Calculate the camera pose of the image; The image to be rendered is obtained by uniform interpolation based on the camera pose of the image.

6. The method according to claim 5, wherein, The object includes multiple images; and the method further includes: The sparse point cloud of the object is determined based on the multiple images; The bounding box of the sparse point cloud is determined as the object space.

7. The method according to claim 4, wherein, The step of sampling within the object space of the object to obtain sampling points corresponding to the pixels of the image to be rendered includes: Using a pinhole camera model, calculate the rays passing through the pixels; Extract a line segment from the ray that falls within the space of the object; The points on the line segment are sampled to obtain the sampling points corresponding to the pixel points.

8. The method according to claim 7, wherein, The step of using a color prediction model to predict the color of the sampling points and obtain the color vector of the sampling points includes: The space of the object is divided into multiple three-dimensional grids; Based on the vertex feature vectors of the three-dimensional grid where the sampling point is located, calculate the feature vector of the sampling point; The feature vector of the sampling point and the direction of the ray are concatenated to form the input vector; The input vector is fed into the color prediction model to obtain the color vector of the sampling point.

9. The method according to claim 5 or 6, wherein, The method further includes: Based on the image and rendered image of the object, a surround video of the object is synthesized.

10. A color prediction model training device, comprising: The acquisition module is configured to acquire training samples, wherein the training samples include sample object images of sample objects and sample color vectors of sample pixels of the sample object images; The sampling module is configured to sample within the sample object space of the sample object to obtain sample sampling points corresponding to sample pixels of the sample object image; The prediction module is configured to use a neural network to predict the color of the sample sampling points and obtain the sample color vector of the sample sampling points. The calculation module is configured to calculate a loss function based on the sample color vector of the sample sampling point and the sample color vector of the corresponding sample pixel point; The update module is configured to update the parameters of the neural network based on the loss function to obtain a color prediction model; The prediction module is further configured to: The space of the sample object is divided into multiple three-dimensional sample grids; Based on the feature vector of the sample vertex of the sample 3D grid where the sample sampling point is located, calculate the sample feature vector of the sample sampling point; The sample feature vector of the sample sampling point and the direction of the sample ray are concatenated to form the sample input vector. The sample ray is the ray that starts from the center point of the camera of the pinhole camera model and passes through the sample pixel. The sample input vector is input into the neural network to obtain the sample color vector of the sample sampling point.

11. The apparatus according to claim 10, wherein, The sampling module is further configured to: Using a pinhole camera model, calculate the sample ray passing through the sample pixels; Extract a sample line segment from the sample ray that falls within the space of the sample object; The points on the sample line segment are sampled to obtain the sample sampling points corresponding to the sample pixels.

12. The apparatus according to any one of claims 10-11, wherein, The sample objects include multiple sample object images; The device also includes: The first determining module is configured to determine the sample sparse point cloud of the sample object based on the multiple sample object images; The second determining module is configured to determine the bounding box of the sample sparse point cloud as the sample object space.

13. An image rendering apparatus, comprising: The acquisition module is configured to acquire the image of the object to be rendered; The sampling module is configured to sample within the object space of the object to obtain sampling points corresponding to the pixels of the image to be rendered; The prediction module is configured to use a color prediction model to predict the color of the sampling points and obtain the color vector of the sampling points, wherein the color prediction model is trained using the apparatus of any one of claims 10-12. The rendering module is configured to render the pixels corresponding to the sampling points based on the color vector of the sampling points to obtain a rendered image.

14. The apparatus according to claim 13, wherein, The acquisition module is further configured to: Images taken around the object; Calculate the camera pose of the image; The image to be rendered is obtained by uniform interpolation based on the camera pose of the image.

15. The apparatus according to claim 14, wherein, The object includes multiple images; and the device further includes: The first determining module is configured to determine the sparse point cloud of the object based on the multiple images; The second determining module is configured to determine the bounding box of the sparse point cloud as the object space.

16. The apparatus according to claim 13, wherein, The sampling module is further configured to: Using a pinhole camera model, calculate the rays passing through the pixels; Extract a line segment from the ray that falls within the space of the object; The points on the line segment are sampled to obtain the sampling points corresponding to the pixel points.

17. The apparatus according to claim 16, wherein, The prediction module is further configured to: The space of the object is divided into multiple three-dimensional grids; Based on the vertex feature vectors of the three-dimensional grid where the sampling point is located, calculate the feature vector of the sampling point; The feature vector of the sampling point and the direction of the ray are concatenated to form the input vector; The input vector is fed into the color prediction model to obtain the color vector of the sampling point.

18. The apparatus according to claim 14 or 15, wherein, The device further includes: The compositing module is configured to compose an ambient video of the object based on the image and rendered image of the object.

19. An electronic device comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-3 or the method of any one of claims 4-9.

20. A non-transitory computer-readable storage medium storing computer instructions for causing the computer to perform the method of any one of claims 1-3 or any one of claims 4-9.

21. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-3 or any one of claims 4-9.

Citation Information

Patent Citations

  • Fast volume rendering three-dimensional ultrasonic image reconstruction algorithm introducing scattering model

    CN110298915A

  • Image rendering method and device of three-dimensional object and electronic equipment

    CN114863007A

  • Three-dimensional reconstruction model training method, three-dimensional reconstruction method and device

    CN115147558A