An image rendering method, apparatus, device, and storage medium

By obtaining the location and color-related attributes of the three-dimensional point cloud, determining the point set and pixel values corresponding to the pixel points, the problem of poor rendering effects in traditional rendering technology is solved, and a more realistic rendering effect is achieved.

CN119444955BActive Publication Date: 2025-07-18IFLYTEK CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510032187.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-07-18
Estimated Expiration
2045-01-09

AI Technical Summary

Technical Problem

The existing traditional rendering technology usually implements rendering based on position information and color information, resulting in insufficient rendering effect.

Method used

Obtain a three-dimensional point cloud. The properties of the three-dimensional point include position attributes and color-related attributes, such as color information, transparency and reflectivity, determine the set of points corresponding to each pixel in the rendered image under the target view, and determine the pixel value based on these attributes.

Benefits of technology

By increasing the participation of attributes such as transparency and reflectivity, richer representation information is provided, making the rendered image closer to the real situation and improving the rendering effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119444955B_ABST
    Figure CN119444955B_ABST
Patent Text Reader

Abstract

The present application discloses an image rendering method, apparatus, device, and storage medium. The method includes: obtaining a three-dimensional point cloud, where the three-dimensional point cloud includes a plurality of three-dimensional points, and the attributes of the three-dimensional points include position attributes and color-related attributes, and the color-related attributes include color information and at least one of transparency and reflectivity; determining, from the three-dimensional point cloud, a point set corresponding to each pixel point in a rendered image from a target perspective; and determining the pixel value of each pixel point based on the attributes of the three-dimensional points in the point set corresponding to each pixel point. The above solution can improve the rendering effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing, and particularly to an image rendering method, apparatus, device, and storage medium. Background Art

[0002] Rendering, in computer graphics, refers to the process of converting a 3D model, animation, or scene into a 2D image or animation through software algorithms.

[0003] Existing traditional rendering techniques usually achieve rendering based on position information and color information, but the fineness of the rendering effect in this way is not high. How to improve the image rendering effect has become an urgent problem to be solved. Summary of the Invention

[0004] This application provides at least one image rendering method, apparatus, device, and storage medium.

[0005] This application provides an image rendering method, including: obtaining a three-dimensional point cloud, the three-dimensional point cloud including a plurality of three-dimensional points, the attributes of the three-dimensional points including position attributes and color-related attributes, the color-related attributes including color information, and further including at least one of transparency and reflectivity; determining, from the three-dimensional point cloud, a point set corresponding to each pixel point in the rendering graph from a target perspective; and determining the pixel value of each pixel point based on the attributes of the three-dimensional points in the point set corresponding to each pixel point.

[0006] This application provides an image rendering apparatus, including an obtaining module, a point selection module, and a pixel value determination module. The obtaining module is used to obtain a three-dimensional point cloud, the three-dimensional point cloud including a plurality of three-dimensional points, the attributes of the three-dimensional points including position attributes and color-related attributes, the color-related attributes including color information, and further including at least one of transparency and reflectivity; the point selection module is used to determine, from the three-dimensional point cloud, a point set corresponding to each pixel point in the rendering graph from a target perspective; and the pixel value determination module is used to determine the pixel value of each pixel point based on the attributes of the three-dimensional points in the point set corresponding to each pixel point.

[0007] This application provides an electronic device, including a memory and a processor, the processor being used to execute program instructions stored in the memory to implement any of the above methods.

[0008] This application provides a computer-readable storage medium, on which program instructions are stored, and when the program instructions are executed by a processor, any of the above methods is implemented.

[0009] In the above solution, position attributes and color-related attributes are attached to the three-dimensional points, where the color-related attributes include color information, and further include at least one of transparency and reflectivity, and at least one of transparency and reflectivity participates in rendering, providing richer characterization information of the object to be rendered, making the rendering graph closer to the real situation, and improving the rendering effect.

[0010] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, rather than limiting the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The accompanying drawings herein are incorporated into and constitute a part of this specification. These drawings illustrate embodiments consistent with the present application and, together with the specification, are used to explain the technical solutions of the present application.

[0012] Figure 1 is a schematic flowchart of an embodiment of the image rendering method of the present application;

[0013] Figure 2 is a schematic diagram of three-dimensional points in an embodiment of the image rendering method of the present application;

[0014] Figure 3 is a schematic diagram of sampling points in an embodiment of the image rendering method of the present application;

[0015] Figure 4 is the present application Figure 1 is a schematic flowchart of another embodiment of step S110 in the present application;

[0016] Figure 5 is a schematic diagram of another embodiment of the image rendering method of the present application;

[0017] Figure 6 is a schematic framework diagram of an embodiment of the image rendering device of the present application;

[0018] Figure 7 is a schematic framework diagram of an embodiment of the electronic device of the present application;

[0019] Figure 8 is a schematic framework diagram of an embodiment of the computer-readable storage medium of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0020] The solutions of the embodiments of the present application will be described in detail below with reference to the accompanying drawings of the specification.

[0021] In the following description, specific details such as specific subsystem structures, interfaces, and technologies are set forth for the purpose of illustration and not limitation in order to provide a thorough understanding of the present application.

[0022] As used herein, the term "and / or" merely describes an associated relationship between associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. Additionally, the character " / " in this text generally indicates that the associated objects before and after are in an "or" relationship. Furthermore, "plurality" in this text means two or more than two. Additionally, the term "at least one" in this text means any one of a plurality or any combination of at least two of a plurality. For example, including at least one of A, B, and C may represent including any one or more elements selected from the set composed of A, B, and C.

[0023] Refer to Figure 1 , Figure 1 is a schematic flowchart of an embodiment of the image rendering method of the present application. This method can be executed by a processing device. Specifically, this method may include:

[0024] Step S110: Obtain a three-dimensional point cloud.

[0025] It can be understood that the above three-dimensional point cloud can represent the object to be rendered. Among them, the three-dimensional point cloud includes a plurality of three-dimensional points. Each three-dimensional point has its attributes, and the attributes of the three-dimensional points include position attributes and color-related attributes. The position attribute can represent the position of the three-dimensional point, and the color-related attribute can represent the information related to the color of the three-dimensional point. The color-related attribute includes color information and at least one of transparency and reflectivity.

[0026] In some embodiments, the manifestation form of the color information is color components under a plurality of color channels.

[0027] In some implementation scenarios, the color-related attribute includes color information and transparency. In some implementation scenarios, the color-related attribute includes color information and reflectivity. In some implementation scenarios, the color-related attribute includes color information, transparency, and reflectivity.

[0028] Step S120: Determine the point sets corresponding to each pixel point in the rendered image from the three-dimensional point cloud under the target view.

[0029] It can be understood that the image rendering method provided by the embodiments of the present application can be used to obtain a rendered image from a certain perspective. This rendered image can be a two-dimensional image, and the rendering process can be understood as determining the pixel values of each pixel point in the two-dimensional image. Each pixel point can independently determine its pixel value.

[0030] The pixel value of a pixel in the rendered image from the target perspective is affected by the three-dimensional points within the imaging range corresponding to the pixel in that perspective. The imaging range can refer to the spatial range that affects the imaging of the pixel. The pixel values of different pixels can be affected by the same or different three-dimensional points. A single three-dimensional point can also affect the pixel values of one or more different pixels. First, a point set corresponding to each pixel can be determined, and this point set contains the three-dimensional points that affect the pixel value of that point.

[0031] Step S130: Determine the pixel value of each pixel based on the attributes of the three-dimensional points in the point set corresponding to each pixel.

[0032] Among them, the pixel value of each pixel is determined separately, and the pixel value of each pixel can be determined according to the attributes of the three-dimensional points in its corresponding point set.

[0033] In some embodiments, the point sets corresponding to different pixels can include the same three-dimensional points.

[0034] It can be understood that the pixel value of a pixel is affected by the attributes of the three-dimensional points within its imaging range. After determining the point set corresponding to the pixel, the pixel value of the pixel can be determined using the attributes of all the three-dimensional points in the point set.

[0035] Using a three-dimensional point cloud to represent the object to be rendered, at a specific perspective, the three-dimensional points within the imaging range of each pixel are selected from the three-dimensional point cloud to obtain the pixel value of the pixel. The three-dimensional points that affect the pixel value of the point are determined according to the imaging range of the pixel, so that the pixel value of each pixel can be closer to the object to be rendered, making the rendered image closer to the real situation and improving the rendering effect.

[0036] Furthermore, by attaching at least one of a position attribute, color information, and transparency and reflectivity to the three-dimensional points, at least one of transparency and reflectivity participates in the rendering as an attribute, providing richer representation information of the object to be rendered, making the rendered image closer to the real situation and improving the rendering effect.

[0037] Among them, the position attribute includes the spatial position information of the three-dimensional point. Among them, the spatial position information can represent the spatial position of the three-dimensional point. Exemplarily, the spatial position information can be represented by three-dimensional spatial coordinates.

[0038] In some embodiments, the position attribute can also include the control range information of the three-dimensional point in space. The control range information can characterize the control range of the three-dimensional point in space.

[0039] It should be noted that the control range of the three-dimensional points in the image rendering method of the present application can indicate that within this control range, the rendering is related to the attributes of the corresponding three-dimensional points. The control range can be a certain range centered on the spatial position of the point.

[0040] This control range can be used to determine the point set corresponding to the pixel point. Specifically, if the imaging light of the pixel point passes through the control range of a three-dimensional point, this three-dimensional point can be added to the point set corresponding to the pixel point and affect its pixel value. It should be noted that at this time, the three-dimensional point itself may not fall on the imaging light, and the attributes of this three-dimensional point can still act on the rendering of this pixel point.

[0041] It can be understood that the specific shape, size, etc. of the control range can be set according to actual application needs. For example, the control range can be set as a sphere, an ellipsoid, a cube, etc.

[0042] In some implementation scenarios, the control range can be reflected by the attitude information of the three-dimensional point coordinate system and the control distances in the directions of each coordinate axis in the three-dimensional point coordinate system. Among them, the representation form of the attitude information can be set according to actual application needs. Exemplarily, the attitude angle of the three-dimensional point coordinate system relative to the world coordinate system, the rotation matrix, or the vector representation of the coordinate axes of the three-dimensional point coordinate system in the world coordinate system can be used, etc.

[0043] In a specific application scenario, the control range information can include the attitude angles of each coordinate axis of the three-dimensional point coordinate system in the world coordinate system, and the control distances in the directions of each coordinate axis. The control range can be defined as an ellipsoid, and the control distances in the directions of each coordinate axis can be used as the three semi-axis lengths of the ellipsoid.

[0044] Refer to Figure 2 , Figure 2 which is a schematic diagram of the three-dimensional points in an embodiment of the image rendering method of the present application.

[0045] The attributes of the three-dimensional points include position attributes and color-related attributes. The position attributes include the spatial position information of the three-dimensional points and the control range information of the three-dimensional points in space. The color-related attributes include color information, transparency, and reflectivity. The color information is the color components under multiple color channels.

[0046] Among them, the spatial position information can be expressed as =(x, y, z). The control range information can include the attitude information of the three-dimensional point coordinate system and the control distances in the directions of each coordinate axis in the three-dimensional point coordinate system, and can be expressed as =(α, β, γ, δ α , δ β , δ γ ). Among them, α, β, γ represent the directions of the three axes of this three-dimensional point, and δα , δ β , δ γ represent the control distances in three axial directions.

[0047] Specifically, α, β, and γ can respectively represent the vectors of three axes. The control range of a three-dimensional point can be constructed as an ellipsoid, and δ α , δ β , δ γ can represent the three semi-axis lengths of the ellipsoid, as shown in Figure 2 .

[0048] The color information can be represented as = (r, g, b). The transparency can be represented as , and the reflectivity can be represented as .

[0049] The meanings of the above properties: The position represents the coordinates of this point in space; the transparency represents the ability of light to pass through this point, with a range between 0 and 1, and the larger the value, the more transparent this point is; the reflectivity represents the ability of light to reflect color after encountering this point, with a range between 0 and 1, and the larger the value, the more the color will be enhanced after the light passes through this point; the color represents the inherent color attribute RGB of this point; the control range information mainly represents the range shape of the spatial area controlled by the point.

[0050] In some embodiments, determining the pixel value of each pixel point based on the attributes of the three-dimensional points in the point set corresponding to each pixel point may include: accumulating the first sub-component and / or the second sub-component corresponding to each three-dimensional point in the point set to obtain the pixel value of the pixel point.

[0051] It should be noted that the color-related attributes may include at least one of transparency and reflectivity, and both of them will affect the pixel value of the pixel point. The specific calculation of the pixel value may be related to the content included in the color-related attributes.

[0052] Specifically, for each three-dimensional point in the point set, the first sub-component and / or the second sub-component can be calculated. The former can be obtained based on the transparency and color information, representing the color affected by the transparency, and the latter can be obtained based on the reflectivity and color information, representing the color affected by the reflectivity. When the color-related attributes include transparency, the first sub-component can be calculated, and when the color-related attributes include reflectivity, the second sub-component can be calculated. The pixel value is obtained by accumulating all the components of all the three-dimensional points.

[0053] In some implementation scenarios, color information can be represented as color components under multiple color channels, and the pixel value of the final pixel point can also be represented as the target color components under the corresponding color channels. For example, they are the three components under the RGB channel. Each color channel is independent, and the target color component under a color channel is calculated through the color components of the three-dimensional points in the point set under that channel.

[0054] Furthermore, transparency and reflectivity can equally act on the color components of each color channel. The first sub-component and / or the second sub-component under this channel can be calculated separately for each color channel to obtain the target color component under that channel.

[0055] In some implementation scenarios, for a color channel, the first sub-component and / or the second sub-component of each three-dimensional point in the point set under that color channel are accumulated to obtain the target color component corresponding to the pixel point.

[0056] Among them, the first sub-component of this color channel is the product of the transparency weight of the three-dimensional point and the color component corresponding to this color channel, and the second sub-component is the product of the reflectivity weight of the three-dimensional point and the color component corresponding to this color channel.

[0057] Furthermore, the transparency weight is inversely proportional to the transparency of the three-dimensional point. The higher the transparency of the three-dimensional point, the lower the weight of its color information, indicating that the influence of this point on the final pixel point is smaller.

[0058] In some application scenarios, transparency can be expressed as , where transparency represents the ability of light to pass through this point, ranging from 0 to 1. The larger the value, the more transparent this point is. The transparency weight can be expressed as . When the transparency is 1, the first sub-component of this point is zero, indicating that it has no influence on the pixel value from the perspective of transparency.

[0059] Furthermore, the reflectivity weight is directly proportional to the reflectivity. Reflectivity represents the ability of light to reflect color after encountering this point, ranging from 0 to 1. The larger the value, the more the color will be enhanced after the light passes through this point.

[0060] In some embodiments, the reflectivity weight can also be associated with the positional relationship between the imaging light of the pixel point and the three-dimensional point. Among them, the imaging light of the pixel point can be understood as the light emitted by the camera and related to the imaging of a pixel point. Each pixel point can correspond to its imaging light, and the imaging light can represent the imaging range of the pixel point. The positional relationship between the imaging light of a pixel point and a three-dimensional point is associated with the reflectivity weight between this three-dimensional point and this pixel point.

[0061] Specifically, the distance between the imaging light ray corresponding to a pixel and the three-dimensional point can be referred to as the target distance, and the reflectivity weight can be proportional to the target distance.

[0062] That is to say, for the same three-dimensional point, the transparency is its fixed attribute, and the transparency weight is only related to the transparency. For any pixel point through which the imaging light ray passes through the control range of the three-dimensional point, the transparency weight is consistent.

[0063] The reflectivity is its fixed attribute, but the reflectivity weight is not only related to the reflectivity but also related to the position of the imaging light ray. The positions of different imaging light rays passing through the control range of the three-dimensional point are different, and the distances from the three-dimensional point can also be different. Accordingly, the reflectivity weights can also be different.

[0064] In some embodiments, the reflectivity weight can be the product of the ratio of the target distance to the target square root and the reflectivity of the three-dimensional point. Wherein, the target square root is the root of the sum of the squares of the control distances in the directions of the respective coordinate axes in the three-dimensional point coordinate system.

[0065] In some embodiments, the target distance can be determined according to the spatial position information of the imaging light ray and the three-dimensional point. For example, the minimum distance between the imaging light ray and the three-dimensional point can be selected.

[0066] In some implementation scenarios, the imaging light ray passes through the control range of the three-dimensional point. The part of the imaging light ray within the control range is a line segment in three-dimensional space. Sampling is performed on this line segment to obtain a number of sampling points. The spatial distances between each sampling point and the three-dimensional point are respectively counted, and the target distance can be determined based on the spatial distances corresponding to all the sampling points to represent the distance between the imaging light ray and the three-dimensional point. Exemplarily, the central tendency statistical value of the spatial distances corresponding to all the sampling points can be used as the target distance.

[0067] Refer to Figure 3 , Figure 3 is a schematic diagram of sampling points in an embodiment of the image rendering method of the present application.

[0068] Figure 3 shows a schematic diagram of an imaging light ray passing through the control range of a three-dimensional point. Uniform sampling is performed on the part of the imaging light ray passing through the control range to obtain three sampling points S1, S2, and S3.

[0069] In this embodiment, taking the uniform sampling of the part of the imaging light ray within the control range of the three-dimensional point as an example. Assuming that the imaging light ray of a pixel point passes through the control ranges of K three-dimensional points, the rendering equation of the final color on this imaging light ray is as follows:

[0070]

[0071] Among them, The target color component representing any one of the three color channels used is also the color component under the corresponding channel. Represents three sampling points, which represents the average distance from the three sampling points to the three-dimensional point. As shown in the above formula, the RGB three channels are calculated separately.

[0072] The above solution provides a differentiable spatial point rendering algorithm.

[0073] In some embodiments, when determining the point sets corresponding to each pixel point in the rendered image from the three-dimensional point cloud under the target view, each pixel point can be processed independently. For a pixel point, it is determined whether the imaging ray of the pixel point passes through the control range of the three-dimensional point, and all the three-dimensional points determined to be "yes" are used to form the point set corresponding to the pixel point.

[0074] Specifically, according to the spatial position information and control range information of the three-dimensional point, the control range of the three-dimensional point can be determined, so as to determine whether the imaging ray of a certain pixel point under a specific view passes through the control range. All the three-dimensional points through which the imaging ray of the pixel point passes constitute the corresponding point set.

[0075] In some implementation scenarios, the position information of the imaging ray of a certain pixel point under a specific view is determined, and the position information of the control range of the three-dimensional point is also determined. Therefore, it is possible to solve whether the two intersect to determine whether the imaging ray passes through the control range of the three-dimensional point.

[0076] It should be noted that the attributes of the three-dimensional points used in the final rendering can be learned from the first acquisition image collected by the camera. First, an initial value is determined for the attributes of each three-dimensional point, and then the current first prediction image is rendered using the current attributes. Based on the difference between the first acquisition image actually collected by the camera and the first prediction image, the attributes of the three-dimensional points are iteratively optimized.

[0077] One or more first acquisition images collected by the camera can be used. All cameras collect the object to be rendered, and the pose information corresponding to different cameras is different, that is, the acquisition perspectives are different.

[0078] The first acquisition image can be regarded as the expected rendering ground truth, while the first prediction image is the prediction result rendered using the current attributes. Based on the difference between the two, the attributes of the three-dimensional points are optimized so that the attributes of the three-dimensional points can more accurately represent the object to be rendered.

[0079] In some cases, there may be a certain deviation in the pose of the camera. The pose information of the camera can also be used as an object for adjustment, and there is a certain tolerance for the deviation of the camera pose.

[0080] Specifically, after one rendering, one of the pose information and the attribute is used as the object of this adjustment: the pose information is fixed and the attribute is adjusted, or the attribute is fixed and the pose information is adjusted.

[0081] It is understandable that the camera's posture information will affect the position information of the imaging light of the pixel point. After the camera's posture information is updated, the three-dimensional points in the point set used by the pixel point will change in the subsequent rendering process.

[0082] See also Figure 4 , Figure 4 This application Figure 1 Schematic diagram of another embodiment of step S110 in FIG. Specifically, this step may include:

[0083] Step S411: respectively determine the initial values of the attributes of each three-dimensional point to generate a three-dimensional point cloud.

[0084] Each 3D point has multiple attributes, and an initial value is determined for each attribute, thereby generating 3D points with attributes to form a 3D point cloud.

[0085] Step S412: Rendering the current first predicted image of each camera based on the current attributes of the multiple three-dimensional points and the current posture information of at least one camera.

[0086] Among them, one or more first acquired images may be used, and different first acquired images are acquired by cameras in different positions and postures.

[0087] Rendering is performed independently for each camera, and the current posture of the camera and the current attributes of multiple three-dimensional points are used to render a current first predicted image of the camera, which is compared with the first acquired image of the camera.

[0088] The step of rendering the first predicted image of the camera in step S412 can refer to the rendering step in the aforementioned embodiment. Specifically, under the current position information of the camera, the imaging light of each pixel point in the expected first predicted image is determined, and a point set corresponding to the pixel point is selected from all three-dimensional points, and the current pixel value is determined using the current attribute of the three-dimensional point in the point set to obtain the current first predicted image.

[0089] Step S413: Based on the difference between each current first predicted image and the first captured image of the corresponding camera, adjust one of the attributes of the multiple three-dimensional points and the position and posture information of each camera.

[0090] After step S413, the process may return to step S412 to repeat the iterative optimization. The adjustment object of one optimization is one of the attributes of the multiple three-dimensional points and the position and posture information of each camera, and the other is fixed.

[0091] Further, the step of taking the attributes of the three-dimensional points as the adjustment object is repeatedly executed for the first execution number of times, and the step of taking the pose information of the camera as the adjustment object is repeatedly executed for the second execution number of times. The first execution number is greater than a preset multiple of the second execution number.

[0092] In some embodiments, the first execution number is greater than ten times the second execution number.

[0093] In a specific application scenario, for every 1000 optimizations of the attributes of the three-dimensional points, 10 optimizations of the camera pose are performed. This realizes fine-tuning of the camera pose.

[0094] In some embodiments, the second execution number can be restricted, for example, a threshold can be set for the second execution number. In some embodiments, it is also possible to determine to stop optimizing the pose information according to the degree of optimization of the pose information of the camera at one time. For example, if the difference in pose information after optimization and before optimization is less than a certain threshold, the optimization can be stopped. The above conditions can also be applied jointly, and the optimization can be stopped as long as any one of them is satisfied.

[0095] Among them, the initial value of the spatial position information can be obtained by image extraction, and the initial values of other attributes can be preset values or randomly generated values. Specifically, the initial value of the transparency can be a preset value, for example, it can be 0. The initial value of the reflectivity can be a randomly generated value. The attitude information of the three-dimensional point coordinate system can be a preset value, for example, it can represent being consistent with the direction of the world coordinate system. The control distance can be a preset value, for example, δ α 、δ β 、δ γ can be preset to 1. The color information can be a randomly generated value.

[0096] In some embodiments, the initial value of the spatial position information of the three-dimensional points can be extracted from the second acquisition images collected by multiple cameras. Specifically, based on the second acquisition images of multiple cameras, point cloud extraction is performed to obtain the spatial position information of multiple initial three-dimensional points; clustering is performed on the multiple initial three-dimensional points to obtain the initial value of the spatial position information of the multiple three-dimensional points.

[0097] Among them, the second acquisition images are obtained by collecting the object to be rendered. The number of the second acquisition images is multiple. The second acquisition images and the first acquisition images can be the same or different.

[0098] In a specific application scenario, a point cloud extraction algorithm can be used to obtain the spatial position information of the initial three-dimensional points. Exemplarily, algorithms such as Patchmatch can be used. The input of the algorithm includes the camera poses of multiple cameras and the corresponding second acquisition images, and the output is the spatial position information of multiple initial three-dimensional points.

[0099] In a specific application scenario, the poses of N cameras and the corresponding second captured images are input, the spatial position information of M initial 3D points is output, and the M initial 3D points are divided into m categories by clustering algorithms such as K-means, and the centroids of the m categories are calculated as m 3D points, thus obtaining the initial values of the spatial position information of the m 3D points.

[0100] It can be understood that the target view can be a new view different from the cameras of the captured images. Before rendering the new view, the number of 3D points included in the point set corresponding to each pixel point can be determined first. If the number of 3D points in the point set corresponding to a pixel point is less than the preset number threshold, the number of 3D points needs to be expanded based on the existing 3D points. Specifically, at least one 3D point in the point set is used as the 3D point to be split, and multiple new 3D points are derived based on one 3D point to be split.

[0101] In some embodiments, for a 3D point to be split, multiple spatial points are selected within the control range of the 3D point to be split as new 3D points, and the attributes of the new 3D points are determined based on the attributes of the 3D point to be split.

[0102] Among them, the number of new 3D points derived from one 3D point to be split can be set according to actual application needs. Where to select the new 3D points within the control range can also be set according to actual application needs.

[0103] Furthermore, the spatial position information of the new 3D points can be determined according to the actual situation. The control range information includes pose information and control distance. The pose information of the new 3D points is the same as that of the 3D point to be split. The control distance of the new 3D points is less than that of the 3D point to be split. The new color-related attributes are the same as those of the 3D point to be split.

[0104] In some embodiments, after multiple new 3D points are derived, the second predicted images of each camera can also be rendered based on the attributes of the current 3D points and the pose information of at least one camera, and the attributes of the current 3D points are adjusted based on the differences between the second predicted images and the third captured images of the corresponding cameras. Among them, the above steps can be executed several times.

[0105] The third captured image, the second captured image, and the first captured image can be the same. They are captured by cameras with the same pose for the object to be rendered.

[0106] Refer to Figure 5 , Figure 5 is a schematic diagram of another embodiment of the image rendering method of the present application.

[0107] In this embodiment, a point cloud generation module with attributes and a new view rendering module are provided to implement the image rendering method.

[0108] Among them, the input of the point cloud generation module with attributes may include the acquired images with camera pose information, and the output is the three-dimensional point cloud with attributes.

[0109] The input of the new view rendering module includes the three-dimensional point cloud with attributes and the pose information of the target view, and the output is the rendering image of the target view.

[0110] It can be understood that the pose information of the camera is different from the pose information of the target view, and the target view can be regarded as a new view.

[0111] The point cloud generation module with attributes first generates a three-dimensional point cloud with initial attribute values. Specifically, the acquired images of multiple cameras are used for point cloud extraction, initial three-dimensional points are extracted and clustered to obtain a three-dimensional point cloud with initial values of spatial position information. Then, initial values are attached to other attributes of the three-dimensional points, so as to obtain a three-dimensional point cloud with initial attribute values.

[0112] After that, the acquired images of at least one camera are used for iterative optimization of attributes and camera pose information to obtain the final three-dimensional point cloud with attributes.

[0113] Specifically, by emitting rays from the camera center, the rendering image corresponding to the camera is obtained , given that the real acquired image corresponding to the camera is denoted as I, thus an optimization objective is constructed: , through this optimization objective, information such as the position, direction, transparency, reflectivity, and color corresponding to each point in space can be obtained.

[0114] In addition, fixing the intrinsic attributes of the points obtained above, the camera parameters R|t are optimized, and the optimization objective is . For the above two optimization steps, generally, the attributes are iterated 1000 times and the camera parameters are iterated 10 times each time.

[0115] The attributes and camera parameters are iteratively looped, and finally the set iteration stop condition is reached. The attributes of these points are recorded and saved to obtain the final three-dimensional point cloud with attributes for use by the new view rendering module.

[0116] The new view rendering module first calculates the number N of point clouds passed through by the rays emitted by each pixel point according to the pose information of the target view. If the number N is less than 100, the points need to be split.

[0117] Assume that the attribute of the point to be split is the position =(x, y, z), the control range information =(α, β, γ, δ α , δ β , δ γ ), transparency , reflectivity , color = (r, g, b). Three new three-dimensional points are derived from one point. Specifically, the midpoints of the point to be split on the three three-dimensional point coordinate axes are used as the new three-dimensional points. The midpoint refers to the midpoint between the point to be split and the point at the control distance on the three-dimensional point coordinate axis. Exemplarily, on the axis corresponding to α, the point at is used as the new three-dimensional point, and accordingly, the spatial position information of the new three-dimensional point is re-determined. The color-related attributes of the new three-dimensional point are the same as those of the point to be split. The attitude information of the coordinate system of the new three-dimensional point is the same as that of the point to be split, and the control distance in each axis direction becomes half of the original.

[0118] Then the positions of the 3 points are the midpoints of the 3 axes, the orientation , transparency , reflectivity , color = (r, g, b).

[0119] After splitting, the attributes can also be optimized. Specifically, the attribute optimization steps in the foregoing module can be referred to, and the attributes of the split three-dimensional point cloud are optimized using the existing acquired images. The iterative optimization steps here can be much fewer than those in the previous module. After splitting is completed, the new three-dimensional point cloud is used for rendering to obtain the rendered image of the target view.

[0120] In the above solution, the attributes of the point cloud are designed to fit the rendering result of the scene. Each point is assigned attributes of position, direction, transparency, reflectivity, and color respectively, which can better represent the spatial color information and take into account the changes of light after passing through each point, meeting the characterization of some semi-transparent and reflective scenes.

[0121] A differentiable spatial point rendering algorithm is adopted, which can conveniently optimize the point attributes. According to the proposed rendering algorithm, the view rendered under each camera pose is obtained, and the consistency between the rendered view and the original view is ensured through optimization, and the attribute values of each point are obtained to accurately represent the information of the object to be rendered. On this basis, the three-dimensional point rendering can well depict the spatial color information. And this information is used to optimize the camera pose in turn, reducing the impact of incorrect camera pose estimation on the final result.

[0122] In this solution, since the position of each point is explicitly recorded, the scene can be conveniently re-edited. Each point can be edited, and the device can respond to the user's editing instructions to delete, add, and change the three-dimensional points and their attributes.

[0123] Adjust the number of points adaptively to ensure the fineness and stability of rendering for new perspectives. After obtaining the attribute points quickly, directly perform rendering. Through the logic of splitting the range of sparse points during the rendering process, reduce the time for calculating attribute points and improve the fineness of the final rendering result.

[0124] See Figure 6 , Figure 6 is a schematic framework diagram of an embodiment of the image rendering device of the present application.

[0125] In this embodiment, the image rendering device 60 includes an acquisition module 61, a point selection module 62, and a pixel value determination module 63. The acquisition module 61 is used to acquire a three-dimensional point cloud, which includes a plurality of three-dimensional points. The attributes of the three-dimensional points include position attributes and color-related attributes. The color-related attributes include color information and at least one of transparency and reflectivity. The point selection module 62 is used to determine the point sets corresponding to each pixel point in the rendering image from the three-dimensional point cloud, and the pixel value determination module 63 is used to determine the pixel values of each pixel point based on the attributes of the three-dimensional points in the point sets corresponding to each pixel point.

[0126] See Figure 7 , Figure 7 is a schematic framework diagram of an embodiment of the electronic device of the present application.

[0127] The electronic device 70 includes a memory 71 and a processor 72. The processor 72 is used to execute the program instructions stored in the memory 71 to implement the steps in any of the above image rendering method embodiments. In a specific implementation scenario, the electronic device 70 may include, but is not limited to: computer devices, electrical devices, microcomputers, desktop computers, servers. In addition, the electronic device 70 may also include mobile devices such as laptop computers and tablet computers, which are not limited here.

[0128] Specifically, the processor 72 is used to control itself and the memory 71 to implement the steps in any of the above-described method embodiments for image rendering. The processor 72 may also be referred to as a CPU (Central Processing Unit). The processor 72 may be an integrated circuit chip with the ability to process signals. The processor 72 may also be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. Additionally, the processor 72 may be implemented jointly by integrated circuit chips.

[0129] Refer to Figure 8 , Figure 8 which is a schematic framework diagram of an embodiment of the computer-readable storage medium of the present application.

[0130] The computer-readable storage medium 80 provided in this embodiment stores program instructions 81 that can be run by a processor. When the program instructions 81 are executed by the processor, they are used to implement the steps in any of the above-described method embodiments for image rendering.

[0131] The descriptions of the above embodiments tend to emphasize the differences between the embodiments. Their similarities can be referred to each other. For the sake of brevity, they will not be elaborated herein.

[0132] In several embodiments provided in the present application, it should be understood that the disclosed methods and apparatuses can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, units or components can be combined or integrated into another subsystem, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the apparatuses or units can be in electrical, mechanical, or other forms.

[0133] In addition, each functional unit in various embodiments of the present application may be integrated into one processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit. If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods in various embodiments of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

Claims

1. An image rendering method, characterized in that, The method includes: Obtaining a three-dimensional point cloud, where the three-dimensional point cloud includes a plurality of three-dimensional points, and the attributes of the three-dimensional points include position attributes and color-related attributes. The color-related attributes include color information and reflectivity; the position attributes include the spatial position information of the three-dimensional points and the control range information of the three-dimensional points in space; the color information includes color components under multiple color channels. For each pixel point, based on the spatial position information and the control range information of the three-dimensional points, determine whether the imaging light of the pixel point in the target view passes through the control range of the three-dimensional points; use all the three-dimensional points determined to be yes to form the point set corresponding to the pixel point. Accumulate the second sub-components of the three-dimensional points in the color channel in the point set corresponding to the pixel point to obtain the pixel value of the pixel point; the second sub-component of the color channel is the product of the reflectivity weight of the three-dimensional point and the color component corresponding to the color channel. The reflectivity weight is proportional to the reflectivity and proportional to the target distance. The target distance is the central tendency statistical value of the spatial distances corresponding to a number of sampling points. The sampling points are obtained by sampling the part of the imaging light within the control range of the three-dimensional points, and the spatial distance corresponding to the sampling point is the distance between the sampling point and the three-dimensional point.

2. The method according to claim 1, wherein The color-related attributes further include transparency.

3. The method according to claim 1, wherein The control range information of the three-dimensional points in space includes the attitude information of the three-dimensional point coordinate system and the control distances in the directions of each coordinate axis in the three-dimensional point coordinate system.

4. The method according to claim 2, wherein The accumulating the second sub-components of the three-dimensional points in the color channel in the point set corresponding to the pixel point to obtain the pixel value of the pixel point includes: For each color channel, accumulate the second sub-components, or the first sub-components and the second sub-components of each three-dimensional point in the color channel in the point set to obtain the target color component corresponding to the pixel point in the color channel. Wherein, the first sub-component of the color channel is the product of the transparency weight of the three-dimensional point and the color component corresponding to the color channel, and the second sub-component of the color channel is the product of the reflectivity weight of the three-dimensional point and the color component corresponding to the color channel; the transparency weight is inversely proportional to the transparency.

5. The method according to claim 4, wherein The transparency weight is a preset value minus the transparency of the three-dimensional point. And / or, the reflectivity weight is the product of the ratio of the target distance to the target square root and the reflectivity of the three-dimensional point, and the control range information includes the control distances in the directions of each coordinate axis in the three-dimensional point coordinate system; the target square root is the root of the sum of the squares of the control distances in the directions of each coordinate axis in the three-dimensional point coordinate system.

6. The method according to claim 1, characterized in that, The obtaining the three-dimensional point cloud includes: Respectively determine the initial values of the attributes of each three-dimensional point to generate the three-dimensional point cloud. Based on the current attributes of multiple three-dimensional points and the current pose information of at least one camera, render the first predicted image of each camera currently. Adjust one of the attributes of the multiple three-dimensional points and the pose information of each camera based on the differences between the current first prediction images and the first captured images of the corresponding cameras; Return to execute the step of rendering the current first prediction images of the cameras based on the current attributes of the multiple three-dimensional points and the current pose information of at least one camera.

7. The method according to claim 6, wherein The adjusting one of the attributes of the multiple three-dimensional points and the pose information of each camera based on the differences between the current first prediction images and the first captured images of the corresponding cameras includes any one of the following steps: Fix the pose information of the camera, and adjust the attributes of the multiple three-dimensional points based on the differences between the current first prediction images and the first captured images of the corresponding cameras. This step is repeatedly executed until the first execution count; Fix the attributes of the three-dimensional points, and adjust the pose information of each camera based on the differences between the first prediction images and the first captured images of the corresponding cameras. This step is repeatedly executed until the second execution count; Wherein, the first execution count is greater than a preset multiple of the second execution count.

8. The method according to claim 6, wherein The position attribute includes the spatial position information of the three-dimensional points; the respectively determining the initial values of the attributes of the three-dimensional points to generate the three-dimensional point cloud includes: Performing point cloud extraction based on the second captured images of multiple cameras to obtain the spatial position information of multiple initial three-dimensional points; Clustering the multiple initial three-dimensional points to obtain the initial values of the spatial position information of the multiple three-dimensional points.

9. The method according to claim 1, characterized in that, Before determining the pixel values of the pixel points based on the attributes of the three-dimensional points in the point sets corresponding to the pixel points, the method further includes: In response to the number of the three-dimensional points included in the point set being less than a preset number threshold, taking at least one of the three-dimensional points in the point set as a three-dimensional point to be split, and deriving multiple new three-dimensional points based on the three-dimensional point to be split.

10. The method according to claim 9, characterized in that, The position attribute includes the spatial position information of the three-dimensional points and the control range information of the three-dimensional points in space; the deriving multiple new three-dimensional points based on the three-dimensional point to be split includes: Selecting multiple spatial points within the control range of the three-dimensional point to be split as the new three-dimensional points, and determining the attributes of the new three-dimensional points based on the attributes of the three-dimensional point to be split; wherein, the color-related attributes of the new three-dimensional points are the same as those of the three-dimensional point to be split.

11. The method according to claim 9, characterized in that, After taking at least one of the three-dimensional points in the point set as a three-dimensional point to be split and deriving multiple new three-dimensional points based on the three-dimensional point to be split, the method further includes: Rendering the second prediction images of the cameras based on the current attributes of the three-dimensional points and the pose information of at least one camera; Adjusting the current attributes of the three-dimensional points based on the differences between the second prediction images and the third captured images of the corresponding cameras.

12. An image rendering device, characterized in that, The device includes: An acquisition module for acquiring a three-dimensional point cloud, the three-dimensional point cloud including a plurality of three-dimensional points, the attributes of the three-dimensional points including position attributes and color-related attributes, the color-related attributes including color information and reflectivity; the position attributes including the spatial position information of the three-dimensional points and the control range information of the three-dimensional points in space; the color information including color components under a plurality of color channels; A point selection module for each pixel point, based on the spatial position information and the control range information of the three-dimensional points, determining whether the imaging light of the pixel point in the target view passes through the control range of the three-dimensional points; using all the three-dimensional points determined to be yes to form a point set corresponding to the pixel point; A pixel value determination module for accumulating the second sub-components of the three-dimensional points in the color channels in the point set corresponding to the pixel point to obtain the pixel value of the pixel point; the second sub-component of the color channel is the product of the reflectivity weight of the three-dimensional point and the color component corresponding to the color channel, the reflectivity weight is proportional to the reflectivity and proportional to the target distance, the target distance is the central tendency statistical value of the spatial distances corresponding to a plurality of sampling points, the sampling points are obtained by sampling a part of the imaging light within the control range of the three-dimensional points, and the spatial distance corresponding to the sampling point is the distance between the sampling point and the three-dimensional point.

13. An electronic device, characterized in that, It includes a memory and a processor, and program instructions are stored on the memory, and when the program instructions are executed by the processor, the method described in any one of claims 1 to 11 above is implemented.

14. A computer-readable storage medium having program instructions stored thereon, characterized in that, When the program instructions are executed by the processor, the method described in any one of claims 1 to 11 above is implemented.

Citation Information

Patent Citations

  • Indoor scene reconstruction method and device, electronic equipment and medium

    CN116805349A

  • Accurate transparency and local volume rendering

    US20080150943A1