Image processing method and device, equipment and storage medium
By obtaining the array of observation area ranges of the implicit network model and processing abnormal pixel points in the candidate image, the three-dimensional object modeling reliability problem caused by observation blind spots in implicit spatial modeling is solved, and a three-dimensional object rendering effect that is closer to reality is achieved.
Patent Information
- Application Number
- CN202510208236.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-06-27
AI Technical Summary
Modeling technology based on implicit spatial expression may have abnormal pixel points in the three-dimensional space within the observation blind spot, resulting in a reduced reliability of three-dimensional object modeling.
By obtaining the observation area range array of pre-trained implicit network models, identifying and processing pixel points in candidate images that are outside the observation area range array, and generating target images to improve the reliability of three-dimensional object modeling.
By eliminating the abnormal pixel points that were not supervised during the training process of the implicit network model, ensuring that each pixel point in the rendering result is an actual object point, which significantly improves the reliability of three-dimensional object modeling.
Smart Images

Figure CN120219612A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and particularly to an image processing method, apparatus, device, and storage medium. Background Art
[0002] When modeling based on implicit spatial expressions such as the Signed Distance Field (SDF) and other techniques, most optimizations and reconstructions are carried out through sparse multi-view image constraints. Objectively, the coverage of sparse multi-view images is limited, and there may be observation blind spots in space. Since the objects in the three-dimensional space within this observation blind spot have not been supervised and trained, the objects within the observation blind spot are randomly generated by the implicit network model, resulting in the generation of abnormal pixel points. Therefore, there is a situation where abnormal pixel points are rendered in the images of the newly rendered perspectives generated by the implicit network model, causing the rendered abnormal pixel points to be regarded as part of the three-dimensional object itself, thereby reducing the reliability of three-dimensional object modeling. Summary of the Invention
[0003] The main purpose of the embodiments of this application is to propose an image processing method, apparatus, device, and storage medium, which can improve the reliability of three-dimensional object modeling.
[0004] To achieve the above objective, in the first aspect of the embodiments of this application, an image processing method is proposed. The method includes: Obtain a pre-trained implicit network model corresponding to the target object, and the training image data set used by the implicit network model during the pre-training process. The training images in the training image data set are images of at least one rendering perspective containing the target object; Determine an observation area range array according to the implicit network model and the training image data set; Generate a candidate image of the target object according to the implicit network model and the expected rendering perspective; Process the pixel points outside the observation area range array in the candidate image as abnormal pixel points to obtain a target image.
[0005] To achieve the above objective, in the second aspect of the embodiments of this application, a modeling apparatus for a three-dimensional model is proposed. The apparatus includes: An obtaining module, configured to obtain a pre-trained implicit network model corresponding to the target object, and the training image data set used by the implicit network model during the pre-training process. The training images in the training image data set are images of at least one rendering perspective containing the target object; A determining module, configured to determine an observation area range array according to the implicit network model and the training image data set; A modeling module, configured to generate a candidate image of the target object according to the implicit network model and the expected rendering perspective; A rendering module, configured to process the pixel points outside the observation area range array in the candidate image as abnormal pixel points to obtain a target image.
[0006] To achieve the above object, a third aspect of the embodiments of the present application proposes an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the image processing method described in any item of the first aspect above is implemented.
[0007] To achieve the above object, a fourth aspect of the embodiments of the present application proposes a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the image processing method described in any item of the first aspect is implemented.
[0008] The image processing method, device, equipment and storage medium proposed in the present application can, by pre-acquiring the observation area range array of the implicit network model, process the abnormal pixel points randomly generated in the candidate image due to the unsupervised learning in the training process of the implicit network model based on the observation area range array. At this time, each pixel point in the target image obtained by rendering and displaying a new perspective based on the expected rendering perspective is an actual object point, and the display result is closer to the actual three-dimensional object. Therefore, compared with the related art, the embodiments of the present application can improve the reliability of image modeling. Description of the Drawings
[0009] Figure 1 is a schematic diagram of the effect of an image model rendered based on an implicit network model in the prior art; Figure 2 is a schematic flowchart of an embodiment of the image processing method provided in the present application; Figure 3 is a schematic flowchart of an embodiment of the image processing method provided in the present application; Figure 4 is a schematic diagram of the effect of a target image of an embodiment of the image processing method provided in the present application; Figure 5 is a schematic diagram of modules of an embodiment of the device corresponding to the image processing method provided in the present application; Figure 6 is a schematic diagram of the hardware structure corresponding to the image processing method provided in the present application. Detailed Embodiments
[0010] In order to make the objectives, technical solutions and advantages of this application more clear and understandable, the following further details this application in combination with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.
[0011] It should be noted that although functional module division is carried out in the device schematic diagram and the logical sequence is shown in the flowchart, in some cases, the steps shown or described can be executed in a different module division in the device or a different sequence in the flowchart. Terms such as "first" and "second" in the description, claims and the above-mentioned drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence.
[0012] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0013] The following are the explanations of the terms involved in the embodiments of this application: SDF, namely Signed Distance Field, is a method for describing geometric shapes. SDF defines the shape by calculating the distance from any point in space to the nearest object. This distance can be positive, negative or zero, depending on the position of the point relative to the object. Its core idea is to calculate the directed distance from each point in space to the nearest object surface. If the point is inside the object, the distance is negative; if the point is on the object surface, the distance is zero; if the point is outside the object, the distance is positive. This representation method is not only applicable to simple geometric shapes such as circles and squares, but also can be used for complex objects and scenes. SDF can be used for graphics rendering. For example, in Ray Marching, an image is generated by gradually approaching the object surface. SDF can also be used for stylized rendering and cartoon rendering to achieve smooth transitions between images. SDF can also be used for collision detection by calculating the distance between objects to determine whether they intersect.
[0014] IDR: The full English name is Implicit Differentiable Renderer, which is a technology that combines differentiable rendering and neural networks and is used to generate high-quality 3D images and videos. The basic principle of Implicit Differentiable Renderer (IDR) is to combine neural networks with differentiable rendering technology. The neural network is used to model the scene, and the differentiable rendering technology is used for image generation and optimization. Specifically, IDR uses a neural network to learn the geometry, lighting and other attributes of the scene, and converts these attributes into images through differentiable rendering technology.
[0015] In the prior art, modeling based on implicit spatial expression, such as modeling based on SDF and other technologies, is mostly optimized and reconstructed through sparse multi-view image constraints. However, objectively, sparse multi-view images have a limited coverage range, and there may be observation blind spots in three-dimensional space, which leads to the existence of some spaces in the observation blind spots that have not been supervised and trained. Therefore, objects in the observation blind spots will be randomly generated by the implicit network model, resulting in the possibility of abnormal pixels in the images of the observation blind spots rendered based on the new perspective. For example, Figure 1 As shown, when implementing scenes such as fusion rendering, it is necessary to generate images of three-dimensional objects at different perspectives for free viewing. When the expected perspective causes foreign objects to appear due to the observation blind spot, the rendering of abnormal pixels will significantly affect the display of the image at that perspective. Based on this, the embodiment of the present application proposes an image processing method, device, equipment and storage medium, which can adaptively remove abnormal pixels in each scene when rendering at a new perspective after scene training is completed, significantly improving the reliability of implicit space modeling.
[0016] Understandably, referring to Figure 2 As shown, according to an embodiment of the present application, an image processing method is provided, the method comprising: Step S100: obtaining a pre-trained implicit network model corresponding to the target object and a training image data set used by the implicit network model in a pre-training process, wherein the training images in the training image data set are images containing at least one rendering perspective of the target object; Step S200, determining an observation area range array according to the implicit network model and the training image data set; Step S300: Generate a candidate image of the target object according to the implicit network model and the expected rendering perspective.
[0017] Step S400: Pixels in the candidate image that are outside the observation area range array are processed as abnormal pixels to obtain a target image.
[0018] Therefore, by pre-acquiring the observation area range array of the implicit network model, the abnormal pixel points in the candidate image that are randomly generated due to the lack of supervised learning of the implicit network model during the training process can be processed based on the observation area range array. At this time, each pixel point in the target image obtained by rendering and displaying a new perspective based on the expected rendering perspective is an actual object point, and the display result is closer to the actual three-dimensional object. Therefore, the embodiment of the present application can improve the reliability of image modeling.
[0019] The observation area range array represents the trained three-dimensional space range in the implicit network model. The embodiment of the present application does not limit the shape of the observation area range array, which can be a rectangular three-dimensional space or an irregular space, etc.
[0020] The implicit network model is a network model constructed based on a neural network. By training an image dataset, the relationship between the known rendering perspectives of the target object and the training images can be constructed. In the embodiments of the present application, the specific network structure of the implicit network model is not limited, nor are the number of rendering perspectives of the target object and the number of training images under the same rendering perspective. Those skilled in the art can selectively set according to the actual situation.
[0021] It can be understood that in the existing sparse-based multi-view image modeling solutions, limited by the coverage range of the multi-view images provided during training (such as limited by the perspectives of the provided images or the coverage range in a single image), it is difficult to avoid unobserved spaces when inferring new perspective images. As a result, when the implicit network model completed based on pre-training is modeled, abnormal pixel points are likely to be generated in the unobserved areas of the generated images. Therefore, in the embodiments of the present application, by providing an observation area range array, it is possible to quickly identify whether each pixel point of the image generated under the new perspective is within the space supervised by the implicit network model during training.
[0022] Pixels located outside the observation area range array indicate that the pixel is an object point that was not involved in the training of the implicit network model during the training process. Therefore, it is outside the three-dimensional space range that has been trained in the implicit network model. The expected three-dimensional coordinates are the coordinates of each pixel point on the candidate image in a preset coordinate system in three-dimensional space. Therefore, it is possible to determine whether a pixel point is located outside the observation area range array by judging whether there is a mapping relationship between the expected three-dimensional coordinates of the pixel points of the candidate image and the pixels involved in the training within the observation area range array. For example, one or more of the ray tracing methods such as Classic Ray Tracing, Path Tracing, Volume Ray Casting, and Ray Marching can be used for judgment.
[0023] The expected rendering perspective is a perspective different from the rendering perspectives corresponding to the training images. In the embodiments of the present application, there is no limitation on how to process the abnormal pixel points in step S400. Those skilled in the art can selectively set according to the actual visual display requirements. For example, the abnormal pixel points can be hidden when rendering the candidate image. The hiding method can be to set the abnormal pixel points not to participate in the rendering or to render the abnormal pixel points with the background color, so as to visually reduce the perception of the abnormal pixel points and achieve the purpose of hiding. In this regard, those skilled in the art can select a suitable processing method for the abnormal pixel points according to actual needs.
[0024] The candidate image is an image model of the expected rendering perspective generated based on the implicit network model.
[0025] Exemplarily, taking object A as the target object for rendering and modeling, and taking the case where abnormal pixel points are rendered as the background color as an example, assuming that images are collected from view 1, view 2, and view 3 of object A currently, then the images from view 1 to view 3 are used as training images to train the initial implicit network model. After training is completed, images of rendering views other than view 1 to view 3 can be generated through this implicit network model. Before new view rendering, an observation area range array is determined based on the images from view 1 to view 3 and the implicit network model. When the expected rendering view to be generated is view 4, view 4 is input into the implicit network model. At this time, the implicit network model can output a candidate image of object A and render the pixel points outside the observation area range array in the candidate image as abnormal pixel points with the background color, and the remaining pixel points are rendered with reference to the colors of the pixel points at the same positions in views 1 to 3 of object A.
[0026] It can be understood that the above steps S100 to S400 can be configured in a plug-in mode. At this time, only the training image dataset and the trained implicit network model need to be configured for this plug-in to achieve the generation of the target image. In some other embodiments, steps S100 to S400 can also be integrated in 3D modeling software. In this regard, the embodiments of the present application do not limit the product form presented by the above steps in practical applications, and those skilled in the art can selectively set according to actual needs.
[0027] It can be understood that taking the pixel points outside the observation area range array in the candidate image as abnormal pixel points for processing to obtain the target image includes: Traverse each pixel point on the candidate image; Determine the expected three-dimensional coordinates of each traversed pixel point through a preset ray tracing algorithm; When the expected three-dimensional coordinates are not in the observation area range array, the traversed pixel point is used as an abnormal pixel point; Perform foreign object elimination processing on each abnormal pixel point; When all abnormal pixel points in the candidate image have completed foreign object elimination processing, the target image is obtained.
[0028] The ray tracing algorithm is a rendering technique that generates realistic images based on simulating the propagation behavior of light in a scene. Its core principle is to emit a large number of rays from a virtual camera (observation point). These rays will perform intersection tests with the objects in the scene (such as geometric objects like spheres and planes), calculate the intersection points of the rays with the objects and the lighting effects at the intersection points (including reflection, refraction, shadows, etc.), and then determine the color of the final pixel based on this information. By tracing numerous rays, the entire image is constructed. Therefore, through the ray tracing algorithm, it can be calculated whether the pixel point can hit an object point in the trained scene, and at this time, the coordinates of the hit object point are used as the expected three-dimensional coordinates of the corresponding pixel point.
[0029] The embodiments of this application do not limit how to select a specific ray tracing algorithm. Those skilled in the art can choose one or more according to actual needs. For example, ray marching or ray traced caustics can be used for ray tracing according to the material and surface complexity of the target object traversed to the pixel point.
[0030] The embodiments of this application do not limit how to traverse the pixel points of the candidate image. For example, the candidate image can be divided into blocks and traversed one by one, or it can be traversed one by one starting from the center point of the candidate image, or it can be traversed according to different materials of the target object A, etc.
[0031] During foreign object elimination processing, it is used to reduce the external perception of abnormal pixel points, and can be eliminated by setting a background color or a transparent color, etc.
[0032] Therefore, by separately judging whether each pixel point in the candidate image belongs to an abnormal pixel point and performing elimination processing for each abnormal pixel point, the probability that the abnormal pixel points of the target image are perceived can be further reduced, thereby improving the reliability of the target object modeling.
[0033] It can be understood that performing foreign object elimination processing on each abnormal pixel point includes: Obtaining a pre-configured foreign object rendering color; Rendering the colors of each abnormal pixel point as the foreign object rendering color.
[0034] The embodiments of this application do not limit how to set the specific color of the foreign object rendering color. In some embodiments, the foreign object rendering color is set as the background color, and in other embodiments, the foreign object rendering color is set as the transparent color.
[0035] In some embodiments, the foreign object rendering color can be dynamically configured or flexibly set through a configuration file or the like. In this regard, the embodiments of the present application do not impose excessive restrictions, and those skilled in the art can selectively set according to actual needs.
[0036] It can be understood that according to the implicit network model and the training image dataset of the target object, an observation area range array is determined, including: Determine the sample three-dimensional coordinates of each sample pixel point in the training image dataset; According to each sample three-dimensional coordinate, determine the boundary coordinates in a preset plurality of three-dimensional space directions; According to each boundary coordinate, obtain the observation area range array.
[0037] Each sample pixel point is a pixel point on a training image in the training image dataset; the embodiments of the present application do not limit how to obtain the sample three-dimensional coordinates of the sample pixel points. For example, the coordinates of the camera origin in the world coordinate system can be obtained through the external parameters of the camera, and the coordinates of each pixel in the world coordinate system can be calculated through the internal and external parameters of the camera. Further, according to the obtained camera origin coordinates and the coordinates of each pixel, the direction vector of the light emitted from the camera origin through each pixel can be obtained. Given the camera origin coordinates and the direction vector of the light emitted through each pixel, the sample three-dimensional coordinates corresponding to different pixel points on the target object can be determined. In this regard, the embodiments of the present application do not limit the algorithm for determining the sample three-dimensional coordinates based on the camera origin coordinates and the direction vector of the light emitted through each pixel, and those skilled in the art can selectively set according to actual needs.
[0038] Exemplarily, taking the IDR as an example of the implicit network model, where IDR uses the SDF based on the unit sphere to represent the geometric structure of the object and the SDF is composed of an MLP network. The input of the SDF is the three-dimensional point coordinates, and the output of the SDF is the distance from this three-dimensional point to the closest point on the object surface. A positive SDF output indicates that the current three-dimensional point is outside the object geometric structure, a negative value indicates that the three-dimensional point is inside the object geometric structure, and when the SDF value is 0, it means that this three-dimensional point is exactly on the object geometric structure surface. Since all geometric structures in IDR are set inside the unit sphere, when obtaining the object point coordinates corresponding to each pixel point, the intersection of the corresponding light ray and the unit sphere can be calculated first to greatly reduce the computational complexity. Assuming that the light ray has no intersection with the unit sphere, this pixel cannot hit the object point and should belong to the background pixel. In addition, if the light ray intersects the object point, the left and right intersection points are (A, B), that is, the object point coordinates corresponding to this pixel should be searched on the line segment AB. By inputting the coordinates of point A and point B into the SDF network, the corresponding values are obtained. If the output value is 0, the 3D object point corresponding to this pixel point has been found. If the output value is not 0, then at (A + SDF(A), B- SDF(B)), that is, at the point of ( ), continue the iterative search until the SDF output value is 0, indicating that the corresponding 3D point is found, or until it indicates that the pixel does not correspond to an object point and belongs to the background. At this time, by performing the above steps on each pixel point of each training image, the sample three-dimensional coordinates of each pixel point can be determined. Thus, the boundary coordinates can be determined based on the three-dimensional coordinates. It can be understood that when the structure of the implicit network model changes, correspondingly, the method for obtaining the sample three-dimensional coordinates will also change.
[0039] The embodiments of the present application do not limit the three-dimensional space directions. For example, it can be at least two of the X direction, Y direction, or Z direction. There are at least two boundary coordinates in each three-dimensional space direction to define the maximum and minimum coordinate values in that direction. For example, if the observation space corresponding to the observation region range array is a regular three-dimensional space, only the maximum and minimum boundary coordinates need to be set in each three-dimensional space direction. Another example is that in some embodiments, if the observation space corresponding to the observation region range array is irregular, the observation space can be divided into multiple regular three-dimensional spaces, and corresponding boundary coordinates are assigned to each three-dimensional space, so as to obtain all the boundary coordinates of the observation space.
[0040] The preset multiple three-dimensional space directions can be pre-configured in the program, can also be configured through a configuration file, or can be configured by providing a user interface. In this regard, the embodiments of the present application do not limit the preset multiple three-dimensional space directions.
[0041] It can be understood that each three-dimensional coordinate is within the region enclosed by the boundary coordinates.
[0042] The determination of the boundary coordinates can be determined by comparing the coordinate values of each three-dimensional coordinate in different spatial dimensions. For example, for the three-dimensional coordinates L1~Ln, the coordinate values of L1~Ln in each dimension of the three-dimensional space are respectively compared, and the maximum and minimum coordinate values in the same dimension are used as the boundary coordinates for that dimension. For example, if L1.x and L4.x are the maximum and minimum values in the X direction respectively, then L1.x and L4.x are the boundary coordinates in the X direction. Similarly, the boundary coordinates in the Y and Z directions can be determined. At this time, based on the boundary coordinates in the X, Y, and Z directions, the observation region range array can be determined.
[0043] It can be understood that determining the boundary coordinates in the preset multiple three-dimensional space directions according to each three-dimensional coordinate includes: Initializing a boundary threshold coordinate array according to the modeling space scale of the implicit network model, and the boundary threshold coordinate array records the boundary threshold coordinates in each three-dimensional space direction; Compare each three-dimensional coordinate with each boundary threshold coordinate respectively, and update the boundary threshold coordinate according to the comparison result; Take the updated boundary threshold coordinate as the boundary coordinate in the corresponding three-dimensional space direction.
[0044] By initializing the boundary threshold coordinate based on the modeling space scale, the number of updates of the boundary threshold coordinate can be reduced, and the efficiency of determining the boundary coordinate can be improved.
[0045] When there is a coordinate value in a spatial dimension of the three-dimensional coordinate that is greater than the largest boundary threshold coordinate or less than the smallest boundary threshold coordinate in the corresponding spatial dimension, update the current corresponding boundary threshold coordinate to this coordinate value.
[0046] The modeling space scale represents a measure of the spatial size adopted by the implicit network model during the modeling process. Different implicit network models can set different measures of the modeling space size. In some embodiments, the boundary threshold coordinate can be set to a single modeling space scale. In other embodiments, the boundary threshold coordinate can also be set to a multiple of the modeling space scale.
[0047] Exemplarily, to create a 3x2 boundary threshold coordinate array , , , , , , the boundary threshold coordinate array is used to record the range interval where the true coordinates of the object points in the three-dimensional space are located in each dimension. When performing implicit modeling based on the SDF of the unit circle, the boundary threshold coordinate array is initialized according to the implicit modeling space scale of the implicit network model, and we get: = = =100, = = =-100.
[0048] It can be understood that multiple three-dimensional space directions include the X three-dimensional space direction, the Y three-dimensional space direction, and the Z three-dimensional space direction. The boundary threshold coordinates include the X boundary threshold coordinate, the Y boundary threshold coordinate, and the Z boundary threshold coordinate. Comparing each sample three-dimensional coordinate with each boundary threshold coordinate respectively and updating the boundary threshold coordinate according to the comparison result includes: Compare the X coordinates of each sample three-dimensional coordinate with the X boundary threshold coordinate in the X three-dimensional space direction to obtain a first comparison result, and determine the X boundary threshold coordinate according to the first comparison result; Compare the Y coordinates of the three-dimensional coordinates of each sample with the Y boundary threshold coordinates in the Y three-dimensional space direction to obtain a second comparison result, and determine the Y boundary threshold coordinates according to the second comparison result; Compare the Z coordinates of the three-dimensional coordinates of each sample with the Z boundary threshold coordinates in the Z three-dimensional space direction to obtain a third comparison result, and determine the Z boundary threshold coordinates according to the third comparison result.
[0049] By respectively comparing the X coordinates, Y coordinates, and Z coordinates of the three-dimensional coordinates of each sample with the X boundary threshold coordinates, Y boundary threshold coordinates, and Z boundary threshold coordinates, and respectively updating the X boundary threshold coordinates, Y boundary threshold coordinates, and Z boundary threshold coordinates, the accuracy of the observation area range array can be further ensured, thereby improving the reliability of modeling.
[0050] It can be understood that the method further includes: When the expected three-dimensional coordinates of the pixel points in the candidate image are within the observation area range array, obtain the rendering color of the original pixel points corresponding to the expected three-dimensional coordinates; Determine the target pixel point rendering color of the corresponding pixel points according to the original pixel point rendering color.
[0051] The target pixel point rendering color can be determined in real time based on the original pixel point rendering color through a ray tracing algorithm, so as to further ensure that the rendering result conforms to the actual effect. The embodiments of the present application do not specifically limit how to determine the target pixel point rendering color. In practical applications, the position point where the ray intersects the line of sight can be determined first through the ray tracing method, and the color value of this point can be calculated using the position, normal, and lighting conditions of the intersection point. Usually, the application of the local lighting model and the contribution of the surrounding environment in the specular reflection and transmission directions to the light brightness at the intersection point can be referred to to determine the specific color value. For this, the embodiments of the present application do not elaborate too much.
[0052] Therefore, in some embodiments, the expected three-dimensional coordinates of each pixel point can be determined through a ray tracing algorithm, and by judging whether the expected three-dimensional coordinates are within the observation area range array, the corresponding rendering color is selected from the observation area range array for rendering to ensure the accuracy of the target image and improve the accuracy of the implicit network model modeling.
[0053] The target pixel point rendering color and foreign object elimination can be processed serially or asynchronously. For this, the embodiments of the present application do not limit whether the processing flow of abnormal pixel points and normal pixel points is set to be synchronous or asynchronous.
[0054] Exemplarily, taking the implicit network model as IDR, and IDR uses the SDF based on the unit sphere to express the geometric structure of the object. Taking the SDF composed of an MLP network as an example, then refer to Figure 3 and Figure 4Describe the image processing method according to the embodiments of the present application as follows: Step 1: Initialize the boundary threshold coordinate array: After the initial implicit network model training is completed, create a 3x2 boundary threshold coordinate array , , , , , , and the boundary threshold coordinate array is used to record the range intervals where the true coordinates of the object points in three-dimensional space are located in each dimension. Initialize the boundary threshold coordinate array according to the implicit modeling space scale of the implicit network model. For example, in the SDF implicit modeling based on the unit circle, it can be initialized as: = = =100, = = =-100.
[0055] Step 2: Determine the observation area range array according to the boundary threshold coordinate array: Traverse the entire training dataset of the implicit network model to obtain the 3D coordinates corresponding to each pixel point in the training dataset. Among them, the specific algorithm for the pixel to obtain the corresponding 3D point coordinates varies depending on the underlying implicit network model. Exemplarily, for example, the 3D coordinates corresponding to whether each pixel can hit the training object in the training scene and the object points of the hit training object can be obtained through the ray tracing algorithm. Specifically, first find the intersection of the ray corresponding to each pixel point with the unit sphere to greatly reduce the computational complexity. Assume that if there is no intersection between the ray and the unit sphere, this pixel cannot hit the object point and should belong to the background pixel. In addition, if the ray intersects with the object point, its left and right intersection points are (A, B), that is, the object point coordinates corresponding to this pixel should be searched on the line segment AB. By inputting the coordinates of point A and point B into the SDF network, the corresponding numerical values are obtained. If the output value is 0, that is, the 3D object point corresponding to this pixel has been found. If the output value is not 0, then in (A + SDF(A), B - SDF(B)) that is ( ), continue the iterative search until the SDF output value is 0 indicating that the corresponding 3D point has been found, or until indicates that this pixel does not correspond to the object point and belongs to the background pixel point. At this time, the 3D coordinates corresponding to each pixel point in the training dataset can be obtained. Compare the 3D coordinates [x, y, z] corresponding to each pixel point with the boundary threshold coordinate array , , , , , Compare the boundary values of each recorded spatial dimension, update the values in the boundary threshold coordinate array through comparison, and finally obtain the trained three-dimensional space range: , , , , , , that is, the observation area range array is , , , , , . Among them, is the minimum value in the X direction, is the maximum value in the X direction. Among them, is the minimum value in the Y direction, is the maximum value in the Y direction. Among them, is the minimum value in the Z direction, is the maximum value in the Z direction.
[0056] Step 3: Abnormal pixel point recognition, specifically: When generating a new view through the implicit network model, for each pixel point in the new view rendered image, it is possible to calculate whether each pixel in the regenerated candidate image can hit the object points in the trained scene through the ray tracing algorithm, where the three-dimensional coordinates of each pixel point in the new view rendered image can be determined by referring to the acquisition method of the 3D coordinates of each pixel point in the training dataset in Step 2.
[0057] Step 4: Rendering and coloring processing: If the pixel point in the view rendered image can hit the object point, that is, when it is determined that the 3D point coordinates of this pixel point are within the closed interval of , , , , , , it is considered that this object point belongs to the rendered scene. If this pixel does not hit the object point, that is, when the 3D point coordinates of this pixel point are not within the closed interval of , , , , , , it is considered that this point belongs to the blind area abnormal pixel point, and the color of this pixel is assigned according to the background color.
[0058] Step 5: Process each pixel in the image to be rendered according to Steps 3-4. Eventually, the abnormal pixel points in the observation blind area can be eliminated, and high-quality new perspective rendering can be achieved.
[0059] At this time, through the above Steps 1-5, an image after foreign object elimination can be obtained as Figure 4 shown. At this time, as Figure 4 shown, since the abnormal pixel points are all rendered with the background color, the surface of the target object can be normally displayed without occlusion.
[0060] It can be understood that, as shown in Figure 5 , an embodiment of the present application provides a modeling device for a 3D model. The device includes: An acquisition module 100, configured to acquire a pre-trained implicit network model corresponding to a target object, and a training image data set used by the implicit network model during pre-training. The training images in the training image data set are images including at least one rendering perspective of the target object; A determination module 200, configured to determine an observation area range array according to the implicit network model and the training image data set; A modeling module 300, configured to generate a candidate image of the target object according to the implicit network model and the expected rendering perspective; A rendering module 400, configured to process the pixel points outside the observation area range array in the candidate image as abnormal pixel points to obtain a target image.
[0061] Therefore, by pre-acquiring the observation area range array of the implicit network model, the abnormal pixel points randomly generated in the candidate image due to the unsupervised learning of the implicit network model during training can be eliminated based on the observation area range array. At this time, each pixel point of the target image is a real existing point. When the new perspective rendering display of the target image is completed, the display result is closer to the actual 3D object. Therefore, the embodiment of the present application can improve the reliability of image modeling.
[0062] It can be understood that, in some embodiments, the rendering module 400 is further configured to, when the expected 3D coordinates of the pixel points in the candidate image are within the observation area range array, acquire the rendering color of the original pixel points corresponding to the expected 3D coordinates; and determine the rendering color of the target pixel points corresponding to the pixel points according to the rendering color of the original pixel points, so as to realize the rendering coloring of the normal pixel points.
[0063] In some embodiments, the modeling device is further provided with a view module, and the view module is used to provide an interactive configuration interface to configure parameters such as the implicit network model and the training image data set.
[0064] In some embodiments, the embodiments of the present application also implement the elimination process of abnormal pixel points through the rendering module 400 to reduce the external perception of abnormal pixel points. Among them, the rendering module 400 is further configured to: Traverse each pixel point on the candidate image; Determine the expected three-dimensional coordinates of each traversed pixel point through a preset ray tracing algorithm; When the expected three-dimensional coordinates are not in the observation area range array, the traversed pixel points are regarded as abnormal pixel points; Perform foreign object elimination processing on each abnormal pixel point; When all the abnormal pixel points in the candidate image have completed the foreign object elimination processing, a target image is obtained.
[0065] In some embodiments, the rendering module 400 further includes an abnormal pixel point rendering sub-module, which implements differential rendering processing of abnormal pixel points through the abnormal pixel point rendering sub-module. Among them, the abnormal pixel point rendering sub-module is specifically configured to: Obtain a pre-configured foreign object rendering color; Render the color of each abnormal pixel point as the foreign object rendering color.
[0066] In some embodiments, the determination module 200 is configured to obtain the observation area range number; among them, the determination module 200 is specifically configured to: Determine the sample three-dimensional coordinates of each sample pixel point in the training image dataset; Determine the boundary coordinates in a preset plurality of three-dimensional space directions according to each sample three-dimensional coordinate; Obtain the observation area range array according to each boundary coordinate.
[0067] In some embodiments, the determination module 200 includes a boundary coordinate determination sub-module, which determines the boundaries of the observation intervals corresponding to the observation area range array through the boundary coordinate determination sub-module. Among them, the boundary coordinate determination sub-module is specifically configured to: Initialize a boundary threshold coordinate array according to the modeling space scale of the implicit network model. The boundary threshold coordinate array records the boundary threshold coordinates in each three-dimensional space direction; Compare each sample three-dimensional coordinate with each boundary threshold coordinate respectively and update the boundary threshold coordinate according to the comparison result; Use the updated boundary threshold coordinate as the boundary coordinate in the corresponding three-dimensional space direction.
[0068] In some embodiments, the plurality of three-dimensional space directions include the X three-dimensional space direction, the Y three-dimensional space direction, and the Z three-dimensional space direction. The boundary threshold coordinates include the X boundary threshold coordinate, the Y boundary threshold coordinate, and the Z boundary threshold coordinate; the boundary coordinate determination sub-module is specifically configured to: Compare the X coordinates of the three-dimensional coordinates of each sample with the X boundary threshold coordinates in the X three-dimensional space direction, and determine the X boundary threshold coordinates according to the comparison result of the coordinates in the X three-dimensional space direction; Compare the Y coordinates of the three-dimensional coordinates of each sample with the Y boundary threshold coordinates in the Y three-dimensional space direction, and determine the Y boundary threshold coordinates according to the comparison result of the coordinates in the Y three-dimensional space direction; Compare the Z coordinates of the three-dimensional coordinates of each sample with the Z boundary threshold coordinates in the Z three-dimensional space direction, and determine the Z boundary threshold coordinates according to the comparison result of the coordinates in the Z three-dimensional space direction; Determine the updated boundary threshold coordinates according to the determined X boundary threshold coordinates, Y boundary threshold coordinates, and Z boundary threshold coordinates.
[0069] The embodiment of the present application does not limit the number of configurable implicit network models in the modeling device. For example, in some embodiments, the modeling device is further provided with a management module to generate different projects for different target objects. At this time, different new perspectives of different target objects can be rendered based on different projects.
[0070] As can be seen from the above, the embodiment of the present application can pre-obtain the observation area range array of the implicit network model, so that the abnormal pixel points randomly generated in the candidate image due to the unsupervised learning of the implicit network model during the training process can be processed. At this time, each pixel point in the target image obtained by rendering and displaying a new perspective based on the expected rendering perspective is an actual object point, and the display result is closer to the actual three-dimensional object. Therefore, compared with the related art, the embodiment of the present application can improve the reliability of image modeling.
[0071] The embodiment of the present application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the above image processing method is implemented. The electronic device can be any intelligent terminal including a tablet computer, an in-vehicle computer, etc.
[0072] Please refer to Figure 6 , Figure 6 schematically shows the hardware structure of an electronic device according to another embodiment. The electronic device includes: A processor 501, which can be implemented in a general-purpose CPU (Central Processing Unit, central processor), a microprocessor, an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present application; A memory 502, which can be a NAND flash. The relevant program code is stored in the memory 502 and is called by the processor 501 to execute the image processing method of the embodiments of the present application; An input / output interface 503 for implementing information input and output; A communication interface 504 for implementing communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or through wireless means (such as mobile network, WIFI, Bluetooth, etc.); A bus 505 for transmitting information between various components of the device (such as the processor 501, the memory 502, the input / output interface 503, and the communication interface 504); Among them, the processor 501, the memory 502, the input / output interface 503, and the communication interface 504 are communicatively connected to each other inside the device through the bus 505.
[0073] It can be understood that the embodiments of the present application also provide a computer-readable storage medium. The computer-readable storage medium is a computer-readable storage medium. This storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned image processing method is implemented.
[0074] Therefore, in the above embodiments of the present application, by pre-obtaining the observation area range array of the implicit network model, the abnormal pixel points randomly generated in the candidate image due to the unsupervised learning of the implicit network model during the training process can be processed based on the observation area range array. At this time, each pixel point in the target image obtained by rendering and displaying a new perspective based on the expected rendering perspective is an actual object point, and the display result is closer to the actual three-dimensional object, thereby improving the reliability of the image modeling of the target object.
[0075] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include high-speed random access memory, and can also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory can optionally include a memory remotely set relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above networks include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0076] The embodiments described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art can know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.
[0077] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than those shown in the figures, or combine certain steps, or different steps.
[0078] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0079] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and their appropriate combinations.
[0080] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0081] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. "At least one (item) of the following" or its similar expression refers to any combination of these items, including any combination of single item (item) or plural items (items). For example, at least one (item) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0082] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical or other forms.
[0083] The units described above as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0084] In addition, in each embodiment of this application, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0085] When an integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of various embodiments of this application. The aforementioned storage medium includes: various media that can store programs, such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs.
[0086] The preferred embodiments of the embodiments of this application have been described above with reference to the accompanying drawings. This does not limit the scope of the rights of the embodiments of this application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of this application shall be within the scope of the rights of the embodiments of this application.
Claims
1. An image processing method, characterized in that: The method comprises: Acquire a pre-trained implicit network model corresponding to the target object, and a training image data set used by the implicit network model in a pre-training process, wherein the training images in the training image data set are images containing at least one rendering perspective of the target object; Determine an observation area range array according to the implicit network model and the training image data set; Generating a candidate image of the target object according to the implicit network model and the expected rendering perspective; The pixel points in the candidate image that are outside the observation area range array are processed as abnormal pixel points to obtain a target image.
2. The image processing method according to claim 1, characterized in that: The step of processing the pixel points in the candidate image that are outside the observation area range array as abnormal pixel points to obtain the target image includes: Traversing each pixel on the candidate image; Determine the expected three-dimensional coordinates of each traversed pixel point by a preset ray tracing algorithm; When the expected three-dimensional coordinates are not in the observation area range array, the traversed pixel point is regarded as an abnormal pixel point; Performing foreign matter elimination processing on each of the abnormal pixels; When all abnormal pixels in the candidate image have completed the foreign matter elimination process, the target image is obtained.
3. The image processing method according to claim 2, characterized in that: The performing foreign matter elimination processing on each of the abnormal pixels includes: Get the preconfigured foreign object rendering color; The color of each of the abnormal pixels is rendered as the foreign object rendering color.
4. The image processing method according to claim 1, characterized in that: Determining an observation area range array according to the implicit network model and the training image data set of the target object includes: Determine the sample three-dimensional coordinates of each sample pixel point in the training image data set; Determining boundary coordinates in a plurality of preset three-dimensional spatial directions according to the three-dimensional coordinates of each sample; According to each of the boundary coordinates, the observation area range array is obtained.
5. The image processing method according to claim 4, characterized in that: The step of determining the boundary coordinates in a plurality of preset three-dimensional spatial directions according to the three-dimensional coordinates of each sample includes: Initializing a boundary threshold coordinate array according to the modeling space scale of the implicit network model, wherein the boundary threshold coordinate array records the boundary threshold coordinates in each three-dimensional space direction; Comparing each of the sample three-dimensional coordinates with each of the boundary threshold coordinates and updating the boundary threshold coordinates according to the comparison results; The updated boundary threshold coordinates are used as the boundary coordinates in the corresponding three-dimensional space direction.
6. The image processing method according to claim 5, characterized in that: The multiple three-dimensional spatial directions include an X three-dimensional spatial direction, a Y three-dimensional spatial direction, and a Z three-dimensional spatial direction, the boundary threshold coordinates include an X boundary threshold coordinate, a Y boundary threshold coordinate, and a Z boundary threshold coordinate, and the comparing each of the sample three-dimensional coordinates with each of the boundary threshold coordinates and updating the boundary threshold coordinates according to the comparison results, comprises: Comparing the X coordinates of each of the sample three-dimensional coordinates with the X boundary threshold coordinates in the X three-dimensional space direction to obtain a first comparison result, and determining the X boundary threshold coordinates according to the first comparison result; Comparing the Y coordinates of the three-dimensional coordinates of each sample with the Y boundary threshold coordinates in the Y three-dimensional space direction to obtain a second comparison result, and determining the Y boundary threshold coordinates according to the second comparison result; The Z coordinates of the three-dimensional coordinates of each sample are compared with the Z boundary threshold coordinates in the Z three-dimensional space direction to obtain a third comparison result, and the Z boundary threshold coordinates are determined according to the third comparison result.
7. The image processing method according to claim 1, characterized in that: The method further comprises: When the expected three-dimensional coordinates of the pixel point in the candidate image are in the observation area range array, obtaining the original pixel point rendering color corresponding to the expected three-dimensional coordinates; According to the original pixel rendering color, a target pixel rendering color of the corresponding pixel is determined.
8. A three-dimensional modeling device, characterized in that: The device comprises: An acquisition module, used to acquire a pre-trained implicit network model corresponding to a target object, and a training image data set used by the implicit network model in a pre-training process, wherein the training images in the training image data set are images containing at least one rendering perspective of the target object; A determination module, used to determine an observation area range array according to the implicit network model and the training image data set; A modeling module, configured to generate a candidate image of the target object according to the implicit network model and the expected rendering perspective; The rendering module is used to process the pixel points in the candidate image that are outside the observation area range array as abnormal pixel points to obtain a target image.
9. An electronic device, characterized in that: The electronic device comprises a memory and a processor, the memory stores a computer program, and the processor implements the image processing method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the image processing method according to any one of claims 1 to 7 is implemented.