Object point cloud reconstruction method based on single image and related device

By using the internal parameter matrix and external parameter matrix as constraints in a single image, the problem of inaccurate reconstruction of three-dimensional point clouds under single-view image is solved, and a more accurate full-view three-dimensional point cloud is generated, which improves the accuracy and consistency of image processing.

CN120495528APending Publication Date: 2025-08-15CRRC TECH INNOVATION (BEIJING) CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510630220.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

In application scenarios where historical data of only single-view images are limited in movement or cost-restricted, the prior art cannot effectively reconstruct the three-dimensional point cloud information of the object, resulting in the generated three-dimensional point cloud being inaccurate or missing.

Method used

By obtaining the internal parameter matrix of a single image and the external parameter matrix of multiple shooting angles, combined with the denoising processing of the diffusion model, the full-view three-dimensional point cloud information of the target object is generated. The specific steps include determining the three-dimensional coordinate set of pixel points, simulating multiple shooting perspectives, using the internal parameter matrix and the external parameter matrix as constraints for denoising, generating image data for each viewing angle, and stitching to form a full-view three-dimensional point cloud.

Benefits of technology

The accuracy and geometric consistency of the three-dimensional point cloud information generated by a single image is improved, the spatial position consistency between multi-view images is ensured, object deformation is reduced, and the generated three-dimensional point clouds are more accurate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120495528A_ABST
    Figure CN120495528A_ABST
Patent Text Reader

Abstract

The invention discloses an object point cloud reconstruction method based on a single image and a related device, and relates to the field of image processing. The method comprises the following steps: determining external parameter matrixes corresponding to a plurality of shooting view angles corresponding to a target object in a single image; determining the depth information of each shooting view angle based on the initial depth information of the single image and the external parameter matrix corresponding to each shooting view angle; calling a diffusion model, and performing denoising processing on a preset noise image in the diffusion model by taking the internal reference matrix corresponding to the single image, the external reference matrix corresponding to each shooting view angle and the depth information as constraint conditions of denoising processing to obtain image data of the target object at each shooting view angle; and generating full-view three-dimensional point cloud information of the target object according to the image data of the target object at each shooting view. According to the method and the device, other view images indeed by the target object in the single image are complemented, and the geometric consistency and the space accuracy among the multi-view images are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to a method and related device for reconstructing an object point cloud based on a single image. Background Art

[0002] In scenarios such as autonomous driving, robotic perception, virtual reality (VR), and the digitization of cultural relics, methods for reconstructing high-precision three-dimensional point clouds from multi-view 2D images of an object are crucial. However, in scenarios where only single-view historical data is available, applications with limited mobility, or industrial applications with cost constraints, these methods are unable to reconstruct the object's 3D point cloud information. Therefore, how to restore an object's 3D point cloud from a single image with a single viewpoint has become a pressing technical challenge.

[0003] To address this problem, existing technologies estimate the 3D shape of objects in a single image based on visual features such as texture, lighting, edges, and perspective distortion, generating corresponding 3D point cloud information. However, due to issues such as line of sight obstruction and missing geometric information, the visual features that can be extracted from a single image are limited, resulting in inaccurate or missing 3D point cloud information.

[0004] Therefore, there is an urgent need for an object point cloud reconstruction method based on a single image to improve the accuracy of three-dimensional point cloud information generated using a single image. Summary of the Invention

[0005] In view of the above problems, this application provides a method and related device for reconstructing object point clouds based on a single image, so as to improve the accuracy of three-dimensional point cloud information generated using a single image. The specific solution is as follows:

[0006] The first aspect of the present application provides a method for reconstructing an object point cloud based on a single image, comprising:

[0007] Acquire a single image containing a target object and an intrinsic parameter matrix corresponding to the single image, wherein the intrinsic parameter matrix is used to represent imaging parameters obtained by a camera photographing the target object to obtain the single image;

[0008] Determine the pixel points of the target object, and obtain a three-dimensional coordinate set obtained by projecting the pixel points of the target object;

[0009] Determining, based on preset shooting rules, multiple shooting angles corresponding to the three-dimensional coordinate set and an extrinsic parameter matrix corresponding to each shooting angle, wherein the extrinsic parameter matrix is used to represent state parameters of the camera when shooting the target object at the corresponding shooting angle, the state parameters including camera position and camera attitude;

[0010] Determining the depth information corresponding to each shooting perspective based on the initial depth information of the single image and the extrinsic parameter matrix corresponding to each shooting perspective;

[0011] Invoking a diffusion model, using the intrinsic parameter matrix, the extrinsic parameter matrix corresponding to each shooting angle, and the depth information as constraints for denoising, performing the denoising process on a preset noise image in the diffusion model to obtain image data of the target object at each shooting angle;

[0012] Full-view three-dimensional point cloud information of the target object is generated based on the image data of the target object at each shooting perspective.

[0013] In a possible implementation, determining, based on a preset shooting rule, multiple shooting perspectives corresponding to the three-dimensional coordinate set and an extrinsic parameter matrix corresponding to each shooting perspective includes:

[0014] Taking the initial shooting angle of view of the single image captured by the camera in the three-dimensional coordinate set as a starting point, a plurality of shooting angles are determined according to a preset angle interval;

[0015] Determining a camera shooting height corresponding to the initial shooting angle of view based on a depth value of an image center point in the initial depth information of the single image;

[0016] determining a shooting distance of the camera when shooting the target object according to the object size of the target object, wherein the object size is determined based on the three-dimensional coordinate set;

[0017] Determining the three-dimensional coordinates of the camera corresponding to each shooting angle based on the shooting distance and the camera shooting height;

[0018] Determining a camera rotation matrix corresponding to each shooting perspective based on differences between the three-dimensional coordinates of the camera at each shooting perspective and the three-dimensional coordinates of the object center pixel point of the target object, wherein the camera rotation matrix is used to represent a rotational posture of the camera at the corresponding shooting perspective relative to the initial shooting perspective;

[0019] Based on the camera three-dimensional coordinates and the camera rotation matrix corresponding to each shooting perspective, an extrinsic parameter matrix for each shooting perspective is generated.

[0020] In a possible implementation, determining the depth information corresponding to each shooting perspective based on the initial depth information of the single image and the extrinsic parameter matrix corresponding to each shooting perspective includes:

[0021] Converting the pixels of the target object into a three-dimensional point cloud based on the initial depth information of the single image to obtain three-dimensional point cloud information of the initial shooting perspective;

[0022] Determining the three-dimensional point cloud information of the target object at each shooting perspective based on the three-dimensional point cloud information of the initial shooting perspective and the extrinsic parameter matrix corresponding to each shooting perspective;

[0023] The three-dimensional point cloud information of each shooting perspective is projected onto a two-dimensional plane to obtain depth information corresponding to each shooting perspective.

[0024] In a possible implementation, generating full-view three-dimensional point cloud information of the target object based on the image data of the target object at each shooting perspective includes:

[0025] Performing depth estimation on the image data of each shooting angle to obtain target depth information corresponding to each shooting angle, the target depth information including: a depth value of each pixel in the image data;

[0026] Generating three-dimensional point cloud information corresponding to each image data according to the depth value of each pixel in the image data;

[0027] The three-dimensional point cloud information corresponding to each two adjacent shooting perspectives is spliced to obtain the full-perspective three-dimensional point cloud information of the target object.

[0028] In a possible implementation, after generating three-dimensional point cloud information corresponding to each image data according to the depth value of each pixel in the image data, the method further includes:

[0029] Filter out three-dimensional point cloud information with a reference distance value greater than a preset threshold, wherein the reference distance value is: an average value of distances between any three-dimensional point cloud coordinate and all other three-dimensional point cloud coordinates in the three-dimensional point cloud information.

[0030] In a possible implementation, the method further includes:

[0031] According to the three-dimensional point cloud information of the initial shooting angle of view, the full-view three-dimensional point cloud information of the target object is moved or rotated to obtain the target full-view three-dimensional point cloud information of the target object with the initial shooting angle of view as the front.

[0032] In a possible implementation, the method further includes:

[0033] The target full-view three-dimensional point cloud information of the target object is embedded into the three-dimensional point cloud information of the scene in the single image to obtain the three-dimensional point cloud information of the single image.

[0034] In a possible implementation, determining the pixel points of the target object and obtaining a three-dimensional coordinate set obtained by projecting the pixel points of the target object includes:

[0035] performing image segmentation processing on the single image to obtain a segmentation mask of the single image;

[0036] Extracting pixels of the target object from the segmentation mask;

[0037] Based on the depth value corresponding to each pixel in the initial depth information of the single image and the intrinsic parameter matrix, each pixel is projected into a three-dimensional space to obtain a three-dimensional coordinate set corresponding to the target object.

[0038] A second aspect of the present application provides a device for reconstructing an object point cloud based on a single image, comprising:

[0039] An image acquisition unit, configured to acquire a single image containing a target object and an intrinsic parameter matrix corresponding to the single image, wherein the intrinsic parameter matrix is used to represent imaging parameters obtained by a camera photographing the target object to obtain the single image;

[0040] A three-dimensional coordinate acquisition unit, configured to determine the pixel points of the target object and acquire a three-dimensional coordinate set obtained by projecting the pixel points of the target object;

[0041] an extrinsic parameter matrix determination unit, configured to determine, based on preset shooting rules, a plurality of shooting angles corresponding to the three-dimensional coordinate set, and an extrinsic parameter matrix corresponding to each shooting angle, wherein the extrinsic parameter matrix is used to represent state parameters of the camera when shooting the target object at the corresponding shooting angle, the state parameters including camera position and camera attitude;

[0042] a depth information determining unit, configured to determine the depth information corresponding to each of the shooting perspectives based on the initial depth information of the single image and the extrinsic parameter matrix corresponding to each of the shooting perspectives;

[0043] an image diffusion unit, configured to call a diffusion model, use the intrinsic parameter matrix, the extrinsic parameter matrix corresponding to each of the shooting angles, and the depth information as constraints for denoising, perform the denoising process on a preset noise image in the diffusion model, and obtain image data of the target object at each of the shooting angles;

[0044] The full-view point cloud generating unit is used to generate full-view three-dimensional point cloud information of the target object based on the image data of the target object at each shooting perspective.

[0045] The third aspect of the present application provides a computer program product, including computer-readable instructions. When the computer-readable instructions are executed on an electronic device, the electronic device implements the object point cloud reconstruction method based on a single image according to the first aspect or any implementation of the first aspect.

[0046] By means of the above technical solution, the present application provides a method for reconstructing an object point cloud based on a single image. The method uses the intrinsic parameter matrix of the camera used to shoot the single image, as well as the extrinsic parameter matrix and depth information of each shooting angle determined based on the single image, as constraints of the diffusion model to perform inverse denoising on the initial noisy image and generate image data corresponding to each shooting angle of the target object. The camera position and posture represented by the extrinsic parameter matrix of each shooting angle provides perspective transformation information for the single image, ensuring that the multi-view images generated after denoising remain consistent in three-dimensional spatial position and avoid spatial misalignment. The depth information of each shooting angle provides geometric constraints for the denoising process, making the generated image of each shooting angle more accurate in three-dimensional structure and reducing the deformation of objects in images with different perspectives. The intrinsic parameter matrix can more accurately represent the imaging principle of the camera used to shoot the single image, constrain the imaging of other perspectives, and thus improve the perspective consistency between multi-view images.

[0047] Based on this, this application adjusts the constraints of the diffusion model's reverse denoising to, on the one hand, complete the other perspective images of the target object in a single image. On the other hand, the generated multi-perspective images not only meet the physical imaging laws, but also improve the geometric consistency and spatial accuracy between the multi-perspective images. Based on this, the multi-perspective images of the target object with higher comprehensive accuracy generate more accurate full-perspective 3D point cloud information of the target object. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] The above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that the originals and elements are not necessarily drawn to scale.

[0049] Figure 1 A flowchart of a method for reconstructing an object point cloud based on a single image provided in an embodiment of the present application;

[0050] Figure 2 An example diagram of a shooting rule for simulating a camera shooting a target object from multiple perspectives provided in an embodiment of the present application;

[0051] Figure 3 A schematic diagram of the structure of an object point cloud reconstruction device based on a single image provided in an embodiment of the present application. DETAILED DESCRIPTION

[0052] The following describes the embodiments of the present application in conjunction with the accompanying drawings. The terms used in the implementation methods of the present application are only used to explain the specific embodiments of the present application and are not intended to limit the present application.

[0053] The embodiments of the present application are described below in conjunction with the accompanying drawings. Those skilled in the art will appreciate that, with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.

[0054] The terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequential order. It should be understood that the terms used in this way can be interchangeable under appropriate circumstances, and this is merely a way of distinguishing the objects of the same attributes when describing them in the embodiments of the present application. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, so that the process, method, system, product or equipment comprising a series of units need not be limited to those units, but may include other units that are not clearly listed or inherent to these processes, methods, products or equipment.

[0055] This application can be applied in the field of image processing. Taking image processing applied in the field of autonomous driving as an example, the following introduces multiple application scenarios that have been implemented in products.

[0056] First, let's introduce an optional application scenario of this application. In the field of autonomous driving, laser radar is usually used to collect and analyze environmental information around the vehicle to provide a reliable three-dimensional environmental model for autonomous driving positioning, path planning, and obstacle detection. However, if the laser radar fails or the viewing angle is limited, such as in rainy or snowy weather, the vehicle's camera can be used to capture images, extract geometric information from the images, and generate a three-dimensional point cloud. However, the vehicle's camera has a fixed shooting angle, and the images of the surrounding environment captured contain limited information. The constructed three-dimensional point cloud is not accurate enough, which may affect driving safety.

[0057] In order to solve the above problems, an embodiment of the present application provides a method for reconstructing an object point cloud based on a single image. The following describes in detail a method for reconstructing an object point cloud based on a single image in accordance with an embodiment of the present application in conjunction with the accompanying drawings.

[0058] Reference Figure 1 , Figure 1 The present invention provides a method for reconstructing a point cloud of an object based on a single image. Figure 1As shown, an embodiment of the present application provides a method for reconstructing an object point cloud based on a single image, which may include steps S110 to S160. These steps are described in detail below.

[0059] First of all, this application can be applied to, but is not limited to, applications with image processing and data processing functions or cloud services provided by cloud-side servers. Optionally, in the field of autonomous driving, this method can be applied to devices such as vehicle controllers or vehicle computers to perform the following processing on a single image captured by the vehicle's camera.

[0060] Step S110 , obtaining a single image containing the target object and an intrinsic parameter matrix corresponding to the single image.

[0061] It can be understood that the target object is the object of the three-dimensional point cloud construction, and a single image contains at least one target object. When a single image contains multiple target objects, the following steps S120-S160 are executed for each target object to construct a three-dimensional point cloud for each target object.

[0062] Among them, a single image can be an image obtained by a user using a camera to capture a target object. In the field of autonomous driving, a single image can be an image captured by a vehicle's monocular camera. Since the camera usually captures driving backgrounds such as bushes and lane lines during vehicle driving, the image does not contain target objects that affect driving safety, such as roadblocks, pedestrians, and vehicles in front. Therefore, whether the captured image contains a target object can be used as a trigger condition. If the captured image contains a target object, the following steps S120-S160 are triggered by the captured single image containing the target object to reconstruct point clouds of target objects such as roadblocks and pedestrians in the image.

[0063] While acquiring a single image containing a target object, the intrinsic parameter matrix corresponding to the single image is acquired. The intrinsic parameter matrix is used to represent the imaging parameters of the single image obtained by the camera shooting the target object. The intrinsic parameter matrix of the camera can reflect the geometric characteristics and imaging laws inside the camera. As long as the same camera, the same lens, resolution and focal length are used, the intrinsic parameter matrix corresponding to the image is the same. If this method is applied to vehicle-assisted autonomous driving to realize point cloud construction, the single images acquired are all from the same camera configured in the vehicle, and the intrinsic parameter matrix of the single image processed each time is the same. When this method is configured with the vehicle-machine system, an intrinsic parameter matrix can be calibrated according to the imaging parameters of the vehicle's camera. If this method is applied to a cloud server to perform point cloud reconstruction on the target image in single images with different data sources, the intrinsic parameter matrix of the single image processed each time is different, and the intrinsic parameter matrix of the camera corresponding to the single image needs to be recalibrated each time.

[0064] Specifically, the intrinsic parameter matrix of the camera that captures a single image of the target object can be calibrated using calibration methods such as optical calibration, deep learning-based calibration, and stereo target-based calibration. The following calibration process is used as an example:

[0065] First, determine the camera used to capture the target object. In this example, an Oak-D-Pro camera with a resolution of 640 x 480 pixels is used. This camera is used to capture the calibration pattern, maintaining a variety of angles, uniform lighting, sharp focus, and a high level of moisture. This helps improve the accuracy of the intrinsic parameter matrix determined based on the captured calibration pattern. The calibration pattern consists of regular geometric patterns, such as a checkerboard pattern.

[0066] Furthermore, the cv2.findChessboardCorners function is used to detect the corners of the checkerboard in the calibration pattern and determine the corner information of the calibration pattern. The corner information includes at least information indicating whether the corner is successfully detected and the pixel coordinates of the detected corner. Optionally, sub-pixel optimization can be further performed using cv2.cornerSubPix to improve the accuracy of corner positioning. Based on the corner information of the calibration pattern, the cv2.calibrateCamera function is called to calculate the intrinsic parameter matrix of the camera, where the intrinsic parameter matrix is a 3*3 matrix, as shown in the following formula (1):

[0067] (1)

[0068] Among them, f x 、f y are the focal lengths on the x-axis and y-axis, respectively, in pixels, indicating the focal length in pixels on the imaging plane, reflecting the camera's field of view and the image's scaling. x 、c y are the x- and y-coordinates of the principal point, in pixels, representing the intersection of Guangzhou on the image plane, which should ideally be located at the center of the image.

[0069] Step S120 , determining the pixel points of the target object, and obtaining a three-dimensional coordinate set obtained by projecting the pixel points of the target object.

[0070] In step S130, based on preset shooting rules, multiple shooting angles corresponding to the three-dimensional coordinate set and the extrinsic parameter matrix corresponding to each shooting angle are determined. The extrinsic parameter matrix is used to represent the state parameters of the camera when shooting the target object at its corresponding shooting angle. The state parameters include the camera position and camera posture.

[0071] Steps S120-S130 can be understood as projecting the target object at the initial shooting angle in a single image into three-dimensional space to obtain an inaccurate three-dimensional state of the target object. Taking the initial shooting angle as the starting point, the camera state when the camera shoots the target object at different shooting angles is simulated, so as to facilitate the subsequent inference of the image data of the target object at the corresponding shooting angle from the camera state.

[0072] In step S120, pixel information of the target object can be extracted from the single image by distinguishing color or texture, or by using image segmentation techniques such as semantic segmentation or instance segmentation. The pixel information may include pixel coordinates and color. Furthermore, based on the pixel coordinates of the target object, using a deep learning model and visual prior knowledge, the pixel coordinates are projected into three-dimensional space to obtain a set of three-dimensional coordinates of the target object.

[0073] In one possible implementation, the pixels of the target object are determined, and a three-dimensional coordinate set obtained by projecting the pixel points of the target object is obtained, including: performing image segmentation processing on a single image to obtain a segmentation mask of the single image; extracting the pixels of the target object from the segmentation mask; and projecting each pixel into three-dimensional space based on the depth value corresponding to each pixel in the initial depth information of the single image and an intrinsic parameter matrix to obtain a three-dimensional coordinate set corresponding to the target object.

[0074] The process of obtaining a set of three-dimensional coordinates is described below. Optionally, first, a single RGB image I is preprocessed, such as normalization and image resizing. Further, an existing deep learning instance segmentation model is used to segment the single image I and obtain the segmentation mask M of the target object output by the model. i , where the segmentation mask M i It is a binary matrix of the same size as the single image I. In the matrix, the pixel (u, v) with a value of 1 belongs to the object i, and the pixel with a value of 0 belongs to the background. By multiplying the value of each pixel in the binary matrix with the single image I, that is, , extract the pixel area I corresponding to the target object in a single image i , and pixel set P.

[0075] Furthermore, the monocular depth estimation model is used to identify the depth of a single image I and obtain the initial depth information D. Referring to the following formula (2), according to the depth information D and the camera's intrinsic parameter matrix K, each pixel point (u, v) in the pixel point set P is projected into the three-dimensional space, and the three-dimensional coordinates (x, y, z) corresponding to each pixel point are determined to obtain the three-dimensional coordinate set of the target object {(x j ,y j ,z j)}, where j represents the j-th pixel of the target object.

[0076] (2)

[0077] Based on this, in step S130, the obtained set of three-dimensional coordinates of the target object is equivalently treated as the target object in a three-dimensional state. Based on preset shooting rules, multiple shooting angles of the camera shooting the three-dimensional target object are simulated, as well as the camera state corresponding to each shooting angle. Preset shooting rules may include: angle transformation rules, such as the interval angle between two shooting angles; shooting paths, such as horizontal or vertical circling of the object; etc. Furthermore, using the three-dimensional coordinates of the target object, the camera's position, posture, and other state information at each shooting angle are determined, thereby determining the extrinsic parameter matrix corresponding to each shooting angle.

[0078] Specifically, based on the preset shooting rules, multiple shooting perspectives corresponding to the three-dimensional coordinate set and the external parameter matrix corresponding to each shooting perspective are determined, including: taking the initial shooting perspective of the camera shooting a single image in the three-dimensional coordinate set as the starting point, and determining multiple shooting perspectives based on the preset perspective interval; determining the camera shooting height corresponding to the initial shooting perspective based on the depth value of the image center point in the initial depth information of the single image; determining the shooting distance when the camera shoots the target object based on the object size of the target object, and the object size is determined based on the three-dimensional coordinate set; determining the camera three-dimensional coordinates corresponding to each shooting perspective based on the shooting distance and the camera shooting height; determining the camera rotation matrix corresponding to each shooting perspective based on the difference between the camera three-dimensional coordinates of each shooting perspective and the three-dimensional coordinates of the object center pixel point of the target object, the camera rotation matrix is used to characterize the line of sight rotation posture of the camera relative to the initial shooting perspective at its corresponding shooting perspective; generating the external parameter matrix of each shooting perspective based on the camera three-dimensional coordinates and the camera rotation matrix corresponding to each shooting perspective.

[0079] Reference Figure 2 The embodiment of the present application provides an example diagram of a shooting rule for simulating a camera shooting a target object from multiple perspectives, which exemplifies the process of determining the external parameter matrix corresponding to each shooting perspective. First, determine the motion path of the camera simulation shooting. Figure 2 In this example, the target object is photographed by circling the target object horizontally with the target object's center pixel point C as the center. Figure 2 In this shooting trajectory, the initial shooting angle P0 corresponding to a single image is used as the shooting starting point, and the preset angle intervals are used. , determine multiple new shooting angles P i .

[0080] Furthermore, the shooting position corresponding to each shooting angle Pi is determined. In three-dimensional space, the shooting position includes the position in the x, y, and z directions, where z is the height of the camera when shooting, and x and y are the horizontal position relationship between the camera and the target object. Specifically, the image center point (u center ,v center ), the corresponding depth value D(u center ,v center ), according to the reference formula (2), the center point of the image is projected into the three-dimensional space to obtain the height z of the center point of the image in the three-dimensional space init , as the camera height.

[0081] In the embodiment of the present application, x and y can be determined based on the distance between the camera and the target object. Specifically, the shooting distance r between the camera and the object center pixel point C of the target object is determined based on the size of the target object reflected by the three-dimensional coordinate set. The shooting distance is usually 1.5 to 2.2 times the diagonal length of the target object size to ensure that the camera shooting range can completely cover the target object.

[0082] Refer to the following formula (3) to determine the object center pixel C(x c ,y c ,z c ) is the center of the circle, the shooting distance r is the radius, and the camera height z init In the shooting path, the camera three-dimensional coordinate P corresponding to each shooting angle i .

[0083] (3)

[0084] in, , i represents the i-th shooting angle, i=1,2,3…,N-1, N is the total number of shooting angles. The geometric center of the target object is the object center pixel C(x c ,y c ,z c ), refer to the following formula (4), where n is the total number of pixels of the target object.

[0085] (4)

[0086] Each path point P of the shooting path obtained above i =(x i ,y i ,z i ) is used as the translation vector Ti to represent the translation of the camera's shooting position relative to the geometric center of the target object under different shooting angles. And, referring to the following formula (5), the line of sight direction vector D is calculated. i, which is used to represent the sight direction of the camera's shooting position relative to the geometric center of the object under different shooting angles.

[0087] (5)

[0088] Refer to the following formula (6) to normalize the above-mentioned sight direction.

[0089] …… (6)

[0090] According to the three coordinate axes of the predefined camera coordinate system, : Line of sight direction, take ; : horizontal axis, perpendicular to and the vector in the world coordinate system ; : The numerical direction of the camera, calculated by the cross product: Based on this, determine the rotation matrix of the camera at each shooting angle The obtained rotation matrix R i and the translation matrix T i Combination, get the external parameter matrix [R i |T i ], referring to the following formula (7), the external parameter matrix describes the posture and position of the camera at each path point corresponding to each shooting perspective in the shooting path.

[0091] (7)

[0092] Step S140 : Determine the depth information corresponding to each shooting perspective based on the initial depth information of the single image and the extrinsic parameter matrix corresponding to each shooting perspective.

[0093] This step can be understood as simulating the depth information of the target object obtained by shooting at different shooting angles based on the depth value of each pixel of the target object in the initial depth information of a single image and the camera shooting position and posture defined by the external parameter matrix.

[0094] In one possible implementation, the depth information corresponding to each shooting perspective is determined based on the initial depth information of a single image and the extrinsic parameter matrix corresponding to each shooting perspective, including: converting the pixel points of the target object into a three-dimensional point cloud according to the initial depth information of the single image to obtain the three-dimensional point cloud information of the initial shooting perspective; determining the three-dimensional point cloud information of the target object at each shooting perspective based on the three-dimensional point cloud information of the initial shooting perspective and the extrinsic parameter matrix corresponding to each shooting perspective; and projecting the three-dimensional point cloud information of each shooting perspective onto a two-dimensional plane to obtain the depth information corresponding to each shooting perspective.

[0095] Using the initial depth information D0 of a single image, the pixels of the target object are converted into a 3D point cloud P0 = {(X, Y, Z) | Z = D0(u, v), (u, v) ∈ M0}. The calculation of (X, Y, Z) can refer to formula (2).

[0096] For the i-th shooting angle, the external parameter matrix [R i |T i ], transform the 3D point cloud P0 of the initial shooting perspective to the i-th shooting perspective, and obtain the 3D point cloud information P of the i-th shooting perspective i =R i P0+T i Then, referring to the following formula (8), the 3D point cloud information of the i-th shooting angle is projected back to the external parameter matrix [R i |T i ] The corresponding image plane is used to obtain the depth information D of the i-th shooting angle i .

[0097] (8)

[0098] Step S150 , calling the diffusion model, using the intrinsic parameter matrix, the extrinsic parameter matrix corresponding to each shooting angle, and the depth information as the constraints of the denoising process, denoising the preset noise image in the diffusion model to obtain the image data of the target object at each shooting angle.

[0099] Step S160 , generating full-view 3D point cloud information of the target object based on the image data of the target object at each shooting viewpoint.

[0100] In the diffusion model, the forward process adds noise to an image until pure noise is generated. In the present embodiment, this forward process is reversed, gradually denoising the pure noise to generate a new image. This is known as the reverse diffusion process. However, in order to ensure that the objects in the denoised new image do not deviate from the geometric form of the target object captured at the initial viewing angle, the present embodiment uses the camera's intrinsic parameter matrix, as well as the extrinsic parameter matrix and depth information corresponding to each viewing angle, as constraints for the reverse diffusion process. The noisy image in the diffusion model is denoised to obtain a new image corresponding to each viewing angle.

[0101] Specifically, the initial noise image X in the diffusion model is predefined T , such as X T ~N(0,1), referring to the inverse diffusion process represented by formula (9), at each denoising step t, the initial noise image X TAfter n rounds of denoising, the initial noise image X T Gradually transform it into a new perspective image to obtain the image data I of the target object at each shooting perspective i .

[0102] (9)

[0103] in, is a pre-trained denoising network, η is the denoising step length, (D i ,K,R i ,T i ) is a constraint condition.

[0104] Furthermore, the image data corresponding to each shooting perspective is spliced, including the image data of a single image of the initial shooting perspective, to obtain the full-perspective image data of the target object. Based on the full-perspective image data, three-dimensional point cloud data of the target object in the full perspective is generated to complete the three-dimensional reconstruction of the target object.

[0105] In summary, the present application provides a method for reconstructing an object point cloud based on a single image. The method uses the intrinsic parameter matrix of the camera used to capture the single image, as well as the extrinsic parameter matrix and depth information of each shooting angle determined based on the single image, as constraints of the diffusion model to perform inverse denoising on the initial noisy image and generate image data corresponding to each shooting angle of the target object. The introduction of the intrinsic parameter matrix of the camera enables the diffusion model to more accurately simulate the camera's imaging process when generating new angle images based on the imaging parameters of the single image obtained by the camera capturing the target object as represented by the intrinsic parameter matrix. The diffusion model in the prior art relies solely on a data-driven learning mechanism, which lacks constraints on physical imaging laws and may cause perspective distortion of the emitted light in the new angle image generated by denoising. The present application introduces the intrinsic parameter matrix of the camera as a constraint and adjusts the transformation calculation in the diffusion process so that the generated new angle image conforms to the real camera projection, reduces the deformation of the object in the image relative to the target object, and improves the perspective consistency of the object in the image.

[0106] The extrinsic parameter matrix corresponding to each shooting angle provides perspective transformation information, ensuring that images from different shooting angles remain consistent in spatial position. Existing diffusion models can cause drift or misalignment between perspectives when generating multi-perspective images. This application constrains the inverse denoising process based on the camera position and camera attitude in the extrinsic parameter matrix, ensuring that the generated new perspective images remain globally consistent across different shooting angles, avoiding misalignment of images from different perspectives.

[0107] The depth information in the constraint conditions further enhances the geometric constraints of the inverse diffusion process. The depth information can characterize the geometric features of the target object under its corresponding shooting perspective. Based on this as a constraint condition, the geometric form of the object in the generated new perspective image does not deviate from the geometric features of the target object, reducing the blur or distortion of the object and improving the accuracy of the generated image.

[0108] Based on this, the present application adjusts the constraints of the reverse denoising of the diffusion model, on the one hand, to complete the other perspective images of the target object in a single image, and on the other hand, the generated multi-perspective images not only meet the physical imaging laws, but also improve the geometric consistency and spatial accuracy between the multi-perspective images. Based on this, the multi-perspective images of the target object with higher comprehensive accuracy, the full-perspective three-dimensional point cloud information of the target object generated is more accurate, wherein the three-dimensional point cloud information can include a three-dimensional point cloud and a three-dimensional point coordinate collection, the three-dimensional point cloud and the three-dimensional point coordinate collection are different representations of a three-dimensional point set, the three-dimensional point cloud is used to describe the density state of the three-dimensional point coordinate collection in space, and the point coordinate collection is used for the formal representation of the point cloud, such as the three-dimensional coordinate representation.

[0109] Next, other possible implementations of the above-mentioned object point cloud reconstruction method based on a single image are explained through the embodiments described below.

[0110] In one possible implementation, step S160 generates full-view three-dimensional point cloud information of the target object from the image data of the target object at each shooting perspective, including: performing depth estimation on the image data of each shooting perspective to obtain target depth information corresponding to each shooting perspective, where the target depth information includes: a depth value of each pixel in the image data; generating three-dimensional point cloud information corresponding to each image data based on the depth value of each pixel in the image data; and splicing the three-dimensional point cloud information corresponding to each two adjacent shooting perspectives to obtain full-view three-dimensional point cloud information of the target object.

[0111] First, the image data I corresponding to each shooting angle i , input monocular depth estimation model f θ Depth estimation / depth prediction is performed in , and image data of each shooting angle is obtained I i Corresponding target depth information D i =f θ (I i ).

[0112] For the target depth information Di corresponding to each shooting angle, the following processing is performed: Referring to the following formula (10), according to the depth value Di(u,v) corresponding to each pixel point (u,v) in the target depth information, the three-dimensional coordinates P of each pixel point are calculated: cFurthermore, the three-dimensional coordinate Pc is converted from the camera coordinate system to the world coordinate system to obtain the three-dimensional point cloud of each pixel under the shooting angle .

[0113] (10)

[0114] Based on this, the three-dimensional point cloud of all pixels in the image data of all shooting angles is collected to obtain the three-dimensional point cloud information of the target object with a complete perspective. .

[0115] Furthermore, the 3D point cloud information P of the target object under the complete viewing angle is processed based on ICP (Iterative Closest Point), aligning the 3D point clouds of different shooting angles, improving the matching accuracy of the 3D point clouds between adjacent viewing angles, and reducing the overlap of point clouds between adjacent shooting angles. Specifically, first, the 3D point cloud information corresponding to the initial shooting angle θ0 is used as the stitching reference, recorded as the reference 3D point cloud information. The 3D point cloud information corresponding to the shooting angles adjacent to the initial shooting angle on the left and right is then ICP-registered with the reference 3D point cloud information. The ICP registration process can be referred to the following equation (11).

[0116] (11)

[0117] Where p is the pixel to be aligned in the shooting angle to be aligned, q is the nearest neighbor of the pixel to be aligned in the reference shooting angle, R represents the rotation of p relative to q, and T represents the translation of p relative to q.

[0118] Based on formula (11), the relative displacement of the point cloud of the shooting perspective to be aligned relative to the 3D point cloud of the initial shooting perspective is calculated, and the 3D point cloud of the pixel points to be aligned in the shooting perspective to be aligned is adjusted to complete the alignment of the nearest points, thereby realizing the alignment and splicing of the 3D point clouds between adjacent perspectives.

[0119] After completing the stitching of the initial shooting perspective and the two adjacent shooting perspectives on the left and right, the three-dimensional point clouds corresponding to the three stitched shooting perspectives are used as new reference shooting perspectives in turn. The three-dimensional point clouds of the new left and right adjacent perspectives corresponding to the new reference shooting perspective are selected, and the above alignment process is repeated to gradually realize the stitching of the three-dimensional point cloud information of all shooting perspectives, ensuring the continuity and consistency of the point cloud data.

[0120] In the above stitching process, there is another possible implementation. After generating the three-dimensional point cloud information corresponding to each image data based on the depth value of each pixel point in the image data, it also includes: filtering out the three-dimensional point cloud information whose reference distance value is greater than a preset threshold. The reference distance value is: the average value of the distance between any three-dimensional point cloud coordinate in the three-dimensional point cloud information and all other three-dimensional point cloud coordinates.

[0121] In the generated three-dimensional point cloud information of the new shooting angle, there will inevitably be prediction errors, which will cause individual three-dimensional point clouds in the three-dimensional point cloud information to be outliers. During the stitching process, it is difficult to stitch the outliers into the point cloud of the target object, resulting in the failure of the above stitching step. Based on this, the embodiment of the present application first filters out the outliers in the three-dimensional point cloud information before stitching the three-dimensional point cloud information. Specifically, referring to the following formula (12), the i-th three-dimensional point cloud P is calculated i The corresponding d i , if d i If the set threshold is exceeded, the three-dimensional point cloud P i Consider them as outliers and remove them to confirm. According to the 3D point cloud information set after removing all outliers, the above stitching process is performed.

[0122] (12)

[0123] In one possible implementation, after obtaining the full-view three-dimensional point cloud information of the target object, it also includes: moving or rotating the full-view three-dimensional point cloud information of the target object based on the three-dimensional point cloud information of the initial shooting perspective, to obtain the target full-view three-dimensional point cloud information of the target object with the initial shooting perspective as the front.

[0124] Based on the global image containing the target object displayed by a single image, the three-dimensional point cloud information of the target object is aligned with the position and posture of the target object in the global image to restore the shape of the target object in the global image.

[0125] Specifically, since the origin of the world coordinate system is the camera position, the center pixel point C(x c ,y c ,z c ) as the translation matrix T', each 3D point cloud in the full-view 3D point cloud information Translate to get , thereby moving the target object of the obtained full-view 3D point cloud information to its original position in the scene of a single image.

[0126] Furthermore, the 3D point cloud information P0 under the initial shooting angle is used as the basis for registration , using ICP (Iterative Closest Point) algorithm to optimize calculation and The rotation matrix R between opt Furthermore, based on the rotation matrix corresponding to each 3D point cloud, the 3D point cloud is calibrated to obtain , so that the target object moved to the original position is rotated to the same posture as the target object in a single image.

[0127] In another possible implementation, after restoring the position and posture of the target object of the full-view three-dimensional point cloud information in a single image, it can also include: embedding the target full-view three-dimensional point cloud information of the target object into the three-dimensional point cloud information of the scene in the single image to obtain the three-dimensional point cloud information of the single image.

[0128] Specifically, the full-view 3D point cloud information of the target object Add to scene point cloud In the example above, we get the global point cloud of a single image. .

[0129] Since the point cloud of the target object may overlap with the scene point cloud, Euclidean distance clustering can be used to delete points that are too close to the existing point clouds. Specifically, refer to the following formula (13) to filter out duplicate point clouds and obtain the final global point cloud. .

[0130] (13)

[0131] Among them, δ is the minimum distance threshold.

[0132] Optionally, to reduce point cloud redundancy, a voxel grid filter can be used. Set the voxel size v = 0.01 meters, divide the point cloud into 3D voxel grids, and retain only the center point of each voxel. Refer to formula (14) to determine the global point cloud .

[0133] (14)

[0134] Where N is the number of point clouds within the current voxel.

[0135] Optionally, a moving least squares (MLS) filter can be used to make the point cloud surface smoother. Referring to the following formula (15), the global three-dimensional point cloud information after the point cloud surface optimization is determined.

[0136] (15)

[0137] Among them, P' represents the optimized three-dimensional point cloud, N(i) is the neighborhood point set of point cloud P, and w j is a distance-based weight.

[0138] The above describes a method for reconstructing an object point cloud based on a single image provided in an embodiment of the present application. The following describes a device for executing the above method for reconstructing an object point cloud based on a single image.

[0139] See also Figure 3 , Figure 3 This is a schematic diagram of a device for reconstructing a point cloud of an object based on a single image provided in an embodiment of the present application. Figure 3 As shown, the object point cloud reconstruction device based on a single image includes:

[0140] An image acquisition unit 100 is configured to acquire a single image containing a target object and an intrinsic parameter matrix corresponding to the single image, wherein the intrinsic parameter matrix is used to represent imaging parameters used by a camera to capture the target object to obtain the single image;

[0141] The three-dimensional coordinate acquisition unit 200 is used to determine the pixel points of the target object and obtain a three-dimensional coordinate set obtained by projecting the pixel points of the target object;

[0142] An extrinsic parameter matrix determination unit 300 is configured to determine, based on a preset shooting rule, multiple shooting angles corresponding to the three-dimensional coordinate set, and an extrinsic parameter matrix corresponding to each shooting angle, wherein the extrinsic parameter matrix is configured to represent state parameters of the camera when shooting the target object at the corresponding shooting angle, the state parameters including camera position and camera attitude.

[0143] A depth information determining unit 400 is configured to determine the depth information corresponding to each shooting perspective based on the initial depth information of the single image and the extrinsic parameter matrix corresponding to each shooting perspective;

[0144] The image diffusion unit 500 is configured to call a diffusion model, use the intrinsic parameter matrix, the extrinsic parameter matrix corresponding to each shooting angle, and the depth information as constraints for denoising, perform the denoising process on the preset noise image in the diffusion model, and obtain image data of the target object at each shooting angle;

[0145] The full-view point cloud generating unit 600 is configured to generate full-view three-dimensional point cloud information of the target object based on the image data of the target object at each of the shooting view angles.

[0146] In summary, the present application uses the intrinsic parameter matrix of the camera used to shoot a single image, and the extrinsic parameter matrix and depth information of each shooting angle determined based on the single image, as the constraints of the diffusion model, to perform reverse denoising on the initial noisy image and generate image data corresponding to each shooting angle of the target object. Among them, the camera position and posture represented by the extrinsic parameter matrix of each shooting angle provides perspective transformation information for the single image, ensuring that the multi-perspective images generated after denoising remain consistent in three-dimensional spatial position to avoid spatial misalignment; the depth information of each shooting angle provides geometric constraints for the denoising process, making the image generated for each shooting angle more accurate in three-dimensional structure and reducing the deformation of objects in images of different perspectives; the intrinsic parameter matrix can more accurately represent the imaging principle of the camera that shoots a single image, constrain the imaging of other perspectives, and thus improve the perspective consistency between multi-perspective images.

[0147] Based on this, this application adjusts the constraints of the diffusion model's reverse denoising to, on the one hand, complete the other perspective images of the target object in a single image. On the other hand, the generated multi-perspective images not only meet the physical imaging laws, but also improve the geometric consistency and spatial accuracy between the multi-perspective images. Based on this, the multi-perspective images of the target object with higher comprehensive accuracy generate more accurate full-perspective 3D point cloud information of the target object.

[0148] In a possible implementation, the extrinsic parameter matrix determining unit 300 includes:

[0149] a shooting angle determination subunit, configured to determine a plurality of shooting angles based on a preset angle interval, taking the initial shooting angle at which the camera captures the single image in the three-dimensional coordinate set as a starting point;

[0150] a shooting height determination subunit, configured to determine a camera shooting height corresponding to the initial shooting angle of view based on a depth value of an image center point in the initial depth information of the single image;

[0151] a three-dimensional coordinate determination subunit, configured to determine a shooting distance of the target object when the camera is photographing the target object according to the object size of the target object, the object size being determined based on the three-dimensional coordinate set;

[0152] A camera coordinate determination subunit, configured to determine the three-dimensional coordinates of the camera corresponding to each of the shooting angles based on the shooting distance and the camera shooting height;

[0153] a rotation attitude determination subunit, configured to determine a camera rotation matrix corresponding to each shooting perspective based on differences between the three-dimensional coordinates of the camera at each shooting perspective and the three-dimensional coordinates of the object center pixel point of the target object, wherein the camera rotation matrix is used to represent the line of sight rotation attitude of the camera at its corresponding shooting perspective relative to the initial shooting perspective;

[0154] The extrinsic parameter matrix determination subunit is configured to generate an extrinsic parameter matrix for each shooting perspective based on the camera three-dimensional coordinates and the camera rotation matrix corresponding to each shooting perspective.

[0155] In a possible implementation, the depth information determining unit 400 includes:

[0156] An initial point cloud acquisition subunit is used to convert the pixel points of the target object into a three-dimensional point cloud based on the initial depth information of the single image, so as to obtain the three-dimensional point cloud information of the initial shooting perspective;

[0157] a multi-view point cloud generating subunit, configured to determine the three-dimensional point cloud information of the target object at each shooting perspective based on the three-dimensional point cloud information of the initial shooting perspective and the extrinsic parameter matrix corresponding to each shooting perspective;

[0158] The depth information determination subunit is used to project the three-dimensional point cloud information of each shooting perspective onto a two-dimensional plane to obtain the depth information corresponding to each shooting perspective.

[0159] In a possible implementation, the full-view point cloud generation unit 600 includes:

[0160] a depth estimation subunit, configured to perform depth estimation on the image data of each shooting angle to obtain target depth information corresponding to each shooting angle, wherein the target depth information includes a depth value of each pixel in the image data;

[0161] a point cloud generating subunit, configured to generate three-dimensional point cloud information corresponding to each image data according to a depth value of each pixel in the image data;

[0162] The point cloud stitching subunit is used to stitch the three-dimensional point cloud information corresponding to each two adjacent shooting perspectives to obtain the full-perspective three-dimensional point cloud information of the target object.

[0163] In a possible implementation, the method further includes:

[0164] The point cloud filtering subunit is used to filter out three-dimensional point cloud information having a reference distance value greater than a preset threshold value after the point cloud generation subunit executes the step of generating three-dimensional point cloud information corresponding to each image data based on the depth value of each pixel point in the image data, wherein the reference distance value is the average value of the distance between any three-dimensional point cloud coordinate in the three-dimensional point cloud information and all other three-dimensional point cloud coordinates.

[0165] In a possible implementation, the method further includes:

[0166] The perspective alignment unit is used to move or rotate the full-perspective three-dimensional point cloud information of the target object based on the three-dimensional point cloud information of the initial shooting perspective, so as to obtain the target full-perspective three-dimensional point cloud information of the target object with the initial shooting perspective as the front.

[0167] In a possible implementation, the method further includes:

[0168] The point cloud embedding unit is used to embed the full-view three-dimensional point cloud information of the target object into the three-dimensional point cloud information of the scene in the single image to obtain the three-dimensional point cloud information of the single image.

[0169] In a possible implementation, the three-dimensional coordinate acquisition unit 200 includes:

[0170] an image segmentation subunit, configured to perform image segmentation processing on the single image to obtain a segmentation mask of the single image;

[0171] a pixel extraction subunit, configured to extract pixel points of the target object from the segmentation mask;

[0172] The three-dimensional coordinate determination subunit is used to project each pixel point into the three-dimensional space based on the depth value corresponding to each pixel point in the initial depth information of the single image and the intrinsic parameter matrix to obtain the three-dimensional coordinate set corresponding to the target object.

[0173] An embodiment of the present application also provides a computer program product including computer-readable instructions. When the computer-readable instructions are executed on an electronic device, the electronic device implements any one of the object point cloud reconstruction methods based on a single image provided in the embodiment of the present application.

[0174] It should also be noted that the device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed across multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment. In addition, in the drawings of the device embodiments provided in this application, the connection relationship between the modules indicates that there is a communication connection between them, which can be specifically implemented as one or more communication buses or signal lines.

[0175] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus necessary general hardware, and of course can also be implemented by special hardware including application-specific integrated circuits, special CPUs, special memories, special components, etc. In general, all functions performed by computer programs can be easily implemented with corresponding hardware, and the specific hardware structures used to implement the same function can also be diverse, such as analog circuits, digital circuits or special circuits, etc. However, for the present application, software program implementation is a better implementation method in most cases. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer's floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk or optical disk, etc., and includes a number of instructions to enable a computer device (which can be a personal computer, training equipment, or network equipment, etc.) to execute the methods described in each embodiment of the present application.

[0176] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware, or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product.

[0177] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, a computer, a training device or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website, a computer, a training device or a data center. The computer-readable storage medium can be any available medium that a computer can store or a data storage device such as a training device, a data center, etc. that includes one or more available media integrations. The available medium can be a magnetic medium, (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)).

Claims

1. A method for reconstructing an object point cloud based on a single image, characterized in that: include: Acquire a single image containing a target object and an intrinsic parameter matrix corresponding to the single image, wherein the intrinsic parameter matrix is used to represent imaging parameters obtained by a camera photographing the target object to obtain the single image; Determine the pixel points of the target object, and obtain a three-dimensional coordinate set obtained by projecting the pixel points of the target object; Determining, based on preset shooting rules, multiple shooting angles corresponding to the three-dimensional coordinate set and an extrinsic parameter matrix corresponding to each shooting angle, wherein the extrinsic parameter matrix is used to represent state parameters of the camera when shooting the target object at the corresponding shooting angle, the state parameters including camera position and camera attitude; Determining the depth information corresponding to each shooting perspective based on the initial depth information of the single image and the extrinsic parameter matrix corresponding to each shooting perspective; Invoking a diffusion model, using the intrinsic parameter matrix, the extrinsic parameter matrix corresponding to each shooting angle, and the depth information as constraints for denoising, performing the denoising process on a preset noise image in the diffusion model to obtain image data of the target object at each shooting angle; Full-view three-dimensional point cloud information of the target object is generated based on the image data of the target object at each shooting perspective.

2. The object point cloud reconstruction method based on a single image according to claim 1, characterized in that: The step of determining, based on a preset shooting rule, a plurality of shooting perspectives corresponding to the three-dimensional coordinate set and an extrinsic parameter matrix corresponding to each shooting perspective includes: Taking the initial shooting angle of view of the single image captured by the camera in the three-dimensional coordinate set as a starting point, a plurality of shooting angles are determined according to a preset angle interval; Determining a camera shooting height corresponding to the initial shooting angle of view based on a depth value of an image center point in the initial depth information of the single image; determining a shooting distance of the target object when the camera shoots the target object according to the object size of the target object, wherein the object size is determined based on the three-dimensional coordinate set; Determining the three-dimensional coordinates of the camera corresponding to each shooting angle based on the shooting distance and the camera shooting height; Determining a camera rotation matrix corresponding to each shooting perspective based on differences between the three-dimensional coordinates of the camera at each shooting perspective and the three-dimensional coordinates of the object center pixel point of the target object, wherein the camera rotation matrix is used to represent a rotational posture of the camera at the corresponding shooting perspective relative to the initial shooting perspective; Based on the camera three-dimensional coordinates and the camera rotation matrix corresponding to each shooting perspective, an extrinsic parameter matrix for each shooting perspective is generated.

3. The object point cloud reconstruction method based on a single image according to claim 1, characterized in that: The determining, based on the initial depth information of the single image and the extrinsic parameter matrix corresponding to each shooting perspective, the depth information corresponding to each shooting perspective includes: Converting the pixels of the target object into a three-dimensional point cloud based on the initial depth information of the single image to obtain three-dimensional point cloud information of the initial shooting perspective; Determining the three-dimensional point cloud information of the target object at each shooting perspective based on the three-dimensional point cloud information of the initial shooting perspective and the extrinsic parameter matrix corresponding to each shooting perspective; The three-dimensional point cloud information of each shooting perspective is projected onto a two-dimensional plane to obtain depth information corresponding to each shooting perspective.

4. The object point cloud reconstruction method based on a single image according to claim 1, characterized in that Generating full-view three-dimensional point cloud information of the target object based on the image data of the target object at each shooting angle includes: Performing depth estimation on the image data of each shooting angle to obtain target depth information corresponding to each shooting angle, the depth information including: a depth value of each pixel in the image data; Generating three-dimensional point cloud information corresponding to each image data according to the depth value of each pixel in the image data; The three-dimensional point cloud information corresponding to each two adjacent shooting perspectives is spliced to obtain the full-perspective three-dimensional point cloud information of the target object.

5. The object point cloud reconstruction method based on a single image according to claim 4, characterized in that: After generating three-dimensional point cloud information corresponding to each image data according to the depth value of each pixel in the image data, the method further includes: Filter out three-dimensional point cloud information with a reference distance value greater than a preset threshold, wherein the reference distance value is: an average value of distances between any three-dimensional point cloud coordinate and all other three-dimensional point cloud coordinates in the three-dimensional point cloud information.

6. The object point cloud reconstruction method based on a single image according to claim 1, characterized in that: Also includes: According to the three-dimensional point cloud information of the initial shooting angle of view, the full-view three-dimensional point cloud information of the target object is moved or rotated to obtain the target full-view three-dimensional point cloud information of the target object with the initial shooting angle of view as the front.

7. The object point cloud reconstruction method based on a single image according to claim 6, characterized in that: Also includes: The target full-view three-dimensional point cloud information of the target object is embedded into the three-dimensional point cloud information of the scene in the single image to obtain the three-dimensional point cloud information of the single image.

8. The object point cloud reconstruction method based on a single image according to any one of claims 1 to 7, characterized in that: The determining of the pixel points of the target object and obtaining a three-dimensional coordinate set obtained by projecting the pixel points of the target object includes: performing image segmentation processing on the single image to obtain a segmentation mask of the single image; Extracting pixels of the target object from the segmentation mask; Based on the depth value corresponding to each pixel in the initial depth information of the single image and the intrinsic parameter matrix, each pixel is projected into a three-dimensional space to obtain a three-dimensional coordinate set corresponding to the target object.

9. A device for reconstructing object point cloud based on a single image, characterized in that: include: An image acquisition unit, configured to acquire a single image containing a target object and an intrinsic parameter matrix corresponding to the single image, wherein the intrinsic parameter matrix is used to represent imaging parameters obtained by a camera photographing the target object to obtain the single image; A three-dimensional coordinate acquisition unit, configured to determine the pixel points of the target object and acquire a three-dimensional coordinate set obtained by projecting the pixel points of the target object; an extrinsic parameter matrix determination unit, configured to determine, based on preset shooting rules, a plurality of shooting angles corresponding to the three-dimensional coordinate set, and an extrinsic parameter matrix corresponding to each shooting angle, wherein the extrinsic parameter matrix is used to represent state parameters of the camera when shooting the target object at the corresponding shooting angle, the state parameters including camera position and camera attitude; a depth information determining unit, configured to determine the depth information corresponding to each of the shooting perspectives based on the initial depth information of the single image and the extrinsic parameter matrix corresponding to each of the shooting perspectives; an image diffusion unit, configured to call a diffusion model, use the intrinsic parameter matrix, the extrinsic parameter matrix corresponding to each of the shooting angles, and the depth information as constraints for denoising, perform the denoising process on a preset noise image in the diffusion model, and obtain image data of the target object at each of the shooting angles; The full-view point cloud generating unit is used to generate full-view three-dimensional point cloud information of the target object based on the image data of the target object at each shooting perspective.

10. A computer program product, characterized in that The method comprises computer-readable instructions, which, when executed on an electronic device, enable the electronic device to implement the object point cloud reconstruction method based on a single image as claimed in any one of claims 1 to 8.