Point cloud construction method and device, electronic equipment and program product

By obtaining the detected depth and predicted depth of multiple frames of scene images, determining the scale compensation parameters, and compensating the detected depth, the problem of large errors in 3D model reconstruction in the existing technology is solved, and a more accurate 3D point cloud construction is achieved.

CN120635170APending Publication Date: 2025-09-12BEIJING AUTONAVI YUNMAP TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510466601.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

When obtaining depth information of objects in a scene, existing technologies are prone to misalignment, stretching, or compression of the reconstructed three-dimensional model, resulting in significant deviations in shape and size between the model and the actual object.

Method used

By obtaining multiple frames of scene images corresponding to the target scene and the detection depth of their target image feature points, the scale compensation parameters are determined using the predicted depth and the detected depth, and the scale compensation is performed on the detected depth of the pixel points in the scene image to obtain a more accurate predicted depth to construct a three-dimensional point cloud.

Benefits of technology

The accuracy of the depth information required to construct the three-dimensional point cloud of the target scene is improved, thereby improving the accuracy and effect of the three-dimensional point cloud.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635170A_ABST
    Figure CN120635170A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a point cloud construction method and device, electronic equipment and a program product. The method comprises the steps that multiple frames of scene images corresponding to a target scene and detection depths corresponding to target image feature points in all the scene images are acquired, and all the target image feature points are projections of target feature points in the target scene in all the scene images; and obtaining a prediction depth corresponding to each target image feature point, and determining a scale compensation parameter corresponding to the target scene based on the prediction depth corresponding to each target image feature point and the detection depth. And performing scale compensation on the detection depth corresponding to each pixel point in each scene image by using the scale compensation parameter to obtain the predicted depth of each pixel point, the predicted depth of each pixel point being used for constructing a three-dimensional point cloud corresponding to the target scene. The method is used for achieving the effect of improving the precision of three-dimensional reconstruction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a point cloud construction method, device, electronic device, and program product. Background Art

[0002] In the fields of computer vision and 3D reconstruction, obtaining depth information of objects in a scene is crucial. Depth information not only helps understand the 3D structure of objects but also provides key data support for robotic navigation, augmented reality, autonomous driving, and other fields.

[0003] Currently, acquiring depth information at different points in a scene primarily relies on a variety of methods, including binocular vision, structured light, time-of-flight, and depth estimation based on single images. However, relying on depth information obtained by these methods to reconstruct objects often results in misalignment, stretching, or compression of the reconstructed 3D model, resulting in significant deviations in shape and size from the actual object. Summary of the Invention

[0004] The embodiments of the present application provide a point cloud construction method, device, electronic device and program product to achieve the desired effect.

[0005] In a first aspect, an embodiment of the present application provides a point cloud construction method, comprising:

[0006] Acquire multiple frames of scene images corresponding to a target scene, and detection depths corresponding to target image feature points in each of the scene images, where each of the target image feature points is a projection of a target feature point in the target scene in each of the scene images;

[0007] Obtaining the predicted depth corresponding to each of the target image feature points;

[0008] Determining a scale compensation parameter corresponding to the target scene based on the predicted depth and the detected depth corresponding to each of the target image feature points;

[0009] The scale compensation parameters are used to perform scale compensation on the detected depth corresponding to each pixel in each of the scene images to obtain a predicted depth of each pixel, wherein the predicted depth of each pixel is used to construct a three-dimensional point cloud corresponding to the target scene.

[0010] In a second aspect, an embodiment of the present application provides a point cloud construction device, comprising:

[0011] A first acquisition module is configured to acquire multiple frames of scene images corresponding to a target scene, and a detection depth corresponding to a target image feature point in each of the scene images, where each of the target image feature points is a projection of a target feature point in the target scene in each of the scene images;

[0012] A second acquisition module is used to obtain the predicted depth corresponding to each target image feature point;

[0013] A first processing module is configured to determine a scale compensation parameter corresponding to the target scene based on the predicted depth and the detected depth corresponding to each of the target image feature points;

[0014] The second processing module is used to use the scale compensation parameter to perform scale compensation on the detected depth corresponding to each pixel in each of the scene images to obtain a predicted depth of each pixel, wherein the predicted depth of each pixel is used to construct a three-dimensional point cloud corresponding to the target scene.

[0015] In a third aspect, an embodiment of the present application provides an electronic device, comprising: a processor and a memory; the processor is communicatively connected to the memory;

[0016] The memory stores computer-executable instructions;

[0017] The processor executes the computer-executable instructions stored in the memory to implement the method as described in any one of the first aspects.

[0018] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, they are used to implement the method as described in any one of the first aspects above.

[0019] In a fifth aspect, an embodiment of the present application provides a computer program product, comprising a computer program, which, when executed by a processor, implements the method as described in any one of the first aspects.

[0020] The point cloud construction method, apparatus, electronic device, and program product provided in the embodiments of the present application obtain multiple frames of scene images corresponding to a target scene, as well as the detected depth corresponding to the target image feature points in each scene image, and obtain the predicted depth corresponding to each target image feature point. Based on the predicted depth and the detected depth corresponding to each target image feature point, a scale compensation parameter corresponding to the target scene is determined. The scale compensation parameter is used to perform scale compensation on the detected depth corresponding to each pixel in each scene image, and the predicted depth of each pixel is obtained. This improves the accuracy of the depth information required to construct a three-dimensional point cloud of the target scene, thereby improving the accuracy and quality of the constructed three-dimensional point cloud. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0022] Figure 1 Schematic diagram of the scene of the point cloud construction method provided in this application;

[0023] Figure 2 A flowchart of another point cloud construction method provided in an embodiment of the present application;

[0024] Figure 3 A flowchart of another point cloud construction method provided in an embodiment of the present application;

[0025] Figure 4 A flowchart of another point cloud construction method provided in an embodiment of the present application;

[0026] Figure 5 A flowchart of another point cloud construction method provided in an embodiment of the present application;

[0027] Figure 6 A triangulated scene diagram provided in an embodiment of the present application;

[0028] Figure 7 A schematic diagram of a process for constructing a point cloud provided in an embodiment of the present application;

[0029] Figure 8 A schematic diagram of the structure of a point cloud construction device provided in an embodiment of the present application;

[0030] Figure 9 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application.

[0031] The above drawings illustrate specific embodiments of the present application, which will be described in more detail below. These drawings and the textual description are not intended to limit the scope of the present application in any way, but rather to illustrate the concepts of the present application to those skilled in the art by reference to specific embodiments. DETAILED DESCRIPTION

[0032] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims.

[0033] The point cloud construction method provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown, the terminal 102 communicates with the server 104 via a network.

[0034] For example, the point cloud construction method is applied to terminal 102, which can obtain multiple frames of scene images corresponding to a target scene and the detected depths corresponding to target image feature points in each scene image. A predicted depth corresponding to each target image feature point is obtained, and a scale compensation parameter corresponding to the target scene is determined based on the predicted depth and the detected depth corresponding to each target image feature point. Using the scale compensation parameter, the detected depth corresponding to each pixel in each scene image is scale-compensated to obtain a predicted depth for each pixel, which is then stored in a data storage system of server 104. Terminal 102 may be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices may include smart speakers, smart TVs, smart air conditioners, smart car devices, etc. Portable wearable devices may include smart watches, smart bracelets, head-mounted devices, etc. Server 104 may be implemented as a standalone server or a server cluster consisting of multiple servers. Terminal 102 and server 104 may be connected directly or indirectly via wired or wireless communication, such as via a network connection.

[0035] For another example, the point cloud construction method is applied to the server 104, and the terminal 102 can obtain multiple frames of scene images corresponding to the target scene, as well as the detection depth corresponding to the target image feature points in each scene image, and send them to the server 104; the server 104 obtains the predicted depth corresponding to each target image feature point, and determines the scale compensation parameter corresponding to the target scene based on the predicted depth and detection depth corresponding to each target image feature point. Using the scale compensation parameter, the detection depth corresponding to each pixel in each scene image is scale-compensated to obtain the predicted depth of each pixel and store it. It is understandable that the data storage system can be an independent storage device, or the data storage system can be located on the server 104, or the data storage system can be located on another terminal.

[0036] In one embodiment, a point cloud construction method is provided. This embodiment uses the point cloud construction method applied to a terminal as an example. It can be understood that the point cloud construction method can also be applied to a server, and can also be applied to a system including a terminal and a server, and implemented through the interaction between the terminal and the server. Figure 2 This is a flow chart of a point cloud construction method provided in an embodiment of the present application. Figure 2 As shown, the point cloud construction method includes:

[0037] S201: Acquire multiple frames of scene images corresponding to a target scene, and detection depths corresponding to target image feature points in each scene image.

[0038] The target image feature points are projections of the target feature points in the target scene in the image of each scene.

[0039] The target scene refers to a specific spatial area where three-dimensional reconstruction or depth information analysis is desired, such as an indoor room, an outdoor street, the surface of an object, etc.

[0040] A multi-frame scene image is a series of images captured from different angles, at different times, or from different viewing angles. These scene images contain visual information about the target scene under different conditions. By analyzing and processing these multi-frame scene images, the three-dimensional structural information of the target scene can be obtained. For example, the multi-frame scene images can be acquired by an image acquisition device installed on a terminal. The image acquisition device can, for example, capture a scene video of the target scene. The terminal can then perform a blur filter on each frame of the scene video to eliminate blurrier images and retain the remaining clear images as the scene image.

[0041] Target feature points are points with significant characteristics in the target scene, such as corners, edges, and points with rich textures. Target feature points have unique locations and characteristics within the scene, allowing them to be identified and matched across different images. By tracking and analyzing target feature points, we can obtain information about the target scene's motion and 3D structure.

[0042] Target image feature points are projections of target feature points in the target scene within each frame of the scene image. Due to different shooting angles, the same target feature point may appear in different locations in different scene images. The detection depth corresponding to the target image feature point can be directly measured by a depth sensor (such as a LiDAR or structured light sensor) installed on the terminal, obtaining the distance information of the target image feature point relative to the image acquisition device.

[0043] S202: Obtain the predicted depth corresponding to each target image feature point.

[0044] The predicted depth refers to the distance information of the target image feature point relative to the image acquisition device obtained by calculation.

[0045] In this step, stereo vision methods can be used to obtain the predicted depth. For example, when capturing multiple frames of scene images, a feature matching algorithm (e.g., a descriptor-based matching algorithm) can be used to find target image feature points corresponding to the same target feature point in different scene images, based on the scene images captured by the image capture device at different times. Then, based on the geometric relationship between these different scene images, the distance between the target feature points and the image capture device can be calculated using triangulation principles, i.e., the predicted depth corresponding to the target image feature points.

[0046] Since the detected depth will be affected by factors such as sensor accuracy and have deviations, the predicted depth is obtained by integrating multi-frame scene image information, combining the posture of the image acquisition device, and calculating through triangulation and other methods. The posture accuracy of the image acquisition device is usually higher, so the predicted depth is more accurate than the detected depth.

[0047] S203 : Determine a scale compensation parameter corresponding to the target scene based on the predicted depth and the detected depth corresponding to each target image feature point.

[0048] In this step, the scale compensation parameters can be determined by establishing an objective function. For example, a linear relationship can be assumed between the detected depth and the predicted depth, for example, predicted depth = scale factor × detected depth + offset. Then, a set of target image feature points is selected, and their detected depth and predicted depth are substituted into the above linear relationship. Based on this set of target image feature points, an error function (e.g., a least squares error function) is constructed. The error function is optimized using an optimization algorithm (such as the least squares method, iterative nearest point algorithm, etc.) to solve for the scale factor and offset. The value of the error function is minimized, that is, the error between the detected depth after scale compensation and the predicted depth is minimized. The corresponding scale factor and offset at this time are used as the scale compensation parameters corresponding to the target scene.

[0049] S204 , using a scale compensation parameter to perform scale compensation on the detected depth corresponding to each pixel in each scene image to obtain a predicted depth for each pixel.

[0050] Among them, the predicted depth of each pixel is used to construct the three-dimensional point cloud corresponding to the target scene.

[0051] In this step, based on the scale compensation parameters obtained in step S203 above, a scale compensation function for the detection depth corresponding to all pixels in each scene image can be constructed. The detection depth corresponding to each pixel is substituted into the scale compensation function to obtain the predicted depth of each pixel.

[0052] Then, after obtaining the predicted depths corresponding to all pixels in all scene images, 3D reconstruction can be performed based on this more accurate predicted depth to construct a 3D point cloud corresponding to the target scene. This improves the accuracy and quality of the constructed 3D point cloud by increasing the accuracy of the depth information required to construct the 3D point cloud. The specific method for constructing the 3D point cloud of the target scene based on depth information can be referenced in existing technologies and will not be elaborated here.

[0053] The method provided in the embodiments of the present application obtains multiple frames of scene images corresponding to a target scene, as well as the detected depth corresponding to the target image feature points in each scene image, and obtains the predicted depth corresponding to each target image feature point. Based on the predicted depth and the detected depth corresponding to each target image feature point, a scale compensation parameter corresponding to the target scene is determined. The scale compensation parameter is used to scale-compensate the detected depth corresponding to each pixel in each scene image, and the predicted depth of each pixel is obtained. This improves the accuracy of the depth information required to construct a three-dimensional point cloud of the target scene, thereby improving the accuracy and quality of the constructed three-dimensional point cloud.

[0054] The following describes in detail how to determine the scale compensation parameter corresponding to the target scene based on the predicted depth and the detected depth corresponding to each target image feature point in the aforementioned step S203. Figure 3 This is a flow chart of another point cloud construction method provided in the embodiment of the present application. Figure 3 As shown, the aforementioned step S203 may specifically include:

[0055] S301: Obtaining the posture information of the image acquisition device when acquiring images of each scene.

[0056] The pose information of an image acquisition device is used to describe the parameters of the image acquisition device's position and posture relative to the world coordinate system. For example, it may include the image acquisition device's rotation matrix, translation vector, position, and posture. The pose information of an image acquisition device when capturing multiple frames of scene images can be used to indicate the corresponding viewing angle of the image acquisition device when capturing multiple frames of scene images.

[0057] The image acquisition device may be provided with, for example, a GPS positioning module, an inertial measurement unit, a laser radar, etc., to obtain the posture information when the image acquisition device captures each frame of scene image.

[0058] S302: Determine a predicted position of a target feature point based on the pose information corresponding to the target scene image and the predicted depth corresponding to the target image feature point in the target scene image.

[0059] In this step, target image feature points can be matched in multiple frames of target scene images captured by the image capture device at different positions, that is, projection points (i.e., target image feature points) corresponding to the same target feature point in different target scene images can be found. For example, a feature extraction operation can be performed on each captured scene image, such as using a feature point detection algorithm (Oriented FAST and Rotated BRIEF, ORB), a scale-invariant feature transform (SIFT), a speeded up robust features (SURF), etc., to extract multiple target image feature points corresponding to the same target feature point contained in the multiple frames of target scene images.

[0060] Then, for each matched target image feature point, the ray from the image acquisition device's optical center through the target image feature point is determined, combined with the corresponding image acquisition device's pose information and the predicted depth of the target image feature point. This ray specifies the direction in space of the target feature point corresponding to the target image feature point, while the predicted depth gives the predicted distance of the target feature point along the ray direction. Therefore, based on this direction and predicted distance, the predicted position of the target feature point can be determined when the image acquisition device is in this pose.

[0061] Optionally, multiple target image feature points of the same target feature point and the corresponding pose information of the target scene image where the multiple target image feature points are located can be used to determine multiple candidate predicted positions of the target feature point. Then, based on the multiple candidate predicted positions, the predicted position of the target feature point is determined. For example, based on the coordinates corresponding to each of the multiple candidate predicted positions, the mean or median of these coordinates can be used as the coordinates of the predicted position of the target feature point to determine the predicted position of the target feature point; or, based on the average distance from each candidate predicted position to all other candidate predicted positions, the weight of each candidate predicted position can be determined (for example, the closer the point is to other candidate predicted positions, the higher its consistency with the majority of candidate positions is, and the higher the weight is), and then, based on the weight of each candidate predicted position and each candidate predicted position, position weighting processing is performed to determine the predicted position of the target feature point, etc.

[0062] S303 : Determine a scale compensation parameter corresponding to the target scene according to the predicted position of the target feature point and the detected position of the target feature point corresponding to the detected depth of the target image feature point.

[0063] Among them, the detection position of the target feature point can be obtained by determining the predicted position of the target feature point in the above-mentioned step S302. The only difference is that the predicted position of the target feature point is obtained based on the posture information corresponding to the target scene image and the predicted depth corresponding to the target image feature point in the target scene image, while the detection position of the target feature point is obtained based on the posture information corresponding to the target scene image and the detection depth corresponding to the target image feature point in the target scene image.

[0064] One possible implementation involves collecting a large number of data samples containing the predicted and detected positions of target feature points and dividing these samples into training, validation, and test sets. Each sample contains the predicted and detected positions, as well as the corresponding true scale compensation parameters. A suitable machine learning model (such as a neural network model) is then used to train the model, using the predicted and detected positions of the target feature points in the samples as input features and the true scale compensation parameters as output labels. The trained model is then used to input the predicted and detected positions of the target feature points, which then output the scale compensation parameters corresponding to the target scene.

[0065] Another possible implementation method is to construct an error function based on the predicted position and the detected position of the target feature point, and determine the scale compensation parameter corresponding to the target scene by solving the method of minimizing the error function.

[0066] Taking the method of determining the scale compensation parameters corresponding to the target scene by solving the minimization error function as an example, this can be specifically achieved through the following sub-steps:

[0067] S3031. Obtain the predicted positions, detected positions, and initial parameters of the objective function of N target feature points in the target feature set.

[0068] Where N is an integer greater than or equal to 2. The target feature set includes multiple target feature points. N target feature points can be randomly selected from the multiple target feature points or selected according to a preset rule, for example, by using a random sample consensus algorithm (RANSAC). After the N target feature points are determined, the predicted positions and detected positions of the N target feature points can be obtained using the aforementioned method, which will not be further described here.

[0069] The objective function can be set according to actual needs. For example, it can be set to a linear relationship between the predicted position and the detected position of the target feature point. For example, the linear relationship is , where k represents the scale factor and b represents the bias. Represents the predicted position of the i-th target feature point, i∈n, represents the detection position of the i-th target feature point, where This is the objective function.

[0070] It should be understood that the above is only for ease of understanding, and the linear relationship between the predicted position and the detected position of the target feature point is used as an example for introduction. In practice, the specific relationship type between the predicted position and the detected position of the target feature point can also be set according to needs, and this application does not impose any restrictions on this.

[0071] In this step, the detected position of each target feature point among the N target feature points can be substituted into the objective function, and the predicted position of the target feature point is subtracted from the objective function substituted into the detected position to construct the corresponding error function. For example, the error function can be .in, To predict the location, To detect the position, after substituting all N target feature points into the objective function, we can obtain N error functions (i.e., error equations). Then, by minimizing the sum of squared errors of all error functions in the error equations using the least squares method, we can obtain the initial parameters of the objective function (i.e., initial scale factor k0 and initial bias b0).

[0072] S3032. Substitute the detected positions of the M target feature points in the target feature set into the target function to obtain the function-predicted positions of the M target feature points.

[0073] Wherein, M is an integer greater than or equal to N.

[0074] Because the N randomly selected target feature points in S3031 using the RANSAC algorithm represent only a small sample of the target feature set, they may not accurately represent the distribution of the entire target feature set. If these N randomly selected points happen to contain a large number of outliers (“outliers”), the objective function parameters obtained through least squares fitting will be affected by these outliers, resulting in inaccurate parameters.

[0075] Therefore, it is necessary to select M target feature points from the target feature set, expand the verification range, and use more data to verify the fitting effect of the objective function under the initial parameters, so as to adjust the initial parameters to obtain more accurate scale compensation parameters.

[0076] In this step, all target feature points in the target feature set can be selected as the M target feature points, or some target feature points in the target feature set can be selected as the M target feature points. The detected positions of the M target feature points are substituted into the objective function based on the initial parameters to calculate the function-predicted positions of the M target feature points.

[0077] S3033. When the number of target feature points whose error between the function predicted position and the predicted position is less than or equal to the preset number threshold meets the preset conditions, the initial parameters are used as scale compensation parameters; otherwise, the initial parameters are adjusted until the preset conditions are met to obtain the scale compensation parameters.

[0078] The preset condition may be, for example, a quantity threshold set according to actual needs.

[0079] When the error between the function predicted position and the predicted position is less than or equal to a preset number threshold, and the number of target feature points is greater than or equal to the number threshold, then the number of internal points corresponding to the initial parameter is greater than or equal to the number threshold, and the initial parameter is accurate. At this time, the initial parameter can be used as a scale compensation parameter.

[0080] When the error between the function predicted position and the predicted position is less than or equal to the preset number threshold, and the number of target feature points is less than the number threshold, it indicates that the number of inliers corresponding to the initial parameter is less than the number threshold. The initial parameter is inaccurate and needs to be adjusted until it is iterated to meet the preset conditions, and then the scale compensation parameter that meets the preset conditions is obtained.

[0081] Optionally, the preset condition may be, for example, an iteration number threshold. When the iteration number reaches the iteration number threshold, the parameter obtained in the last iteration is used as the scale compensation parameter.

[0082] The method provided in the embodiment of the present application obtains the posture information of the image acquisition device when capturing images of each scene, determines the predicted position of the target feature point based on the posture information corresponding to the target scene image and the predicted depth corresponding to the target image feature point in the target scene image, determines the scale compensation parameter corresponding to the target scene according to the predicted position of the target feature point and the detected position of the target feature point corresponding to the detected depth of the target image feature point, improves the accuracy of the scale step parameter, and provides an accurate compensation basis for subsequent compensation of the depth of all points in the target scene based on the scale compensation parameter, thereby improving the accuracy of the depth information used to construct the three-dimensional point cloud of the target scene, and further improving the effect and accuracy of the constructed three-dimensional point cloud of the target scene.

[0083] Next, a detailed description is given of how to use the scale compensation parameters in the aforementioned step S204 to perform scale compensation on the detected depth corresponding to each pixel in each scene image and obtain the predicted depth of each pixel. Figure 4 This is a flow chart of another point cloud construction method provided in the embodiment of the present application. Figure 4 As shown, the aforementioned step S204 may specifically include:

[0084] S401: Substitute the scale compensation parameter into the objective function to obtain a scale compensation function.

[0085] After obtaining the scale compensation parameters, the accurate scale factor and bias of the objective function are obtained. Within the same target scene, the spatial distribution and depth relationships of objects follow certain physical laws. For example, factors such as light propagation and the relative positions of objects within the target scene are relatively stable throughout the scene. Target feature points are representative points extracted from the scene. The scale relationships they reflect represent, to a certain extent, the scale characteristics of the entire scene. Therefore, the scale compensation function derived from these feature points can be generalized to the entire scene.

[0086] Therefore, for the target scene, the scale compensation function corresponding to the position of the target feature point can be regarded as the scale compensation function of all points in the target scene. Among them, when the points in the target scene are mapped into each scene image, they are the pixels in each scene image.

[0087] S402: Substitute the detected depth corresponding to each pixel into the scale compensation function to calculate and obtain the predicted depth of each pixel.

[0088] Since the scale compensation function corresponding to the position of the target feature point can be regarded as the scale compensation function of all points in the target scene, and the points in the target scene are mapped into each scene image, they are the pixels in each scene image. Therefore, each pixel can also be scale-compensated for its corresponding detected depth using this scale step function, thereby obtaining the predicted depth of each pixel, thereby improving the accuracy of the depth information of each pixel.

[0089] Next, how to obtain the predicted depth corresponding to each target image feature point in the aforementioned step S202 is described in detail. Figure 5 This is a flow chart of another point cloud construction method provided in the embodiment of the present application. Figure 5 As shown, the aforementioned step S202 may specifically include:

[0090] S501: Acquire the posture information of the image acquisition device when acquiring images of each scene.

[0091] One possible implementation method is to directly obtain the posture information of the image acquisition device when it captures each frame of scene image through data collected by the GPS positioning module, inertial measurement unit, lidar, etc. of the image acquisition device.

[0092] Another possible implementation method is to obtain the pose information of the image acquisition device when capturing images of each scene through motion estimation. In this implementation method, the following sub-steps can be used:

[0093] S5011. Obtain the position information of the image acquisition device when acquiring each frame of scene image.

[0094] Among them, multiple frames of scene images are sorted based on position information.

[0095] When capturing scene images, the image capture device's location information can be obtained in a variety of ways. For example, the Global Positioning System (GPS) can be used to receive satellite signals in real time to calculate the image capture device's precise location coordinates, including longitude, latitude, and altitude, thereby obtaining the image capture device's location information when capturing each frame of the scene image. For indoor scenes, an inertial measurement unit can also be used to measure the device's acceleration and angular velocity, combined with an integration algorithm to infer the device's position and attitude changes, thereby obtaining the image capture device's location information when capturing each frame of the scene image.

[0096] After obtaining the position information corresponding to each frame of scene imagery, the multiple frames can be sorted. For example, this can be done based on the spatial relationship of the positional information, such as distance. This can be done by calculating the Euclidean distance between the acquisition positions of each two frames and placing frames with closer distances adjacent to each other. Alternatively, the image sequence can be sorted based on direction, determining a reference direction and sorting the images based on the angle of the image acquisition position relative to the reference direction. This sorted image sequence can better reflect the motion trajectory of the image acquisition device within the scene, facilitating subsequent analysis and processing.

[0097] S5012. For any two adjacent frames of scene images, perform motion estimation on target image feature points corresponding to the same target feature point to obtain posture transformation information when the image acquisition device acquires the two adjacent frames of scene images.

[0098] First, we need to find target image feature points that correspond to the same target feature point in two adjacent scene image frames. Feature extraction algorithms, such as SIFT or SURF, can be used to extract unique and stable feature points from the two adjacent scene image frames. Then, feature matching algorithms, such as nearest neighbor matching or bidirectional matching, are used to find the correspondence between the feature points in the two scene image frames. For example, the Euclidean distance between the descriptors of the feature points in the two image frames can be calculated, and feature points with a distance less than a certain threshold are considered matching points.

[0099] After determining the correspondence between the target image's feature points, mathematical tools such as epipolar geometry, fundamental matrices, and essential matrices can be used for motion estimation to obtain the pose transformation information between two consecutive frames of the scene captured by the image acquisition device. Epipolar geometry describes the geometric relationship between the two images. By solving the fundamental matrix or essential matrix, the rotation and translation information of the image acquisition device between different positions can be obtained.

[0100] S5013. Determine the posture information of the image acquisition device when acquiring each frame of the scene image based on the initial posture information of the image acquisition device and the posture transformation information when the image acquisition device acquires two adjacent frames of scene images.

[0101] After obtaining the initial pose information and the pose transformation information of two adjacent frames of scene images, the pose information of each frame can be calculated iteratively. Specifically, the initial pose matrix can be multiplied by the pose transformation matrix between the first and second frames to obtain the pose matrix of the image acquisition device when capturing the second frame. Then, the pose matrix of the second frame is multiplied by the pose transformation matrix between the second and third frames to obtain the pose matrix of the third frame, and so on, until the pose information of all frames is calculated.

[0102] Optionally, during the iterative calculation process, the pose information may become inaccurate due to accumulated errors. Therefore, the pose information can be corrected and optimized. For example, a filtering algorithm such as a Kalman filter or an extended Kalman filter can be used to filter the calculated pose information to reduce the impact of errors.

[0103] S502: Based on the positions of target image feature points in at least two target scene images and the pose information corresponding to the at least two target scene images, obtain the predicted depth corresponding to each target image feature point.

[0104] In each target scene image, the two-dimensional projection point position of the target feature point can be determined, that is, the pixel coordinates of the image feature point of the target feature point on the target scene image. Since the at least two target scene images are acquired by the image acquisition device at different positions, that is, the image feature points are the two-dimensional projections obtained by the image acquisition device observing the target feature points at different observation positions. Therefore, a triangular geometric relationship can be constructed based on the pose information (image acquisition device coordinates) corresponding to the at least two target scene images and the pixel coordinates of the image feature points to obtain the predicted depth corresponding to each target image feature point. For example, the predicted depth corresponding to each target image feature point can be obtained by the Perspective-n-Point (PnP) algorithm.

[0105] For example, Figure 6 A triangulated scene diagram provided in an embodiment of the present application. Figure 6As shown, a corresponding triangular geometric relationship can be established based on the positions P1 and P2 of the image feature points in the two target scene images and the positions O1 and O2 of the image acquisition devices corresponding to the two target scene images. Based on this triangular geometric relationship, since the pixel coordinates of the positions P1 and P2 of the image feature points are known, and the positions O1 and O2 of the image acquisition devices are known, the predicted depth corresponding to each target image feature point can be obtained based on the triangular geometric relationship, that is, the distance between O1 and the target feature point P, and the distance between O2 and the target feature point P.

[0106] The method provided in the embodiment of the present application obtains the posture information of the image acquisition device when capturing images of each scene, and utilizes the triangulation principle to obtain the predicted depth corresponding to each target image feature point based on the positions of the target image feature points in at least two target scene images and the posture information corresponding to at least two target scene images. This provides a data basis for subsequently adding scale constraints based on the predicted depths corresponding to the target image feature points, so that scale compensation parameters can be subsequently determined based on the predicted depths corresponding to the target image feature points, and the depth information of all points in the target scene can be calibrated based on the scale compensation parameters, thereby improving the accuracy and reliability of the subsequent construction of the three-dimensional point cloud of the target scene.

[0107] Figure 7 This is a flow chart of another point cloud construction method provided in the embodiment of the present application. Figure 7 As shown, the method may further include:

[0108] S701. Obtain the initial position of the target feature point corresponding to the target pixel point in the target coordinate system according to the posture information of the image acquisition device corresponding to each scene image and the position of the target pixel point in each scene image.

[0109] The target coordinate system may be, for example, a world coordinate system, or a reference coordinate system that can be set arbitrarily.

[0110] In this step, the position of the target feature point corresponding to the target pixel can be obtained based on the triangulation method described in the previous embodiment, according to the pose information of the image acquisition device corresponding to each scene image and the position of the target pixel in each scene image. Then, based on the coordinates of the position and the conversion relationship between the coordinate system of the position coordinates and the target coordinate system, the coordinates of the position are converted to the target coordinate system to obtain the initial position of the target feature point.

[0111] S702 : Determine the actual position of the target feature point corresponding to the target pixel point in the target coordinate system according to the initial positions of the target feature points corresponding to the target pixel point in at least two target coordinate systems.

[0112] Since the target pixel point is mapped onto a two-dimensional plane (i.e., the scene image where the target pixel point is located) when the image acquisition device observes the target feature point corresponding to the target pixel point at different positions, and since the depth information collected by the image acquisition device contains noise, even if the accuracy of the depth of the target pixel point is improved by the aforementioned scale compensation method, further denoising processing can still be performed based on the initial positions of the target feature points corresponding to the target pixel point in at least two target coordinate systems to further improve the accuracy of the depth information of the target feature point.

[0113] In this step, the coordinates of the initial positions of the target feature points corresponding to the target pixel points in the target coordinate system can be obtained, and then the coordinates of the initial positions of multiple target feature points in the target coordinate system can be fused to reduce the noise of the positions of the target feature points.

[0114] For example, the median of the coordinates of the initial positions of multiple target feature points in the target coordinate system can be taken to fuse the coordinates of the initial positions of multiple target feature points in the target coordinate system. For example, if there are j target pixels, the coordinates of the j target pixels on the x-axis of the target coordinate system are x1 to x2. j , the coordinates on the y-axis of the target coordinate system are y1 to y j , the coordinates on the z-axis of the target coordinate system are z1 to z j , then the actual position of the target feature point corresponding to the target pixel point in the target coordinate system is:

[0115]

[0116]

[0117]

[0118] S703: Determine the predicted depth of the target pixel point according to the actual position.

[0119] After determining the actual position, the position of the image acquisition device in the target coordinate system can be determined based on the image acquisition device's posture. Then, based on the position of the image acquisition device in the target coordinate system and the actual position of the target feature point corresponding to each target pixel in the target scene in the target coordinate system, the distance between the two is calculated to obtain the predicted depth of each target pixel.

[0120] The method provided in the embodiment of the present application obtains the initial position of the target feature point corresponding to the target pixel in the target coordinate system based on the posture information of the image acquisition device corresponding to each scene image and the position of the target pixel in each scene image, determines the actual position of the target feature point corresponding to the target pixel in the target coordinate system based on the initial position of the target feature point corresponding to the target pixel in at least two target coordinate systems, and determines the predicted depth of the target pixel based on the actual position. This method effectively fuses the position information of the target feature point corresponding to the target pixel in multiple frames of scene images, reduces the noise impact on the position information, improves the accuracy of the actual position of the target feature point corresponding to the target pixel, and further improves the accuracy of the predicted depth of the target pixel determined based on the actual position, providing a high-precision data foundation for the subsequent construction of an accurate and reliable three-dimensional point cloud of the target scene.

[0121] Figure 8 This is a schematic diagram of the structure of a point cloud construction device provided in an embodiment of the present application. Figure 8 As shown, the device may include: a first acquisition module 11 , a second acquisition module 12 , a first processing module 13 , and a second processing module 14 .

[0122] The first acquisition module 11 is used to acquire multiple frames of scene images corresponding to the target scene and the detection depth corresponding to the target image feature points in each scene image, where each target image feature point is the projection of the target feature point in the target scene in each scene image.

[0123] The second acquisition module 12 is used to obtain the predicted depth corresponding to each target image feature point.

[0124] The first processing module 13 is configured to determine a scale compensation parameter corresponding to the target scene based on the predicted depth and the detected depth corresponding to each target image feature point.

[0125] The second processing module 14 is used to use the scale compensation parameter to perform scale compensation on the detected depth corresponding to each pixel in each scene image to obtain the predicted depth of each pixel, wherein the predicted depth of each pixel is used to construct a three-dimensional point cloud corresponding to the target scene.

[0126] Optionally, the first processing module 13 is specifically configured to obtain pose information of each scene image captured by the image acquisition device. A predicted position of a target feature point is determined based on the pose information corresponding to the target scene image and the predicted depth corresponding to the target image feature point in the target scene image. A scale compensation parameter corresponding to the target scene is determined based on the predicted position of the target feature point and the detected position of the target feature point corresponding to the detected depth of the target image feature point.

[0127] Optionally, the first processing module 13 is specifically used to obtain the predicted positions, detected positions, and initial parameters of the objective function of N target feature points in the target feature set. The detected positions of the M target feature points in the target feature set are substituted into the objective function to obtain the function-predicted positions of the M target feature points. When the number of target feature points whose function-predicted positions have an error with the predicted positions that is less than or equal to a preset number threshold meets the preset conditions, the initial parameters are used as scale compensation parameters; otherwise, the initial parameters are adjusted until the preset conditions are met to obtain the scale compensation parameters. Wherein, M is an integer greater than or equal to N, and N is an integer greater than or equal to 2.

[0128] Optionally, the second processing module 14 is specifically configured to substitute the scale compensation parameter into the objective function to obtain a scale compensation function, substitute the detected depth corresponding to each pixel into the scale compensation function, and calculate and obtain the predicted depth of each pixel.

[0129] Optionally, the second acquisition module 12 is specifically configured to obtain pose information of each scene image captured by the image acquisition device. Based on the positions of target image feature points in at least two target scene images and the pose information corresponding to the at least two target scene images, a predicted depth corresponding to each target image feature point is obtained.

[0130] Optionally, the second acquisition module 12 is specifically configured to obtain position information when the image acquisition device captures each frame of the scene image. For any two adjacent frames of the scene image, motion estimation is performed on target image feature points corresponding to the same target feature point to obtain position transformation information when the image acquisition device captures the two adjacent frames of the scene image. The position information when the image acquisition device captures each frame of the scene image is determined based on the initial position information of the image acquisition device and the position transformation information when the image acquisition device captures the two adjacent frames of the scene image. The multiple frames of the scene image are sorted based on the position information.

[0131] Optionally, the second processing module 14 is further configured to obtain an initial position of a target feature point corresponding to a target pixel in a target coordinate system based on the pose information of the image acquisition device corresponding to each scene image and the position of the target pixel in each scene image. Based on the initial positions of the target feature points corresponding to the target pixel in at least two target coordinate systems, determine an actual position of the target feature point corresponding to the target pixel in the target coordinate system. Based on the actual position, determine a predicted depth of the target pixel.

[0132] The point cloud construction device provided in the embodiment of the present application can execute the point cloud construction method in the above method embodiment. Its implementation principle and technical effects are similar and will not be repeated here.

[0133] Figure 9This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. The electronic device can be used to execute the aforementioned point cloud construction method. Figure 9 As shown, the electronic device 900 may include: at least one processor 901 and a memory 902. In a possible implementation, the electronic device 900 may further include a communication interface 903.

[0134] The memory 902 is used to store programs. Specifically, the programs may include program codes, and the program codes include computer operation instructions.

[0135] The memory 902 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.

[0136] The processor 901 is configured to execute computer-executable instructions stored in the memory 902 to implement the method described in the aforementioned method embodiment. The processor 901 may be a CPU, an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.

[0137] Processor 901 can communicate with external devices via communication interface 903. In a specific implementation, if communication interface 903, memory 902, and processor 901 are implemented independently, communication interface 903, memory 902, and processor 901 can be interconnected via a bus to facilitate communication. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. Buses can be categorized as address buses, data buses, and control buses, but this does not necessarily mean there is only one bus or only one type of bus.

[0138] Optionally, in a specific implementation, if the communication interface 903, the memory 902 and the processor 901 are integrated on a chip, the communication interface 903, the memory 902 and the processor 901 can complete communication through an internal interface.

[0139] The present application also provides a computer-readable storage medium, which may include: a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk, and other media that can store program codes. Specifically, the computer-readable storage medium stores program instructions, and the program instructions are used for the methods in the above embodiments.

[0140] The present application also provides a program product, comprising execution instructions stored in a readable storage medium. At least one processor of an electronic device can read the execution instructions from the readable storage medium, and the at least one processor executes the execution instructions so that the electronic device implements the point cloud construction method provided in the various embodiments described above.

[0141] The term "plurality" in this article refers to two or more. The term "and / or" in this article is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent three situations: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the previous and next associated objects are in an "or" relationship; in the formula, the character " / " indicates that the previous and next associated objects are in a "division" relationship. In addition, it should be understood that in the description of this application, words such as "first" and "second" are only used for the purpose of distinguishing the description, and cannot be understood as indicating or implying relative importance, nor can they be understood as indicating or implying order.

[0142] It will be understood that the various numerical numbers involved in the embodiments of the present application are merely distinctions for the convenience of description and are not intended to limit the scope of the embodiments of the present application.

[0143] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some or all of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A point cloud construction method, characterized in that: include: Acquire multiple frames of scene images corresponding to a target scene, and detection depths corresponding to target image feature points in each of the scene images, where each of the target image feature points is a projection of a target feature point in the target scene in each of the scene images; Obtaining the predicted depth corresponding to each of the target image feature points; Determining a scale compensation parameter corresponding to the target scene based on the predicted depth and the detected depth corresponding to each of the target image feature points; The scale compensation parameters are used to perform scale compensation on the detected depth corresponding to each pixel in each of the scene images to obtain a predicted depth of each pixel, wherein the predicted depth of each pixel is used to construct a three-dimensional point cloud corresponding to the target scene.

2. The method according to claim 1, characterized in that The determining of the scale compensation parameter corresponding to the target scene based on the predicted depth and the detected depth corresponding to each target image feature point includes: Acquiring the posture information of the image acquisition device when acquiring the image of each scene; Determining a predicted position of a target feature point based on pose information corresponding to a target scene image and a predicted depth corresponding to a target image feature point in the target scene image; A scale compensation parameter corresponding to the target scene is determined according to the predicted position of the target feature point and the detected position of the target feature point corresponding to the detected depth of the target image feature point.

3. The method according to claim 2, characterized in that The determining of the scale compensation parameter corresponding to the target scene according to the predicted position of the target feature point and the detected position of the target feature point corresponding to the detected depth of the target image feature point includes: Obtain predicted positions, detected positions, and initial parameters of the objective function for N target feature points in the target feature set, where N is an integer greater than or equal to 2; Substituting the detected positions of M target feature points in the target feature set into the objective function to obtain the function-predicted positions of the M target feature points, where M is an integer greater than or equal to N; When the number of target feature points whose error between the function predicted position and the predicted position is less than or equal to a preset number threshold meets the preset conditions, the initial parameters are used as the scale compensation parameters; otherwise, the initial parameters are adjusted until the preset conditions are met to obtain the scale compensation parameters.

4. The method according to claim 3, characterized in that The using the scale compensation parameter to perform scale compensation on the detected depth corresponding to each pixel in each of the scene images to obtain the predicted depth of each pixel includes: Substituting the scale compensation parameter into the objective function to obtain a scale compensation function; The detected depth corresponding to each pixel point is substituted into the scale compensation function to calculate and obtain the predicted depth of each pixel point.

5. The method according to claim 1, characterized in that The obtaining of the predicted depth corresponding to each target image feature point includes: Acquiring the posture information of the image acquisition device when acquiring the image of each scene; Based on the positions of target image feature points in at least two target scene images and the pose information corresponding to at least two of the target scene images, the predicted depth corresponding to each of the target image feature points is obtained.

6. The method according to claim 5, characterized in that The acquiring of the posture information of each scene image acquired by the image acquisition device includes: Acquiring position information of each frame of the scene image captured by the image acquisition device; wherein the multiple frames of the scene image are sorted based on the position information; For any two adjacent frames of scene images, motion estimation is performed on target image feature points corresponding to the same target feature point to obtain pose transformation information when the image acquisition device acquires the two adjacent frames of scene images; The posture information of the image acquisition device when acquiring each frame of the scene image is determined based on the initial posture information of the image acquisition device and the posture transformation information when the image acquisition device acquires the two adjacent frames of scene images.

7. The method according to claim 6, characterized in that Also includes: Obtaining the initial position of the target feature point corresponding to the target pixel point in the target coordinate system according to the posture information of the image acquisition device corresponding to each of the scene images and the position of the target pixel point in each of the scene images; determining, based on the initial positions of the target feature points corresponding to the target pixel points in at least two of the target coordinate systems, the actual positions of the target feature points corresponding to the target pixel points in the target coordinate systems; Determine the predicted depth of the target pixel point according to the actual position.

8. A point cloud construction device, characterized in that: The device comprises: A first acquisition module is configured to acquire multiple frames of scene images corresponding to a target scene, and a detection depth corresponding to a target image feature point in each of the scene images, where each of the target image feature points is a projection of a target feature point in the target scene in each of the scene images; A second acquisition module is used to obtain the predicted depth corresponding to each target image feature point; A first processing module is configured to determine a scale compensation parameter corresponding to the target scene based on the predicted depth and the detected depth corresponding to each of the target image feature points; The second processing module is used to use the scale compensation parameter to perform scale compensation on the detected depth corresponding to each pixel in each of the scene images to obtain a predicted depth of each pixel, wherein the predicted depth of each pixel is used to construct a three-dimensional point cloud corresponding to the target scene.

9. An electronic device, characterized in that: include: processor, and memory; The processor is communicatively connected to the memory; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1 to 7.

10. A computer program product, characterized in that The method comprises a computer program, which implements the method according to any one of claims 1 to 7 when the computer program is executed by a processor.