Scene reconstruction method, apparatus, storage medium, and electronic device
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-06-26
- Publication Date
- 2026-08-13
AI Technical Summary
【0011】 本開示の上記実施例に係るシーン再構築方法、装置、媒体及び電子機器によれば、カメラが収集した多視点画像(即ち、第1の画像シーケンス)に対して後景分割を行い、後景分割によって取得された第2の画像シーケンスを利用して点群のスパース再構築を行って、第1の再構築点群及び第1のカメラ位置姿勢シーケンスを取得することができる。また、多視点画像に対して動的静的分割を行って、静的シーンに対応する第3の画像シーケンス及び動的シーンに対応する第4の画像シーケンスを取得し、第3の画像シーケンスを第1の再構築点群及び第1のカメラ位置姿勢シーケンスと共に点群のデンス再構築に用いることで、静的シーンを効果的に表現する第2の再構築点群を取得し、第4の画像シーケンスを第1のカメラ位置姿勢シーケンスと共に点群のスパース再構築及びデンス再構築に用いることで、動的シーンを効果的に表現する第3の再構築点群を取得し、第2の再構築点群と第3の再構築点群とを結合して、第1の画像シーケンスに対応するシーン再構築結果を決定することができる。以上から分かるように、本開示の実施例では、カメラが収集した多視点画像に対する一連の処理により、シーン再構築を実現し、レーザレーダに比べ、カメラの価格が比較的低いため、シーン再構築の実現コストを低減することができる。また、本開示の実施例では、静的シーン及び動的シーンを分けて処理するポリシーを採用し、即ち静的シーン及び動的シーンをそれぞれ再構築することにより、再構築の難易度を低下させ、再構築速度を向上させることができる。
Smart Images

Figure 0007905009000001 
Figure 0007905009000002 
Figure 0007905009000003
Abstract
Description
[Technical Field]
[0001] [Cross-reference of related applications] This disclosure claims priority to a Chinese patent application filed with the China National Intellectual Property Administration on June 30, 2023, application number CN202310806421.2, with the title of the invention "Scene reconstruction method, apparatus, storage medium and electronic device," the entirety of which is incorporated herein by reference.
[0002] This disclosure relates to driving technology, and in particular to scene reconstruction methods, apparatus, storage media, and electronic equipment. [Background technology]
[0003] In some cases, scene reconstruction is necessary, for example, for scenes outside the unmanned control room. The results of scene reconstruction are of crucial value for scene reproduction and data annotation. Currently, achieving scene reconstruction often requires reliance on laser radar. [Overview of the project] [Problems that the invention aims to solve]
[0004] Current scene reconstruction methods rely on laser radar, which makes them relatively expensive to implement.
[0005] This disclosure provides a scene reconstruction method, apparatus, medium, and electronic equipment to solve the above problems and reduce the cost of realizing scene reconstruction. [Means for solving the problem]
[0006] A scene reconstruction method according to one embodiment of the embodiments of this disclosure is: The method includes the steps of: performing background segmentation on a first image sequence collected by a camera mounted on a movable device to obtain a second image sequence; performing sparse reconstruction of a point cloud using the second image sequence to obtain a first reconstructed point cloud and a first camera position and orientation sequence of the camera; performing dynamic static segmentation on the first image sequence to obtain a third image sequence corresponding to a static scene and a fourth image sequence corresponding to a dynamic scene; performing dense reconstruction of a point cloud using the first reconstructed point cloud, the first camera position and orientation sequence, and the third image sequence to obtain a second reconstructed point cloud; sequentially performing sparse and dense reconstruction of a point cloud using the first camera position and orientation sequence and the fourth image sequence to obtain a third reconstructed point cloud; and determining a scene reconstruction result corresponding to the first image sequence based on the second and third reconstructed point clouds.
[0007] A scene reconstruction device according to another embodiment of the embodiments of this disclosure is: A first segmentation module for performing background segmentation on a first image sequence collected by a camera provided in a movable device to obtain a second image sequence, and using the second image sequence obtained by the first segmentation module to perform sparse reconstruction of a point cloud to obtain a first reconstructed point cloud and a first camera position and orientation sequence; a second segmentation module for performing dynamic / static segmentation on the first image sequence to obtain a third image sequence corresponding to a static scene and a fourth image sequence corresponding to a dynamic scene; a second reconstruction module for performing dense reconstruction of the point cloud using the first reconstructed point cloud and the first camera position and orientation sequence obtained by the first reconstruction module and the third image sequence obtained by the second segmentation module to obtain a second reconstructed point cloud; a third reconstruction module for sequentially performing sparse reconstruction and dense reconstruction of the point cloud using the first camera position and orientation sequence obtained by the first reconstruction module and the fourth image sequence obtained by the second segmentation module to obtain a third reconstructed point cloud; and a determination module for determining a scene reconstruction result corresponding to the first image sequence based on the second reconstructed point cloud obtained by the second reconstruction module and the third reconstructed point cloud obtained by the third reconstruction module.
[0008] A computer-readable storage medium according to yet another aspect of the embodiments of the present disclosure stores a computer program for executing the above scene reconstruction method.
[0009] An electronic device according to yet another aspect of the embodiments of the present disclosure includes a processor, a memory for storing instructions executable by the processor, and by the processor reading and executing the executable instructions from the memory, the above scene reconstruction method is realized.
[0010] Another computer program product according to another aspect of the embodiments of the present disclosure, when the instructions in this computer program product are executed by a processor, the above scene reconstruction method is executed.
Advantages of the Invention
[0011] According to the scene reconstruction method, apparatus, medium, and electronic device according to the above embodiments of the present disclosure, background segmentation is performed on the multi-viewpoint images (i.e., the first image sequence) collected by the camera, and the sparse reconstruction of the point cloud is performed using the second image sequence obtained by the background segmentation to obtain the first reconstructed point cloud and the first camera position and orientation sequence. In addition, dynamic and static segmentation is performed on the multi-viewpoint images to obtain a third image sequence corresponding to the static scene and a fourth image sequence corresponding to the dynamic scene. By using the third image sequence together with the first reconstructed point cloud and the first camera position and orientation sequence for the dense reconstruction of the point cloud, a second reconstructed point cloud that effectively represents the static scene is obtained. By using the fourth image sequence together with the first camera position and orientation sequence for the sparse and dense reconstruction of the point cloud, a third reconstructed point cloud that effectively represents the dynamic scene is obtained. By combining the second reconstructed point cloud and the third reconstructed point cloud, the scene reconstruction result corresponding to the first image sequence can be determined. As can be seen from the above, in the embodiments of the present disclosure, scene reconstruction is realized through a series of processes on the multi-viewpoint images collected by the camera. Compared with lidar, the price of the camera is relatively low, so the implementation cost of scene reconstruction can be reduced. In addition, in the embodiments of the present disclosure, a policy of separately processing the static scene and the dynamic scene is adopted, that is, by reconstructing the static scene and the dynamic scene respectively, the difficulty of reconstruction can be reduced and the reconstruction speed can be improved.
Brief Description of the Drawings
[0012] [Figure 1-1] It is a schematic diagram of the motion model of the static scene. [Figure 1-2] It is a schematic diagram of the motion model of the dynamic scene. [Figure 2]This is a flowchart of a scene reconstruction method according to one exemplary embodiment of the present disclosure. [Figure 3] This is flowchart 1 of a dynamic-static partitioning method according to an exemplary embodiment of the present disclosure. [Figure 4] This is flowchart 2 of a dynamic-static partitioning method according to one exemplary embodiment of the present disclosure. [Figure 5-1] This is flowchart 3 of a dynamic-static partitioning method according to an exemplary embodiment of the present disclosure. [Figure 5-2] This is flowchart 4 of a dynamic-static partitioning method according to an exemplary embodiment of the present disclosure. [Figure 6] This is a flowchart of a dynamic scene reconstruction method according to one exemplary embodiment of the present disclosure. [Figure 7-1] This is a flowchart of an object position and orientation sequence correction method according to one exemplary embodiment of the present disclosure. [Figure 7-2] This is a flowchart of an object position and orientation sequence correction method according to another exemplary embodiment of the present disclosure. [Figure 8] This is a flowchart of a dynamic scene density reconstruction method according to one exemplary embodiment of the present disclosure. [Figure 9] This is a flowchart of a dense reconstruction method for a static scene according to one exemplary embodiment of the present disclosure. [Figure 10] This is a flowchart of a dense reconstruction method for a static scene according to another exemplary embodiment of the present disclosure. [Figure 11] This is a flowchart of a dense reconstruction method for a static scene according to yet another exemplary embodiment of the present disclosure. [Figure 12] This is a flowchart of a scene reconstruction method relating to another exemplary embodiment of the present disclosure. [Figure 13] This is a schematic diagram of the structure of a scene reconstruction device according to one exemplary embodiment of the present disclosure. [Figure 14] This is a schematic diagram of the structure of a scene reconstruction device according to another exemplary embodiment of the present disclosure. [Figure 15-1]This is a schematic diagram of the structure of a scene reconstruction device according to yet another exemplary embodiment of the present disclosure. [Figure 15-2] This is a schematic diagram of the structure of a scene reconstruction device according to yet another exemplary embodiment of the present disclosure. [Figure 16] This is a schematic diagram of the structure of a scene reconstruction device according to yet another exemplary embodiment of the present disclosure. [Figure 17] This is a schematic diagram of the structure of a scene reconstruction device according to yet another exemplary embodiment of the present disclosure. [Figure 18] This is a schematic diagram of the structure of a scene reconstruction device according to yet another exemplary embodiment of the present disclosure. [Figure 19] This is a structural diagram of an electronic device according to an exemplary embodiment of the present application. [Modes for carrying out the invention]
[0013] Hereinafter, exemplary embodiments of the Disclosure will be described in detail with reference to the drawings in order to interpret the Disclosure, and the embodiments described are not all embodiments but merely a part of the embodiments of the Disclosure, and the Disclosure is not limited to exemplary embodiments.
[0014] Unless otherwise specifically stated, the relative deployment of components and steps, formulas, and numerical values described in these embodiments do not limit the scope of this disclosure.
[0015] [Summary of the application] Laser radar (LiDAR) is a radar system that detects characteristic quantities such as the position, velocity, direction, attitude, and shape of a target by irradiating it with a laser beam.
[0016] Laser radar can be applied to scene reconstruction, for example, to the reconstruction of scenes outside the unmanned operating cabin. However, because laser radar is expensive, it results in a relatively high cost of implementing scene reconstruction. How to reduce the cost of implementing scene reconstruction is a matter of interest to those skilled in the art.
[0017] [Example System] Later in this text, we will discuss various models, including optical flow estimation models, depth estimation models, instance partitioning models, and motion models for static and dynamic scenes. For the sake of clarity, we will first briefly introduce these models.
[0018] The optical flow estimation model, depth estimation model, and instance partitioning model can all be convolutional neural network (CNN) models. The optical flow estimation model can be provided with a single frame of image as input, and it can estimate the optical flow value of each pixel point in that frame of image, where the optical flow value of any pixel point can indicate the pixel position in the next frame of image corresponding to the pixel position where that pixel point is located. The depth estimation model can be provided with a single frame of image as input, and it can estimate the depth value of each pixel point in that frame of image, where the depth value of any pixel point can determine the spatial position of that pixel point in the camera coordinate system. The instance partitioning model can be provided with a single frame of image as input, and it can estimate the category of each pixel point in that frame of image and position different instances in that frame of image.
[0019] A motion model for a static scene can be a mathematical model for describing the laws of motion of objects in a static scene, and the motion model for a static scene can be seen in Figure 1-1. A motion model for a dynamic scene can be a mathematical model for describing the laws of motion of objects in a dynamic scene, and the laws of motion for a dynamic scene can be seen in Figure 1-2. As shown in Figure 1-1, the camera is in motion, and the optical center of the camera at time t-1, time t, and time t+1 are O, respectively. t-1 , Ot and O t+1 is such that the object (e.g., a lane, a building, etc.) is stationary, and the k-th point on the object is the starting and ending point P k represented by, and in the images collected by the camera at time t - 1, time t, and time t + 1, the projection points of the k-th point on the object are respectively P t-1 k , P t k , P t+1 k respectively. The main difference between FIG. 1-2 and FIG. 1-1 is that both the camera and the object (e.g., a vehicle, a pedestrian, etc.) are in motion, and the k-th point on the object at time t - 1, time t, and time t + 1 are respectively P k t-1 , P k t , P k <00,00016>represented by, and the poses of the object at time t - 1, time t, and time t + 1 are respectively T t-1 k , T t k , T t+1 k represented by.
[0020] [Exemplary method] FIG. 2 is a flowchart of a scene reconstruction method according to an exemplary embodiment of the present disclosure. The method shown in FIG. 2 can include steps 210, step 220, step 23, step 240, step 250, and step 260, which will be described below for each step respectively.
[0021] In step 210, background segmentation is performed on the first image sequence collected by a camera provided on a movable device to obtain a second image sequence.
[0022] Optionally, the movable device includes, but is not limited to, vehicles, trains, ships, etc., and the camera provided on the movable device includes, but is not limited to, a front camera, a rear camera, a side camera, etc.
[0023] A camera mounted on a mobile device can collect images at a constant frame rate to obtain a first image sequence containing multiple frames of collected images. In the embodiments of this disclosure, for ease of understanding, the case in which the multiple frames of collected images in the first image sequence consist of N frames of collected images will be described as an example.
[0024] For each frame of the N frames of collected images included in the first image sequence, the instance partitioning model can be used to partition the collected image and determine the estimated category of each pixel point in the collected image. This allows us to determine which areas of the collected image contain objects that may be in motion (e.g., vehicles, pedestrians, etc.), generate masks over these areas, and thereby remove objects that may be in motion from the collected image, leaving only stationary objects (e.g., the ground, walls, buildings, trees, etc.) and separating the background image from the collected image.
[0025] The method described in the paragraph above allows for obtaining N frames of background images that correspond one-to-one with the collected N frames of images, and by arranging the N frames of background images in chronological order, a second image sequence can be formed.
[0026] In summary, by using the instance partitioning model, background partitioning of the first image sequence is achieved, and the second image sequence is obtained. In some cases, by using the semantic partitioning model, foreground and background separation of the collected images for each frame in the first image sequence can also be achieved, and the second image sequence can be obtained.
[0027] In an optional example, step 210 may be performed by the processor calling a corresponding instruction stored in memory, or by a first partitioned module operated by the processor.
[0028] In step 220, the second image sequence is used to perform a sparse reconstruction of the point cloud to obtain the first reconstructed point cloud and the first camera position and posture sequence.
[0029] For each background image in the N frames of background images included in the second image sequence, keypoint detection can be performed on the background image using a keypoint detection algorithm. Here, the keypoint detection algorithm includes, but is not limited to, the scale-invariant feature transform (SIFT) algorithm and deep learning-based algorithms.
[0030] If the keypoints of each of the N frames of background images are known, keypoint matching and screening can be performed, and in combination with the operation of the Incremental Motion Recovery Structure (SFM) algorithm, the camera position and orientation of the camera at the acquisition time corresponding to each of the N frames of acquired images can be determined, thereby obtaining N camera position and orientations. These N camera position and orientations can then be arranged in chronological order to form a first camera position and orientation sequence, and multiple three-dimensional points in the world coordinate system can be determined to construct a first reconstructed point cloud.
[0031] Naturally, the algorithms on which the first reconstructed point cloud and the first camera position and orientation sequence are generated are not limited to incremental SFM algorithms, but may also be global SFM algorithms, hybrid SFM algorithms, etc., and will not be listed individually here.
[0032] In an optional example, step 220 may be performed by the processor calling a corresponding instruction stored in memory, or by a first reconstruction module operated by the processor.
[0033] In step 230, dynamic static segmentation is performed on the first image sequence to obtain a third image sequence corresponding to the static scene and a fourth image sequence corresponding to the dynamic scene.
[0034] For each frame of the N frames of collected images included in the first image sequence, it is possible to determine which objects in the collected image are definitely moving objects and which are not. This allows the collected image to be divided into a segmented image containing only moving objects and a segmented image containing only non-moving objects, thereby achieving dynamic-static segmentation of the collected image. The segmented image containing only moving objects can be considered as a segmented image corresponding to a dynamic scene. The segmented image containing only non-moving objects can be considered as a segmented image corresponding to a static scene.
[0035] Using the method described in the paragraph above, dynamic static segmentation can be performed on all N frames of collected images to obtain N-frame segmented images corresponding to static scenes and N-frame segmented images corresponding to dynamic scenes. The N-frame segmented images corresponding to static scenes can be arranged in chronological order to form a third image sequence, and the N-frame segmented images corresponding to dynamic scenes can be arranged in chronological order to form a fourth image sequence.
[0036] In an optional example, step 230 may be performed by the processor calling a corresponding instruction stored in memory, or by a second partitioned module operated by the processor.
[0037] In step 240, the first reconstructed point cloud, the first camera position and orientation sequence, and the third image sequence are used to perform density reconstruction of the point cloud and obtain a second reconstructed point cloud.
[0038] If the first reconstructed point cloud, the first camera position and orientation sequence, and the third image sequence are known, the Multi-View Stereo (MVS) algorithm, when combined with the motion model shown in Figure 1-1, searches for pixel points with matching luminosity in different images in the third image sequence. Stereo matching is then performed to determine multiple points in the world coordinate system and construct a second reconstructed point cloud. The second reconstructed point cloud is denser than the first reconstructed point cloud and can effectively represent static scenes.
[0039] In an optional example, step 240 may be performed by the processor calling a corresponding instruction stored in memory, or by a second reconstruction module operated by the processor.
[0040] In step 250, the first camera position and orientation sequence and the fourth image sequence are used to sequentially perform sparse and dense reconstruction of the point cloud to obtain a third reconstructed point cloud.
[0041] If the first camera position and orientation sequence and the fourth image sequence are known, a third reconstructed point cloud can be obtained by first performing a sparse reconstruction using the SFM algorithm, followed by a dense reconstruction using the MVS algorithm. This third reconstructed point cloud can effectively represent the dynamic scene.
[0042] In an optional example, step 250 may be performed by the processor calling a corresponding instruction stored in memory, or by a third reconstruction module operated by the processor.
[0043] In step 260, the scene reconstruction result corresponding to the first image sequence is determined based on the second and third reconstructed point clouds.
[0044] Optionally, the second and third reconstructed point clouds can be modeled to obtain a first three-dimensional model corresponding to the second reconstructed point cloud and a second three-dimensional model corresponding to the third reconstructed point cloud. Subsequently, the first and second three-dimensional models can be texture-mapped, and the scene reconstruction result corresponding to the first image sequence can include the texture-mapped first three-dimensional model and the texture-mapped second three-dimensional model. In some embodiments, the scene reconstruction result corresponding to the first image sequence can include the motion trajectory of a movable device, the motion trajectory of a moving object around the movable device, and so on.
[0045] In an optional example, step 260 may be performed by the processor calling a corresponding instruction stored in memory, or by a decision module executed by the processor.
[0046] In the embodiments of this disclosure, background segmentation can be performed on multi-view images (i.e., a first image sequence) collected by a camera, and sparse reconstruction of the point cloud can be performed using the second image sequence obtained by background segmentation to obtain a first reconstructed point cloud and a first camera position and orientation sequence. Furthermore, dynamic static segmentation can be performed on the multi-view images to obtain a third image sequence corresponding to a static scene and a fourth image sequence corresponding to a dynamic scene. By using the third image sequence together with the first reconstructed point cloud and the first camera position and orientation sequence for dense reconstruction of the point cloud, a second reconstructed point cloud that effectively represents the static scene can be obtained. By using the fourth image sequence together with the first camera position and orientation sequence for sparse and dense reconstruction of the point cloud, a third reconstructed point cloud that effectively represents the dynamic scene can be obtained. By combining the second and third reconstructed point clouds, the scene reconstruction result corresponding to the first image sequence can be determined. As can be seen from the above, in the embodiments of this disclosure, scene reconstruction is achieved by a series of processes on multi-view images collected by the camera, and because the price of the camera is very low compared to laser radar, the cost of achieving scene reconstruction can be reduced. Furthermore, in the embodiments of this disclosure, a policy is adopted to process static scenes and dynamic scenes separately, that is, by reconstructing static scenes and dynamic scenes respectively, the difficulty of reconstruction can be reduced and the reconstruction speed can be improved.
[0047] Figure 3 is a flowchart of a dynamic-static partitioning method according to an exemplary embodiment of the present disclosure. The method shown in Figure 3 may include steps 310, 320, and 330, each of which will be described below.
[0048] In step 310, optical flow estimation is performed on the first image sequence using the optical flow estimation model, and the optical flow estimation results are obtained.
[0049] For each frame of the N frames of collected images included in the first image sequence, the collected image can be provided as input to the optical flow estimation model. The optical flow estimation model can perform calculations based on the collected image to obtain the optical flow value for each pixel point in the collected image. The optical flow values for all pixel points in the collected image can constitute the optical flow estimation data corresponding to the collected image.
[0050] The method described in the paragraph above allows for the acquisition of optical flow estimation data corresponding to each of the N frames of collected images, and this optical flow estimation data corresponding to each of the N frames of collected images can constitute the optical flow estimation result corresponding to the first image sequence.
[0051] In an optional example, step 310 may be performed by the processor calling a corresponding instruction stored in memory, or by a first estimated module operated by the processor.
[0052] In step 320, the depth estimation model is used to perform depth estimation on the first image sequence and obtain the depth estimation result.
[0053] For each frame of the N frames of collected images included in the first image sequence, the collected image can be provided as input to a depth estimation model. The depth estimation model can perform calculations based on the collected image to obtain the depth value of each pixel point in the collected image. The depth values of all pixel points in the collected image can constitute the depth estimation data corresponding to the collected image. Optionally, the depth estimation data corresponding to the collected image may be in the form of a depth map.
[0054] The method described in the paragraph above allows for obtaining depth estimation data corresponding to each of the N frames of collected images, and the depth estimation data corresponding to each of the N frames of collected images can constitute the depth estimation result corresponding to the first image sequence.
[0055] In an optional example, step 320 may be performed by the processor calling a corresponding instruction stored in memory, or by a second estimated module operated by the processor.
[0056] In step 330, dynamic static segmentation is performed on the first image sequence based on the optical flow estimation result, depth estimation result, and the first camera position and orientation sequence to obtain a third image sequence corresponding to a static scene and a fourth image sequence corresponding to a dynamic scene.
[0057] In an optional example, step 330 may be performed by the processor calling a corresponding instruction stored in memory, or by a second partitioned module operated by the processor.
[0058] Optionally, step 330 may be an optional embodiment of step 230 in the embodiment shown in Figure 2.
[0059] In some optional embodiments of this disclosure, as shown in Figure 4, step 330 may include steps 3301, 3303, and 3305.
[0060] In step 3301, based on the optical flow estimation results, a first pixel position correspondence between images in adjacent frames in the first image sequence is determined.
[0061] If we assume that the N acquired images in the first image sequence are sequentially represented as image 1, image 2, image 3, ..., image N, and that the optical flow value corresponding to pixel point 1 in image 1 indicates that the pixel position of pixel point 2 in image 2 corresponds to the pixel position of pixel point 1 in image 1, then the first pixel position correspondence can include the correspondence between the pixel position of pixel point 1 and the pixel position of pixel point 2. In this manner, the first pixel position correspondence can include multiple pairs of matching. pixels This can include the correspondence between the pixel positions of points, where pixel point 1 and pixel point 2 are a pair of matching pixels It is possible to construct a point.
[0062] In an optional example, step 3301 may be performed by calling a corresponding instruction stored in memory by the processor, or by a first decision submodule operated by the processor.
[0063] In step 3303, a second pixel position correspondence between images of adjacent frames in the first image sequence is determined based on the depth estimation result and the first camera position and orientation sequence.
[0064] The first camera position and orientation sequence may include a camera position and orientation corresponding to image 1 and a camera position and orientation corresponding to image 2. Using the camera position and orientation corresponding to image 1 and the camera position and orientation corresponding to image 2, the camera's translation matrix and rotation matrix from the acquisition time corresponding to image 1 to the acquisition time corresponding to image 2 can be calculated.
[0065] Referring to the motion model shown in Figure 1-1, if the depth value corresponding to pixel point 1 and the camera's translation and rotation matrices from the acquisition time corresponding to image 1 to the acquisition time corresponding to image 2 are known, then by combining them with pre-calibrated camera intrinsic parameters, the pixel point 2' corresponding to pixel point 1 in image 2 can be determined by the following equation. u(P t k )=K*[Rt c *K -1 *P t k *D t (P t k )+t t c ) Here, u(P t k ) represents the pixel position of pixel point 2', K represents the pre-calibrated camera intrinsic parameters and is in matrix form, and R t c This represents the camera rotation matrix from the acquisition time corresponding to image 1 to the acquisition time corresponding to image 2, and K -1 This represents the inverse matrix of K, and P t k This represents the pixel position of pixel point 1, and D t (P t k ) represents the depth value of pixel point 1 in the depth estimation result, and t t c This represents the camera's translation matrix from the acquisition time corresponding to image 1 to the acquisition time corresponding to image 2.
[0066] The above equation can be understood as follows: Pixel point 1 in image 2 is transformed from the pixel coordinate system to the camera coordinate system, then from the camera coordinate system to the world coordinate system to obtain spatial point 1 corresponding to pixel point 1 in the world coordinate system; then spatial point 1 is transformed from the world coordinate system to the camera coordinate system, and then from the camera coordinate system to the pixel coordinate system to determine pixel point 2' corresponding to pixel point 1 in image 2. The second pixel position correspondence relationship can include the correspondence relationship between the pixel position of pixel point 1 and the pixel position of pixel point 2'. In this manner, the second pixel position correspondence relationship can be a matching of multiple pairs. pixels This can include the correspondence between the pixel positions of points, where pixel point 1 and pixel point 2' are a pair of matching pixels It is possible to construct a point.
[0067] In an optional example, step 3303 may be performed by calling a corresponding instruction stored in memory by the processor, or by a second decision submodule operated by the processor.
[0068] In step 3305, dynamic static partitioning is performed on the first image sequence based on the first pixel position correspondence and the second pixel position correspondence to obtain a third image sequence corresponding to a static scene and a fourth image sequence corresponding to a dynamic scene.
[0069] In an optional example, step 3305 may be performed by the processor calling a corresponding instruction stored in memory, or by a divided submodule operated by the processor.
[0070] In some optional embodiments of the present disclosure, as shown in Figure 5-1, step 3305 includes steps 33051a, 33052a, 33053a, 33054a, 33055a, and 33056a.
[0071] In step 33051a, for a first target pixel point in the first target image in the first image sequence, the first target pixel position corresponding to the pixel position of the first target pixel point is determined from the second target image using the first pixel position correspondence relationship, where the second target image is the image of the frame immediately following the first target image in the first image sequence.
[0072] In step 33052a, the second pixel position Based on the correspondence, the second target pixel position corresponding to the pixel position of the first target pixel point is determined from the second target image.
[0073] The first target image can be any frame in the first image sequence other than the last frame, and the first target pixel point can be any pixel point in the first target image.
[0074] Assuming that the first target image is Image 1, the second target image is Image 2, and the first target pixel point is pixel point 1, then the first target pixel position determined from the second target image by the first pixel position correspondence can be the pixel position of pixel point 2 in Image 2, and the second target pixel position determined from the second target image can be the pixel position of pixel point 2' in Image 2.
[0075] In step 33053a, the pixel point attribute of the first target pixel point is determined based on the positional error between the first target pixel position and the second target pixel position, where the pixel point attribute of any pixel point represents whether or not that pixel point belongs to a dynamic object.
[0076] Assuming that the first target pixel position is represented by f(p) and the second target pixel position is represented by u(p), the positional error between the first and second target pixel positions can be expressed as |f(p)-u(p)|. For the sake of explanation, the positional error between the first and second target pixel positions will be simply referred to as the pixel position error. Based on the pixel position error and a preset position error, the pixel point attributes of the first target pixel point can be determined.
[0077] Optionally, if the pixel position error is greater than a preset position error, the pixel attribute of the first target pixel point can be determined as a first-class attribute indicating that the first target pixel point belongs to a dynamic object; if the pixel position error is less than or equal to the preset position error, the pixel attribute of the first target pixel point can be determined as a second-class attribute indicating that the first target pixel point does not belong to a dynamic object (i.e., the first target pixel point belongs to a static object).
[0078] Naturally, the method for determining pixel point attributes is not limited to this. For example, the difference between the pixel position error and a preset position error can be calculated, and the ratio of the absolute value of the difference to the preset position error can be calculated. If the calculated ratio is greater than the preset ratio, the pixel point attribute of the first target pixel point can be determined as a first-class attribute indicating that the first target pixel point belongs to a dynamic object. If the calculated ratio is less than or equal to the preset ratio, the pixel point attribute of the first target pixel point can be determined as a second-class attribute indicating that the first target pixel point does not belong to a dynamic object (i.e., the first target pixel point belongs to a static object).
[0079] In step 33054a, dynamic static segmentation is performed on the first target image based on the pixel point attributes of the first target pixel point to obtain a first segmented image corresponding to a static scene and a second segmented image corresponding to a dynamic scene.
[0080] The pixel point attribute determination method described above allows us to obtain the respective pixel point attributes of each pixel point in the first target image. By generating a mask in the region where pixel points with attribute type 1 are distributed in the first target image, dynamic objects can be removed from the first target image, thereby dividing the first segmented image corresponding to the static scene. Similarly, by generating a mask in the region where pixel points with attribute type 2 are distributed in the first target image, static objects can be removed from the first target image, thereby dividing the second segmented image corresponding to the dynamic scene.
[0081] In step 33055a, a third image sequence is determined based on the first segmented image.
[0082] In step 33056a, the fourth image sequence is determined based on the second segmented image.
[0083] The above describes the dynamic static segmentation method for the first target image. Each frame in the first image sequence can be dynamically segmented using the same method, thereby obtaining N segmented images corresponding to static scenes and N segmented images corresponding to dynamic scenes. The N segmented images corresponding to static scenes can be arranged chronologically to form a third image sequence, and the N segmented images corresponding to dynamic scenes can be arranged chronologically to form a fourth image sequence.
[0084] In some embodiments, after obtaining N-frame segmented images corresponding to a static scene and N-frame segmented images corresponding to a dynamic scene, these segmented images can first be corrected (for example, filtered) using a certain algorithm, and then a third image sequence and a fourth image sequence can be determined based on these corrected segmented images.
[0085] In an optional example, steps 33051a to 33056a above may be performed by the processor calling corresponding instructions stored in memory, or by corresponding units in a divided submodule operated by the processor.
[0086] In such an embodiment, the first pixel position Correspondence and second pixel position By comparing the positional errors between two target pixel positions obtained based on the correspondence, it is possible to determine which pixel points in the first target image belong to dynamic objects and which belong to static objects, using the positional errors as reference information. Then, dynamic-static segmentation can be performed efficiently and reliably on the first target image, thereby obtaining the third and fourth image sequences for use in subsequent steps.
[0087] In some optional embodiments of the present disclosure, the method according to the embodiment of the present disclosure further includes step 3304, as shown in Figure 5-2.
[0088] In step 3304, the first image sequence is subjected to instance partitioning using the instance partitioning model, and the instance partitioning results are obtained.
[0089] For each frame of the N frames of collected images included in the first image sequence, the collected image can be provided as input to an instance partitioning model. The instance partitioning model can perform calculations based on the collected image to obtain the estimated category of each pixel point in the collected image, and the instance region of each instance (which may be a static object or a dynamic object) in the collected image. This information can be used to construct instance partitioning data corresponding to the collected image.
[0090] In an optional example, step 3304 may be performed by the processor calling a corresponding instruction stored in memory, or by a third partitioned module operated by the processor.
[0091] The method described in the paragraph above allows for obtaining instance partitioning data corresponding to each of the N frames of collected images, and the instance partitioning data corresponding to each of the N frames of collected images can constitute the instance partitioning result corresponding to the first image sequence.
[0092] Step 3305 includes steps 33051b, 33052b, 33053b, 33054b, 33055b, and 33056b.
[0093] In an optional example, step 3305 may be performed by the processor calling a corresponding instruction stored in memory, or by a divided submodule operated by the processor.
[0094] In step 33051b, the instance region of each instance in the first target image is determined based on the instance partitioning result for the first target image in the first image sequence.
[0095] Optionally, instance partitioning data corresponding to the first target image can be directly extracted from the instance partitioning results, and the respective instance regions of each instance in the first target image can be positioned based on the instance partitioning data corresponding to the first target image.
[0096] In step 33052b, based on the first and second pixel position correspondences, the respective pixel point attributes of each pixel point in the first target image are determined, where the pixel point attribute of any pixel point represents whether or not that pixel point belongs to a dynamic object.
[0097] The method for determining the respective pixel point attributes of each pixel point in the first target image based on the first and second pixel position correspondence relationships can be found in the related explanation in the embodiment shown in Figure 5-1, and redundant explanations are omitted here.
[0098] In step 33053b, for each instance region in the first target image, the pixel point attributes of each pixel point in that instance region are statistically analyzed to obtain the statistical results.
[0099] Optionally, for each instance region in the first target image, the first quantity and first occupancy ratio of pixel points whose pixel point attribute is of type 1, and the second quantity and second occupancy ratio of pixel points whose pixel point attribute is of type 2 can be statistically calculated. The first quantity, first occupancy ratio, second quantity, and second occupancy ratio corresponding to each instance region in the first target image can constitute the statistical result.
[0100] In some cases, the statistical results may include only the first and second quantities corresponding to each instance region in the first target image, or only the first and second occupancy ratios corresponding to each instance region in the first target image.
[0101] In step 33054b, dynamic static segmentation is performed on the first target image based on the statistical results corresponding to each instance region in the first target image to obtain a third segmented image corresponding to a static scene and a fourth segmented image corresponding to a dynamic scene.
[0102] In some optional embodiments of this disclosure, step 33054b is: The method includes the steps of: determining the occupancy ratio of pixel points in each instance region of the first target image based on statistical results corresponding to the instance region, where the pixel point attribute in the instance region represents that the current pixel point does not belong to a dynamic object; and performing dynamic-static subdivision on the first target image based on the occupancy ratio corresponding to each instance region of the first target image to obtain a third subdivision image corresponding to a static scene and a fourth subdivision image corresponding to a dynamic scene.
[0103] Assuming that the statistical results include a second occupancy ratio corresponding to each instance region in the first target image, the second occupancy ratio corresponding to any instance region can be considered as the occupancy ratio of pixel points in that instance region where the pixel point attribute represents that the current pixel point does not belong to a dynamic object. In this way, dynamic static partitioning can be performed on the target image directly based on the second occupancy ratio corresponding to each instance region in the first target image and a preset occupancy ratio.
[0104] Assuming that the statistical results do not include the second occupancy ratio corresponding to each instance region in the first target image, but do include the first occupancy ratio corresponding to each instance region in the first target image, the second occupancy ratio corresponding to each instance region can be obtained by calculating the difference between a preset value of 1 and the first occupancy ratio corresponding to that instance region for each instance region. Subsequently, dynamic static partitioning is performed on the target image based on the second occupancy ratio corresponding to each instance region in the first target image and the preset occupancy ratio.
[0105] The pre-set occupancy ratios can be arbitrarily chosen, such as 70%, 75%, or 80%, and we will not list them all here.
[0106] If the second occupancy ratio corresponding to any instance region in the first target image is greater than a preset occupancy ratio, it can be determined that the instance region belongs to a static object. If the second occupancy ratio corresponding to the instance region is less than or equal to the preset occupancy ratio, it can be determined that the instance region belongs to a dynamic object. By generating a mask for each instance region belonging to a dynamic object in the first target image, the dynamic object in the first target image can be removed, thereby dividing the third segmented image corresponding to the static scene. Similarly, by generating a mask for each instance region belonging to a static object in the first target image, the static object in the first target image can be removed, thereby dividing the fourth segmented image corresponding to the dynamic scene.
[0107] The above paragraph describes the case where dynamic static partitioning is performed by referring to a second occupancy ratio corresponding to each instance region. However, in some cases, dynamic static partitioning may be performed by referring to a first occupancy ratio corresponding to each instance region. For example, for any instance region in the first target image, if the first occupancy ratio corresponding to that instance region is less than or equal to a preset occupancy ratio, it can be determined that the instance region belongs to a static object. If the first occupancy ratio corresponding to that instance region is greater than the preset occupancy ratio, it can be determined that the instance region belongs to a dynamic object. Based on this, a third partitioned image corresponding to the static scene and a fourth partitioned image corresponding to the dynamic scene can be created by generating a mask.
[0108] In step 33055b, a third image sequence is determined based on the third segmented image.
[0109] In step 33056b, the fourth image sequence is determined based on the fourth segmented image.
[0110] For specific embodiments of steps 33055b and 33056b, please refer to the related descriptions of steps 33055a and 33056a described above, and redundant explanations will be omitted here.
[0111] Furthermore, both the depth estimation model and the optical flow estimation model may have a certain degree of estimation error, and a certain degree of estimation error also exists in the acquisition process of the first camera position and orientation sequence. In order to reduce the impact of these errors on the scene reconstruction results, the instance division model is used to determine the instance region of each instance in the first target image, and statistical analysis of pixel point attributes is performed for each instance region. The statistical results are then used for the dynamic static division of the first target image. For example, dynamic static division can be performed by referring to the occupancy ratio of pixel points with type 2 attributes in each instance region. This is equivalent to introducing a certain constraint to the dynamic static division process (the motion state of each point in the same object is always the same, i.e., these points are all fixed and do not move, or they move together), which makes the dynamic static division results more robust, and the dynamic static division results can be used in subsequent steps to improve the effect of the final scene reconstruction.
[0112] In an optional example, steps 33051b to 33056b above may be performed by the processor calling corresponding instructions stored in memory, or by corresponding units in a divided submodule operated by the processor.
[0113] In the embodiments of this disclosure, the optical flow estimation model can be used to efficiently and reliably determine the optical flow estimation result corresponding to the first image, and the optical flow estimation result can be used to determine the first pixel position correspondence between images of adjacent frames. The depth estimation model can be used to efficiently and reliably determine the depth estimation result corresponding to the first image sequence, and the depth estimation result can be used together with the first camera position and orientation sequence to determine the second pixel position correspondence between images of adjacent frames. The first pixel position correspondence can be considered as an object tracking result obtained by one method, and the second pixel position correspondence can be considered as an object tracking result obtained by another method. The first pixel position correspondence and the second pixel position correspondence can be combined to determine the position error, and the determined position error can be considered as a residual stream. Using the residual stream as reference information, dynamic objects and static objects in the image can be distinguished relatively accurately, thereby improving the effect of dynamic-static partitioning of the first image sequence. In concrete implementation, several more complex tracking algorithms may be introduced, and the tracking results obtained by these algorithms may be used for dynamic static partitioning.
[0114] Figure 6 is a flowchart of a dynamic scene reconstruction method according to an exemplary embodiment of the present disclosure. The method shown in Figure 6 includes steps 610, 620, 630, 640, and 650. Optionally, the combination of steps 620 to 650 may constitute an optional embodiment of step 250.
[0115] In step 610, optical flow estimation is performed on the first image sequence using the optical flow estimation model, and the optical flow estimation results are obtained.
[0116] For specific embodiments of step 610, please refer to the related explanation of step 310 described above, and redundant explanations will be omitted here.
[0117] In an optional example, step 610 may be performed by the processor calling a corresponding instruction stored in memory, or by a first estimated module operated by the processor.
[0118] In step 620, based on the optical flow estimation results, a third pixel position correspondence between images in adjacent frames in the fourth image sequence is determined.
[0119] For specific embodiments of step 620, please refer to the related description in step 3301 above, and redundant explanations will be omitted here.
[0120] In an optional example, step 620 may be performed by calling a corresponding instruction stored in memory by the processor, or by a third decision submodule operated by the processor.
[0121] In step 630, the first camera position and orientation sequence, the fourth image sequence, and the third pixel position correspondence are used to perform a sparse reconstruction of the point cloud to obtain the fourth reconstructed point cloud and the object position and orientation sequence of the dynamic objects in the fourth image sequence.
[0122] If the first camera position / orientation sequence, the fourth image sequence, and the third pixel position correspondence are known, the SFM algorithm can be used to further combine with the motion model shown in Figure 1-2 to obtain the fourth reconstructed point cloud and the object position / orientation sequence of the dynamic object in the fourth image sequence, where the object position / orientation sequence may include the object position / orientation at each acquisition time of the N frames of acquired images of the dynamic object in the fourth image sequence.
[0123] Furthermore, referring to the motion model shown in Figure 1-1, the formula for determining pixel point 2' corresponding to pixel point 1 in image 2 is presented, and the basic principle of the formula is explained. Here, when sparse reconstruction of the point cloud is performed using the first camera position / orientation sequence, the fourth image sequence, and the third pixel position correspondence, multiple pairs of matching are performed. pixels It is also necessary to find points, and the main difference from the basic principles explained above is that multiple pairs of matching are involved. pixels It is necessary to consider the motion of objects when searching for a point.
[0124] Assuming that pixel point 1 in image 1 belongs to moving object c, pixel point 1 can be a point where the features of moving object c are relatively prominent, for example, a point with a relatively large gradient. Here, first, we can obtain spatial point 1 corresponding to pixel point 1 in the world coordinate system at the acquisition time corresponding to image 1 by first converting pixel point 1 in image 1 from the pixel coordinate system to the camera coordinate system, and then from the camera coordinate system to the world coordinate system. Next, by referring to the translation matrix of moving object c from the acquisition time corresponding to image 1 to the acquisition time corresponding to image 2, we can obtain spatial point 1'. After that, we can obtain spatial point 1' by first converting spatial point 1' from the world coordinate system to the camera coordinate system, and then from the camera coordinate system to the pixel coordinate system to determine pixel point 2'' corresponding to pixel point 1 in image 2, and pixel point 1 and pixel point 2'' are a pair matching. pixels It is possible to construct a point.
[0125] In an optional example, step 630 may be performed by the processor calling a corresponding instruction stored in memory, or by a first reconstructed submodule operated by the processor.
[0126] In step 640, the object position and orientation sequence is corrected using a pre-set correction method.
[0127] In an optional example, step 640 may be performed by the processor calling a corresponding instruction stored in memory, or by a compensatory submodule operated by the processor.
[0128] In some optional embodiments of the present disclosure, the method according to the embodiment of the present disclosure further includes step 632, as shown in Figure 7-1.
[0129] In step 632, depth estimation is performed on the first image sequence using the depth estimation model, and the depth estimation result is obtained.
[0130] For specific embodiments of step 632, please refer to the related explanation of step 320 described above, and redundant explanations will be omitted here.
[0131] In an optional example, step 632 may be performed by the processor calling a corresponding instruction stored in memory, or by a second estimated module operated by the processor.
[0132] Step 640 includes steps 6401, 6402, and 6403.
[0133] In step 6401, based on the depth estimation result, the first camera position and orientation sequence, and the object position and orientation sequence, the fourth pixel position correspondence between images of adjacent frames in the fourth image sequence is determined.
[0134] The fourth method for determining the pixel position correspondence is similar to the second method for determining the pixel position correspondence described above, with the main differences being as follows. The determination of the fourth pixel position correspondence requires referring to the motion model shown in Figure 1-2 and considering the motion of the object in the determination process. Specifically, one can refer to the introduction of the method for determining pixel point 2'' corresponding to pixel point 1 in image 2 described above. Therefore, it is necessary to operate an object position and orientation sequence. The determination of the second pixel position correspondence requires referring to the motion model shown in Figure 1-1 and considering the motion of the object in the determination process.
[0135] In step 6402, the reprojection error is determined based on the third pixel position correspondence and the fourth pixel position correspondence.
[0136] Furthermore, the above describes a method for determining the position error for a first target pixel point based on the first and second pixel position correspondence relationships. Here, a similar method can be used to determine the position error for each of multiple target pixel points based on the third and fourth pixel position correspondence relationships. Subsequently, the reprojection error can be determined based on the position errors for each of the multiple target pixel points. For example, the average value of the position errors for each of the multiple target pixel points can be directly calculated, and the calculated average value can be used as the reprojection error. Alternatively, for example, position errors with significantly abnormal values (e.g., position errors that are too large or too small) can be removed from the position errors for each of the multiple target pixel points, the average value of the remaining position errors can be calculated, and the calculated average value can be used as the reprojection error.
[0137] In step 6403, the object position and orientation sequence is corrected by minimizing the reprojection error.
[0138] The reprojection error can be minimized by an optional optimization algorithm, which may be a linear optimization algorithm or a nonlinear optimization algorithm, for example, a linear least squares method or a nonlinear least squares method.
[0139] Calculating the reprojection error requires the use of a fourth pixel position correspondence, and determining the fourth pixel position correspondence requires the use of an object position and orientation sequence. In other words, the object position and orientation sequence and the reprojection error are interrelated. Thus, by minimizing the reprojection error, the object position and orientation sequence can be optimized, and the corrected object position and orientation sequence can be used in subsequent point cloud reconstruction to improve the point cloud reconstruction effect.
[0140] In an optional example, steps 6401 to 6403 above may be performed by the processor calling corresponding instructions stored in memory, or by corresponding units in a correction submodule operated by the processor.
[0141] In some other optional embodiments of the present disclosure, the method according to the embodiment of the present disclosure further includes step 634, as shown in Figure 7-2.
[0142] In step 634, the first image sequence is subjected to instance partitioning using the instance partitioning model, and the instance partitioning results are obtained.
[0143] For specific embodiments of step 634, please refer to the related explanation in step 320 above, and redundant explanations will be omitted here.
[0144] In an optional example, step 634 may be performed by the processor calling a corresponding instruction stored in memory, or by a third partitioned module operated by the processor.
[0145] Step 640 includes steps 6404, 6405, 6406, and 6407.
[0146] In step 6404, the motion trajectory of the dynamic object in the fourth image sequence is determined based on the object position and orientation sequence.
[0147] For any dynamic object in the fourth image sequence, the object position sequence of that dynamic object can be extracted from the object position and orientation sequence, and each position in the extracted object position sequence can constitute the motion trajectory of that dynamic object. Alternatively, after extracting the object position sequence, positions that are significantly abnormal in the extracted object position sequence can be corrected, and each position in the corrected object position sequence can constitute the motion trajectory of that dynamic object.
[0148] In step 6405, the first estimated category of the dynamic object in the fourth image sequence is determined based on the instance partitioning results.
[0149] Optionally, an estimated category of any dynamic object in the fourth image sequence can be extracted from the instance partitioning results, and this estimated category may be the first estimated category.
[0150] In step 6406, a pre-defined trajectory corresponding to the first estimated category is determined.
[0151] The correspondence between categories and trajectories can be optionally set in advance. In this way, the trajectory corresponding to the first estimated category can be efficiently and reliably determined based on the pre-set correspondence, and the determined trajectory can be the pre-set trajectory in step 6406.
[0152] In step 6407, the object position and orientation sequence is corrected by minimizing the trajectory error between the motion trajectory and the preset trajectory.
[0153] The trajectory error between the motion trajectory and a preset trajectory can be calculated at will, and the trajectory error can be minimized using the specified optimization algorithm described above.
[0154] In an optional example, steps 6404 to 6407 above may be performed by the processor calling corresponding instructions stored in memory, or by corresponding units in a correction submodule operated by the processor.
[0155] Calculating trajectory errors requires the use of motion trajectories, and determining motion trajectories requires the use of object position and orientation sequences. In other words, object position and orientation sequences and trajectory errors are interrelated. Thus, minimizing trajectory errors can optimize object position and orientation sequences, and the corrected object position and orientation sequences can be used in subsequent point cloud reconstruction to improve the effectiveness of point cloud reconstruction.
[0156] In step 650, the point cloud is densely reconstructed using the first camera position / orientation sequence, the fourth image sequence, the third pixel position correspondence, the fourth reconstructed point cloud, and the corrected object position / orientation sequence to obtain the third reconstructed point cloud.
[0157] If the first camera position / orientation sequence, the fourth image sequence, the third pixel position correspondence, the fourth reconstructed point cloud, and the corrected object position / orientation sequence are known, then by combining them with the motion model shown in Figure 1-2 and applying the MVS algorithm, dense reconstruction of the point cloud can be achieved, and the third reconstructed point cloud can be obtained.
[0158] In an optional example, step 650 may be performed by the processor calling a corresponding instruction stored in memory, or by a second reconstruction submodule operated by the processor.
[0159] In the embodiments of this disclosure, by operating an optical flow estimation model, optical flow estimation results corresponding to the first image can be obtained efficiently and reliably. The optical flow estimation results are used to determine a third pixel position correspondence between images of adjacent frames. The third pixel position correspondence is used together with the first camera position / pose sequence and the fourth image sequence for sparse reconstruction of the dynamic scene, thereby obtaining a fourth reconstructed point cloud and an object position / pose sequence. Subsequently, the object position / pose sequence is corrected using a certain method, and the corrected object position / pose sequence can be used together with the first camera position / pose sequence, the fourth image sequence, the third pixel position correspondence, and the fourth reconstructed point cloud for dense reconstruction of the dynamic scene. In this way, the accuracy and reliability of the data on which the dense reconstruction of the dynamic scene depends can be improved, thereby improving the final scene reconstruction effect.
[0160] Figure 8 is a flowchart of a dynamic scene density reconstruction method according to an exemplary embodiment of the present disclosure. The method shown in Figure 8 includes steps 810, 820, 830, 840, and 850. Optionally, the combination of steps 820 to 840 may be an optional embodiment of step 630, and step 850 may be an optional embodiment of step 650.
[0161] In step 810, the first image sequence is subjected to instance partitioning using the instance partitioning model, and the instance partitioning results are obtained.
[0162] For specific embodiments of step 810, please refer to the related explanation in step 3304 described above, and redundant explanations will be omitted here.
[0163] In step 820, based on the instance division results and the third pixel position correspondence relationship, it is determined whether the estimated categories of the corresponding pixel positions in adjacent frames of the fourth image sequence match, and the determination result is obtained.
[0164] For a pixel position x1 in any image other than the last frame in the fourth image sequence, the pixel position x2 corresponding to pixel position x1 in the next frame of that image is determined based on the third pixel position correspondence relationship. Based on the instance division result, the estimated category y1 for pixel position x1 and the estimated category y2 for pixel position x2 are determined, and the determination result can be obtained by comparing whether the estimated category y1 and estimated category y2 match.
[0165] In step 830, the third pixel position correspondence is corrected based on the determination result.
[0166] If pixel positions x1 and x2 belong to the same moving object, theoretically, the estimated categories for pixel positions x1 and x2 should match. In light of this, if the determination result indicates that estimated categories y1 and y2 do not match, it can be determined that pixel positions x1 and x2 do not substantially belong to the same moving object and that there is no substantial correspondence between them. Therefore, the third pixel position correspondence can be corrected by either removing the correspondence between pixel positions x1 and x2 from the third pixel position correspondence or by adding an invalid flag to the correspondence between pixel positions x1 and x2 in the third pixel position correspondence.
[0167] In step 840, a sparse reconstruction of the point cloud is performed using the first camera position / orientation sequence, the fourth image sequence, and the corrected third pixel position correspondence to obtain the fourth reconstructed point cloud and the object position / orientation sequence of the dynamic objects in the fourth image sequence.
[0168] The specific embodiment of step 840 is similar to the specific embodiment of step 630 described above, the main difference being that the sparse reconstruction effect of the dynamic scene is improved by using the corrected third pixel position correspondence in step 840.
[0169] In step 850, the point cloud is densely reconstructed using the first camera position / orientation sequence, the fourth image sequence, the corrected third pixel position correspondence, the fourth reconstructed point cloud, and the corrected object position / orientation sequence to obtain the third reconstructed point cloud.
[0170] The specific embodiment of step 850 is similar to the specific embodiment of step 650 described above, the main difference being that the density reconstruction effect of the dynamic scene is improved by using the corrected third pixel position correspondence in step 850.
[0171] In an optional example, steps 810 to 850 above may be performed by the processor calling corresponding instructions stored in memory, or by corresponding modules operated by the processor.
[0172] In the embodiment of the present invention, the instance division results of the first image sequence can be efficiently and reliably obtained by operating an instance division model, and the third pixel position correspondence can be corrected by organically utilizing the instance division results and the third pixel position correspondence. By using the corrected third pixel position correspondence for sparse and dense reconstruction of the point cloud, the reconstruction effect of the dynamic scene can be improved.
[0173] Figure 9 is a flowchart of a dense reconstruction method for a static scene according to an exemplary embodiment of the present disclosure. The method shown in Figure 9 includes steps 910, 920, 930, 940, 950, and 960. Optionally, the combination of steps 920 to 960 may be an optional embodiment of step 240.
[0174] In step 910, the first image sequence is subjected to instance partitioning using the instance partitioning model, and the instance partitioning results are obtained.
[0175] For specific embodiments of step 910, please refer to the related explanation in step 3304 described above, and redundant explanations will be omitted here.
[0176] In step 920, for the second target pixel point in the third target image in the third image sequence, the second estimated category of the second target pixel point is determined based on the instance division results.
[0177] The third target image can be any of the images included in the third image sequence, and the second target pixel point can be any of the pixel points in the third target image.
[0178] Optionally, the estimated category of the second target pixel point can be directly extracted from the instance division result, and the extracted estimated category can be used as the second estimated category.
[0179] In step 930, the local region size matching the second estimated category is determined.
[0180] The correspondence between categories and window sizes can be optionally pre-set. Based on this pre-set correspondence, the window size corresponding to the second estimated category can be determined, and this window size can be used as the local region size that matches the second estimated category.
[0181] In step 940, the target local region is determined from the third target image based on the pixel position of the second target pixel point and the local region size.
[0182] Optionally, the local region size may include a width size and a height size, in which case the target local region can be obtained by cutting out a region from the third target image centered on the pixel position of the second target pixel point and having the said width size and height size. In some cases, the target local region can also be obtained by cutting out a region from the third target image with the pixel position of the second target pixel point as the upper left corner point and having the said width size and height size.
[0183] In step 950, a dense reconstruction of the point cloud is performed on the target local region using the first reconstructed point cloud, the first camera position and orientation sequence, and other images in the third image sequence other than the third target image, to obtain a local point cloud corresponding to the target local region.
[0184] After determining the target local region, the spatial region corresponding to the target local region in the world coordinate system can be found by using the first reconstructed point cloud, the first camera position and orientation sequence, and pre-calibrated camera intrinsic parameters, and by referring to the motion model shown in Figure 1-1. This spatial region can be a rectangular region. Subsequently, spatial points can be sampled within this spatial region to obtain multiple spatial points; for example, M spatial points can be obtained.
[0185] For each of the M spatial points, the pixel point corresponding to that spatial point in the third image sequence, other than the third target image, is determined using the first camera position / attitude sequence and pre-calibrated camera intrinsic parameters, and further by referring to the motion model shown in Figure 1-1. The pixel position of this pixel point can be represented by pixel position x3. Because the spatial region and the target local region are related, the pixel point corresponding to that spatial point in the target local region can also be calculated. The pixel position of this pixel point can be represented by pixel position x4. By matching between images, the pixel position corresponding to pixel position x4 in other images can also be determined. This pixel position can be represented by pixel position x5. Subsequently, the corresponding position error is obtained by subtracting (deducting) pixel position x5 and pixel position x4. The reprojection error is obtained by calculating the average value based on the position error corresponding to each of the M spatial points. Based on the obtained reprojection error, the normal vector of the spatial region and the depth values of the M spatial points are optimized. The M spatial points with optimized depth values can constitute a local point group corresponding to the target local region.
[0186] In step 960, a second reconstructed point cloud is determined based on the local point cloud.
[0187] The above describes a method for reconstructing a local point cloud for a target local region corresponding to a second target pixel point. Using this method, a local point cloud can be reconstructed for multiple pixel points in the third image sequence, and these local point clouds can constitute a second reconstructed point cloud. Alternatively, after obtaining these local point clouds, a certain method can be employed to optimize these reconstructed point clouds, and these optimized reconstructed point clouds can constitute a second reconstructed point cloud.
[0188] In an optional example, steps 910 to 960 above may be performed by the processor calling corresponding instructions stored in memory, or by corresponding modules operated by the processor.
[0189] In the embodiments of this disclosure, the instance partitioning model can be used to efficiently and reliably obtain the instance partitioning results of the first image sequence. The instance partitioning results determine the second estimated category of the second target pixel point in the third target image in the third image sequence. By referring to the second estimated category, an appropriate local region size is determined. Based on the local region size, a corresponding local region is determined, and this local region can be used to reconstruct the local point cloud corresponding to the second target pixel point. For example, if the second target pixel point is located in a low-texture area such as the ground, the local point cloud can be reconstructed using a large window, thereby reconstructing more scene detail. If the second target pixel point is located in a high-texture area, the local point cloud can be reconstructed using a small window, thereby improving the reconstruction accuracy. Therefore, the embodiments of this disclosure can improve the dense reconstruction effect of static scenes.
[0190] Figure 10 is a flowchart of a static scene density reconstruction method according to another exemplary embodiment of the present disclosure. The method shown in Figure 10 includes steps 1010, 1020, and 1030. Optionally, the combination of steps 1020 and 1030 may be an optional embodiment of step 240.
[0191] In step 1010, optical flow estimation is performed on the first image sequence using the optical flow estimation model, and the optical flow estimation result is obtained.
[0192] For specific embodiments of step 1010, please refer to the related explanation of step 310 described above, and redundant explanations will be omitted here.
[0193] In step 1020, based on the optical flow estimation results, a fifth pixel position correspondence between images in adjacent frames in the third image sequence is determined.
[0194] For specific embodiments of step 1020, please refer to the related explanation in step 3301 described above, and redundant explanations will be omitted here.
[0195] In step 1030, the point cloud is densely reconstructed using the first reconstructed point cloud, the first camera position and orientation sequence, the third image sequence, and the fifth pixel position correspondence to obtain the second reconstructed point cloud.
[0196] In an optional example, steps 1010 to 1030 above may be performed by the processor calling corresponding instructions stored in memory, or by corresponding modules operated by the processor.
[0197] Furthermore, the specific embodiment of step 1030 is similar to the specific embodiment of step 240 described above, the main difference being that the density reconstruction in step 1031 further utilizes a fifth pixel position correspondence, the fifth pixel position correspondence can be considered as an object tracking result obtained by a certain method, and the guidance of this object tracking result makes it possible to clarify which pixel points in adjacent frames correspond to the same spatial point, thereby providing an effective reference during the reconstruction process and improving the density reconstruction effect of static scenes.
[0198] Figure 11 is a flowchart of a static scene density reconstruction method according to yet another exemplary embodiment of the present disclosure. The method shown in Figure 11 includes steps 1110, 1120, and 1130. Optionally, the combination of steps 1110 to 1130 may be another optional embodiment of step 240.
[0199] In step 1110, a second camera position and orientation sequence for the camera is determined based on data collected by non-visual sensors installed on the mobile device.
[0200] Optionally, non-visual sensors include, but are not limited to, position sensors, inertial measurement units, and wheel speedometers, where position sensors include, but are not limited to, position sensors based on the Global Positioning System (GPS) and position sensors based on the Global Navigation Satellite System (GNSS).
[0201] Furthermore, for sensors located in the chassis system of mobile devices that are included in the non-visual sensors, the data collected by these sensors can be obtained from the Controller Area Network (CAN) bus.
[0202] Non-visual sensors are used for position sensors and inertial measurement. unit Assuming that the system includes the following, the camera position at each acquisition time of the N frames of collected images can be determined based on the position data collected by the position sensor, and the camera orientation at each acquisition time of the N frames of collected images can be determined based on the inertial measurement data collected by the inertial measurement unit. By integrating these camera positions and camera position orientations, a second camera position orientation sequence can be obtained.
[0203] In step 1120, a similarity transformation alignment is performed on the first camera position and orientation sequence, using the second camera position and orientation sequence as the alignment reference.
[0204] Furthermore, a similarity transformation alignment can be performed on the first camera position and orientation sequence by employing any feasible similarity transformation alignment algorithm.
[0205] In step 1130, a dense reconstruction of the point cloud is performed using the first reconstructed point cloud, the similar-transformed and aligned first camera position and orientation sequence, and the third image sequence to obtain a second reconstructed point cloud.
[0206] When performing sparse reconstruction using the SFM algorithm to obtain the first camera position and orientation sequence, each position and orientation in the first camera position and orientation sequence has an undetermined scale, i.e., it does not have a true scale. Therefore, by using data collected by non-visual sensors to determine a second camera position and orientation sequence that has a true scale, and by using a similarity transformation alignment algorithm to align the first camera position and orientation sequence with the second camera position and orientation sequence, the similarity transformation aligned first camera position and orientation sequence can have a true scale, and the similarity transformation aligned first camera position and orientation sequence can be used for dense reconstruction of a static scene to improve the dense reconstruction effect of the static scene.
[0207] In an optional example, steps 1110 to 1130 above may be performed by the processor calling corresponding instructions stored in memory, or by corresponding modules operated by the processor.
[0208] In some optional embodiments of the present disclosure, Figure 12 is a flowchart of a scene reconstruction method according to another exemplary embodiment of the present disclosure, where a forward monocular image can be acquired by a forward camera mounted on a movable device, and the flowchart of the scene reconstruction method of the embodiment of the present disclosure may include the following steps 1210 to 1290.
[0209] In step 1210, the first image sequence is determined. That is, the first image sequence can be obtained based on the forward monocular image.
[0210] In step 1220, the first image sequence is processed by a large-scale visual perception model including an instance partitioning model, a depth estimation model, and an optical flow estimation model to obtain instance partitioning results, depth estimation results, and optical flow estimation results.
[0211] In step 1230, the second image sequence is obtained by performing background segmentation on the first image sequence.
[0212] In step 1240, sparse reconstruction is performed using the second image sequence to obtain the sparse reconstruction result of the static scene. The sparse reconstruction result of the static scene may include the first reconstructed point cloud and the first camera position and orientation sequence.
[0213] In step 1250, data collected by non-visual sensors installed on a movable device is used to perform a similarity transformation alignment on the first camera position and orientation sequence to obtain the similarity transformation aligned first camera position and orientation sequence.
[0214] In step 1260, dynamic static partitioning is performed on the first image sequence to obtain a third image sequence corresponding to a static scene and a fourth image sequence corresponding to a dynamic scene. For example, based on the instance partitioning results, optical flow estimation results, and depth estimation results obtained in step 1220, dynamic static partitioning can be performed on the first image sequence to obtain a third image sequence corresponding to a static scene and a fourth image sequence corresponding to a dynamic scene. Specifically, the aforementioned related embodiments (for example, specific extended embodiments of each step such as steps 3301, 3303, 3304, 3305, and steps 33051a to 33056a, and steps 33051b to 33056b) can be referenced.
[0215] In step 1270, dense reconstruction is performed on the static scene. That is, based on the similarity-transformed first camera position and orientation sequence, the first reconstructed point cloud, and the third image sequence, dense reconstruction is performed on the static scene to obtain the second reconstructed point cloud.
[0216] In step 1280, sparse reconstruction is performed on the dynamic scene. That is, sparse reconstruction can be performed on the dynamic scene first, based on the fourth image sequence.
[0217] In step 1290, dense reconstruction is performed on the dynamic scene to finally obtain the scene reconstruction result corresponding to the first image sequence. That is, the scene reconstruction result corresponding to the first image sequence can be obtained from the third reconstructed point cloud obtained by sequentially performing sparse reconstruction and dense reconstruction on the dynamic scene, and the second reconstructed point cloud obtained in step 1270.
[0218] Based on the above, in the embodiments of this disclosure, scene reconstruction is achieved by employing a visual form, and specifically, scene reconstruction can be performed by employing a method of dividing the process from simple to difficult. For example, by first reconstructing a static scene, and then reconstructing a dynamic scene after the reconstruction of the static scene is completed, reconstruction costs and difficulty can be reduced, and a highly accurate, dense point cloud can be obtained.
[0219] Any scene reconstruction method according to the embodiments of this disclosure can be executed by any suitable device having data processing capabilities, including but not limited to terminal devices and servers. Alternatively, any scene reconstruction method according to the embodiments of the present invention can be executed by a processor, for example, by calling a corresponding instruction stored in memory, which can execute any scene reconstruction method referred to in the embodiments of this disclosure. Repetitive explanations are omitted below.
[0220] As those skilled in the art will understand, all or some of the steps of the embodiments of the above method can be completed by a program instructing the relevant hardware, the program can be stored in a computer-readable storage medium, and when the program is executed, it performs the steps including the embodiments of the above method, the storage medium can include various media capable of storing program code, such as ROM, RAM, magnetic disks or optical disks.
[0221] [Example device] Figure 13 is a schematic diagram of the structure of a scene reconstruction apparatus according to an exemplary embodiment of the present disclosure. The apparatus shown in Figure 12 includes a first splitting module 1310, a first reconstruction module 1320, a second splitting module 1330, a second reconstruction module 1340, a third reconstruction module 1350, and a determination module 1360.
[0222] The first segmentation module 1310 performs background segmentation on the first image sequence collected by a camera mounted on a mobile device to acquire a second image sequence. The first reconstruction module 1320 performs sparse reconstruction of the point cloud using the second image sequence acquired by the first segmentation module 1310 to acquire the first reconstructed point cloud and the first camera position and orientation sequence. The second division module 1330 performs dynamic-static division on the first image sequence to obtain a third image sequence corresponding to a static scene and a fourth image sequence corresponding to a dynamic scene. The second reconstruction module 1340 performs dense reconstruction of the point cloud using the first reconstructed point cloud and the first camera position and orientation sequence acquired by the first reconstruction module 1320, and the third image sequence acquired by the second division module 1330, to acquire a second reconstructed point cloud. The third reconstruction module 1350 sequentially performs sparse reconstruction and dense reconstruction of the point cloud using the first camera position and orientation sequence acquired by the first reconstruction module 1320 and the fourth image sequence acquired by the second segmentation module 1330 to acquire the third reconstructed point cloud. The decision module 1360 determines the scene reconstruction result corresponding to the first image sequence based on the second reconstructed point cloud acquired by the second reconstruction module 1340 and the third reconstructed point cloud acquired by the third reconstruction module 1350.
[0223] In some optional examples, as shown in Figure 14, the scene reconstruction apparatus according to the embodiment of the present disclosure is An optical flow estimation model is used to perform optical flow estimation on the first image sequence, and a first estimation module 1370 is used to obtain the optical flow estimation result. The system further includes a second estimation module 1380 for performing depth estimation on a first image sequence using a depth estimation model and obtaining depth estimation results. Specifically, the second division module 1330 performs dynamic static division on the first image sequence based on the optical flow estimation results obtained by the first estimation module 1370, the depth estimation results obtained by the second estimation module 1380, and the first camera position and orientation sequence obtained by the first reconstruction module 1320, thereby obtaining a third image sequence corresponding to a static scene and a fourth image sequence corresponding to a dynamic scene.
[0224] In some optional examples, as shown in Figure 14, the second segmented module 1330 is: Based on the optical flow estimation results obtained by the first estimation module 1370, a first determination submodule 13301 is used to determine the first pixel position correspondence between images of adjacent frames in the first image sequence, A second determination submodule 13303 for determining a second pixel position correspondence between images of adjacent frames in a first image sequence, based on the depth estimation results obtained by the second estimation module 1380 and the first camera position and orientation sequence obtained by the first reconstruction module 1320, The system includes a division submodule 13305 for performing dynamic static division on a first image sequence based on a first pixel position correspondence determined by a first determination submodule 13301 and a second pixel position correspondence determined by a second determination submodule 13303, in order to obtain a third image sequence corresponding to a static scene and a fourth image sequence corresponding to a dynamic scene.
[0225] In some optional examples, the split submodule 13305 is, A first determination unit for determining a first target pixel position corresponding to the pixel position of a first target pixel point in a first target image in a first image sequence, based on a first pixel position correspondence relationship determined by a first determination submodule 13301, wherein the second target image is the image of the frame immediately following the first target image in the first image sequence. The second pixel determined by the second determination submodule 13303 position A second determination unit for determining the second target pixel position corresponding to the pixel position of the first target pixel point from the second target image based on the point correspondence relationship, A third decision unit for determining the pixel point attributes of a first target pixel point based on the positional error between a first target pixel position determined by a first decision unit and a second target pixel position determined by a second decision unit, wherein the pixel point attributes of any pixel point include a third decision unit that expresses whether or not the pixel point belongs to a dynamic object, A first division unit performs dynamic static division on the first target image based on the pixel point attributes of the first target pixel point determined by the third decision unit, in order to obtain a first division image corresponding to a static scene and a second division image corresponding to a dynamic scene. A fourth determination unit for determining a third image sequence based on the first segmented image acquired by the first segmentation unit, The system includes a fifth determination unit for determining a fourth image sequence based on a second segmented image acquired by a first segmentation unit.
[0226] In an optional example, as shown in Figure 14, the scene reconstruction apparatus according to the embodiment of the present disclosure is The system further includes a third partitioning module 1390 for performing instance partitioning on a first image sequence using an instance partitioning model and obtaining the instance partitioning results. The split submodule 13305 is, A fifth decision unit for determining the respective instance regions of each instance in the first target image based on the instance partitioning results obtained by the third partitioning module 1390 for the first target image in the first image sequence, A sixth determination unit for determining the respective pixel point attributes of each pixel point in a first target image, based on the first pixel position correspondence determined by the first determination submodule 13301 and the second pixel position correspondence determined by the second determination submodule 13303, wherein the pixel point attribute of any pixel point includes a sixth determination unit that expresses whether or not the pixel point belongs to a dynamic object, A statistical unit for obtaining statistical results by statistically analyzing the pixel point attributes of each pixel point in the instance region determined by the sixth decision unit for each instance region in the first target image, A second division unit performs dynamic static division on the first target image based on the statistical results corresponding to each instance region in the first target image obtained by the statistical unit, in order to obtain a third division image corresponding to a static scene and a fourth division image corresponding to a dynamic scene. A seventh decision unit for determining a third image sequence based on the third segmented image acquired by the second segmentation unit, The system includes an eighth determination unit for determining a fourth image sequence based on a fourth segmented image acquired by a second segmentation unit.
[0227] In some optional examples, the second division unit is: For each instance region in the first target image, a determination subunit for determining the occupancy ratio of pixel points in the instance region that represent the pixel point attribute in that instance region not belonging to a dynamic object, based on the statistical results corresponding to that instance region obtained by the statistical unit, The system includes a division subunit for performing dynamic static division on the first target image based on the occupancy ratio corresponding to each instance region in the first target image determined by the decision subunit, in order to obtain a third division image corresponding to a static scene and a fourth division image corresponding to a dynamic scene.
[0228] In some optional examples, as shown in Figures 15-1 and 15-2, the scene reconstruction apparatus according to the embodiment of the present disclosure is The system further includes a first estimation module 1370 for performing optical flow estimation on a first image sequence using an optical flow estimation model and obtaining optical flow estimation results. The third reconstruction module 1350 is, Based on the optical flow estimation results obtained by the first estimation module 1370, a third determination submodule 13501 is used to determine the third pixel position correspondence between images of adjacent frames in the fourth image sequence obtained by the division submodule 13305, A first reconstruction submodule 13503 performs sparse reconstruction of the point cloud using the first camera position and orientation sequence acquired by the first reconstruction module 1320, the fourth image sequence acquired by the second division module 1330, and the third pixel position correspondence determined by the third determination submodule 13501, in order to acquire the fourth reconstructed point cloud and the object position and orientation sequence of dynamic objects in the fourth image sequence. A correction submodule 13507 for correcting the object position and orientation sequence acquired by the first reconstruction submodule 13503 using a pre-set correction method, The system includes a second reconstruction submodule 13509 for performing dense reconstruction of the point cloud using a first camera position and orientation sequence acquired by a first reconstruction module 1320, a fourth image sequence acquired by a second division module 1330, a third pixel position correspondence determined by a third determination submodule 13501, a fourth reconstructed point cloud acquired by a first reconstruction submodule 13503, and an object position and orientation sequence corrected by a correction submodule 13507, in order to acquire a third reconstructed point cloud.
[0229] In some optional examples, as shown in Figure 15-1, the scene reconstruction apparatus according to the embodiment of the present disclosure is The system further includes a second estimation module 1380 for performing depth estimation on a first image sequence using a depth estimation model and obtaining depth estimation results. The correction submodule 13507 is, A ninth determination unit for determining the fourth pixel position correspondence between images of adjacent frames in the fourth image sequence acquired by the second division module 1330, based on the depth estimation result acquired by the second estimation module 1380, the first camera position and orientation sequence acquired by the first reconstruction module 1320, and the object sequence acquired by the first reconstruction submodule 13503, A tenth determination unit for determining the reprojection error based on the third pixel position correspondence determined by the third determination submodule 13501 and the fourth pixel position correspondence determined by the ninth determination unit, The system includes a first correction subunit for correcting the object position and orientation sequence acquired by the first reconstruction submodule 13503 by minimizing the reprojection error determined by the tenth decision unit.
[0230] In some optional examples, as shown in Figure 15-2, the scene reconstruction apparatus according to the embodiment of the present disclosure is The system further includes a third partitioning module 1390 for performing instance partitioning on a first image sequence using an instance partitioning model and obtaining the instance partitioning results. The correction submodule 13507 is, An eleventh determination unit for determining the motion trajectory of a dynamic object in a fourth image sequence acquired by a second division module 1330, based on the object position and orientation sequence acquired by the first reconstruction submodule 13503, A twelfth decision unit for determining a first estimated category of dynamic objects in a fourth image sequence acquired by the second division module 1330, based on the instance division results obtained by the third division module 1390, A 13th decision unit for determining a pre-set trajectory corresponding to the first estimated category determined by the 12th decision unit, The system includes a second correction subunit for correcting the object position and orientation sequence acquired by the first reconstruction submodule 13503 by minimizing the trajectory error between the motion trajectory determined by the 11th decision unit and the preset trajectory determined by the 13th decision unit.
[0231] In some optional examples, as shown in Figure 16, the scene reconstruction apparatus according to the embodiment of the present disclosure is The system further includes a third partitioning module 1390 for performing instance partitioning on a first image sequence using an instance partitioning model and obtaining the instance partitioning results. The first reconstruction submodule 13503 is, A 14th decision unit for obtaining a decision result, which determines whether the estimated categories of corresponding pixel positions in adjacent frames of the fourth image sequence obtained by the second division module 1330 match, based on the instance division result obtained by the third division module 1390 and the third pixel position correspondence relationship determined by the third decision submodule 13501, and for obtaining a decision result, A correction unit for correcting the third pixel position correspondence determined by the third decision submodule 13501 based on the decision result obtained by the 14th decision unit, The system includes a reconstruction unit for performing sparse reconstruction of a point cloud using a first camera position and orientation sequence acquired by a first reconstruction module 1320, a fourth image sequence acquired by a second division module 1330, and a third pixel position correspondence corrected by a correction unit, in order to acquire a fourth reconstructed point cloud and an object position and orientation sequence of dynamic objects in the fourth image sequence. The second reconstruction submodule 13509 specifically performs dense reconstruction of the point cloud using the first camera position / orientation sequence, the fourth image sequence, the corrected third pixel position correspondence, the fourth reconstructed point cloud, and the corrected object position / orientation sequence to obtain the third reconstructed point cloud.
[0232] In some optional examples, as shown in Figure 16, the scene reconstruction apparatus according to the embodiment of the present disclosure is The system further includes a third partitioning module 1390 for performing instance partitioning on the first image sequence using an instance partitioning model and obtaining the instance partitioning result. The second reconstruction module 1340 is, A fifth determination submodule 13401 determines the second estimated category of the second target pixel point in the third target image in the third image sequence acquired by the second division module 1330, based on the instance division results acquired by the third division module 1390, A sixth decision submodule 13403 for determining the local region size that matches the second estimated category determined by the fifth decision submodule 13401, Based on the pixel position of the second target pixel point and the local region size determined by the sixth determination submodule 13403, a seventh determination submodule 13405 is used to determine the target local region from the third target image, A third reconstruction submodule 13407 is used to perform dense reconstruction of the point cloud for the target local region determined by the seventh determination submodule 13405, using the first reconstructed point cloud acquired by the first reconstruction module 1320, the first camera position and orientation sequence acquired by the first reconstruction module 1320, and other images in the third image sequence acquired by the second division module 1330, excluding the third target image, in order to obtain a local point cloud corresponding to the target local region. The system includes an eighth determination submodule 13409 for determining a second reconstructed point cloud based on the local point cloud obtained by a third reconstruction submodule 13407.
[0233] In some optional examples, as shown in Figure 17, the scene reconstruction apparatus according to the embodiment of the present disclosure is The optical flow estimation model further includes a first estimation module 1370 for performing optical flow estimation on a first image sequence and obtaining the optical flow estimation result. The second reconstruction module 1340 is, Based on the optical flow estimation results obtained by the first estimation module 1370, a ninth determination submodule 13411 is used to determine the fifth pixel position correspondence between images of adjacent frames in the third image sequence obtained by the second division module 1330, A fourth reconstruction sub-module 13413 that performs dense reconstruction of the point cloud using the first reconstructed point cloud obtained by the first reconstruction module 1320, the first camera position and orientation sequence obtained by the first reconstruction module 1320, the third image sequence obtained by the second segmentation module 1330, and the fifth pixel position correspondence relationship determined by the ninth determination sub-module 13411, to obtain a second reconstructed point cloud, is included.
[0234] In some optional examples, as shown in FIG. 18, the second reconstruction module 1340 A tenth determination sub-module 13415 for determining a second camera position and orientation sequence of the camera based on data collected by a non-visual sensor provided on a movable device, A similarity transformation alignment sub-module 13417 for performing a similarity transformation alignment on the first camera position and orientation sequence obtained by the first reconstruction module 1320 with the second camera position and orientation sequence determined by the tenth determination sub-module 13415 as an alignment reference, A fifth reconstruction sub-module 13419 that performs dense reconstruction of the point cloud using the first reconstructed point cloud obtained by the first reconstruction module 1320, the similarity transformation aligned first camera position and orientation sequence obtained by the similarity transformation alignment sub-module 13417, and the third image sequence obtained by the second segmentation module 1330, to obtain a second reconstructed point cloud, is included.
[0235] In the apparatus of the embodiments of the present disclosure, each of the various optional examples, optional embodiments, and optional examples disclosed above can be flexibly selected and combined as needed, thereby realizing the corresponding functions and effects, and enumerating them one by one in the present disclosure is omitted here.
[0236] The beneficial technical effects corresponding to the exemplary embodiments of the present device can be referred to the corresponding beneficial technical effects of the above-mentioned exemplary method part, and here, the overlapping descriptions are omitted.
[0237] [Exemplary Electronic Device] FIG. 19 shows a block diagram of an electronic device according to an embodiment of the present disclosure. The electronic device 1900 includes one or more processors 1910 and a memory 1920.
[0238] The processor 1910 can be a central processing unit (CPU) or other forms of processing units having data processing capabilities and / or instruction execution capabilities, and can also control other components in the electronic device 1900 to execute desired functions.
[0239] The memory 1920 can include one or more computer program products, and the computer program products can include various forms of computer-readable storage media such as, for example, volatile memory and / or non-volatile memory. Volatile memory can include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory can include, for example, read-only memory (ROM), hard disk, flash memory, etc. The computer-readable storage medium can store one or more computer program instructions, and the processor 1910 can execute one or more computer program instructions to implement the methods according to each embodiment of the present disclosure above and / or other desired functions.
[0240] As an example, the electronic device 1900 can further include an input device 1930 and an output device 194 that are interconnected via a bus system and / or other forms of connection mechanisms (not shown).
[0241] This input device 1930 can further include, for example, a keyboard, a mouse, etc.
[0242] This output device 1940 can output various types of information to the outside. This output device 1940 may include, for example, a display, a speaker, a printer, a communication network, and remote output devices connected thereto.
[0243] For simplicity, Figure 19 shows only some of the components relevant to this disclosure in the electronic device 1900, omitting components such as buses and input / output interfaces. Beyond this, the electronic device 1900 may further include any other appropriate components depending on the specific application.
[0244] [Examples of computer program products and computer-readable storage media] Embodiments of this disclosure further provide computer program products, including computer program instructions, in addition to the methods and apparatus described above. When the computer program instructions are executed by a processor, the processor is caused to perform steps in the scene reconstruction methods relating to various embodiments of this disclosure as described in the “Exemplary Methods” portion of this specification.
[0245] Computer program products can be created using one or any combination of programming languages to produce program code for performing the operations of the embodiments of this disclosure, including object-oriented programming languages such as Java® and C++, and conventional procedural programming languages such as the C language or similar programming languages. The program code may run entirely on a user computing device, partially on a user device, run as a standalone software package, run partially on a user computing device and partially on a remote computing device, or run entirely on a remote computing device or a server.
[0246] Furthermore, embodiments of this disclosure further provide a computer-readable storage medium in which computer program instructions are stored. When the computer program instructions are executed by a processor, the processor is caused to perform steps in the scene reconstruction method relating to various embodiments of the disclosure described in the “Exemplary Methods” portion of this specification.
[0247] The computer-readable storage medium may employ any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable media may include, but are not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples (non-exclusive list) of readable storage media include electrical connections with one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, compact disk read-only memory (CD-ROM), optical memory elements, magnetic memory elements, or any suitable combination of the above.
[0248] While the basic principles of this disclosure have been explained above with reference to specific examples, the advantages, merits, and effects mentioned in this disclosure are merely illustrative and not limiting, and these advantages, merits, and effects are not necessarily present in every example of this disclosure. Furthermore, the specific details of the above disclosure are merely illustrative and easy-to-understand effects and are not limiting, and the above details do not necessarily limit this disclosure to being realized by the above specific details.
[0249] Those skilled in the art can make various modifications and alterations to this disclosure without departing from the spirit and scope of the present application. Thus, if such modifications and alterations of the present application fall within the claims of this disclosure and the equivalent art thereto, this disclosure also includes such modifications and alterations.
Claims
1. A scene reconstruction method in which each step is performed by a scene reconstruction device, A step of obtaining a second image sequence by performing background segmentation on a first image sequence collected by a camera mounted on a movable device, The steps include: performing sparse reconstruction of the point cloud using the second image sequence to obtain the first reconstructed point cloud and the first camera position and orientation sequence of the camera; The steps include performing dynamic static segmentation on the first image sequence to obtain a third image sequence corresponding to a static scene and a fourth image sequence corresponding to a dynamic scene, The steps include: performing dense reconstruction of the point cloud using the first reconstructed point cloud, the first camera position and orientation sequence, and the third image sequence to obtain a second reconstructed point cloud; The steps include sequentially performing sparse and dense reconstruction of the point cloud using the first camera position and orientation sequence and the fourth image sequence to obtain a third reconstructed point cloud, A scene reconstruction method characterized by comprising the step of determining a scene reconstruction result corresponding to the first image sequence based on the second reconstructed point cloud and the third reconstructed point cloud.
2. The aforementioned scene reconstruction method is: The steps include performing optical flow estimation on the first image sequence using an optical flow estimation model and obtaining the optical flow estimation result, The method further includes the step of performing depth estimation on the first image sequence using a depth estimation model and obtaining the depth estimation result, The step of performing dynamic static segmentation on the first image sequence to obtain a third image sequence corresponding to a static scene and a fourth image sequence corresponding to a dynamic scene is: The scene reconstruction method according to claim 1, characterized in that it includes the step of performing dynamic static segmentation on the first image sequence based on the optical flow estimation result, the depth estimation result, and the first camera position and orientation sequence to obtain a third image sequence corresponding to a static scene and a fourth image sequence corresponding to a dynamic scene.
3. The step of performing dynamic static segmentation on the first image sequence based on the optical flow estimation result, the depth estimation result, and the first camera position and orientation sequence to obtain a third image sequence corresponding to a static scene and a fourth image sequence corresponding to a dynamic scene is as follows: Based on the optical flow estimation results, the steps include determining a first pixel position correspondence between images of adjacent frames in the first image sequence, The steps include determining a second pixel position correspondence between images of adjacent frames in the first image sequence based on the depth estimation result and the first camera position and orientation sequence, The scene reconstruction method according to claim 2, comprising the step of performing dynamic static partitioning on the first image sequence based on the first pixel position correspondence relationship and the second pixel position correspondence relationship to obtain a third image sequence corresponding to a static scene and a fourth image sequence corresponding to a dynamic scene.
4. The step of performing dynamic static segmentation on the first image sequence based on the first pixel position correspondence relationship and the second pixel position correspondence relationship to obtain a third image sequence corresponding to a static scene and a fourth image sequence corresponding to a dynamic scene is: A step of determining a first target pixel position corresponding to the pixel position of a first target pixel point in a first target image in the first image sequence, from a second target image, based on the first pixel position correspondence relationship, wherein the second target image is the image of the frame immediately following the first target image in the first image sequence. The steps include determining a second target pixel position corresponding to the pixel position of the first target pixel point from the second target image based on the second pixel position correspondence relationship, A step of determining the pixel point attributes of the first target pixel point based on the positional error between the first target pixel position and the second target pixel position, wherein the pixel point attributes of any of the pixel points express whether or not the pixel point belongs to a dynamic object. The steps include: performing dynamic static segmentation on the first target image based on the pixel point attributes of the first target pixel point to obtain a first segmented image corresponding to a static scene and a second segmented image corresponding to a dynamic scene; The steps include determining the third image sequence based on the first segmented image, The scene reconstruction method according to claim 3, comprising the step of determining the fourth image sequence based on the second segmented image.
5. The aforementioned scene reconstruction method is: The method further includes the step of performing instance partitioning on the first image sequence using an instance partitioning model and obtaining the instance partitioning result, The step of performing dynamic static segmentation on the first image sequence based on the first pixel position correspondence relationship and the second pixel position correspondence relationship to obtain a third image sequence corresponding to a static scene and a fourth image sequence corresponding to a dynamic scene is: A step of determining the respective instance regions of each instance in the first target image based on the instance division result, with respect to the first target image in the first image sequence, A step of determining the respective pixel point attributes of each pixel point in the first target image based on the first pixel position correspondence relationship and the second pixel position correspondence relationship, wherein the pixel point attribute of any pixel point expresses whether or not that pixel point belongs to a dynamic object. The steps include: obtaining statistical results by statistically analyzing the pixel point attributes of each pixel point in each instance region of the first target image; The steps include: performing dynamic static segmentation on the first target image based on the statistical results corresponding to each instance region in the first target image to obtain a third segmented image corresponding to a static scene and a fourth segmented image corresponding to a dynamic scene; The steps include determining the third image sequence based on the third segmented image, The scene reconstruction method according to claim 3, comprising the step of determining the fourth image sequence based on the fourth segmented image.
6. The step of performing dynamic static segmentation on the first target image based on the statistical results corresponding to each instance region in the first target image to obtain a third segmented image corresponding to a static scene and a fourth segmented image corresponding to a dynamic scene is as follows: For each instance region in the first target image, the steps include determining the occupancy ratio of pixel points in the instance region that represent the pixel point attribute in the instance region that the current pixel point does not belong to a dynamic object, based on the statistical results corresponding to the instance region, The scene reconstruction method according to claim 5, comprising the step of performing dynamic static division on the first target image based on the occupancy ratio corresponding to each instance region in the first target image to obtain a third divided image corresponding to a static scene and a fourth divided image corresponding to a dynamic scene.
7. The aforementioned scene reconstruction method is: The process further includes the step of performing optical flow estimation on the first image sequence using an optical flow estimation model and obtaining the optical flow estimation result, The step of obtaining a third reconstructed point cloud by sequentially performing sparse reconstruction and dense reconstruction of the point cloud using the first camera position and orientation sequence and the fourth image sequence is as follows: Based on the optical flow estimation results, the third step of determining the pixel position correspondence between images of adjacent frames in the fourth image sequence, The steps include: performing sparse reconstruction of the point cloud using the first camera position / orientation sequence, the fourth image sequence, and the third pixel position correspondence to obtain the fourth reconstructed point cloud and the object position / orientation sequence of the dynamic object in the fourth image sequence; The steps include correcting the object position and orientation sequence using a pre-set correction method, The scene reconstruction method according to claim 1, comprising the step of performing a dense reconstruction of the point cloud using the first camera position and orientation sequence, the fourth image sequence, the third pixel position correspondence, the fourth reconstructed point cloud, and the corrected object position and orientation sequence to obtain the third reconstructed point cloud.
8. The aforementioned scene reconstruction method is: The process further includes the step of performing depth estimation on the first image sequence using a depth estimation model and obtaining the depth estimation result, The step of correcting the object position and orientation sequence using a pre-set correction method is: The steps include determining a fourth pixel position correspondence between images of adjacent frames in the fourth image sequence based on the depth estimation result, the first camera position and orientation sequence, and the object position and orientation sequence, The steps include determining the reprojection error based on the third pixel position correspondence and the fourth pixel position correspondence, The scene reconstruction method according to claim 7, comprising the step of correcting the object position and orientation sequence by minimizing the reprojection error.
9. The aforementioned scene reconstruction method is: The method further includes the step of performing instance partitioning on the first image sequence using an instance partitioning model and obtaining the instance partitioning result, The step of correcting the object position and orientation sequence using a pre-set correction method is: The steps include determining the motion trajectory of the dynamic object in the fourth image sequence based on the object position and orientation sequence, The steps include determining a first estimated category of the dynamic object in the fourth image sequence based on the instance partitioning result, The steps include determining a pre-set trajectory corresponding to the first estimated category, The scene reconstruction method according to claim 7, comprising the step of correcting the object position and orientation sequence by minimizing the trajectory error between the motion trajectory and the preset trajectory.
10. The aforementioned scene reconstruction method is: The method further includes the step of performing instance partitioning on the first image sequence using an instance partitioning model and obtaining the instance partitioning result, The step of performing sparse reconstruction of the point cloud using the first camera position / orientation sequence, the fourth image sequence, and the third pixel position correspondence relationship to obtain the fourth reconstructed point cloud and the object position / orientation sequence of the dynamic object in the fourth image sequence is as follows: Based on the instance division result and the third pixel position correspondence relationship, the steps include determining whether the estimated categories of corresponding pixel positions in adjacent frames of the fourth image sequence match and obtaining the determination result, The steps include correcting the third pixel position correspondence relationship based on the aforementioned determination result, The process includes the step of performing a sparse reconstruction of the point cloud using the first camera position and orientation sequence, the fourth image sequence, and the corrected third pixel position correspondence to obtain a fourth reconstructed point cloud and an object position and orientation sequence of dynamic objects in the fourth image sequence, The step of performing a dense reconstruction of the point cloud using the first camera position / orientation sequence, the fourth image sequence, the third pixel position correspondence, the fourth reconstructed point cloud, and the corrected object position / orientation sequence to obtain the third reconstructed point cloud is: The scene reconstruction method according to claim 7, characterized by including the step of performing a dense reconstruction of the point cloud using the first camera position and orientation sequence, the fourth image sequence, the corrected third pixel position correspondence, the fourth reconstructed point cloud, and the corrected object position and orientation sequence to obtain the third reconstructed point cloud.
11. The aforementioned scene reconstruction method is: The method further includes the step of performing instance partitioning on the first image sequence using an instance partitioning model and obtaining the instance partitioning result, The step of obtaining a second reconstructed point cloud by performing dense reconstruction of the point cloud using the first reconstructed point cloud, the first camera position and orientation sequence, and the third image sequence is as follows: The steps include determining a second estimated category for a second target pixel point in a third target image in the third image sequence based on the instance division result, The steps include determining the local region size that matches the second estimated category, A step of determining a target local region from the third target image based on the pixel position of the second target pixel point and the size of the local region, The steps include: performing a dense reconstruction of the point cloud for the target local region using the first reconstructed point cloud, the first camera position and orientation sequence, and other images in the third image sequence other than the third target image, to obtain a local point cloud corresponding to the target local region; The scene reconstruction method according to claim 1, characterized by comprising the step of determining the second reconstruction point cloud based on the local point cloud.
12. A first segmentation module for performing background segmentation on a first image sequence collected by a camera mounted on a mobile device to acquire a second image sequence, A first reconstruction module for performing sparse reconstruction of a point cloud using the second image sequence acquired by the first division module, and for acquiring a first reconstructed point cloud and a first camera position and orientation sequence, A second division module for performing dynamic static division on the first image sequence to obtain a third image sequence corresponding to a static scene and a fourth image sequence corresponding to a dynamic scene, A second reconstruction module for performing dense reconstruction of the point cloud using the first reconstructed point cloud and the first camera position and orientation sequence acquired by the first reconstruction module, and the third image sequence acquired by the second division module, in order to acquire a second reconstructed point cloud, A third reconstruction module for sequentially performing sparse reconstruction and dense reconstruction of a point cloud using the first camera position and orientation sequence acquired by the first reconstruction module and the fourth image sequence acquired by the second division module to acquire a third reconstructed point cloud, A scene reconstruction apparatus comprising: a determination module for determining a scene reconstruction result corresponding to the first image sequence based on the second reconstructed point cloud acquired by the second reconstruction module and the third reconstructed point cloud acquired by the third reconstruction module.
13. A computer-readable storage medium, A computer-readable storage medium characterized in that, when executed by a processor, the storage medium stores a computer program that causes the processor to execute the scene reconstruction method described in any one of claims 1 to 11.
14. An electronic device comprising a processor and a memory for storing instructions that the processor can execute, An electronic device characterized in that the processor reads and executes the executable instructions from the memory to realize the scene reconstruction method described in any one of claims 1 to 11.
Citation Information
Patent Citations
Combined scene reconstruction method and device based on model segmentation
CN107909643A
Dynamic scene three-dimensional reconstruction method based on semantic information assistance
CN114332394A
Image depth marking method and device, equipment and storage medium
CN114926485A
Static environment and dynamic object dense reconstruction method and system and storage medium
CN115578435A