Three-dimensional reconstruction method, device, computer equipment and storage medium
By using the signal data supervision of the associated image pairs in three-dimensional reconstruction, image pairs with appropriate signal distance are screened out, the problem of the impact of repeated scenes in large scenarios is solved, the accuracy and efficiency of three-dimensional reconstruction is improved, and the automatic positioning of signal sources is realized.
Patent Information
- Application Number
- CN202210589283.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-27
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2042-05-27
AI Technical Summary
Vision-based three-dimensional reconstruction is susceptible to repeated scenes when constructing large scenes, resulting in reconstruction errors, making it difficult to accurately distinguish scenes with similar structures.
By using the signal data supervision of the associated image pairs during the three-dimensional reconstruction process, image pairs with signal distances that meet preset conditions are selected to reduce the impact of repeated scenes, and image pairs are screened in combination with signal source identification and intensity information to improve the accuracy of three-dimensional reconstruction.
It effectively reduces the impact of repeated scenarios on three-dimensional reconstruction, improves the accuracy and efficiency of the three-dimensional model, enhances the robustness of visual three-dimensional reconstruction, and realizes the automatic positioning of signal sources.
Smart Images

Figure CN114972645B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer vision technology, and in particular to a three-dimensional reconstruction method, apparatus, computer equipment, and storage medium. Background Art
[0002] Vision-based 3D reconstruction takes a collection of input images and recovers the position and pose parameters of each image, as well as a point cloud representing the 3D structure of the captured scene. Vision-based 3D reconstruction is widely used in 3D digitization, high-precision map construction, and augmented reality. However, when constructing large scenes, vision-based 3D reconstruction is susceptible to repetitive scenes, leading to reconstruction errors. Summary of the Invention
[0003] The embodiments of the present disclosure at least provide a three-dimensional reconstruction method, apparatus, computer equipment, and storage medium.
[0004] In a first aspect, an embodiment of the present disclosure provides a three-dimensional reconstruction method, comprising:
[0005] Acquire multiple videos captured for a target scene, and signal data corresponding to each of the multiple videos; wherein the signal data is obtained by capturing a signal of a target signal source located within the target scene when the videos are captured; each video includes multiple frames of images;
[0006] determining a plurality of associated image pairs from the plurality of videos based on the signal data;
[0007] Based on the associated image pair, the target scene is three-dimensionally reconstructed to obtain a three-dimensional model of the target scene.
[0008] In this way, multiple associated image pairs can be determined through the signal data corresponding to multiple videos respectively; since the signal data corresponding to the two frames of images in the associated image pair can represent that the signal distance corresponding to the two frames of images meets the preset distance condition, when the target scene is three-dimensionally reconstructed based on the associated image pair, the signal distance between different images is used to supervise the three-dimensional reconstruction process. Therefore, in a target scene with a large number of repeated scenes, the influence of repeated scenes in different positions on the three-dimensional reconstruction effect can also be reduced, thereby reducing reconstruction errors in the three-dimensional reconstruction process.
[0009] In an optional embodiment,
[0010] The determining, based on the signal data, a plurality of associated image pairs from the plurality of videos comprises:
[0011] Extract feature data of each frame image in multiple videos;
[0012] Determining a plurality of candidate image pairs based on similarities between feature data of different images; each candidate image pair includes a first image and a second image;
[0013] The associated image pairs are screened from the candidate image pairs based on signal data corresponding to the first image and the second image in a plurality of candidate image pairs.
[0014] In this way, by initially selecting alternative image pairs based on image feature information, the efficiency of subsequent screening of related image pairs from the alternative image pairs based on signal data can be improved. Then, related image pairs can be screened from the alternative image pairs through the signal data corresponding to the first image and the second image in each alternative image pair, so that alternative image pairs with similar image features but not the same scene can be more accurately eliminated.
[0015] In an optional embodiment, the selecting the associated image pair from the candidate image pairs based on the signal data corresponding to the first image and the second image in the plurality of candidate image pairs includes:
[0016] For each candidate image pair among the plurality of candidate image pairs, determining a signal distance between the first image and the second image based on signal data corresponding to the first image in each candidate image pair and signal data corresponding to the second image in each candidate image pair;
[0017] In response to a signal distance between the first image and the second image being less than a preset distance threshold, the candidate image pair is determined as the associated image pair.
[0018] In this way, by adding signal observations at the same or nearby moments to the three-dimensional reconstructed image data, during the three-dimensional reconstruction process using associated image pairs, by comparing the distance between the signal data of the image pairs, it is determined whether the image pairs represent the same scene or nearby scenes, thereby reducing the impact of similar scenes on the three-dimensional reconstruction.
[0019] In an optional implementation manner, the signal data includes: signal source identification information corresponding to each frame image in the target video;
[0020] The determining of the signal distance between the first image and the second image based on the signal data corresponding to the first image in each candidate image pair and the signal data corresponding to the second image in each candidate image pair includes:
[0021] determining, based on the signal source identification information corresponding to the first image and the signal source identification information corresponding to the second image, the number of common signal sources corresponding to the first image and the second image;
[0022] Based on the number, a signal distance corresponding to the first image and the second image is determined; wherein the signal distance is negatively correlated with the number of the common signal sources.
[0023] In this way, by characterizing the signal distance between the two images through the common signal source information of the corresponding signal data of the two images, the distance between the two signal data is judged, thereby quickly and simply screening the associated image pairs and improving the efficiency of three-dimensional reconstruction.
[0024] In an optional implementation, the signal data includes: signal strength information corresponding to each frame image in the target video, and signal source identification information;
[0025] The determining of the signal distance between the first image and the second image based on the signal data corresponding to the first image in each candidate image pair and the signal data corresponding to the second image in each candidate image pair includes:
[0026] determining a common signal source and a non-common signal source corresponding to the first image and the second image based on a signal source identifier corresponding to the first image and a signal source identifier corresponding to the second image;
[0027] determining a strength difference of the common signal source based on signal strength information corresponding to the first image and the second image respectively;
[0028] A signal distance between the first image and the second image is determined based on the intensity difference and a preset signal distance corresponding to a non-shared signal source.
[0029] In this way, the signal distance between the two frames of image can be determined more accurately, the data accuracy in the 3D re-addition process can be improved, and the interference of similar scenes on 3D reconstruction can be reduced.
[0030] In an optional embodiment, the performing three-dimensional reconstruction of the target scene based on the associated image pair to obtain a three-dimensional model of the target scene includes:
[0031] For each of the multiple videos, performing three-dimensional reconstruction based on multiple frames of images in each video to obtain a three-dimensional sub-model corresponding to each video;
[0032] Based on the associated image pairs, the three-dimensional sub-models corresponding to the multiple videos are spliced together to obtain the three-dimensional model of the target scene.
[0033] In this way, the correlation relationship between multiple videos can be determined by associating image pairs to perform three-dimensional model splicing. The three-dimensional model splicing based on associated image pairs improves the accuracy of three-dimensional reconstruction.
[0034] In an optional embodiment, the step of splicing the three-dimensional sub-models corresponding to the plurality of videos based on the associated image pairs to obtain the three-dimensional model of the target scene includes:
[0035] For each two videos, determining whether there is a target associated image pair between the two videos; the first associated image and the second associated image in the target associated image pair belong to the two videos respectively;
[0036] In response to the target-associated image pair existing between the two videos, determining, based on the target-associated image pair between the two videos, conversion relationship information between the three-dimensional sub-models corresponding to the two videos, and performing splicing processing on the three-dimensional sub-models corresponding to the two videos based on the conversion relationship information;
[0037] Based on the stitching results corresponding to the multiple videos, a three-dimensional model of the target scene is obtained.
[0038] In this way, based on two images belonging to different video segments in the same associated image pair, the conversion relationship between the two 3D sub-models can be determined, so that the splicing effect of the 3D sub-models is better.
[0039] In an optional embodiment, determining, based on the target associated image pair between the two videos, the conversion relationship information between the three-dimensional sub-models corresponding to the two videos, and splicing the three-dimensional sub-models corresponding to the two videos based on the conversion relationship information includes:
[0040] Perform at least one iteration of the following:
[0041] Determining a first target-associated image pair corresponding to a current iteration cycle from the target-associated image pairs;
[0042] Determining, based on the first target associated image pair, current conversion relationship information between the three-dimensional sub-models corresponding to the two videos respectively;
[0043] Based on the current conversion relationship information, the three-dimensional sub-models corresponding to the two videos are spliced to obtain a spliced model corresponding to the current iteration cycle;
[0044] Verifying the stitching correctness of the stitching model based on physical positions represented by other target-associated image pairs except the first target-associated image in the target-associated image pairs;
[0045] In response to the splicing correctness verification being passed, the splicing model corresponding to the current iteration cycle is used as the result model of splicing the three-dimensional sub-models corresponding to the two videos, and the iteration process is terminated;
[0046] In response to the correctness verification failing, entering the next iteration cycle;
[0047] If all target-associated image pairs taken as the first target-associated image pairs fail to pass the splicing correctness verification, it indicates that there is no splicing relationship between the three-dimensional sub-models corresponding to the two videos.
[0048] In this way, all the related image pairs in the two videos are checked for stitching correctness in an iterative judgment manner, which effectively avoids errors in model stitching and improves the effect of 3D reconstruction.
[0049] In an optional embodiment, the performing three-dimensional reconstruction of the target scene based on the associated image pair to obtain a three-dimensional model of the target scene includes:
[0050] Traversing each associated image pair in the associated image pairs, and performing three-dimensional reconstruction based on the traversed associated image pairs to obtain a three-dimensional model corresponding to the traversed associated image pairs;
[0051] The three-dimensional models corresponding to the multiple associated image pairs are fused to obtain the three-dimensional model of the target scene.
[0052] In this way, by associating the association relationships between the images, the 3D sub-models corresponding to the multiple videos are spliced together to obtain the 3D model corresponding to the target scene, which can reduce the impact of repeated scenes on the reconstruction of the 3D model of the target scene.
[0053] In an optional embodiment, performing three-dimensional reconstruction based on the traversed associated image pairs to obtain a three-dimensional model corresponding to the traversed associated image pairs includes:
[0054] Performing three-dimensional reconstruction based on the traversed associated image pairs, the camera poses corresponding to the first image and the second image, and the first image and the second image to obtain a first three-dimensional model;
[0055] determining a reprojection error of the first three-dimensional model based on the first image, the second image, the camera poses corresponding to the first image and the second image, and the first three-dimensional model; and
[0056] determining a signal distance error between the first image and the second image based on signal data corresponding to the first image and the second image, respectively, and a signal source position of a target signal source corresponding to the signal data;
[0057] Calibrate the camera poses corresponding to the first image and the second image, respectively, and the signal source position of the target signal source based on the reprojection error and the signal distance error to obtain the target poses corresponding to the first image and the second image, respectively, and the target position of the target signal source;
[0058] Three-dimensional reconstruction is performed based on the first image, the second image, and the target posture to obtain a three-dimensional model corresponding to the traversed associated image pair.
[0059] In this way, through the reprojection error function and the signal distance error function, the position of the signal source, the three-dimensional point position and the camera position can be optimized at the same time, thereby avoiding reconstruction errors and obtaining a three-dimensional reconstruction including the camera position posture and three-dimensional point data.
[0060] In an optional implementation, the method further includes: constructing a signal source map based on target positions of target signal sources corresponding to the multiple associated images.
[0061] In this way, the signal information and visual 3D reconstruction can be fused, and the signal source map and 3D reconstruction can be constructed simultaneously.
[0062] In a second aspect, the present disclosure further provides a three-dimensional reconstruction device, including:
[0063] an acquisition module, configured to acquire a plurality of videos captured for a target scene, and signal data corresponding to each of the plurality of videos; wherein the signal data is obtained by capturing a signal of a target signal source located within the target scene when the videos are captured; and each video includes a plurality of frames of images;
[0064] a processing module for determining a plurality of associated image pairs from the plurality of videos based on the signal data;
[0065] A reconstruction module is used to perform three-dimensional reconstruction on the target scene based on the associated image pair to obtain a three-dimensional model of the target scene.
[0066] In an optional embodiment, the processing module, when determining a plurality of associated image pairs from the plurality of videos based on the signal data, is configured to:
[0067] Extract feature data of each frame image in multiple videos;
[0068] Determining a plurality of candidate image pairs based on similarities between feature data of different images; each candidate image pair includes a first image and a second image;
[0069] The associated image pairs are screened from the candidate image pairs based on signal data corresponding to the first image and the second image in a plurality of candidate image pairs.
[0070] In an optional embodiment, the processing module, when screening the associated image pair from a plurality of candidate image pairs based on the signal data corresponding to the first image and the second image in the plurality of candidate image pairs, is configured to:
[0071] For each candidate image pair among the plurality of candidate image pairs, determining a signal distance between the first image and the second image based on signal data corresponding to the first image in each candidate image pair and signal data corresponding to the second image in each candidate image pair;
[0072] In response to a signal distance between the first image and the second image being less than a preset distance threshold, the candidate image pair is determined as the associated image pair.
[0073] In an optional implementation manner, the signal data includes: signal source identification information corresponding to each frame image in the target video;
[0074] The processing module, when determining the signal distance between the first image and the second image based on the signal data corresponding to the first image in each candidate image pair and the signal data corresponding to the second image in each candidate image pair, is configured to:
[0075] determining, based on the signal source identification information corresponding to the first image and the signal source identification information corresponding to the second image, the number of common signal sources corresponding to the first image and the second image;
[0076] Based on the number, a signal distance corresponding to the first image and the second image is determined; wherein the signal distance is negatively correlated with the number of the common signal sources.
[0077] In an optional implementation, the signal data includes: signal strength information corresponding to each frame image in the target video, and signal source identification information;
[0078] The processing module, when determining the signal distance between the first image and the second image based on the signal data corresponding to the first image in each candidate image pair and the signal data corresponding to the second image in each candidate image pair, is configured to:
[0079] determining a common signal source and a non-common signal source corresponding to the first image and the second image based on a signal source identifier corresponding to the first image and a signal source identifier corresponding to the second image;
[0080] determining a strength difference of the common signal source based on signal strength information corresponding to the first image and the second image respectively;
[0081] A signal distance between the first image and the second image is determined based on the intensity difference and a preset signal distance corresponding to a non-shared signal source.
[0082] In an optional embodiment, the reconstruction module, when performing three-dimensional reconstruction of the target scene based on the associated image pair to obtain the three-dimensional model of the target scene, is configured to:
[0083] For each of the multiple videos, performing three-dimensional reconstruction based on multiple frames of images in each video to obtain a three-dimensional sub-model corresponding to each video;
[0084] Based on the associated image pairs, the three-dimensional sub-models corresponding to the multiple videos are spliced together to obtain the three-dimensional model of the target scene.
[0085] In an optional embodiment, the reconstruction module, when splicing the three-dimensional sub-models corresponding to the plurality of videos based on the associated image pairs to obtain the three-dimensional model of the target scene, is configured to:
[0086] For each two videos, determining whether there is a target associated image pair between the two videos; the first associated image and the second associated image in the target associated image pair belong to the two videos respectively;
[0087] In response to the target-associated image pair existing between the two videos, determining, based on the target-associated image pair between the two videos, conversion relationship information between the three-dimensional sub-models corresponding to the two videos, and performing splicing processing on the three-dimensional sub-models corresponding to the two videos based on the conversion relationship information;
[0088] Based on the stitching results corresponding to the multiple videos, a three-dimensional model of the target scene is obtained.
[0089] In an optional embodiment, the reconstruction module, when determining, based on the target associated image pair between the two videos, the transformation relationship information between the three-dimensional sub-models corresponding to the two videos, and performing splicing processing on the three-dimensional sub-models corresponding to the two videos based on the transformation relationship information, is configured to:
[0090] Perform at least one iteration of the following:
[0091] Determining a first target-associated image pair corresponding to a current iteration cycle from the target-associated image pairs;
[0092] Determining, based on the first target associated image pair, current conversion relationship information between the three-dimensional sub-models corresponding to the two videos respectively;
[0093] Based on the current conversion relationship information, the three-dimensional sub-models corresponding to the two videos are spliced to obtain a spliced model corresponding to the current iteration cycle;
[0094] Verifying the stitching correctness of the stitching model based on physical positions represented by other target-associated image pairs except the first target-associated image in the target-associated image pairs;
[0095] In response to the splicing correctness verification being passed, the splicing model corresponding to the current iteration cycle is used as the result model of splicing the three-dimensional sub-models corresponding to the two videos, and the iteration process is terminated;
[0096] In response to the correctness verification failing, entering the next iteration cycle;
[0097] If all target-associated image pairs taken as the first target-associated image pairs fail to pass the splicing correctness verification, it indicates that there is no splicing relationship between the three-dimensional sub-models corresponding to the two videos.
[0098] In an optional embodiment, the reconstruction module, when performing three-dimensional reconstruction of the target scene based on the associated image pair to obtain the three-dimensional model of the target scene, is configured to:
[0099] Traversing each associated image pair in the associated image pairs, and performing three-dimensional reconstruction based on the traversed associated image pairs to obtain a three-dimensional model corresponding to the traversed associated image pairs;
[0100] The three-dimensional models corresponding to the multiple associated image pairs are fused to obtain the three-dimensional model of the target scene.
[0101] In an optional embodiment, the reconstruction module, when performing three-dimensional reconstruction based on the traversed associated image pairs to obtain a three-dimensional model corresponding to the traversed associated image pairs, is configured to:
[0102] Performing three-dimensional reconstruction based on the traversed associated image pairs, the camera poses corresponding to the first image and the second image, and the first image and the second image to obtain a first three-dimensional model;
[0103] determining a reprojection error of the first three-dimensional model based on the first image, the second image, the camera poses corresponding to the first image and the second image, and the first three-dimensional model; and
[0104] determining a signal distance error between the first image and the second image based on signal data corresponding to the first image and the second image, respectively, and a signal source position of a target signal source corresponding to the signal data;
[0105] Calibrate the camera poses corresponding to the first image and the second image, respectively, and the signal source position of the target signal source based on the reprojection error and the signal distance error to obtain the target poses corresponding to the first image and the second image, respectively, and the target position of the target signal source;
[0106] Three-dimensional reconstruction is performed based on the first image, the second image, and the target posture to obtain a three-dimensional model corresponding to the traversed associated image pair.
[0107] In an optional implementation, the reconstruction module is further configured to construct a signal source map based on target positions of target signal sources corresponding to the plurality of associated images.
[0108] In a third aspect, an optional implementation of the present disclosure further provides a computer device, a processor, and a memory, wherein the memory stores machine-readable instructions executable by the processor, and the processor is used to execute the machine-readable instructions stored in the memory. When the machine-readable instructions are executed by the processor, the machine-readable instructions perform the steps of the above-mentioned first aspect, or any possible implementation of the first aspect.
[0109] In a fourth aspect, an optional implementation of the present disclosure further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed, it executes the steps of the above-mentioned first aspect or any possible implementation of the first aspect.
[0110] The effects of the above-mentioned 3D reconstruction apparatus, computer device, and computer-readable storage medium are described in the description of the above-mentioned 3D reconstruction method, which will not be repeated here.
[0111] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and are not intended to limit the technical solutions of the present disclosure.
[0112] In order to make the above-mentioned objectives, features and advantages of the present disclosure more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0113] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the following briefly introduces the drawings required for use in the embodiments. The drawings herein are incorporated into and constitute a part of the specification. These drawings illustrate embodiments consistent with the present disclosure and, together with the specification, are used to illustrate the technical solutions of the present disclosure. It should be understood that the following drawings only illustrate certain embodiments of the present disclosure and should not be regarded as limiting the scope. For those of ordinary skill in the art, other relevant drawings can be obtained based on these drawings without inventive effort.
[0114] Figure 1 A flowchart of a three-dimensional reconstruction method provided by an embodiment of the present disclosure is shown;
[0115] Figure 2 A specific example of a repetitive scenario provided by an embodiment of the present disclosure is shown;
[0116] Figure 3 A specific example of a 3D reconstruction error caused by repeated scenes provided by an embodiment of the present disclosure is shown;
[0117] Figure 4 A specific example of correct three-dimensional reconstruction after fusion of signal data provided by an embodiment of the present disclosure is shown;
[0118] Figure 5 A specific example of the relative position relationship between a signal source and signal data provided by an embodiment of the present disclosure is shown;
[0119] Figure 6 shows a specific example of the relative position relationship between another signal source and signal data provided by an embodiment of the present disclosure;
[0120] Figure 7 A schematic diagram illustrating the principle of a three-point positioning method provided by an embodiment of the present disclosure is shown;
[0121] Figure 8 A specific example of the projection relationship between a two-dimensional image and a three-dimensional object provided by an embodiment of the present disclosure is shown;
[0122] Figure 9 A schematic diagram of a three-dimensional reconstruction device provided by an embodiment of the present disclosure is shown;
[0123] Figure 10 A schematic diagram of a computer device provided by an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0124] In order to make the purpose, technical solutions and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only part of the embodiments of the present disclosure, rather than all of the embodiments. The components of the embodiments of the present disclosure generally described and shown here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present disclosure is not intended to limit the scope of the present disclosure for protection, but merely represents the selected embodiments of the present disclosure. Based on the embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of the present disclosure.
[0125] Research has found that when constructing large scenes, vision-based 3D reconstruction is easily affected by repeated scenes, leading to reconstruction errors. For example, scenes such as large airports, stadiums, and warehouses are prone to repeated similar structures that are difficult to distinguish with the naked eye. In other words, because the design of the building may be two scenes with completely different actual locations, they are indeed very similar. When reconstructing the target scene based on the image, the camera position and posture information is reconstructed. Figure 3 As shown, this will cause confusion in the camera position and posture information of each scene (the correct one should be Figure 4 shown).
[0126] Based on the above research, the present disclosure provides a three-dimensional modeling method, which adds signal observations at the same or adjacent moments to the three-dimensional reconstructed image data. In the process of three-dimensional reconstruction through associated image pairs, the distance between the signal data of the associated image pairs is supervised to determine whether the image pairs represent the same scene or adjacent scenes, thereby reducing the influence of similar scenes on the three-dimensional reconstruction and enhancing the robustness of visual three-dimensional reconstruction.
[0127] Furthermore, the disclosed embodiments can integrate signal observations with image observations, leveraging accurate camera position and posture information constructed through visual 3D reconstruction to automatically and accurately recover the signal source's location. This simultaneously combines visual 3D reconstruction with the calculation of the signal source's position within the scene, streamlining the algorithm and improving efficiency.
[0128] The defects in the above solutions are the results obtained by the inventors after practice and careful research. Therefore, the process of discovering the above problems and the solutions proposed by this disclosure for the above problems below should be the contributions made by the inventors to this disclosure during the disclosure process.
[0129] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings.
[0130] To facilitate understanding of this embodiment, a detailed description of a 3D reconstruction method disclosed in an embodiment of the present disclosure is first provided. The 3D reconstruction method provided in an embodiment of the present disclosure is generally executed by a computer device with certain computing capabilities, such as a terminal device, a server, or other processing device. The terminal device may be a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, an in-vehicle device, a wearable device, etc. In some possible implementations, the 3D reconstruction method may be implemented by a processor invoking computer-readable instructions stored in a memory.
[0131] See also Figure 1 FIG. 1 is a flow chart of a three-dimensional reconstruction method provided by an embodiment of the present disclosure, wherein the method includes steps S101 to S103, wherein:
[0132] S101: Acquire multiple videos captured for a target scene, and signal data corresponding to the multiple videos respectively; wherein the signal data is obtained by capturing a signal of a target signal source located in the target scene when the videos are captured; each video includes multiple frames of images.
[0133] S102: Determine a plurality of associated image pairs from the plurality of videos based on the signal data.
[0134] S103: Based on the associated image pair, perform three-dimensional reconstruction on the target scene to obtain a three-dimensional model of the target scene.
[0135] In the disclosed example, after obtaining multiple videos captured for a target scene and signal data corresponding to each of the multiple videos, multiple associated image pairs are determined based on the signal data, and based on the associated image pairs, a three-dimensional reconstruction of the target scene is performed to obtain a three-dimensional model of the target scene. In this process, because the signal data corresponding to the two frames of images in the associated image pairs can indicate that the signal distance corresponding to the two frames of images meets a preset distance condition, when the target scene is three-dimensionally reconstructed based on the associated image pairs, the signal distance between different images is used to supervise the three-dimensional reconstruction process. Therefore, in a target scene with a large number of repeated scenes, the impact of repeated scenes in different locations on the three-dimensional reconstruction effect can be reduced, thereby reducing reconstruction errors during the three-dimensional reconstruction process.
[0136] The above S101 to S103 are described below respectively.
[0137] Regarding the above S101: In the example of the present disclosure, the target scene may be a scene having the following characteristics: Figure 2 Large venues with multiple repetitive scenes include, but are not limited to, stadiums, airports, and warehouses. For example, large stadiums typically have multiple exits within the auditorium, and the layouts near these exits are very similar, making it difficult for humans to distinguish them without the signage. Vision-based 3D reconstruction is even more difficult for scenes with similar features to distinguish them.
[0138] In addition, this method can also be used for other types of scenes. The specific scene types are not limited here. The accuracy and efficiency of the three-dimensional reconstruction performed by this method can be improved.
[0139] In practice, the capture of the target scene's corresponding video can be accomplished using devices with video capture capabilities, such as cameras, video recorders, and mobile phones. The target scene video can be captured by a tester holding a handheld camera while moving within the target scene, capturing the video as they move. This can be done by multiple testers simultaneously or by a single tester, as long as the entire target scene is fully captured. Alternatively, a robot carrying the camera can be used to capture the video as it moves within the target scene.
[0140] Multiple target signal sources are pre-deployed in the target scene. Since the area covered by the signal emitted by each signal source is limited, the spacing between different signal sources will be set in advance to ensure that the signals emitted by multiple signal sources can cover the entire target scene.
[0141] For example, different signal types can be used depending on the desired 3D scene type, and this disclosure does not limit this. For example, in large, open spaces, a signal type with wide coverage can be used, while in dense multipath environments such as indoor spaces, a signal type with low interception capability and high positioning accuracy can be used. Signal types suitable for this disclosure include, but are not limited to, mobile hotspots (Wi-Fi), Bluetooth, and Ultra Wide Band (UWB).
[0142] The signal source can emit a signal. When acquiring the videos corresponding to the plurality of videos, the signal data corresponding to each video can also be obtained by using the signal of the target signal source located in the target scene captured when the video is collected.
[0143] This disclosure also does not limit signal acquisition. For example, a signal collector can be used for signal acquisition. The collector is an automated device capable of real-time on-site data acquisition and processing. It provides real-time acquisition, automatic storage, instant display, instant feedback, automatic processing, and automatic transmission. This ensures the authenticity, validity, real-time nature, and usability of on-site data. The collector includes an operating system and a built-in wireless communication module (Wi-Fi, GPRS, or Bluetooth).
[0144] For example, video and signal acquisition can be performed simultaneously, so that the two are one-to-one corresponding during acquisition. After the video and signal acquisition are completed, the image information and signal data can be matched using the timestamps of the images in the video and the timestamps of the signal data. Full coverage of the target scene requires multiple videos, each of which includes multiple frames, and each frame has corresponding signal data.
[0145] Regarding the above S102: the embodiment of the present disclosure provides a specific method for determining multiple associated image pairs from the multiple videos based on the signal data, including the following steps 21 to 23:
[0146] Step 21: extracting feature data of each frame image in the multiple videos;
[0147] Step 22: Determine a plurality of candidate image pairs based on similarities between feature data of different images; each candidate image pair includes a first image and a second image;
[0148] Step 23: Based on the signal data corresponding to the first image and the second image in a plurality of candidate image pairs, select the associated image pair from the candidate image pairs.
[0149] Regarding step 21: Features are the corresponding (essential) characteristics or properties that distinguish one class of objects from other classes of objects, or a collection of these characteristics and properties. Features are data that can be extracted through measurement or processing. For images, each image has its own unique characteristics that distinguish it from other images. Some are natural features that can be intuitively perceived, such as brightness, edges, texture, and color; others require transformation or processing, such as moments, histograms, and principal components.
[0150] Multiple or various characteristics of a class of objects are combined to form a feature vector to represent that class of objects. If there is only a single numerical feature, the feature vector is a one-dimensional vector. If it is a combination of n characteristics, it is an n-dimensional feature vector. This type of feature vector is often used as the input of the recognition system. In fact, an n-dimensional feature is a point located in n-dimensional space, and the task of recognition and classification is to find a partition of this n-dimensional space. For example, to distinguish between three different flowers, you can choose their petal length and petal width as features. In this way, a plant object is represented by a two-dimensional feature, such as (5.1, 3.5). If the leaf length and leaf width are added, each flower object is represented by a four-dimensional feature vector, such as (5.1, 3.5, 1.4, 0.2).
[0151] In addition, a neural network may be used to extract features from each frame image in the video to obtain feature data corresponding to each frame image.
[0152] In practice, this disclosure aims to integrate visual information with signal data. Therefore, the method for extracting image feature data is not limited. As long as the feature data of each frame in multiple videos can be successfully extracted for subsequent processing, it can be used. For example, methods based on shape features, spatial relationships, and texture features can be used.
[0153] Regarding step 22: after obtaining the feature data of each frame image in multiple videos, multiple candidate image pairs can be determined based on the similarity between the feature data of different images. Each of the multiple candidate image pairs includes a first image and a second image, and the feature data of these two images are similar. The present disclosure does not limit the method for calculating image similarity. For example, whether two images are similar can be determined by methods such as structural similarity measurement, cosine similarity, histogram, and mutual information.
[0154] For example, after determining multiple candidate image pairs, the first image and the second image in each candidate image pair are the two most similar images calculated based on the similarity of feature data. However, it is worth noting that although the feature data of the first image and the second image are very similar, it does not mean that they represent the same or adjacent scenes. For example, when the target scene is an airport, different images may represent Gate 1 and Gate 13 of the airport, or Figure 2 The East and West Gates of the large stadium shown.
[0155] Regarding step 23: multiple associated image pairs can be obtained by screening candidate image pairs using the signal data corresponding to the first image and the second image in the image pair.
[0156] The two images in the associated image pair not only have similar feature data, but also have the same or similar corresponding actual description scenes.
[0157] In an optional implementation, the present disclosure further provides a specific method for selecting the associated image pair from a plurality of candidate image pairs based on signal data corresponding to the first image and the second image, respectively, including the following steps 231 to 232:
[0158] Step 231: For each candidate image pair in a plurality of candidate image pairs, determine a signal distance between the first image and the second image based on signal data corresponding to the first image in each candidate image pair and signal data corresponding to the second image in each candidate image pair.
[0159] Step 232: In response to the signal distance between the first image and the second image being less than a preset distance threshold, determining the candidate image pair as the associated image pair.
[0160] In practice, the distance between the signal data corresponding to two images can indicate whether the two signal data were collected from similar locations. Signal distance is determined based on received signal strength (RSS), sometimes also called received signal strength indicator (RSSI), a key metric used by the wireless transmission layer to determine link quality. Typically, RSS is expressed in watts (W). However, wireless signals are relatively weak, typically in the milliwatt (mW) range. Therefore, a common practice is to use 1mW as a baseline and express signal strength in logarithmic form, known as RSSI, in decibel-milliwatts (dBm). Therefore, in wireless signals, 1mW is equivalent to 0dBm. Signals with less than 1mW have a negative RSSI, while signals with greater than 1mW have a positive RSSI. Signal strength RSSI is related to distance; intuitively, the greater the distance, the lower the signal strength.
[0161] In a specific implementation, the signal measuring device can reflect the signal strength corresponding to multiple signal sources at the measurement location. It should be noted that the signal distance of multiple signal data cannot be determined based on the signal strength of a single signal source, because what can be obtained based on the signal strength is not the specific position of the signal data relative to the signal source, but the range information. For example, it can be known that the distance between the signal data and the signal source is 10m, but the direction cannot be determined. It may be 10m in the southeast direction or 10m in the southwest direction. In the embodiment of the present disclosure, a collaborative ranging method of multiple signal sources is adopted to determine the signal distance of multiple signal sources.
[0162] For example, there are two signal collection locations A and B, and both locations are covered by signal source 1, signal source 2, and signal source 3. If the signal strength of signal source 1 measured at A is -10dBm, the signal strength of signal source 2 is -15dBm, and the signal strength of signal source 3 is -12dBm, and the signal strength of signal source 1 measured at B is -10dBm, the signal strength of signal source 2 is -14dBm, and the signal strength of signal source 3 is -12dBm, then based on the preset judgment condition of similar signal data, it can be determined that the signal collection locations A and B are similar.
[0163] In a specific implementation, for example, the signal distance may be determined by using any of the following methods A or B:
[0164] A: The signal data includes: signal source identification information corresponding to each frame image in the target video.
[0165] In a specific implementation, each frame of the target video has corresponding signal data. When a signal is detected at a location, it must be emitted by a certain signal source. In addition, the signal detected at each location can also come from multiple signal sources. The signal sources of a particular signal data can be directly determined by the signal acquisition device.
[0166] The determining of the signal distance between the first image and the second image based on the signal data corresponding to the first image in each candidate image pair and the signal data corresponding to the second image in each candidate image pair includes:
[0167] determining, based on the signal source identification information corresponding to the first image and the signal source identification information corresponding to the second image, the number of common signal sources corresponding to the first image and the second image;
[0168] In a specific implementation, as described above, the signal data collected at each signal collection location includes signal source identification information, that is, signal data from which signal sources can be detected at the collection location.
[0169] For example, the image shooting location of the first image, that is, the signal data collection location, is A, and the signal data collection location corresponding to the second image is B. The signal sources in the signal data collected at A are identified as signal source 1, signal source 3, and signal source 6, and the signal sources in the signal data collected at B are identified as signal source 1, signal source 3, and signal source 12. Then, signal source 1 and signal source 3 are the common signal sources of the first image and the second image, and the number of common signal sources is 2 at this time.
[0170] Based on the number, a signal distance corresponding to the first image and the second image is determined; wherein the signal distance is negatively correlated with the number of the common signal sources.
[0171] For example, when the number of common signal sources between two signal data collection locations is greater, it means that the two signal data collection locations are closer, and when the number of common signal sources between two locations is smaller, it means that the two signal data collection locations are far apart and have no correlation.
[0172] B: The signal data includes: signal strength information corresponding to each frame image in the target video, and signal source identification information;
[0173] The determining of the signal distance between the first image and the second image based on the signal data corresponding to the first image in each candidate image pair and the signal data corresponding to the second image in each candidate image pair includes:
[0174] Determining a shared signal source and a non-shared signal source corresponding to the first and second images based on a signal source identifier corresponding to the first image and a signal source identifier corresponding to the second image; and determining a strength difference of the shared signal source based on signal strength information corresponding to the first and second images, respectively.
[0175] A signal distance between the first image and the second image is determined based on the intensity difference and a preset signal distance corresponding to a non-shared signal source.
[0176] In a specific implementation, for example, if the signal source identification information of the first image includes: signal source 1, signal source 2, signal source 3, and the signal source identification information of the second image includes: signal source 1, signal source 2, signal source 4, then signal source 1 and signal source 2 are the common signal sources in the signal data corresponding to the first image and the second image respectively, and signal source 3 and signal source 4 are the non-common signal sources in the signal data corresponding to the first image and the second image respectively.
[0177] The signal strength information corresponding to each image can be obtained through weighted fitting within a certain time range. Taking Bluetooth signals as an example, the signal acquisition device does not receive Bluetooth signal data immediately but has a certain delay.
[0178] For example, if one wants to obtain the signal strength information of a certain signal source corresponding to the image at the 100th second in a 200-second video (one of the frames can be taken if there are multiple frames per second), since there is a delay in signal reception, in order to make the signal strength information easier to collect, the signal data collected in the interval of 91-109 seconds in the 200-second video can be weighted fitted to represent the signal strength information at the 100th second. For example, if the signal strength weight coefficient collected in the interval of 91-100 seconds is set to 0.1-1.0, and the signal strength weight coefficient collected in the interval of 100-109 seconds is set to 1.0-0.1, if the signal strength information collected at the 100th second is 20dBm, the signal strength information collected at the 92nd second is 23dBm, the signal strength information collected at the 101st second is 25dBm, and no signal data is collected at other times, then the signal strength information of the image at the 100th second corresponding to a certain signal source can be obtained by weighted fitting; For another example, if the signal strength weight coefficient collected between 91 and 100 seconds is set to 0.1-1.0, and the signal strength weight coefficient collected between 100 and 109 seconds is set to 1.0-0.1, if the signal strength information collected at the 92nd second is 19dBm, the signal strength information collected at the 104th second is 22dBm, the signal strength information collected at the 101st second is 18dBm, and no signal data is collected at other times, then the image at the 100th second can be weightedly fitted to obtain the signal strength information corresponding to a certain signal source. This disclosure does not limit the weighted fitting algorithm for data. As a data processing method, any algorithm that can perform weighted fitting on the above data can be used.
[0179] In specific implementation, for the intensity difference of common signal sources, the above example illustrates the signal intensity information of a certain image corresponding to a certain signal source. Based on this method, the signal intensity information of all signal sources included in the signal data matched by the image can be obtained. Similarly, the weighted fitting of the data can be used to obtain the intensity difference of the common signal source of the two images.
[0180] For example, Figure 5As shown, the common signal sources of the first image and the second image are signal source 1, signal source 2, and signal source 13. The signal coverage range of each signal source in the figure is represented by a circle. In the corresponding signal data of the first image, the intensity of signal source 1 is 30dBm, the intensity of signal source 2 is 40dBm, and the intensity of signal source 13 is 15dBm. In the corresponding signal data of the second image, the intensity of signal source 1 is 29dBm, the intensity of signal source 2 is 16dBm, and the intensity of signal source 13 is 40dBm. The above is the signal data strength information. Based on the signal data strength information corresponding to each signal source, it can be determined that for the same common signal source, the signal strength difference between the first image and the second image can be determined. For example, based on the above data, it can be determined that for signal source 1, the signal strength difference between the first image and the second image is 1dBm, for signal source 2, the signal strength difference between the first image and the second image is 24dBm, and for signal source 13, the signal strength difference between the first image and the second image is 25dBm. Then, the signal strength difference data of each signal source can be weighted, and is given by Figure 5 It can be seen that the signal strengths of the first and second images relative to signal source 1 are almost equal, but their actual signal sampling locations are not close. In other words, when there are multiple shared signal sources, if the difference in signal strength between a particular signal source is small, it can be assigned a lower weight; if the difference in signal strength between a particular signal source is large, it can be assigned a higher weight. Based on the obtained signal strength difference data and the weight of the signal source to which the signal strength difference belongs, a weighted fit can be used to calculate the total signal strength difference between the first and second images.
[0181] In a specific implementation, after determining the signal strength difference between the first image and the second image, the correlation between the two images can be judged based on a preset signal strength difference threshold. If the signal strength difference between the two images is greater than the preset threshold, the two images are considered unrelated; if the signal strength difference between the two images is less than the preset threshold, the two images are considered related.
[0182] In a specific implementation, a non-public signal source can determine the signal distance.
[0183] For example, Figure 6 As shown, signal source 1 and signal source 2 are non-common signal sources of the first image and the second image. When arranging the signal sources in the target scene, the distance between the signal sources can be determined in advance. If the two signal sources are far apart, the signal data within their signal coverage range will also be far apart.
[0184] The method disclosed in 102 above can screen the associated image pair from the candidate image pairs based on the signal data corresponding to the first image and the second image in the plurality of image pairs.
[0185] The images in the associated image pair are images with similar signal distances, that is, images with similar actual shooting distances.
[0186] Regarding S103 above: the embodiment of the present disclosure further provides a specific method for performing three-dimensional reconstruction on the target scene based on the associated image pair to obtain a three-dimensional model of the target scene, including:
[0187] For each of the multiple videos, performing three-dimensional reconstruction based on multiple frames of images in each video to obtain a three-dimensional sub-model corresponding to each video;
[0188] Based on the associated image pairs, the three-dimensional sub-models corresponding to the multiple videos are spliced together to obtain the three-dimensional model of the target scene.
[0189] In practice, the goal of image-based 3D reconstruction is to infer the 3D geometry and structure of objects and scenes from one or more 2D images, and to infer missing dimensions from 2D images. This disclosure does not limit the specific 3D reconstruction algorithm; instead, it aims to provide an optimization method for 3D reconstruction.
[0190] In a specific implementation, as described above, in order to improve the efficiency of video data acquisition, the acquisition of the target scene video does not utilize a whole video to completely shoot the target scene, but rather simultaneously shoots multiple videos of the target scene. In multiple videos, the scenes in each video are continuous, and multiple frames of images in the same video describe the same or adjacent target scene segments. When performing three-dimensional reconstruction, three-dimensional reconstruction can be performed on each video segment to obtain a three-dimensional sub-model corresponding to the partial target scene captured in each video segment, and the association relationship between each video segment can be determined through the associated image pairs obtained in the above S102. For example, the last frame image in video segment A is the first image in the associated image pair, and the first frame image in video segment B is the second image in the same associated image pair, then the three-dimensional sub-models corresponding to the two images from different video segments in the same associated image can be spliced.
[0191] For example, a Perspective-n-Point (PNP) algorithm is performed on the triangulated 3D points within the image sequence reconstruction and the 2D points associated with the data to calculate the 3D transformation between the image sequences, including the 3D rotation R and the 3D translation t. Applying this 3D transformation to one of the image sequences allows the 3D reconstructions of the two image sequences to be stitched together.
[0192] PNP is a method for solving 3D-to-2D point pair motion, aiming to determine the pose of the camera coordinate system relative to the world coordinate system. It describes how to estimate the camera pose (i.e., solving the rotation matrix R and translation vector t from the world coordinate system to the camera coordinate system) when the coordinates of n 3D points (relative to the world coordinate system) and the pixel coordinates of these points are known.
[0193] In an optional embodiment, in addition to obtaining a three-dimensional model of the target scene by splicing the corresponding three-dimensional models of multiple video segments, each frame image in the multiple video segments can be spliced into an image set that can fully describe the target scene based on the associated image pairs, and the image set can be three-dimensionally reconstructed.
[0194] In a specific implementation, a group of associated image pairs can describe the same scene or adjacent scenes, and by splicing the pictures described by all associated image pairs, an image that can describe the entire target scene can be obtained. The picture described by the image is similar to a panoramic photo.
[0195] In an optional implementation, the embodiment of the present disclosure further provides a specific method for splicing the three-dimensional sub-models corresponding to the plurality of videos based on the associated image pairs to obtain the three-dimensional model of the target scene, including:
[0196] For every two videos, determine whether there is a target-associated image pair between the two videos; the first associated image and the second associated image in the target-associated image pair belong to the two videos respectively; in response to the existence of the target-associated image pair between the two videos, based on the target-associated image pair between the two videos, determine the conversion relationship information between the three-dimensional sub-models corresponding to the two videos respectively, and based on the conversion relationship information, splice the three-dimensional sub-models corresponding to the two videos respectively; based on the splicing results corresponding to multiple videos, obtain the three-dimensional model of the target scene.
[0197] Exemplarily, each of the multiple groups of associated image pairs has two associated images, namely the first image and the second image. These two images may come from the same video or from different segments of video. For example, if a video consists of 1,000 frames of images (in a specific implementation, dozens or hundreds of them can be sampled to perform the three-dimensional reconstruction method in the present disclosure), these 1,000 frames of images can constitute 500 groups of associated image pairs. However, since the scenes represented by the same video are most likely similar or continuous, the first image and the second image in 498 groups of associated image pairs may come from the same video, and the 1,000th frame image and the first frame image can form a group of associated image segments with a frame image in other segments of video. Then this type of associated image pair is what we need.
[0198] The 3D sub-model corresponding to a video is reconstructed using information from each frame within that video. This means that the information in each frame can be found in the reconstructed 3D sub-model. When two videos are linked using a set of associated image pairs, the corresponding positions of the first image in the reconstructed 3D sub-model can be joined with the corresponding positions of the second image in the reconstructed 3D sub-model due to the association between the images.
[0199] Among them, since the camera parameters of the two videos may be different when they are shot, the splicing of the corresponding models of the two videos requires first determining the conversion relationship between the two 3D sub-models.
[0200] For example, the essence of spatial coordinate transformation is to use two sets of coordinates for common points and one set of coordinates for non-common points to estimate another set of coordinates for non-common points. The coordinate transformation process is typically divided into two steps: first, calculating the transformation parameters from the common point coordinates, and then using the transformation parameters to transform the non-common points. The transformation parameters are generally divided into rotation, translation, and scale parameters, with the determination of the rotation parameters being the core of coordinate transformation. Traditional three-dimensional coordinate transformation models use three rotation angles as rotation parameters. The resulting model is nonlinear and often requires linearization using Taylor series expansion, which is computationally complex. For small-angle rotations, the rotation matrix can be approximated to obtain a linear model, such as the commonly used Bursa model. For coordinate transformation problems involving large rotation angles, the Rodriguez matrix is often used to represent the rotation matrix. This method uses only three rotation parameters, eliminates the need for linearization in the calculation process, and is applicable to large rotation angle transformations. The present disclosure does not limit the coordinate transformation algorithm. Through coordinate transformation, the two three-dimensional sub-models to be spliced can be converted to the same coordinate system, allowing the model to be spliced.
[0201] After splicing the 3D sub-models corresponding to all videos, the 3D reconstruction result of the target scene can be obtained.
[0202] In an optional implementation, the disclosed embodiment further provides a specific method for determining, based on a target-related image pair between the two videos, conversion relationship information between the three-dimensional sub-models corresponding to the two videos, and splicing the three-dimensional sub-models corresponding to the two videos based on the conversion relationship information, including:
[0203] Perform at least one iteration of the following:
[0204] Determining a first target-associated image pair corresponding to a current iteration cycle from the target-associated image pairs;
[0205] Determining, based on the first target associated image pair, current conversion relationship information between the three-dimensional sub-models corresponding to the two videos respectively;
[0206] Based on the current conversion relationship information, the three-dimensional sub-models corresponding to the two videos are spliced to obtain a spliced model corresponding to the current iteration cycle;
[0207] Verifying the stitching correctness of the stitching model based on physical positions represented by other target-associated image pairs except the first target-associated image in the target-associated image pairs;
[0208] In a specific implementation, there may be multiple groups of associated image pairs that can be successfully matched in the two videos. Therefore, in the process of splicing the three-dimensional sub-models, in order to ensure the accuracy of the final three-dimensional model splicing results, the three-dimensional model can be spliced first using the three-dimensional coordinate transformation relationship corresponding to one group of associated image pairs. The multiple groups of associated image pairs at the splicing point of the two three-dimensional sub-models can use the three-point positioning method to determine the actual shooting position of the image, and compare the distance between the actual shooting position of the first image and the second image in the same group of associated image pairs with the signal distance. If the two distances are close, it can be considered that the splicing is correct.
[0209] Among them, Figure 7 As shown in the figure, the three-point positioning method requires three points, which are generally signal sources or signal base stations. The location to be located is generally the signal receiving end. By measuring the distance between the signal receiving end and the signal transmitting end, using this as the radius of the three circles, a diagram is drawn, and the intersection of the three circles is ultimately determined. The intersection is the terminal's location, achieving positioning. In the figure, points A, B, and C are the three signal sources, and point O is the signal receiving end to be located.
[0210] For example, the actual physical position of the signal receiving end can be obtained through the three-point positioning method, so the physical distance between the first image and the second image in the same group of associated image pairs can be determined, and the signal distance of the two images can be obtained based on the signal distance determination method disclosed in listing S102. The physical distance and the signal distance are compared. If the two distances are close, it can be considered that the splicing is correct.
[0211] In response to the splicing correctness verification being passed, the splicing model corresponding to the current iteration cycle is used as the result model of splicing the three-dimensional sub-models corresponding to the two videos, and the iteration process ends.
[0212] In a specific implementation, if there are n groups of related image pairs that can be successfully matched among all the images in the two videos, then there are n-1 groups of other related image pairs that need to undergo the above-mentioned splicing correctness test.
[0213] For example, when all n-1 groups of associated image pairs pass the stitching correctness test, it can be considered that the stitching between the three-dimensional sub-models corresponding to the two videos is correct, and the stitching result can be output.
[0214] In response to the correctness verification failing, entering the next iteration cycle.
[0215] If all target-associated image pairs taken as the first target-associated image pairs fail to pass the splicing correctness verification, it indicates that there is no splicing relationship between the three-dimensional sub-models corresponding to the two videos.
[0216] In practice, the aforementioned verification of the correctness of the 3D model stitching is based on a pre-determined 3D coordinate transformation relationship, namely the initial first target associated image pair. However, because the determined 3D coordinate transformation relationship itself may not be accurate, if the correctness verification fails under a certain coordinate transformation relationship, the current 3D coordinate transformation relationship can be replaced and the aforementioned verification can be repeated. Finally, after finding the relationship that can pass the correctness verification among multiple 3D coordinate transformation relationships, a 3D reconstruction containing the camera position and posture (R, t) and 3D point (x, y, z) data can be obtained. Based on the accurate camera position and its corresponding signal data, three-point positioning is performed to calculate an accurate signal source location map.
[0217] For example, if there are n sets of target-associated image pairs, and a 3D coordinate transformation relationship is determined based on one set, i.e., a 3D coordinate transformation matrix A is determined. After performing stitching between 3D models based on A, if the remaining n-1 sets of target-associated image pairs fail the correctness test, a new 3D coordinate transformation relationship matrix B can be re-determined based on one of the n sets of target-associated image pairs, and the correctness test can be re-performed on the remaining n-1 sets of target-associated image pairs. If none of the n 3D coordinate transformation matrices determined based on the n sets of target-associated image pairs pass the correctness test, then the two videos being stitched are unrelated.
[0218] In an optional implementation, the present disclosure also provides another specific method for performing three-dimensional reconstruction of the target scene based on the associated image pair to obtain a three-dimensional model of the target scene, including:
[0219] Traversing each associated image pair in the associated image pairs, and performing three-dimensional reconstruction based on the traversed associated image pairs to obtain a three-dimensional model corresponding to the traversed associated image pairs;
[0220] The three-dimensional models corresponding to the multiple associated image pairs are fused to obtain the three-dimensional model of the target scene.
[0221] For example, if the target scene corresponds to n videos, and there are m sets of associated image pairs obtained using the method described in S102, then the m sets of associated image pairs are traversed, and a 3D reconstruction is performed for each traversed image pair to obtain a 3D sub-model corresponding to that associated image pair. Then, based on whether the image pairs have identical images, the 3D models corresponding to different associated image pairs are spliced together to ultimately obtain a 3D reconstruction of the target scene.
[0222] In an optional embodiment, performing three-dimensional reconstruction based on the traversed associated image pairs to obtain a three-dimensional model corresponding to the traversed associated image pairs includes:
[0223] Performing three-dimensional reconstruction based on the traversed associated image pairs, the camera poses corresponding to the first image and the second image, and the first image and the second image to obtain a first three-dimensional model;
[0224] determining a reprojection error of the first three-dimensional model based on the first image, the second image, the camera poses corresponding to the first image and the second image, and the first three-dimensional model; and
[0225] determining a signal distance error between the first image and the second image based on signal data corresponding to the first image and the second image, respectively, and a signal source position of a target signal source corresponding to the signal data;
[0226] Among them, Figure 8 As shown in the figure, the first projection refers to the projection of three-dimensional space points onto the image when the camera takes a picture, and then these images are used to triangulate some feature points, that is, triangles are constructed using geometric information to determine the position of the three-dimensional space points.
[0227] Finally, the calculated coordinates of the three-dimensional point (not real) and the camera matrix (not real) are used for the second projection, that is, reprojection.
[0228] The reprojection error refers to the difference between the projection of a real 3D space point on the image plane (that is, the pixel point on the image) and the reprojection (the calculated virtual pixel point). Due to various reasons, the calculated value will not be completely consistent with the actual situation, that is, this difference cannot be exactly 0. At this time, it is necessary to minimize the sum of these differences to obtain the optimal camera parameters and the coordinates of the 3D space point.
[0229] The pose is the transformation from the world coordinate system to the camera coordinate system, including rotation and translation. The pose is essentially a transformation matrix, which transforms the world coordinate system into the camera coordinate system. As can be seen from the above pose description, the position and posture of an object are uniquely determined by the reference coordinate system. When the pose of the same object is described in different reference coordinate systems, the pose representation will also change. Therefore, coordinate transformation is required to link the descriptions of the object in different reference coordinate systems. Coordinate transformations between different coordinate systems include translation transformation, rotation transformation, and compound transformation.
[0230] In a specific implementation, the error function of the reprojection error is as follows:
[0231]
[0232] Among them, E reproject is the reprojection error, m is the number of 3D points, n is the number of images, P is the reprojection function, x is the reprojected 3D point, R and t are the rotation and translation of the current image, and u and v are the 2D observations of the current feature point.
[0233] For example, after obtaining the reprojection error using the above method, the displacement deviation of the projection of a two-dimensional point to a three-dimensional point can be obtained. Based on the coordinate deviation corresponding to the displacement deviation, the transformation matrix between the world coordinate system and the camera coordinate system can be optimized to achieve a more accurate coordinate conversion effect. As for minimizing the reprojection error, mathematical methods such as the least-squares method can be used. This disclosure aims to combine signal data with visual information and is therefore not limited to this.
[0234] In order to avoid repeated scene reconstruction errors caused by visual information association errors, the signal data corresponding to each image is also added as observations to the overall error function for calculation.
[0235] In a specific implementation, the signal distance error function is as follows:
[0236]
[0237] Among them, E SignalDistance is the signal distance error, m is the number of signal points, n is the number of image pairs, i and j represent image pairs with data association, T s,i With T s,j is the relative pose of the current image to the signal s. The initial position index can be obtained by triangulation, T i,j is the relative pose of the two images in the associated image pair. By minimizing this error function, the 3D position of the signal point can be optimized. The specific optimization method is the same as the method for minimizing the reprojection error, both using mathematical tools.
[0238] Based on the reprojection error and the signal distance error, the camera poses corresponding to the first image and the second image, and the signal source position of the target signal source are calibrated to obtain the target poses corresponding to the first image and the second image, and the target position of the target signal source.
[0239] For example, in the three-dimensional reconstruction process, a three-dimensional scene can be reconstructed based on two adjacent frames of images in a video and the corresponding postures of the two frames of images. The three-dimensional scene is then reprojected at the corresponding posture, and the reprojection error is determined based on the reprojection result and the original image. At the same time, the signal distance error is determined based on the signal data corresponding to the two frames of images. The reprojection error and the signal distance error are used to optimize the signal source position, the three-dimensional point position, and the camera position.
[0240] The final optimization error can be calculated as the weighted sum of these two errors. This optimization method allows for simultaneous optimization of the signal source position, 3D point position, and camera position. This avoids reconstruction errors and yields a 3D reconstruction that includes both the camera position and pose, as well as the 3D point data.
[0241] Three-dimensional reconstruction is performed based on the first image, the second image, and the target posture to obtain a three-dimensional model corresponding to the traversed associated image pair.
[0242] In an optional implementation, the method further includes: constructing a signal source map based on target positions of target signal sources corresponding to the multiple associated images.
[0243] In a specific implementation, after optimizing the signal source position using the above method, there is no need to recalculate the signal source position, because the specific position of each optimized signal source can be directly output based on the optimization algorithm used, and a signal source map can be constructed based on multiple signal source positions.
[0244] The signal source map can be used for more accurate signal positioning.
[0245] Those skilled in the art will understand that in the above-mentioned method of the specific implementation method, the writing order of each step does not mean a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.
[0246] Based on the same inventive concept, the embodiments of the present disclosure also provide a three-dimensional reconstruction device corresponding to the three-dimensional reconstruction method. Since the principle of solving the problem by the device in the embodiments of the present disclosure is similar to the above-mentioned three-dimensional reconstruction method in the embodiments of the present disclosure, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be repeated.
[0247] Reference Figure 9FIG. 1 is a schematic diagram of a three-dimensional reconstruction device provided by an embodiment of the present disclosure, wherein the device includes: an acquisition module 91, a processing module 92, and a reconstruction module 93; wherein,
[0248] The acquisition module 91 is configured to acquire a plurality of videos captured for a target scene, and signal data corresponding to each of the plurality of videos; wherein the signal data is obtained by capturing a signal of a target signal source located within the target scene when the videos are captured; and each video includes a plurality of frames of images;
[0249] The processing module 92 is configured to determine a plurality of associated image pairs from the plurality of videos based on the signal data;
[0250] The reconstruction module 93 is configured to perform three-dimensional reconstruction on the target scene based on the associated image pair to obtain a three-dimensional model of the target scene.
[0251] In an optional embodiment, the processing module 92, when determining a plurality of associated image pairs from the plurality of videos based on the signal data, is configured to:
[0252] Extract feature data of each frame image in multiple videos;
[0253] Determining a plurality of candidate image pairs based on similarities between feature data of different images; each candidate image pair includes a first image and a second image;
[0254] The associated image pairs are screened from the candidate image pairs based on signal data corresponding to the first image and the second image in a plurality of candidate image pairs.
[0255] In an optional embodiment, the processing module 92, when screening the associated image pair from a plurality of candidate image pairs based on the signal data corresponding to the first image and the second image in the plurality of candidate image pairs, is configured to:
[0256] For each candidate image pair among the plurality of candidate image pairs, determining a signal distance between the first image and the second image based on signal data corresponding to the first image in each candidate image pair and signal data corresponding to the second image in each candidate image pair;
[0257] In response to a signal distance between the first image and the second image being less than a preset distance threshold, the candidate image pair is determined as the associated image pair.
[0258] In an optional implementation manner, the signal data includes: signal source identification information corresponding to each frame image in the target video;
[0259] The processing module 92, when determining the signal distance between the first image and the second image based on the signal data corresponding to the first image in each candidate image pair and the signal data corresponding to the second image in each candidate image pair, is configured to:
[0260] determining, based on the signal source identification information corresponding to the first image and the signal source identification information corresponding to the second image, the number of common signal sources corresponding to the first image and the second image;
[0261] Based on the number, a signal distance corresponding to the first image and the second image is determined; wherein the signal distance is negatively correlated with the number of the common signal sources.
[0262] In an optional implementation, the signal data includes: signal strength information corresponding to each frame image in the target video, and signal source identification information;
[0263] The processing module 92, when determining the signal distance between the first image and the second image based on the signal data corresponding to the first image in each candidate image pair and the signal data corresponding to the second image in each candidate image pair, is configured to:
[0264] determining a common signal source and a non-common signal source corresponding to the first image and the second image based on a signal source identifier corresponding to the first image and a signal source identifier corresponding to the second image;
[0265] determining a strength difference of the common signal source based on signal strength information corresponding to the first image and the second image respectively;
[0266] A signal distance between the first image and the second image is determined based on the intensity difference and a preset signal distance corresponding to a non-shared signal source.
[0267] In an optional embodiment, the reconstruction module 93, when performing three-dimensional reconstruction on the target scene based on the associated image pair to obtain the three-dimensional model of the target scene, is configured to:
[0268] For each of the multiple videos, performing three-dimensional reconstruction based on multiple frames of images in each video to obtain a three-dimensional sub-model corresponding to each video;
[0269] Based on the associated image pairs, the three-dimensional sub-models corresponding to the multiple videos are spliced together to obtain the three-dimensional model of the target scene.
[0270] In an optional embodiment, the reconstruction module 93, when splicing the three-dimensional sub-models corresponding to the plurality of videos based on the associated image pairs to obtain the three-dimensional model of the target scene, is configured to:
[0271] For each two videos, determining whether there is a target associated image pair between the two videos; the first associated image and the second associated image in the target associated image pair belong to the two videos respectively;
[0272] In response to the target-associated image pair existing between the two videos, determining, based on the target-associated image pair between the two videos, conversion relationship information between the three-dimensional sub-models corresponding to the two videos, and performing splicing processing on the three-dimensional sub-models corresponding to the two videos based on the conversion relationship information;
[0273] Based on the stitching results corresponding to the multiple videos, a three-dimensional model of the target scene is obtained.
[0274] In an optional embodiment, the reconstruction module 93, when determining, based on the target associated image pair between the two videos, the conversion relationship information between the three-dimensional sub-models corresponding to the two videos, and performing splicing processing on the three-dimensional sub-models corresponding to the two videos based on the conversion relationship information, is configured to:
[0275] Perform at least one iteration of the following:
[0276] Determining a first target-associated image pair corresponding to a current iteration cycle from the target-associated image pairs;
[0277] Determining, based on the first target associated image pair, current conversion relationship information between the three-dimensional sub-models corresponding to the two videos respectively;
[0278] Based on the current conversion relationship information, the three-dimensional sub-models corresponding to the two videos are spliced to obtain a spliced model corresponding to the current iteration cycle;
[0279] Verifying the stitching correctness of the stitching model based on physical positions represented by other target-associated image pairs except the first target-associated image in the target-associated image pairs;
[0280] In response to the splicing correctness verification being passed, the splicing model corresponding to the current iteration cycle is used as the result model of splicing the three-dimensional sub-models corresponding to the two videos, and the iteration process is terminated;
[0281] In response to the correctness verification failing, entering the next iteration cycle;
[0282] If all target-associated image pairs taken as the first target-associated image pairs fail to pass the splicing correctness verification, it indicates that there is no splicing relationship between the three-dimensional sub-models corresponding to the two videos.
[0283] In an optional embodiment, the reconstruction module 93, when performing three-dimensional reconstruction on the target scene based on the associated image pair to obtain the three-dimensional model of the target scene, is configured to:
[0284] Traversing each associated image pair in the associated image pairs, and performing three-dimensional reconstruction based on the traversed associated image pairs to obtain a three-dimensional model corresponding to the traversed associated image pairs;
[0285] The three-dimensional models corresponding to the multiple associated image pairs are fused to obtain the three-dimensional model of the target scene.
[0286] In an optional embodiment, the reconstruction module 93, when performing three-dimensional reconstruction based on the traversed associated image pairs to obtain a three-dimensional model corresponding to the traversed associated image pairs, is configured to:
[0287] Performing three-dimensional reconstruction based on the traversed associated image pairs, the camera poses corresponding to the first image and the second image, and the first image and the second image to obtain a first three-dimensional model;
[0288] determining a reprojection error of the first three-dimensional model based on the first image, the second image, the camera poses corresponding to the first image and the second image, and the first three-dimensional model; and
[0289] determining a signal distance error between the first image and the second image based on signal data corresponding to the first image and the second image, respectively, and a signal source position of a target signal source corresponding to the signal data;
[0290] Calibrate the camera poses corresponding to the first image and the second image, respectively, and the signal source position of the target signal source based on the reprojection error and the signal distance error to obtain the target poses corresponding to the first image and the second image, respectively, and the target position of the target signal source;
[0291] Three-dimensional reconstruction is performed based on the first image, the second image, and the target posture to obtain a three-dimensional model corresponding to the traversed associated image pair.
[0292] In an optional implementation, the reconstruction module 93 is further configured to construct a signal source map based on target positions of target signal sources corresponding to the plurality of associated images.
[0293] For descriptions of the processing flow of each module in the device and the interaction flow between each module, reference can be made to the relevant descriptions in the above method embodiment, which will not be described in detail here.
[0294] The present disclosure also provides a computer device, such as Figure 10 FIG. 1 is a schematic diagram of a computer device structure provided by an embodiment of the present disclosure, including:
[0295] Processor 101 and memory 102; the memory 102 stores machine-readable instructions executable by the processor 101, and the processor 101 is configured to execute the machine-readable instructions stored in the memory 102. When the machine-readable instructions are executed by the processor 101, the processor 101 performs the following steps:
[0296] Acquire multiple videos captured for a target scene, and signal data corresponding to each of the multiple videos; wherein the signal data is obtained by capturing a signal of a target signal source located within the target scene when the videos are captured; each video includes multiple frames of images;
[0297] determining a plurality of associated image pairs from the plurality of videos based on the signal data;
[0298] Based on the associated image pair, the target scene is three-dimensionally reconstructed to obtain a three-dimensional model of the target scene.
[0299] The above-mentioned memory 102 includes internal memory 1021 and external memory 1022; the memory 1021 here is also called internal memory, which is used to temporarily store the calculation data in the processor 101, as well as the data exchanged with the external memory 1022 such as the hard disk. The processor 101 exchanges data with the external memory 1022 through the memory 1021.
[0300] The specific execution process of the above instructions can refer to the steps of the three-dimensional reconstruction method described in the embodiment of the present disclosure, and will not be repeated here.
[0301] The present disclosure also provides a computer-readable storage medium having a computer program stored thereon. When executed by a processor, the computer program executes the steps of the 3D reconstruction method described in the above method embodiment. The storage medium may be a volatile or non-volatile computer-readable storage medium.
[0302] The embodiments of the present disclosure also provide a computer program product, which carries program code. The instructions included in the program code can be used to execute the steps of the three-dimensional reconstruction method described in the above method embodiment. For details, please refer to the above method embodiment and will not be repeated here.
[0303] The computer program product may be implemented in hardware, software, or a combination thereof. In one embodiment, the computer program product is implemented as a computer storage medium. In another embodiment, the computer program product is implemented as a software product, such as a software development kit (SDK).
[0304] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems and devices described above can refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here. In the several embodiments provided in the present disclosure, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0305] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0306] In addition, each functional unit in each embodiment of the present disclosure may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0307] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium that is executable by a processor. Based on this understanding, the technical solution of the present disclosure, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present disclosure. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0308] Finally, it should be noted that the above-described embodiments are only specific implementation methods of the present disclosure, which are used to illustrate the technical solutions of the present disclosure, rather than to limit them. The scope of protection of the present disclosure is not limited thereto. Although the present disclosure has been described in detail with reference to the above-described embodiments, those skilled in the art should understand that any person skilled in the art can modify or easily conceive of changes to the technical solutions described in the above-described embodiments within the technical scope disclosed in the present disclosure, or replace some of the technical features therein with equivalents. Such modifications, changes, or replacements do not deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure, and should be included in the scope of protection of the present disclosure. Therefore, the scope of protection of the present disclosure shall be subject to the scope of protection of the claims.
Claims
1. A three-dimensional reconstruction method, characterized in that: include: Acquire multiple videos captured for a target scene, and signal data corresponding to each of the multiple videos; wherein the signal data is obtained by capturing a signal of a target signal source located within the target scene when the videos are captured; each video includes multiple frames of images; determining a plurality of associated image pairs from the plurality of videos based on the signal data; Based on the associated image pair, the target scene is three-dimensionally reconstructed to obtain a three-dimensional model of the target scene; Wherein, determining a plurality of associated image pairs from the plurality of videos based on the signal data includes: Extract feature data of each frame image in multiple videos; Determining a plurality of candidate image pairs based on similarities between feature data of different images; each candidate image pair includes a first image and a second image; For each candidate image pair among the plurality of candidate image pairs, determining a signal distance between the first image and the second image based on signal data corresponding to the first image in each candidate image pair and signal data corresponding to the second image in each candidate image pair; In response to a signal distance between the first image and the second image being less than a preset distance threshold, the candidate image pair is determined as the associated image pair.
2. The method according to claim 1, characterized in that The signal data includes: signal source identification information corresponding to each frame image in the video; The determining of the signal distance between the first image and the second image based on the signal data corresponding to the first image in each candidate image pair and the signal data corresponding to the second image in each candidate image pair includes: determining, based on the signal source identification information corresponding to the first image and the signal source identification information corresponding to the second image, the number of common signal sources corresponding to the first image and the second image; Based on the number, a signal distance corresponding to the first image and the second image is determined; wherein the signal distance is negatively correlated with the number of the common signal sources.
3. The method according to claim 1, characterized in that The signal data includes: signal strength information corresponding to each frame image in the video, and signal source identification information; The determining of the signal distance between the first image and the second image based on the signal data corresponding to the first image in each candidate image pair and the signal data corresponding to the second image in each candidate image pair includes: determining a common signal source and a non-common signal source corresponding to the first image and the second image based on a signal source identifier corresponding to the first image and a signal source identifier corresponding to the second image; determining a strength difference of the common signal source based on signal strength information corresponding to the first image and the second image respectively; A signal distance between the first image and the second image is determined based on the intensity difference and a preset signal distance corresponding to a non-shared signal source.
4. The method according to any one of claims 1 to 3, characterized in that The step of performing three-dimensional reconstruction on the target scene based on the associated image pair to obtain a three-dimensional model of the target scene includes: For each of the multiple videos, performing three-dimensional reconstruction based on multiple frames of images in each video to obtain a three-dimensional sub-model corresponding to each video; Based on the associated image pairs, the three-dimensional sub-models corresponding to the multiple videos are spliced together to obtain the three-dimensional model of the target scene.
5. The method according to claim 4, characterized in that The step of splicing the three-dimensional sub-models corresponding to the plurality of videos based on the associated image pairs to obtain the three-dimensional model of the target scene includes: For each two videos, determining whether there is a target associated image pair between the two videos; the first associated image and the second associated image in the target associated image pair belong to the two videos respectively; In response to the target-associated image pair existing between the two videos, determining, based on the target-associated image pair between the two videos, conversion relationship information between the three-dimensional sub-models corresponding to the two videos, and performing splicing processing on the three-dimensional sub-models corresponding to the two videos based on the conversion relationship information; Based on the stitching results corresponding to the multiple videos, a three-dimensional model of the target scene is obtained.
6. The method according to claim 5, characterized in that The determining, based on the target associated image pair between the two videos, the conversion relationship information between the three-dimensional sub-models corresponding to the two videos, and performing splicing processing on the three-dimensional sub-models corresponding to the two videos based on the conversion relationship information, includes: Perform at least one iteration of the following: Determining a first target-associated image pair corresponding to a current iteration cycle from the target-associated image pairs; Determining, based on the first target associated image pair, current conversion relationship information between the three-dimensional sub-models corresponding to the two videos respectively; Based on the current conversion relationship information, the three-dimensional sub-models corresponding to the two videos are spliced to obtain a spliced model corresponding to the current iteration cycle; Verifying the stitching correctness of the stitching model based on physical positions represented by other target-associated image pairs except the first target-associated image in the target-associated image pairs; In response to the splicing correctness verification being passed, the splicing model corresponding to the current iteration cycle is used as the result model of splicing the three-dimensional sub-models corresponding to the two videos, and the iteration process is terminated; In response to the correctness verification failing, entering the next iteration cycle; If all target-associated image pairs taken as the first target-associated image pairs fail to pass the splicing correctness verification, it indicates that there is no splicing relationship between the three-dimensional sub-models corresponding to the two videos.
7. The method according to any one of claims 1 to 3, characterized in that The step of performing three-dimensional reconstruction on the target scene based on the associated image pair to obtain a three-dimensional model of the target scene includes: Traversing each associated image pair in the associated image pairs, and performing three-dimensional reconstruction based on the traversed associated image pairs to obtain a three-dimensional model corresponding to the traversed associated image pairs; The three-dimensional models corresponding to the multiple associated image pairs are fused to obtain the three-dimensional model of the target scene.
8. The method according to claim 7, characterized in that The three-dimensional reconstruction based on the traversed associated image pairs to obtain a three-dimensional model corresponding to the traversed associated image pairs includes: Performing three-dimensional reconstruction based on the traversed associated image pairs, the camera poses corresponding to the first image and the second image, and the first image and the second image to obtain a first three-dimensional model; determining a reprojection error of the first three-dimensional model based on the first image, the second image, the camera poses corresponding to the first image and the second image, and the first three-dimensional model; and determining a signal distance error between the first image and the second image based on signal data corresponding to the first image and the second image, respectively, and a signal source position of a target signal source corresponding to the signal data; Calibrate the camera poses corresponding to the first image and the second image, respectively, and the signal source position of the target signal source based on the reprojection error and the signal distance error to obtain the target poses corresponding to the first image and the second image, respectively, and the target position of the target signal source; Three-dimensional reconstruction is performed based on the first image, the second image, and the target posture to obtain a three-dimensional model corresponding to the traversed associated image pair.
9. The method according to claim 8, characterized in that Also includes: A signal source map is constructed based on the target positions of the target signal sources corresponding to the multiple associated images.
10. A three-dimensional reconstruction device, characterized in that: include: an acquisition module, configured to acquire a plurality of videos captured for a target scene, and signal data corresponding to each of the plurality of videos; wherein the signal data is obtained by capturing a signal of a target signal source located within the target scene when the videos are captured; and each video includes a plurality of frames of images; a processing module for determining a plurality of associated image pairs from the plurality of videos based on the signal data; A reconstruction module, configured to perform three-dimensional reconstruction of the target scene based on the associated image pair to obtain a three-dimensional model of the target scene; The processing module, when determining a plurality of associated image pairs from the plurality of videos based on the signal data, is configured to: Extract feature data of each frame image in multiple videos; Determining a plurality of candidate image pairs based on similarities between feature data of different images; each candidate image pair includes a first image and a second image; For each candidate image pair among the plurality of candidate image pairs, determining a signal distance between the first image and the second image based on signal data corresponding to the first image in each candidate image pair and signal data corresponding to the second image in each candidate image pair; In response to a signal distance between the first image and the second image being less than a preset distance threshold, the candidate image pair is determined as the associated image pair.
11. A computer device, characterized in that: include: A processor and a memory, wherein the memory stores machine-readable instructions executable by the processor, and the processor is used to execute the machine-readable instructions stored in the memory. When the machine-readable instructions are executed by the processor, the processor performs the steps of the three-dimensional reconstruction method according to any one of claims 1 to 9.
12. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program. When the computer program is executed by a computer device, the computer device performs the steps of the three-dimensional reconstruction method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Vehicle driving safety detection method and system based on machine vision
CN103522970A
Pose determination method, pose determination device, storage medium and electronic equipment
CN112270710A