3D Reconstruction Data Generation Method, Medium and System
By performing frame-by-frame decimation, feature point matching and filtering of panoramic videos, the problems of inefficiency and low success rate caused by redundant frames in three-dimensional reconstruction are solved, and more efficient and more accurate three-dimensional reconstruction is achieved.
Patent Information
- Application Number
- CN202510360759.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-03-26
AI Technical Summary
The prior art has problems in the process of three-dimensional reconstruction that cause inefficiency and low success rate.
By acquiring panoramic videos, extracting image sequences frame by frame, extracting local feature points, computing adjacent image feature points matching pairs, and performing filtering and loopback marking, removing redundant and occlusion or motion blur images, performing camera viewing angle alignment, and finally generating three-dimensional reconstruction data.
It improves the efficiency and success rate of three-dimensional reconstruction, reduces redundant data, and enhances the accuracy and efficiency of reconstruction.
Smart Images

Figure CN119888092B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of three-dimensional reconstruction, and particularly relates to a method, medium, and system for generating three-dimensional reconstruction data. Background Art
[0002] Three-dimensional reconstruction is a process of processing input image or video data and performing structural reconstruction in a three-dimensional space based on the processed data.
[0003] In related technologies, when generating three-dimensional reconstruction data, video frames in the input data are often directly extracted at fixed intervals for three-dimensional reconstruction based on the extracted video frames. The input data generated in this way has a large number of redundant video frames, reducing the efficiency of three-dimensional reconstruction; moreover, the success rate of three-dimensional reconstruction is relatively low. Summary of the Invention
[0004] The present invention aims to solve at least one of the technical problems in related technologies to some extent. To this end, an object of the present invention is to propose a method for generating three-dimensional reconstruction data, which can effectively improve the efficiency of three-dimensional reconstruction; at the same time, improve the success rate of three-dimensional reconstruction.
[0005] In a first aspect, an embodiment of the present invention proposes a method for generating three-dimensional reconstruction data, including the following steps: obtaining a panoramic video, and extracting frames from the panoramic video one by one to obtain a corresponding first image sequence; extracting local feature points corresponding to each video frame in the first image sequence, and calculating feature point matching pairs between adjacent images based on the local feature points; filtering the first image sequence according to the feature point matching pairs to obtain a second image sequence; performing loop closure marking on the video frames in the second image sequence to obtain three-dimensional reconstruction data.
[0006] According to the method for generating three-dimensional reconstruction data of the embodiment of the present invention, first, a panoramic video is obtained, and frames are extracted from the panoramic video one by one to obtain a corresponding first image sequence; then, local feature points corresponding to each video frame in the first image sequence are extracted, and feature point matching pairs between adjacent images are calculated based on the local feature points; then, the first image sequence is filtered according to the feature point matching pairs to obtain a second image sequence; then, loop closure marking is performed on the video frames in the second image sequence to obtain three-dimensional reconstruction data. Thereby effectively improving the efficiency of three-dimensional reconstruction; at the same time, improving the success rate of three-dimensional reconstruction.
[0007] In some embodiments, the first image sequence is filtered for redundant images according to the following formula:
[0008] ;
[0009] where denotes the number of feature point matching pairs between the th frame and the th frame in the first image sequence, denotes the first preset threshold;
[0010] If the filtering condition is satisfied, then the th frame is removed.
[0011] In some embodiments, the first image sequence is filtered for occluded or motion-blurred images according to the following formula:
[0012] ;
[0013] wherein, denotes the number of feature point matching pairs between the th frame and the th frame in the first image sequence, denotes the second preset threshold;
[0014] If the filtering condition is satisfied, then the video frames between the th frame and the th frame are removed.
[0015] In some embodiments, the second image sequence includes ordinary video frames and reference images. After filtering the first image sequence according to the feature points to obtain the second image sequence, it further includes: calculating the rotation matrix and translation vector of all adjacent images in the second image sequence; for any ordinary video frame, calculating the rotation matrix of the ordinary video frame relative to the reference image according to the rotation matrix of the adjacent images; converting the rotation matrix of the ordinary video frame relative to the reference image into Euler angles, and taking the Euler angles as the yaw angles corresponding to the ordinary video frames; rotating the ordinary video frames according to the yaw angles to align the camera perspectives of the second image sequence.
[0016] In some embodiments, loop labeling the video frames in the second image sequence includes: establishing a landmark relationship graph according to the feature point matching relationship between the video frames in the second image sequence to characterize the co-visibility relationship of the video frames in the second image sequence through the landmark relationship graph; performing clustering based on the landmark relationship graph to obtain multiple image categories; for each video frame in the second image sequence, calculating the distance between the current video frame and each of the image categories, and taking the image category with the smallest distance as the target image category; selecting corresponding video frames from the target image category to perform feature point matching with the current video frame to determine the loop frame corresponding to the current video frame according to the matching result; establishing the adjacent relationship between the current video frame and the loop frame.
[0017] In some embodiments, establishing a landmark relationship graph based on the feature point matching relationships between video frames in the second image sequence includes: for any pair of images formed by two video frames in the second image sequence, initializing a landmark according to the feature point pair of the image pair; for a feature point in any one video frame, determining whether the feature point matches feature points in multiple images; if so, merging the feature points that match the feature point into the same landmark; constructing a graph structure, where the nodes in the graph structure represent video frames or landmarks, and the edges represent the connection relationships between video frames and landmarks.
[0018] In some embodiments, performing clustering based on the landmark relationship graph to obtain multiple image categories includes: for any video frame in the second image sequence, calculating the global feature corresponding to the current video frame; obtaining, according to the global feature, the similar video frames corresponding to the current video frame in the second image sequence; performing clustering based on the landmark relationship graph and the similar video frames to obtain multiple image categories.
[0019] In some embodiments, the method further includes:
[0020] S201, setting the video frames in the second image sequence to the initial state;
[0021] S202, traversing the second image sequence to determine whether there is a video frame with the initial state in the second image sequence; if so, execute step S203; if not, execute step S211;
[0022] S203, taking the video frame with the largest landmark length among the video frames with the initial state as the current video frame;
[0023] S204, obtaining the adjacent images corresponding to the current video frame according to the landmark relationship graph;
[0024] S205, determining whether the number of the adjacent images is 1; if so, execute step S206; if not, execute step S207;
[0025] S206, modifying the state of the current video frame to the reserved state and returning to step S202;
[0026] S207, pre-deleting the current video frame, updating the landmark relationship graph according to the pre-deletion result, and calculating the average landmark length of the updated landmark relationship graph;
[0027] S208, determining whether the average landmark length is greater than or equal to a preset length threshold; if so, execute step S209; if not, execute step S210;
[0028] S209. Delete the current video frame and return to step S202;
[0029] S210. Modify the status of the current video frame to the reserved status and return to step S202;
[0030] S211. Use the video frames with the reserved status as key frames.
[0031] In a second aspect, an embodiment of the present invention provides a computer-readable storage medium, on which a three-dimensional reconstruction data generation program is stored. When the three-dimensional reconstruction data generation program is executed by a processor, the above-mentioned three-dimensional reconstruction data generation method is implemented.
[0032] In a third aspect, an embodiment of the present invention provides a three-dimensional reconstruction data generation system, including: an acquisition module, which is used to acquire a panoramic video and extract frames from the panoramic video one by one to obtain a corresponding first image sequence; an extraction module, which is used to extract local feature points corresponding to each video frame in the first image sequence and calculate feature point matching pairs between adjacent images according to the local feature points; a filtering module, which is used to filter the first image sequence according to the feature point matching pairs to obtain a second image sequence; a marking module, which is used to perform loop marking on the video frames in the second image sequence to obtain three-dimensional reconstruction data.
[0033] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be understood through the practice of the present invention. Description of the Drawings
[0034] Figure 1 is a flowchart of the three-dimensional reconstruction data generation method according to an embodiment of the present invention;
[0035] Figure 2 is a schematic diagram of the camera view alignment effect according to an embodiment of the present invention;
[0036] Figure 3 is a flowchart of key frame extraction according to an embodiment of the present invention;
[0037] Figure 4 is a block diagram of the three-dimensional reconstruction data generation system according to an embodiment of the present invention. Detailed Embodiments
[0038] Embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where like or similar reference numerals denote like or similar elements or elements having like or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are intended to explain the present invention, and should not be construed as limiting the present invention.
[0039] The method for generating three-dimensional reconstruction data according to an embodiment of the present invention will be described below with reference to the accompanying drawings.
[0040] Please refer to Figure 1 , Figure 1 which is a flowchart of the method for generating three-dimensional reconstruction data according to an embodiment of the present invention. As shown in Figure 1 , the method for generating three-dimensional reconstruction data includes the following steps:
[0041] S101: Obtain a panoramic video and extract each frame of the panoramic video to obtain a corresponding first image sequence.
[0042] As an example, extract each frame of the obtained panoramic video to obtain a corresponding first image sequence , where represents the th video frame, and represents the number of video frames in the panoramic video. According to the extraction of each frame of the panoramic video, a first image sequence is obtained and named in sequence. For example, the th frame, and the next frame of the th frame is the th frame; that is, the naming of the images represents the temporal relationship between the video frames.
[0043] S102: Extract local feature points corresponding to each video frame in the first image sequence, and calculate feature point matching pairs between adjacent images according to the local feature points.
[0044] As an example, first, local feature points corresponding to each video frame can be extracted by methods such as SIFT and SuperPoint; then, the feature point matching pairs between adjacent images are the same feature points in adjacent images; the feature point matching pairs between adjacent images can be searched by mutual nearest neighbor search; for example, implemented by methods such as superglue. The number of feature point matching pairs between the th frame and the th frame is represented by .
[0045] S103: Filter the first image sequence according to the feature point matching pairs to obtain a second image sequence.
[0046] In some embodiments, redundant image filtering is performed on the first image sequence according to the following formula:
[0047] ;
[0048] wherein, represents the number of feature point matching pairs between the -th frame and the -th frame in the first image sequence, represents a first preset threshold;
[0049] If the filtering condition is satisfied, the -th frame is removed.
[0050] As an example, first, set the redundant judgment condition as: ;
[0051] Preferably, set the determination threshold (i.e., the first preset threshold) to 5%. If the redundant judgment condition is satisfied, it means that the number of feature matching points between the -th frame and the -th frame is similar; it indicates that the feature points extracted from the -th frame can be replaced by the -th frame. Therefore, it is determined that the -th frame is redundant, and the -th frame is removed. For the case where the redundant judgment condition is satisfied, the -th frame is proposed:
[0052] ;
[0053] For the case where the redundant judgment condition is not satisfied, the -th frame is retained:
[0054] .
[0055] In some embodiments, occlusion or motion blur image filtering is performed on the first image sequence according to the following formula:
[0056] ;
[0057] wherein, represents the number of feature point matching pairs between the -th frame and the -th frame in the first image sequence, represents a second preset threshold;
[0058] If the filtering condition is satisfied, the video frames between the -th frame and the -th frame are removed.
[0059] As an example, first, set the judgment condition for the occluded or motion-blurred image as follows: ;
[0060] Preferably, set the threshold corresponding to this judgment condition (i.e., the second preset threshold) to 90%. Then, for the current frame , , if there exists a frame interval that satisfies: , indicating is much smaller than , which satisfies the judgment condition for the occluded or motion-blurred image, and perform filtering to filter out the video frames between the current frame and ;
[0061] If , indicating and are similar and do not satisfy the filtering condition, retain them. Specifically, assume , and ;
[0062] ...
[0063] ;
[0064] ;
[0065] It can be known that the video frames , ... satisfy the determination condition for the occluded or motion-blurred image and are filtered.
[0066] In some embodiments, the second image sequence includes ordinary video frames and reference images. After filtering the first image sequence according to feature point matching to obtain the second image sequence, it further includes: calculating the rotation matrix and translation vector of all adjacent images in the second image sequence; for any ordinary video frame, calculating the rotation matrix of the ordinary video frame relative to the reference image according to the rotation matrix of the adjacent images; converting the rotation matrix of the ordinary video frame relative to the reference image into Euler angles, and using the Euler angles as the yaw angle corresponding to the ordinary video frame; rotating the ordinary video frame according to the yaw angle to align the camera perspectives of the second image sequence.
[0067] That is to say, align the camera perspectives of the video frames in the second image sequence. In this way, it is beneficial for subsequent image retrieval and matching.
[0068] As an example, first, calculate the rotation matrix and translation vector between all adjacent images in the second image sequence. It should be noted that the rotation matrix and translation vector can be calculated from the feature point matching pairs through the epipolar geometry. The translation vector and rotation matrix are used to recover the rotation matrix and translation vector of adjacent images. Then, denote as the rotation matrix relative to . This rotation matrix represents the relative relationship of the orientations of two video frames in three-dimensional space. Denote as the translation vector relative to . This translation vector represents the relative position relationship of two video frames in three-dimensional space. Then, calculate the rotation matrix of any ordinary video frame relative to the reference image based on the rotation matrix and translation vector between adjacent images: , ; Next, calculate the yaw angle of the ordinary video frame relative to the reference image according to the rotation matrix of the ordinary video frame relative to the reference image. Specifically, convert the rotation matrix of the ordinary video frame relative to the reference image into Euler angles, and the rotation order is ryxz (i.e., rotate in the inner rotation order from y to z to x), and take the yaw angle yaw of the Euler angles to rotate the image according to the yaw angle to achieve camera view alignment. Among them, the view alignment effect is as Figure 2 shown.
[0069] S104. Perform loop closure labeling on the video frames in the second image sequence to obtain three-dimensional reconstruction data.
[0070] In some embodiments, performing loop closure labeling on the video frames in the second image sequence includes: establishing a landmark relationship graph according to the feature point matching relationship between the video frames in the second image sequence to characterize the co-visibility relationship of the video frames in the second image sequence; performing clustering based on the landmark relationship graph to obtain multiple image categories; for each video frame in the second image sequence, calculate the distance between the current video frame and each image category, and take the image category with the smallest distance as the target image category; select the corresponding video frame from the target image category to perform feature point matching with the current video frame to determine the loop closure frame corresponding to the current video frame according to the matching result; establish the adjacent relationship between the current video frame and the loop closure frame.
[0071] In some embodiments, a landmark relationship graph is established based on the feature point matching relationship between video frames in the second image sequence, including: for any pair of video frames in the second image sequence, initialize landmarks according to the feature point pairs of the image pair; for a feature point in any video frame, determine whether the feature point matches the feature points in multiple images; if so, merge the feature points that match the feature point into the same landmark; construct a graph structure, where the nodes in the graph structure represent video frames or landmarks, and the edges represent the connection relationships between video frames and landmarks.
[0072] As an example, first, the landmark Track represents a landmark in three-dimensional space, which can be projected onto feature points in different images to characterize the co-visibility relationship of different images. Since there are multiple feature points in a video frame, define as the th feature point of video frame , and define as the th feature point of video frame . If and are a matching pair of video frame and video frame , then we map them to the same landmark; if and , and are matching pairs, and at the same time, the feature points , and do not have any other feature point threshold matches, then they map to a landmark with a length of 3, that is, landmark track_len = 3, and track_len is the number of feature points corresponding to the landmark. Then, using the definition of the landmark, the landmark relationship graph of all video frames in the second image sequence can be constructed; it should be noted that the landmark relationship graph includes the following information: the connection relationship between video frames and landmarks and the feature point matching pairs of adjacent images.
[0073] Specifically, first, create waypoints for each pair of matched feature points, and traverse all image pairs; for each pair of matched feature points, check whether the pair of feature points belongs to a certain waypoint; if they do not belong to any waypoint, create a new waypoint and add the pair of feature points to the newly added waypoint; if one of the feature points in the pair of feature points already belongs to a certain waypoint, add the other feature point to that waypoint as well; use a hash table or union-find data structure for the data structure; then, expand the waypoints, that is, merge the matched feature points into the same waypoint to ensure that the waypoint contains all co-visible feature points; specifically, traverse all video frames, for each feature point in the current video frame, check whether the feature point matches feature points in multiple images; if the feature point matches multiple feature points, merge all the feature points that match this feature point into the same waypoint, and use the union-find data structure in the data structure to efficiently merge waypoints. Then, construct a waypoint relationship graph; that is, construct a graph structure that describes the connection relationship between images and waypoints; specifically, first initialize a graph structure G, where the nodes represent video frames and waypoints, and the edges represent the connection relationship between video frames and waypoints; then, traverse all waypoints, for each waypoint, find all the feature points connected to it, and connect the video frame and the waypoint to form an edge in graph G; then, traverse all image pairs, if two images share the same waypoint, add an edge in graph G to represent the feature point matching relationship between the two images; use a graph data structure to represent the waypoint relationship graph in the data structure, where the nodes can be images and waypoints, and the edges represent the connection relationship.
[0074] As an example, assume there is the following video frame sequence and pairs of matched feature points: Feature points of image match those of image ; Feature points of image match those of image ; Feature points of image match those of image ; Feature points of image match those of image ; Feature points of image match those of image ; Feature points of image match those of image ; Then, initialize the waypoints, create waypoint , and add and to ; Create waypoint , and add and to ; Create waypoint , and add and Add to ; then, expand the road punctuation points and check if there are feature points to be merged; assume that and belong to the same road punctuation point, then merge and ; then, construct a road punctuation point relationship graph; Image is connected to and ; Image is connected to the road punctuation point and ; Image is connected to the road punctuation point and .
[0075] In some embodiments, clustering is performed based on the road punctuation point relationship graph to obtain multiple image categories, including: for any video frame in the second image sequence, calculate the global feature corresponding to the current video frame; obtain the similar video frames corresponding to the current video frame in the second image sequence according to the global feature; perform clustering based on the road punctuation point relationship graph and the similar video frames to obtain multiple image categories.
[0076] As an example, first, for any video frame, if there are image frames that are far apart in time but have similar spatial positions and orientations, then such an image frame is the loopback frame of the current video frame; it can be understood that the loopback frame can give some constraints with a longer time interval in addition to adjacent frames, reducing the cumulative error problem caused by long-term movement.
[0077] Specifically, first, calculate the global features of all video frames. The global feature is the overall attribute of the image, and images with similar colors, textures, and shapes will also have relatively similar global features. The global feature can be calculated using methods such as BOW and Netvald; then, using the global feature, calculate the K most similar images of each video frame in the second image sequence (excluding the K nearest neighbors of the current video frame, and the specific value of K can be selected according to needs); then, cluster the K most similar images according to the road punctuation point relationship graph, and cluster the images with common road punctuation points into one category.
[0078] Assume that the value of K is 8, and the eight most similar images are f1, f3, f4, f6, f20, f25, f27, and f45, and the associated road punctuation points are 1, 2, 3, 100, 102, and 350. First, initialize the image class names with the serial numbers of all road punctuation points connected to each image in the K most similar images. Each image can correspond to multiple class names (for example, f1 corresponds to two class names, 1 and 2); then, sort the class names of each image and map all the class names (except the smallest class name) corresponding to the image to the smallest class name corresponding to the image (such as f3, and the corresponding class names 2 and 3 are both mapped to 1).); then, merge all the class name mappings until there is no mapping value for the corresponding class name; then, for the mapping of 3->2, the mapping value is 2. Since there is a mapping of 2->1 in the class name mapping, merge them to get 3->1, and its mapping value is 1. At this time. There is no mapping value of 1 in the class name mapping, and the merging ends. The merged class name mapping is used as the final class name until each image corresponds to only one class name. At this time, this class name is the cluster to which the image belongs (such as the class names 1, 2, and 3 corresponding to f3 are modified to 1, 1, and 1, so the cluster name of f3 is 1); then, calculate the distance of each cluster for the current frame to:
[0079] ;
[0080] wherein, represents the set of all images in each cluster, represents the set the number of images in;
[0081] Then, from the cluster with the smallest distance to the current frame , take out the five images with the highest similarity and perform feature point matching with the current frame respectively; then, record the image with the most matching points as the loopback frame of the current frame and establish the adjacent relationship between the current frame and the loopback frame.
[0082] Specifically, first, increase the feature point matching pairs between the current frame and the loopback frame; then, replace the road punctuation points of the loopback frame , that is, for and all matching pairs of and , delete the corresponding road punctuation point and connect it to the corresponding road punctuation point.
[0083] In some embodiments, as Figure 3 shown, the method further includes:
[0084] S201. Set the video frames in the second image sequence to the initial state;
[0085] S202. Traverse the second image sequence to determine whether there is a video frame with the initial state in the second image sequence; if so, execute step S203; if not, execute step S211;
[0086] S203. Take the video frame with the largest road punctuation length among the video frames with the initial state as the current video frame;
[0087] S204. Obtain the adjacent image corresponding to the current video frame according to the road punctuation relationship graph;
[0088] S205. Determine whether the number of adjacent images is 1; if so, execute step S206; if not, execute step S207;
[0089] S206. Modify the state of the current video frame to the reserved state and return to step S202;
[0090] S207. Perform pre-deletion on the current video frame, update the road punctuation relationship graph according to the pre-deletion result, and calculate the average road punctuation length of the updated road punctuation relationship graph;
[0091] S208. Determine whether the average road punctuation length is greater than or equal to the preset length threshold; if so, execute step S209; if not, execute step S210;
[0092] S209. Delete the current video frame and return to step S202;
[0093] S210. Modify the state of the current video frame to the reserved state and return to step S202;
[0094] S211. Take the video frames with the reserved state as key frames.
[0095] As an example, first, for the image frame , the number of its local feature points is , is the length of the road punctuation corresponding to the local feature point . The road punctuation length of the image frame is:
[0096] ;
[0097] It should be noted that the road punctuation length threshold TL can be set according to the mapping fineness requirement. The larger the value of TL, the finer the mapping result.
[0098] Next, set all images to the initialization state, which indicates that the images are not processed; then, select the image fp with the largest road punctuation length from the images in the initialization state, and the adjacent images of fp. The adjacent images are the images with a matching pair relationship in the road punctuation relationship graph, and the number is greater than one; then, if the number of adjacent images is only one, mark fp as the reserved state; if the number of adjacent images is greater than one, remove the image with the largest road punctuation length from the road punctuation relationship graph and regenerate the road punctuation relationship graph; that is to say, this step attempts to delete the image with the largest road punctuation length; if the road punctuation relationship graph can meet the set conditions after deletion, discard the image; if the road punctuation relationship graph cannot meet the set conditions, keep the image. Specifically, perform image feature point matching on pairwise combinations of the image frames in the adjacent images. Define the connection condition: for any one of the matching pairs , if , then the image meets the condition for establishing a feature point matching pair relationship; select the two images with the largest number of matches from the adjacent images, and if the connection condition is met, connect the two images with the largest number of matches as described above, that is, establish a feature point matching pair relationship between the two images and mark these two as the connected state. If the connection condition is not met, mark the extracted fp as the reserved (keep) state; then, select the image with the largest number of matches from the unconnected adjacent images and the images in the connected state. If the connection condition is met, connect them and mark them as connected. If the connection condition is not met, mark the extracted fp as the reserved (keep) state; repeat the above steps until all adjacent images are marked as connected; calculate the average road punctuation length of the newly generated road punctuation relationship graph; if the average road punctuation length is greater than the road punctuation length threshold, set the extracted fp to the discarded state. Finally, for all images in the reserved state, use them as the key frames for the three-dimensional reconstruction input.
[0099] In summary, according to the method for generating three-dimensional reconstruction data of the embodiments of the present invention, first, obtain a panoramic video and extract each frame of the panoramic video to obtain a corresponding first image sequence; then, extract the local feature points corresponding to each video frame in the first image sequence and calculate the feature point matching pairs between adjacent images according to the local feature points; then, filter the first image sequence according to the feature point matching pairs to obtain a second image sequence; then, perform loop closure marking on the video frames in the second image sequence to obtain three-dimensional reconstruction data. Thereby effectively improving the efficiency of three-dimensional reconstruction; at the same time, improving the success rate of three-dimensional reconstruction.
[0100] In a second aspect, an embodiment of the present invention proposes a computer-readable storage medium, on which a three-dimensional reconstruction data generation program is stored. When the three-dimensional reconstruction data generation program is executed by a processor, the above-mentioned method for generating three-dimensional reconstruction data is implemented.
[0101] In a third aspect, an embodiment of the present invention provides a three-dimensional reconstruction data generation system. As Figure 4 shown, the three-dimensional reconstruction data generation system includes: an acquisition module 10, an extraction module 20, a filtering module 30, and a marking module 40.
[0102] Among them, the acquisition module 10 is configured to acquire a panoramic video and extract each frame of the panoramic video to obtain a corresponding first image sequence;
[0103] The extraction module 20 is configured to extract local feature points corresponding to each video frame in the first image sequence and calculate feature point matching pairs between adjacent images according to the local feature points;
[0104] The filtering module 30 is configured to filter the first image sequence according to the feature point matching pairs to obtain a second image sequence;
[0105] The marking module 40 is configured to perform loop marking on the video frames in the second image sequence to obtain three-dimensional reconstruction data.
[0106] In addition, it should be noted that the above description of the three-dimensional reconstruction data generation method also applies to this three-dimensional reconstruction data generation system and will not be elaborated here.
[0107] It should be noted that the logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch instructions from the instruction execution system, apparatus, or device and execute the instructions), or in combination with these instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in combination with an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection portion with one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). In addition, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other suitable processing as necessary, and then stored in a computer memory.
[0108] It should be understood that each part of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0109] In the description of this specification, the descriptions referring to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0110] In the description of the present invention, it should be understood that the orientation or positional relationships indicated by the terms "center", "longitudinal", "transverse", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", "axial", "radial", "circumferential", etc. are based on the orientation or positional relationships shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation of the present invention.
[0111] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the present invention, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise specifically and clearly defined.
[0112] In the present invention, unless otherwise clearly specified or limited, the terms "mounted", "connected", "coupled", "fixed", etc. shall be construed in a broad sense. For example, it may be a fixed connection, a detachable connection, or integrated; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the internal connection of two components or the interaction relationship between two components, unless otherwise clearly defined. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0113] In the present invention, unless otherwise clearly specified or limited, the first feature being "on" or "under" the second feature may be that the first and second features are in direct contact, or the first and second features are indirectly in contact through an intermediate medium. Moreover, the first feature being "above", "over" and "on top of" the second feature may be that the first feature is directly above or obliquely above the second feature, or merely indicates that the first feature has a higher horizontal height than the second feature. The first feature being "under", "below" and "beneath" the second feature may be that the first feature is directly below or obliquely below the second feature, or merely indicates that the first feature has a lower horizontal height than the second feature.
[0114] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.
Claims
1. A method for generating three-dimensional reconstruction data, characterized in that, It includes the following steps: Obtain a panoramic video, and extract each frame of the panoramic video to obtain a corresponding first image sequence; Extract local feature points corresponding to each video frame in the first image sequence, and calculate feature point matching pairs between adjacent images according to the local feature points; Filter the first image sequence according to the feature point matching pairs to obtain a second image sequence; Perform loop closure labeling on the video frames in the second image sequence to obtain 3D reconstruction data; Among them, redundant image filtering of the first image sequence is performed according to the following formula: ; Among them, represents the number of feature point matching pairs between the -th frame and the -th frame in the first image sequence, represents the first preset threshold; If the filtering condition is satisfied, the frame is removed; Among them, occluded or motion-blurred image filtering of the first image sequence is performed according to the following formula: ; Among them, represents the number of feature point matching pairs between the -th frame and the -th frame in the first image sequence, represents the second preset threshold; If the filtering condition is satisfied, the video frames between the th frame and the th frame are removed; Among them, performing loop closure labeling on the video frames in the second image sequence includes: Establish a landmark relationship graph according to the feature point matching relationship between video frames in the second image sequence, so as to characterize the co-visibility relationship of video frames in the second image sequence through the landmark relationship graph; Perform clustering based on the landmark relationship graph to obtain multiple image categories; For each video frame in the second image sequence, calculate the distance between the current video frame and each of the image categories, and take the image category with the smallest distance as the target image category; Select corresponding video frames from the target image category to perform feature point matching with the current video frame, so as to determine the loop closure frame corresponding to the current video frame according to the matching result; Establish the adjacent relationship between the current video frame and the loop closure frame.
2. The three-dimensional reconstruction data generation method according to claim 1, wherein The second image sequence includes ordinary video frames and reference images. After filtering the first image sequence according to the feature point matching pairs to obtain the second image sequence, it further includes: Calculate the rotation matrix and translation vector of all adjacent images in the second image sequence; For any ordinary video frame, calculate the rotation matrix of the ordinary video frame relative to the reference image according to the rotation matrix of the adjacent images; Convert the rotation matrix of the ordinary video frame relative to the reference image into Euler angles, and use the Euler angles as the yaw angles corresponding to the ordinary video frame; Rotate the ordinary video frame according to the yaw angle to align the camera perspectives of the second image sequence.
3. The three-dimensional reconstruction data generation method according to claim 1, wherein, Establishing a landmark relationship graph according to the feature point matching relationship between video frames in the second image sequence includes: For an image pair formed by any two video frames in the second image sequence, initialize landmarks according to the feature point pairs of the image pair; For a feature point in any video frame, determine whether the feature point matches feature points in multiple images; If so, merge the feature points that match the feature point into the same landmark; Construct a graph structure, where the nodes in the graph structure represent video frames or landmarks, and the edges represent the connection relationships between video frames and landmarks.
4. The three-dimensional reconstruction data generation method according to claim 1, wherein Performing clustering based on the landmark relationship graph to obtain multiple image categories includes: For any video frame in the second image sequence, calculate the global feature corresponding to the current video frame; Obtain the similar video frames corresponding to the current video frame in the second image sequence according to the global feature; Cluster based on the road punctuation relationship graph and the similar video frames to obtain multiple image categories.
5. The three-dimensional reconstruction data generation method according to claim 3, wherein The method further includes: S201, set the video frames in the second image sequence to the initial state; S202, traverse the second image sequence to determine whether there is a video frame with the initial state in the second image sequence; if yes, execute step S203; if no, execute step S211; S203, take the video frame with the largest road punctuation length among the video frames with the initial state as the current video frame; S204, obtain the adjacent image corresponding to the current video frame according to the road punctuation relationship graph; S205, determine whether the number of the adjacent images is 1; if yes, execute step S206; if no, execute step S207; S206, modify the state of the current video frame to the reserved state and return to step S202; S207, perform pre-deletion on the current video frame, update the road punctuation relationship graph according to the pre-deletion result, and calculate the average road punctuation length of the updated road punctuation relationship graph; S208, determine whether the average road punctuation length is greater than or equal to the preset length threshold; if yes, execute step S209; if no, execute step S210; S209, delete the current video frame and return to step S202; S210, modify the state of the current video frame to the reserved state and return to step S202; S211, use the video frames with the reserved state as key frames.
6. A computer-readable storage medium, characterized in that, A three-dimensional reconstruction data generation program is stored thereon, and when the three-dimensional reconstruction data generation program is executed by a processor, the three-dimensional reconstruction data generation method according to any one of claims 1-5 is implemented.
7. A three-dimensional reconstruction data generation system, characterized in that Including: An acquisition module, which is used to acquire a panoramic video and extract each frame of the panoramic video to obtain a corresponding first image sequence; An extraction module, which is used to extract the local feature points corresponding to each video frame in the first image sequence and calculate the feature point matching pairs between adjacent images according to the local feature points; A filtering module, which is used to filter the first image sequence according to the feature point matching pairs to obtain a second image sequence; A marking module, which is used to perform loop marking on the video frames in the second image sequence to obtain three-dimensional reconstruction data; Wherein, the first image sequence is filtered for redundant images according to the following formula: ; Among them, represents the number of feature point matching pairs between the -th frame and the -th frame in the first image sequence, represents the first preset threshold; If the filtering condition is satisfied, the frame is removed; Wherein, the first image sequence is filtered for occluded or motion-blurred images according to the following formula: ; Among them, represents the number of feature point matching pairs between the th frame and the th frame in the first image sequence, represents the second preset threshold; If the filtering condition is satisfied, the video frames between the th frame and the th frame are removed; Wherein, performing loop marking on the video frames in the second image sequence includes: Establish a road punctuation relationship graph according to the feature point matching relationship between the video frames in the second image sequence, so as to characterize the co-visibility relationship of the video frames in the second image sequence through the road punctuation relationship graph; Cluster based on the road punctuation relationship graph to obtain multiple image categories; For each video frame in the second image sequence, calculate the distance between the current video frame and each of the image categories, and take the image category with the smallest distance as the target image category; Select corresponding video frames from the target image categories and perform feature point matching with the current video frame to determine the loopback frame corresponding to the current video frame according to the matching result; Establish the adjacent relationship between the current video frame and the loopback frame.
Citation Information
Patent Citations
Image matching method and device, equipment and storage medium
CN114511719A
Electric power robot operation method based on real-time three-dimensional environment reconstruction and related equipment
CN116152455A