Self-adaptive image matching pair selection method and device based on global description
Through the adaptive image matching pairing selection method based on global description, the problem of redundant and wrong matching pairing in three-dimensional reconstruction is solved, the efficiency and accuracy of three-dimensional reconstruction are improved, and high-quality three-dimensional reconstruction is achieved.
Patent Information
- Application Number
- CN202510085763.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-16
AI Technical Summary
In large and complex environments, there are a large number of redundant and incorrect image matching pairs in three-dimensional reconstruction technology, making it difficult for the accuracy and efficiency of three-dimensional reconstruction to meet the high-precision needs.
Adaptive image matching pair selection method based on global description is adopted, and match pairs are selected adaptively through NETVLAD global description extraction, global description distance calculation, distance matrix binarization, corrosion expansion and clustering processing, so as to reduce the proportion of wrong match pairs and retain the effective match pairs.
The efficiency and accuracy of 3D reconstruction are improved, the calculation amount is reduced, and the simplicity of matching pairs is enhanced, thereby improving the quality of 3D reconstruction.
Smart Images

Figure CN120014302A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular to a method and device for selecting adaptive image matching pairs based on global description. Background Art
[0002] At present, 3D reconstruction technology (SFM) is an important branch of computer vision. It is the process of converting real-world objects or scenes into 3D models through computer vision or sensor technology. It has a wide range of applications, including virtual reality, autonomous driving, robot navigation, industrial inspection and other fields. The SFM method extracts feature points from multiple images and reconstructs 3D points using the relative position and posture between cameras. The process includes feature point detection, image pair matching, camera posture estimation, 3D point triangulation and global optimization. Although it is now possible to obtain good reconstruction effects in static ordinary scenes, in large and complex environments, the accuracy and efficiency are often difficult to meet the needs of high-precision 3D reconstruction due to a large number of image matching pairs and false matches.
[0003] One of the core of 3D reconstruction is to establish a rich, complete and accurate image matching relationship. For SFM, if all image pairs can be matched once, the established matching relationship is undoubtedly the richest, but there are: a. Huge amount of calculation, many redundant matching pairs. b. A large number of wrong matches will interfere with 3D reconstruction. How to reasonably select matches becomes a core key issue. When there are omissions, multiple sub-maps will appear, there is no effective closed loop, resulting in low map construction accuracy; when there is a large amount of redundancy, the amount of calculation will increase out of thin air, resulting in very slow matching, and a large number of wrong matches may also occur.
[0004] The Colmap algorithm is currently recognized as a better 3D reconstruction strategy in the industry. Most of the 3D reconstruction software in use, including commercial software, is developed based on it. The Colmap algorithm provides 6 image pair matching strategies: 1. Brute force matching: any image is matched once, and the matching pairs are rich and complete, but there are huge computational complexity and many useless matching pairs; a large number of wrong matches occur, especially the problem of individual repeated textures; 2. Sequence matching: when the data is a sequence image, matching and association are performed with the previous and next frames based on the sequence relationship. This matching strategy can ensure linear association, but loops cannot be found and the trajectory is prone to "drift"; 3. Bag of words matching: bag of words matching is performed based on voc_ab_tree to find association relationships. This method can find potential matching relationships and is relatively stable, but a large number of matching pairs will still appear when the textures are similar; 4. Spatial matching: if there is a priori such as GPS in advance, the matching pairs can be constrained within a certain range based on this, for example: each frame is only matched with images within a range of 10m around it, which requires GPS priori and is not satisfied in most cases; 5. Expansion matching: based on the existing matched images, enrich and expand, for example, if there are existing AB and BC matches, try AC matching to enrich the image. This method is based on the expansion of existing image pairs and cannot be used alone; 6. User-defined: purely manual design, requiring a lot of user intervention. Summary of the invention
[0005] In response to the above-mentioned defects, an embodiment of the present invention discloses a method and device for selecting adaptive image matching pairs based on global description, which can extract enough matching pairs, retain valid matches while controlling the final number of matching pairs, ensure the streamlining of matching pairs, and greatly improve the efficiency of three-dimensional reconstruction.
[0006] A first aspect of an embodiment of the present invention discloses a method for selecting adaptive image matching pairs based on global description, comprising:
[0007] Obtain a target image set, and perform NETVLAD global description extraction on each image in the target image set to obtain a global description definition of each image;
[0008] Performing global description distance calculation on all image matching pairs in the target image set according to the global description definition, and constructing a distance matrix diagram according to the global description distance calculation;
[0009] Binarizing the distance matrix graph;
[0010] The binary distance matrix graph is corroded, expanded and clustered to obtain different matching pair clusters;
[0011] Adaptively pick matching pairs from different matching pair clusters.
[0012] As an optional implementation, in the first aspect of the embodiment of the present invention, performing global description distance calculation on all image matching pairs in the target image set according to the global description definition includes:
[0013] Define the NETVLAD description of each image as a vector D i , the similarity between any two images in the target image set is defined as S ij ;
[0014] By formula The similarity between any two images in the target image set is calculated as the global description distance, where D i The transpose of .
[0015] As an optional implementation manner, in the first aspect of the embodiment of the present invention, the step of calculating and constructing a distance matrix diagram according to the global description distance includes:
[0016] The total number of all images in the target image set is obtained as N, and the similarity between any two images is used as the matrix element to construct an N*N distance matrix diagram: Among them, 1 to N represent the image sequence number, S NN Represents the image at row N and column N.
[0017] As an optional implementation manner, in the first aspect of the embodiment of the present invention, binarizing the distance matrix graph includes:
[0018] A binarization threshold is set, and the distance matrix image is binarized based on the binarization threshold, so that the distance matrix image is processed into a black and white image.
[0019] As an optional implementation, in the first aspect of the embodiment of the present invention, the binarization threshold is 1.2.
[0020] As an optional implementation manner, in the first aspect of the embodiment of the present invention, adaptively selecting matching pairs from different image matching pair clusters includes:
[0021] Define any image as frame F i , select the image range as [F i-m ,F i-1 ],[F i+1 ,F i+m ] and the image frame F i Select matching pairs, F i-m represents the im-th frame image, that is, the m-th frame image on the left side of the i-th frame image; F i+m represents the i+mth frame image, that is, the mth frame image on the right side of the ith frame image; F i-1represents the i-1th frame image, that is, the first frame image to the left of the i-th frame image; F i+1 It represents the i+1th frame image, that is, the first frame image to the right of the i-th frame image.
[0022] As an optional implementation manner, in the first aspect of the embodiment of the present invention, it also includes:
[0023] Set the maximum number of matching pairs for each image to N max , the category of the image matching pair cluster is Q, and the weight of the category Q of the image matching pair cluster is the number of images in the category of the image matching pair cluster w i , then the total weight of the image matching pair cluster is
[0024] The number of image pairs is selected from each image category according to the formula:
[0025] Select the class P with the highest similarity from each image matching pair cluster i image pairs.
[0026] A second aspect of an embodiment of the present invention discloses an adaptive image matching pair selection device based on global description, comprising:
[0027] Image acquisition module: used to acquire the target image set, and perform NETVLAD global description extraction on each image in the target image set to obtain the global description definition of each image;
[0028] Image calculation module: used for performing global description distance calculation on all image matching pairs in the target image set according to the global description definition, and constructing a distance matrix diagram according to the global description distance calculation;
[0029] Image processing module: used for binarizing the distance matrix graph;
[0030] Category division module: used to perform corrosion, expansion and clustering processing on the binarized distance matrix graph to obtain different image matching pair clusters;
[0031] Image matching module: used to adaptively select matching pairs from different image matching pair clusters.
[0032] As an optional implementation, in the second aspect of the embodiment of the present invention, performing global description distance calculation on all image matching pairs in the target image set according to the global description definition includes:
[0033] Define the NETVLAD description of each image as a vector D i , the similarity between any two images in the target image set is defined as S ij ;
[0034] By formula The similarity between any two images in the target image set is calculated as the global description distance, where D i The transpose of .
[0035] As an optional implementation manner, in the second aspect of the embodiment of the present invention, the step of calculating and constructing a distance matrix diagram according to the global description distance includes:
[0036] Get the total number of all images in the target image set as N, and use the similarity between any two images as the matrix element to construct an N*N distance matrix diagram Among them, 1 to N represent the image sequence number, S NN Represents the image at row N and column N.
[0037] As an optional implementation manner, in the second aspect of the embodiment of the present invention, binarizing the distance matrix graph includes:
[0038] A binarization threshold is set, and the distance matrix image is binarized based on the binarization threshold, so that the distance matrix image is processed into a black and white image.
[0039] As an optional implementation, in the second aspect of the embodiment of the present invention, the binarization threshold is 1.2.
[0040] As an optional implementation manner, in the second aspect of the embodiment of the present invention, adaptively selecting matching pairs from different image matching pair clusters includes:
[0041] Define any image as frame F i , select the image range as [F i-m ,F i-1 ],[F i+1 ,F i+m ] and the image frame F i Select matching pairs, F i-m represents the im-th frame image, that is, the m-th frame image on the left side of the i-th frame image; F i+m represents the i+mth frame image, that is, the mth frame image on the right side of the ith frame image; F i-1 represents the i-1th frame image, that is, the first frame image to the left of the i-th frame image; F i+1 It represents the i+1th frame image, that is, the first frame image to the right of the i-th frame image.
[0042] As an optional implementation manner, in the second aspect of the embodiment of the present invention, it also includes:
[0043] Set the maximum number of matching pairs for each image to N max, the category of the image matching pair cluster is Q, and the weight of the category Q of the image matching pair cluster is the number of images in the image matching pair cluster w i , then the total weight of the image matching pair in the cluster is
[0044] The number of image pairs is selected from each image category according to the formula:
[0045] Select the category P with the highest similarity from each image matching pair cluster i image matching pairs.
[0046] A third aspect of an embodiment of the present invention discloses an electronic device, comprising: a memory storing executable program code; a processor coupled to the memory; the processor calls the executable program code stored in the memory to execute the adaptive image matching pair selection method based on global description disclosed in the first aspect of the embodiment of the present invention.
[0047] A fourth aspect of an embodiment of the present invention discloses a computer-readable storage medium storing a computer program, wherein the computer program enables a computer to execute the adaptive image matching pair selection method based on global description disclosed in the first aspect of an embodiment of the present invention.
[0048] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:
[0049] The adaptive image matching pair selection method based on global description provided by the embodiment of the present invention obtains a target image set, extracts a global description for each image matching pair and obtains a global description definition, performs global description distance calculation on all images according to the global description definition to construct a distance matrix diagram, binarizes, erodes, dilates and clusters the distance matrix diagram to obtain different image matching pair clusters, and finally adaptively selects matching pairs from different image matching pair clusters; the embodiment adaptively selects matching pairs from all image matches through global description, reduces the proportion of erroneous matching pairs to a certain extent, can extract enough matching pairs, retain valid matches, and control the number of final matching pairs, ensure the streamlining of matching pairs, accelerate the three-dimensional reconstruction process, and improve the quality of three-dimensional reconstruction in terms of efficiency and accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0051] Figure 1It is a flowchart of a method for selecting adaptive image matching pairs based on global description disclosed in an embodiment of the present invention;
[0052] Figure 2 is a structural schematic diagram of a global description-based adaptive image matching pair selection device provided by an embodiment of the present invention;
[0053] Figure 3 is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention;
[0054] Figure 4 It is a schematic diagram of the distance matrix diagram after being binarized provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0055] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0056] It should be noted that the terms "first", "second", "third", "fourth", etc. in the specification and claims of the present invention are used to distinguish different objects rather than to describe a specific order. The terms "including" and "having" in the embodiments of the present invention and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device including a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0057] The embodiment of the present invention discloses a method, device, electronic device and storage medium for adaptively selecting image matching pairs based on global description. The method comprises the following steps: obtaining a target image set, extracting a global description of each image and obtaining a global description definition; performing global description distance calculation on all image matching pairs according to the global description definition to construct a distance matrix diagram; binarizing, corroding, dilating and clustering the distance matrix diagram to obtain different image matching pair clusters; and finally adaptively selecting matching pairs from different image matching pair clusters. The embodiment adaptively selects matching pairs from all image matches through global description, thereby reducing the proportion of erroneous matching pairs to a certain extent, being able to extract enough matching pairs, retaining valid matching, and controlling the number of final matching pairs, thereby ensuring the streamlining of matching pairs, accelerating the three-dimensional reconstruction process, and improving the quality of three-dimensional reconstruction in terms of efficiency and accuracy.
[0058] Embodiment 1
[0059] See also Figure 1 , Figure 1 It is a flow chart of an adaptive image matching pair selection method based on global description disclosed in an embodiment of the present invention. Among them, the execution subject of the method described in the embodiment of the present invention is an execution subject composed of software and / or hardware, and the execution subject can receive relevant information by wired or / and wireless means, and can send certain instructions. Of course, it can also have certain processing functions and storage functions. The execution subject can control multiple devices, such as a remote physical server or cloud server and related software, or it can be a local host or server and related software that performs related operations on a device placed somewhere. In some scenarios, multiple storage devices can also be controlled, and the storage devices can be placed in the same place or different places as the devices. For example Figure 1 As shown, the adaptive image matching pair selection method based on global description includes the following steps:
[0060] 101. Obtain a target image set, and perform NETVLAD global description extraction on each image in the target image set to obtain a global description definition of each image.
[0061] The target image set is the set from which the present invention extracts images and performs image matching. In the embodiment, firstly, a series of images related to the target scene are collected or generated to form the target image set, and these images may be from different viewing angles, lighting conditions or time points.
[0062] NETVLAD (Network of Vector of Locally Aggregated Descriptors) is a deep learning method that is particularly suitable for visual clustering tasks of images and videos, as well as applications such as scene recognition and location revisiting. It works by locally aggregating descriptors in feature space and then combining these aggregated results into a global vector to capture the global characteristics of a scene or object. The core of NETVLAD is its network architecture, which combines convolutional neural networks (CNNs) with the VLAD method.
[0063] The embodiment uses the NETVLAD algorithm to extract a global description for each image in the target image set, which has particular flexibility and ease of use.
[0064] 102. Perform global description distance calculation on all image matching pairs in the target image set according to the global description definition, and construct a distance matrix diagram according to the global description distance calculation.
[0065] In the embodiment, the global description is used to calculate the distances between all images in the target image set, and a distance matrix diagram is constructed based on the calculated global description distances. In this matrix diagram, the i-th row and the j-th column represent the distance between the i-th frame image and the j-th frame image.
[0066] In this step, the global description distance calculation is performed on all image matching pairs in the target image set according to the global description definition, including: defining the NETVLAD description of each image as a vector D i , the similarity between any two images in the target image set is defined as S ij ; Through the formula The similarity between any two images in the target image set is calculated as the global description distance, where D i The transpose of . The calculation result is the distance, and the similarity means that the smaller the distance, the higher the similarity.
[0067] Furthermore, the total number of all images in the target image set is obtained as N, and the similarity between any two images is used as the matrix element to construct an N*N distance matrix diagram: Among them, 1 to N represent the image sequence number, S NN Represents the image at row N and column N.
[0068] 103. Binarize the distance matrix graph.
[0069] Binarize the distance matrix, that is, convert the elements in the distance matrix to 0 or 1 according to a preset threshold. Binarization is an image processing technique, the core idea of which is to set the grayscale value of each pixel in the image to 0 or 255 (in an 8-bit grayscale image), thereby obtaining an image with only black and white colors. Binarization divides the pixels in the image into two categories by setting a threshold. Pixels with grayscale values greater than the threshold are set to white (255), while pixels with grayscale values less than or equal to the threshold are set to black (0).
[0070] In this step, a binarization threshold is set, and the distance matrix is binarized based on the binarization threshold, so that the distance matrix is processed into a black and white image. Specifically, the binarization threshold is 1.2. That is, after binarization by the binarization threshold, values greater than the binarization threshold are white, and values less than the binarization threshold are black. That is, in the image, 1 corresponds to white, and 0 corresponds to black. Figure 4 shown.
[0071] 104. The binarized distance matrix graph is corroded, expanded and clustered to obtain different image matching pair clusters.
[0072] Erosion, dilation, and clustering operations are all implemented based on OpenCV. Erosion is an operation that erodes the foreground boundary of an image, and is usually used to remove small noise points or separate objects. Dilation is an operation that expands the foreground boundary of an image, and is usually used to fill small holes or connect adjacent objects. OpenCV (Open Source Computer Vision Library) is an open source computer vision and machine learning software library that provides a large number of algorithms for computer vision, image processing, and pattern recognition, covering real-time image processing, video analysis, feature detection, target tracking, face recognition, object recognition, image segmentation, optical flow, stereo vision, motion estimation, machine learning, and deep learning. Specifically, the erosion parameter of the embodiment is 5px and the dilation parameter is 10px, where px is a pixel.
[0073] 105. Adaptively select matching pairs from different image matching degree clusters.
[0074] In this step, matching pairs are selected from the perspectives of sequence matching and loop matching. Specifically, all diagonal adjacent neighborhoods are matching pairs. Define any image as frame F i , from the image frame F i The image range selected from the images in the cluster of the same image matching pair is [F i-m ,F i-1 ],[F i+1 ,F i+m ] and the image frame F i Select matching pairs. In this step, 2*m matching pairs are selected in the diagonal line according to a fixed selection method, that is, the image of the i-th frame and the images of the left m frames and the right m frames all constitute matching pairs. i-m represents the im-th frame image, that is, the m-th frame image on the left side of the i-th frame image; F i+m represents the i+mth frame image, that is, the mth frame image on the right side of the ith frame image; F i-1 represents the i-1th frame image, that is, the first frame image to the left of the i-th frame image; F i+1 Represents the i+1th frame image, that is, the first frame image to the right of the i-th frame image
[0075] Then, set the maximum number of matching pairs for each image to N max , the category of the image matching pair cluster is Q, and the weight of the category Q of the image matching pair cluster is the number of images in the image matching pair cluster w i , then the total weight of the image matching pair cluster is The number of image pairs is selected from each image category according to the formula: Select the class P with the highest similarity from each image matching pair cluster iThe goal of this step is to select a controllable number of evenly distributed and best-quality matching pairs from different clusters. Calculate P i The number of matching pairs that can be selected in each weighted cluster is then the P with the highest similarity is selected from the cluster based on similarity considerations. i Matching pairs, similarity The similarity between any two images in the target image set is calculated as the global description distance.
[0076] The embodiment performs efficient similar global retrieval of image pairs through global description matching, and performs effective image pair selection strategy based on erosion and dilation clustering of images, which has the characteristics of high efficiency, stability and robustness.
[0077] Embodiment 2
[0078] See also Figure 2 , Figure 2 Schematic diagram of a global description-based adaptive image matching pair selection device disclosed in an embodiment of the present invention. Figure 2 As shown, the adaptive image matching pair selection device based on global description may include: an image acquisition module 201, an image calculation module 202, an image processing module 203, a category division module 204 and an image matching module 205; wherein, the image acquisition module 201 is used to acquire a target image set, and perform NETVLAD global description extraction on each image in the target image set to obtain a global description definition of each image; the image calculation module 202 is used to perform global description distance calculation on all image matching pairs in the target image set according to the global description definition, and construct a distance matrix diagram according to the global description distance calculation; the image processing module 203 is used to binarize the distance matrix diagram; the category division module 204 is used to perform corrosion, expansion and clustering processing on the binarized distance matrix diagram to obtain different image matching pair clusters; the image matching module 205 is used to adaptively select matching pairs from different image matching pair clusters.
[0079] In the image calculation module 202, a global description distance calculation is performed on all image matching pairs in the target image set according to the global description definition, including: defining the NETVLAD description of each image as a vector D i , the similarity between any two images in the target image set is defined as S ij ; Through the formula The similarity between any two images in the target image set is calculated as the global description distance, where D i The transpose of .
[0080] Further, constructing a distance matrix diagram according to the global description distance calculation includes: obtaining the total number of all images in the target image set as N, and constructing an N*N distance matrix diagram using the similarity between any two images as matrix elements: Among them, 1 to N represent the image sequence number, S NN Represents the image at row N and column N.
[0081] In the image processing module 203, the distance matrix image is binarized, including: setting a binarization threshold, and binarizing the distance matrix image based on the binarization threshold, so that the distance matrix image is processed into a black and white image. Further, the binarization threshold is 1.2.
[0082] In the image matching module 205, matching pairs are adaptively selected from different image matching pair clusters, including: defining any image as frame F i , from the image frame F i The image range selected from the clustered images of the same image matching pair is [F i-m ,F i-1 ],[F i+1 ,F i+m ] and the image frame F i Select matching pairs. Further, it also includes: setting the maximum number of matching pairs for each image to N max , the category of the image matching pair cluster is Q, and the weight of the category Q of the image matching pair cluster is the number of images in the category of the image matching pair cluster w i , then the total weight of the image matching cluster is The number of image pairs is selected from each image category according to the formula: Select the class P with the highest similarity from each image matching pair cluster i image pairs. ,F i-m represents the im-th frame image, that is, the m-th frame image on the left side of the i-th frame image; F i+m represents the i+mth frame image, that is, the mth frame image on the right side of the ith frame image; F i-1 represents the i-1th frame image, that is, the first frame image to the left of the i-th frame image; F i+1 Represents the i+1th frame image, that is, the first frame image to the right of the i-th frame image
[0083] Embodiment 3
[0084] See also Figure 3 , Figure 3 Schematic diagram of the structure of an electronic device disclosed in an embodiment of the present invention. The electronic device may be a computer, a server, etc. Of course, in certain circumstances, it may also be a smart device such as a mobile phone, a tablet computer, a monitoring terminal, and an image acquisition device with processing functions. Figure 3 As shown, the electronic device may include:
[0085] A memory 301 storing executable program codes;
[0086] a processor 302 coupled to the memory 301;
[0087] The processor 302 calls the executable program code stored in the memory 301 to execute part or all of the steps in the adaptive image matching pair selection method based on global description in the first embodiment.
[0088] An embodiment of the present invention discloses a computer-readable storage medium storing a computer program, wherein the computer program enables a computer to execute part or all of the steps in the adaptive image matching pair selection method based on global description in the first embodiment.
[0089] The embodiment of the present invention further discloses a computer program product, wherein when the computer program product is run on a computer, the computer is enabled to execute part or all of the steps in the method for selecting adaptive image matching pairs based on global description in the first embodiment.
[0090] An embodiment of the present invention further discloses an application publishing platform, wherein the application publishing platform is used to publish a computer program product, wherein when the computer program product runs on a computer, the computer executes part or all of the steps in the adaptive image matching pair selection method based on global description in embodiment one.
[0091] In various embodiments of the present invention, it should be understood that the size of the serial numbers of the processes does not necessarily mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0092] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, i.e., they may be located in one place or distributed over multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0093] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0094] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-accessible memory. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product, which is stored in a memory and includes several requests for a computer device (which can be a personal computer, a server or a network device, etc., specifically a processor in a computer device) to perform some or all of the steps of the method described in each embodiment of the present invention.
[0095] In the embodiments provided by the present invention, it should be understood that "B corresponding to A" means that B is associated with A, and B can be determined according to A. However, it should also be understood that determining B according to A does not mean determining B only according to A, and B can also be determined according to A and / or other information.
[0096] A person of ordinary skill in the art can understand that some or all of the steps in the various methods of the embodiments can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium, and the storage medium includes a read-only memory (ROM), a random access memory (RAM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electronically erasable rewritable read-only memory (EEPROM), a compact disc (CD-ROM) or other optical disc storage, magnetic disk storage, magnetic tape storage, or any other computer-readable medium that can be used to carry or store data.
[0097] The above is a detailed introduction to the adaptive image matching pair selection method, device, electronic device and storage medium based on global description disclosed in the embodiments of the present invention. Specific examples are used in this article to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea; at the same time, for general technical personnel in this field, according to the ideas of the present invention, there will be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be understood as a limitation on the present invention.
Claims
1. An adaptive image matching pair selection method based on global description, characterized in that: include: Obtain a target image set, and perform NETVLAD global description extraction on each image in the target image set to obtain a global description definition of each image; Performing global description distance calculation on all image matching pairs in the target image set according to the global description definition, and constructing a distance matrix diagram according to the global description distance calculation; Binarizing the distance matrix graph; The binary distance matrix is corroded, expanded and clustered to obtain different image matching pair clusters; Adaptively pick matching pairs from different clusters of image matching pairs.
2. The adaptive image matching pair selection method according to claim 1, characterized in that: The performing global description distance calculation on all image matching pairs in the target image set according to the global description definition includes: Define the NETVLAD description of each image as a vector D i , the similarity between any two images in the target image set is defined as S ij ; By formula The similarity between any two images in the target image set is calculated as the global description distance, where D i The transpose of .
3. The adaptive image matching pair selection method according to claim 2, characterized in that: The step of calculating and constructing a distance matrix diagram according to the global description distance includes: The total number of all images in the target image set is obtained as N, and the similarity between any two images is used as the matrix element to construct an N*N distance matrix diagram: Among them, 1 to N represent the image sequence number, S NN Represents the image at row N and column N.
4. The adaptive image matching pair selection method according to claim 1, characterized in that: Binarizing the distance matrix graph includes: A binarization threshold is set, and the distance matrix image is binarized based on the binarization threshold, so that the distance matrix image is processed into a black and white image.
5. The adaptive image matching pair selection method according to claim 4, characterized in that: The binarization threshold is 1.
2.
6. The adaptive image matching pair selection method according to claim 1, characterized in that: Adaptively select matching pairs from different image matching pair clusters, including: Define any image as frame F i , select the image range as [F i-m ,F i-1 ],[F i+1 ,F i+m ] and the image frame F i Select matching pairs, F i-m represents the im-th frame image, that is, the m-th frame image on the left side of the i-th frame image; F i+m represents the i+mth frame image, that is, the mth frame image on the right side of the ith frame image; F i-1 represents the i-1th frame image, that is, the first frame image to the left of the i-th frame image; F i+1 It represents the i+1th frame image, that is, the first frame image to the right of the i-th frame image.
7. The adaptive image matching pair selection method according to claim 6, characterized in that: Also includes: Set the maximum number of matching pairs for each image to N max , the category of the image matching pair cluster is Q, and the weight of the category Q of the image matching pair cluster is the number of images in the image matching pair cluster w i , then the total weight of the image matching cluster is The number of image pairs is selected from each image category according to the formula: Select the class P with the highest similarity from each image matching pair cluster i image matching pairs.
8. An adaptive image matching pair selection device based on global description, characterized in that: include: Image acquisition module: used to acquire the target image set, and perform NETVLAD global description extraction on each image in the target image set to obtain the global description definition of each image; Image calculation module: used for performing global description distance calculation on all image matching pairs in the target image set according to the global description definition, and constructing a distance matrix diagram according to the global description distance calculation; Image processing module: used for binarizing the distance matrix graph; Category division module: used to perform corrosion, expansion and clustering processing on the binarized distance matrix graph to obtain different image matching pair clusters; Image matching module: used to adaptively select matching pairs from different image matching pair clusters.
9. An electronic device, characterized in that: include: A memory storing executable program code; a processor coupled to the memory; The processor calls the executable program code stored in the memory to execute the global description-based adaptive image matching pair selection method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program enables a computer to execute the global description-based adaptive image matching pair selection method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Image data processing method and device, electronic equipment and readable storage medium
CN113515659A
Three-dimensional reconstruction method, system and device in large scene
CN115205489A
Cervical liquid-based cell segmentation method and system based on watershed
CN115511815A
Global three-dimensional shape and temperature field reconstruction method for high-temperature object
CN118463848A
Visualization framework based on document representation learning
US20180196873A1