Method and apparatus for matching images

By acquiring and updating the depth value and three primary color information of the image and combining this information for image matching, the problems of low image matching accuracy and poor robustness are solved, and higher matching accuracy and robustness are achieved.

CN118968100BActive Publication Date: 2025-06-03ZHEJIANG DAHUA TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411440094.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-15
Publication Date
2025-06-03
Estimated Expiration
2044-10-15

AI Technical Summary

Technical Problem

During the image matching process, due to the differences in sources, time, location and viewing angles of different imaging devices, the image color texture, clarity and other differences are large, resulting in low matching accuracy and poor robustness.

Method used

By acquiring the set of target depth values ​​and the set of three primary color information of the image, the image is matched using this information. The specific steps include updating the initial depth value and three primary color information, determining the number of matching pixels, and judging image matching based on the preset threshold.

Benefits of technology

By integrating the texture information and structure information of the image, the robustness and accuracy of image matching are improved, and the problem of low matching accuracy of multi-source image data is solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118968100B_ABST
    Figure CN118968100B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention provides a method and apparatus for matching images, including: obtaining a first set of target depth values of a first image and a second set of target depth values of a second image, where the first set of target depth values includes depth values obtained by updating initial depth values of the first image through edge values of the first image, and the second set of target depth values includes depth values obtained by updating initial depth values of the second image through edge values of the second image; obtaining a first set of trichromatic information of the first image and a second set of trichromatic information of the second image; and matching the first image and the second image through the first set of target depth values and the second set of target depth values, as well as the first set of trichromatic information and the second set of trichromatic information. Through the present invention, the problem of low matching accuracy of multi-source image data in the related art is solved, and thus the effect of improving the matching accuracy of multi-source image data is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to the field of image processing, and more particularly, to a method and apparatus for matching images. Background Art

[0002] With the development of artificial intelligence technology, image matching technology has been widely applied in various fields. For example, the security field, the medical field, the retail field, etc. Image matching technology refers to using computer vision and machine learning technologies to identify and compare the content in images, thereby achieving image matching and similarity comparison. The development of image matching technology not only improves the efficiency and accuracy of image processing, but also brings more application possibilities to various industries.

[0003] However, during the process of image acquisition, factors such as different imaging device sources, different imaging times, different imaging positions and perspectives will result in significant differences in the color texture, clarity, visual confusion, content, etc. of the acquired images, thus making the image data matching solution have low accuracy and poor robustness.

[0004] In view of the above problems, there is currently no effective solution. Summary of the Invention

[0005] Embodiments of the present invention provide a method and apparatus for matching images to at least solve the problem of low matching accuracy of multi-source image data in related technologies.

[0006] According to an embodiment of the present invention, a method for matching images is provided, including: obtaining a first set of target depth values of a first image and a second set of target depth values of a second image, where the first set of target depth values includes: depth values obtained by updating the initial depth value of the first image through the edge value of the first image, and the second set of target depth values includes: depth values obtained by updating the initial depth value of the second image through the edge value of the second image; obtaining a first set of primary color information of the first image and a second set of primary color information of the second image, where the first set of primary color information includes: primary color information of each pixel point of the first image, and the second set of primary color information includes: primary color information of each pixel point of the second image; and matching the first image and the second image through the first set of target depth values and the second set of target depth values, and the first set of primary color information and the second set of primary color information.

[0007] In an exemplary embodiment, matching the first image and the second image by means of the first set of target depth values, the second set of target depth values, the first set of trichromatic information, and the second set of trichromatic information includes: determining the number of mutually matching pixel points in the first image and the second image by means of the first set of target depth values, the second set of target depth values, the first set of trichromatic information, and the second set of trichromatic information; determining that the first image and the second image match if the number of mutually matching pixel points in the first image and the second image reaches a first preset number threshold, and otherwise determining that the first image and the second image do not match.

[0008] In an exemplary embodiment, whether the first pixel point in the first image matches the second pixel point in the second image is determined by the following method: determining a first matching probability by means of the first target depth value of the first pixel point and the second target depth value of the second pixel point, where the pixel coordinates of the first pixel point in the first image correspond to the pixel coordinates of the second pixel point in the second image, the first set of target depth values includes the first target depth value, and the second set of target depth values includes the second target depth value; determining a second matching probability by means of the first trichromatic information of the first pixel point and the second trichromatic information of the second pixel point, where the first set of trichromatic information includes the first trichromatic information and the second set of trichromatic information includes the second trichromatic information; determining the weighted sum of the first matching probability and the second matching probability as the target matching probability of the first pixel point and the second pixel point; determining that the first pixel point matches the second pixel point if the target matching probability is greater than or equal to a preset probability threshold, and otherwise determining that the first pixel point and the second pixel point do not match.

[0009] In an exemplary embodiment, before obtaining the first set of target depth values of the first image and the second set of target depth values of the second image, the method further includes: determining a direction vector of the i-th pixel point according to the pixel coordinates of the i-th pixel point; determining the intersection coordinates of the direction vector and the target plane where the i-th pixel point is located as the target intersection coordinates; determining the distance value between the target intersection coordinates and the origin coordinates as the target depth value of the i-th pixel point, where the i-th pixel point is any pixel point in the first image or the i-th pixel point is any pixel point in the second image.

[0010] In an exemplary embodiment, before determining the intersection coordinates of the direction vector and the target plane where the i-th pixel point is located as the target intersection coordinates, the method further includes: determining a difference value between each pixel point and its adjacent pixel points through the initial depth values of the pixel points of the target image and the edge values of the pixel points of the target image, where the target image is the first image or the second image; performing region segmentation on the target image through the difference values between each pixel point and its adjacent pixel points to obtain a plurality of segmented regions; determining the segmented regions with the number of pixel points greater than or equal to a second preset number threshold among the plurality of segmented regions as target segmented regions; obtaining the plane of the target image through the target segmented regions, where the plane of the target image includes the target plane.

[0011] In an exemplary embodiment, determining a difference value between each pixel point and its adjacent pixel points through the initial depth values of the pixel points of the target image and the edge values of the pixel points of the target image includes: obtaining a first parameter value through the initial depth value of the j-th pixel point and the initial depth value of the k-th pixel point, where the target image includes the j-th pixel point and the k-th pixel point, and the j-th pixel point and the k-th pixel point are adjacent; obtaining a second parameter value through the pixel coordinates of the j-th pixel point and the pixel coordinates of the k-th pixel point; obtaining a third parameter value through the edge value of the j-th pixel point and the edge value of the k-th pixel point; determining the sum of the first parameter value, the second parameter value, and the third parameter value as the difference value between the j-th pixel point and the k-th pixel point.

[0012] In an exemplary embodiment, performing region segmentation on the target image through the difference values between each pixel point and its adjacent pixel points to obtain a plurality of segmented regions includes: in a case where the difference value between the pixel point and its adjacent pixel points is less than or equal to a preset difference threshold, determining that the pixel point and its adjacent pixel points belong to the same segmented region.

[0013] In an exemplary embodiment, obtaining the plane of the target image through the target segmented regions includes: performing plane fitting on the target segmented regions to obtain a fitted plane; performing smoothing optimization on the initial depth values of the pixel points in the fitted plane to obtain the plane.

[0014] According to another embodiment of the present invention, there is provided a device for matching images, including: a first acquisition module, configured to acquire a first set of target depth values of a first image and a second set of target depth values of a second image, wherein the first set of target depth values includes: depth values obtained by updating initial depth values of the first image through edge values of the first image, and the second set of target depth values includes: depth values obtained by updating initial depth values of the second image through edge values of the second image; a second acquisition module, configured to acquire a first set of primary color information of the first image and a second set of primary color information of the second image, wherein the first set of primary color information includes: primary color information of each pixel point of the first image, and the second set of primary color information includes: primary color information of each pixel point of the second image; a matching module, configured to match the first image and the second image through the first set of target depth values and the second set of target depth values, and the first set of primary color information and the second set of primary color information.

[0015] According to still another embodiment of the present invention, there is also provided a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the steps of the method described in any one of the above are implemented.

[0016] According to still another embodiment of the present invention, there is also provided an electronic device, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.

[0017] According to still another embodiment of the present invention, there is also provided a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the method described in any one of the above are implemented.

[0018] Through the present invention, the depth value information and primary color information of the first image and the second image are respectively acquired, and the first image and the second image are matched based on the depth value information and the primary color information. Since the structural information of the image can be acquired through the primary color information of the image, and the texture information of the image can be acquired through the depth information, the image is matched by integrating the texture information and the structural information, effectively solving problems such as weak texture and weak texture correlation of heterogeneous images, and improving the robustness and accuracy of image matching. Therefore, the problem of low matching accuracy of multi-source image data in the related art can be solved, and the effect of improving the matching accuracy of multi-source image data can be achieved. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1It is a block diagram of the hardware structure of a mobile terminal for a method of matching images according to an embodiment of the present invention;

[0020] Figure 2 It is a flowchart of a method of matching images according to an embodiment of the present invention;

[0021] Figure 3 It is a flowchart of multi-source image data matching according to an embodiment of the present invention;

[0022] Figure 4 It is a block diagram of the structure of a device for matching images according to an embodiment of the present invention. Detailed implementation manners

[0023] In the following, embodiments of the present invention will be described in detail with reference to the accompanying drawings and in conjunction with the embodiments.

[0024] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence.

[0025] The method embodiments provided in the embodiments of the present application can be executed on a mobile terminal, a computer terminal or a similar computing device. Taking running on a mobile terminal as an example, Figure 1 It is a block diagram of the hardware structure of a mobile terminal for a method of matching images according to an embodiment of the present invention. As Figure 1 shown, the mobile terminal may include one or more ( Figure 1 only one is shown in the figure) processors 102 (the processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data. Among them, the above-mentioned mobile terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that, Figure 1 the structure shown in the figure is only schematic and does not limit the structure of the above-mentioned mobile terminal. For example, the mobile terminal may further include more or fewer components than Figure 1 shown in the figure, or have a different configuration from Figure 1 shown in the figure.

[0026] The memory 104 can be used to store computer programs, such as software programs and modules of application software, such as the computer program corresponding to the method of matching images in the embodiments of the present invention. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, the above-mentioned method is implemented. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely disposed relative to the processor 102, and these remote memories may be connected to the mobile terminal through a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0027] The transmission device 106 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wireless network provided by a communication provider of the mobile terminal. In one instance, the transmission device 106 includes a network adapter (abbreviated as NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one instance, the transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0028] In this embodiment, a method running on the above-mentioned mobile terminal is provided. Figure 2 It is a flowchart of the method for matching images according to the embodiments of the present invention, as Figure 2 shown, and the process includes the following steps:

[0029] Step S202, obtaining a first set of target depth values of the first image and a second set of target depth values of the second image, where the first set of target depth values includes: the depth values obtained by updating the initial depth value of the first image through the edge value of the first image, and the second set of target depth values includes: the depth values obtained by updating the initial depth value of the second image through the edge value of the second image;

[0030] The above-mentioned first image may be an original image or an image to be matched, etc.; the above-mentioned first set of target depth values may be a set representing the depth information of the first image, and this depth information is the depth value or distance information of each pixel point in the image, and these depth values may represent the distance between an object and a camera, or the distance difference between objects.

[0031] The second image mentioned above can be a reference image or an image for continued matching with the image to be matched, etc.; the second set of target depth values mentioned above can be a set representing the depth information of the second image, and this depth information is the depth value or distance information of each pixel point in the image. These depth values can represent the distance of an object from the camera or the distance difference between objects.

[0032] The above-mentioned edge value can be the pixel value at the edge or boundary of a pixel in the image. This edge value can be any value that is distinguishable from the sub-edge region, such as 1, 2, 3, etc. Usually, the pixel value of the edge region can be set to 1, and the pixel value of other regions can be set to 0, so as to distinguish the edge of the image and help identify the edge features in the image. Commonly used edge detection algorithms include the Sobel operator, the Prewitt operator, the Canny edge detection algorithm, etc. Through these algorithms, the edge values in the image can be effectively extracted.

[0033] The preliminary depth value of the first image is updated through the edge value of the first image to obtain the first set of target depths. The preliminary depth value of the second image is updated through the edge value of the second image to obtain the second set of target depths, thereby realizing the depth estimation and depth optimization of the image, improving the accuracy and stability of the depth estimation, ensuring the precision of the depth information, and laying a solid foundation for subsequent feature extraction and matching.

[0034] Step S204: Obtain the first set of primary color information of the first image and the second set of primary color information of the second image. Among them, the first set of primary color information includes the primary color information of each pixel point of the first image, and the second set of primary color information includes the primary color information of each pixel point of the second image.

[0035] The above-mentioned set of primary color information can be a set used to represent the color information of an image, that is, a set of information on the three color channels of R (red), G (green), and B (blue) of the image. This set of primary color information can be represented in the form of a digital matrix, and a value on each color channel represents the pixel value of the corresponding pixel point on that channel. The sets of primary color information of the first image and the second image are obtained respectively.

[0036] Step S206: Match the first image and the second image through the first set of target depth values, the second set of target depth values, the first set of primary color information, and the second set of primary color information.

[0037] Match the first image and the second image based on the depth information and color information of the first image, as well as the depth information and color information of the second image, integrating the texture information and structural information of the images, realizing end-to-end feature extraction and matching of composite texture features and depth information features, effectively solving problems such as weak texture, weak texture correlation of heterogeneous images, large imaging perspective span, and small overlapping area, and improving the robustness and accuracy of the feature extraction and matching process.

[0038] Specifically, determine the number of mutually matching pixel points in the first image and the second image through the first set of target depth values, the second set of target depth values, the first set of primary color information, and the second set of primary color information; when the number of mutually matching pixel points in the first image and the second image reaches a first preset quantity threshold, determine that the first image and the second image match; otherwise, determine that the first image and the second image do not match.

[0039] The above-mentioned first preset quantity threshold can be a threshold for determining the number of pixel points for image matching, and this threshold can be set according to actual situations; the above-mentioned pixel points can be the smallest unit in a digital image, which is composed of a combination of red, green, and blue primary colors and is used to represent the color and brightness in the image. That is to say, an image is composed of thousands of pixel points, and each pixel point has its own color value and brightness value. When these pixel points are combined together, a complete image is formed. Generally, the higher the density of pixel points, the higher the clarity of the image; the above-mentioned mutually matching pixel points can be pairs of points that match each other between the first image and the second image, and such matching point pairs can be determined by mapping the pixel point local area of a small-resolution feature image to a large-resolution feature image and calculating the pixel point confidence distribution.

[0040] Through the number of matching point pairs and the depth information of the corresponding points, the point position mapping relationship can be iteratively established in the form of RANSAC (Random Sample Consensus), obtaining the matching point pairs between the images and filtering out the mis-matching points again. Finally, based on the number of mutually matching pixel points between the first image and the second image, when the number of mutually matching pixel points is greater than or equal to the first preset quantity threshold, determine that the first image and the second image match; otherwise, determine that the first image and the second image do not match. Thus, by setting the threshold as the matching credibility threshold for the matching point pairs, the higher the threshold, the higher the probability that the output matching point pairs are successfully matched, and the determination of whether the images match is measured by specifying the number of matches under the matching threshold.

[0041] Optionally, the execution entity of the above steps may be a background processor, or other devices with similar processing capabilities, or may also be a machine integrated with at least an image acquisition device and a data processing device. Among them, the image acquisition device may include a graphics acquisition module such as a camera, and the data processing device may include terminals such as a computer and a mobile phone, but is not limited thereto.

[0042] Through the above steps, the depth value information and the three primary color information of the first image and the second image are respectively obtained, and the first image and the second image are matched based on the depth value information and the three primary color information. Since the structural information of the image can be obtained through the three primary color information of the image, and the texture information of the image can be obtained through the depth information, the image is matched by integrating the texture information and the structural information, effectively solving problems such as weak texture and weak texture correlation of heterologous images. The problem of low matching accuracy of multi-source image data in the related art is solved, and the robustness and accuracy of image matching are improved.

[0043] The execution order of step S202 and step S204 can be interchanged, that is, step S204 can be executed first, and then S202 can be executed.

[0044] As an optional implementation manner, determining the number of mutually matching pixel points in the first image and the second image through the first target depth value set and the second target depth value set, and the first three primary color information set and the second three primary color information set includes: determining whether the first pixel point in the first image and the second pixel point in the second image match in the following manner: determining a first matching probability through the first target depth value of the first pixel point and the second target depth value of the second pixel point, where the pixel coordinates of the first pixel point in the first image correspond to the pixel coordinates of the second pixel point in the second image, the first target depth value set includes the first target depth value, and the second target depth value set includes the second target depth value; determining a second matching probability through the first three primary color information of the first pixel point and the second three primary color information of the second pixel point, where the first three primary color information set includes the first three primary color information, and the second three primary color information set includes the second three primary color information; determining the weighted sum of the first matching probability and the second matching probability as the target matching probability of the first pixel point and the second pixel point; and determining that the first pixel point and the second pixel point match when the target matching probability is greater than or equal to a preset probability threshold, otherwise determining that the first pixel point and the second pixel point do not match.

[0045] The above first matching probability may be the confidence probability distribution among pixel points within a local area in the depth information dimension, and the pixel points may be the pixel points corresponding to the same position in two different images; the above second matching probability may be the confidence probability distribution among pixel points within a local area in the RGB information dimension, and the pixel points may be the pixel points corresponding to the same position in two different images; the above target matching probability may be the confidence distribution probability among pixel points within a local area, and the weight for calculating the target matching probability may be set according to the actual situation; the above preset probability threshold may be the threshold for determining whether pixel points are a matching point pair. Wherein, when the target matching probability is greater than or equal to the preset probability threshold, it is determined that the first pixel point and the second pixel point match, that is, the first pixel point and the second pixel point are a matching point pair, otherwise it is determined that the first pixel point and the second pixel point do not match. The target matching probability can be determined by the following formula:

[0046]

[0047] Wherein, is the target matching probability, is the second matching probability, is the first matching probability, is the first pixel point, is the second pixel point, is the preset weight, 0 ≤ θ ≤ 1, and can be specifically adjusted according to the depth information situation.

[0048] As an optional implementation manner, before obtaining the first set of target depth values of the first image and the second set of target depth values of the second image, the method further includes: determining the direction vector of the i-th pixel point according to the pixel coordinates of the i-th pixel point; determining the intersection coordinate of the direction vector and the target plane where the i-th pixel point is located as the target intersection coordinate; determining the distance value between the target intersection coordinate and the origin coordinate as the target depth value of the i-th pixel point, where the i-th pixel point is any pixel point in the first image, or the i-th pixel point is any pixel point in the second image.

[0049] The above-mentioned $i$-th pixel can be any pixel in the image, and the image can be the first image or the second image; the pixel coordinates of the above-mentioned $i$-th pixel can be the position of the $i$-th pixel in the image, and the pixel coordinates can be determined according to the pixel origin in the image, where the pixel origin can be set according to the actual situation, such as the central pixel of the image, the bottom-left pixel of the image, or the top-left pixel of the image, etc. For example, if the resolution of an image is $50\times50$, then the image has 2500 pixels. Set the bottom-left pixel of the image as the pixel origin, and the origin coordinates are $(0, 0)$, then the pixel coordinates of the top-right pixel of the image are $(49, 49)$.

[0050] The above-mentioned direction vector can be a vector starting from the $i$-th pixel and pointing to another pixel. In a two-dimensional image, horizontal and vertical direction vectors are usually used to represent the direction of pixels. The horizontal direction vector is usually represented as $(1, 0)$, and the vertical direction vector is usually represented as $(0, 1)$. By combining horizontal and vertical direction vectors, pixels in any direction can be represented.

[0051] The target plane where the above-mentioned $i$-th pixel is located depends on the internal and external parameters of the camera and the coordinates of the $i$-th pixel. Usually, the target plane where the $i$-th pixel is located can be calculated through the camera parameters, pixel coordinates, and the position of the camera. Specifically, the ray equation corresponding to the $i$-th pixel can be calculated through the internal parameter matrix of the camera, the external parameter matrix, and the coordinates of the pixel. Then, through the position and direction of the camera, the intersection of the ray and the target plane is calculated to determine the target plane where the $i$-th pixel is located.

[0052] Determine the direction vector of the pixel according to the internal parameter matrix of the camera and the pixel coordinates of the $i$-th pixel, so as to determine the target intersection between the direction vector and the target plane where the $i$-th pixel is located; determine the coordinates of the target intersection through the pixel coordinates of the $i$-th pixel, and calculate the distance between the target intersection coordinates and the origin coordinates, which is the target depth value of the $i$-th pixel. Through plane correction for post-processing of the depth map, the depth of the image is ensured to be on a unified plane, improving the result accuracy of common depth estimation methods, and more accurately obtaining the depth information of the image, laying a solid foundation for subsequent depth feature extraction and matching.

[0053] As an optional implementation, before determining the intersection coordinates of the direction vector and the target plane where the i-th pixel is located as the target intersection coordinates, the method further includes: determining the difference value between each pixel and its adjacent pixel through the initial depth value of each pixel of the target image and the edge value of each pixel of the target image, where the target image is the first image or the second image; performing region segmentation on the target image through the difference value between each pixel and its adjacent pixel to obtain a plurality of segmentation regions; determining the segmentation regions with the number of pixels greater than or equal to a second preset number threshold in the plurality of segmentation regions as the target segmentation regions; obtaining the plane of the target image through the target segmentation regions, where the plane of the target image includes the target plane.

[0054] The above initial depth value may be the original depth value of the target image, and the target image may be the first image or the second image; the above edge value may be the gray value of the pixels at the edge or contour part of the target image, and the edge of the target image may be obtained through an edge operator, such as the sobel operator, the Prewitt operator, the Roberts operator, etc.; the above difference value may be the difference degree value between the i-th pixel and its adjacent pixel, and the difference value may be determined according to the unit normal vector of the pixel, the distance between the plane corresponding to the pixel and the camera origin, and the edge value of the target image where the pixel is located; the above segmentation region may be the region obtained after pixel merging, and the segmentation region may be iteratively merged according to the connectivity score; the above second preset number threshold may be the number of pixels for determining the target segmentation region, where the segmentation region with the number of pixels greater than or equal to the second preset number threshold in the segmentation region is the target segmentation region.

[0055] Performing edge extraction on the target image, calculating the difference value between each pixel and its adjacent pixel in the target image according to the extracted edge value and the initial depth value of the target image, and at the same time, based on obtaining the segmentation regions in the target image, determining the target segmentation regions, and calculating the plane of the target image according to the target segmentation regions, that is, calculating the target plane, so as to smooth and optimize the depth values of the relevant regions according to the fitted plane of the target image to ensure that the depth is on the same plane.

[0056] Optionally, obtaining the plane of the target image through the target segmentation region includes: performing plane fitting on the target segmentation region to obtain a fitted plane; performing smoothing optimization on the initial depth values of each pixel in the fitted plane to obtain the plane.

[0057] As an optional implementation, the difference value between each pixel point and its adjacent pixel points is determined by the initial depth value of each pixel point of the target image and the edge value of each pixel point of the target image, including: obtaining a first parameter value through the initial depth value of the j-th pixel point and the initial depth value of the k-th pixel point, where the target image includes the j-th pixel point and the k-th pixel point, and the j-th pixel point and the k-th pixel point are adjacent; obtaining a second parameter value through the pixel coordinates of the j-th pixel point and the pixel coordinates of the k-th pixel point; obtaining a third parameter value through the edge value of the j-th pixel point and the edge value of the k-th pixel point; and determining the sum of the first parameter value, the second parameter value, and the third parameter value as the difference value between the j-th pixel point and the k-th pixel point.

[0058] Specifically, the difference value can be determined by the following formula :

[0059]

[0060] where is the first parameter value, is the second parameter value, is the third parameter value.

[0061] Optionally, the above first parameter value can be determined according to the unit normal vector of the initial depth value of the pixel point:

[0062]

[0063]

[0064] where is the unit normal vector corresponding to the j-th pixel point in the depth map, that is, the unit normal vector of the initial depth value of the j-th pixel point, and this unit normal vector is perpendicular to the plane corresponding to this pixel point; is the unit normal vector corresponding to the k-th pixel point in the depth map, that is, the unit normal vector of the initial depth value of the k-th pixel point, and this unit normal vector is perpendicular to the plane corresponding to this pixel point; is the length operation symbol; is the distance between the unit normal vector corresponding to the j-th pixel point and the unit normal vector corresponding to the k-th pixel point; is the minimum value of the distances between the unit normal vector corresponding to the j-th pixel point and the unit normal vectors corresponding to all pixel points in its neighborhood; is the maximum value of the distances between the unit normal vector corresponding to the j-th pixel point and the unit normal vectors corresponding to all pixel points in its neighborhood.

[0065] Optionally, the above second parameter value can be determined according to the distance between the plane corresponding to the pixel point and the camera origin:

[0066]

[0067]

[0068] wherein, is the distance between the plane corresponding to the j-th pixel point in the depth map and the camera origin, that is, the distance from the point to the plane; is the distance between the plane corresponding to the k-th pixel point in the depth map and the camera origin, that is, the distance from the point to the plane; is the length operation symbol; is the difference between the distance between the plane corresponding to the j-th pixel point and the camera origin and the distance between the plane corresponding to the k-th pixel point and the camera origin; is the minimum difference between the distance between the plane corresponding to the j-th pixel point and the camera origin and the distances between the planes corresponding to all pixel points in its neighborhood and the camera origin; is the maximum difference between the distance between the plane corresponding to the j-th pixel point and the camera origin and the distances between the planes corresponding to all pixel points in its neighborhood and the camera origin.

[0069] Optionally, the above third parameter value can be determined according to the edge value of the target image:

[0070]

[0071] wherein, is the edge value in the edge image corresponding to the j-th pixel point. In this edge image, the edge value of the edge part is 0, and the edge value of the rest part is only non-zero; is the edge value in the edge image corresponding to the k-th pixel point. In this edge image, the edge value of the edge part is 0, and the edge value of the rest part is only non-zero; is a preset value, which can be adjusted according to the sharpness of the actual scene edge, such as 0.05, etc.

[0072] As an optional implementation manner, the target image is segmented into multiple segmentation regions by the difference values between each of the pixel points and its adjacent pixel points, including: when the difference value between the pixel point and its adjacent pixel points is less than or equal to a preset difference threshold, the pixel point and its adjacent pixel points are determined to belong to the same segmentation region.

[0073] Optionally, the segmentation region can also be determined by the connectivity score, that is, when the connectivity score between the pixel and its adjacent pixels is greater than or equal to the preset connectivity threshold, the pixel and its adjacent pixels are determined to belong to the same segmentation region. Specifically, the segmentation region can be determined by the following formula:

[0074]

[0075] where is the region merging between the pixel and its adjacent pixels, where true indicates that the pixel and its adjacent pixels merge regions, that is, the pixel and its adjacent pixels belong to the same segmentation region, and false indicates that the pixel and its adjacent pixels do not merge regions, that is, the pixel and its adjacent pixels do not belong to the same segmentation region; is the connectivity score between the pixel and its adjacent pixels; is the preset connectivity threshold, which can be set according to the actual situation.

[0076] As an optional implementation manner, Figure 3 is the flowchart of multi-source image data matching according to the embodiments of the present invention, as shown in Figure 3 shown, the specific process is as follows:

[0077] S1, respectively determine whether the original image (the first image) and the reference image (the second image) contain depth information, and obtain the RGBD data of the original image and the reference image. Among them, when depth information is included, directly jump to S3; otherwise, jump to S2;

[0078] S2, perform image depth estimation on the original image and / or the reference image:

[0079] S21, perform initial depth information estimation on the original image and / or the reference image through a common monocular depth estimation network, including networks trained in supervised, semi-supervised, self-supervised manners, etc.

[0080] S22, perform post-processing optimization on the initially estimated depth information:

[0081] S221, perform edge extraction on the original image and / or the reference image, and common edge operators such as the sobel operator can be used to obtain the edge image;

[0082] S222, calculate the per-pixel difference value based on the initial depth information estimation and the edge image;

[0083] S223, according to the connectivity description, when the connectivity score is higher than the threshold, perform region merging to obtain multiple segmentation regions;

[0084] S224. The above-mentioned region merging process is iteratively performed until there are no merging edges that meet the conditions;

[0085] S225. Determine the regions with the number of pixels greater than the threshold C (for example, C is taken as 600) in the multiple segmented regions as the target segmented regions, and perform plane fitting on the target segmented regions to obtain the plane. The specific fitting method can adopt a common scheme: define the plane equation and solve the least squares to obtain the plane parameters;

[0086] S226. Smooth and optimize the depth values of the relevant regions according to the fitted plane to ensure that the depths are on the same plane. Specifically: calculate the direction vector according to the camera internal parameter matrix and the pixel coordinates of the i-th pixel point, solve the intersection coordinates of the direction vector and the plane obtained above, and the distance from the intersection point to the origin is the depth value corresponding to the pixel coordinates;

[0087] S3. Extract composite features from the original image and the reference image, including texture features and depth structure features of the image, etc.:

[0088] S31. Use a standard CNN of the ResNet class to extract a coarse feature map with a resolution of 1 / 8 (small resolution) and a fine feature map with a resolution of 1 / 2 (large resolution) of the RGBD image, and use DINVOv2 to extract a fine depth feature map with a resolution of 1 / 2;

[0089] S32. First perform position encoding on the above-mentioned coarse feature map, and then send it to the transformer module to extract matching features. Among them, the transformer module includes two types of attention layers: self-attention layer and cross-attention layer;

[0090] S4. Perform composite feature matching on the original image and the reference image, fuse texture and depth structure feature information, and establish a sub-pixel level matching mapping relationship between the images:

[0091] S41. Calculate the score matrix between the above-mentioned features, then multiply after applying softmax in two dimensions to obtain the matching probability of the nearest neighbor, and finally filter out some outlier matching pairs through the mutual nearest neighbor algorithm;

[0092] S42. On the basis of the above-mentioned coarse matching, the coarse matching point pairs are first mapped to the corresponding fine feature map and depth feature map, and then local feature matching is performed on the fine feature map and the depth feature map respectively to calculate the target matching probability between the first pixel point in the first image and the second pixel point in the second image.

[0093] It should be noted that the training process of the above entire network is a supervised training mode, and the RGBD matching image pair data and the pixel-by-pixel matching relationship are obtained through simulation data or labeled data. Finally, the loss function includes a coarse-grained matching loss and a fine matching loss. , specifically define the reuse of the original definition of LoFTR (LayoutLM Fine-Tuning for Text Recognition, a model for text recognition).

[0094] S5, output the matching result:

[0095] Based on the above network output, a semi-dense matching point pair relationship of the image (i.e., the matching relationship between the first pixel point and the second pixel point) can be obtained. Then, through the number of matching point pairs and the depth information of the corresponding points, a point position mapping relationship can be established by the RANSAC method, and mis-matched points can be filtered again. Finally, when the threshold matching point pairs are set to be greater than the number N, it can be determined that the previous matching of the image is successful, otherwise the matching fails.

[0096] In the above process, through depth estimation and depth optimization processing of the RGB image, texture features and depth information features are obtained respectively, and the image matching result is output end-to-end. Thus, it can adapt to any RGB image and can also adapt to RGBD images, omitting the depth information acquisition step, with high matching result accuracy and good adaptability.

[0097] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc), and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of the present invention.

[0098] In this embodiment, a device for matching images is also provided. This device is used to implement the above embodiments and preferred implementation manners, and those that have been described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0099] Figure 4 is a structural block diagram of the device for matching images according to an embodiment of the present invention, as Figure 4As shown, the device includes: a first acquisition module 402, configured to acquire a first set of target depth values of a first image and a second set of target depth values of a second image, where the first set of target depth values includes: depth values obtained by updating initial depth values of the first image through edge values of the first image, and the second set of target depth values includes: depth values obtained by updating initial depth values of the second image through edge values of the second image; a second acquisition module 404, configured to acquire a first set of primary color information of the first image and a second set of primary color information of the second image, where the first set of primary color information includes: primary color information of each pixel point of the first image, and the second set of primary color information includes: primary color information of each pixel point of the second image; a matching module 406, configured to match the first image and the second image through the first set of target depth values and the second set of target depth values, and the first set of primary color information and the second set of primary color information.

[0100] In an exemplary embodiment, the device is further configured to determine the number of mutually matching pixel points in the first image and the second image through the first set of target depth values and the second set of target depth values, and the first set of primary color information and the second set of primary color information; in a case where the number of mutually matching pixel points in the first image and the second image reaches a first preset number threshold, determine that the first image and the second image match, otherwise determine that the first image and the second image do not match.

[0101] In an exemplary embodiment, the device is further configured to determine whether a first pixel point in the first image matches a second pixel point in the second image in the following manner: determine a first matching probability through the first target depth value of the first pixel point and the second target depth value of the second pixel point, where the pixel coordinates of the first pixel point in the first image correspond to the pixel coordinates of the second pixel point in the second image, the first set of target depth values includes the first target depth value, and the second set of target depth values includes the second target depth value; determine a second matching probability through the first primary color information of the first pixel point and the second primary color information of the second pixel point, where the first set of primary color information includes the first primary color information, and the second set of primary color information includes the second primary color information; determine the weighted sum of the first matching probability and the second matching probability as the target matching probability of the first pixel point and the second pixel point; in a case where the target matching probability is greater than or equal to a preset probability threshold, determine that the first pixel point matches the second pixel point, otherwise determine that the first pixel point and the second pixel point do not match.

[0102] In an exemplary embodiment, the device is further configured to determine a direction vector of the i-th pixel point according to the pixel coordinates of the i-th pixel point; determine the intersection coordinates of the direction vector and the target plane where the i-th pixel point is located as the target intersection coordinates; determine the distance value between the target intersection coordinates and the origin coordinates as the target depth value of the i-th pixel point, where the i-th pixel point is any pixel point in the first image, or the i-th pixel point is any pixel point in the second image.

[0103] In an exemplary embodiment, the device is further configured to determine a difference value between each pixel point and its adjacent pixel points through the initial depth value of each pixel point of the target image and the edge value of each pixel point of the target image, where the target image is the first image or the second image; perform region segmentation on the target image through the difference value between each pixel point and its adjacent pixel points to obtain a plurality of segmentation regions; determine a segmentation region with the number of pixel points greater than or equal to a second preset number threshold in the plurality of segmentation regions as the target segmentation region; obtain the plane of the target image through the target segmentation region, where the plane of the target image includes the target plane.

[0104] In an exemplary embodiment, the device is further configured to obtain a first parameter value through the initial depth value of the j-th pixel point and the initial depth value of the k-th pixel point, where the target image includes the j-th pixel point and the k-th pixel point, and the j-th pixel point and the k-th pixel point are adjacent; obtain a second parameter value through the pixel coordinates of the j-th pixel point and the pixel coordinates of the k-th pixel point; obtain a third parameter value through the edge value of the j-th pixel point and the edge value of the k-th pixel point; determine the sum of the first parameter value, the second parameter value, and the third parameter value as the difference value between the j-th pixel point and the k-th pixel point.

[0105] In an exemplary embodiment, the device is further configured to determine that the pixel point and its adjacent pixel point belong to the same segmentation region when the difference value between the pixel point and its adjacent pixel point is less than or equal to a preset difference threshold.

[0106] In an exemplary embodiment, the device is further configured to perform plane fitting on the target segmentation region to obtain a fitting plane; perform smoothing optimization on the initial depth value of each pixel point in the fitting plane to obtain the plane.

[0107] It should be noted that the above-mentioned various modules can be implemented by software or hardware. For the latter, it can be implemented in the following ways, but not limited to this: the above-mentioned modules are all located in the same processor; or, the above-mentioned various modules are respectively located in different processors in any combination form.

[0108] An embodiment of the present invention also provides a computer-readable storage medium, in which a computer program is stored. Wherein, when the computer program is executed by a processor, the steps of the method described in any one of the above are implemented.

[0109] In an exemplary embodiment, the above-mentioned computer-readable storage medium may include but is not limited to: USB flash drives, read-only memories (ROM for short), random access memories (RAM for short), mobile hard disks, magnetic disks or optical disks, etc., various media that can store computer programs.

[0110] An embodiment of the present invention also provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.

[0111] In an exemplary embodiment, the above-mentioned electronic device may further include a transmission device and an input / output device. Wherein, the transmission device is connected to the above-mentioned processor, and the input / output device is connected to the above-mentioned processor.

[0112] An embodiment of the present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the method described in each embodiment of the present application are implemented.

[0113] The specific examples in this embodiment may refer to the examples described in the above embodiments and exemplary embodiments, and will not be repeated here.

[0114] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the present invention can be implemented by a general-purpose computing device. They can be concentrated on a single computing device, or distributed on a network composed of multiple computing devices. They can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order than here, or they can be separately made into individual integrated circuit modules, or multiple modules or steps among them can be made into a single integrated circuit module to implement. In this way, the present invention is not limited to any specific combination of hardware and software.

[0115] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for matching images, characterized in that: include: Acquire a first target depth value set of a first image and a second target depth value set of a second image, wherein the first target depth value set includes: depth values ​​obtained by updating initial depth values ​​of the first image by edge values ​​of the first image, and the second target depth value set includes: depth values ​​obtained by updating initial depth values ​​of the second image by edge values ​​of the second image; Acquire a first three-primary color information set of the first image and a second three-primary color information set of the second image, wherein the first three-primary color information set includes: three-primary color information of each pixel point of the first image, and the second three-primary color information set includes: three-primary color information of each pixel point of the second image; Matching the first image and the second image by using the first target depth value set and the second target depth value set, and the first three-primary color information set and the second three-primary color information set; Before acquiring the first target depth value set of the first image and the second target depth value set of the second image, the method further includes: Determine the direction vector of the ith pixel point according to the pixel coordinates of the ith pixel point; Determine the intersection coordinates of the direction vector and the target plane where the i-th pixel point is located as the target intersection coordinates; Determine the distance value between the target intersection coordinates and the origin coordinates as the target depth value of the i-th pixel point, wherein the i-th pixel point is any pixel point in the first image, or the i-th pixel point is any pixel point in the second image; Before determining the intersection coordinates of the direction vector and the target plane where the i-th pixel point is located as the target intersection coordinates, the method further includes: Determining a difference value between each pixel and its adjacent pixel by using an initial depth value of each pixel of a target image and an edge value of each pixel of the target image, wherein the target image is the first image or the second image; Performing region segmentation on the target image according to the difference value between each pixel and its adjacent pixel to obtain a plurality of segmented regions; Determine a segmented area in which the number of pixel points is greater than or equal to a second preset number threshold value as a target segmented area; The plane of the target image is obtained through the target segmentation area, wherein the plane of the target image includes the target plane.

2. The method according to claim 1, characterized in that Matching the first image and the second image by using the first target depth value set and the second target depth value set, and the first three-primary color information set and the second three-primary color information set, includes: Determine the number of mutually matching pixels in the first image and the second image by using the first target depth value set and the second target depth value set, and the first three-primary color information set and the second three-primary color information set; When the number of mutually matching pixels in the first image and the second image reaches a first preset number threshold, it is determined that the first image and the second image match; otherwise, it is determined that the first image and the second image do not match.

3. The method according to claim 2, characterized in that Determining the number of mutually matching pixels in the first image and the second image by using the first target depth value set and the second target depth value set, and the first three-primary color information set and the second three-primary color information set includes: Determine whether a first pixel in the first image matches a second pixel in the second image by: Determining a first matching probability by a first target depth value of the first pixel and a second target depth value of the second pixel, wherein a pixel coordinate of the first pixel in the first image corresponds to a pixel coordinate of the second pixel in the second image, the first target depth value set includes the first target depth value, and the second target depth value set includes the second target depth value; Determine a second matching probability by using first three primary color information of the first pixel and second three primary color information of the second pixel, wherein the first three primary color information set includes the first three primary color information, and the second three primary color information set includes the second three primary color information; Determine a weighted sum of the first matching probability and the second matching probability as a target matching probability of the first pixel point and the second pixel point; When the target matching probability is greater than or equal to a preset probability threshold, it is determined that the first pixel point matches the second pixel point; otherwise, it is determined that the first pixel point does not match the second pixel point.

4. The method according to claim 1, characterized in that Determining the difference value between each pixel and its adjacent pixel by using the initial depth value of each pixel of the target image and the edge value of each pixel of the target image, including: Obtaining a first parameter value through an initial depth value of a j-th pixel and an initial depth value of a k-th pixel, wherein the target image includes: the j-th pixel and the k-th pixel, and the j-th pixel and the k-th pixel are adjacent; Obtaining a second parameter value through the pixel coordinates of the j-th pixel point and the pixel coordinates of the k-th pixel point; Obtain a third parameter value through the edge value of the j-th pixel point and the edge value of the k-th pixel point; The sum of the first parameter value, the second parameter value and the third parameter value is determined as a difference value between the j-th pixel point and the k-th pixel point.

5. The method according to claim 4, characterized in that The target image is segmented by the difference value between each pixel and its adjacent pixel to obtain a plurality of segmented regions, including: When the difference value between the pixel point and its adjacent pixel point is less than or equal to a preset difference threshold, the pixel point and its adjacent pixel point are determined to belong to the same segmented area.

6. The method according to claim 1, characterized in that Obtaining the plane of the target image through the target segmentation area includes: Performing plane fitting on the target segmentation area to obtain a fitting plane; The initial depth value of each pixel point in the fitting plane is smoothly optimized to obtain the plane.

7. A device for matching images, characterized in that: include: A first acquisition module is used to acquire a first target depth value set of a first image and a second target depth value set of a second image, wherein the first target depth value set includes: depth values ​​obtained by updating initial depth values ​​of the first image by edge values ​​of the first image, and the second target depth value set includes: depth values ​​obtained by updating initial depth values ​​of the second image by edge values ​​of the second image; A second acquisition module is used to acquire a first three-primary color information set of the first image and a second three-primary color information set of the second image, wherein the first three-primary color information set includes: three-primary color information of each pixel point of the first image, and the second three-primary color information set includes: three-primary color information of each pixel point of the second image; a matching module, configured to match the first image and the second image using the first target depth value set and the second target depth value set, and the first three-primary color information set and the second three-primary color information set; The device is further used to determine the direction vector of the i-th pixel point according to the pixel coordinates of the i-th pixel point before acquiring the first target depth value set of the first image and the second target depth value set of the second image; determine the intersection coordinates of the direction vector and the target plane where the i-th pixel point is located as the target intersection coordinates; and determine the distance value between the target intersection coordinates and the origin coordinates as the target depth value of the i-th pixel point, wherein the i-th pixel point is any pixel point in the first image, or the i-th pixel point is any pixel point in the second image; The device is also used to determine the difference value between each pixel point and its adjacent pixel point through the initial depth value of each pixel point of the target image and the edge value of each pixel point of the target image before determining the intersection coordinates of the direction vector and the target plane where the i-th pixel point is located as the target intersection coordinates, wherein the target image is the first image or the second image; perform region segmentation on the target image through the difference value between each pixel point and its adjacent pixel point to obtain multiple segmentation regions; determine the segmentation region in which the number of pixels in the multiple segmentation regions is greater than or equal to a second preset number threshold as the target segmentation region; and obtain the plane of the target image through the target segmentation region, wherein the plane of the target image includes the target plane.

8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program implements the steps of the method described in any one of claims 1 to 6 when executed by a processor.

9. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to run the computer program to perform the method according to any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method described in any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Target detection method based on template matching algorithm

    CN117197501A