Image Similarity Matching Method, Device, and Storage Medium
By separating the image into color channels, the histogram similarity and structural similarity index are calculated, and the feature point matching algorithm is used to solve the problem of low text similarity matching efficiency, achieving efficient and accurate text recognition and matching.
Patent Information
- Application Number
- CN202210441209.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-25
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2042-04-25
AI Technical Summary
In the prior art, text similarity matching efficiency is low and error-prone. Especially when processing text in different countries, it is difficult to efficiently identify punctuation marks and special symbols, and the calculation amount is large, making it difficult to apply to Android embedded devices.
By separating the image to be tested and the standard image into color channel images, the histogram similarity is calculated, the structural similarity index of the character string area image is extracted, and the image similarity is judged using the feature point matching algorithm, including Harris corner point detection, SIFT, SURF, FAST, BRIEF and ORB algorithms.
It improves the efficiency and accuracy of text similarity matching, can effectively identify text and symbols in different countries, reduces the amount of calculation, and makes it suitable for Android embedded devices.
Smart Images

Figure CN114972817B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly to an image similarity matching method, device, and storage medium. Background Art
[0002] With the continuous growth of the economy and the rapid development of the Internet, the development of terminal devices based on application programs emerges in an endless stream, and these intelligent application programs bring convenience and fun to people's lives. Before the mass production and release of application program products, a large number of software tests are often required to ensure the reliability and stability of the use of software products.
[0003] In terms of the development of language text translation software, during the software testing stage, when selecting texts in different countries, they need to be compared with standard template texts. If there are differences, the texts in that country need to be re-translated and displayed. Currently, testers mainly verify whether the texts are the same by comparing them one by one line by line. The work is cumbersome, and long-term work will cause eye fatigue, with low work efficiency and easy to make mistakes. Summary of the Invention
[0004] The main purpose of the present invention is to provide an image similarity matching method, device, and storage medium, aiming to solve the problem of low efficiency in text similarity matching.
[0005] To achieve the above object, the present invention provides an image similarity matching method, which includes:
[0006] Obtain a to-be-tested image and a standard image, separate the to-be-tested image and the standard image into color channel images, and determine the histogram similarity of the color channel images;
[0007] If the histogram similarity is higher than a first preset threshold, extract a string region image from the color channel image, and determine a structural similarity index of the string region image;
[0008] If the structural similarity index is higher than a second preset threshold, extract feature point information from the string region image;
[0009] Perform feature point matching according to the feature point information to determine whether the to-be-tested image and the standard image are the same.
[0010] Optionally, the steps of obtaining a to-be-tested image and a standard image, separating the to-be-tested image and the standard image into color channel images, and calculating the histogram similarity of the color channel images include:
[0011] Use a preset text control to obtain a standard image, locate the text box position of the to-be-tested text on the test page, and intercept the to-be-tested image according to the text box position;
[0012] Separate the image to be measured and the standard image into three-channel images;
[0013] Calculate the frequency score of each pixel point in the three-channel images using a preset frequency formula;
[0014] In each channel, take the average value of all pixel points as the channel score of a single channel, and take the average value of all the channel scores as the histogram similarity.
[0015] Optionally, before the step of separating the image to be measured and the standard image into three-channel images, it further includes:
[0016] Extract the height values and width values of the image to be measured and the standard image, and compare the height value difference and width value difference between the image to be measured and the standard image;
[0017] If both the height value difference and the width value difference are within a preset range, then execute the step of separating the image to be measured and the standard image into three-channel images;
[0018] If either the height value difference or the width value difference is not within the preset range, then execute the step of extracting the string region image in the color channel image and calculating the structural similarity index of the string region image.
[0019] Optionally, the step of extracting the string region image in the color channel image and calculating the structural similarity index of the string region image includes:
[0020] Perform grayscale conversion and binarization processing on the color channel image in sequence;
[0021] Locate the region covering all characters in the binarized color channel image, and intercept the string region image in the region;
[0022] Perform normalization processing on the string region image to make the sizes of the string region images consistent;
[0023] Process the normalized string region image using a preset structural similarity algorithm to obtain the structural similarity index of the string region image.
[0024] Optionally, the step of locating the region covering all characters in the binarized color channel image and intercepting the string region image in the region includes:
[0025] Perform preset row cutting and column cutting on the binarized color channel image to obtain the coordinate positions of each character;
[0026] Obtain the start character coordinate position and the end character coordinate position in the said coordinate positions, and intercept the string area image according to the start character coordinate position and the end character coordinate position.
[0027] Optionally, the step of performing feature point matching according to the said feature point information to determine whether the image to be measured and the standard image are the same includes:
[0028] Obtain the initial matching points of the feature points according to the descriptors in the said feature point information;
[0029] If the number of the initial matching points is less than the preset number, then use the initial matching points as the final matching points;
[0030] Calculate the difference degree and the coincidence degree of the feature points according to the number of the final matching points;
[0031] If the difference degree is less than the preset difference degree threshold and the coincidence degree is greater than the preset coincidence degree threshold, then determine that the image to be measured and the standard image are the same.
[0032] Optionally, the step of obtaining the initial matching points of the feature points according to the descriptors in the said feature point information includes:
[0033] Establish an index tree according to the multi-dimensional data of the feature points;
[0034] Compare the node feature vectors in the index tree with the descriptors, traverse the nodes of the index tree, and obtain the nearest neighbor points and the second nearest neighbor points of the feature points;
[0035] Calculate the distance ratio of the nearest neighbor points and the second nearest neighbor points. If the distance ratio is less than the preset distance threshold, then use the nearest neighbor points and the second nearest neighbor points as the initial matching points.
[0036] Optionally, before the step of calculating the difference degree and the coincidence degree of the feature points according to the number of the final matching points, it further includes:
[0037] If the number of the initial matching points is greater than or equal to the preset number, then use the preset random sample consensus algorithm to remove the mis-matched feature points in the initial matching points;
[0038] Use the initial matching points after the removal process as the final matching points.
[0039] In addition, to achieve the above object, the present invention further provides an electronic device, and the electronic device includes: a memory, a processor, and an image similarity matching program stored on the memory and operable on the processor. The image similarity matching program is configured to implement the steps of the image similarity matching method as described above.
[0040] In addition, to achieve the above object, the present invention further provides a computer-readable storage medium, on which an image similarity matching program is stored. When the image similarity matching program is executed by a processor, the steps of the image similarity matching method described above are implemented.
[0041] The present invention obtains a to-be-tested image and a standard image, determines the histogram similarity between the to-be-tested image and the standard image, and overall evaluates the similarity between the two from the aspect of the image background of the to-be-tested image and the standard image. If the histogram similarity is higher than a first preset threshold, a string region image is extracted, and the structural similarity index between the two string region images is determined. The similarity between the two is further overall evaluated from the aspect of the image structure. If the structural similarity index is higher than a second preset threshold, the feature point information in the string region image is extracted, and the local similarity is evaluated through the matching degree of feature point matching, so as to obtain the final similarity matching result. By means of the matching method from the whole to the part, the efficiency of text similarity matching is improved. Description of the Drawings
[0042] Figure 1 It is a schematic structural diagram of an electronic device for the hardware operating environment involved in the embodiment solution of the present invention;
[0043] Figure 2 It is a schematic flowchart of the first embodiment of the image similarity matching method of the present invention;
[0044] Figure 3 It is a schematic flowchart of the second embodiment of the image similarity matching method of the present invention;
[0045] Figure 4 It is a schematic flowchart of the third embodiment of the image similarity matching method of the present invention;
[0046] Figure 5 It is a schematic flowchart of the fourth embodiment of the image similarity matching method of the present invention.
[0047] The realization, functional features and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. Detailed Embodiment
[0048] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0049] Current text similarity matching technologies are all based on existing data. If text content is detected and recognized and then the differences between the two are compared, the following problems will be faced: First, text outside the dataset cannot be recognized; second, punctuation marks or special symbols cannot be recognized, and these special characters have a great impact on the accuracy of text detection; finally, due to the large storage and computing volume of the existing dataset model, it is difficult to be applied to Android embedded devices.
[0050] The main technical solution of the present invention is: obtain a to-be-tested image and a standard image, separate the to-be-tested image and the standard image into color channel images, and determine the histogram similarity of the color channel images; if the histogram similarity is higher than a first preset threshold, extract the string region image in the color channel image, and determine the structural similarity index of the string region image; if the structural similarity index is higher than a second preset threshold, extract the feature point information in the string region image; perform feature point matching according to the feature point information to determine whether the to-be-tested image and the standard image are the same.
[0051] Refer to Figure 1 , Figure 1 It is a schematic structural diagram of an electronic device for the hardware operating environment involved in the embodiment solution of the present invention.
[0052] As Figure 1 shown, the electronic device may include: a processor 1001, such as a central processing unit (CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. Among them, the communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display screen (Display) and an input unit such as a keyboard (Keyboard). Optionally, the user interface 1003 may further include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a wireless fidelity (WI-FI) interface). The memory 1005 may be a high-speed random access memory (Random Access Memory, RAM) memory, or a stable non-volatile memory (Non-Volatile Memory, NVM), such as a disk memory. Optionally, the memory 1005 may also be a storage device independent of the foregoing processor 1001.
[0053] Those skilled in the art can understand that Figure 1 the structure shown in
[0054] As Figure 1 shown, the memory 1005 as a storage medium may include an operating system, a network communication module, a user interface module, and an image similarity matching program.
[0055] In Figure 1 the electronic device shown, the network interface 1004 is mainly used for data communication with other devices; the user interface 1003 is mainly used for data interaction with users; the processor 1001 and the memory 1005 in the electronic device of the present invention may be arranged in the electronic device, and the electronic device calls the image similarity matching program stored in the memory 1005 through the processor 1001 and executes the image similarity matching method provided by the embodiments of the present invention.
[0056] The embodiments of the present invention provide an image similarity matching method. Referring to Figure 2 , Figure 2 it is a schematic flowchart of the first embodiment of an image similarity matching method of the present invention.
[0057] In this embodiment, the image similarity matching method includes:
[0058] Step S10, obtaining a to-be-tested image and a standard image, separating the to-be-tested image and the standard image into color channel images, and determining the histogram similarity of the color channel images;
[0059] When text is displayed on a terminal device, the text can be regarded as a kind of "image", that is, the text and the area where the text is located are taken as a whole, having measurement indexes such as hue, background, and structure. When performing text similarity matching, a text box can be selected according to the display position of the to-be-tested text to obtain the to-be-tested image, and the standard image of the known standard text can be taken out.
[0060] The color channel refers to the channel that stores the color information of the image. Each image has one or more channels, and the default number of color channels in the image depends on its color mode, that is, the color mode of each image will determine the number of its color channels. A CYMK (Cyan Magenta Yellow blacK) image has 4 channels by default, and an RGB (Red Green Blue) image and a Lab image have 3 channels by default.
[0061] After separating the to-be-tested image and the standard image into color channel images, calculate the histogram similarity of each single channel. Drawing a histogram can count the frequency of all pixels in a digital image according to the size of the gray value. For the frequency difference of the pixel points in the to-be-tested image and the standard image, the larger the frequency difference, the greater the difference between the two images.
[0062] Step S20, if the histogram similarity is higher than the first preset threshold, extract the string region image in the color channel image and determine the structural similarity index of the string region image.
[0063] According to the similarity requirements in the actual scenario, if the similarity requirements are high, the first preset threshold can be set higher to exclude the test images that do not match the standard image as a whole and improve the matching efficiency. In a possible implementation, the first preset threshold is set to 0.9. If the histogram similarity is higher than or equal to 0.9, extract the string region image from the color channel image. If the histogram similarity is lower than 0.9, it can be directly determined that the test image is different from the standard image.
[0064] When extracting the string region image, the method of text cutting can be used to locate all character regions, perform horizontal and vertical projections on the character regions to obtain the coordinate positions of each character, and obtain the string region image according to the coordinate positions. Characters include letters and symbols. The string region image can remove the intervals between characters.
[0065] The structural similarity index can be characterized by SSIM (Structural Similarity). The structural similarity index defines the structural information from the perspective of the image composition as independent of brightness and contrast and reflects the attributes of the object structure in the scene. When the SSIM value is 1, it means that the two images are exactly the same.
[0066] Step S30, if the structural similarity index is higher than the second preset threshold, extract the feature point information in the string region image.
[0067] The structural similarity index can evaluate the similarity between two images from the dimension of the overall image structure. In a scenario with high similarity requirements, the second preset threshold can also be set higher to exclude the test images that do not match the standard image from the overall structure. In a possible implementation, the second preset threshold is set to 0.5. If the structural similarity index is greater than or equal to 0.5, the feature point information in the string region image can be extracted. If the structural similarity index is less than 0.5, it is directly determined that the test image is different from the standard image.
[0068] The feature point information in the string region can be extracted through an image feature point extraction algorithm. Examples of image feature point extraction algorithms include Harris corner detection, SIFT (Scale Invariant Feature Transform), SURF (Speeded Up Robust Features), FAST corner detection, BRIEF (Binary Robust Independent Elementary Features) descriptor, ORB (Oriented FAST and Rotated BRIEF), etc.
[0069] Step S40: Perform feature point matching based on the feature point information to determine whether the image to be tested is the same as the standard image.
[0070] When performing feature point matching, initial matching can be carried out through the descriptors in the feature point information to obtain initial matching points. If the number of initial matching points is large, elimination processing can be performed on the initial matching points to eliminate the mis-matched points among the initial matching points and obtain the final matching points. Evaluate the matching degrees of the final matching points in terms of difference degree and coincidence degree. If both the difference degree and the coincidence degree meet the requirements, it can be determined that the image to be tested is the same as the standard image.
[0071] In this embodiment, the image to be tested and the standard image are obtained, and the image to be tested and the standard image are separated into color channel images. Determine the histogram similarity of the color channel images. If the histogram similarity is higher than the first preset threshold, extract the string region image in the color channel image and determine the structural similarity index of the string region image. If the structural similarity index is higher than the second preset threshold, extract the feature point information in the string region image. Perform feature point matching based on the feature point information to determine whether the image to be tested is the same as the standard image.
[0072] Further, in the second embodiment of the image similarity matching method of the present invention, refer to Figure 3 , the method includes:
[0073] Step S11: Use a preset text control to obtain the standard image, locate the position of the text box of the text to be tested on the test page, and intercept the image to be tested according to the position of the text box.
[0074] When designing a text control, the text control is enabled with the functions of positioning the specific position of each paragraph of text in the image, intercepting the text, and finding the standard text in the corresponding national language, and obtaining the image to be tested and the standard image. During the actual testing process, the tester pauses the page of the display device, and the text control can intercept the text box area in the test page, and the image in the text box area is the image to be tested.
[0075] Step S12: Extract the height values and width values of the image to be tested and the standard image, and compare the height value difference and width value difference between the image to be tested and the standard image.
[0076] The height value and width value of the image can be represented by pixel values, that is, the pixel value in the height direction is the height value, and the pixel value in the width direction is the width value. Divide the height value of the image to be tested by the height value of the standard image to obtain the height value difference, and divide the width value of the image to be tested by the width value of the standard image to obtain the width value difference.
[0077] Step S13: If both the height value difference and the width value difference are within the preset interval, separate the image to be tested and the standard image into three-channel images.
[0078] In an implementable embodiment, the preset interval is set to 0.8 - 1.2. If both the height value difference and the width value difference are within the range of 0.8 - 1.2, it means that the sizes of the image to be tested and the standard image are basically the same. The sizes of the two images are basically the same, which can be regarded as the number of pixel points and the density distribution in the images are basically the same, facilitating subsequent image processing steps.
[0079] After separating the image to be tested and the standard image into three-channel images, count the pixel point frequencies in each single-channel image.
[0080] Step S14: Calculate the frequency score of each pixel point in the three-channel image using a preset frequency formula.
[0081] The preset frequency formula can be:
[0082] Socre = (1 - |hist1 - hist2|) / Max(hist1, hist2)
[0083] Among them, Score represents the score, hist1 represents the frequency of a single pixel point appearing in the image to be tested, hist2 represents the frequency of a single pixel point appearing in the standard image, and Max represents taking the maximum value. It can be seen from the frequency formula that the greater the difference in the frequencies of a single pixel point appearing, the larger the denominator, and the smaller the frequency score, that is, the greater the frequency difference, the lower its score.
[0084] Step S15: In each channel, take the average value of all pixel points as the channel score of a single channel, and take the average value of all the channel scores as the histogram similarity.
[0085] In a single channel, the average value of all pixel points is the single-channel score, and the average value of the three-channel scores is the frequency score. The histogram similarity is represented by the frequency score. The higher the frequency score, the smaller the background difference between the image to be tested and the standard image.
[0086] In this embodiment, a text control is used to obtain the image to be tested and the standard image, and the difference in the overall background between the two is judged by the histogram similarity between the image to be tested and the standard image, which can effectively identify the background difference of the text and the deformation of the text.
[0087] Further, in the third embodiment of the image similarity matching method of the present invention, referring to Figure 4 , the method includes:
[0088] Step S21: If the height value difference and the width value difference are not within the preset interval, perform graying and binarization processing on the color channel image in sequence;
[0089] The height value difference or the width value difference not being within the preset interval indicates that there is a large size difference between the image to be tested and the standard image, and there are large differences in the number and distribution density of pixel points. The size of the image to be tested and the standard image can be unified in the subsequent process, and then the structural similarity between the two is judged.
[0090] First perform graying and then binarization processing on the color channel image, converting the color channel image into a grayscale image and then into a black-and-white image. When actually comparing the image to be tested and the standard image, there are often large gray differences between the two, such as differences in brightness and contrast, etc., and the standard template of the text is easily affected by noise. The shape of the image generated by adaptive binarization remains basically unchanged at the corresponding positions, which can effectively eliminate the interference caused by the gray difference, and at the same time reduce the interference for subsequent cutting characters to extract the horizontal and vertical coordinate pixel points.
[0091] Step S22: Perform preset row cutting and column cutting on the binarized color channel image to obtain the coordinate positions of each character;
[0092] When cutting, a completely black background image can be defined, and the number of white pixel points in each row is cyclically counted to obtain the horizontal projection. The vertical segmentation position is obtained according to the horizontal projection, and then the rows are segmented. The column segmentation is obtained in the same way as the row segmentation, and the coordinate positions of each character are obtained through row segmentation and column segmentation.
[0093] Step S23: Obtain the start character coordinate position and the end character coordinate position in the said coordinate position, and intercept the string area image according to the start character coordinate position and the end character coordinate position.
[0094] Obtain the coordinate positions of the start character and the end character, extract the upper left coordinate point of the start character and the lower right coordinate point of the end character, draw a rectangle through these two coordinate points, and intercept the string area to obtain the string area image. The string area image only contains text and symbols.
[0095] Step S24: Perform normalization processing on the string area image so that the sizes of the string area images are consistent.
[0096] Enlarge the intercepted string area image proportionally. In the actual application scenario, since the image size intercepted by the text control is small, too small a size will result in a small number of subsequent extracted feature points or even no feature points can be extracted. Therefore, it is necessary to enlarge the image to increase the writing width of the text and increase the number of detected feature points. Then normalize the sizes of the two images, enlarge the smaller-sized image to the same size as the larger-sized image, so that the string area image sizes of the image to be measured and the standard image are consistent.
[0097] Step S25: Use a preset structural similarity algorithm to process the normalized string area image to obtain the structural similarity index of the string area image.
[0098] The structural similarity index can be the SSIM index, and it is calculated using the following structural similarity formula:
[0099]
[0100] where the image to be measured is x, the standard image is y, μ x is the average value of x, μ y is the average value of y, is the variance of x, is the variance of y, σ xy is the covariance of x and y, c1 and c2 are constants used to maintain stability, c1 = (k1L) 2 , c2 = (k2L) 2 , L is the dynamic range of pixel values, k1 = 0.01, k2 = 0.03.
[0101] In this embodiment, the structural similarity index is used to judge the overall structural difference between the image to be measured and the standard image. The texts used for the same semantic content in different countries will be completely different. Through the structural difference, it can be effectively identified whether the texts in the image to be measured and the standard image are in the same language.
[0102] Further, in the fourth embodiment of the image similarity matching method of the present invention, referring to Figure 5 , the method includes:
[0103] Step S31, establishing an index tree according to the multi-dimensional data of the feature points;
[0104] The algorithm used for extracting feature points can be the SIFT algorithm. During the extraction process, first, extreme point detection is performed in the scale space to find feature points, then the stability of the candidate feature points is detected, and the feature points that pass the detection are SIFT feature points. Then, sampling is performed within the neighborhood window to determine the main direction of the feature points, and finally, a SIFT feature point descriptor, that is, a 128-dimensional feature vector, is generated.
[0105] The index tree can be a KD (K-Dimensional) tree. The KD tree is a query index structure widely used in database indexing. First, the variance of the multi-dimensional data of each feature point in the feature point set is calculated, and the median with the largest variance is selected to divide the feature point set into two subsets. The above operation is repeated until all set partitions are completed, and a KD tree is generated.
[0106] Step S32, comparing the node feature vectors in the index tree with the descriptor, traversing the nodes of the index tree, and obtaining the nearest neighbor point and the second nearest neighbor point of the feature points;
[0107] Query the KD tree from the root node, compare the feature vector of the query node with the descriptor generated on the KD tree, and continuously traverse to find the nearest neighbor point and the second nearest neighbor point.
[0108] Step S33, calculating the ratio of the nearest neighbor point and the second nearest neighbor point. If the ratio is less than a preset distance threshold, then use the nearest neighbor point and the second nearest neighbor point as the initial matching points;
[0109] Obtain the nearest neighbor point and the second nearest neighbor point of the feature points, calculate the distance ratio of the nearest neighbor point and the second nearest neighbor point. When the ratio is less than the preset distance threshold, it can be considered that the nearest neighbor point and the second nearest neighbor point match. In a possible implementation manner, the preset distance threshold can be 0.75. If the ratio is greater than or equal to the preset distance threshold, then discard the nearest neighbor point and the second nearest neighbor point. Store all the matching feature points in a list, and the length of the list is the number of initial matching points.
[0110] Step S34, if the number of the initial matching points is less than a preset number, then use the initial matching points as the final matching points;
[0111] After the initial matching, a relatively large number of initial matching points may be obtained. By setting a preset number of the initial matching points, the mis-matched feature points among the initial matching points can be further removed. The preset number can be 8. If the number of the initial matching points is less than 8, the initial matching points can be used as the final matching points.
[0112] Step S35, if the number of the initial matching points is greater than or equal to the preset number, use the preset Random Sample Consensus (RANSAC) algorithm to remove the mis-matched feature points among the initial matching points;
[0113] The Random Sample Consensus (RANSAC) algorithm can better remove the mis-matched feature points. Its ideal result is that the lines connecting the matching points are parallel, and those mis-matched feature points with inclined connecting lines of the matching points are removed. During the removal process, randomly initialize eight inliers from the initial matching points, perform normalization processing on the matching feature points, perform singular value decomposition on the linear equations corresponding to the eight pairs of initial matching points to obtain the fundamental matrix, calculate the distance from other points to the epipolar line of the fundamental matrix. If the distance is less than the preset value, it is an inlier; otherwise, it is an outlier. Repeat the iteration, and take the feature points at the time when the number of inliers is the largest as the final result.
[0114] Step S36, use the initial matching points after the removal process as the final matching points;
[0115] After the removal process, the mis-matched feature points among the initial matching points can be removed, and the matching points at the time when the number of inliers is the largest are used as the final matching points.
[0116] Step S37, calculate the difference degree and the coincidence degree of the feature points according to the number of the final matching points;
[0117] The formula for calculating the difference degree can be: D = |P1 - P2| / Max(P1, P2)
[0118] Wherein, P1 represents the number of the initial matching points in the image to be measured, P2 represents the number of the initial matching points in the standard image, Max represents taking the maximum value, and D represents the difference degree.
[0119] The formula for calculating the coincidence degree can be: C = R / P
[0120] Wherein, R represents the number of the final matching points, P represents the number of the initial matching points, and C represents the coincidence degree.
[0121] Step S38, if the difference degree is less than the preset difference degree threshold and the coincidence degree is greater than the preset coincidence degree threshold, determine that the image to be measured and the standard image are the same.
[0122] The difference degree represents the overall difference degree of the feature points extracted from two images. When the difference between the two is too large, it indicates a large difference in image features, and it can be determined that the two images are different. The coincidence degree represents the amount of similarity of the extracted feature points. The average value of the coincidence degrees of the two images is taken as the final result. If it is determined that the images are the same, the more feature points that match, the better. The preset difference degree threshold can be 0.1, and the preset coincidence degree threshold can be 0.8.
[0123] If the difference degree is greater than the preset difference degree threshold or the coincidence degree is less than the preset coincidence degree threshold, it is determined that the image to be measured is different from the standard image.
[0124] In this embodiment, the feature point information in the string region image is extracted, the feature points are initially matched and re-matched, and the difference degree and the coincidence degree are used to comprehensively evaluate the feature point matching degree between the image to be measured and the standard image, improving the accuracy of the similarity matching of text images.
[0125] The embodiment of the present invention also provides an electronic device, which includes: a memory, a processor, and an image similarity matching program stored on the memory and executable on the processor. The image similarity matching program is configured to implement the steps of the image similarity matching method as described above.
[0126] The embodiment of the present invention also provides a computer-readable storage medium, which is characterized in that an image similarity matching program is stored on the computer-readable storage medium. When the image similarity matching program is executed by a processor, the steps of the image similarity matching method as described above are implemented.
[0127] It should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or system. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of another identical element in the process, method, article or system including the element.
[0128] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.
[0129] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described example methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium as described above (such as ROM / RAM, magnetic disk, optical disc), and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of the present invention.
[0130] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the description of the present invention and the content of the drawings, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.
Claims
1. An image similarity matching method, characterized in that, The described image similarity matching method includes the following steps: Use a preset text control to obtain a standard image, locate the text box position of the text to be tested on the test page, and intercept the image to be tested according to the text box position. The text control obtains the standard text in the corresponding national language; Separate the image to be tested and the standard image into three-channel images, and the three-channel images are color channel images; Use a preset frequency formula to calculate the frequency score of each pixel point in the three-channel image; In each channel, take the average value of all pixel points as the channel score of a single channel, and take the average value of all the channel scores as the histogram similarity; If the histogram similarity is higher than the first preset threshold, extract the string region image in the color channel image, and determine the structural similarity index of the string region image, where the intervals between characters in the string region image are removed; If the structural similarity index is higher than the second preset threshold, extract the feature point information in the string region image; Obtain the initial matching points of the feature points according to the descriptors in the feature point information; If the number of the initial matching points is less than the preset number, take the initial matching points as the final matching points; Calculate the difference degree and coincidence degree of the feature points according to the number of the final matching points; If the difference degree is less than the preset difference degree threshold and the coincidence degree is greater than the preset coincidence degree threshold, determine that the image to be tested and the standard image are the same.
2. The image similarity matching method according to claim 1, characterized in that Before the step of separating the image to be tested and the standard image into three-channel images, it further includes: Extract the height values and width values of the image to be tested and the standard image, and compare the height value difference and width value difference between the image to be tested and the standard image; If both the height value difference and the width value difference are within the preset range, execute the step of separating the image to be tested and the standard image into three-channel images; If the height value difference or the width value difference is not within the preset range, execute the step of extracting the string region image in the color channel image and calculating the structural similarity index of the string region image.
3. The image similarity matching method according to claim 1, wherein The step of extracting the string region image in the color channel image and calculating the structural similarity index of the string region image includes: Perform grayscale processing and binarization processing on the color channel image in sequence; Locate the region covering all characters in the binarized color channel image, and intercept the string region image in the region; Perform normalization processing on the string region image to make the sizes of the string region images consistent; Use a preset structural similarity algorithm to process the normalized string region image to obtain the structural similarity index of the string region image.
4. The image similarity matching method according to claim 3, wherein The step of locating the region covering all characters in the binarized color channel image and intercepting the string region image in the region includes: Perform preset row cutting and column cutting on the binarized color channel image to obtain the coordinate positions of each character; Obtain the start character coordinate position and the end character coordinate position in the coordinate positions, and intercept the string region image according to the start character coordinate position and the end character coordinate position.
5. The image similarity matching method according to claim 1, characterized in that The step of obtaining the initial matching points of the feature points according to the descriptors in the feature point information includes: Establish an index tree based on the multi-dimensional data of the feature points; Compare the node feature vectors in the index tree with the descriptors, traverse the nodes of the index tree, and obtain the nearest neighbor point and the second nearest neighbor point of the feature points; Calculate the distance ratio between the nearest neighbor point and the second nearest neighbor point. If the distance ratio is less than a preset distance threshold, use the nearest neighbor point and the second nearest neighbor point as the initial matching points.
6. The image similarity matching method according to claim 1, characterized in that Before the step of calculating the difference degree and the coincidence degree of the feature points according to the number of the final matching points, it further includes: If the number of the initial matching points is greater than or equal to a preset number, use a preset random sample consensus algorithm to remove the mis-matched feature points in the initial matching points; Use the initial matching points after the removal process as the final matching points.
7. An electronic device, characterized in that, The electronic device includes: a memory, a processor, and an image similarity matching program stored on the memory and executable on the processor. The image similarity matching program is configured to implement the steps of the image similarity matching method according to any one of claims 1 to 6.
8. A computer-readable storage medium, characterized in that, An image similarity matching program is stored on the computer-readable storage medium. When the image similarity matching program is executed by a processor, it implements the steps of the image similarity matching method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Image multifeature extraction and fusion method and system
CN102663391A
Remote sensing image registration method of multi-source sensor
CN103020945A
Similar character determination method and device, computer equipment and storage medium
CN110097002A