Character string fuzzy matching method and system
By extracting the image feature vectors of characters and constructing the glyph similarity matrix, combined with the weighted editing distance method, the problem of character errors in similar results in OCR technology is solved, and more accurate character similarity matching is achieved.
Patent Information
- Application Number
- CN202510132426.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-06
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-02-06
AI Technical Summary
When existing OCR technology recognizes characters, similar results have character errors and cannot effectively distinguish similar characters.
By obtaining multiple characters of different font size fonts, extracting the image feature vector of each character, calculating the angle cosine distance between the image feature vectors of two characters as similarity, constructing a glyph similarity matrix, and using the weighted editing distance method, improving the accuracy of character similarity matching based on the cost of the character similarity weighted replacement operation.
It effectively reduces character errors in OCR recognition results, improves the accuracy of character similarity matching, and ensures the correctness of recognition results.
Smart Images

Figure CN120014654A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of character recognition, and particularly to a method and system for fuzzy matching of strings. Background Art
[0002] Optical Character Recognition (OCR) technology can convert different types of documents, such as scanned paper documents, PDF files or images, into editable and searchable data.
[0003] In the prior art, OCR performs fuzzy comparison and search between the recognition result and the standard string by means of image preprocessing, text detection, character segmentation, character recognition and post-processing to find characters similar to the correct characters. In the judgment of similarity, for example, the edit distance method is adopted, which is an algorithm for measuring the similarity between two strings.
[0004] However, although the characters obtained by fuzzy comparison and search are similar to the characters to be recognized, such as "紫" and "柴", they are not the same character, and there is no further distinction between the two similar characters, resulting in incorrect recognition results. Summary of the Invention
[0005] Embodiments of the present invention provide a method and system for fuzzy matching of strings, which can solve the problem of character errors in the similar results recognized by OCR technology in the prior art.
[0006] The embodiment of the present invention provides a string fuzzy matching method, comprising the following steps: obtaining a plurality of characters of different font sizes; selecting any two characters from the plurality of characters, extracting an image feature vector of each character by using an optical character recognition technology (OCR), and obtaining an angle cosine distance between the image feature vectors of the two characters as a similarity; constructing a glyph similarity matrix according to the similarity between any two characters in the plurality of characters; performing character similarity weighting on the cost of describing the cost of replacing the character in the replacement operation according to a replacement operation for replacing one character of a string with another character in an edit distance method, so as to construct a weighted edit distance method; the character similarity weighting includes: when the similarity of the replaced character is higher than When the similarity threshold is set, the cost of the replacement operation is the difference between the first set value and the similarity; otherwise, the cost of the replacement operation is the second set value, and the second set value indicates that a penalty value higher than the difference between the first set value and the similarity is assigned to the cost of the replacement operation; the weighted edit distance method is used for the two strings to be matched, and during the replacement operation, the similarity between the two characters that need to be replaced in the two strings to be matched is queried according to the glyph similarity matrix, and the cost of the replacement operation is weighted by the character similarity to obtain the result of the weighted edit distance; the result of the weighted edit distance is normalized to obtain the similarity of the two strings to be matched; the matching result is determined according to the relationship between the similarity of the two strings to be matched and the preset value.
[0007] Furthermore, the weighted edit distance method is used for the two strings to be matched, and the specific steps include: Create a ( m +1)×( n +1); in, m is the length of the first string to be matched, n is the length of the second string to be matched, and the row index of the matrix i The number of characters from left to right corresponding to the first string, the column label of the matrix j Corresponding to the number of characters in the second string from left to right, each matrix element in the matrix D [ i ][ j ] indicates the first character string i Converts the first character of the second string to j The minimum edit distance among the edit distances required for characters; Fill the matrix: in, D [ i -1][ j ]+1 is the delete operation, D [ i][ j -1]+1 is the insertion operation, D [ i -1][ j -1]+cos t For the replacement operation, cos t is the cost of the replacement operation; Query the character similarity in the similarity matrix between the character to be replaced in the first string and the character used as a replacement reference in the second string, and weight the cost by character similarity: in, if Indicates conditions, similarity Indicates the character similarity, threshold Indicates the set similarity threshold; Filled, the matrix D [ m ][ n ] represents the weighted edit distance between the first character string and the second character string.
[0008] Furthermore, the step of obtaining the angle cosine distance between the image feature vectors of two characters as the similarity comprises: When two characters are full-width characters, convert the full-width characters to half-width characters; The HOG descriptor of each character is obtained by using the Histogram of Oriented Gradients (HOG) algorithm in the optical character recognition technology (OCR), wherein the HOG descriptor is a one-dimensional array; the similarity comparison between two characters is converted into the similarity comparison between the HOG descriptors of the two characters; Get the cosine distance of the angle between the HOG descriptors of two characters as the similarity.
[0009] Furthermore, the query method for the character similarity between the characters to be replaced in the first character string and the characters used as a replacement reference in the second character string in the similarity matrix needs to be optimized to increase the query speed. The optimization specifically includes: if the first character string and the second character string are both company names, the character similarity of all characters in the company name is screened in the similarity matrix to obtain a similarity matrix subset.
[0010] Furthermore, the optimization also includes: when each character in each character string is queried in the similarity matrix, the number of similar characters queried is no more than 15, and the similarity threshold is set to be no less than 0.85.
[0011] Furthermore, the optimization further includes: simultaneously querying the character similarity of each character in the first character string and each character in the second character string in the similarity matrix on multiple threads.
[0012] Furthermore, the matching result is determined based on the relationship between the similarity of the two strings to be matched and a preset value, and the specific steps include: when the similarity is 1, it indicates a successful match; when the similarity is greater than a set similarity threshold and less than 1, it indicates a suspicious match result; when the similarity is less than the set similarity threshold, it indicates a failed match.
[0013] Furthermore, the suspicious matching result requires an in-depth comparison of the two character strings to be matched, and the specific steps of the in-depth comparison include: When the two character strings to be matched are company names, the company names of the two character strings to be matched are searched separately on a website that can query the unified social credit code, and the search results are obtained and compared; If the comparison results are the same, it means the match is successful, and the company names of the two strings to be matched are the same company; When the comparison results are different, it indicates that the matching fails, and it means that the corporate names of the two character strings to be matched have similar glyphs.
[0014] The embodiment of the present invention provides a string fuzzy matching system, including: A matrix construction module is used to obtain multiple characters of different font sizes; select any two characters from the multiple characters, extract the image feature vector of each character by optical character recognition technology OCR, and obtain the angle cosine distance between the image feature vectors of the two characters as the similarity; construct a glyph similarity matrix according to the similarity between any two characters in the multiple characters; A weighted edit distance construction and use module is used to perform character similarity weighting on the cost of describing the cost of replacing characters in the replacement operation according to the replacement operation used to replace one character of a string with another character in the edit distance method, so as to construct a weighted edit distance method; the character similarity weighting includes: when the similarity of the replaced character is higher than the set similarity threshold, the cost of the replacement operation is the difference between the first set value and the similarity; otherwise, the cost of the replacement operation is the second set value, and the second set value indicates that a penalty value higher than the difference between the first set value and the similarity is assigned to the cost of the replacement operation; the weighted edit distance method is used for the two strings to be matched, and during the replacement operation, the similarity between the two characters to be replaced in the two strings to be matched is queried according to the glyph similarity matrix, and the cost of the replacement operation is weighted by character similarity to obtain the result of the weighted edit distance; the result of the weighted edit distance is normalized to obtain the similarity of the two strings to be matched; The similarity matching module is used to determine the matching result according to the relationship between the similarity of two character strings to be matched and a preset value.
[0015] An embodiment of the present invention provides a method and system for fuzzy matching of strings. Compared with the prior art, the beneficial effects are as follows: Select any two characters from multiple characters, extract the image feature vectors of each character through the optical character recognition technology OCR, and obtain the cosine distance of the included angle between the image feature vectors of the two characters as the similarity; construct a glyph similarity matrix according to the similarity between any two characters in the multiple characters; according to the replacement operation in the edit distance method for replacing one character of a string with another character, perform character similarity weighting on the cost for describing the replacement character in the replacement operation to construct a weighted edit distance method; the character similarity weighting includes: when the similarity of the character to be replaced is higher than a set similarity threshold, the cost of the replacement operation is the difference between the first set value and the similarity; otherwise, the cost of the replacement operation is the second set value, and the second set value represents a penalty value given to the cost of the replacement operation that is higher than the difference between the first set value and the similarity; use the weighted edit distance method for the two strings to be matched, and when performing the replacement operation, query the similarity between the two characters that need to be replaced in the two strings to be matched according to the glyph similarity matrix, and perform character similarity weighting on the cost of the replacement operation to obtain the result of the weighted edit distance; normalize the result of the weighted edit distance to obtain the similarity of the two strings to be matched; determine the matching result according to the relationship between the similarity of the two strings to be matched and the preset value.
[0016] Among them, the cost of the replacement operation takes into account the similarity, so that the replacement operation is affected by the similarity. When the similarity of the character to be replaced is higher than the set similarity threshold, the cost of the replacement operation is expressed as the difference between the first set value and the similarity, and the cost is small; otherwise, the cost of the replacement operation is expressed as the second set value, and the cost is high; then, use the weighted edit distance for the two strings to be matched, and the replacement operation will be affected by the similarity of the characters when performing the replacement operation; normalize the obtained weighted edit distance to obtain the similarity of the two strings to be matched. Finally, the improvement of the edit distance method is realized, so that when similar characters are encountered in the comparison process, the similarity of the two characters is considered to distinguish the similar two characters, and the correct recognition result is obtained. Description of the Drawings
[0017] Figure 1 It is a schematic diagram of the glyph similarity matrix of a method for fuzzy matching of strings provided by an embodiment of the present invention. Among them, (a) represents the glyph images of the four characters "zǐ chái dīng xiāng" (font = FangSong_GB2312, font size = 14), and (b) represents the glyph similarity matrix of the four characters "zǐ chái dīng xiāng" (font = FangSong_GB2312, font size = 14); Figure 2 It is a heat map of the glyph similarity matrix of common characters of a method for fuzzy matching of strings provided by an embodiment of the present invention; Figure 3 A method flow chart of a string fuzzy matching method provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0018] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific embodiments of the present invention are described in detail below in conjunction with the accompanying drawings. In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without violating the connotation of the present invention, so the present invention is not limited by the specific embodiments disclosed below.
[0019] The embodiment of the present invention provides a string fuzzy matching method, comprising the following steps: Step 1: Obtain multiple characters in different font sizes; select any two characters from the multiple characters, extract the image feature vector of each character through optical character recognition technology OCR, and obtain the cosine distance of the angle between the image feature vectors of the two characters as the similarity; construct a glyph similarity matrix based on the similarity between any two characters in the multiple characters.
[0020] Step 2: According to the replacement operation used to replace one character of a string with another character in the edit distance method, the cost of describing the cost of replacing the character in the replacement operation is weighted by character similarity to construct a weighted edit distance method. Character similarity weighting includes: when the similarity of the replaced character is higher than the set similarity threshold, the cost of the replacement operation is the difference between the first set value and the similarity; otherwise, the cost of the replacement operation is the second set value, and the second set value indicates that a penalty value higher than the difference between the first set value and the similarity is assigned to the cost of the replacement operation; the weighted edit distance method is used for the two strings to be matched. During the replacement operation, the similarity between the two characters to be replaced in the two strings to be matched is queried according to the glyph similarity matrix, and the cost of the replacement operation is weighted by character similarity to obtain the result of the weighted edit distance; the result of the weighted edit distance is normalized to obtain the similarity of the two strings to be matched.
[0021] Step 3: Determine the matching result based on the relationship between the similarity of the two strings to be matched and the preset value: when the similarity is 1, it indicates a successful match; when the similarity is greater than the set similarity threshold and less than 1, it indicates a suspicious match result; when the similarity is less than the set similarity threshold, it indicates a failed match.
[0022] The specific operation process is as follows: The present invention first implements a simple data cleaning function of a list of specialized, special and new enterprises in Excel by using a method of customizing VBA functions and formulas. Then, based on the HOG algorithm of the histogram of directional gradients commonly used in OCR recognition, the image features of each character are extracted, and the basic data of the 3606×3606 glyph similarity matrix between any two characters in common Chinese characters and symbols (2500 common characters, 1000 times common characters, and 106 common symbols) is formed and stored. Next, by customizing the weighted edit distance Levenshtein, the replacement cost between two similar glyphs is reduced (reducing the distance), and a penalty factor is applied to the replacement of non-similar glyph characters (increasing the distance), so as to realize the fuzzy matching of the enterprise name string in the OCR recognition scenario. In order to improve the matching efficiency, the multiprocessing library capable of process parallelization is used to realize the parallel computing function. For non-perfect corner matching samples (the edit distance is not 0, but small), a code-free visual crawler is designed through the process automation platform Microsoft Power Automate, which searches for the unified social credit code suspected of matching the enterprise name on the Tianyancha website to determine whether the match is successful or whether it is indeed two different companies with particularly similar glyphs.
[0023] The final result is: 760 specialized, sophisticated and new listed companies were completely matched (string similarity = 1, weighted edit distance = 0), 10 specialized, sophisticated and new listed companies were suspiciously matched (string similarity was close to 1, weighted edit distance was close to 0), and by crawling the unified social credit code, it was finally confirmed that 7 were listed companies, 3 were different companies with names very similar to those of listed companies, and there were still 4152 companies that had no matching results (non-listed specialized, sophisticated and new companies).
[0024] In addition, we also conducted a robustness test of the glyph similarity algorithm, an improvement in the results of the character fuzzy matching algorithm based on the traditional edit distance, and a performance test for common font and font size combinations.
[0025] 1. Data cleaning and full-width / half-width character conversion.
[0026] To avoid the situation where the mixed use of full-width and half-width characters causes the inability to search and match, we convert common full-width characters into half-width characters, as shown in Table 1.
[0027] Table 1 Correspondence between full-width and half-width characters In order to convert all English letters in the company name string to uppercase, use all half-width characters, and remove titles such as "Stock Co., Ltd.", "Joint Stock Company", and "Limited Company", etc., to reduce the difficulty of subsequent matching, use the following Excel formula for processing (assuming that cell A2 is the original data): =UPPER(TRIM(SUBSTITUTE(SUBSTITUTE(SUBSTITUTE(ToHalfWidth(A2),"Co., Ltd.",""),"Joint Stock Co., Ltd.",""),"Limited by Share Ltd",""))).
[0028] 2. Definition of glyph similarity and calculation of similarity matrix.
[0029] To define glyph similarity, we try to understand it from the perspective of images. If two Chinese characters (with the same font and font size), after being converted into images, the higher the similarity of the images, the higher the glyph similarity of these two Chinese characters. Therefore, the HOG algorithm (Histogram of Oriented Gradients) can be used to extract the features of character images. The HOG algorithm returns a one-dimensional array, called the HOG descriptor, which contains the feature information of the entire image. Therefore, comparing the similarity of two images is transformed into comparing the similarity of HOG descriptors.
[0030] If the resolutions of two images are the same and the HOG parameters used are consistent, then the dimensions (lengths) of the HOG descriptors are also the same. This means that glyph similarity can ultimately be transformed into calculating the cosine of the angle between two high-dimensional vectors (i.e., HOG descriptors). As shown in Equation 1-1.
[0031] (1-1) The core code for the above calculation is as follows: The function generate_character_image is used to convert the given text into the corresponding image (the font and font size need to be specified), the extract_features function is used to calculate the HOG feature vector of the given image, and the calculate_similarity_matrix function uses the cosine of the angle to calculate the glyph similarity matrix between all pairs of character sets.
[0032] Taking the four characters "紫", "柴", "丁", and "香" given as examples in the competition questions, the above algorithm can generate images of the four characters, as shown in Figure 1 (a) of. It can be seen that "紫" and "柴" are more similar in glyph.
[0033] Further calculate the glyph similarity matrix of the above characters, as shown in Figure 1 (b) of. Except that the self-similarity of each character is 1 (diagonal elements), the glyph similarity between "紫" and "柴" is significantly higher than other character combinations, reaching 0.845, which is consistent with the conclusion of subjective observation. This shows that the glyph similarity algorithm designed by the present invention is effective and reasonable.
[0034] Comparing glyph similarity requires specifying characters, fonts, and font sizes. Since it is impossible to exhaust all possible situations, the present invention selects commonly used Chinese characters and characters, and font and font size combinations in common scenarios to calculate the glyph similarity matrix for subsequent use. The specific reasons for the selection are shown in Table 2.
[0035] Table 2 Scenarios and reasons for selecting characters, fonts, and font sizes In Table 2, the following points should be noted: (1) The full character list used by the present invention to generate the similarity matrix is in "\Final Work Submission\1 Enterprise Name String Fuzzy Matching\Step 1_Generate Character Similarity Matrix\Chinese Character Table.xlsx". (2) The file name after the font is the file name corresponding to the font in the submitted source code. (3) In order to reduce the size of the submitted content, the submitted work file only includes the similarity matrix of Microsoft YaHei No. 14 font generated by the program.
[0036] The corresponding file names are similarity_matrix_msyh.ttc_14.xlsx (matrix format) and similarity_list_msyh.ttc_14.csv (list format), and are used for subsequent fuzzy matching. Other result files can be generated by the program if necessary.
[0037] The 3606×3606 similarity matrix heat map of common characters under typical configuration (font = Microsoft YaHei, font size = 14) was calculated, as shown in Figure 2 As shown in the figure, except for the self-perfected similarity on the main diagonal, most characters have low similarity (the blue-green part of the heat map), but some characters have high similarity (reflected in the red dots outside the main diagonal in the figure), which is the root cause of OCR misrecognition and the basis for subsequent fuzzy matching of glyphs.
[0038] 3. Weighted Levenshtein distance calculation and string fuzzy matching.
[0039] The traditional Levenshtein distance (also known as the edit distance) is an algorithm for measuring the similarity between two strings. Its basic principle is to measure the similarity by calculating the minimum number of edit operations required to transform one string into another. Edit operations include inserting a character, deleting a character, or replacing a character. The steps to calculate the Levenshtein distance include:
[0040] The first step is to initialize the matrix: create a matrix of size (m+1)×(n+1), where m and n are the lengths of the two strings, respectively. The row and column labels of the matrix correspond to the number of characters from the left of the first and second strings, respectively. The matrix element D[i][j] represents the minimum edit distance required to convert the first i characters of the first string into the first j characters of the second string, and the first row and first column of the matrix are initialized to increasing values from 0 to m and from 0 to n, respectively. This represents the number of edit operations required to convert a string to an empty string.
[0041] The second step is to fill the matrix: for each element in the matrix) D[i][j], fill it using formula 1-2: Among them, cost is the cost of replacement. After filling, the value of the last element in the matrix) D[m][n] is the Levenshtein distance between the two strings. For the purpose of comparability, we often use the normalized edit distance and map it to the interval [0,1] using formula 1-3.
[0042] (1-3) In the traditional Levenshtein distance, if two characters are the same, the cost is 0; otherwise, the cost is 1, which does not take into account the similarity of characters. Therefore, the present invention introduces character similarity and defines cost as formula 1-4, that is, when the replaced characters are similar (similarity is higher than the threshold), cost = 1-character similarity, otherwise (similarity is lower than the threshold) a punitive large distance measurement (cost = 2) will be given, which is more conducive to widening the distinction compared to the traditional Levenshtein distance that assigns different character replacement costs to 1.
[0043] Among them, similarity represents the similarity between two characters. Similarity is queried from a pre-calculated similarity list. If similarity can be found, the similarity value is returned. If not, it means that the similarity is too low, and -1 is returned. At this time, cost will increase the penalty weight to 2.
[0044] In addition, since each character comparison requires querying a huge 3606×3606 matrix, the operation speed is slow. When implementing the algorithm, the present invention also makes optimizations in the following aspects to improve the query speed: (1) Retaining subsets: From the complete similarity matrix, pre-screen the similarity matrix subset of all characters involved in the names of "specialized, sophisticated, and innovative" enterprises and listed companies, discard other useless information, and reduce the resource overhead of the query.
[0045] (2) Sparse matrix: Set the number of similar characters for each character to no more than 15, and the similarity threshold to not less than 0.85, and remove most of the low-value elements in the matrix).
[0046] (3) Parallel computing: Through the multiprocessing library, implement the comparison between multiple strings at one time to improve the running efficiency.
[0047] 4. Confirmation of the matching results based on the Microsoft Power Automate crawler.
[0048] The fuzzy matching of enterprise names cannot assert whether the minor differences between two strings are due to OCR recognition errors or two different enterprises with similar names and glyphs (although humans can make most judgments based on experience), as shown in Table 3.
[0049] Table 3 Fuzzy matching results considering glyph similarity (incomplete matching part) Considering that the unified social credit code can uniquely identify an enterprise, the idea of this part is: Send the matched company name to Tianyancha website for query, obtain the unified social credit code of the first company (the company with the most similar name to the query string) in each returned query result, and compare them. If the two are equal, it is the same company and the matching is successful; otherwise, the matching fails.
[0050] Use Microsoft Power Automate to build this code-free crawler. Since Tianyancha has a strict anti-crawler mechanism and it is difficult for crawlers based on headless browsers to achieve the expected results, Power Automate is used to simulate human mouse and keyboard operations to access the website, and its built-in visual recognition function is used to determine the position of web page clicks (to avoid triggering the anti-crawler mechanism by traditional crawlers to obtain web page elements) to obtain the value of the unified social credit code.
[0051] Run the crawler, and the final results are shown in Table 4. It can be seen that 3 enterprises finally failed to match, and the listed companies with very similar relevant characters in their names are actually other enterprises. For example, "尔" and "乐", "裕" and "榕", etc.
[0052] Table 4 Reconfirmation of the incomplete matching results Compare the designed standard Levenshtein distance with the optimized Levenshtein distance of the present invention. The time spent on finally matching the list of 4922 specialized, sophisticated, distinctive and innovative enterprises is as follows: Table 5 Efficiency comparison between improved algorithm and traditional algorithm It can be seen from Table 5 that without using the library function to calculate the standard Levenshtein distance, the improved algorithm achieves a significant improvement in matching accuracy by increasing the time consumption.
[0053] 5. Robustness analysis of glyph similarity algorithm.
[0054] When calculating the glyph similarity matrix, you need to tell the program which font and size to use. However, in the fuzzy matching process, information such as which font and size the string to be matched comes from is generally missing. This requires that the similarity matrix calculated by the algorithm under different fonts and sizes has a certain stability.
[0055] The correlation coefficient is used to measure the difference between two similarity matrices. The closer the correlation coefficient is to 1, the smaller the effect of font size on the calculated similarity, otherwise it is greater.
[0056] The correlation coefficients of the glyph similarity matrices under different font sizes are all above 0.4. When the font size is small (size 9), there are certain differences in the glyph similarity matrices generated by different font sizes. In this case, it is necessary to clearly specify the font size used in the optical character image to obtain better results. When the font size is moderate and large, the correlation coefficient values under various combinations are large, which means that the font and font size have little effect on the glyph similarity matrix.
[0057] In general, under different font and size combinations, especially when the font size is larger, the glyph similarity matrix is highly consistent, and the algorithm is generally robust.
[0058] 6. Comparison between the algorithm of the present invention and the traditional string fuzzy matching algorithm.
[0059] The main difference between the algorithm of the present invention and the traditional string fuzzy matching algorithm is that the similarity of the glyphs is taken into account when calculating the Levenshtein distance. In order to analyze the effect of this improvement, we compare the matching accuracy and algorithm efficiency.
[0060] Comparison of matching results: Table 6 lists the matching results of the traditional Levenshtein distance when the glyph ss similarity is not considered (only the incomplete corner matching part is shown), and the unified social credit code verification is performed at the same time. It can be seen that the traditional algorithm has a lot of misrecognition (FALSE results), while the algorithm of the present invention avoids such misrecognition to a considerable extent after considering the glyph similarity.
[0061] Table 6 String matching results without considering glyph similarity (incomplete matching part) However, it is worth noting that the following three company names that were successfully matched by the traditional algorithm were not matched by the improved algorithm of the present invention, as shown in Table 7. The reasons involved mainly include: the company name was changed, the error was not caused by OCR recognition, and uncommon Chinese characters appeared.
[0062] Table 7 Samples and reasons for successful matching of the traditional algorithm and failed matching of the improved algorithm The present invention improves the Levenshtein distance measurement between character strings based on weighted glyph similarity, and proposes a new character string fuzzy matching algorithm, which is particularly suitable for fuzzy matching of character strings that may be misidentified by OCR. By matching the listed companies in the list of specialized, sophisticated and innovative enterprises identified by OCR, it is found that compared with the character string fuzzy matching algorithm based on the traditional Levenshtein distance measurement, the improved algorithm of the present invention effectively reduces the misidentification situation and balances the algorithm efficiency and accuracy. The robustness test shows that the algorithm of the present invention can adapt to the results of OCR recognition of different fonts and font sizes.
[0063] The embodiment of the present invention provides a string fuzzy matching system, including: The matrix construction module is used to obtain multiple characters in different font sizes; select any two characters from the multiple characters, extract the image feature vector of each character through optical character recognition technology OCR, and obtain the cosine distance of the angle between the image feature vectors of the two characters as the similarity; and construct a glyph similarity matrix according to the similarity between any two characters in the multiple characters.
[0064] A weighted edit distance construction and use module is used to weight the cost of describing the cost of replacing characters in the replacement operation according to the replacement operation used to replace one character of a string with another character in the edit distance method, so as to construct a weighted edit distance method; the character similarity weighting includes: when the similarity of the replaced character is higher than the set similarity threshold, the cost of the replacement operation is the difference between the first set value and the similarity; otherwise, the cost of the replacement operation is the second set value, and the second set value indicates that a penalty value higher than the difference between the first set value and the similarity is assigned to the cost of the replacement operation; the weighted edit distance method is used for the two strings to be matched, and during the replacement operation, the similarity between the two characters that need to be replaced in the two strings to be matched is queried according to the glyph similarity matrix, and the cost of the replacement operation is weighted by character similarity to obtain the result of the weighted edit distance; the result of the weighted edit distance is normalized to obtain the similarity of the two strings to be matched.
[0065] The similarity matching module is used to determine the matching result according to the relationship between the similarity of two character strings to be matched and a preset value.
[0066] Innovations of the present invention: First, simple data cleaning. This is mainly achieved through custom VBA functions and formulas in Excel. First, use Excel formulas to remove the original company names and the names of listed companies such as "company", "limited liability company", "limited liability company", and "joint stock company", and focus on matching company names. Secondly, write a custom VBA function ToHalfWidth to convert all full-width characters in the string into corresponding half-width characters in EXCEL to handle the inconsistencies in full-width / half-width brackets and numbers that may appear in the company name list.
[0067] Second, calculate the character similarity. Based on the HOG algorithm commonly used in OCR recognition, the image feature vector of each character is extracted, and the angle cosine distance is used to define the character similarity between two characters. On this basis, four commonly used fonts (Microsoft YaHei, Kaiti GB2312, Fangsong GB2312, Songti) and four commonly used font sizes (Small Five, Four, Two, and First) are selected to calculate and store the 3606×3606 glyph similarity matrix basic data between any two characters in the commonly used Chinese characters and symbols (2500 commonly used characters, 1000 commonly used characters, and 106 common symbols) under different character size combinations.
[0068] Third, a custom weighted Levenshtein distance is used to calculate string similarity. Based on the above basic character similarity data, by reducing the replacement cost between two similar glyphs (reducing the distance), a penalty factor is imposed on the replacement of non-similar glyph characters (increasing the distance), so that the string that is incorrectly recognized by OCR has a higher similarity with the correct string under the algorithm of the present invention, while other strings (different characters and far-flung glyphs) have a lower similarity with the correct string, even if the traditional Levenshtein distances of these two types of strings are equal. In this step of calculation, due to the large amount of calculation, in order to improve the matching efficiency, the multiprocessing library is used to implement parallel computing functions.
[0069] Crawling the unified social credit code finally confirms the fuzzy matching results. For non-perfect matching samples (weighted Levenshtein distance is not 0, but small), due to the strict anti-crawling mechanism of the Tianyancha website, it is almost impossible to implement crawlers based on headless browsers. Therefore, a code-free visual crawler is designed through Microsoft Power Automate to directly simulate human keyboard and mouse operations, search for the unified social credit code on the Tianyancha website that is suspected to match the company name, and determine whether the match is successful or whether it is indeed two different companies with particularly similar names.
[0070] Practicality of the present invention: The present invention is used in the list of "specialized, sophisticated and innovative" enterprises that may have OCR misidentification, and in combination with the standard character string list of listed company names, to find listed companies in the "specialized, sophisticated and innovative" enterprises. Since it is essentially a string fuzzy matching problem, it can also be further extended to other string fuzzy matching problems that mainly deal with scenes with possible OCR misidentification and other glyph similarity.
[0071] A specific embodiment is as follows: Step 1: Pre-calculate the glyph similarity matrix (3606×3606) of common Chinese characters and symbols under different font and size combinations based on the HOG algorithm and angle cosine distance, and store it on the hard disk for future use.
[0072] Step 2: Perform simple data cleaning on the list of matching strings and the list of standard strings to remove different company names and convert full-width characters to half-width characters.
[0073] Step 3: For each string to be matched, calculate its similarity with each standard string, and return the result of the maximum similarity of string matching and the similarity value. Four inputs are required: Input 1: the string to be matched, which in this scenario is the name of a specialized, sophisticated and innovative enterprise identified by OCR; Input 2: a list of standard strings, which in this scenario is a list of listed companies; Input 3: a glyph similarity matrix; Input 4: a string similarity threshold, below which a match is considered unsuccessful.
[0074] Step 4: For matches with string similarity = 1, the match is considered successful, and the matched standard string is the matching result; for matches with string similarity ∈ [similarity threshold, 1], the matched standard string is included in the suspicious result for further verification; for matches with string similarity < similarity threshold, the match is considered failed.
[0075] Step 5: For suspicious matching samples, the crawler automatically queries whether the unified social credit code of the first search result of the two matching strings on the Tianyancha website is equal. If they are equal, it is judged as a successful match (the same company); if they are not equal, it is judged as a failed match (two different companies). It also means that the characters of the names of the two companies are very similar.
[0076] The above-mentioned embodiments only express several implementation methods of the present invention, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the attached claims.
Claims
1. A string fuzzy matching method, characterized in that: The following steps are involved: Get multiple characters in different font sizes; Select any two characters from the multiple characters, extract the image feature vector of each character by optical character recognition technology (OCR), and obtain the angle cosine distance between the image feature vectors of the two characters as the similarity; construct a glyph similarity matrix according to the similarity between any two characters from the multiple characters; According to a replacement operation used to replace one character of a string with another character in the edit distance method, the cost describing the cost of replacing the character in the replacement operation is weighted by character similarity to construct a weighted edit distance method; The character similarity weighting includes: when the similarity of the replaced character is higher than the set similarity threshold, the cost of the replacement operation is the difference between the first set value and the similarity; otherwise, the cost of the replacement operation is the second set value, and the second set value indicates that a penalty value higher than the difference between the first set value and the similarity is assigned to the cost of the replacement operation; The weighted edit distance method is used for the two strings to be matched. During the replacement operation, the similarity between the two characters to be replaced in the two strings to be matched is queried according to the glyph similarity matrix, and the cost of the replacement operation is weighted by the character similarity to obtain the result of the weighted edit distance; the result of the weighted edit distance is normalized to obtain the similarity of the two strings to be matched; The matching result is determined based on the relationship between the similarity of the two character strings to be matched and a preset value.
2. A string fuzzy matching method as claimed in claim 1, characterized in that: The weighted edit distance method is used for the two strings to be matched, and the specific steps include: Create a ( m +1)×( n +1); in, m is the length of the first string to be matched, n is the length of the second string to be matched, and the row index of the matrix i The number of characters from left to right corresponding to the first string, the column label of the matrix j Corresponding to the number of characters in the second string from left to right, each matrix element in the matrix D [ i ][ j ] indicates the first character string i Converts the first character of the second string to j The minimum edit distance among the edit distances required for characters; Fill the matrix: in, D [ i -1][ j ]+1 is the delete operation, D [ i ][ j -1]+1 is the insertion operation, D [ i -1][ j -1]+cos t For the replacement operation, cos t is the cost of the replacement operation; Query the character similarity in the similarity matrix between the character to be replaced in the first string and the character used as a replacement reference in the second string, and weight the cost by character similarity: in, if Indicates conditions, similarity Indicates the character similarity, threshold Indicates the set similarity threshold; Filled, the matrix D [ m ][ n ] represents the weighted edit distance between the first character string and the second character string.
3. A string fuzzy matching method as claimed in claim 1, characterized in that: The step of obtaining the angle cosine distance between the image feature vectors of two characters as the similarity comprises: When two characters are full-width characters, convert the full-width characters to half-width characters; The HOG descriptor of each character is obtained by using the Histogram of Oriented Gradients (HOG) algorithm in the optical character recognition technology (OCR), wherein the HOG descriptor is a one-dimensional array; the similarity comparison between two characters is converted into the similarity comparison between the HOG descriptors of the two characters; Get the cosine distance of the angle between the HOG descriptors of two characters as the similarity.
4. A string fuzzy matching method as claimed in claim 2, characterized in that: The query method needs to be optimized to increase the query speed in the similarity matrix between the characters to be replaced in the first character string and the characters used as replacement reference in the second character string. The optimization specifically includes: If the first character string and the second character string are both company names, the character similarities of all characters in the company name are screened in the similarity matrix to obtain a similarity matrix subset.
5. A string fuzzy matching method as claimed in claim 4, characterized in that: The optimization also includes: when each character in each character string is queried in the similarity matrix, the number of similar characters queried is no more than 15, and the similarity threshold is set to be no less than 0.
85.
6. A string fuzzy matching method as claimed in claim 4, characterized in that: The optimization further includes: simultaneously querying the character similarity of each character in the first character string and each character in the second character string in the similarity matrix on multiple threads.
7. A string fuzzy matching method as claimed in claim 1, characterized in that: The step of determining the matching result according to the relationship between the similarity of the two character strings to be matched and the preset value comprises: When the similarity is 1, it means the match is successful; When the similarity is greater than the set similarity threshold and less than 1, it indicates a suspicious match result; When the similarity is less than the set similarity threshold, it means that the matching fails.
8. A string fuzzy matching method as claimed in claim 7, characterized in that: The suspicious matching result requires an in-depth comparison of the two strings to be matched, and the specific steps of the in-depth comparison include: When the two character strings to be matched are company names, the company names of the two character strings to be matched are searched separately on a website that can query the unified social credit code, and the search results are obtained and compared; If the comparison results are the same, it means the match is successful, and the company names of the two strings to be matched are the same company; When the comparison results are different, it indicates that the matching fails, and it means that the corporate names of the two character strings to be matched have similar glyphs.
9. A string fuzzy matching system, characterized in that: include: A matrix construction module is used to obtain multiple characters of different font sizes; select any two characters from the multiple characters, extract the image feature vector of each character by optical character recognition technology OCR, and obtain the angle cosine distance between the image feature vectors of the two characters as the similarity; construct a glyph similarity matrix according to the similarity between any two characters in the multiple characters; A weighted edit distance construction and use module is used to perform character similarity weighting on the cost of describing the cost of replacing characters in the replacement operation according to the replacement operation used to replace one character of a string with another character in the edit distance method, so as to construct a weighted edit distance method; character similarity weighting includes: when the similarity of the replaced character is higher than the set similarity threshold, the cost of the replacement operation is the difference between the first set value and the similarity; otherwise, the cost of the replacement operation is the second set value, and the second set value indicates that a penalty value higher than the difference between the first set value and the similarity is assigned to the cost of the replacement operation; the weighted edit distance method is used for the two strings to be matched, and during the replacement operation, the similarity between the two characters to be replaced in the two strings to be matched is queried according to the glyph similarity matrix, and the cost of the replacement operation is weighted by character similarity to obtain the result of the weighted edit distance; the result of the weighted edit distance is normalized to obtain the similarity of the two strings to be matched; The similarity matching module is used to determine the matching result according to the relationship between the similarity of two character strings to be matched and a preset value.
Citation Information
Patent Citations
Correction method and apparatus for OCR result
CN107220639A
An optical character recognition and error correction method based on a dictionary
CN109711412A
Character string matching method and device based on similarity measurement and storage medium
CN116911253A
Error correction method for OCR medical record text
CN119337866A
Cited By
Password analysis method and device, electronic equipment and storage medium
CN121030730A
Text-to-instruction system based on fuzzy pinyin matching and parameter robust analysis
CN121659895A