Method and System for Address Information Matching Based on Similarity Algorithm
Through the address information matching method based on the similarity algorithm, the problem that the address filled in by the customer cannot match the cell level is solved, and efficient and accurate address matching effect is achieved.
Patent Information
- Application Number
- CN202411155736.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-22
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2044-08-22
AI Technical Summary
The prior art cannot match the address filled in by customers to the cell level, and there are problems of low matching efficiency and insufficient accuracy.
The address information matching method based on the similarity algorithm is adopted. By obtaining the address filled in by the customer and the set of cells to be matched, the non-numbers in the address are filtered using regular expressions, and the set of cells are arranged in non-ascending order according to the length of the string. Then, based on the preset similarity algorithm, the matching score between the cell to be matched and the filtered address is calculated. If the first match score is obtained as 1, the corresponding cell is used as the standardization result.
The efficiency and accuracy of address information matching are improved, and the addresses filled in by customers can be accurately matched to the cell level, solving the problems of low matching efficiency and insufficient accuracy in the prior art.
Smart Images

Figure CN119066441B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of big data, and specifically to a method, system, storage medium, and electronic device for address information matching based on a similarity algorithm. Background Art
[0002] In order to better serve customers, financial institutions use digital means to conduct business activities in-depth in communities and need to standardize the addresses filled in by customers to complete the address information matching work.
[0003] In the related art, Patent CN117610560A discloses a method for supplementing and standardizing administrative divisions of Chinese addresses, which uses the classic term frequency-inverse document TF-IDF algorithm to standardize addresses and outputs the matching results at the administrative planning levels of provinces, cities, and districts, but does not involve the community level.
[0004] In view of this, it is necessary to provide a technical solution that can match the addresses filled in by customers to communities. Summary of the Invention
[0005] (1) Technical Problems to be Solved
[0006] Aiming at the deficiencies of the prior art, the present invention provides a method, system, storage medium, and electronic device for address information matching based on a similarity algorithm, which solves the technical problem that the addresses filled in by customers cannot be matched to communities.
[0007] (2) Technical Solutions
[0008] To achieve the above objectives, the present invention is realized through the following technical solutions:
[0009] A method for address information matching based on a similarity algorithm includes:
[0010] Obtain the address filled in by the customer and the set of communities to be matched;
[0011] Use a regular expression to filter out non-alphanumeric Chinese and English in the address, and arrange all the communities to be matched in the community set in non-ascending order according to the string length;
[0012] Based on a preset similarity algorithm, calculate the matching scores between the arranged communities to be matched and the filtered address one by one. If the matching score obtained for the first time in the calculation process is 1, then use the corresponding community to be matched as the standardized result of the address; otherwise, after the calculation traversal is completed, use the community to be matched corresponding to the highest matching score obtained for the first time as the standardized result of the address.
[0013] Preferably, define the filtered address as the first string, and define any cell to be matched as the second string; based on a preset similarity algorithm, calculate the matching scores between the arranged cells to be matched and the filtered address one by one, including:
[0014] Count the number of common characters between the first string and the second string, and calculate the character ratio score in combination with the number of characters in the second string; calculate the position similarity score based on the positional distance relationship of all common characters in the first string and the second string respectively; and calculate the character similarity score through weighted calculation based on the character ratio score and the position similarity score.
[0015] Based on a preset conversion rule, convert the first string into a first pinyin sequence, and convert the second string into a second pinyin sequence.
[0016] Count the number of common pinyins between the first pinyin sequence and the second pinyin sequence, and calculate the pinyin ratio score in combination with the number of pinyins in the second pinyin sequence; count the longest common subsequence of the first pinyin sequence and the second pinyin sequence, and calculate the common sequence ratio score in combination with the number of pinyins in the first pinyin sequence; and calculate the pinyin similarity score through weighted calculation based on the pinyin ratio score and the common sequence ratio score.
[0017] Based on the character similarity score and the pinyin similarity score, obtain the final matching score through weighting.
[0018] Preferably, define the position subscripts of all common characters in the first string S 1 as the first position subscript set x 1 , and define the position subscripts of all common characters in the second string S 2 as the second position subscript set x 2 ; define the distance between every two adjacent position subscripts in the first position subscript set x 1 as the first distance set n 1 , and define the distance between every two adjacent position subscripts in the second position subscript set x 2 as the second distance set n 2 ; the position similarity score is expressed as:
[0019]
[0020] where sim(S 1 , S 2 ) represents the position similarity score between S 1 and S 2 ; len(x 1 ) represents x1 The number of middle position subscripts; k 1 , k 2 are the corresponding control coefficients respectively; abs represents the absolute value function; n 1,i , n 2,i respectively represent 1 the i-th distance in n 2 and n
[0021] Preferably, the conversion rule includes:
[0022] (a) Convert one Chinese character into one pinyin;
[0023] (b) Unify the pinyins corresponding to all Chinese characters into front nasal sounds or back nasal sounds;
[0024] (c) Unify the pinyins starting with L and N corresponding to all Chinese characters into those starting with L or N;
[0025] (d) Retain English characters without any processing;
[0026] (e) First convert numbers into Chinese characters, and then convert the corresponding Chinese characters into pinyins.
[0027] A system for address information matching based on a similarity algorithm, comprising:
[0028] An information acquisition module, configured to acquire the address filled in by the customer and the set of communities to be matched;
[0029] An information processing module, configured to filter non-numeric Chinese and English in the address by using a regular expression, and arrange all the communities to be matched in the community set in non-ascending order according to the string length;
[0030] An information matching module, configured to calculate the matching scores between the arranged communities to be matched and the filtered address one by one based on a preset similarity algorithm. If the matching score obtained for the first time in the calculation process is 1, the corresponding community to be matched is used as the standardized result of the address; otherwise, after the calculation traversal is completed, the community to be matched corresponding to the highest matching score obtained for the first time is used as the standardized result of the address.
[0031] Preferably, the filtered address is defined as the first string, and any community to be matched is defined as the second string; calculating the matching scores between the arranged communities to be matched and the filtered address one by one based on the preset similarity algorithm; includes:
[0032] Count the number of common characters between the first string and the second string, and calculate the character ratio score in combination with the number of characters in the second character; calculate the position similarity score based on the position distance relationship of all common characters in the first string and the second string respectively; and calculate the character similarity score by weighted calculation based on the character ratio score and the position similarity score;
[0033] Based on a preset conversion rule, convert the first string into a first pinyin sequence, and convert the second string into a second pinyin sequence;
[0034] Count the number of common pinyins between the first pinyin sequence and the second pinyin sequence, and calculate the pinyin ratio score in combination with the number of pinyins in the second pinyin sequence; count the longest common subsequence of the first pinyin sequence and the second pinyin sequence, and calculate the common sequence ratio score in combination with the number of pinyins in the first pinyin sequence; and calculate the pinyin similarity score by weighted calculation based on the pinyin ratio score and the common sequence ratio score;
[0035] Based on the character similarity score and the pinyin similarity score, obtain the final matching score by weighting.
[0036] Preferably, define the position subscripts of all common characters in the first string S 1 as the first position subscript set x 1 , and define the position subscripts of all common characters in the second string S 2 as the second position subscript set x 2 ; define the distance between every two adjacent position subscripts in the first position subscript set x 1 as the first distance set n 1 , and define the distance between every two adjacent position subscripts in the second position subscript set x 2 as the second distance set n 2 ; the position similarity score is expressed as:
[0037]
[0038] where sim(S 1 , S 2 ) represents the position similarity score between S 1 and S 2 ; len(x 1 ) represents the number of position subscripts in x 1 ; k 1 , k 2 are the corresponding control coefficients respectively; abs represents the absolute value function; n 1,i , n 2,i represent n 1and n 2 the i-th distance in
[0039] Preferably, the conversion rule includes:
[0040] (a) Convert one Chinese character into one pinyin;
[0041] (b) Unify the pinyins corresponding to all Chinese characters into front nasal sounds or back nasal sounds;
[0042] (c) Unify the pinyins starting with L and N corresponding to all Chinese characters into those starting with L or N;
[0043] (d) Retain English characters without any processing;
[0044] (e) First convert numbers into Chinese characters, and then convert the corresponding Chinese characters into pinyins.
[0045] A storage medium stores a computer program for address information matching based on a similarity algorithm, wherein the computer program causes a computer to execute the method for address information matching as described above.
[0046] An electronic device includes:
[0047] One or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the programs include a method for executing the address information matching as described above.
[0048] (III) Beneficial effects
[0049] The present invention provides a method, a system, a storage medium and an electronic device for address information matching based on a similarity algorithm. Compared with the prior art, the following beneficial effects are achieved:
[0050] In the present invention: First, obtain the address filled in by the customer and the set of communities to be matched. Then, use a regular expression to filter out non-numeric Chinese and English in the address, and arrange all the communities to be matched in the set of communities in non-ascending order according to the string length. Finally, based on a preset similarity algorithm, calculate the matching scores between the arranged communities to be matched and the filtered address one by one. If the matching score obtained for the first time in the calculation process is 1, the corresponding community to be matched is used as the standardized result of the address to improve the matching efficiency; otherwise, after the calculation traversal is completed, the community to be matched corresponding to the highest matching score obtained for the first time is used as the standardized result of the address. By introducing the set of communities to be matched and based on the similarity algorithm, the present invention can match the address filled in by the customer to the community level. Description of the drawings
[0051] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0052] Figure 1 It is a block diagram of a method for address information matching based on a similarity algorithm provided by an embodiment of the present invention;
[0053] Figure 2 It is a flowchart of a method for address information matching based on a similarity algorithm provided by an embodiment of the present invention;
[0054] Figure 3 It is a structural block diagram of a system for address information matching based on a similarity algorithm provided by an embodiment of the present invention. Detailed implementation manners
[0055] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention are clearly and completely described below. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0056] The embodiments of the present application solve the technical problem of being unable to match the address filled in by the customer to the community by providing a method, system, storage medium, and electronic device for address information matching based on a similarity algorithm.
[0057] The overall idea of the technical solutions in the embodiments of the present application to solve the above technical problems is as follows:
[0058] In the present invention, regular expressions are used to filter non-alphanumeric Chinese and English in the address, and all the to-be-matched communities in the community set are sorted in non-ascending order according to the string length. Based on a preset similarity algorithm, the matching scores between the sorted to-be-matched communities and the filtered address are calculated one by one, and the to-be-matched community in the most matching result is used as the standardized result of the address.
[0059] That is, the embodiments of the present invention introduce a set of to-be-matched communities and, based on a similarity algorithm, are able to match the address filled in by the customer to the community level; in addition, the community corresponding to the first obtained matching score of 1 is used as the standardized result, greatly improving the matching efficiency.
[0060] Furthermore, considering that the addresses filled in by customers in practice often have many defects such as typos, homophones, address abbreviations, and incomplete filling, it brings considerable technical difficulties to the address information matching work at the community level.
[0061] In order to accurately standardize the addresses filled in by customers of banks or other financial institutions to the community level, the embodiments of the present invention are based on the similarity algorithm of traditional string matching, and innovatively propose similarity algorithms for character matching degree and pinyin matching degree, so that the finally calculated matching score includes factors such as the ratio of the number of matching characters and pinyin, the relative position of the match, and the continuity of the match, greatly improving the matching accuracy in practical applications.
[0062] In order to better understand the above technical solution, the above technical solution will be described in detail below in conjunction with the accompanying drawings of the specification and specific implementation manners.
[0063] Embodiment 1:
[0064] As Figure 1 shown, the embodiments of the present invention provide a method for address information matching based on a similarity algorithm, including:
[0065] S1. Obtain the address filled in by the customer and the set of communities to be matched;
[0066] S2. Use a regular expression to filter out non-alphanumeric Chinese and English in the address, and arrange all the communities to be matched in the set of communities in non-increasing order according to the string length;
[0067] S3. Based on a preset similarity algorithm, calculate the matching score between each arranged community to be matched and the filtered address one by one. If the matching score obtained for the first time in the calculation process is 1, use the corresponding community to be matched as the standardized result of the address; otherwise, after the calculation traversal is completed, use the community to be matched corresponding to the highest matching score obtained for the first time as the standardized result of the address.
[0068] The embodiments of the present invention introduce the set of communities to be matched and, based on the similarity algorithm, are able to match the address filled in by the customer to the community level; in addition, using the community corresponding to the matching score of 1 obtained for the first time as the standardized result greatly improves the matching efficiency.
[0069] As Figure 2 shown, Figure 2 the flowchart of the above method for address information matching is given. Next, each step of the solution will be described in detail in conjunction with Figure 2 :
[0070] In step S1, obtain the address filled in by the customer and the set of communities to be matched.
[0071] In this step, the original address filled in by the customer is collected. This original address can be collected online or offline and finally converted into a data format suitable for computer processing.
[0072] And a set of communities to be matched is pre-constructed. This set of communities is used to store multiple communities to be matched in the same data format as the aforementioned one. Any community to be matched is stored in the form of a string. According to the actual situation, the strings corresponding to different communities may include Chinese characters, English characters, numbers, etc., but special characters similar to "#" and "$" are excluded. In addition, this set of communities needs to be updated and maintained regularly to ensure that the information is as accurate and complete as possible, providing basic support for the subsequent address information matching link.
[0073] In step S2, regular expressions are used to filter out non-alphanumeric Chinese and English in the address, and all the communities to be matched in the set of communities are sorted in non-ascending order according to the string length.
[0074] In this step, necessary preprocessing work is carried out on the obtained address filled in by the customer and the set of communities to be matched. Specifically:
[0075] Regular expressions are used to filter out non-alphanumeric Chinese and English in the address to exclude special characters similar to "#" and "$" in the address. It should be noted that regular expressions (Regular Expression) are a powerful tool for matching and operating on text. It is a pattern composed of a series of characters and special characters used to describe the text pattern to be matched.
[0076] In order to cooperate with the subsequent calculation process to obtain a matching score of 1 for the first time, the corresponding community to be matched is used as the standardized result of the address to improve the matching efficiency. This step also sorts all the communities to be matched in the set of communities in non-ascending order according to the string length. The sorting result is used to determine the participation order of each community to be matched in the subsequent matching score calculation process.
[0077] In step S3, based on a preset similarity algorithm, the matching scores between the sorted communities to be matched and the filtered address are calculated one by one. If the matching score of 1 is obtained for the first time in the calculation process, the corresponding community to be matched is used as the standardized result of the address; otherwise, after the calculation traversal is completed, the community to be matched corresponding to the highest matching score obtained for the first time is used as the standardized result of the address.
[0078] In this step, based on a similarity algorithm composed of character matching degree and pinyin matching degree, the matching scores between the sorted communities to be matched and the filtered address are calculated one by one. As Figure 2 shown, the community corresponding to the first obtained matching score of 1 is used as the standardized result, and other communities to be matched are no longer traversed, greatly improving the matching efficiency.
[0079] When the matching score is not calculated to be 1, then as Figure 2 shown, traverse all the cells to be matched, and use the cell to be matched corresponding to the highest matching score obtained for the first time as the standardized result of the address. In other words, if there are multiple identical highest matching scores, in this step, select the cell to be matched with the longest string length that participates in the matching calculation for the first time as the standardized result.
[0080] It can be understood that when only cells with very low scores are matched (in this case, it is very likely that the customer's cell is not accurately matched), in this case, reasonable thresholds can be set for differential processing. For example, set the threshold to 0.7 - 0.9, and add a label to the cell corresponding to the highest matching score lower than the threshold to remind the relevant staff to review it in time.
[0081] In order to accurately standardize the address filled in by customers of banks or other financial institutions to the cell level, the embodiment of the present invention is based on the similarity algorithm of traditional string matching and innovatively proposes a similarity algorithm of character matching degree and pinyin matching degree.
[0082] Correspondingly, the process of calculating the matching score in this step includes:
[0083] First, execute S100: Define the filtered address as the first string S 1 , and define any cell to be matched as the second string S 2 .
[0084] Secondly, execute S200: Calculate the character similarity score:
[0085] (1) Count the number of common characters between the first string S 1 and the second string S 2 , and combine the number of characters of the second string S 2 to calculate the character ratio score; expressed as:
[0086] P(S 1 , S 2 ) = p * slen(S 1 , S 2 ) / len(S 2 )
[0087] where P(S 1 , S 2 ) represents the character ratio score; p represents the penalty coefficient, and slen(S 1 , S 2 ) represents the number of common characters between S 1 and S 2 ; len(S 2) is S 2 the number of characters.
[0088] (2) Based on the position distance relationship of all common characters in the first string S 1 and the second string S 2 , calculate the position similarity score; including:
[0089] Define the position subscripts of all common characters in the first string S 1 as the first position subscript set x 1 , and define the position subscripts of all common characters in the second string S 2 as the second position subscript set x 2 .
[0090] Define the distance between every two adjacent position subscripts in the first position subscript set x 1 as the first distance set n 1 , and define the distance between every two adjacent position subscripts in the second position subscript set x 2 as the second distance set n 2 .
[0091] For example, S 1 = [Shi, Ji, Jin, Yuan, Guang, Chang] and S 2 = [Shi, Ji, Guang, Chang], then x 1 is [1, 2, 6], x 2 is [1, 2, 4], n 1 is [1, 4], n 2 is [1, 2].
[0092] Therefore, the position similarity score is expressed as:
[0093]
[0094] where sim(S 1 , S 2 ) represents the position similarity score between S 1 and S 2 ; len(x 1 ) represents the number of position subscripts in x 1 ; k 1 , k 2 are the corresponding control coefficients respectively; abs represents the absolute value function; n 1,i , n 2,i respectively represent the i-th distance in n 1 and n 2 .
[0095] It should be noted that in the embodiments of the present invention, on the basis of character correspondence, the position similarity score fully considers the position distance relationship between common characters, so that when there are several identical characters in the filtered address string or the cell string to be matched, it can be accurately recognized, thereby avoiding taking a cell that only has character correspondence but different character positions compared with the address as a matching result.
[0096] (3) Based on the character ratio score and the position similarity score, calculate the character similarity score by weighted calculation; including:
[0097] Define k as the weight ratio of the character ratio score, and the value of k is dynamically adjusted according to the v string length of S, which is expressed as follows:
[0098] k = min(len(S1) / (len(S1)+2), 0.75)
[0099] where min represents the minimization function.
[0100] Then the character similarity score is expressed as:
[0101] A = k * P(S 1 , S 2 ) + (1 - k) * sim(S 1 , S 2 )
[0102] Execute S300 again to calculate the pinyin similarity score:
[0103] (1) Based on the preset conversion rules, convert the first string S 1 into the first pinyin sequence Q 1 , and convert the second string S 2 into the second pinyin sequence Q 2 .
[0104] Exemplarily, the conversion rules here include:
[0105] (a) Convert one Chinese character into one pinyin;
[0106] (b) Unify the pinyins corresponding to all Chinese characters into front nasal sounds or back nasal sounds;
[0107] (c) Unify the pinyins starting with L and N corresponding to all Chinese characters into starting with L or N;
[0108] (d) Keep English characters without any processing;
[0109] (e) First convert numbers into Chinese characters, and then convert the corresponding Chinese characters into pinyins.
[0110] By applying the above transformation rules, the obtained pinyin sequence at least eliminates the adverse effects caused by homophones, typos, etc. in the corresponding string, and greatly improves the matching accuracy.
[0111] (2) Count the number of common pinyins in the first pinyin sequence Q 1 and the second pinyin sequence Q 2 , and calculate the pinyin ratio score in combination with the number of pinyins in the second pinyin sequence Q 1 ; expressed as:
[0112] P(Q 1 ,Q 1 ) = p * slen(Q 1 ,Q 2 ) / len(Q 2 )
[0113] where P(Q 1 ,Q v ) represents the pinyin ratio score.
[0114] It can be understood that the calculation process of the pinyin ratio score here completely refers to the calculation process of the character similarity score.
[0115] (3) Count the longest common subsequence of the first pinyin sequence Q 1 and the second pinyin sequence Q 2 ; calculate the common sequence ratio score in combination with the number of pinyins in the first pinyin sequence Q 1 ;
[0116] (4) Based on the pinyin ratio score and the common sequence ratio score, calculate the pinyin similarity score by weighted calculation.
[0117] Similarly, define k ′ as the weight ratio of the pinyin ratio score, and the value of k ′ is dynamically adjusted according to the string length of Q 1 , expressed as follows:
[0118] k ′ = min(len(Q1) / (len(Q1) + 2), 0.75)
[0119] Then the pinyin similarity score is expressed as:
[0120]
[0121] Finally, execute S400 to calculate the matching score:
[0122] Based on the character similarity score and the pinyin similarity score, a weighted final matching score is obtained; expressed as:
[0123] Score = j * A+(1 - j)*B
[0124] Where j and 1 - j respectively represent the weighting coefficients of the character similarity score and the pinyin similarity score.
[0125] Thus, based on the similarity algorithm composed of the above character matching degree and pinyin matching degree, the embodiments of the present invention enable the finally calculated matching score to include factors such as the ratio of the number of matching characters and pinyins, the relative position of the match, and the matching continuity, greatly improving the matching accuracy in practical applications.
[0126] Embodiment 2:
[0127] As Figure 3 shown, the embodiments of the present invention provide a system for matching address information based on a similarity algorithm, including:
[0128] An information acquisition module, configured to acquire the address filled in by the customer and the set of communities to be matched;
[0129] An information processing module, configured to filter non - numeric Chinese and English in the address using a regular expression, and arrange all the communities to be matched in the set of communities in non - ascending order according to the string length;
[0130] An information matching module, configured to calculate the matching score between the arranged communities to be matched and the filtered address one by one based on a preset similarity algorithm. If the matching score obtained for the first time in the calculation process is 1, the corresponding community to be matched is used as the standardized result of the address; otherwise, after the calculation traversal is completed, the community to be matched corresponding to the highest matching score obtained for the first time is used as the standardized result of the address.
[0131] In an optional implementation manner, the filtered address is defined as the first string, and any community to be matched is defined as the second string; calculating the matching score between the arranged communities to be matched and the filtered address one by one based on the preset similarity algorithm; includes:
[0132] Count the number of common characters between the first string and the second string, and calculate the character ratio score in combination with the number of characters of the second string; calculate the position similarity score based on the position distance relationship of all common characters in the first string and the second string respectively; and calculate the character similarity score by weighted calculation based on the character ratio score and the position similarity score;
[0133] Based on a preset conversion rule, convert the first string into a first pinyin sequence and convert the second string into a second pinyin sequence;
[0134] Count the number of common pinyins in the first pinyin sequence and the second pinyin sequence, and calculate a pinyin ratio score in combination with the number of pinyins in the second pinyin sequence; count the longest common subsequence of the first pinyin sequence and the second pinyin sequence, and calculate a common sequence ratio score in combination with the number of pinyins in the first pinyin sequence; and calculate a pinyin similarity score through weighted calculation based on the pinyin ratio score and the common sequence ratio score;
[0135] Based on the character similarity score and the pinyin similarity score, obtain a final matching score through weighting.
[0136] In an optional embodiment, define the position subscripts of all common characters in the first string S 1 as a first position subscript set x 1 , and define the position subscripts of all common characters in the second string S 2 as a second position subscript set x 2 ; define the distance between every two adjacent position subscripts in the first position subscript set x 1 as a first distance set n 1 , and define the distance between every two adjacent position subscripts in the second position subscript set x 2 as a second distance set n 2 ; the position similarity score is expressed as:
[0137]
[0138] wherein, sim(S 1 , S 2 ) represents the position similarity score between S 1 and S 2 ; len(x 1 ) represents the number of position subscripts in x 1 ; k 1 , k 2 are respectively corresponding control coefficients; abs represents the absolute value function; n 1,i , n 2,i respectively represent the i-th distance in n 1 and n 2 .
[0139] In an optional embodiment, the conversion rule includes:
[0140] (a) Convert one Chinese character into one pinyin;
[0141] (b) Unify the pinyin corresponding to all Chinese characters into either front nasal sounds or back nasal sounds;
[0142] (c) Unify the pinyin starting with L and N corresponding to all Chinese characters to start with either L or N;
[0143] (d) Retain English characters without any processing;
[0144] (e) First convert numbers into Chinese characters, and then convert the corresponding Chinese characters into pinyin.
[0145] Example 3:
[0146] An embodiment of the present invention provides a storage medium that stores a computer program for address information matching based on a similarity algorithm. Among them, the computer program enables a computer to execute the address information matching method as described in Example 1.
[0147] Example 4:
[0148] An embodiment of the present invention provides an electronic device, including:
[0149] One or more processors; a memory; and one or more programs, where the one or more programs are stored in the memory and are configured to be executed by the one or more processors. The program includes a method for executing address information matching as described in Example 1.
[0150] It can be understood that the system, storage medium, and electronic device for address information matching based on the similarity algorithm provided by the embodiments of the present invention correspond to the method for address information matching based on the similarity algorithm provided by the embodiments of the present invention. For the explanations, examples, beneficial effects, and other parts of the relevant content, reference can be made to the corresponding parts in the method for address information matching, which will not be elaborated here.
[0151] In summary, compared with the prior art, the following beneficial effects are achieved:
[0152] 1. In the embodiments of the present invention, by introducing a set of cells to be matched and based on a similarity algorithm, the address filled in by the customer can be matched to the cell level.
[0153] 2. In the embodiments of the present invention, the cell corresponding to the first obtained matching score of 1 is used as the standardized result, and other cells to be matched are no longer traversed, greatly improving the matching efficiency.
[0154] 3. In the embodiments of the present invention, based on a similarity algorithm composed of character matching degree and pinyin matching degree, the finally calculated matching score includes factors such as the ratio of the number of matching characters and pinyin, the relative position of the match, and the continuity of the match, greatly improving the matching accuracy in practical applications.
[0155] 4. In the embodiments of the present invention, on the basis of character correspondence, the position similarity score also fully considers the position distance relationship between common characters, so that the filtered address string or the cell string to be matched can be accurately recognized even when there are several identical characters in itself, thereby avoiding taking a cell that only has character correspondence but different character positions compared with the address as a matching result.
[0156] 5. In the embodiments of the present invention, the pinyin sequence obtained based on the conversion rule at least eliminates the adverse effects caused by homophones, typos and other defects in the corresponding string, and greatly improves the matching accuracy.
[0157] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.
[0158] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for matching address information based on a similarity algorithm, characterized in that: include: Get the address filled in by the customer and the set of communities to be matched; Using regular expressions to filter non-numeric Chinese and English characters in the address, and arranging all the cells to be matched in the cell set in non-ascending order according to the length of the character string; Based on a preset similarity algorithm, the matching scores between the arranged cells to be matched and the filtered addresses are calculated one by one. If the matching score obtained for the first time in the calculation process is 1, the corresponding cell to be matched is used as the standardized result of the address; Otherwise, after the calculation and traversal are completed, the to-be-matched cell corresponding to the highest matching score obtained for the first time is used as the standardized result of the address; The filtered address is defined as a first character string, and any cell to be matched is defined as a second character string; The method of calculating the matching scores between the arranged cells to be matched and the filtered addresses one by one based on a preset similarity algorithm includes: Counting the number of common characters between the first character string and the second character string, and calculating a character ratio score in combination with the number of characters in the second character string; calculating a position similarity score based on the position distance relationship between all common characters in the first character string and the second character string; and performing a weighted calculation based on the character ratio score and the position similarity score to obtain a character similarity score; Based on a preset conversion rule, convert the first character string into a first pinyin sequence, and convert the second character string into a second pinyin sequence; Counting the number of common pinyins between the first pinyin sequence and the second pinyin sequence, and combining the number of pinyins in the second pinyin sequence to calculate a pinyin ratio score; counting the longest common subsequence between the first pinyin sequence and the second pinyin sequence, and combining the number of pinyins in the first pinyin sequence to calculate a common sequence ratio score; and performing weighted calculation based on the pinyin ratio score and the common sequence ratio score to obtain a pinyin similarity score; Based on the character similarity score and the pinyin similarity score, a final matching score is obtained by weighting.
2. The method for matching address information according to claim 1, characterized in that: The position subscripts of all common characters in the first string S1 are defined as the first position subscript set x1, and the position subscripts of all common characters in the second string S2 are defined as the second position subscript set x2; the distances between all adjacent position subscripts in the first position subscript set x1 are defined as the first distance set n1, and the distances between all adjacent position subscripts in the second position subscript set x2 are defined as the second distance set n2; the position similarity score is expressed as: Where sim(S1, S2) represents the position similarity score between S1 and S2; len(x1) represents the number of position subscripts in x1; k1 and k2 are the corresponding control coefficients; abs represents the absolute value function; n 1,i 、n 2,i Represent the i-th distance in n1 and n2 respectively.
3. The method for matching address information according to claim 1, characterized in that: The conversion rules include: (a) One Chinese character is converted into one pinyin; (b) unify the pinyin corresponding to all Chinese characters into front nasal or back nasal; (c) All Chinese characters beginning with L and N are unified into one with L or N; (d) English characters are retained without any processing; (e) Convert the numbers into Chinese characters first, and then convert the corresponding Chinese characters into pinyin.
4. A system for matching address information based on a similarity algorithm, characterized in that: include: The information acquisition module is used to obtain the address filled in by the customer and the set of communities to be matched; An information processing module, used to filter non-numeric Chinese and English characters in the address using regular expressions, and to arrange all the cells to be matched in the cell set in non-ascending order according to the length of the character string; An information matching module, used to calculate the matching scores between the arranged cells to be matched and the filtered addresses one by one based on a preset similarity algorithm, and if the matching score obtained for the first time in the calculation process is 1, the corresponding cell to be matched is used as the standardized result of the address; Otherwise, after the calculation and traversal are completed, the to-be-matched cell corresponding to the highest matching score obtained for the first time is used as the standardized result of the address; The filtered address is defined as a first character string, and any cell to be matched is defined as a second character string; The method of calculating the matching scores between the arranged cells to be matched and the filtered addresses one by one based on a preset similarity algorithm includes: Counting the number of common characters between the first character string and the second character string, and calculating a character ratio score in combination with the number of characters in the second character string; calculating a position similarity score based on the position distance relationship between all common characters in the first character string and the second character string; and performing a weighted calculation based on the character ratio score and the position similarity score to obtain a character similarity score; Based on a preset conversion rule, convert the first character string into a first pinyin sequence, and convert the second character string into a second pinyin sequence; Counting the number of common pinyins between the first pinyin sequence and the second pinyin sequence, and combining the number of pinyins in the second pinyin sequence to calculate a pinyin ratio score; counting the longest common subsequence between the first pinyin sequence and the second pinyin sequence, and combining the number of pinyins in the first pinyin sequence to calculate a common sequence ratio score; and performing weighted calculation based on the pinyin ratio score and the common sequence ratio score to obtain a pinyin similarity score; Based on the character similarity score and the pinyin similarity score, a final matching score is obtained by weighting.
5. The address information matching system according to claim 4, characterized in that: The position subscripts of all common characters in the first string S1 are defined as the first position subscript set x1, and the position subscripts of all common characters in the second string S2 are defined as the second position subscript set x2; the distances between all adjacent position subscripts in the first position subscript set x1 are defined as the first distance set n1, and the distances between all adjacent position subscripts in the second position subscript set x2 are defined as the second distance set n2; the position similarity score is expressed as: Where sim(S1, S2) represents the position similarity score between S1 and S2; len(x1) represents the number of position subscripts in x1; k1 and k2 are the corresponding control coefficients; abs represents the absolute value function; n 1,i 、n 2,i Represent the i-th distance in n1 and n2 respectively.
6. The address information matching system according to claim 4, characterized in that: The conversion rules include: (a) One Chinese character is converted into one pinyin; (b) unify the pinyin corresponding to all Chinese characters into front nasal or back nasal; (c) All Chinese characters beginning with L and N are unified into one with L or N; (d) English characters are retained without any processing; (e) Convert the numbers into Chinese characters first, and then convert the corresponding Chinese characters into pinyin.
7. A storage medium, characterized in that: The computer program for address information matching based on a similarity algorithm is stored therein, wherein the computer program enables a computer to execute the address information matching method according to any one of claims 1 to 3.
8. An electronic device, characterized in that: include: one or more processors; Memory; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the programs comprising a method for executing the address information matching method according to any one of claims 1 to 3.
Citation Information
Patent Citations
Address information processing method and device, mobile terminal and storage medium
CN116501834A