A method, device, equipment and storage medium for list scanning

By filtering the list string using list index to obtain candidate strings and only perform similarity calculations on the candidate strings, the problem of low list scanning efficiency in the prior art is solved, and the saving of computing resources and the improvement of comparison efficiency is achieved.

CN114299505BActive Publication Date: 2025-06-20SHANGHAI PUDONG DEVELOPMENT BANK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111588243.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-23
Publication Date
2025-06-20
Estimated Expiration
2041-12-23

AI Technical Summary

Technical Problem

In the prior art, list scanning efficiency is inefficient, resulting in waste of computing resources, because 99.9% of the comparisons are comparing two strings that are completely unavailable to reach the preset similarity threshold.

Method used

By obtaining the string to be scanned and the list index, the list string is filtered to obtain the candidate string associated with the string to be scanned, and only the candidate string is similarly calculated.

Benefits of technology

Improve the efficiency of list scanning, save computing resources, and reduce unnecessary comparisons.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114299505B_ABST
    Figure CN114299505B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention discloses a method, device, equipment and storage medium for list scanning. The method includes: when a list scanning event is detected, obtaining a to-be-scanned string corresponding to the list scanning event and a list index of the to-be-scanned list, where the list index is an index pre-constructed according to the list string features of each list string in the to-be-scanned list; filtering each list string according to the to-be-scanned string features of the to-be-scanned string and the list index to obtain candidate strings related to the to-be-scanned string; for each candidate string, calculating the similarity between the candidate string and the to-be-scanned string, and determining whether to use the candidate string as a target string according to the similarity. The technical solution of the embodiment of the present invention first filters each list string in the to-be-scanned list through the list index, and then calculates the similarity between the to-be-scanned string and the candidate strings obtained after filtering, saving computing resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the technical field of information processing, and in particular, to a list scanning method, apparatus, device, and storage medium. Background Art

[0002] List scanning is applied in many application scenarios. It is to compare the string to be scanned with each list string in the list to be scanned one by one (that is, calculate the similarity between the two), so as to match the target string whose similarity with the string to be scanned reaches the preset similarity threshold from each list string.

[0003] In the process of implementing the present invention, the inventors found that the following technical problems exist in the prior art: 99.9% of the comparisons are between two strings that are completely impossible to reach the preset similarity threshold, which causes a great waste of computing resources and low list scanning efficiency. Summary of the Invention

[0004] The embodiments of the present invention provide a list scanning method, apparatus, device, and storage medium, which solve the problem of low list scanning efficiency caused by a great waste of computing resources.

[0005] In a first aspect, the embodiments of the present invention provide a list scanning method, which may include:

[0006] When a list scanning event is detected, obtain the string to be scanned corresponding to the list scanning event and the list index of the list to be scanned, where the list index is an index pre-constructed according to the list string features of each list string in the list to be scanned;

[0007] Filter each list string according to the string features of the string to be scanned and the list index, and obtain candidate strings related to the string to be scanned;

[0008] For each candidate string, calculate the similarity between the candidate string and the string to be scanned, and determine whether to use the candidate string as the target string according to the similarity.

[0009] In a second aspect, the embodiments of the present invention further provide a list scanning apparatus, which may include:

[0010] A list index obtaining module, configured to obtain the string to be scanned corresponding to the list scanning event and the list index of the list to be scanned when a list scanning event is detected, where the list index is an index pre-constructed according to the list string features of each list string in the list to be scanned;

[0011] A candidate string obtaining module, configured to filter each list string according to the characteristics of the string to be scanned and the list index of the list to be scanned, so as to obtain candidate strings related to the string to be scanned;

[0012] A target string determining module, configured to calculate the similarity between a candidate string and the string to be scanned for each candidate string, and determine whether to use the candidate string as the target string according to the similarity.

[0013] In a third aspect, an embodiment of the present invention further provides a list scanning device, which may include:

[0014] One or more processors;

[0015] A memory, configured to store one or more programs;

[0016] When the one or more programs are executed by the one or more processors, the one or more processors implement the list scanning method provided in any embodiment of the present invention.

[0017] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the list scanning method provided in any embodiment of the present invention is implemented.

[0018] The technical solution of the embodiment of the present invention obtains the string to be scanned corresponding to the detected list scanning event and the list index of the list to be scanned, and the list index may be an index pre-constructed according to the characteristics of each list string in the list to be scanned; since the similarity between many list strings and the string to be scanned cannot reach the preset similarity threshold at all, in order to solve the problem of waste of computing resources caused by calculating the similarity of these list strings, each list string can be filtered according to the characteristics of the string to be scanned and the list index, so as to obtain candidate strings related to the string to be scanned (that is, the similarity between the candidate strings and the string to be scanned may reach the preset similarity threshold), so that only the similarity between the candidate strings and the string to be scanned needs to be calculated subsequently; furthermore, for each candidate string, calculate the similarity between the candidate string and the string to be scanned, and determine whether to use the candidate string as the target string according to the similarity. The above technical solution filters each list string in the list to be scanned through the list index, so as to obtain candidate strings whose similarity with the string to be scanned may reach the preset similarity threshold. In this way, only the similarity between these candidate strings and the string to be scanned needs to be calculated subsequently, thereby improving the scanning efficiency of the list to be scanned by saving computing resources. Description of the Drawings

[0019] Figure 1 It is a flowchart of a list scanning method in Embodiment 1 of the present invention;

[0020] Figure 2 It is a flowchart of a list scanning method in Embodiment 2 of the present invention;

[0021] Figure 3 It is a schematic structural diagram of a list index in a list scanning method in Embodiment 2 of the present invention;

[0022] Figure 4 It is a schematic diagram of the characteristics of a string to be scanned in a list scanning method in Embodiment 2 of the present invention;

[0023] Figure 5 It is a schematic diagram of filtering each list string based on a list index in a list scanning method in Embodiment 2 of the present invention;

[0024] Figure 6 It is a flowchart of a list scanning method in Embodiment 3 of the present invention;

[0025] Figure 7 It is a flowchart of a list scanning method in Embodiment 4 of the present invention;

[0026] Figure 8 It is a schematic structural diagram of a list index in a list scanning method in Embodiment 4 of the present invention;

[0027] Figure 9 It is a schematic diagram of the update of a list index in a list scanning method in Embodiment 4 of the present invention;

[0028] Figure 10 It is a schematic diagram of filtering each list string based on a list index in a list scanning method in Embodiment 4 of the present invention;

[0029] Figure 11 It is a structural block diagram of a list scanning device in Embodiment 5 of the present invention;

[0030] Figure 12 It is a schematic structural diagram of a list scanning device in Embodiment 6 of the present invention. Detailed implementation manners

[0031] The present invention will be further described in detail below with reference to the drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the present invention, rather than limiting the present invention. In addition, it should be noted that, for the convenience of description, only parts related to the present invention rather than all structures are shown in the drawings.

[0032] Embodiment 1

[0033] Figure 1It is a flowchart of a list scanning method provided in Embodiment 1 of the present invention. This embodiment is applicable to the situation of efficiently scanning a list, especially applicable to the situation of improving the scanning efficiency of the list to be scanned by filtering each list string in the list to be scanned. This method can be executed by the list scanning device provided in the embodiments of the present invention. The device can be implemented in software and / or hardware, and can be integrated on a list scanning device, which can be various user terminals or servers.

[0034] See Figure 1 , the method of the embodiment of the present invention specifically includes the following steps:

[0035] S110. When a list scanning event is detected, obtain the string to be scanned corresponding to the list scanning event and the list index of the list to be scanned, where the list index is an index pre-constructed according to the list string features of each list string in the list to be scanned.

[0036] Among them, the list scanning event may be an event for scanning each list string in the list to be scanned to match a target string whose similarity with the string to be scanned reaches a preset similarity threshold; the list index may be an index of the list to be scanned pre-constructed according to the list string features of each list string, and the list string feature is a feature of the list string, which may be the list string length, the number of list words, the list character frequency fingerprint, the list hash signature, the list word segmentation, the global identifier, etc., which are not specifically limited here.

[0037] S120. Filter each list string according to the list index and the string features of the string to be scanned to obtain candidate strings related to the string to be scanned.

[0038] Among them, the string features of the string to be scanned are the features of the string to be scanned, which may be the length of the string to be scanned, the number of words to be scanned, the character frequency fingerprint to be scanned, the maximum tolerance distance, the word segmentation to be scanned, etc., which are not specifically limited here. It should be noted that taking the list string length and the string length of the string to be scanned as an example, the essence of both is the string length, but the former is the string length of the list string and the latter is the string length of the string to be scanned. Therefore, here it is only for distinguishing whose string length it is and different names are used, rather than a specific limitation on its essential content. The situations of the remaining list string features and the string features of the string to be scanned are similar and will not be elaborated here.

[0039] Since the list index is constructed based on the characteristics of each list string, based on the list index and the characteristics of the string to be scanned, each list string can be filtered to filter out the list strings that are not relevant to the string to be scanned from these list strings, that is, filter out the list strings that have no possibility of reaching the preset similarity threshold with the string to be scanned. Thus, the list strings that are retained after filtering and are relevant to the string to be scanned (that is, the similarity between the list strings and the string to be scanned may reach the preset similarity threshold) are obtained, and such list strings are used as candidate strings.

[0040] S130. For each candidate string, calculate the similarity between the candidate string and the string to be scanned, and determine whether to use the candidate string as the target string according to the similarity.

[0041] Among them, the following steps are performed for each candidate string: calculate the similarity between the candidate string and the string to be scanned, such as calculating the similarity between the two based on algorithms such as edit distance algorithm, partial alignment algorithm, and pronunciation alignment algorithm; then, determine whether to use the candidate string as the target string according to the similarity. For example, when the similarity is greater than or equal to the preset similarity threshold, the candidate string can be used as the target string, and furthermore, the target string can be output to give a warning effect.

[0042] The technical solution of the embodiment of the present invention is to obtain the string to be scanned corresponding to the detected list scanning event and the list index of the list to be scanned. The list index can be an index pre-constructed according to the list string characteristics of each list string in the list to be scanned. Since the similarity between many list strings and the string to be scanned is completely impossible to reach the preset similarity threshold, in order to solve the problem of waste of computing resources caused by calculating the similarity of these list strings, each list string can be filtered according to the characteristics of the string to be scanned and the list index, so as to obtain candidate strings that are relevant to the string to be scanned (that is, the similarity between the candidate strings and the string to be scanned may reach the preset similarity threshold), so that only the similarity between the candidate strings and the string to be scanned needs to be calculated subsequently. Furthermore, for each candidate string, calculate the similarity between the candidate string and the string to be scanned, and determine whether to use the candidate string as the target string according to the similarity. The above technical solution filters each list string in the list to be scanned through the list index, so as to obtain candidate strings whose similarity with the string to be scanned may reach the preset similarity threshold. In this way, only the similarity between these candidate strings and the string to be scanned needs to be calculated subsequently, thereby improving the scanning efficiency of the list to be scanned by saving computing resources.

[0043] Embodiment 2

[0044] Figure 2 It is a flowchart of a list scanning method provided in the second embodiment of the present invention. This embodiment is optimized based on the above technical solutions. In this embodiment, optionally, the list string features may include the list string length, the number of list words, and the list character frequency fingerprint. The number of list words is the number of words included in the list string, and the list character frequency fingerprint is determined according to whether each English letter appears in the list string. The list index is pre-constructed through the following steps: constructing the first dimension in the list index according to each list string length; for each list string length in the first dimension, constructing the second dimension in the list index according to the number of list words of the list string with the list string length; for each number of list words in the second dimension, constructing the third dimension in the list index according to the list character frequency fingerprint of the list string with the number of list words and the list string length at the upper level of the number of list words; obtaining the list index according to the first dimension, the second dimension, and the third dimension. Among them, the explanations of the same or corresponding terms as those in the above embodiments will not be repeated here.

[0045] See Figure 2 , the method of this embodiment may specifically include the following steps:

[0046] S210. Obtain the list string features of each list string in the list to be scanned, where the list string features include the list string length, the number of list words, and the list character frequency fingerprint. The number of list words is the number of words included in the list string, and the list character frequency fingerprint is determined according to whether each English letter appears in the list string.

[0047] Among them, obtain the list string features of each list string in the list to be scanned respectively. The list string features may include the list string length, the number of list words, and the list character frequency fingerprint. Specifically, the list string length may be the number of English letters and various symbols included in the list string; the number of list words may be the number of words included in the list string; the list character frequency fingerprint may be a fingerprint determined according to whether each English letter appears in the list string. For example, it may be composed of 26-bit binary digits. Each bit in the binary digits corresponds to one of the 26 English letters respectively. When a certain English letter appears in the list string, the binary digit at the position corresponding to the English letter may be set to 1, otherwise it is set to 0. Exemplarily, for the list string "Michael Jackson", its list string length is 15 (i.e., M, i, c, h, a, e, l, space symbol, J, a, c, k, s, o, and n), the number of list words is 2 (i.e., Michael and Jackson), and the list character frequency fingerprint is 10101001111111100010000000.

[0048] S220. Construct the first dimension in the list index according to the lengths of each list string.

[0049] Among them, the list index includes a first dimension, a second dimension, and a third dimension. The second dimension is the next-level dimension of the first dimension, and the third dimension is the next-level dimension of the second dimension. The superior-subordinate relationship between the dimensions can represent the order of application of each dimension when filtering using the list index.

[0050] Construct the first dimension according to the lengths of each list string. Specifically, obtain the length of each list string, determine how many different list string lengths there are among these lengths, and then construct the first dimension based on these different list string lengths. Exemplarily, assume that the lengths of each list string include 13, 15, 15, and 18. Among them, there are three different list string lengths: 13, 15, and 18. Therefore, the first dimension can be constructed based on 13, 15, and 18.

[0051] S230. For each list string length under the first dimension, construct the second dimension in the list index according to the number of list words in the list strings with that list string length.

[0052] Among them, taking a certain list string length under the first dimension as an example, determine the list strings with that list string length from each list string, and then construct the second dimension under that list string length according to the number of list words in these list strings. Thus, when each list string length under the first dimension has been processed, the second dimension is constructed. Exemplarily, assume that the list strings with a certain list string length include A, B, and C. Among them, the number of list words in A is 2, the number of list words in B is 3, and the number of list words in C is 2. Among them, there are two different numbers of list words: 2 and 3. Therefore, the second dimension under that list string length can be constructed based on 2 and 3.

[0053] S240. For each number of list words under the second dimension, construct the third dimension in the list index according to the list word frequency fingerprints of the list strings with that number of list words and the list string length at the level above the number of list words.

[0054] Among them, taking the number of words in a certain list in the second dimension as an example, determine the list strings with this number of list words and the length of the list string at the level above this number of list words from each list string, and then construct the third dimension under this number of list words based on the list word frequency fingerprints of these list strings. Thus, when each number of list words in the second dimension is processed, the third dimension is constructed. Exemplarily, taking the number of words in a certain list under a certain list string length as an example, the list strings corresponding to the number of words in a certain list under this certain list string length include D, E, and F. Among them, the list word frequency fingerprints of D and E are both 10101001111111100010000000, and the list word frequency fingerprint of F is 10101001111111100010000001, that is, there are these two list word frequency fingerprints of 10101001111111100010000000 and 10101001111111100010000001. Therefore, the third dimension under the number of words in a certain list under this certain list string length can be constructed based on these two list word frequency fingerprints.

[0055] S250. Obtain a list index according to the first dimension, the second dimension, and the third dimension.

[0056] Among them, obtain the list index according to the construction results of the first dimension, the second dimension, and the third dimension. It should be noted that the reason for constructing the list index with the list string length as the first dimension, the number of list words as the second dimension, and the list word frequency fingerprint as the third dimension is that, since the number of lengths of the list string length is relatively small, when filtering with it as the first dimension, the number of comparisons to be performed is the number of lengths. Thus, at the beginning, more list strings can be filtered out with fewer comparisons; correspondingly, since the number of fingerprints of the list word frequency fingerprint is relatively large, when filtering with it as the first dimension, the number of comparisons to be performed is the number of fingerprints, that is, at the beginning, more list strings cannot be filtered out with fewer comparisons, which cannot save computing resources to the greatest extent.

[0057] S260. When a list scanning event is detected, obtain the string to be scanned corresponding to the list scanning event, and filter each list string according to the list index and the characteristics of the string to be scanned of the string to be scanned, so as to obtain candidate strings related to the string to be scanned.

[0058] S270. For each candidate string, calculate the similarity between the candidate string and the string to be scanned, and determine whether to use the candidate string as the target string according to the similarity.

[0059] Among them, in practical applications, optionally, the multi-dimensional list index constructed based on the above steps can be applied to the edit distance algorithm. Therefore, this step can calculate the similarity based on the edit distance algorithm. Of course, the similarity can also be calculated based on other algorithms, which is not specifically limited herein.

[0060] The technical solution of the embodiment of the present invention constructs a list index with the own characteristics of the list string (i.e., the list string feature) as the dimension. Subsequently, when filtering each list string based on such a list index, the effective filtering effect of the list string is achieved.

[0061] An optional technical solution is that the list string feature may further include a list hash signature. The above list scanning method may further include: for each list word frequency fingerprint in the third dimension, constructing a fourth dimension in the list index according to the list hash signature of the dimension string; where the dimension string includes the list string having the list word frequency fingerprint, the number of list words at the upper level of the list word frequency fingerprint, and the length of the list string at the upper level of the number of list words at the upper level of the list word frequency fingerprint; obtaining the list index according to the first dimension, the second dimension, and the third dimension, including: obtaining the list index according to the first dimension, the second dimension, the third dimension, and the fourth dimension.

[0062] Among them, in order to filter out more list strings irrelevant to the string to be scanned and further save computing resources, a fourth dimension can be continuously constructed under the third dimension. Specifically, taking a certain list word frequency fingerprint under the third dimension as an example, determine the dimension string from each list string, and construct the fourth dimension under this list word frequency fingerprint based on the list hash signature of these dimension strings, where the dimension string can be the list string having this list word frequency fingerprint, the number of list words at the upper level of this list word frequency fingerprint, and the length of the list string at the upper level of the number of list words at the upper level of this list word frequency fingerprint. In this way, when each number of list words under the third dimension is processed, the fourth dimension is constructed. Through experimental verification, the fourth dimension is very effective in filtering long texts. Exemplarily, taking a certain word frequency fingerprint feature under a certain number of list words under a certain list string length as an example, the list strings corresponding to this certain word frequency fingerprint feature under a certain number of list words under a certain list string length include G and H. The list hash signature of G is [101011], and the list hash signature of H is [110011], that is, there are these two list hash signatures of [101011] and [110011]. Therefore, the fourth dimension under this certain word frequency fingerprint feature under a certain number of list words under a certain list string length can be constructed based on these two list hash signatures.

[0063] To better understand the construction result of the above list index, an exemplary illustration is provided below in combination with specific examples. Exemplarily, refer to Figure 3 the structural schematic diagram of the blacklist index shown in FIG. Among them, the first dimension is the length of the list string. Taking the length of the list string as N as an example, the second dimension - the number of list words is introduced. Then, taking the number of list words as 2 as an example, the third dimension - the list word frequency fingerprint is introduced. Continuing with the list word frequency fingerprint 10101101101110011011101111 as an example, the fourth dimension - the list hash signature is introduced. For any list hash signature, there is one or more list strings corresponding to it below. The list string features of these list strings can be a series of features connected to the list hash signature in the list index.

[0064] On this basis, optionally, the above list scanning method may further include: for each list string, segment the list string, and respectively determine the hash value and weight value of each resulting list segment; for each list segment, obtain the weighted result of the list segment according to the hash value and weight value of the list segment; merge the weighted results of each list segment, and perform dimensionality reduction on the obtained merged result to obtain the list hash signature of the list string. In other words, the generation process of the list hash signature (i.e., the SimHash signature) can be divided into the following 5 steps: segmentation, hashing, weighting, merging, and dimensionality reduction. Specifically,

[0065] The first step, segmentation: Segment the list string to obtain valid list segments, and determine the weight value of each list segment. This weight value can represent the importance of the list segment in the list string. In practical applications, optionally, if the list string is a text, the weight value of the list segment can also be determined according to the number of times the list segment appears in the text. Exemplarily, for the list string This money was used for arms, the list segments can include This, money, was, used, for, and arms, and their weight values can be This(1)money(2)was(1)used(1)for(1)arms(5). The larger the number, the more important the list segment.

[0066] The second step, hashing: Calculate the hash value of each list segment through a hash function, where the hash value is an n-bit signature composed of binary numbers 0 and 1, and n is the number of list words. Exemplarily, the hash value of This, Hash(This), is 100101, and the hash value of arms, Hash(arms), is 101011.

[0067] Step 3, Weighting: Based on the hash value, perform weighting on each word segmentation of the list, that is, W = Hash(hash value) * weight(weight value). When encountering 1, the hash value and the weight value are multiplied positively, and when encountering 0, the hash value and the weight value are multiplied negatively. Exemplarily, weighting the hash value 100101 of "money" gives: W(money) = 100101 × 2 = 2 -2 -2 2 -2 2, and weighting the hash value 101011 of "arms" gives: W(arms) = 101011 × 5 = 5 -5 5 -5 5 5. The word segmentations of the remaining lists are operated similarly.

[0068] Step 4, Merging: Accumulate the weighted results of each word segmentation of the above lists to obtain a sequence string. Exemplarily, taking "money" and "arms" as an example, accumulate 2 -2 -2 2 -2 2 of "money" and 5 -5 5 -5 5 5 of "arms", that is, 2 + 5 -2 + -5 -2 + 5 2 + -5 -2 + 5 2 + 5, thus obtaining 7 -7 3 -3 7.

[0069] Step 5, Dimensionality Reduction: For the accumulated result (i.e., the merging result) of the n-bit signature, if it is greater than 0, set it to 1, otherwise set it to 0, thereby obtaining the list hash signature of the list string. Exemplarily, assuming the final merging result is 9 -9 1 -1 19, then the list hash signature is [101011].

[0070] Another alternative technical solution is that the characteristics of the string to be scanned may include the length of the string to be scanned, the number of words to be scanned, the word frequency fingerprint to be scanned, and the maximum tolerance distance. The number of words to be scanned is the number of words contained in the string to be scanned. The word frequency fingerprint to be scanned is determined according to whether each English letter appears in the string to be scanned. The maximum tolerance distance includes the distance determined according to the length of the string to be scanned and a preset similarity threshold; filtering each list string according to the characteristics of the string to be scanned and the list index to obtain candidate strings related to the string to be scanned may include: for the lengths of each list string in the first dimension of the list index, determining the candidate string length from the lengths of each list string in the first dimension based on the maximum tolerance distance and the length of the string to be scanned; for the number of list words associated with the candidate string length in the second dimension of the list index, determining the candidate word number from the number of list words associated with the candidate string length based on the maximum tolerance distance and the number of words to be scanned; for the list word frequency fingerprint associated with the candidate word number in the third dimension of the list index, determining the candidate word frequency fingerprint from the list word frequency fingerprints associated with the candidate word number based on the maximum tolerance distance and the word frequency fingerprint to be scanned; taking the list string with the candidate string length, candidate word number, and candidate word frequency fingerprint as the candidate string related to the string to be scanned.

[0071] Among them, the maximum tolerance distance can be a distance determined according to the length of the string to be scanned and a preset similarity threshold. Exemplarily, represents rounding up. When calculating the similarity between two strings based on the edit distance algorithm, it can be calculated based on the following formula: Similarity = 100% - (minimum edit distance ÷ length of the string to be scanned), where the minimum edit distance can be the edit distance calculated based on the edit distance algorithm. It should be noted that since the maximum tolerance distance can represent the maximum number of letters that can be different between two strings, when at least one of the string lengths, the number of words, and the similarity between two character frequency fingerprints (which can be represented by the Hamming distance between two character frequency fingerprints) of the two strings differs by 1, the number of different letters between the two strings must be greater than or equal to 1. Of course, the similarity between the hash signatures of the two strings (which can be represented by the Hamming distance between two hash signatures) can also be considered here, and no specific limitation is made here. In other words, if the string lengths, the number of words, and the Hamming distance between two character frequency fingerprints (the Hamming distance between two hash signatures can also be added here) of the two strings are all greater than the maximum tolerance distance, the calculated similarity will necessarily be less than the preset similarity threshold. Therefore, based on the maximum tolerance distance and the characteristics of the string to be scanned of the string to be scanned, candidate strings can be matched from the list index based on the above steps. It should be noted that the number of bits with different symbol element values in the corresponding bits between any two codewords in a codeword combination can be called the Hamming distance between these two codewords, which can be expressed as In addition to representing similarity through the Hamming distance, the character frequency fingerprint can also be converted into a vector, and then the similarity can be represented by the angle between the vectors, and so on.

[0072] To better understand the filtering process of the above list index, an exemplary description will be given below in combination with a specific example. Exemplarily, as Figure 4 shown, this is a schematic diagram of the characteristics of the string to be scanned extracted for the string to be scanned Michael Jackson. On this basis, assuming that the calculated maximum tolerance distance is 1, see Figure 5, the absolute value of the difference between the length of the string to be scanned and the length of the list string is less than or equal to the maximum tolerance distance, that is, |length of the string to be scanned - length of the list string| ≤ maximum tolerance distance. Thus, after filtering under the first dimension of the list index, list string lengths 14, 15, and 16 are obtained. Such list string lengths can be called candidate string lengths. Further, the absolute value of the difference between the number of words to be scanned and the number of words in the list is less than or equal to the maximum tolerance distance, that is, |number of words to be scanned - number of words in the list| ≤ maximum tolerance distance. Taking the list string length of 15 as an example, after filtering under the second dimension of the list index, the number of words in the list 1, 2, and 3 are obtained. Such numbers of words in the list can be called candidate word numbers. Further, the Hamming distance between the word frequency fingerprint to be scanned and the word frequency fingerprint of the list is less than or equal to the maximum tolerance distance, that is, d(word frequency fingerprint to be scanned, word frequency fingerprint of the list) ≤ maximum tolerance distance. Taking the number of words in the list being 2 as an example, after filtering under the third dimension of the list index, the word frequency fingerprints of the list 10101001111111100010000000 and … (i.e., the bold part in the box) are obtained. Such word frequency fingerprints of the list can be called candidate word frequency fingerprints. Finally, the Hamming distance between the hash signature to be scanned and the hash signature of the list is less than or equal to the maximum tolerance distance, that is, d(hash signature to be scanned, hash signature of the list) ≤ maximum tolerance distance. Taking the word frequency fingerprint of the list 10101001111111100010000000 as an example, after filtering under the fourth dimension of the list index, the hash signatures of the list [011011] and [101101] are obtained. Such hash signatures of the list can be called candidate hash signatures. Thus, the list string corresponding to the candidate hash signature in the list index can be used as the candidate string. The candidate string obtained through the maximum tolerance distance may have redundant situations, but it will not miss the list strings whose similarity with the string to be scanned may reach the preset similarity threshold. Thus, the efficiency and accuracy of list scanning can be taken into account simultaneously.

[0073] Embodiment III

[0074] Figure 6It is a flowchart of a list scanning method provided in the third embodiment of the present invention. This embodiment is optimized based on the above technical solutions. In this embodiment, optionally, calculating the similarity between a candidate string and a string to be scanned may include: obtaining the maximum tolerance distance of the string to be scanned, where the maximum tolerance distance is a distance determined according to the length of the string to be scanned and a preset similarity threshold; when calculating the minimum edit distance between the candidate string and the string to be scanned, when a newly calculated current edit distance is obtained, if the current edit distance is greater than the maximum tolerance distance, then use the current edit distance as the minimum edit distance and stop the calculation process of the minimum edit distance; determining the similarity between the candidate string and the string to be scanned according to the calculated minimum edit distance and the length of the string to be scanned. Among them, the explanations of the same or corresponding terms as those in the above embodiments are not repeated here.

[0075] See Figure 6 , the method of this embodiment may specifically include the following steps:

[0076] S310. When a list scanning event is detected, obtain the string to be scanned corresponding to the list scanning event and the list index of the list to be scanned, where the list index is an index pre-constructed according to the list string features of each list string in the list to be scanned.

[0077] S320. Filter each list string according to the list index and the string features of the string to be scanned to obtain candidate strings related to the string to be scanned.

[0078] S330. For each candidate string, obtain the maximum tolerance distance of the string to be scanned, where the maximum tolerance distance is a distance determined according to the length of the string to be scanned and a preset similarity threshold.

[0079] S340. When calculating the minimum edit distance between the candidate string and the string to be scanned, when a newly calculated current edit distance is obtained, if the current edit distance is greater than the maximum tolerance distance, then use the current edit distance as the minimum edit distance and stop the calculation process of the minimum edit distance.

[0080] Among them, the minimum edit distance, also known as the minimum Levenshtein distance, can represent the minimum number of edit operations required to transform one string into another between two strings. If the minimum edit distance between the two is larger, it indicates that they are more different. The permitted edit operations include replacing one character with another, inserting a character, and deleting a character. For example, when converting "kitten" to "sitting", the following edit operations need to be performed in sequence: "sitten" (k→s), "sittin" (e→i), and "sitting" (→g). The minimum edit distance can be calculated through the following dynamic programming equation:

[0081]

[0082] Exemplarily, taking str1 = "kitten" and str2 = "sitting" as an example, m = strlen(str1) = 6, n = strlen(str2) = 7, where m is the string length of str1 and n is the string length of str2. Then the derivation process of the minimum edit distance is as follows. Refer to Table 1. First, calculate dp[0][0], dp[0][1]...dp[0][7], then calculate dp[1][0]...dp[1][7]... Finally, the minimum edit distance dp[m][n] = 3 (i.e., the 3 in the lower right corner of Table 1) is calculated.

[0083] Table 1 Derivation process of the minimum edit distance

[0084]

[0085]

[0086] Table 2 Change process of the current edit distance

[0087] # k i t t e n # 0 1 2 3 4 5 6 s 1 1 2 3 4 5 6 i 2 2 1 2 3 4 5 t 3 3 2 1 2 3 4 t 4 4 3 2 1 2 3 i 5 5 4 3 2 2 3 n 6 6 5 4 3 3 2 g 7 7 7 5 4 4 3

[0088] Refer to Table 2, which can show the change process of the current edit distance (i.e., the latest calculated edit distance, the underlined numbers on the diagonal in Table 1. Only the numbers on the diagonal are called edit distances, and the remaining numbers are only intermediate results calculated for calculating the edit distance) during the matrix operation process. Since the current edit distance during the matrix calculation process is cumulative, as can be seen from the above embodiments, when a certain current edit distance is greater than the maximum tolerance distance, the finally calculated similarity must be lower than the preset similarity threshold. Exemplarily, taking the preset similarity threshold as 90% as an example, When the matrix operation reaches dp[5][5], the current edit distance increases to 2, which is already greater than the maximum tolerance distance. Then, there is no need to continue the subsequent matrix operations because the similarity calculated based on the minimum edit distance obtained from the continued operations will surely be less than the preset similarity threshold, and such candidate strings will not be alerted and output.

[0089] Therefore, to further save computing resources, when calculating the minimum edit distance between a candidate string and a string to be scanned, when the newly calculated current edit distance is obtained, if the current edit distance is greater than the maximum tolerance distance, the current edit distance can be used as the minimum edit distance, and the calculation process of the minimum edit distance can be stopped. Of course, when calculating the minimum edit distance, if the current edit distance is always less than or equal to the maximum tolerance distance, the matrix operation can be continuously performed, and the last current edit distance can be output as the minimum edit distance.

[0090] S350. Determine the similarity between the candidate string and the string to be scanned according to the calculated minimum edit distance and the length of the string to be scanned, and determine whether to use the candidate string as the target string according to the similarity.

[0091] The technical solution of the embodiment of the present invention, when the newly calculated current edit distance obtained is greater than the maximum tolerance distance, uses the current edit distance as the minimum edit distance and stops the calculation process of the minimum edit distance. Thus, by timely stopping the calculation process of the minimum edit distance of candidate strings whose similarity with the string to be scanned will surely be less than the preset similarity threshold, computing resources are further saved.

[0092] Embodiment 4

[0093] Figure 7 is a flowchart of a list scanning method provided in Embodiment 4 of the present invention. This embodiment is optimized based on the above technical solutions. In this embodiment, optionally, the list string features include list word segmentation, global identifier, and the number of list words. The global identifier is the unique identifier of the list string in the list to be scanned, and the number of list words is the number of words included in the list string. The list index is pre-constructed through the following steps: construct the key dimension in the list index according to each list word segmentation; for each list word segmentation under the key dimension, construct the value dimension in the list index according to the global identifier and the number of list words of the list string having the list word segmentation; obtain the list index according to the key dimension and the value dimension. The explanations of the same or corresponding terms as those in the above embodiments are not repeated here.

[0094] See Figure 7 , the method of this embodiment may specifically include the following steps:

[0095] S410. Obtain the list string features of each list string in the list to be scanned. The list string features include list word segmentation, global identifier, and the number of list words. The global identifier is the unique identifier of the list string in the list to be scanned, and the number of list words is the number of words included in the list string.

[0096] Among them, the global identifier (UID) is the unique identifier of the list string in the list to be scanned. That is to say, the global identifiers of each list string in the list to be scanned are different from each other.

[0097] S420. Construct the key dimension in the list index according to each list word segmentation.

[0098] Among them, constructing the key dimension in the list index according to each list word segmentation. Specifically, obtain the list word segmentation of each list string, and determine how many types of list word segmentation there are in total among these list word segmentations, and then construct the key dimension based on these types of list word segmentation. Exemplarily, assume that each list word segmentation includes This, money, money, and Michael, among which there are three types of list word segmentation: This, money, and Michael. Therefore, the key dimension can be constructed based on This, money, and Michael.

[0099] S430. For each list word segmentation under the key dimension, construct the value dimension in the list index according to the global identifier and the number of list words of the list string with the list word segmentation.

[0100] Among them, taking a certain list word segmentation under the key dimension as an example, determine the list strings with this list word segmentation from each list string, and then construct the value dimension under this list word segmentation according to the global identifier and the number of list words of these list strings. Thus, when all the list word segmentations under the key dimension are processed, the value dimension is constructed. Exemplarily, assume that the list strings with a certain list word segmentation include I and J, where the UID of I is 001, the number of list words (TokenNo) is 1, and the UID of J is 011, and TokenNo is 3. Therefore, the value dimension under this list word segmentation can be constructed based on 001 and 1, and 011 and 3.

[0101] To better understand the construction result of the above list index, the following is an exemplary description with a specific example. Exemplarily, see Figure 8 , where the key dimension is the list word segmentation, and the UID and TokenNo under their respective value dimensions are derived from each list word segmentation, and each UID corresponds to its own list string.

[0102] S440. Obtain the list index according to the key dimension and the value dimension.

[0103] Among them, according to the construction results of the key dimension and the value dimension, a list index is obtained.

[0104] It should be noted that Embodiment 2 and Embodiment 4 give two construction schemes for the list index. In practical applications, optionally, corresponding list indexes can be constructed based on these two construction schemes respectively, and each list string can be filtered based on the two list indexes respectively, and the two are executed in parallel without interference.

[0105] S450. When a list scanning event is detected, obtain the string to be scanned corresponding to the list scanning event, and filter each list string according to the characteristics of the string to be scanned of the list index and the string to be scanned, so as to obtain candidate strings related to the string to be scanned.

[0106] S460. For each candidate string, calculate the similarity between the candidate string and the string to be scanned, and determine whether to use the candidate string as the target string according to the similarity.

[0107] Among them, in practical applications, optionally, the list index constructed based on the above steps can be applied to some comparison algorithms. Therefore, this step can calculate the similarity based on some comparison algorithms. Of course, the similarity can also be calculated based on other algorithms, which is not specifically limited here.

[0108] The technical solution of the embodiment of the present invention constructs the key dimension in the list index through each list word segmentation, and for each list word segmentation under the key dimension, constructs the value dimension in the list index according to the global identifier and the number of list words of the list string with the list word segmentation. Subsequently, when filtering each list string based on the obtained list index, the effect of effective filtering of the list string is achieved.

[0109] An optional technical solution, the above list scanning method may further include: when an index update event is detected, obtain the new string corresponding to the index update event, and perform word segmentation on the new string to obtain new word segments; for each new word segment, determine whether there is a list word segment in the key dimension of the list index that is the same as the new word segment; if so, use the global identifier and the number of list words of the new string as the value corresponding to the list word segment in the value dimension that is the same as the new word segment; otherwise, add the new word segment to the key dimension, and use the global identifier and the number of list words of the new string as the value corresponding to the new word segment in the value dimension; update the list index according to the result. Among them, the new string can be a list string newly added to the list to be scanned, and the new word segment can be the list word segment of the new string. To better understand the above process of updating the list index, the following will give an exemplary description with a specific example. Suppose there is Figure 8The list index shown currently has a new string "Michael Jackson" added, with its UID = 123. The new word segments are "Michael" and "Jackson", and the TokenNo is 2. See Figure 9 , compare each new word segment one by one with each list word segment in the key dimension of the list index. Since "Michael" already exists among the list word segments in this key dimension, UID:123 TokenNo:2 can be directly added to the value corresponding to "Michael" in the value dimension ( Figure 9 the value indicated by the underscore below); since "Jackson" does not exist among the list word segments in this key dimension, "Jackson" can be added to the key dimension ( Figure 9 the key indicated by the underscore below), and UID:123 TokenNo:2 is added to the value corresponding to "Jackson" ( Figure 9 the value indicated by the underscore below). The above technical solution achieves the effect of effective update of the list index.

[0110] Another alternative technical solution is that the characteristics of the string to be scanned include the word segments to be scanned obtained after segmenting the string to be scanned; according to the characteristics of the string to be scanned and the list index, each list string is filtered to obtain candidate strings related to the string to be scanned, which may include: for each word segment to be scanned, taking the list word segment in the key dimension of the list index that is the same as the word segment to be scanned as a candidate word segment; for each candidate word segment, obtaining each global identifier corresponding to the candidate word segment in the value dimension of the list index, and taking the list string corresponding to each global identifier as a candidate string related to the string to be scanned, thereby achieving the effect of accurate filtering of candidate strings. To better understand the specific implementation process of the above technical solution, the following will give an exemplary description with a specific example. Exemplarily, see Figure 10 , assume the string to be scanned is "Michael Jackson", and the word segments to be scanned obtained after segmenting it include "Michael" and "Jackson". For each word segment to be scanned, traverse it in the key dimension of the list index to obtain the list word segment in this key dimension that is the same as the word segment to be scanned, and take such a list word segment as a candidate word segment. The boxes where the global identifiers of each candidate word segment are located in the value dimension are in Figure 10 bold. Take the list string with this global identifier as a candidate string.

[0111] Another optional technical solution is to calculate the similarity between each candidate string and the string to be scanned, which may include: taking the set containing each global identifier as the global identifier set to obtain the global identifier sets corresponding to the respective candidate word segmentations; for each candidate string, taking the number of occurrences of the global identifier of the candidate string in each global identifier set as the number of common words, where the number of common words represents the number of common words between the string to be scanned and the candidate string; calculating the similarity between the candidate string and the string to be scanned according to the number of common words, the number of words to be scanned of the string to be scanned, and the number of single words of the candidate string. Among them, some comparison algorithms can be applied to the algorithm for calculating the similarity between two strings in the application scenario where strings contain each other. It calculates the similarity between the common parts of the two strings and the two strings respectively, and finally takes the one with the higher score as the final similarity. After filtering each list string based on the list index constructed in the embodiment of the present invention, the similarity can be calculated based on this technical solution. Specifically, for each candidate word segmentation, taking the set containing each global identifier corresponding to the candidate word segmentation under the value dimension of the list index as the global identifier set; then, for each candidate string, determining the number of common words of its global identifier in each global identifier set, and thus calculating the similarity between the candidate string and the string to be scanned according to the number of common words, the number of words to be scanned of the string to be scanned, and the number of single words of the candidate string, such as similarity = max{(number of common words÷number of words to be scanned), (number of common words÷number of words)}*100%, achieving the effect of accurately calculating the similarity between two strings in the application scenario where strings contain each other.

[0112] To better understand the above process of calculating similarity, an exemplary illustration is provided below with specific examples. Exemplarily, taking Figure 10 the filtered candidate strings and candidate word segmentations as examples, the global identifier set corresponding to the candidate word segmentation Michael is {021, 211, 123}, and the global identifier set corresponding to the candidate word segmentation Jackson is {123}. It can be seen that the number of common words between the string to be scanned Michael Jackson and UID: 021 is 1, the number of common words with UID: 211 is 1, and the number of common words with UID: 123 is 2. Therefore, the similarities between Michael Jackson and the candidate strings with UIDs 021, 211, and 123 are respectively: 021: max{(1÷2), (1÷1)}×100% = 100%; 211: max{(1÷2), (1÷3)}×100% = 50%; 123: max{(2÷2), (2÷2)}×100% = 100%.

[0113] Embodiment 5

[0114] Figure 11 The structural block diagram of the list scanning device provided in the fifth embodiment of the present invention. This device is used to execute the list scanning method provided in any of the above embodiments. This device and the list scanning methods of the above embodiments belong to the same inventive concept. For the details not described in detail in the embodiments of the list scanning device, reference can be made to the embodiments of the above list scanning methods. Refer to Figure 11 , this device may specifically include: a list index acquisition module 510, a candidate string obtaining module 520, and a target string determination module 530. Among them,

[0115] The list index acquisition module 510 is used to obtain the string to be scanned corresponding to the list scanning event and the list index of the list to be scanned when detecting the list scanning event. Among them, the list index is an index pre-constructed according to the list string features of each list string in the list to be scanned;

[0116] The candidate string obtaining module 520 is used to filter each list string according to the string features of the string to be scanned and the list index, and obtain candidate strings related to the string to be scanned;

[0117] The target string determination module 530 is used to calculate the similarity between the candidate string and the string to be scanned for each candidate string, and determine whether to use the candidate string as the target string according to the similarity.

[0118] Optionally, the list string features may include the list string length, the number of list words, and the list character frequency fingerprint. The number of list words is the number of words contained in the list string, and the list character frequency fingerprint is determined according to whether each English letter appears in the list string;

[0119] The list index is pre-constructed through the following modules:

[0120] The first dimension construction module is used to construct the first dimension in the list index according to the lengths of each list string;

[0121] The second dimension construction module is used to construct the second dimension in the list index according to the number of list words of the list strings with the list string length for each list string length in the first dimension;

[0122] The third dimension construction module is used to construct the third dimension in the list index according to the list character frequency fingerprints of the list strings with the number of list words and the list string length at the previous level of the number of list words for each number of list words in the second dimension;

[0123] The list index first obtaining module is used to obtain the list index according to the first dimension, the second dimension, and the third dimension.

[0124] On this basis, optionally, the list string feature further includes a list hash signature, and the above list scanning method may further include:

[0125] A fourth dimension construction module, configured to, for each list word frequency fingerprint in the third dimension, construct a fourth dimension in the list index according to the list hash signature of the dimension string, where the dimension string includes the list strings having list word frequency fingerprints, the number of list words at the level above the list word frequency fingerprint, and the length of the list string at the level above the number of list words at the level above the list word frequency fingerprint;

[0126] The first list index obtaining module may specifically be configured to:

[0127] Obtain a list index according to the first dimension, the second dimension, the third dimension, and the fourth dimension.

[0128] On this basis, optionally, the above list scanning device may further include:

[0129] A weight value determining module, configured to, for each list string, perform word segmentation on the list string, and respectively determine the hash value and the weight value of each obtained list word segment;

[0130] A weighted result obtaining module, configured to, for each list word segment, obtain a weighted result of the list word segment according to the hash value and the weight value of the list word segment;

[0131] A list hash signature obtaining module, configured to merge the weighted results of each list word segment, and perform dimensionality reduction on the obtained merged result to obtain the list hash signature of the list string.

[0132] Another optionally, the characteristics of the string to be scanned include the length of the string to be scanned, the number of words to be scanned, the word frequency fingerprint to be scanned, and the maximum tolerance distance. The number of words to be scanned includes the number of words contained in the string to be scanned. The word frequency fingerprint to be scanned is determined according to whether each English letter appears in the string to be scanned. The maximum tolerance distance is a distance determined according to the length of the string to be scanned and a preset similarity threshold;

[0133] The candidate string obtaining module 520 may include:

[0134] A candidate string length determining unit, configured to, for each list string length in the first dimension of the list index, determine a candidate string length from each list string length in the first dimension based on the maximum tolerance distance and the length of the string to be scanned;

[0135] A candidate word quantity determination unit, which is used to determine the candidate word quantity from the list word quantity associated with the candidate string length under the second dimension of the list index, based on the maximum tolerance distance and the number of words to be scanned.

[0136] A candidate character frequency fingerprint determination unit, which is used to determine the candidate character frequency fingerprint from the list character frequency fingerprints associated with the candidate word quantity under the third dimension of the list index, based on the maximum tolerance distance and the character frequency fingerprints to be scanned.

[0137] A candidate string first determination unit, which is used to use the list string with the candidate string length, candidate word quantity, and candidate character frequency fingerprint as the candidate string related to the string to be scanned.

[0138] Optionally, the target string determination module 530 may include:

[0139] A maximum tolerance distance acquisition unit, which is used to acquire the maximum tolerance distance of the string to be scanned. The maximum tolerance distance is a distance determined according to the length of the string to be scanned and the preset similarity threshold.

[0140] A minimum edit distance calculation unit, which is used to calculate the minimum edit distance between the candidate string and the string to be scanned. When the newly calculated current edit distance is obtained, if the current edit distance is greater than the maximum tolerance distance, the current edit distance is used as the minimum edit distance, and the calculation process of the minimum edit distance is stopped.

[0141] A similarity determination unit, which is used to determine the similarity between the candidate string and the string to be scanned according to the calculated minimum edit distance and the length of the string to be scanned.

[0142] Optionally, the list string features include list word segmentation, global identifier, and list word quantity. The global identifier is the unique identifier of the list string in the list to be scanned, and the list word quantity is the number of words included in the list string. The list index is pre-constructed through the following modules:

[0143] A key dimension construction module, which is used to construct the key dimension in the list index according to each list word segmentation.

[0144] A value dimension construction module, which is used to construct the value dimension in the list index for each list word segmentation under the key dimension according to the global identifier and list word quantity of the list string with the list word segmentation.

[0145] A list index second acquisition module, which is used to obtain the list index according to the key dimension and the value dimension.

[0146] On this basis, optionally, the above list scanning device may further include:

[0147] A new word segmentation obtaining module is used to obtain an added string corresponding to an index update event when an index update event is detected, perform word segmentation on the added string, and obtain added word segments;

[0148] A list word segment determination module is used to determine whether there is a list word segment identical to the added word segment in the key dimension of the list index for each added word segment;

[0149] A value first addition module is used to, if so, use the global identifier of the added string and the number of list words as the value corresponding to the list word segment identical to the added word segment in the value dimension;

[0150] A value second addition module is used to, otherwise, add the added word segment to the key dimension, and use the global identifier of the added string and the number of list words as the value corresponding to the added word segment in the value dimension;

[0151] A list index update module is used to update the list index according to the result.

[0152] Optionally, the characteristics of the string to be scanned may include the word segments to be scanned obtained after performing word segmentation on the string to be scanned. The candidate string obtaining module 520 may include:

[0153] A candidate word segment obtaining unit is used to, for each word segment to be scanned, use the list word segment identical to the word segment to be scanned in the key dimension of the list index as the candidate word segment;

[0154] A candidate string second determination unit is used to, for each candidate word segment, obtain each global identifier corresponding to the candidate word segment in the value dimension of the list index, and use the list strings corresponding to the respective global identifiers as the candidate strings related to the string to be scanned.

[0155] On this basis, optionally, the target string determination module 530 may include:

[0156] A global identifier set obtaining unit is used to use the set containing each global identifier as the global identifier set, and obtain the global identifier sets respectively corresponding to the respective candidate word segments;

[0157] A common word number determination unit is used to, for each candidate string, use the number of occurrences of the global identifier of the candidate string in each global identifier set as the common word number, where the common word number represents the number of common words between the string to be scanned and the candidate string;

[0158] A similarity calculation unit is used to calculate the similarity between the candidate string and the string to be scanned according to the common word number, the number of words to be scanned of the string to be scanned, and the number of list words of the candidate string.

[0159] The list scanning device provided in the fifth embodiment of the present invention obtains a string to be scanned corresponding to a detected list scanning event and a list index of the list to be scanned through a list index obtaining module. The list index can be an index pre-constructed according to the list string features of each list string in the list to be scanned. Since the similarity between many list strings and the string to be scanned cannot reach the preset similarity threshold at all, in order to solve the problem of waste of computing resources caused by calculating the similarity of these list strings, a candidate string obtaining module filters each list string according to the string feature of the string to be scanned and the list index, thereby obtaining candidate strings related to the string to be scanned (that is, the similarity between the candidate strings and the string to be scanned may reach the preset similarity threshold), so that only the similarity between the candidate strings and the string to be scanned needs to be calculated subsequently. Furthermore, a target string determining module calculates the similarity between each candidate string and the string to be scanned, and determines whether to use the candidate string as a target string according to the similarity. The above device filters each list string in the list to be scanned through a list index, thereby obtaining candidate strings whose similarity with the string to be scanned may reach the preset similarity threshold. In this way, only the similarity between these candidate strings and the string to be scanned needs to be calculated subsequently, thereby improving the scanning efficiency of the list to be scanned by saving computing resources.

[0160] The target detection device provided in the embodiment of the present invention can execute the target detection method provided in any embodiment of the present invention, and has function modules and beneficial effects corresponding to the execution of the method.

[0161] It should be noted that in the embodiment of the above target detection device, the included units and modules are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of the functional units are only for the convenience of mutual distinction and do not limit the protection scope of the present invention.

[0162] Embodiment Six

[0163] Figure 12 is a schematic structural diagram of a list scanning device provided in the sixth embodiment of the present invention. Refer to Figure 12 , the device includes a memory 610, a processor 620, an input device 630, and an output device 640. The number of processors 620 in the device can be one or more, Figure 12 taking one processor 620 as an example; the memory 610, the processor 620, the input device 630, and the output device 640 in the device can be connected through a bus or other means, Figure 12 taking the connection through the bus 650 as an example.

[0164] The memory 610, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to the list scanning method in the embodiments of the present invention (for example, the list index acquisition module 510, the candidate string obtaining module 520, and the target string determination module 530 in the list scanning device). The processor 620 executes various functional applications and data processing of the device by running the software programs, instructions, and modules stored in the memory 610, that is, implements the above-mentioned list scanning method.

[0165] The memory 610 may mainly include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the device, etc. In addition, the memory 610 may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some instances, the memory 610 may further include a memory remotely set relative to the processor 620, and these remote memories can be connected to the device through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0166] The input device 630 can be used to receive input digital or character information, and generate key signal inputs related to the user settings and function controls of the device. The output device 640 may include a display device such as a display screen.

[0167] Embodiment VII

[0168] Embodiment VII of the present invention provides a storage medium containing computer-executable instructions, and the computer-executable instructions are used to execute a list scanning method when executed by a computer processor. The method includes:

[0169] When a list scanning event is detected, obtain a to-be-scanned string corresponding to the list scanning event and a list index of the to-be-scanned list, where the list index is an index pre-constructed according to the list string features of each list string in the to-be-scanned list;

[0170] Filter each list string according to the to-be-scanned string features of the to-be-scanned string and the list index to obtain candidate strings related to the to-be-scanned string;

[0171] For each candidate string, calculate the similarity between the candidate string and the to-be-scanned string, and determine whether to use the candidate string as the target string according to the similarity.

[0172] Of course, for the storage medium containing computer-executable instructions provided in the embodiments of the present invention, the computer-executable instructions are not limited to the method operations described above, and can also execute the related operations in the list scanning method provided in any embodiment of the present invention.

[0173] From the above description of the embodiments, those skilled in the art can clearly understand that the present invention can be implemented by means of software and the necessary general-purpose hardware. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as a floppy disk, a read-only memory (ROM), a random access memory (RAM), a flash memory (FLASH), a hard disk, or an optical disc of a computer, and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0174] Note that the above are only the preferred embodiments of the present invention and the technical principles applied. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and various obvious changes, re-adjustments, and substitutions can be made by those skilled in the art without departing from the protection scope of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments. Without departing from the concept of the present invention, it can also include more other equivalent embodiments, and the scope of the present invention is determined by the scope of the appended claims.

Claims

1. A method for scanning a list, characterized in that, Including: When a list scanning event is detected, obtain a string to be scanned corresponding to the list scanning event and a list index of the list to be scanned, where the list index is an index pre-constructed according to the list string features of each list string in the list to be scanned; Filter each of the list strings according to the string features of the string to be scanned and the list index to obtain candidate strings related to the string to be scanned; For each of the candidate strings, calculate the similarity between the candidate string and the string to be scanned, and determine whether to use the candidate string as a target string according to the similarity; Wherein, calculating the similarity between the candidate string and the string to be scanned includes: Obtain the maximum tolerance distance of the string to be scanned, where, represents rounding up; When calculating the minimum edit distance between the candidate string and the string to be scanned, when a newly calculated current edit distance is obtained, if the current edit distance is greater than the maximum tolerance distance, use the current edit distance as the minimum edit distance and stop the calculation process of the minimum edit distance; Determine the similarity between the candidate string and the string to be scanned according to the calculated minimum edit distance and the length of the string to be scanned.

2. The method according to claim 1, characterized in that, The list string features include the list string length, the number of list words, and the list character frequency fingerprint. The number of list words is the number of words included in the list string. The list character frequency fingerprint is determined according to whether each English letter appears in the list string. The list index is pre-constructed through the following steps: Construct the first dimension in the list index according to each of the list string lengths; For each of the list string lengths in the first dimension, construct the second dimension in the list index according to the number of list words of the list string having the list string length; For each of the number of list words in the second dimension, construct the third dimension in the list index according to the list character frequency fingerprint of the list string having the number of list words and the list string length at the upper level of the number of list words; Obtain the list index according to the first dimension, the second dimension, and the third dimension.

3. The method according to claim 2, characterized in that, The list string features further include a list hash signature, and the method further includes: For each of the list character frequency fingerprints in the third dimension, construct the fourth dimension in the list index according to the list hash signature of the dimension string; Wherein, the dimension string includes the list strings having the list character frequency fingerprint, the number of list words at the upper level of the list character frequency fingerprint, and the list string length at the upper level of the number of list words at the upper level of the list character frequency fingerprint in each of the list strings; Obtaining the list index according to the first dimension, the second dimension, and the third dimension includes: obtaining the list index according to the first dimension, the second dimension, the third dimension, and the fourth dimension.

4. The method according to claim 3, characterized in that, Also including: For each of the said list strings, perform word segmentation on the list string, and respectively determine the hash value and weight value of each resulting list word segment; For each of the said list word segments, obtain the weighted result of the list word segment according to the hash value and the weight value of the list word segment; Merge the weighted results of the said list word segments, and perform dimensionality reduction on the obtained merged result to obtain the list hash signature of the list string; 5. The method according to claim 2, characterized in that, The characteristics of the string to be scanned include the length of the string to be scanned, the number of words to be scanned, the word frequency fingerprint to be scanned, and the maximum tolerance distance. The number of words to be scanned is the number of words contained in the string to be scanned, and the word frequency fingerprint to be scanned is determined according to whether each English letter appears in the string to be scanned; Filtering each of the said list strings according to the characteristics of the string to be scanned and the said list index to obtain candidate strings related to the string to be scanned includes: For each length of the said list strings in the first dimension of the list index, determine the candidate string length from each length of the said list strings in the first dimension based on the maximum tolerance distance and the length of the string to be scanned; For each number of words in the list associated with the candidate string length in the second dimension of the list index, determine the candidate number of words from each number of words in the list associated with the candidate string length based on the maximum tolerance distance and the number of words to be scanned; For each word frequency fingerprint in the list associated with the candidate number of words in the third dimension of the list index, determine the candidate word frequency fingerprint from each word frequency fingerprint in the list associated with the candidate number of words based on the maximum tolerance distance and the word frequency fingerprint to be scanned; Use the list string having the candidate string length, the candidate number of words, and the candidate word frequency fingerprint as the candidate string related to the string to be scanned; 6. The method according to claim 1, characterized in that, The characteristics of the list string include list word segments, global identifiers, and the number of words in the list. The global identifier is the unique identifier of the list string in the list to be scanned, and the number of words in the list is the number of words contained in the list string. The list index is pre-constructed through the following steps: Construct the key dimension in the list index according to each of the said list word segments; For each of the said list word segments in the key dimension, construct the value dimension in the list index according to the global identifier and the number of words in the list string having the list word segment; Obtain the list index according to the key dimension and the value dimension; 7. The method according to claim 6, characterized in that, It further includes: When detecting an index update event, obtain the new string corresponding to the index update event, perform word segmentation on the new string to obtain new word segments; For each of the said new word segments, determine whether there is the same list word segment as the new word segment in the key dimension of the list index; If so, use the global identifier of the new string and the number of list words as the value corresponding to the list word that is the same as the new word segment under the value dimension; Otherwise, add the new word segment to the key dimension, and use the global identifier of the new string and the number of list words as the value corresponding to the new word segment under the value dimension; Update the list index according to the result.

8. The method according to claim 6, characterized in that, The string feature to be scanned includes the word segments to be scanned obtained by segmenting the string to be scanned; Filter each list string according to the string feature to be scanned of the string to be scanned and the list index to obtain candidate strings related to the string to be scanned, including: For each word segment to be scanned, use the list word that is the same as the word segment to be scanned under the key dimension of the list index as the candidate word segment; For each candidate word segment, obtain each global identifier corresponding to the candidate word segment under the value dimension of the list index, and use the list strings corresponding to the respective global identifiers as candidate strings related to the string to be scanned.

9. The method according to claim 8, characterized in that, For each candidate string, calculate the similarity between the candidate string and the string to be scanned, including: Use the set containing the respective global identifiers as the global identifier set to obtain the global identifier sets corresponding to the respective candidate word segments; For each candidate string, use the number of occurrences of the global identifier of the candidate string in each global identifier set as the number of common words, where the number of common words represents the number of common words between the string to be scanned and the candidate string; Calculate the similarity between the candidate string and the string to be scanned according to the number of common words, the number of words to be scanned of the string to be scanned, and the number of list words of the candidate string.

10. A list scanning device, characterized in that, Include: A list index acquisition module, configured to, when detecting a list scanning event, acquire a string to be scanned corresponding to the list scanning event and a list index of a list to be scanned, where the list index is an index pre-constructed according to the list string features of each list string in the list to be scanned; A candidate string obtaining module, configured to filter each list string according to the string feature to be scanned of the string to be scanned and the list index to obtain candidate strings related to the string to be scanned; A target string determination module, configured to, for each candidate string, calculate the similarity between the candidate string and the string to be scanned, and determine whether to use the candidate string as the target string according to the similarity; The target string determination module includes: The maximum tolerance distance acquisition unit is used to acquire the maximum tolerance distance of the string to be scanned, where represents rounding up; A minimum edit distance calculation unit, configured to, when calculating the minimum edit distance between a candidate string and a string to be scanned, if a newly calculated current edit distance is obtained and the current edit distance is greater than the maximum tolerance distance, use the current edit distance as the minimum edit distance and stop the calculation process of the minimum edit distance; A similarity determination unit, configured to determine the similarity between a candidate string and a string to be scanned according to the calculated minimum edit distance and the length of the string to be scanned.

11. A list scanning device, characterized in that, Comprising: One or more processors; A memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the list scanning method according to any one of claims 1-9.

12. A computer-readable storage medium, on which a computer program is stored, characterized in that, When the computer program is executed by a processor, it implements the list scanning method according to any one of claims 1-9.

Citation Information

Patent Citations

  • Address correction method, device and equipment and storage medium

    CN111008625A