A short text matching method, device and equipment and storage medium

CN117349487BActive Publication Date: 2026-10-09HANGZHOU DBAPPSECURITY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311528792.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-11-16
Publication Date
2026-10-09
Estimated Expiration
2043-11-16

AI Technical Summary

Technical Problem

[0003]然而,通过上述方式进行短文本匹配的过程中,部分关键词在领域知识库中从未出现过,针对这种情况,目前通常是通过编写表达式的方式进行短文本匹配,但随着业务的不断积累,不同人员对表达式的写法和理解各异,导致表达式的数量剧增,如10万条时,而大数据量的表达式在使用过程中,会使短文本匹配效率变得很低,并且相同类别的表达式之间存在冗余,不同类别的表达式存在冲突,另外,存在误报的表达式难以被发现等问题,从而导致了匹配时间过长,尽管业务员在部署时可以通过调整表达式来缓解误报,但效果不佳依然会导致生产环境中的低效率和高误报,进而导致业务延期甚至业务失败

Benefits of technology

[0042]可见,本申请先采集为目标业务数据编写的正则表达式得到正则表达式集合,并分别对所述正则表达式集合中的各个正则表达式进行预处理,得到多个预处理后表达式,然后对所有所述预处理后表达式进行分类,得到多个分类后表达式组,并分别对每个所述分类后表达式组中的正则表达式进行相似度计算,得到每个所述分类后表达式组对应的第一相似度值,接着判断所述第一相似度值是否超过第一阈值,若是则从超过所述第一阈值的所述第一相似度值对应的所述分类后表达式组中确定出任意一个正则表达式,得到目标表达式,并分别删除各所述分类后表达式组中除所述目标表达式外的所有正则表达式,得到只包含所述目标表达式的第一删除后表达式组,再利用优化后的DBSCAN算法对所有所述第一删除后表达式组中的所述目标表达式进行聚类,得到多个聚类后表达式簇;当接收到短文本匹配请求时,利用所述聚类后表达式簇对目标短文本进行匹配得到匹配结果。本申请从影响匹配效率和准确率的根本原因,即正则表达式出发,先对正则表达式集合中的正则表达式进行分类,并删除分类后表达式组中相似度较高的表达式,从而降低了表达式的数量,即去除冗余的表达式,并利用优化后的DBSCAN算法对所有删除后表达式组中的表达式进行聚类得到多个聚类后表达式簇,从而降低了匹配的类别数量,进而提高了短文本匹配的准确率和效率,并降低了误报率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117349487B_ABST
    Figure CN117349487B_ABST
Patent Text Reader

Abstract

The application discloses a short text matching method and device, equipment and storage medium, and relates to the technical field of text classification, which comprises the following steps: preprocessing each regular expression written for target business data, classifying the preprocessed expressions to obtain a plurality of classified expression groups, and calculating the similarity of the regular expressions in each classified expression group to obtain a first similarity value; determining whether the first similarity value exceeds a first threshold value, if yes, determining a target expression from any regular expression in the classified expression group corresponding to the first similarity value exceeding the first threshold value, and deleting all expressions in each classified expression group except the target expression to obtain a first deleted expression group; and clustering the expressions in all first deleted expression groups by using an optimized DBSCAN algorithm to obtain a clustered expression cluster for matching the short text. The application can improve the accuracy and efficiency of short text matching and reduce the false positive rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a short text matching method, apparatus, device and storage medium. Background Technology

[0002] In the context of data security, there is a demand for short text classification, which is widely used in various industries, such as news classification, spam detection, user sentiment classification, and intelligent product recommendation. The process of classifying short text typically involves matching the text against a domain knowledge base. Specifically, the text content within the domain is segmented and stop words are removed to obtain a large number of words. After deduplication, a domain knowledge base is created. After removing stop words, the keywords in the short text to be classified are compared with the keywords in the domain knowledge base, and the comparisons are sorted according to the degree of similarity. Finally, the keyword with the highest similarity is taken as the matching result.

[0003] However, during the short text matching process described above, some keywords have never appeared in the domain knowledge base. To address this, short text matching is currently usually performed by writing expressions. However, as business operations accumulate, different personnel have different ways of writing and understanding expressions, leading to a surge in the number of expressions, such as 100,000. When using expressions with such a large number of data, the efficiency of short text matching becomes very low. Furthermore, there is redundancy between expressions of the same category, conflicts between expressions of different categories, and difficulty in detecting false positives. This results in excessively long matching times. Although business personnel can adjust expressions during deployment to mitigate false positives, the effect is still unsatisfactory, leading to low efficiency and high false positives in the production environment, which in turn can cause business delays or even business failures.

[0004] In summary, how to solve the above problems is a problem that needs to be solved by those skilled in the art. Summary of the Invention

[0005] In view of this, the purpose of this application is to provide a short text matching method, apparatus, device, and storage medium, which can improve the accuracy and efficiency of short text matching and reduce the false alarm rate. The specific solution is as follows:

[0006] Firstly, this application discloses a short text matching method, including:

[0007] Collect regular expressions written for the target business data to obtain a set of regular expressions, and preprocess each regular expression in the set of regular expressions to obtain multiple preprocessed expressions;

[0008] All the preprocessed expressions are classified to obtain multiple groups of classified expressions, and the similarity of the regular expressions in each group of classified expressions is calculated to obtain the first similarity value corresponding to each group of classified expressions.

[0009] Determine whether the first similarity value exceeds the first threshold. If so, determine any regular expression from the group of classified expressions corresponding to the first similarity value that exceeds the first threshold to obtain the target expression. Then, delete all regular expressions in each group of classified expressions except the target expression to obtain a first group of deleted expressions containing only the target expression.

[0010] The optimized DBSCAN algorithm is used to cluster all the target expressions in the first group of deleted expressions to obtain multiple clustered expression clusters.

[0011] When a short text matching request is received, the target short text is matched using the clustered expression cluster to obtain the matching result.

[0012] Optionally, the step of preprocessing each regular expression in the set of regular expressions to obtain multiple preprocessed expressions includes:

[0013] Each regular expression in the set of regular expressions is segmented into words to obtain multiple segmented expressions.

[0014] Stop words in each of the segmented expressions are filtered out to obtain multiple preprocessed expressions.

[0015] Optionally, the optimized DBSCAN algorithm is used to cluster the target expressions in all the first deleted expression groups to obtain multiple clustered expression clusters, including:

[0016] The target expression in each of the first deleted expression groups is converted into a vector representation to obtain the converted expression;

[0017] The optimized DBSCAN algorithm is used to cluster the transformed expression to obtain multiple clustered expression clusters.

[0018] Optionally, converting the target expression in each of the first deleted expression groups into a vector representation to obtain the converted expression includes:

[0019] The target expression in each of the first deleted expression groups is converted into a vector representation using the bert4vec vector generation tool to obtain the converted expression.

[0020] Optionally, the optimized DBSCAN algorithm is used to cluster the transformed expression to obtain multiple clustered expression clusters, including:

[0021] A second similarity value is obtained by calculating the similarity of the transformed expressions in all the first deleted expression groups of different categories.

[0022] Determine whether the second similarity value exceeds the second threshold. If so, delete the transformed expression in the first deleted expression group corresponding to the second similarity value that exceeds the second threshold to obtain the second deleted expression group.

[0023] The optimized DBSCAN algorithm is used to cluster all the transformed expressions in the second group of deleted expressions to obtain multiple clustered expression clusters.

[0024] Optionally, the optimized DBSCAN algorithm is used to cluster the target expressions in all the first deleted expression groups to obtain multiple clustered expression clusters, including:

[0025] The optimized DBSCAN algorithm is used to cluster the target expressions in all the first deleted expression groups to obtain multiple clustered expression clusters. The regular expressions contained in each clustered expression cluster are then aggregated to obtain aggregated expressions.

[0026] Accordingly, the step of matching the target short text using the clustered expression clusters to obtain the matching results includes:

[0027] The target short text is matched using the aggregated expression to obtain the matching result.

[0028] Optionally, the aggregation of the regular expressions contained in each clustered expression cluster to obtain the aggregated expression includes:

[0029] The regular expressions contained in each clustered expression cluster are aggregated using natural language processing algorithms to obtain aggregated expressions.

[0030] Secondly, this application discloses a short text matching device, comprising:

[0031] The expression acquisition module is used to collect regular expressions written for the target business data and obtain a set of regular expressions;

[0032] The preprocessing module is used to preprocess each regular expression in the set of regular expressions to obtain multiple preprocessed expressions;

[0033] The expression classification module is used to classify all the preprocessed expressions to obtain multiple groups of classified expressions;

[0034] The similarity calculation module is used to calculate the similarity of the regular expressions in each of the classified expression groups to obtain the first similarity value corresponding to each of the classified expression groups.

[0035] The judgment module is used to determine whether the first similarity value exceeds the first threshold;

[0036] The determination module is used to determine any regular expression from the group of classified expressions corresponding to the first similarity values ​​that exceed the first threshold if the first similarity value exceeds the first threshold, so as to obtain the target expression;

[0037] The deletion module is used to delete all regular expressions except the target expression in each of the classified expression groups, so as to obtain a first deleted expression group containing only the target expression.

[0038] The clustering module is used to cluster the target expressions in all the first deleted expression groups using the optimized DBSCAN algorithm to obtain multiple clustered expression clusters.

[0039] The short text matching module is used to match the target short text using the clustered expression cluster when a short text matching request is received, and to obtain the matching result.

[0040] Thirdly, this application discloses an electronic device, including a processor and a memory; wherein, when the processor executes a computer program stored in the memory, it implements the aforementioned short text matching method.

[0041] Fourthly, this application discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the aforementioned short text matching method.

[0042] As can be seen, this application first collects regular expressions written for the target business data to obtain a set of regular expressions, and preprocesses each regular expression in the set to obtain multiple preprocessed expressions. Then, it classifies all the preprocessed expressions to obtain multiple groups of categorized expressions, and calculates the similarity of the regular expressions in each group to obtain a first similarity value for each group. Next, it determines whether the first similarity value exceeds a first threshold. If so, it selects any regular expression from the groups of categorized expressions corresponding to the first similarity values ​​that exceed the first threshold to obtain the target expression. Then, it deletes all regular expressions except the target expression from each group of categorized expressions to obtain a first group of deleted expressions containing only the target expression. Finally, it uses the optimized DBSCAN algorithm to cluster the target expressions in all the first groups of deleted expressions to obtain multiple clustered expression clusters. When a short text matching request is received, the clustered expression clusters are used to match the target short text to obtain a matching result. This application addresses the root cause affecting matching efficiency and accuracy—regular expressions. It first categorizes regular expressions in the set and then removes highly similar expressions from the categorized expression groups, thereby reducing the number of expressions (i.e., removing redundant expressions). Then, it uses an optimized DBSCAN algorithm to cluster all the deleted expressions into multiple clustered expression clusters, further reducing the number of matching categories. This improves the accuracy and efficiency of short text matching while reducing the false positive rate. Attached Figure Description

[0043] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0044] Figure 1 This is a flowchart of a short text matching method disclosed in this application;

[0045] Figure 2 This is a schematic diagram of a specific short text matching method disclosed in this application;

[0046] Figure 3 Here is a flowchart of a specific short text matching method disclosed in this application;

[0047] Figure 4 This application discloses a schematic diagram of a short text matching device;

[0048] Figure 5 This application discloses a structural diagram of an electronic device. Detailed Implementation

[0049] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0050] This application discloses a short text matching method, see [link to relevant documentation]. Figure 1 As shown, the method includes:

[0051] Step S11: Collect regular expressions written for the target business data to obtain a set of regular expressions, and preprocess each regular expression in the set of regular expressions to obtain multiple preprocessed expressions.

[0052] Understandably, see Figure 2 As shown, for different business operations, business personnel can tag business data based on industry background and knowledge, as well as existing industry template information, such as industry guidelines and their own knowledge reserves. They can then write corresponding expressions (i.e., regular expressions) based on the tags. As business data increases dramatically, the number of expressions also surges. At the same time, the different writing methods of various business personnel lead to low matching efficiency, easy conflicts, high redundancy, and high false alarms in the surged expressions. The industry template information can contain several or even hundreds of categories.

[0053] In this embodiment, to address the aforementioned issues, a set of regular expressions can first be collected from the regular expressions written by different business personnel for the target business data. The target business data can originate from a database, and its specific format can be documents, images, etc. Next, each regular expression in the collected set is preprocessed to obtain multiple preprocessed expressions.

[0054] Specifically, the step of preprocessing each regular expression in the regular expression set to obtain multiple preprocessed expressions may include: performing word segmentation on each regular expression in the regular expression set to obtain multiple word-segmented expressions; and filtering stop words in each word-segmented expression to obtain multiple preprocessed expressions. In this embodiment, to eliminate noisy expression data, the regular expressions in the regular expression set can be segmented one by one according to the dimension of the expression to obtain multiple word-segmented expressions. Then, stop words (i.e., stopwords) in the word-segmented expressions are filtered, such as filtering words like "a", "an", "the", "and", "is", and "of".

[0055] Step S12: Classify all the preprocessed expressions to obtain multiple groups of classified expressions, and calculate the similarity of the regular expressions in each group of classified expressions to obtain the first similarity value corresponding to each group of classified expressions.

[0056] In this embodiment, after preprocessing each regular expression in the regular expression set to obtain multiple preprocessed expressions, in order to eliminate expression redundancy, all the preprocessed expressions can be classified according to categories to obtain multiple categorized expression groups C, that is, the expressions are divided into multiple category spaces. Then, the similarity is calculated for the regular expressions in each of the categorized expression groups, that is, matching is performed, so as to obtain multiple similarity values ​​corresponding to each categorized expression group.

[0057] Step S13: Determine whether the first similarity value exceeds the first threshold. If so, determine any regular expression from the group of classified expressions corresponding to the first similarity value that exceeds the first threshold to obtain the target expression, and delete all regular expressions except the target expression in each group of classified expressions to obtain a first group of deleted expressions containing only the target expression.

[0058] In this embodiment, the similarity of the regular expressions in each group of classified expressions is calculated to obtain the first similarity value corresponding to each group of classified expressions. Then, it is determined whether the first similarity value exceeds the first threshold, i.e., the preset similarity threshold_redundant_sim. If the first similarity value exceeds the threshold, any regular expression is determined from the group of classified expressions corresponding to the first similarity value that exceeds the threshold to obtain the target expression. Then, all regular expressions in each group of classified expressions except the target expression are deleted. That is, when the similarity threshold is exceeded, any expression in the expression group is retained, thereby obtaining a deleted expression group that only contains the target expression.

[0059] Step S14: Use the optimized DBSCAN algorithm to cluster the target expressions in all the first deleted expression groups to obtain multiple clustered expression clusters.

[0060] In this embodiment, after obtaining the first group of deleted expressions containing only the target expression, considering the large number of expressions in the expression group and the unknown number of actual categories, the DBSCAN (Density-Based Spatial Clustering of Applications with Noise) algorithm, which can divide regions with sufficiently high density into clusters and can discover arbitrary shapes in a noisy spatial database, can be used for clustering. Specifically, the optimized DBSCAN algorithm can be used to cluster all the target expressions in the first group of deleted expressions to obtain multiple clustered expression clusters. However, it should be noted that the DBSCAN algorithm in this application is an optimized DBSCAN algorithm. By using a custom loss function, an automated mechanism is formed to perform clustering operations. The formula for the custom loss function is:

[0061]

[0062]

[0063] Where k represents all parameter combinations of hyperparameters in the DBSCAN algorithm, used for automated parameter finding; N is the total number of expressions, and n is the number of clusters; a i b i z iα1, β1, β2, λ1, λ2, and λ3 are the number of expressions, the number of categories, and the number of expressions corresponding to the category with the highest frequency within the cluster, respectively. α1, β1, β2, λ1, λ2, and λ3 are all artificial parameters. α1 ≥ 1 is used to penalize clusters with a large number of categories. β1 and β2 are both integers used to penalize clusters that are too large or too small after clustering. λ1, λ2, and λ3 are all 1 by default and are used to balance the loss values ​​of each part. It should be noted that k mainly includes the neighborhood radius ε and the minimum number of expressions num for a single cluster, where ε∈{0.5, 1, 2, ..., 10} and num∈{1, 2, ..., 20}. By combining the parameters of k, the optimal parameters ε and num can be obtained automatically, and the clustering algorithm can be executed with this set of parameters to obtain n clusters.

[0064] Specifically, the step of clustering the target expressions in all the first deleted expression groups using the optimized DBSCAN algorithm to obtain multiple clustered expression clusters may include: converting the target expressions in each of the first deleted expression groups into vector representations to obtain converted expressions; and then using the optimized DBSCAN algorithm to cluster the converted expressions to obtain multiple clustered expression clusters. In this embodiment, the target expressions in each of the first deleted expression groups can first be converted into vector representations (vec, vector), that is, the expression text can be converted into vector form to obtain converted expressions. Then, the optimized DBSCAN algorithm can be used to cluster the converted expressions to obtain multiple clustered expression clusters.

[0065] In this embodiment, the step of clustering the transformed expressions using the optimized DBSCAN algorithm to obtain multiple clustered expression clusters may specifically include: calculating the similarity of the transformed expressions in all the first deleted expression groups of different categories to obtain a second similarity value; determining whether the second similarity value exceeds a second threshold, and if so, deleting the transformed expressions in the first deleted expression groups corresponding to the second similarity values ​​that exceed the second threshold to obtain a second deleted expression group; and using the optimized DBSCAN algorithm to cluster the transformed expressions in all the second deleted expression groups to obtain multiple clustered expression clusters. In this embodiment, to eliminate expression conflicts, a second similarity value can be calculated by category dimension for the transformed expressions in all first deleted expression groups of different categories, i.e., matching and similarity calculation is performed on expressions of different categories. Then, it is determined whether the second similarity value exceeds the second threshold threshold_conflict_sim. If the second similarity value exceeds the second threshold threshold_conflict_sim, the transformed expressions in the first deleted expression group corresponding to the second similarity value exceeding the second threshold threshold_conflict_sim are directly deleted to obtain a second deleted expression group. Furthermore, the character length of the expression groups contained in the second deleted expression group can be calculated, and expressions with a character length lower than the threshold threshold_char can be sent to relevant technical personnel for manual confirmation to retain the usable parts of the expressions. Then, the optimized DBSCAN algorithm is used to cluster all the transformed expressions in the second deleted expression group to obtain multiple clustered expression clusters. Furthermore, the clustering results of the optimized DBSCAN algorithm can be measured using euclid (euclidean distance), with the following formula:

[0066]

[0067] In the formula, n is the dimension of vec, x i y i vec x vec y The value of the i-th dimension.

[0068] Step S15: When a short text matching request is received, the target short text is matched using the clustered expression cluster to obtain the matching result.

[0069] In this embodiment, when a short text matching request is received from the user terminal, the target short text can be directly matched using the clustered expression clusters described above to obtain the corresponding matching results, and these matching results are then returned to the user terminal. It should be noted that when there are multiple matching results, either a preset number or all matching results can be returned to the user terminal; the choice can be made based on actual application needs, and no specific limitation is made here.

[0070] As can be seen, in this embodiment, regular expressions written for target business data are first collected to obtain a set of regular expressions. Each regular expression in the set is preprocessed to obtain multiple preprocessed expressions. Then, all preprocessed expressions are classified to obtain multiple groups of categorized expressions. The similarity of regular expressions in each group of categorized expressions is calculated to obtain a first similarity value for each group of categorized expressions. Next, it is determined whether the first similarity value exceeds a first threshold. If so, any regular expression is selected from the groups of categorized expressions corresponding to the first similarity values ​​that exceed the first threshold to obtain the target expression. All regular expressions except the target expression are deleted from each group of categorized expressions to obtain a first group of deleted expressions containing only the target expression. The optimized DBSCAN algorithm is then used to cluster the target expressions in all the first groups of deleted expressions to obtain multiple clustered expression clusters. When a short text matching request is received, the clustered expression clusters are used to match the target short text to obtain a matching result. This application's embodiments start from the root cause affecting matching efficiency and accuracy, namely regular expressions. First, the regular expressions in the regular expression set are classified, and expressions with high similarity in the classified expression groups are deleted, thereby reducing the number of expressions, i.e., removing redundant expressions. Then, the optimized DBSCAN algorithm is used to cluster the expressions in all the deleted expression groups to obtain multiple clustered expression clusters, thereby reducing the number of matching categories, thus improving the accuracy and efficiency of short text matching, and reducing the false alarm rate.

[0071] This application discloses a specific short text matching method. See [link to relevant documentation]. Figure 3 As shown, the method includes:

[0072] Step S21: Collect regular expressions written for the target business data to obtain a set of regular expressions, and perform word segmentation on each regular expression in the set of regular expressions to obtain multiple word segmented expressions.

[0073] Step S22: Filter out the stop words in each of the segmented expressions to obtain multiple preprocessed expressions.

[0074] Step S23: Classify all the preprocessed expressions to obtain multiple groups of classified expressions, and calculate the similarity of the regular expressions in each group of classified expressions to obtain the first similarity value corresponding to each group of classified expressions.

[0075] Step S24: Determine whether the first similarity value exceeds the first threshold. If so, determine any regular expression from the group of classified expressions corresponding to the first similarity value that exceeds the first threshold to obtain the target expression, and delete all regular expressions except the target expression in each group of classified expressions to obtain a first group of deleted expressions containing only the target expression.

[0076] Step S25: Use the bert4vec vector generation tool to convert the target expression in each of the first deleted expression groups into a vector representation to obtain the converted expression.

[0077] In this embodiment, after obtaining the first group of deleted expressions containing only the target expression, the target expression in each of the first group of deleted expressions can be converted into a vector representation using the bert4vec (a pre-trained sentence vector generation tool) vector generation tool to obtain the corresponding converted expression.

[0078] Step S26: Calculate the similarity of the transformed expressions in all the first deleted expression groups of different categories to obtain a second similarity value.

[0079] Step S27: Determine whether the second similarity value exceeds the second threshold. If so, delete the transformed expression in the first deleted expression group corresponding to the second similarity value that exceeds the second threshold to obtain the second deleted expression group.

[0080] Step S28: Use the optimized DBSCAN algorithm to cluster all the transformed expressions in the second deleted expression group to obtain multiple clustered expression clusters, and aggregate the regular expressions contained in each clustered expression cluster to obtain aggregated expressions.

[0081] In this embodiment, the optimized DBSCAN algorithm can be used to cluster the transformed expressions in all the above-mentioned second deleted expression groups to obtain multiple clustered expression clusters. Then, the regular expressions contained in each of the above-mentioned clustered expression clusters are aggregated to obtain aggregated expressions. That is, the expressions are aggregated by cluster. For example, for the i-th cluster, the number of expressions corresponding to the category with the highest frequency in the cluster is z. i , for zi The expressions are merged to obtain an aggregated expression, such as A|B, where the aggregated state of expressions A and B is A|B. Furthermore, the aggregated expression can be manually validated and adjusted or supplemented according to the actual application scenario.

[0082] In one specific implementation, the aggregation of regular expressions contained in each clustered expression cluster to obtain an aggregated expression may specifically include: using a natural language processing (NLP) algorithm to aggregate the regular expressions contained in each clustered expression cluster to obtain an aggregated expression. In this embodiment, a NLP algorithm can be used to aggregate the regular expressions contained in each clustered expression cluster to obtain the aggregated expression; wherein, the NLP algorithm includes, but is not limited to, Bag of Words, TF-IDF (Term Frequency-Inverse Document Frequency, a weighted technique for information retrieval and data mining), Word2Vec (a group of related models used to generate word vectors), etc.

[0083] Step S29: When a short text matching request is received, the target short text is matched using the aggregated expression to obtain the matching result.

[0084] In this embodiment, when a short text matching request is received, the target short text can be directly matched using the above-mentioned aggregated expression A|B to obtain the corresponding matching result.

[0085] For more detailed processing procedures regarding steps S21 to S24, S26, and S27, please refer to the corresponding content disclosed in the foregoing embodiments, which will not be repeated here.

[0086] As can be seen, the embodiments of this application address the root cause affecting matching efficiency and accuracy by starting with regular expressions, cleaning large amounts of expressions, then vectorizing and clustering the expression text, and finally using the optimized DBSCAN algorithm to automatically generate a controllable number of expressions, thus fundamentally solving problems such as inefficiency, conflict, high redundancy, and high false alarms in expressions.

[0087] Accordingly, embodiments of this application also disclose a short text matching device, see [link to relevant documentation]. Figure 4 As shown, the device includes:

[0088] The expression acquisition module 11 is used to acquire regular expressions written for the target business data and obtain a set of regular expressions;

[0089] Preprocessing module 12 is used to preprocess each regular expression in the set of regular expressions to obtain multiple preprocessed expressions;

[0090] The expression classification module 13 is used to classify all the preprocessed expressions to obtain multiple groups of classified expressions.

[0091] Similarity calculation module 14 is used to calculate the similarity of the regular expressions in each of the classified expression groups to obtain the first similarity value corresponding to each of the classified expression groups;

[0092] The judgment module 15 is used to determine whether the first similarity value exceeds the first threshold;

[0093] The determining module 16 is used to determine any regular expression from the group of classified expressions corresponding to the first similarity values ​​that exceed the first threshold if the first similarity value exceeds the first threshold, so as to obtain the target expression;

[0094] Deletion module 17 is used to delete all regular expressions except the target expression in each of the classified expression groups, to obtain a first deleted expression group containing only the target expression;

[0095] Clustering module 18 is used to cluster the target expressions in all the first deleted expression groups using the optimized DBSCAN algorithm to obtain multiple clustered expression clusters;

[0096] The short text matching module 19 is used to match the target short text using the clustered expression cluster when a short text matching request is received, and to obtain the matching result.

[0097] The specific workflow of each of the above modules can be found in the relevant content disclosed in the foregoing embodiments, and will not be repeated here.

[0098] As can be seen, in this embodiment, regular expressions written for the target business data are first collected to obtain a set of regular expressions. Each regular expression in the set is preprocessed to obtain multiple preprocessed expressions. Then, all the preprocessed expressions are classified to obtain multiple groups of classified expressions. The similarity of the regular expressions in each group of classified expressions is calculated to obtain a first similarity value for each group of classified expressions. Next, it is determined whether the first similarity value exceeds a first threshold. If so, any regular expression is selected from the groups of classified expressions corresponding to the first similarity values ​​that exceed the first threshold to obtain the target expression. All regular expressions except the target expression are deleted from each group of classified expressions to obtain a first group of deleted expressions containing only the target expression. Then, the optimized DBSCAN algorithm is used to cluster the target expressions in all the first groups of deleted expressions to obtain multiple clustered expression clusters. When a short text matching request is received, the clustered expression clusters are used to match the target short text to obtain a matching result. This application's embodiments start from the root cause affecting matching efficiency and accuracy, namely regular expressions. First, the regular expressions in the regular expression set are classified, and expressions with high similarity in the classified expression groups are deleted, thereby reducing the number of expressions, i.e., removing redundant expressions. Then, the optimized DBSCAN algorithm is used to cluster the expressions in all the deleted expression groups to obtain multiple clustered expression clusters, thereby reducing the number of matching categories, thus improving the accuracy and efficiency of short text matching, and reducing the false alarm rate.

[0099] In some specific embodiments, the preprocessing module 12 may specifically include:

[0100] The word segmentation processing unit is used to perform word segmentation processing on each regular expression in the set of regular expressions to obtain multiple word-segmented expressions;

[0101] The stop word filtering unit is used to filter stop words in each of the segmented expressions to obtain multiple preprocessed expressions.

[0102] In some specific embodiments, the clustering module 18 may specifically include:

[0103] The first conversion unit is used to convert the target expression in each of the first deleted expression groups into a vector representation to obtain the converted expression.

[0104] The first clustering unit is used to cluster the transformed expression using the optimized DBSCAN algorithm to obtain multiple clustered expression clusters.

[0105] In some specific embodiments, the first conversion unit may specifically include:

[0106] The second conversion unit is used to convert the target expressions in each of the first deleted expression groups into vector representations using the bert4vec vector generation tool, thereby obtaining the converted expressions.

[0107] In some specific embodiments, the first clustering unit may specifically include:

[0108] A similarity calculation unit is used to calculate the similarity of the transformed expressions in all the first deleted expression groups of different categories to obtain a second similarity value;

[0109] A judgment unit is used to determine whether the second similarity value exceeds a second threshold.

[0110] The deletion unit is configured to delete the transformed expression in the first deleted expression group corresponding to the second similarity value that exceeds the second threshold if the second similarity value exceeds the second threshold, thereby obtaining a second deleted expression group.

[0111] The second clustering unit is used to cluster the transformed expressions in all the second deleted expression groups using the optimized DBSCAN algorithm, resulting in multiple clustered expression clusters.

[0112] In some specific embodiments, the clustering module 18 may specifically include:

[0113] A clustering unit is used to cluster the target expressions in all the first deleted expression groups using the optimized DBSCAN algorithm to obtain multiple clustered expression clusters.

[0114] The first aggregation unit is used to aggregate the regular expressions contained in each clustered expression cluster to obtain aggregated expressions.

[0115] Accordingly, the short text matching module 19 may specifically include:

[0116] The short text matching unit is used to match the target short text using the aggregated expression to obtain the matching result.

[0117] In some specific embodiments, the first aggregation unit may specifically include:

[0118] The second aggregation unit is used to aggregate the regular expressions contained in each clustered expression cluster using a natural language processing algorithm to obtain aggregated expressions.

[0119] Furthermore, embodiments of this application also disclose an electronic device, Figure 5 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content of the diagram should not be construed as limiting the scope of this application.

[0120] Figure 5 This is a schematic diagram of the structure of an electronic device 20 provided in an embodiment of this application. Specifically, the electronic device 20 may include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 stores a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the short text matching method disclosed in any of the foregoing embodiments. Alternatively, the electronic device 20 in this embodiment may specifically be an electronic computer.

[0121] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and is not specifically limited here; the input / output interface 25 is used to acquire external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs, and is not specifically limited here.

[0122] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or optical disk, etc. The resources stored thereon can include operating system 221, computer program 222, etc., and the storage method can be temporary storage or permanent storage.

[0123] The operating system 221 is used to manage and control the various hardware devices on the electronic device 20 and the computer program 222, which may be Windows Server, Netware, Unix, Linux, etc. In addition to including a computer program capable of performing the short text matching method executed by the electronic device 20 as disclosed in any of the foregoing embodiments, the computer program 222 may further include a computer program capable of performing other specific tasks.

[0124] Furthermore, this application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the aforementioned short text matching method. Specific steps of this method can be found in the corresponding content disclosed in the foregoing embodiments, and will not be repeated here.

[0125] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.

[0126] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0127] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0128] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0129] The above provides a detailed description of a short text matching method, apparatus, device, and storage medium provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A short text matching method, characterized in that, include: Collect regular expressions written for the target business data to obtain a set of regular expressions, and preprocess each regular expression in the set of regular expressions to obtain multiple preprocessed expressions; All the preprocessed expressions are classified to obtain multiple groups of classified expressions, and the similarity of the regular expressions in each group of classified expressions is calculated to obtain the first similarity value corresponding to each group of classified expressions. Determine whether the first similarity value exceeds the first threshold. If so, determine any regular expression from the group of classified expressions corresponding to the first similarity value that exceeds the first threshold to obtain the target expression. Then, delete all regular expressions in each group of classified expressions except the target expression to obtain a first group of deleted expressions containing only the target expression. The optimized DBSCAN algorithm is used to cluster all the target expressions in the first group of deleted expressions to obtain multiple clustered expression clusters. When a short text matching request is received, the target short text is matched using the clustered expression cluster to obtain the matching result; The step of clustering the target expressions in all the first deleted expression groups using the optimized DBSCAN algorithm to obtain multiple clustered expression clusters includes: converting the target expressions in each of the first deleted expression groups into vector representations to obtain converted expressions; and clustering the converted expressions using the optimized DBSCAN algorithm to obtain multiple clustered expression clusters. The step of clustering the transformed expressions using the optimized DBSCAN algorithm to obtain multiple clustered expression clusters includes: calculating the similarity of the transformed expressions in all first deleted expression groups of different categories to obtain a second similarity value; determining whether the second similarity value exceeds a second threshold, and if so, deleting the transformed expressions in the first deleted expression groups corresponding to the second similarity values ​​that exceed the second threshold to obtain a second deleted expression group; and clustering the transformed expressions in all second deleted expression groups using the optimized DBSCAN algorithm to obtain multiple clustered expression clusters. The step of using the optimized DBSCAN algorithm to cluster the transformed expressions in all the second deleted expression groups to obtain multiple clustered expression clusters includes: calculating the character length of the expression groups contained in the second deleted expression group, sending expressions with character lengths lower than a third threshold to relevant technical personnel for manual confirmation, retaining the usable parts of the expressions, and then using the optimized DBSCAN algorithm to cluster the transformed expressions in all the second deleted expression groups to obtain multiple clustered expression clusters.

2. The short text matching method according to claim 1, characterized in that, The process involves preprocessing each regular expression in the set of regular expressions to obtain multiple preprocessed expressions, including: Each regular expression in the set of regular expressions is segmented into words to obtain multiple segmented expressions. Stop words in each of the segmented expressions are filtered out to obtain multiple preprocessed expressions.

3. The short text matching method according to claim 1, characterized in that, The step of converting the target expression in each of the first deleted expression groups into a vector representation to obtain the converted expression includes: The target expression in each of the first deleted expression groups is converted into a vector representation using the bert4vec vector generation tool to obtain the converted expression.

4. The short text matching method according to any one of claims 1 to 3, characterized in that, The optimized DBSCAN algorithm is used to cluster the target expressions in all the first deleted expression groups to obtain multiple clustered expression clusters, including: The optimized DBSCAN algorithm is used to cluster the target expressions in all the first deleted expression groups to obtain multiple clustered expression clusters. The regular expressions contained in each clustered expression cluster are then aggregated to obtain aggregated expressions. Accordingly, the step of matching the target short text using the clustered expression clusters to obtain the matching results includes: The target short text is matched using the aggregated expression to obtain the matching result.

5. The short text matching method according to claim 4, characterized in that, The aggregation of regular expressions contained in each clustered expression cluster to obtain aggregated expressions includes: The regular expressions contained in each clustered expression cluster are aggregated using natural language processing algorithms to obtain aggregated expressions.

6. A short text matching device, characterized in that, include: The expression acquisition module is used to collect regular expressions written for the target business data and obtain a set of regular expressions; The preprocessing module is used to preprocess each regular expression in the set of regular expressions to obtain multiple preprocessed expressions; The expression classification module is used to classify all the preprocessed expressions to obtain multiple groups of classified expressions; The similarity calculation module is used to calculate the similarity of the regular expressions in each of the classified expression groups to obtain the first similarity value corresponding to each of the classified expression groups. The judgment module is used to determine whether the first similarity value exceeds the first threshold; The determination module is used to determine any regular expression from the group of classified expressions corresponding to the first similarity values ​​that exceed the first threshold if the first similarity value exceeds the first threshold, so as to obtain the target expression; The deletion module is used to delete all regular expressions except the target expression in each of the classified expression groups, so as to obtain a first deleted expression group containing only the target expression. The clustering module is used to cluster the target expressions in all the first deleted expression groups using the optimized DBSCAN algorithm to obtain multiple clustered expression clusters. The short text matching module is used to match the target short text using the clustered expression clusters when a short text matching request is received, and to obtain the matching result. The clustering module is specifically used to convert the target expression in each of the first deleted expression groups into a vector representation to obtain the converted expression; and to cluster the converted expression using the optimized DBSCAN algorithm to obtain multiple clustered expression clusters. The clustering module is also used to calculate the similarity of the transformed expressions in all the first deleted expression groups of different categories to obtain a second similarity value; Determine whether the second similarity value exceeds the second threshold. If so, delete the transformed expression in the first deleted expression group corresponding to the second similarity value that exceeds the second threshold to obtain the second deleted expression group. The optimized DBSCAN algorithm is used to cluster all the transformed expressions in the second group of deleted expressions to obtain multiple clustered expression clusters. The clustering module is also used to calculate the character length of the expression groups contained in the second deleted expression group, and send the expressions with character lengths lower than the third threshold to relevant technical personnel for manual confirmation, retain the usable expressions, and then use the optimized DBSCAN algorithm to cluster all the transformed expressions in the second deleted expression group to obtain multiple clustered expression clusters.

7. An electronic device, characterized in that, It includes a processor and a memory; wherein, when the processor executes a computer program stored in the memory, it implements the short text matching method as described in any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that, Used to store a computer program; wherein, when the computer program is executed by a processor, it implements the short text matching method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Method and apparatus for performing decision-making on request according to rules

    CN108764726A

  • Static check rule set generation method and device

    CN116756004A