Text matching method and device, medium and program product

Through the text matching method based on natural language processing, using feature vectors and similarity thresholds, the efficiency and accuracy of interpretation and matching work in traditional compliance inspections are solved, and automated text matching is achieved, and efficiency and accuracy are improved.

CN120216677APending Publication Date: 2025-06-27E-CAPITAL TRANSFER CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510686201.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In traditional information system compliance and risk control, supervision inspection, external audit and internal self-inspection work need to be completed within a limited time, and the interpretation and matching of laws and regulations is time-consuming and laborious, and the accuracy of manual interpretation and matching is low.

Method used

A text matching method based on natural language processing is provided. By obtaining multiple feature vectors associated with the query text, the matching candidate matching text is determined from it based on the similarity in the vector database of the candidate text, and the matching result is optimized using the target similarity threshold.

Benefits of technology

It improves the efficiency and accuracy of text matching, reduces the time and cost of manual interpretation and matching, and can automatically parse and understand the text content of the inspection items and task items, achieving accurate and fast matching of the two.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216677A_ABST
    Figure CN120216677A_ABST
Patent Text Reader

Abstract

The invention relates to a text matching method and device, a non-transitory computer readable storage medium and a computer program product. The text matching method according to one aspect of the invention comprises the following steps: acquiring a plurality of first feature vectors associated with a query text; determining feature vectors of a plurality of candidate matching texts matched with each first feature vector from the vector database of the candidate texts based on a first similarity between the feature vectors of the candidate texts in the vector database of the candidate texts and each first feature vector in a plurality of first feature vectors of the query text; sorting the feature vectors of the plurality of candidate matching texts based on a second similarity between the second feature vector of the query text and the feature vector of each candidate matching text in the feature vectors of the plurality of candidate matching texts; and processing the feature vectors of the plurality of sorted candidate matching texts by using a target similarity threshold and determining a target matching text of the query text based on a processing result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of natural language processing, and more particularly to a method for text matching, an apparatus for text matching, a non-transitory computer-readable storage medium, and a computer program product. Background Art

[0002] With the development of society and the continuous change of the economic environment, financial supervision has become increasingly complex and strict. As one of the institutions subject to strict supervision, financial institutions have to undergo multiple regulatory inspections, external audits, and internal self-inspections every year.

[0003] The main challenge faced by traditional information system compliance and risk control is that regulatory inspections, external audits, and internal self-inspections usually need to be completed within a limited time. However, the work of interpreting and matching laws and regulations (for example, matching daily task items according to inspection items to assist financial institutions in filling in compliance inspection items) is time-consuming and laborious, and the accuracy of manual interpretation and matching work is relatively low.

[0004] In view of this, it is desirable to propose a text matching method based on natural language processing. Summary of the Invention

[0005] In order to solve or at least alleviate one or more of the above problems, the following technical solutions are provided.

[0006] According to a first aspect of the present application, there is provided a method for text matching, the method comprising the following steps: obtaining a plurality of first feature vectors associated with a query text; determining, from a vector database of candidate texts, a plurality of feature vectors of candidate matching texts that match each of the first feature vectors among the plurality of first feature vectors of the query text based on a first similarity between the feature vectors of the candidate texts in the vector database of candidate texts and each of the first feature vectors of the query text; sorting the plurality of feature vectors of the candidate matching texts based on a second similarity between a second feature vector of the query text and each of the feature vectors of the plurality of candidate matching texts; and processing the sorted plurality of feature vectors of the candidate matching texts using a target similarity threshold and determining a target matching text of the query text based on the processing result.

[0007] According to the method for text matching according to an embodiment of the present application, wherein obtaining a plurality of first feature vectors associated with a query text includes: performing word segmentation processing on the query text to obtain query text units of the query text; replacing the query text units with synonymous text units of the query text units to obtain a plurality of synonymous query texts corresponding to the query text; and encoding the plurality of synonymous query texts to obtain a plurality of first feature vectors associated with the query text.

[0008] The method for text matching according to an embodiment of the present application or any one of the above embodiments, wherein the vector database of the candidate texts is obtained by: acquiring a plurality of candidate texts to be matched with the query text; and performing vectorization processing on the plurality of candidate texts to obtain the vector database of the candidate texts.

[0009] The method for text matching according to an embodiment of the present application or any one of the above embodiments, wherein determining the feature vectors of a plurality of candidate matching texts that match each first feature vector among the plurality of first feature vectors of the query text based on the first similarity between the feature vectors of the candidate texts in the vector database of the candidate texts and each first feature vector includes: selecting, in descending order of the first similarity, the vectors of a predetermined number of candidate texts from the vector database of the candidate texts as the feature vectors of the plurality of candidate matching texts that match each first feature vector.

[0010] The method for text matching according to an embodiment of the present application or any one of the above embodiments, wherein the target similarity threshold is determined by: acquiring the label data of a plurality of check item texts and the candidate texts in the vector database of the candidate texts; determining the feature vectors of a plurality of matching texts that match the feature vectors of each check item text from the vector database of the candidate texts; and determining the target similarity threshold based at least on the label data of the candidate texts and the feature vectors of the plurality of matching texts that match the feature vectors of each check item text.

[0011] The method for text matching according to an embodiment of the present application or any one of the above embodiments, wherein determining the target similarity threshold based at least on the label data of the candidate texts and the feature vectors of the plurality of matching texts that match the feature vectors of each check item text includes: determining a plurality of similarities between the feature vectors of the plurality of check item texts and the feature vectors of the plurality of matching texts that match the feature vectors of each check item text; taking one or more of the plurality of similarities as a preset similarity threshold; using the preset similarity threshold to determine a mis-match rate based on the label data of the candidate texts and the feature vectors of the plurality of matching texts that match the feature vectors of each check item text; and determining the target similarity threshold based on the mis-match rate.

[0012] The method for text matching according to an embodiment of the present application or any one of the above embodiments, wherein determining the target similarity threshold based on the mis-match rate includes: determining the preset similarity threshold that makes the mis-match rate less than the mis-match rate threshold as the target similarity threshold.

[0013] The method for text matching according to an embodiment of the present application or any one of the above embodiments, wherein sorting the feature vectors of the multiple candidate matching texts based on the second similarity between the second feature vector of the query text and the feature vector of each candidate matching text among the feature vectors of the multiple candidate matching texts includes: encoding the query text to obtain the second feature vector of the query text; determining the second similarity between the second feature vector of the query text and the feature vector of each candidate matching text among the feature vectors of the multiple candidate matching texts; and sorting the feature vectors of the multiple candidate matching texts in descending order according to the second similarity.

[0014] The method for text matching according to an embodiment of the present application or any one of the above embodiments, wherein processing the feature vectors of the sorted multiple candidate matching texts by using a target similarity threshold and determining the target matching text of the query text based on the processing result includes: removing the feature vectors of one or more candidate matching texts from the feature vectors of the sorted multiple candidate matching texts based on the target similarity threshold; and determining the target matching text of the query text based on the feature vectors of the sorted multiple candidate matching texts after removing the feature vectors of one or more candidate matching texts.

[0015] The method for text matching according to an embodiment of the present application or any one of the above embodiments, wherein the method further includes: determining the first similarity corresponding to the feature vector of each sorted candidate matching text; removing the feature vectors of one or more candidate matching texts whose first similarity is less than the target similarity threshold from the feature vectors of the sorted multiple candidate matching texts; and among the feature vectors of the sorted multiple candidate matching texts after removing the feature vectors of one or more candidate matching texts, determining the candidate matching text corresponding to the feature vector of the candidate matching text ranked first as the target matching text of the query text.

[0016] According to a second aspect of the present application, there is provided a text matching device, the device includes: a memory; a processor coupled to the memory; and a computer program stored on the memory, which causes the steps of the text matching method according to the first aspect of the present application to be executed when the computer program runs on the processor.

[0017] According to a third aspect of the present application, there is provided a non-transitory computer-readable storage medium, characterized in that the non-transitory computer-readable storage medium includes instructions, and the instructions execute the steps of the text matching method according to the first aspect of the present application when running.

[0018] According to a fourth aspect of the present application, there is provided a computer program product, which includes instructions that, when executed by a processor, implement the steps of the method for text matching according to the first aspect of the present application.

[0019] The text matching solution according to one or more embodiments of the present application can determine the feature vectors of multiple candidate matching texts from the vector database of candidate texts based on the first similarity between the feature vectors of the candidate texts in the vector database of candidate texts and the first feature vector of the query text, and sort the feature vectors of the multiple candidate matching texts based on the second similarity between the second feature vector of the query text and the feature vectors of the multiple candidate matching texts, and use the target similarity threshold to determine the target matching text of the query text. Thus, vector search can be performed through the first similarity to improve the text matching efficiency, and at the same time, the feature vectors of the multiple candidate matching texts can be sorted through the second similarity to optimize the vector search result based on the first similarity according to the text semantics, so as to make full use of the semantic information of the query text and the multiple candidate matching texts, improve the text matching accuracy while improving the text matching efficiency, and save the labor and material costs. The text matching solution according to one or more embodiments of the present application can be applied to enterprise compliance inspections, for example, applied to the compliance inspection work of matching daily task items according to inspection items in financial institutions, automatically parsing and understanding the text content of inspection items and task items, so as to achieve accurate and rapid matching of the two, and save the labor and material costs. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The above and / or other aspects and advantages of the present application will become clearer and easier to understand through the following descriptions of various aspects in conjunction with the drawings, in which the same or similar units are denoted by the same reference numerals. In the drawings: Figure 1 A flowchart of the method for text matching according to one or more embodiments of the present application is shown.

[0021] Figure 2 A flowchart of the method for matching task items according to inspection items according to an embodiment of the present application is shown.

[0022] Figure 3 A schematic block diagram of the text matching device according to one or more embodiments of the present application is shown. DETAILED DESCRIPTION

[0023] Exemplary embodiments of the present application are described in detail below, and examples of these embodiments are illustrated in the accompanying drawings. It should be noted that the following description is for explanation and illustration purposes only and should not be construed as a limitation of the present application. Without departing from the principles of the present application, those skilled in the art can make electrical, mechanical, logical, and structural changes to these embodiments according to actual needs without departing from the scope of the present application. In addition, those skilled in the art can understand that one or more features of different embodiments described below can be combined according to any specific application scenario or actual needs.

[0024] Terms such as "including" and "comprising" indicate that in addition to the units and steps directly and explicitly stated in the specification, the technical solutions of the present application do not exclude the situation of having other units and steps that are not directly or explicitly stated. Terms such as "first" and "second" do not indicate the order of the units in terms of time, space, size, etc., but are only used to distinguish the units.

[0025] Hereinafter, various exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings.

[0026] Figure 1 A flowchart showing a method for text matching according to one or more embodiments of the present application is shown.

[0027] As Figure 1 shown, in step S101, a plurality of first feature vectors associated with the query text are obtained.

[0028] Optionally, in step S101, the query text can be segmented to obtain query text units of the query text, the query text units can be replaced with synonymous text units of the query text units to obtain a plurality of synonymous query texts corresponding to the query text, and the plurality of synonymous query texts can be encoded to obtain a plurality of first feature vectors associated with the query text.

[0029] In one embodiment, tokenizing the query text may include splitting the query text according to predefined rules (e.g., punctuation marks, spaces, etc.) to obtain query text units of the query text. In one embodiment, tokenizing the query text may include determining the tokenization boundaries of the text through statistical methods (e.g., maximum matching method, Markov model, etc.), and splitting the query text according to the tokenization boundaries to obtain query text units of the query text. In one embodiment, a machine learning model may be used to obtain the tokenization boundaries of the text, and split the query text according to the tokenization boundaries to obtain query text units of the query text. Exemplarily, the query text may be "Whether to review the functions and permissions of the risk management system". After tokenizing this query text, the query text units of this query text may include "Whether", "review", "risk", "risk management", "management", "management system", "system", "of", "function", "and", "permission".

[0030] In one embodiment, for each query text unit of the query text, a synonym list thereof may be looked up in a thesaurus, and one or more synonyms may be selected from the synonym list as synonymous text units to replace the query text units of the query text. The replaced synonymous text units are combined to obtain multiple synonymous query texts corresponding to the query text. In one embodiment, the query text may be tokenized and synonymously replaced to expand each query text into multiple (e.g., 3) synonymous query texts, increasing the subsequent vector search scope, ensuring the integrity of the subsequent vector search, avoiding the problem of missing sentence meaning of the query text caused by text vectorization, and improving the vector search recall rate.

[0031] In step S103, based on the first similarity between the feature vectors of the candidate texts in the vector database of the candidate texts and each first feature vector among the multiple first feature vectors of the query text, multiple candidate matching text feature vectors that match each first feature vector are determined from the vector database of the candidate texts.

[0032] In one embodiment, multiple candidate texts to be matched with the query text may be obtained, and the multiple candidate texts are vectorized to obtain a vector database of the candidate texts. It should be noted that the candidate texts can be understood as the texts collected to be matched with the query text.

[0033] In one embodiment, when querying for inspection items in the compliance inspection of a financial institution, daily task items can be collected as multiple candidate texts. For example, daily task items can be obtained from the laws and regulations involved in various business processes of the financial institution. It should be noted that daily task items can be understood as matters that the financial institution actively executes daily in accordance with the requirements in the laws and regulations. For example, according to the provisions of Chapter 5 of the "Guidelines for Information Technology Management of Futures Companies", the daily task item can be "whether an effective mechanism is established to ensure that the headquarters understands the operation status of each business department". In one embodiment, the inspection items in the compliance inspection of a financial institution can include inspection items issued by the national financial supervision and management institution for the compliance review of the financial institution's development strategy, internal norms, new product and new business plans, and major decision-making matters, which can formulate the basic framework and specific requirements for the compliance management of the financial institution, such as including but not limited to aspects such as compliance management architecture, responsibility assignment, compliance review, reporting mechanism, supervision and management, and legal liability. Exemplarily, the inspection items in the compliance inspection of a financial institution can include key control points for activities such as information system planning, construction, operation and maintenance, and emergency response, which can be used to judge the security of information system operation, the compliance of information system construction, and the application performance of information systems, etc.

[0034] In one embodiment, the similarity distance between the feature vector of a candidate text in the vector database of candidate texts and each of the multiple first feature vectors of the query text can be calculated, and this similarity distance can be used as the first similarity to determine the feature vectors of multiple candidate matching texts that match each first feature vector from the vector database of candidate texts. Exemplarily, the cosine similarity between the feature vector of a candidate text in the vector database of candidate texts and each of the multiple first feature vectors of the query text can be calculated as the first similarity.

[0035] Optionally, in step S103, for each first feature vector of the query text, a predetermined number of vector of candidate texts can be selected from the vector database of candidate texts as the feature vectors of multiple candidate matching texts that match this first feature vector in the order from the largest to the smallest of the first similarity. For example, for each first feature vector among the multiple first feature vectors associated with the query text, the vector of the candidate text ranked in the top five in terms of the first similarity between its feature vector and the feature vector of the candidate text in the vector database of candidate texts can be selected from the vector database of candidate texts as the feature vectors of multiple candidate matching texts that match this first feature vector.

[0036] In step S105, the feature vectors of the multiple candidate matching texts are sorted based on the second similarity between the second feature vector of the query text and the feature vector of each candidate matching text among the feature vectors of the multiple candidate matching texts.

[0037] Optionally, in step S105, the query text can be directly encoded to obtain the second feature vector of the query text. Exemplarily, the BERT (Bidirectional Encoder Representations from Transformers) model can be used to encode the query text to obtain the second feature vector of the query text, and the feature vectors of the multiple candidate matching texts are sorted in descending order according to the second similarity between the second feature vector of the query text and the feature vector of each candidate matching text among the feature vectors of the multiple candidate matching texts. By sorting the feature vectors of the multiple candidate matching texts, the vector search result based on the first similarity can be optimized according to the text semantics, so as to make full use of the semantic information of the query text and the multiple candidate matching texts, and improve the matching accuracy of the target matching text while improving the text matching efficiency.

[0038] As an example, assume that in step S103, based on the first similarity between the feature vector of the candidate text in the vector database of candidate texts and each first feature vector of the query text, the feature vectors of the multiple candidate matching texts matching each first feature vector determined from the vector database of candidate texts are A, B, C, D, and the first similarities corresponding to the feature vectors A, B, C, D are 0.9, 0.8, 0.6, 0.3 respectively. In step S105, the second similarities between the second feature vector of the query text and the feature vectors A, B, C, D of the multiple candidate matching texts are determined to be 0.5, 0.6, 0.4, 0.7 respectively. Based on the descending order of the second similarities 0.5, 0.6, 0.4, 0.7, the feature vectors A, B, C, D of the multiple candidate matching texts can be sorted as D, B, A, C.

[0039] In step S107, the target similarity threshold is used to process the sorted feature vectors of the multiple candidate matching texts and the target matching text of the query text is determined based on the processing result.

[0040] In one embodiment, it is possible to obtain the label data of candidate texts in a vector database of multiple inspection item texts and candidate texts, determine the feature vectors of multiple matching texts that match the feature vectors of each inspection item text from the vector database of candidate texts, and determine a target similarity threshold based at least on the label data of the candidate texts and the feature vectors of the multiple matching texts that match the feature vectors of each inspection item text. In one embodiment, the label data of the candidate texts can be used to indicate whether the candidate text matches each inspection item text. For example, it can indicate "match" or "no match" and is labeled as 1 or 0. Exemplarily, 2000 inspection items in real scenarios can be randomly collected from the daily operations of a financial institution as the above-mentioned multiple inspection item texts. For each inspection item text, determine the feature vectors of multiple matching texts that match the feature vectors of each inspection item text from the vector database of candidate texts. For example, the feature vectors of multiple matching texts that match the feature vectors of each inspection item text can be determined from the vector database of candidate texts through the similarity distance between vectors.

[0041] In one embodiment, it is possible to determine multiple similarities between the feature vectors of multiple inspection item texts and the feature vectors of multiple matching texts that match the feature vectors of each inspection item text, use one or more of the multiple similarities as a preset similarity threshold, use the preset similarity threshold to determine the False Accept Rate (FAR) based on the label data of the candidate texts and the feature vectors of the multiple matching texts that match the feature vectors of each inspection item text, and determine the target similarity threshold based on the FAR. In one embodiment, the preset similarity threshold that makes the FAR less than the FAR threshold can be determined as the target similarity threshold. Exemplarily, after determining the multiple similarities between the feature vectors of multiple inspection item texts and the feature vectors of multiple matching texts that match the feature vectors of each inspection item text, one or more similarities can be selected from high to low as the preset similarity threshold, calculate the FAR under each preset similarity threshold, and select the highest preset similarity threshold that makes the FAR less than 10% as the target similarity threshold.

[0042] Optionally, in step S107, one or more feature vectors of candidate matching texts can be removed from the sorted feature vectors of multiple candidate matching texts based on a target similarity threshold, and the target matching text of the query text can be determined based on the sorted feature vectors of multiple candidate matching texts after removing one or more feature vectors of candidate matching texts. In one embodiment, a first similarity corresponding to the feature vector of each sorted candidate matching text can be determined, one or more feature vectors of candidate matching texts with a first similarity less than the target similarity threshold can be removed from the feature vectors of multiple sorted candidate matching texts, and among the feature vectors of multiple sorted candidate matching texts after removing one or more feature vectors of candidate matching texts, the candidate matching text corresponding to the feature vector of the candidate matching text ranked first can be determined as the target matching text of the query text.

[0043] As an example, assume that in step S103, based on the first similarity between the feature vector of each candidate text in the vector database of candidate texts and each first feature vector of the query text, the feature vectors of multiple candidate matching texts that match each first feature vector of the query text determined from the vector database of candidate texts are A, B, C, D, and the first similarities corresponding to the feature vectors A, B, C, D are 0.9, 0.8, 0.6, 0.3 respectively. In step S105, the second similarities between the second feature vector of the query text and the feature vectors A, B, C, D of multiple candidate matching texts are determined to be 0.5, 0.6, 0.4, 0.7. Based on the order from largest to smallest of the second similarities 0.5, 0.6, 0.4, 0.7, the feature vectors A, B, C, D of multiple candidate matching texts can be sorted as D, B, A, C. As an example, assume that the target similarity threshold is 0.5. In step S107, the feature vectors of candidate matching texts corresponding to candidate matching texts with a first similarity less than 0.5 can be removed from the sorted feature vectors D, B, A, C of multiple candidate matching texts, that is, the feature vector D of the candidate matching text (with a first similarity of 0.3) is removed. After removing the feature vector D of the candidate matching text, the candidate matching text corresponding to the feature vector B of the candidate matching text ranked first among the sorted feature vectors B, A, C of multiple candidate matching texts is determined as the target matching text of the query text. By using the target similarity threshold to process the feature vectors of multiple sorted candidate matching texts and determining the target matching text of the query text based on the processing result, candidate matching texts irrelevant to the query text can be removed and the target matching text strongly relevant to the query text can be screened out, thereby improving the text matching accuracy.

[0044] The text matching method according to one or more embodiments of the present application can determine the feature vectors of multiple candidate matching texts from the vector database of candidate texts based on the first similarity between the feature vectors of the candidate texts in the vector database of candidate texts and the first feature vector of the query text, sort the feature vectors of the multiple candidate matching texts based on the second similarity between the second feature vector of the query text and the feature vectors of the multiple candidate matching texts, and use the target similarity threshold to determine the target matching text of the query text. Thus, vector search can be performed through the first similarity to improve text matching efficiency, and at the same time, the feature vectors of the multiple candidate matching texts can be sorted through the second similarity to optimize the vector search result based on the first similarity according to text semantics, so as to make full use of the semantic information of the query text and the multiple candidate matching texts, improve text matching accuracy while improving text matching efficiency, and save labor and material costs. The text matching method according to one or more embodiments of the present application can be applied to enterprise compliance inspections, for example, in the compliance inspection work of financial institutions to match daily task items according to inspection items, automatically parse and understand the text content of inspection items and task items, so as to achieve accurate and rapid matching of the two, and save labor and material costs.

[0045] Figure 2 The flowchart of a method for matching task items according to inspection items according to an embodiment of the present application is shown. Figure 2 The method shown in can be applied to the compliance inspection work of financial institutions to match daily task items according to inspection items.

[0046] As Figure 2 As shown in, in step S201, multiple first feature vectors associated with the inspection item are obtained.

[0047] Optionally, in step S201, the inspection item can be segmented to obtain text units of the inspection item, the text units of the inspection item can be replaced with synonymous text units of the text units of the inspection item to obtain multiple synonymous inspection texts corresponding to the inspection item, and the multiple synonymous inspection texts can be encoded to obtain multiple first feature vectors associated with the inspection item. Exemplarily, the synonymous text units of the text units of the inspection item can be obtained from a thesaurus.

[0048] In one embodiment, the inspection items in the compliance inspection of a financial institution may include the inspection items issued by the State Financial Supervision and Administration Institution for the compliance review of the financial institution's development strategy, internal norms, new product and new business plans, and major decision-making matters. It can formulate the basic framework and specific requirements for the compliance management of financial institutions, such as including but not limited to aspects such as compliance management architecture, responsibility assignment, compliance review, reporting mechanism, supervision and management, and legal liability. Exemplarily, the inspection items in the compliance inspection of a financial institution may include the key control points of activities such as information system planning, construction, operation and maintenance, and emergency response, which can be used to judge the security of information system operation, the compliance of information system construction, and the application performance of information systems, etc.

[0049] As an example, the inspection item can be "whether to review the functions and permissions of the risk management system". After performing word segmentation on this inspection item, the text units of this inspection item may include "whether", "review", "risk", "risk management", "management", "management system", "system", "of", "function", "and", "permission". Exemplarily, the text unit "review" of the inspection item can be replaced with its synonymous text unit "inspect", and the text unit "risk management" of the inspection item can be replaced with its synonymous text unit "risk control". Exemplarily, the inspection item can be segmented and synonymously replaced to expand each inspection item into multiple (for example, 3) synonymous inspection texts, and the multiple synonymous inspection texts are encoded to obtain multiple first feature vectors associated with the inspection item, increasing the subsequent vector search range, ensuring the integrity of the subsequent vector search, avoiding the problem of missing sentence meaning of the inspection item caused by text vectorization, and improving the vector search recall rate.

[0050] In step S203, based on the first similarity between the feature vectors of the candidate task items in the task item database and each of the multiple first feature vectors of the inspection items, determine the feature vectors of the multiple candidate matching task items that match each first feature vector.

[0051] In one embodiment, multiple daily task items to be matched with the inspection items can be obtained, and the multiple daily task types are vectorized to obtain a task item database. Exemplarily, the daily task items can be obtained from the laws and regulations involved in the various business processes of the financial institution. It should be noted that the daily task items can be understood as the matters that the financial institution actively executes daily according to the requirements in the laws and regulations. For example, according to the provisions of Chapter 5 of the "Guidelines for Information Technology Management of Futures Companies", the daily task item can be "whether to establish an effective mechanism to ensure that the headquarters understands the operation of each business department".

[0052] In one embodiment, the similarity distance between the feature vector of a candidate task item in the task item database and each of the multiple first feature vectors of the inspection item can be calculated, and this similarity distance can be used as the first similarity to determine, in the task item database, the feature vectors of multiple candidate matching task items that match each first feature vector. Exemplarily, the cosine similarity between the feature vector of a candidate task item in the task item database and each of the multiple first feature vectors of the inspection item can be calculated as the first similarity.

[0053] Optionally, in step S203, for each first feature vector of the inspection item, the vectors of a predetermined number of candidate task items can be selected from the task item database in descending order of the first similarity as the feature vectors of multiple candidate matching task items that match this first feature vector. For example, for each of the multiple first feature vectors associated with the inspection item, the vectors of the top five candidate task items in terms of the first similarity between their feature vectors and the feature vectors of the candidate task items in the task item database can be selected from the task item database as the feature vectors of multiple candidate matching task items that match this first feature vector.

[0054] In step S205, the feature vectors of multiple candidate matching task items are sorted based on the second similarity between the second feature vector of the inspection item and the feature vector of each candidate matching task item among the feature vectors of multiple candidate matching task items.

[0055] Optionally, in step S205, the inspection item can be directly encoded to obtain the second feature vector of this inspection item. Exemplarily, the BERT model can be used to encode the inspection item to obtain the second feature vector of this inspection item, and the feature vectors of the multiple candidate matching task items are sorted in descending order of the second similarity between the second feature vector of the inspection item and the feature vector of each candidate matching task item among the feature vectors of multiple candidate matching task items. By sorting the feature vectors of multiple candidate matching task items, the vector search result based on the first similarity can be optimized according to the text semantics, so as to make full use of the semantic information of the inspection item and multiple candidate matching task items, and improve the matching accuracy of the target matching task item while improving the text matching efficiency.

[0056] In step S207, the target similarity threshold is used to process the sorted feature vectors of multiple candidate matching task items, and the target matching task item of the inspection item is determined based on the processing result.

[0057] Optionally, in step S207, one or more feature vectors of candidate matching task items can be removed from the sorted feature vectors of multiple candidate matching task items based on a target similarity threshold, and the target matching task item of the inspection item can be determined based on the sorted feature vectors of multiple candidate matching task items after removing one or more feature vectors of candidate matching task items. In one embodiment, a first similarity corresponding to the feature vector of each sorted candidate matching task item can be determined, one or more feature vectors of candidate matching task items with a first similarity less than the target similarity threshold can be removed from the sorted feature vectors of multiple candidate matching task items, and among the sorted feature vectors of multiple candidate matching task items after removing one or more feature vectors of candidate matching task items, the candidate matching task item corresponding to the feature vector of the candidate matching task item ranked first can be determined as the target matching task item of the inspection item.

[0058] The text matching method according to one or more embodiments of the present application can be applied to enterprise compliance inspections, for example, to the compliance inspection work of matching daily task items according to inspection items in financial institutions, automatically parsing and understanding the text content of inspection items and task items, so as to achieve accurate and rapid matching of the two, saving manpower and material costs.

[0059] Figure 3 The schematic block diagram of a text matching device according to one or more embodiments of the present application is shown.

[0060] As Figure 3 shown in, the text matching device 300 includes a memory 310, a processor 320, and a computer program 330 stored on the memory 310 and executable on the processor 320. The processor 320 runs the computer program 330 to implement the text matching method according to one aspect of the present application.

[0061] In addition, the present application can also be implemented as a non-transitory computer-readable storage medium, in which a program for causing a computer to execute the text matching method according to one aspect of the present application is stored.

[0062] Here, as the non-transitory computer-readable storage medium, various non-transitory computer-readable storage media such as disk types (e.g., magnetic disks, optical disks, etc.), card types (e.g., memory cards, optical cards, etc.), semiconductor memory types (e.g., ROM, non-volatile memories, etc.), tape types (e.g., magnetic tapes, cassette tapes, etc.) can be adopted.

[0063] In applicable cases, various embodiments provided by this application can be implemented using hardware, software, or a combination of hardware and software. Moreover, in applicable cases, without departing from the scope of this application, the various hardware components and / or software components described herein can be combined into composite components including software, hardware, and / or both. In applicable cases, without departing from the scope of this application, the various hardware components and / or software components described herein can be divided into sub-components including software, hardware, or both. Additionally, in applicable cases, it is contemplated that software components can be implemented as hardware components and vice versa.

[0064] The software according to this application (such as program code and / or data) can be stored on one or more non-transitory computer-readable storage media. It is also contemplated that one or more general-purpose or special-purpose computers and / or computer systems, either networked and / or otherwise, can be used to implement the software identified herein. In applicable cases, the order of the various steps described herein can be changed, combined into composite steps, and / or divided into sub-steps to provide the features described herein.

[0065] The embodiments and examples presented herein are provided to best illustrate the embodiments in accordance with this application and its specific applications, and thereby enable those skilled in the art to implement and use this application. However, those skilled in the art will know that the above description and examples are provided for the sake of illustration and example only. The presented description is not intended to cover all aspects of this application or to limit this application to the precise form disclosed.

Claims

1. A method for text matching, characterized in that, The method includes the following steps: Obtain a plurality of first feature vectors associated with the query text; Based on the first similarity between the feature vectors of the candidate texts in the vector database of the candidate texts and each of the plurality of first feature vectors of the query text, determine, from the vector database of the candidate texts, the feature vectors of a plurality of candidate matching texts that match each of the first feature vectors; Based on the second similarity between the second feature vector of the query text and the feature vector of each of the plurality of candidate matching texts among the feature vectors of the plurality of candidate matching texts, sort the feature vectors of the plurality of candidate matching texts; And Use a target similarity threshold to process the sorted feature vectors of the plurality of candidate matching texts and determine the target matching text of the query text based on the processing result, where obtaining a plurality of first feature vectors associated with the query text includes: Perform word segmentation on the query text to obtain query text units of the query text; Replace the query text units with synonymous text units of the query text units to obtain a plurality of synonymous query texts corresponding to the query text; and Encode the plurality of synonymous query texts to obtain a plurality of first feature vectors associated with the query text.

2. The method according to claim 1, wherein the vector database of the candidate texts is obtained by: Obtain a plurality of candidate texts to be matched with the query text; and Perform vectorization processing on the plurality of candidate texts to obtain the vector database of the candidate texts.

3. The method according to claim 1, wherein based on the first similarity between the feature vectors of the candidate texts in the vector database of the candidate texts and each of the plurality of first feature vectors of the query text, determining, from the vector database of the candidate texts, the feature vectors of a plurality of candidate matching texts that match each of the first feature vectors includes: Select, in descending order of the first similarity, a predetermined number of vector of candidate texts from the vector database of the candidate texts as the feature vectors of a plurality of candidate matching texts that match each of the first feature vectors.

4. The method according to claim 1, wherein the target similarity threshold is determined by: Obtain a plurality of inspection item texts and label data of the candidate texts in the vector database of the candidate texts; Determine, from the vector database of the candidate texts, the feature vectors of a plurality of matching texts that match the feature vectors of each of the inspection item texts; and Determine the target similarity threshold based at least on the label data of the candidate texts and the feature vectors of a plurality of matching texts that match the feature vectors of each of the inspection item texts.

5. The method according to claim 4, wherein determining the target similarity threshold based at least on the label data of the candidate texts and the feature vectors of a plurality of matching texts that match the feature vectors of each of the inspection item texts includes: Determine a plurality of similarities between the feature vectors of the plurality of inspection item texts and the feature vectors of a plurality of matching texts that match the feature vectors of each of the inspection item texts; Use one or more of the multiple similarities as a preset similarity threshold; Use the preset similarity threshold to determine a mismatch rate based on the label data of the candidate text and the feature vectors of multiple matching texts that match the feature vectors of each inspection item text; And Determine the target similarity threshold based on the mismatch rate.

6. The method according to claim 5, wherein determining the target similarity threshold based on the mismatch rate includes: Determine the preset similarity threshold that makes the mismatch rate less than the mismatch rate threshold as the target similarity threshold.

7. The method according to claim 1, wherein sorting the feature vectors of the multiple candidate matching texts based on the second similarity between the second feature vector of the query text and the feature vector of each candidate matching text among the feature vectors of the multiple candidate matching texts includes: Encode the query text to obtain a second feature vector of the query text; Determine the second similarity between the second feature vector of the query text and the feature vector of each candidate matching text among the feature vectors of the multiple candidate matching texts; and Sort the feature vectors of the multiple candidate matching texts in descending order of the second similarity.

8. The method according to claim 1, wherein using the target similarity threshold to process the feature vectors of the sorted multiple candidate matching texts and determining the target matching text of the query text based on the processing result includes: Remove the feature vectors of one or more candidate matching texts from the feature vectors of the sorted multiple candidate matching texts based on the target similarity threshold; And Determine the target matching text of the query text based on the feature vectors of the sorted multiple candidate matching texts after removing the feature vectors of one or more candidate matching texts.

9. The method according to claim 8, wherein the method further includes: Determine the first similarity corresponding to the feature vector of each sorted candidate matching text; Remove the feature vectors of one or more candidate matching texts from the feature vectors of the sorted multiple candidate matching texts whose first similarity is less than the target similarity threshold; And Among the feature vectors of the sorted multiple candidate matching texts after removing the feature vectors of one or more candidate matching texts, determine the candidate matching text corresponding to the feature vector of the candidate matching text ranked first as the target matching text of the query text.

10. A text matching device, characterized in that, The apparatus includes: A memory; A processor coupled to the memory; and A computer program stored on the memory and running on the processor, the running of the computer program causes the execution of the text matching method according to any one of claims 1-9.

11. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium includes instructions that, when running, execute the text matching method according to any one of claims 1-9.

12. A computer program product, characterized in that, The computer program product includes instructions that, when executed by a processor, implement the text matching method according to any one of claims 1-9.

Citation Information

Patent Citations

  • A method for implementing a question answering system based on a question-answer pair

    CN109271505A

  • Text retrieval method and device

    CN116578693A

  • FAQ intelligent question-answering method and system in financial field

    CN116628146A

  • Text retrieval method and device, electronic equipment and readable storage medium

    CN117149956A

  • Malicious encrypted traffic detection method and device, storage medium and electronic equipment

    CN117579379A