Text fuzzy matching method and device and computer readable storage medium

Through the method of word segmentation processing and editing distance redundancy setting, the problem of fuzzy matching inaccurate caused by ignoring the order of words in the prior art is solved, and efficient and accurate text fuzzy matching is achieved.

CN120031031APending Publication Date: 2025-05-23中国邮政储蓄银行股份有限公司
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510064700.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

Existing text fuzzy matching methods ignore word order, resulting in inaccurate search results, and high-precision matching requires longer search time when processing large amounts of text data, and increasing speed may sacrifice accuracy.

Method used

By using word segmentation tool to process the preset statements, a fuzzy word collection is obtained, and the edit distance redundancy is set according to the word part-of-speech, a fuzzy search statement is constructed, and the text fragment matching the preset statement is inputted to the full-text search engine for searching, and a text fragment matching the preset statement is identified.

Benefits of technology

It realizes efficient and accurate text fuzzy matching while taking into account word order and text deformation, improving the accuracy of search results and reducing search time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120031031A_ABST
    Figure CN120031031A_ABST
Patent Text Reader

Abstract

The invention provides a text fuzzy matching method and device and a computer readable storage medium. According to the scheme, firstly, a word segmentation tool is used for carrying out word segmentation processing on a preset statement, and a fuzzy word set composed of multiple words is obtained; then, corresponding editing distance redundancy is set for each word in the fuzzy word set according to the part-of-speech of the words, the editing distance redundancy allows each word to change in a corresponding preset deformation range, and the original sequence relation among the multiple words is not changed after each word is deformed; and finally, constructing a fuzzy search statement according to the editing distance redundancy, inputting the fuzzy search statement into a full-text search engine for retrieval, and identifying a text fragment matched with a preset statement. According to the scheme, the problem of how to efficiently and accurately perform text fuzzy matching under the condition of considering a word sequence and text deformation is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of natural language processing, and in particular to a text fuzzy matching method, device, computer-readable storage medium and electronic device. Background Art

[0002] Text fuzzy matching requires the ability to process texts with similar semantics but different expressions, which is mainly reflected in the changes in word form, synonym replacement, sentence structure adjustment, etc.

[0003] Although existing text fuzzy matching methods can handle deformation and structural adjustments in text, they often ignore the order of words in the sentence after word segmentation, which can lead to inaccurate search results because the meaning of a sentence often depends on the order of words.

[0004] In addition, when processing large amounts of text data, high-precision fuzzy matching often requires longer search times, while increasing search speed may sacrifice matching accuracy. Summary of the invention

[0005] The main purpose of the present application is to provide a text fuzzy matching method, device, computer-readable storage medium and electronic device to at least solve the problem of how to perform text fuzzy matching efficiently and accurately while considering word order and text deformation.

[0006] In order to achieve the above-mentioned purpose, according to one aspect of the present application, a text fuzzy matching method is provided, comprising: using a word segmentation tool to perform word segmentation processing on a preset sentence to obtain a fuzzy word set consisting of multiple words; setting a corresponding edit distance redundancy for each word in the fuzzy word set according to the word's part of speech, wherein the edit distance redundancy allows each word to change within a corresponding preset deformation range, and the original order relationship between the multiple words remains unchanged after the deformation of each word; constructing a fuzzy search statement according to the edit distance redundancy, and inputting the fuzzy search statement into a full-text search engine for retrieval, and identifying a text fragment that matches the preset sentence.

[0007] Optionally, before using a word segmentation tool to perform word segmentation processing on a preset sentence to obtain a fuzzy word set consisting of multiple words, the method also includes: converting a non-text format to-be-matched file into a text format, and storing the converted to-be-matched file in the full-text search engine, and constructing a search database containing the converted to-be-matched file; using the indexing mechanism of the full-text search engine, establishing an inverted index for each text in the converted to-be-matched file, wherein the inverted index is used to record the text where the target word is located and the position information in the text.

[0008] Optionally, a corresponding edit distance redundancy is set for each word in the fuzzy word set according to the part of speech of the word, including: performing part-of-speech tagging on each word in the fuzzy word set, and performing word classification processing based on the part of speech of the word to obtain multiple sub-word sets, each sub-word set corresponding to a part of speech; determining the weight of each word in the sub-word set according to the part of speech of each sub-word set, wherein the weight reflects the degree of influence of the word on the matching result of text fuzzy matching; and determining the edit distance redundancy of each word according to the weight of each word.

[0009] Optionally, the edit distance redundancy of each of the words is determined according to the weight of each of the words, including: when the weight of the word is greater than or equal to a preset weight, setting the edit distance redundancy of the word to be within a first redundancy range; when the weight of the word is less than the preset weight, setting the edit distance redundancy of the word to be within a second redundancy range, wherein the minimum value of the second redundancy range is greater than the maximum value of the first redundancy range.

[0010] Optionally, constructing a fuzzy search statement according to the edit distance redundancy includes: arranging the words in the fuzzy word set in the order of the original sentence, and adding deformation range marks allowed by the edit distance redundancy before and after each word to obtain the fuzzy search statement.

[0011] Optionally, the fuzzy search statement is input into a full-text search engine for retrieval to identify a text segment that matches the preset statement, including: inputting the fuzzy search statement into the full-text search engine for retrieval, calculating the minimum edit distance between each word in each retrieved text segment and the corresponding word of the preset statement; accumulating the minimum edit distances of all words in each of the text segments to obtain the cumulative edit distance of the text segment, and comparing the cumulative edit distance with a preset threshold; when the cumulative edit distance is less than or equal to the preset threshold, determining that the text segment is a text segment that matches the preset statement.

[0012] Optionally, calculating the minimum edit distance between each word in each retrieved text segment and a corresponding word in the preset sentence comprises: using the formula Calculate the minimum edit distance between each word in each retrieved text segment and the corresponding word of the preset sentence, where sim is the minimum edit distance, and Represents the string s 1 and the string s 2 Length, Indicates that the string s 1 Converted into the string s 2 Minimum number of operations required.

[0013] According to another aspect of the present application, a text fuzzy matching device is provided, comprising: a word segmentation unit, used to use a word segmentation tool to perform word segmentation processing on a preset sentence to obtain a fuzzy word set consisting of multiple words; a setting unit, used to set a corresponding edit distance redundancy for each word in the fuzzy word set according to the word's part of speech, wherein the edit distance redundancy allows each word to change within a corresponding preset deformation range, and the original order relationship between the multiple words remains unchanged after the deformation of each word; an identification unit, used to construct a fuzzy search statement according to the edit distance redundancy, and input the fuzzy search statement into a full-text search engine for retrieval, and identify a text segment matching the preset statement.

[0014] According to another aspect of the present application, a computer-readable storage medium is provided, wherein the computer-readable storage medium includes a stored program, wherein when the program is executed, the device where the computer-readable storage medium is located is controlled to execute any one of the text fuzzy matching methods.

[0015] According to another aspect of the present application, an electronic device is provided, comprising: one or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include a method for executing any one of the text fuzzy matching methods.

[0016] The technical solution of the present application is applied, and a word segmentation tool is used to perform word segmentation processing on a preset sentence to obtain a fuzzy word set consisting of multiple words; a corresponding edit distance redundancy is set for each word in the fuzzy word set according to the word's part of speech, wherein the edit distance redundancy allows each word to change within a corresponding preset deformation range, and the original order relationship between multiple words after the deformation remains unchanged; a fuzzy search statement is constructed according to the edit distance redundancy, and the fuzzy search statement is input into a full-text search engine for retrieval to identify text fragments that match the preset statement. This solution solves the problem of how to efficiently and accurately perform text fuzzy matching while considering word order and text deformation. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The drawings constituting part of the present application are used to provide a further understanding of the present application. The exemplary embodiments and descriptions of the present application are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0018] Figure 1A hardware structure block diagram of a mobile terminal for executing a text fuzzy matching method provided in an embodiment of the present application is shown;

[0019] Figure 2 A flowchart of a text fuzzy matching method provided according to an embodiment of the present application is shown;

[0020] Figure 3 A schematic diagram of a specific text fuzzy matching method provided according to an embodiment of the present application is shown;

[0021] Figure 4 A structural block diagram of a text fuzzy matching device provided according to an embodiment of the present application is shown.

[0022] The above drawings include the following reference numerals:

[0023] 102, processor; 104, memory; 106, transmission device; 108, input and output devices. DETAILED DESCRIPTION

[0024] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0025] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present application.

[0026] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present application described here. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0027] As introduced in the background technology, although the existing technology can handle deformation and structural adjustment in the text, it often ignores the order of words in the sentence after word segmentation, which will lead to inaccurate search results. In order to solve the problem of how to perform text fuzzy matching efficiently and accurately while considering word order and text deformation, the embodiments of the present application provide a text fuzzy matching method, device, computer-readable storage medium and electronic device.

[0028] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the accompanying drawings in the embodiments of the present invention.

[0029] The method embodiments provided in the embodiments of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Taking running on a mobile terminal as an example, Figure 1 1 is a hardware structure block diagram of a mobile terminal of a text fuzzy matching method according to an embodiment of the present invention. Figure 1 As shown, the mobile terminal may include one or more ( Figure 1 Only one is shown in the figure) a processor 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data, wherein the mobile terminal may also include a transmission device 106 and an input / output device 108 for communication functions. It can be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the mobile terminal. Figure 1 More or fewer components as shown, or with Figure 1 Different configurations are shown.

[0030] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the display method of device information in the embodiment of the present invention. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, the above method is implemented. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 may further include a memory remotely arranged relative to the processor 102, and these remote memories can be connected to the mobile terminal via a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof. The transmission device 106 is used to receive or send data via a network. The above-mentioned specific examples of the network may include a wireless network provided by a communication provider of the mobile terminal. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, referred to as NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0031] In this embodiment, a text fuzzy matching method running on a mobile terminal, a computer terminal or a similar computing device is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0032] Figure 2 is a flowchart of a text fuzzy matching method according to an embodiment of the present application. Figure 2 As shown, the method comprises the following steps:

[0033] Step S201, using a word segmentation tool to perform word segmentation processing on a preset sentence to obtain a fuzzy word set consisting of multiple words;

[0034] Specifically, preset sentence refers to the target sentence predefined when performing text fuzzy matching, which is usually a specific sentence that needs to be found or compared in a large amount of text materials. Word segmentation is a basic task in natural language processing (Natural Language Processing, NLP), and its goal is to cut a continuous text sequence into meaningful words or phrases. In Chinese, due to the lack of clear word separators, word segmentation is particularly important to ensure the accuracy of subsequent semantic analysis, retrieval and other operations. Commonly used word segmentation tools include jieba, THULAC, etc., which use statistical models and corpora to effectively process the word segmentation tasks of text. In the embodiment of the present application, jieba word segmentation can be used to divide the preset sentence into multiple words to obtain a fuzzy word set consisting of multiple words.

[0035] Step S202, setting a corresponding edit distance redundancy for each of the above words in the above fuzzy word set according to the part of speech of the word, wherein the above edit distance redundancy allows each of the above words to change within a corresponding preset deformation range, and the original order relationship between the plurality of the above words remains unchanged after the deformation of each of the above words;

[0036] Specifically, part of speech refers to the grammatical function of words in a language, such as nouns (indicating people, things, places, etc.), verbs (indicating actions, states), adjectives (describing the properties of people or things), etc. Different parts of speech play different roles and importance in sentences, so when performing fuzzy matching, their edit distance redundancy should also be different.

[0037] Edit distance redundancy refers to the degree of allowed word deformation, that is, the difference between a word and its original form that can be tolerated during the retrieval process. For example, a word can accept the insertion, deletion or replacement of a certain number of characters and still be considered a match. The setting of edit distance redundancy is directly related to the part of speech of the word. In the set of fuzzy words obtained by word segmentation, each word is assigned a different edit distance redundancy according to its part of speech. That is to say, core words (such as key nouns and verbs) have a lower edit distance redundancy because the position and meaning of core words in the sentence are crucial, and too much deformation is not allowed to ensure the accuracy of the match; while auxiliary words (such as some adjectives and adverbs) may have a higher edit distance redundancy because auxiliary words may have more variations in expression, and at the same time, the deformation of auxiliary words has less impact on the overall meaning of the sentence.

[0038] When constructing fuzzy query statements, it is crucial to retain the original order relationship between words. Even if word deformation is allowed, it is necessary to ensure that the deformed words still appear in the order of the original sentence as much as possible. This is because the order relationship of words directly affects the meaning of the sentence. In addition, maintaining the order of words helps the matching algorithm understand the context in which the words appear, thereby more accurately judging the relevance between the text fragment and the preset sentence.

[0039] Step S203: construct a fuzzy search statement according to the edit distance redundancy, and input the fuzzy search statement into a full-text search engine for retrieval to identify text segments matching the preset statement.

[0040] Specifically, since the edit distance redundancy allows words to be transformed within a preset range but keep the original order unchanged, when constructing a fuzzy search statement, it is necessary to retain the original order of the words in the fuzzy word set, that is, the fuzzy search statement will be arranged in the order in which the words appear in the preset statement. A fuzzy query statement includes all the words in the preset statement and the corresponding deformation range indication.

[0041] The constructed fuzzy search statement is input into a full-text search engine for retrieval, such as Whoosh or other similar search engines. The embodiment of the present application may use the Whoosh full-text search engine. A full-text search engine is a software that can index and search text data and can handle complex queries, including fuzzy matching. When the full-text search engine finds a series of words whose order and variation range match the fuzzy search statement, it is considered that a matching text segment has been found.

[0042] Through this embodiment, first, a word segmentation tool is used to perform word segmentation processing on a preset sentence to obtain a fuzzy word set consisting of multiple words; then, a corresponding edit distance redundancy is set for each word in the fuzzy word set according to the word's part of speech, wherein the edit distance redundancy allows each word to change within a corresponding preset deformation range, and the original order relationship between multiple words remains unchanged after each word is deformed; finally, a fuzzy search statement is constructed according to the edit distance redundancy, and the fuzzy search statement is input into a full-text search engine for retrieval, and a text segment matching the preset sentence is identified, thereby solving the problem of how to efficiently and accurately perform text fuzzy matching while considering word order and text deformation.

[0043] In one embodiment of the present application, before using a word segmentation tool to perform word segmentation processing on a preset sentence to obtain a fuzzy word set consisting of multiple words, the method further includes: converting a non-text format file to be matched into a text format, and storing the converted file to be matched in the full-text search engine, and constructing a search database containing the converted file to be matched; using the indexing mechanism of the full-text search engine, establishing an inverted index for each text in the converted file to be matched, and the inverted index is used to record the text where the target word is located and the position information in the text.

[0044] Specifically, non-text formats usually refer to file formats such as PDF, Word documents, pictures or scans. Although these files contain text information, their formats do not directly support text retrieval. Therefore, before performing text fuzzy matching, these non-text files to be matched need to be converted into text formats that can be directly processed by full-text search engines. This process can be achieved using optical character recognition (OCR) technology. OCR technology can recognize text in images or scans and convert it into an editable text format. For example, if the file to be matched is a picture or PDF document, the OCR tool will analyze the image in the document, recognize the characters in it, and then convert it into a string in text format.

[0045] The text converted by OCR technology is stored in the full-text search engine. The full-text search engine is a search system designed for text data, which can handle the indexing and retrieval of large amounts of text data. Storing text in the full-text search engine indexes the text content, making it the basis for subsequent text fuzzy matching queries. When storing, each document or material will be processed separately to ensure the independence and retrievability of each text content.

[0046] In the full-text search engine, all converted text materials will be integrated into a search database. The search database is a collection of all texts to be retrieved. Through the indexing mechanism, the retrieval of text content becomes fast and efficient. Building a search database is the core task of the full-text retrieval system, which converts massive document content into a structured index to facilitate subsequent query operations.

[0047] An inverted index is an optimized index data structure used to quickly locate documents containing specific words. Unlike a forward index (which builds an index based on document content), an inverted index points to documents based on words, i.e., the key of the index is the word, and the value is a list of all documents containing the word. By using the indexing mechanism of a full-text search engine to build an inverted index for each document, the specific text where each target word is located, as well as the position information of the word in the text, can be recorded. This allows the search engine to quickly locate text fragments containing potential matching words when performing fuzzy matching searches without having to traverse the entire document, significantly improving search speed and efficiency.

[0048] The above text processing of non-text format files, storage in full-text search engines, construction of search databases, and creation of inverted indexes are all for the purpose of more accurately and quickly locating and retrieving text fragments similar to preset sentences in a large number of documents when performing text fuzzy matching. Through the above processing, the full-text search engine can better understand the document structure and word position, so that subsequent text fuzzy matching has higher accuracy and wider coverage.

[0049] In another embodiment of the present application, a corresponding edit distance redundancy is set for each of the above-mentioned words in the above-mentioned fuzzy word set according to the part of speech of the word, including: marking the part of speech of each of the above-mentioned words in the above-mentioned fuzzy word set, and performing word classification processing based on the part of speech of the above-mentioned words to obtain multiple sub-word sets, each of the above-mentioned sub-word sets corresponding to a part of speech; determining the weight of each word in the above-mentioned sub-word set according to the part of speech of each of the above-mentioned sub-word sets, wherein the above-mentioned weight reflects the influence of the above-mentioned word on the matching result of text fuzzy matching; and determining the above-mentioned edit distance redundancy of each of the above-mentioned words according to the weight of each of the above-mentioned words.

[0050] Specifically, part-of-speech tagging of each word in the fuzzy word set is a key step in understanding the role of the word in the sentence. Part-of-speech tagging is a basic task in natural language processing, which assigns a label to each word to indicate its grammatical function in the sentence, such as noun, verb, adjective, etc. Through part-of-speech tagging, the word set can be further subdivided into multiple sub-word sets, each of which contains words of the same part of speech, for example, a noun subset, a verb subset, etc.

[0051] After the words are classified, the next task is to determine the weight of the words, that is, to measure the relative importance of the words in the fuzzy matching of the text. The weight of a word is usually based on the role it plays in the sentence and its contribution to the meaning of the sentence. For example, some key nouns and verbs can be given higher weights because these words are directly related to the meaning of the sentence, while some descriptive adverbs or prepositions can be given lower weights because these words play a relatively small role in expressing the core meaning.

[0052] Edit distance redundancy is the degree of deformation allowed for words during the matching process. Based on the weight of the words, different edit distance redundancies can be set for words of different parts of speech. Usually, words with higher weights, such as core words, have smaller edit distance redundancy, which means that these words require higher precision when matching and do not allow too much deformation. On the other hand, words with lower weights, such as auxiliary words, can have larger edit distance redundancy, allowing more deformation to increase the flexibility of matching. This mechanism ensures that while maintaining the core meaning of the sentence, subtle differences or recognition errors that may exist in the text material are also taken into account.

[0053] In practical applications, multiple factors can be considered when setting the edit distance redundancy, such as common variations of words: if words based on part of speech have common spelling variants, abbreviations or synonyms, these should be considered in the setting of edit distance redundancy. Complexity of text: if the text material to be matched contains a lot of professional terms or language in a specific field, the edit distance redundancy for these terms or words can be set lower to ensure the accuracy of the match.

[0054] In summary, the embodiments of the present application can finely adjust the redundancy of the edit distance according to the word's part of speech and importance, thereby improving the efficiency and accuracy of text fuzzy matching while ensuring the core meaning of the sentence. This is crucial for processing large amounts of text materials, especially unstructured documents converted by OCR.

[0055] In a specific embodiment, the edit distance redundancy of each of the above-mentioned words is determined according to the weight of each of the above-mentioned words, including: when the weight of the above-mentioned words is greater than or equal to the preset weight, the edit distance redundancy of the above-mentioned words is set to be within a first redundancy range; when the weight of the above-mentioned words is less than the above-mentioned preset weight, the edit distance redundancy of the above-mentioned words is set to be within a second redundancy range, wherein the minimum value of the above-mentioned second redundancy range is greater than the maximum value of the above-mentioned first redundancy range.

[0056] Specifically, the preset weight is a threshold used to distinguish the importance of words in the word set in text fuzzy matching. Usually, words with higher weights are key components in the sentence, such as core nouns and verbs. These words have a greater impact on the text matching results and should be treated with importance to maintain the accuracy of the core meaning; on the contrary, words with lower weights, such as some adjectives, adverbs or conjunctions, play an auxiliary role in the sentence and have less impact on the overall meaning. They can be given a larger deformation space to increase the flexibility of matching.

[0057] When the weight of a word is higher than or equal to the preset weight, the edit distance redundancy of the word will be set within the first redundancy range. The first redundancy range is usually small, which means that these words are less deformed when fuzzy matching is performed. For example, in a text about financial management, if the main keyword of the sentence is "financial products", then the keyword will be given a smaller edit distance redundancy, and too many characters will not be inserted, deleted or replaced to ensure that the semantics of this keyword in the matching results does not change. This is to protect the core meaning of the sentence from being destroyed by deformation during the fuzzy matching process.

[0058] When the weight of a word is lower than the preset weight, the edit distance redundancy of the word will be set within the second redundancy range. The minimum value of the second redundancy range is greater than the maximum value of the first redundancy range, which means that auxiliary words can withstand greater deformation during fuzzy matching, but the degree of deformation will not be enough to change the semantics. For example, the descriptive adjective "accurate" can be given a larger edit distance redundancy, allowing a certain degree of deformation, such as deformation to "precise", which helps to match synonymous expressions in text fragments more widely without deviating from the meaning of the original sentence.

[0059] When constructing fuzzy query statements, the full-text search engine will consider these variation ranges when querying, based on the weight of the words and the corresponding edit distance redundancy range. This involves intelligent adjustment of the query statement to ensure that both the accuracy of the words and the flexibility of the context can be considered when searching. For example, for words with high weights, the query will be more stringent, only searching for very close matches to the original words; for words with low weights, the query will be more relaxed, covering more semantically similar words with slightly changed forms. This operation achieves refined control of text fuzzy matching by dynamically adjusting the edit distance redundancy of words.

[0060] In another embodiment of the present application, a fuzzy search statement is constructed based on the above-mentioned edit distance redundancy, including: arranging the words in the above-mentioned fuzzy word set in the order of the original sentence, and adding deformation range marks allowed by the above-mentioned edit distance redundancy before and after each of the above-mentioned words to obtain the above-mentioned fuzzy search statement.

[0061] Specifically, when constructing a fuzzy search statement, it is important to maintain the original order of the words in the fuzzy word set, because the order of the words directly affects the semantics of the statement. Even if the words are the same, if the order changes, the meaning of the statement will change. In order to perform fuzzy matching in a full-text search engine, it is necessary to add deformation range marks allowed by the edit distance redundancy before and after each word. These marks can be special characters or codes that are used to indicate how the search engine can deform the word while maintaining the accuracy of the match. For example, if a word allows the deformation of one character (insertion, deletion or replacement), then when constructing a fuzzy search statement, corresponding marks such as "~" or "*" can be added to the appropriate positions before and after the word to form a format similar to "keyword~1keyword~1", which means to indicate to the search engine that the word allows the edit distance of one character to change, but it cannot exceed this range.

[0062] According to the words in the preset sentence and the edit distance redundancy set for each word, the process of constructing a fuzzy search sentence is as follows: Preserve order: Ensure that the order of the words in the search sentence is consistent with the order of the words in the preset sentence. Add tags: Add the deformation range tags allowed by the edit distance redundancy before and after each word, which can include instructions for inserting, deleting, and replacing characters. Combine sentences: Combine the marked words in order to form a complete fuzzy search sentence. This sentence contains all the words in the preset sentence and the deformation range.

[0063] This construction process ensures that the full-text search engine can take into account the range of word variations when performing text fuzzy matching, while maintaining the original order of words. Therefore, when processing large amounts of text materials, it can not only find text fragments that are highly similar to the preset sentences, but also handle small errors or synonyms in the text, thereby improving the accuracy and efficiency of matching.

[0064] In another embodiment of the present application, the fuzzy search statement is input into a full-text search engine for retrieval to identify a text segment that matches the preset statement, including: inputting the fuzzy search statement into the full-text search engine for retrieval, calculating the minimum edit distance between each word in each retrieved text segment and the corresponding word of the preset statement; accumulating the minimum edit distances of all words in each of the text segments to obtain the cumulative edit distance of the text segment, and comparing the cumulative edit distance with a preset threshold; when the cumulative edit distance is less than or equal to the preset threshold, determining that the text segment is a text segment that matches the preset statement.

[0065] Specifically, the constructed fuzzy search statement is input into the full-text search engine for retrieval. The fuzzy search statement contains all the words in the preset sentence, as well as the edit distance redundancy markers set before and after each word, which are used to indicate the range of deformation allowed by the search engine when matching. After the full-text search engine retrieves the text fragments, it calculates the minimum edit distance between each word in these fragments and the corresponding word in the preset sentence. The minimum edit distance is an indicator to measure the similarity between two strings, indicating the minimum number of single-character editing operations (including insertion, deletion, and replacement) required to convert from one string to another. By calculating the minimum edit distance, the similarity between the words in the text fragment and the words in the preset sentence can be quantified.

[0066] For each retrieved text segment, the full-text search engine will cumulatively calculate the minimum edit distance between all the words in it and the corresponding words of the preset sentence to obtain the cumulative edit distance of the text segment. The cumulative edit distance is a comprehensive indicator that reflects the similarity between the entire text segment and the preset sentence. If there are multiple words in a text segment that have edit distances with the words in the preset sentence, their edit distances will be added up to form the cumulative edit distance of the segment. Finally, the cumulative edit distance of each text segment is compared with a preset threshold. The preset threshold is a pre-set value used to determine whether the text segments are similar enough to be considered a match for the preset sentence. If the cumulative edit distance of a text segment is less than or equal to the preset threshold, then the text segment is considered to match the preset sentence. For example, if the preset threshold is 5 and the cumulative edit distance of a text segment is 4, then the text segment will be confirmed as a match because the similarity between the text segment and the preset sentence is within the allowed deformation range.

[0067] This method based on the cumulative calculation of the minimum edit distance and the comparison of the threshold ensures that even if there are small deformations or recognition errors in the text fragment, as long as the overall similarity meets the preset conditions, it can still be identified as a match, thereby improving the accuracy and efficiency of fuzzy matching. At the same time, by reasonably setting the preset threshold, the strictness and flexibility of the match can be balanced.

[0068] In a specific embodiment, calculating the minimum edit distance between each word in each retrieved text segment and the corresponding word of the above-mentioned preset sentence includes: using the formula Calculate the minimum edit distance between each word in each retrieved text segment and the corresponding word in the above preset sentence, where sim is the above minimum edit distance, and Represents the string s 1 and the string s 2 Length, Indicates that the above string s1 Converted into the above string s 2 Minimum number of operations required.

[0069] Specifically, in the full-text search and matching process, the formula It is used to calculate the similarity between each word in the text fragment and the corresponding word of the preset sentence. The minimum edit distance of each word will be calculated and adjusted according to the weight of the word and the redundancy of the edit distance. By calculating the edit distance, combined with the weight of the word and the allowed deformation range, the full-text search engine can comprehensively consider the exact matching and semantic similarity of the words when performing text fuzzy matching. This makes it possible to identify text fragments similar to the preset sentence even if there are OCR recognition errors, synonym replacements or slight sentence deformations in the text material. Not only can it intelligently identify similar content in the text, but it also avoids the strict limitations of completely accurate matching, making the text fuzzy matching method of the embodiment of the present application more practical and efficient in the processing of large amounts of text data.

[0070] The embodiments of the present application can be applied in a variety of scenarios, such as content review, academic duplication detection and public opinion monitoring, intelligent translation and language processing, etc. In the banking field, the embodiments of the present application can be applied in consumer rights protection review scenarios. A specific example will be shown below. Figure 3 It is a flowchart of a specific text fuzzy matching method, such as Figure 3 As shown, it includes key steps such as word segmentation to construct a fuzzy word set, setting the edit distance redundancy, and constructing a fuzzy search statement. The specific explanation is as follows:

[0071] If the review point is "This service promises that your account information is safe and will not share personal information with third parties without your consent", the default sentence is "This service promises that your account information is safe and will not share personal information with third parties without your consent."

[0072] First, the text content of the consumer rights protection review materials to be reviewed is converted into a searchable text format through OCR technology, and stored in the whoosh full-text search engine to build a retrieval library for subsequent fuzzy search.

[0073] The above review points are segmented using the Chinese word segmentation algorithm based on jieba, and the review points are decomposed into a set of fuzzy words: "this service", "commitment", "your", "account", "information", "security", "without", "your", "consent", "will not", "to", "third party", "institution", "share", "individual", "information".

[0074] Each word in the fuzzy word set is labeled. For example, "this service", "account", "information", "consent", "third party", "institution", "individual" and "information" are labeled as nouns; "commitment", "without", "will not", "to", and "share" are labeled as verbs, "your" is labeled as a pronoun, and "safe" is labeled as an adjective. Word classification is performed based on the word's part of speech to obtain multiple sub-word sets, each of which corresponds to a part of speech, such as a noun sub-set, a verb sub-set, a pronoun sub-set, and an adjective sub-set.

[0075] The weight of each word in the sub-word set is determined according to the part of speech of the sub-word set, wherein the weight reflects the influence of the above words on the matching results of the text fuzzy matching; generally, the information represented by nouns is more critical, and the weight of nouns can be set higher, while the information represented by adjectives is relatively less important, for example, the weight range of words in the noun sub-set is set to 7-9, the weight range of words in the adjective sub-set is set to 1-2, the weight range of words in the pronoun sub-set is set to 3-4, and the weight range of words in the verb sub-set is set to 5-6. For example, the weights of "this service", "account", "information", "consent", "third party", "institution", "individual" and "information" can be set to 7, 9, 8, 8, 7, 7, 8, 8 respectively; the weight of "safety" is set to 2; the weight of "your" is set to 3; the weights of "commitment", "without", "will not", "to", and "share" are set to 5, 5, 6, 5, 6 respectively.

[0076] Assuming the preset weight is 6, if the weight of a word is greater than or equal to 6, it means that the word is more critical, then the edit distance redundancy of the word with a weight greater than or equal to 6 can be set to a smaller range, such as 1 to 3, that is, the word is allowed to have a small range of deformation; if the weight of a word is less than 6, it means that the word is relatively not so important, then the edit distance redundancy of the word with a weight less than 6 can be set to a larger range, such as 4 to 7, that is, the word is allowed to have a large range of deformation, and there can be more deformation or replacement space. For example, the edit distance redundancy of "this service", "account", "information", "consent", "third party", "institution", "individual" and "information" can be set to 2, 1, 2, 2, 2, 3, 2, 2 respectively; the edit distance redundancy of "safety" can be set to 6; the edit distance redundancy of "your" can be set to 5; the edit distance redundancy of "commitment", "without", "will not", "to", and "share" can be set to 6, 5, 3, 5, 3 respectively.

[0077] Based on the redundancy of the edit distance, the fuzzy search statement constructed is: "~2This service~6promises~5that~5your~1account~2information~6is~6secure,~5and~5will~5not~5share~2personal~2information~3with~2third-party~3institutions~3without~5your~2consent."

[0078] The constructed fuzzy search statement is input into the whoosh full-text search engine for retrieval to find text fragments of consumer rights protection review materials that are similar to the preset review points.

[0079] For each retrieved text segment, such as "This service guarantees the confidentiality of customer account information and does not disclose private information to third-party entities without the customer's permission.", calculate the minimum edit distance between each word and the corresponding word of the preset review point. For example, calculate the minimum edit distance between "commitment" and "guarantee", "safety" and "confidentiality". Then add up the minimum edit distances of all words in each text segment to get the cumulative edit distance of the text segment. Assuming the preset threshold is 10, if the cumulative edit distance is less than or equal to 10, the text segment is considered to match the review point, which is the text fuzzy matching result.

[0080] Of course, the above embodiment is only an example of the present application, and the edit distance redundancy can be determined according to the actual situation. If it is a noun that has little influence on the preset sentence, a larger edit distance redundancy can be set. If an adjective has a greater influence on the preset sentence, a smaller edit distance redundancy can be set for the adjective.

[0081] The embodiment of the present application also provides a text fuzzy matching device. It should be noted that the text fuzzy matching device of the embodiment of the present application can be used to execute the text fuzzy matching method provided by the embodiment of the present application. The device is used to implement the above-mentioned embodiments and preferred implementation modes, and the descriptions thereof will not be repeated. As used below, the term "module" can implement a combination of software and / or hardware of a predetermined function. Although the device described in the following embodiments is preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceived.

[0082] The text fuzzy matching device provided in the embodiment of the present application is introduced below.

[0083] Figure 4 is a structural block diagram of a text fuzzy matching device according to an embodiment of the present application. Figure 4 As shown, the device includes a word segmentation unit 10, a setting unit 20 and a recognition unit 30.

[0084] The word segmentation unit 10 is used to perform word segmentation processing on the preset sentence using a word segmentation tool to obtain a fuzzy word set consisting of multiple words;

[0085] Specifically, preset sentence refers to the target sentence predefined when performing text fuzzy matching, which is usually a specific sentence that needs to be found or compared in a large amount of text materials. Word segmentation is a basic task in natural language processing (Natural Language Processing, NLP), and its goal is to cut a continuous text sequence into meaningful words or phrases. In Chinese, due to the lack of clear word separators, word segmentation is particularly important to ensure the accuracy of subsequent semantic analysis, retrieval and other operations. Commonly used word segmentation tools include jieba, THULAC, etc., which use statistical models and corpora to effectively process the word segmentation tasks of text. In the embodiment of the present application, jieba word segmentation can be used to divide the preset sentence into multiple words to obtain a fuzzy word set consisting of multiple words.

[0086] A setting unit 20 is used to set a corresponding edit distance redundancy for each of the above-mentioned words in the above-mentioned fuzzy word set according to the part of speech of the word, wherein the above-mentioned edit distance redundancy allows each of the above-mentioned words to change within a corresponding preset deformation range, and the original order relationship between the plurality of above-mentioned words remains unchanged after the deformation of each of the above-mentioned words;

[0087] Specifically, part of speech refers to the grammatical function of words in a language, such as nouns (indicating people, things, places, etc.), verbs (indicating actions, states), adjectives (describing the properties of people or things), etc. Different parts of speech play different roles and importance in sentences, so when performing fuzzy matching, their edit distance redundancy should also be different.

[0088] Edit distance redundancy refers to the degree of allowed word deformation, that is, the difference between a word and its original form that can be tolerated during the retrieval process. For example, a word can accept the insertion, deletion or replacement of a certain number of characters and still be considered a match. The setting of edit distance redundancy is directly related to the part of speech of the word. In the set of fuzzy words obtained by word segmentation, each word is assigned a different edit distance redundancy according to its part of speech. That is to say, core words (such as key nouns and verbs) have a lower edit distance redundancy because the position and meaning of core words in the sentence are crucial, and too much deformation is not allowed to ensure the accuracy of the match; while auxiliary words (such as some adjectives and adverbs) may have a higher edit distance redundancy because auxiliary words may have more variations in expression, and at the same time, the deformation of auxiliary words has less impact on the overall meaning of the sentence.

[0089] When constructing fuzzy query statements, it is crucial to retain the original order relationship between words. Even if word deformation is allowed, it is necessary to ensure that the deformed words still appear in the order of the original sentence as much as possible. This is because the order relationship of words directly affects the meaning of the sentence. In addition, maintaining the order of words helps the matching algorithm understand the context in which the words appear, thereby more accurately judging the relevance between the text fragment and the preset sentence.

[0090] The identification unit 30 is used to construct a fuzzy search statement according to the edit distance redundancy, and input the fuzzy search statement into a full-text search engine for retrieval, so as to identify a text segment matching the preset statement.

[0091] Specifically, since the edit distance redundancy allows words to be transformed within a preset range but keep the original order unchanged, when constructing a fuzzy search statement, it is necessary to retain the original order of the words in the fuzzy word set, that is, the fuzzy search statement will be arranged in the order in which the words appear in the preset statement. A fuzzy query statement includes all the words in the preset statement and the corresponding deformation range indication.

[0092] The constructed fuzzy search statement is input into a full-text search engine for retrieval, such as Whoosh or other similar search engines. The embodiment of the present application may use the Whoosh full-text search engine. A full-text search engine is a software that can index and search text data and can handle complex queries, including fuzzy matching. When the full-text search engine finds a series of words whose order and variation range match the fuzzy search statement, it is considered that a matching text segment has been found.

[0093] In one embodiment of the present application, the above-mentioned device further includes:

[0094] A conversion unit, used to convert the non-text format to-be-matched files into text format, store the converted to-be-matched files into the above-mentioned full-text search engine, and construct a search database containing the above-mentioned converted to-be-matched files;

[0095] The construction unit is used to use the indexing mechanism of the full-text search engine to create an inverted index for each text in the converted to-be-matched file, wherein the inverted index is used to record the text where the target word is located and the position information in the text.

[0096] Specifically, non-text formats usually refer to file formats such as PDF, Word documents, pictures or scans. Although these files contain text information, their formats do not directly support text retrieval. Therefore, before performing text fuzzy matching, these non-text files to be matched need to be converted into text formats that can be directly processed by full-text search engines. This process can be achieved using optical character recognition (OCR) technology. OCR technology can recognize text in images or scans and convert it into an editable text format. For example, if the file to be matched is a picture or PDF document, the OCR tool will analyze the image in the document, recognize the characters in it, and then convert it into a string in text format.

[0097] The text converted by OCR technology is stored in the full-text search engine. The full-text search engine is a search system designed for text data, which can handle the indexing and retrieval of large amounts of text data. Storing text in the full-text search engine indexes the text content, making it the basis for subsequent text fuzzy matching queries. When storing, each document or material will be processed separately to ensure the independence and retrievability of each text content.

[0098] In the full-text search engine, all converted text materials will be integrated into a search database. The search database is a collection of all texts to be retrieved. Through the indexing mechanism, the retrieval of text content becomes fast and efficient. Building a search database is the core task of the full-text retrieval system, which converts massive document content into a structured index to facilitate subsequent query operations.

[0099] An inverted index is an optimized index data structure used to quickly locate documents containing specific words. Unlike a forward index (which builds an index based on document content), an inverted index points to documents based on words, i.e., the key of the index is the word, and the value is a list of all documents containing the word. By using the indexing mechanism of a full-text search engine to build an inverted index for each document, the specific text where each target word is located, as well as the position information of the word in the text, can be recorded. This allows the search engine to quickly locate text fragments containing potential matching words when performing fuzzy matching searches without having to traverse the entire document, significantly improving search speed and efficiency.

[0100] In another embodiment of the present application, the above-mentioned setting unit includes:

[0101] A tagging module is used to tag the part of speech of each of the above-mentioned words in the above-mentioned fuzzy word set, and perform word classification processing based on the part of speech of the above-mentioned words to obtain multiple sub-word sets, and each of the above-mentioned sub-word sets corresponds to a part of speech;

[0102] A first determination module is used to determine the weight of each word in the above sub-word set according to the part of speech of each sub-word set, wherein the weight reflects the influence of the above word on the matching result of text fuzzy matching;

[0103] The second determining module is used to determine the edit distance redundancy of each of the above words according to the weight of each of the above words.

[0104] Specifically, part-of-speech tagging of each word in the fuzzy word set is a key step in understanding the role of the word in the sentence. Part-of-speech tagging is a basic task in natural language processing, which assigns a label to each word to indicate its grammatical function in the sentence, such as noun, verb, adjective, etc. Through part-of-speech tagging, the word set can be further subdivided into multiple sub-word sets, each of which contains words of the same part of speech, for example, a noun subset, a verb subset, etc.

[0105] After the words are classified, the next task is to determine the weight of the words, that is, to measure the relative importance of the words in the fuzzy matching of the text. The weight of a word is usually based on the role it plays in the sentence and its contribution to the meaning of the sentence. For example, some key nouns and verbs can be given higher weights because these words are directly related to the meaning of the sentence, while some descriptive adverbs or prepositions can be given lower weights because these words play a relatively small role in expressing the core meaning.

[0106] Edit distance redundancy is the degree of deformation allowed for words during the matching process. Based on the weight of the words, different edit distance redundancies can be set for words of different parts of speech. Usually, words with higher weights, such as core words, have smaller edit distance redundancy, which means that these words require higher precision when matching and do not allow too much deformation. On the other hand, words with lower weights, such as auxiliary words, can have larger edit distance redundancy, allowing more deformation to increase the flexibility of matching. This mechanism ensures that while maintaining the core meaning of the sentence, subtle differences or recognition errors that may exist in the text material are also taken into account.

[0107] In practical applications, multiple factors can be considered when setting the edit distance redundancy, such as common variations of words: if words based on part of speech have common spelling variants, abbreviations or synonyms, these should be considered in the setting of edit distance redundancy. Complexity of text: if the text material to be matched contains a lot of professional terms or language in a specific field, the edit distance redundancy for these terms or words can be set lower to ensure the accuracy of the match.

[0108] In a specific embodiment, the second determining module includes:

[0109] A first setting submodule, configured to set the edit distance redundancy of the word to within a first redundancy range when the weight of the word is greater than or equal to a preset weight;

[0110] The second setting submodule is used to set the edit distance redundancy of the above-mentioned word to within a second redundancy range when the weight of the above-mentioned word is less than the above-mentioned preset weight, wherein the minimum value of the above-mentioned second redundancy range is greater than the maximum value of the above-mentioned first redundancy range.

[0111] Specifically, the preset weight is a threshold used to distinguish the importance of words in the word set in text fuzzy matching. Usually, words with higher weights are key components in the sentence, such as core nouns and verbs. These words have a greater impact on the text matching results and should be treated with importance to maintain the accuracy of the core meaning; on the contrary, words with lower weights, such as some adjectives, adverbs or conjunctions, play an auxiliary role in the sentence and have less impact on the overall meaning. They can be given a larger deformation space to increase the flexibility of matching.

[0112] When the weight of a word is higher than or equal to the preset weight, the edit distance redundancy of the word will be set within the first redundancy range. The first redundancy range is usually small, which means that these words are less deformed when fuzzy matching is performed. For example, in a text about financial management, if the main keyword of the sentence is "financial products", then the keyword will be given a smaller edit distance redundancy, and too many characters will not be inserted, deleted or replaced to ensure that the semantics of this keyword in the matching results does not change. This is to protect the core meaning of the sentence from being destroyed by deformation during the fuzzy matching process.

[0113] When the weight of a word is lower than the preset weight, the edit distance redundancy of the word will be set within the second redundancy range. The minimum value of the second redundancy range is greater than the maximum value of the first redundancy range, which means that auxiliary words can withstand greater deformation during fuzzy matching, but the degree of deformation will not be enough to change the semantics. For example, the descriptive adjective "accurate" can be given a larger edit distance redundancy, allowing a certain degree of deformation, such as deformation to "precise", which helps to match synonymous expressions in text fragments more widely without deviating from the meaning of the original sentence.

[0114] When constructing fuzzy query statements, the full-text search engine will consider these variation ranges when querying, based on the weight of the words and the corresponding edit distance redundancy range. This involves intelligent adjustment of the query statement to ensure that both the accuracy of the words and the flexibility of the context can be considered when searching. For example, for words with high weights, the query will be more stringent, only searching for very close matches to the original words; for words with low weights, the query will be more relaxed, covering more semantically similar words with slightly changed forms. This operation achieves refined control of text fuzzy matching by dynamically adjusting the edit distance redundancy of words.

[0115] In another embodiment of the present application, the identification unit includes:

[0116] The arrangement module is used to arrange the words in the fuzzy word set according to the original sentence order, and add the deformation range mark allowed by the edit distance redundancy before and after each of the words to obtain the fuzzy search sentence.

[0117] Specifically, when constructing a fuzzy search statement, it is important to maintain the original order of the words in the fuzzy word set, because the order of the words directly affects the semantics of the statement. Even if the words are the same, if the order changes, the meaning of the statement will change. In order to perform fuzzy matching in a full-text search engine, it is necessary to add deformation range marks allowed by the edit distance redundancy before and after each word. These marks can be special characters or codes that are used to indicate how the search engine can deform the word while maintaining the accuracy of the match. For example, if a word allows the deformation of one character (insertion, deletion or replacement), then when constructing a fuzzy search statement, corresponding marks such as "~" or "*" can be added to the appropriate positions before and after the word to form a format similar to "keyword~1keyword~1", which means to indicate to the search engine that the word allows the edit distance of one character to change, but it cannot exceed this range.

[0118] According to the words in the preset sentence and the edit distance redundancy set for each word, the process of constructing a fuzzy search sentence is as follows: Preserve order: Ensure that the order of the words in the search sentence is consistent with the order of the words in the preset sentence. Add tags: Add the deformation range tags allowed by the edit distance redundancy before and after each word, which can include instructions for inserting, deleting, and replacing characters. Combine sentences: Combine the marked words in order to form a complete fuzzy search sentence. This sentence contains all the words in the preset sentence and the deformation range.

[0119] In another embodiment of the present application, the identification unit further includes:

[0120] A first calculation module, configured to input the fuzzy search sentence into the full-text search engine for retrieval, and calculate the minimum edit distance between each word in each retrieved text segment and the corresponding word of the preset sentence;

[0121] A second calculation module is used to accumulate and calculate the minimum edit distances of all words in each of the text segments to obtain a cumulative edit distance of the text segments, and compare the cumulative edit distance with a preset threshold;

[0122] The third determination module is used to determine that the text segment is a text segment matching the preset sentence when the cumulative edit distance is less than or equal to the preset threshold.

[0123] Specifically, the constructed fuzzy search statement is input into the full-text search engine for retrieval. The fuzzy search statement contains all the words in the preset sentence, as well as the edit distance redundancy markers set before and after each word, which are used to indicate the range of deformation allowed by the search engine when matching. After the full-text search engine retrieves the text fragments, it calculates the minimum edit distance between each word in these fragments and the corresponding word in the preset sentence. The minimum edit distance is an indicator to measure the similarity between two strings, indicating the minimum number of single-character editing operations (including insertion, deletion, and replacement) required to convert from one string to another. By calculating the minimum edit distance, the similarity between the words in the text fragment and the words in the preset sentence can be quantified.

[0124] For each retrieved text segment, the full-text search engine will cumulatively calculate the minimum edit distance between all the words in it and the corresponding words of the preset sentence to obtain the cumulative edit distance of the text segment. The cumulative edit distance is a comprehensive indicator that reflects the similarity between the entire text segment and the preset sentence. If there are multiple words in a text segment that have edit distances with the words in the preset sentence, their edit distances will be added up to form the cumulative edit distance of the segment. Finally, the cumulative edit distance of each text segment is compared with a preset threshold. The preset threshold is a pre-set value used to determine whether the text segments are similar enough to be considered a match for the preset sentence. If the cumulative edit distance of a text segment is less than or equal to the preset threshold, then the text segment is considered to match the preset sentence. For example, if the preset threshold is 5 and the cumulative edit distance of a text segment is 4, then the text segment will be confirmed as a match because the similarity between the text segment and the preset sentence is within the allowed deformation range.

[0125] In a specific embodiment, the first calculation module includes:

[0126] Using formula Calculate the minimum edit distance between each word in each retrieved text segment and the corresponding word in the above preset sentence, where sim is the above minimum edit distance, and Represents the string s 1 and the string s 2 Length, Indicates that the above string s 1 Converted into the above string s 2 Minimum number of operations required.

[0127] Specifically, in the full-text search and matching process, the formula It is used to calculate the similarity between each word in the text fragment and the corresponding word in the preset sentence. The minimum edit distance of each word is calculated and adjusted according to the weight of the word and the redundancy of the edit distance. By calculating the edit distance, combined with the weight of the word and the allowed deformation range, the full-text search engine can comprehensively consider the exact match and semantic similarity of the words when performing fuzzy matching of text. This makes it possible to identify text fragments similar to the preset sentence even when there are OCR recognition errors, synonym replacements or slight sentence deformations in the text material. It can not only intelligently identify similar content in the text, but also avoid the strict limitations of completely accurate matching.

[0128] The above-mentioned text fuzzy matching device includes a processor and a memory. The above-mentioned word segmentation unit, setting unit, recognition unit, etc. are all stored in the memory as program units, and the processor executes the above-mentioned program units stored in the memory to realize corresponding functions. The above-mentioned modules are all located in the same processor; or, the above-mentioned modules are located in different processors in the form of any combination.

[0129] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM, and the memory includes at least one memory chip.

[0130] An embodiment of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium includes a stored program, wherein when the program is executed, the device where the computer-readable storage medium is located is controlled to execute the text fuzzy matching method.

[0131] Specifically, the text fuzzy matching method includes:

[0132] Step S201, using a word segmentation tool to perform word segmentation processing on a preset sentence to obtain a fuzzy word set consisting of multiple words;

[0133] Step S202, setting a corresponding edit distance redundancy for each of the above words in the above fuzzy word set according to the part of speech of the word, wherein the above edit distance redundancy allows each of the above words to change within a corresponding preset deformation range, and the original order relationship between the plurality of the above words remains unchanged after the deformation of each of the above words;

[0134] Step S203: construct a fuzzy search statement according to the edit distance redundancy, and input the fuzzy search statement into a full-text search engine for retrieval to identify text segments matching the preset statement.

[0135] An embodiment of the present invention provides an electronic device, the device comprising a processor, a memory, and a program stored in the memory and executable on the processor, and when the processor executes the program, at least the following steps are implemented:

[0136] Step S201, using a word segmentation tool to perform word segmentation processing on a preset sentence to obtain a fuzzy word set consisting of multiple words;

[0137] Step S202, setting a corresponding edit distance redundancy for each of the above words in the above fuzzy word set according to the part of speech of the word, wherein the above edit distance redundancy allows each of the above words to change within a corresponding preset deformation range, and the original order relationship between the plurality of the above words remains unchanged after the deformation of each of the above words;

[0138] Step S203: construct a fuzzy search statement according to the edit distance redundancy, and input the fuzzy search statement into a full-text search engine for retrieval to identify text segments matching the preset statement.

[0139] Obviously, those skilled in the art should understand that the above modules or steps of the present invention can be implemented by a general computing device, they can be concentrated on a single computing device, or distributed on a network composed of multiple computing devices, they can be implemented by a program code executable by a computing device, so that they can be stored in a storage device and executed by the computing device, and in some cases, the steps shown or described can be executed in a different order than here, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. Thus, the present invention is not limited to any specific combination of hardware and software.

[0140] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.

[0141] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0142] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0143] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0144] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0145] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0146] Computer readable media include permanent and non-permanent, removable and non-removable media that can be used to store information by any method or technology. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0147] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0148] The above description is only the preferred embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A text fuzzy matching method, characterized in that: include: Use the word segmentation tool to segment the preset sentence to obtain a fuzzy word set consisting of multiple words; Setting a corresponding edit distance redundancy for each of the words in the fuzzy word set according to the part of speech of the word, wherein the edit distance redundancy allows each of the words to change within a corresponding preset deformation range, and the original order relationship between the multiple words remains unchanged after the deformation of each of the words; A fuzzy search statement is constructed according to the edit distance redundancy, and the fuzzy search statement is input into a full-text search engine for retrieval to identify a text segment matching the preset statement.

2. The method according to claim 1, characterized in that Before using a word segmentation tool to perform word segmentation processing on a preset sentence to obtain a fuzzy word set consisting of a plurality of words, the method further includes: Converting the non-text files to be matched into text format, storing the converted files to be matched into the full-text search engine, and constructing a search database including the converted files to be matched; By using the indexing mechanism of the full-text search engine, an inverted index is created for each text in the converted to-be-matched file, wherein the inverted index is used to record the text where the target word is located and the position information in the text.

3. The method according to claim 1, characterized in that Setting a corresponding edit distance redundancy for each word in the fuzzy word set according to the part of speech of the word includes: Performing part-of-speech tagging on each of the words in the fuzzy word set, and performing word classification processing based on the part-of-speech of the words to obtain multiple sub-word sets, each of the sub-word sets corresponding to a part-of-speech; Determining the weight of each word in the sub-word set according to the part of speech of each sub-word set, wherein the weight reflects the influence of the word on the matching result of the text fuzzy matching; The edit distance redundancy of each of the words is determined according to the weight of each of the words.

4. The method according to claim 3, characterized in that Determining the edit distance redundancy of each of the words according to the weight of each of the words includes: When the weight of the word is greater than or equal to a preset weight, setting the edit distance redundancy of the word to be within a first redundancy range; When the weight of the word is less than the preset weight, the edit distance redundancy of the word is set to be within a second redundancy range, The minimum value of the second redundancy range is greater than the maximum value of the first redundancy range.

5. The method according to claim 1, characterized in that Constructing a fuzzy search statement according to the edit distance redundancy includes: The words in the fuzzy word set are arranged in the order of the original sentence, and deformation range marks allowed by the edit distance redundancy are added before and after each word to obtain the fuzzy search sentence.

6. The method according to claim 1, characterized in that Input the fuzzy search statement into a full-text search engine for retrieval, and identify text segments that match the preset statement, including: Inputting the fuzzy search statement into the full-text search engine for retrieval, and calculating the minimum edit distance between each word in each retrieved text segment and the corresponding word of the preset statement; Accumulating and calculating the minimum edit distances of all words in each of the text segments to obtain a cumulative edit distance of the text segment, and comparing the cumulative edit distance with a preset threshold; When the cumulative edit distance is less than or equal to the preset threshold, the text segment is determined to be a text segment matching the preset sentence.

7. The method according to claim 6, characterized in that Calculating the minimum edit distance between each word in each retrieved text segment and a corresponding word of the preset sentence, including: Using formula Calculate the minimum edit distance between each word in each retrieved text segment and the corresponding word of the preset sentence, where sim is the minimum edit distance, and Represent the length of string s1 and string s2 respectively, Indicates the minimum number of operations required to convert the string s1 into the string s2.

8. A text fuzzy matching device, characterized in that: include: A word segmentation unit is used to perform word segmentation processing on a preset sentence using a word segmentation tool to obtain a fuzzy word set consisting of multiple words; A setting unit, configured to set a corresponding edit distance redundancy for each of the words in the fuzzy word set according to the part of speech of the word, wherein the edit distance redundancy allows each of the words to change within a corresponding preset deformation range, and the original order relationship between the multiple words remains unchanged after the deformation of each word; The identification unit is used to construct a fuzzy search statement according to the edit distance redundancy, and input the fuzzy search statement into a full-text search engine for retrieval to identify a text segment matching the preset statement.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein when the program is executed, the device where the computer-readable storage medium is located is controlled to execute the text fuzzy matching method according to any one of claims 1 to 7.

10. An electronic device, characterized in that: include: One or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include a method for executing the text fuzzy matching method described in any one of claims 1 to 7.

Citation Information

Cited By

  • Document verification method and device, electronic equipment and storage medium

    CN121257551A