Keyword matching method and device, storage medium and electronic equipment

By using the subword filtering mechanism in the jieba word segmentation algorithm, we can detect whether the target words in the target text only contain preset keywords, which solves the problem of false positives in English keyword matching and improves the accuracy of English keyword matching.

CN120146045APending Publication Date: 2025-06-13HILLSTONE NETWORKS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510308189.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

When using the jieba word segmentation algorithm to match keywords in the prior art, there are false positives problems for accurate identification and matching of English keywords, resulting in reduced accuracy.

Method used

By detecting whether the target word in the target text includes other characters other than the preset keywords, the matching result between the target word and the preset keyword is determined, and a subword filtering mechanism is used to accurately identify independent English keywords.

Benefits of technology

It improves the accuracy of English keyword matching, reduces false positives, and achieves accurate recognition of English keywords.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146045A_ABST
    Figure CN120146045A_ABST
Patent Text Reader

Abstract

The invention discloses a keyword matching method and device, a storage medium and electronic equipment, and relates to the field of natural language processing. The method comprises the steps that when it is detected that a kth target word in a target text is a word of a target language and the kth target word comprises a preset keyword, whether the kth target word comprises other characters except the preset keyword or not is detected, and k is an integer larger than or equal to 1; when it is detected that the kth target word comprises other characters except the preset keyword, it is determined that matching between the kth target word and the preset keyword fails; and when it is detected that the kth target word does not include other characters except the preset keyword, determining that the kth target word is successfully matched with the preset keyword. The technical problem that in the prior art, when a jieba word segmentation algorithm is used for keyword matching, English keywords cannot be effectively recognized is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of natural language processing, and in particular, to a method, apparatus, storage medium, and electronic device for keyword matching. Background Art

[0002] In the fields of information retrieval, natural language processing, and data security, keyword matching is one of the crucial basic technologies, which is used to quickly locate and filter texts containing specific information. Conventional word segmentation algorithms (such as the jieba algorithm, a main algorithm for segmenting Chinese texts) are widely used in many application scenarios due to their efficient and accurate word segmentation capabilities.

[0003] However, when processing texts containing English keywords, the jieba algorithm shows certain limitations, mainly reflected in the inaccurate recognition and matching of English keywords, especially when facing the situation where English words are sub-words. Since the jieba algorithm is designed to focus on the processing of Chinese texts, its English word segmentation function is relatively basic and often achieved through simple character matching. This results in that in practical applications, the matching results of English keywords often contain false positives, that is, the algorithm will misidentify some keywords as parts of other longer keywords. This misrecognition seriously reduces the accuracy of English keyword matching. Therefore, when using the jieba word segmentation algorithm for keyword matching in the prior art, the accurate recognition and matching of English keywords have become a key technical problem to be solved urgently.

[0004] In response to the above problems, no effective solution has been proposed yet. Summary of the Invention

[0005] The present application provides a method, apparatus, storage medium, and electronic device for keyword matching, so as to at least solve the technical problem that English keywords cannot be effectively recognized when using the jieba word segmentation algorithm for keyword matching in the prior art.

[0006] According to one aspect of the present application, a method for keyword matching is provided, including: when it is detected that the k-th target word in the target text is a word in the target language and the k-th target word includes a preset keyword, detecting whether the k-th target word includes other characters except the preset keyword, where k is an integer greater than or equal to 1; when it is detected that the k-th target word includes other characters except the preset keyword, determining that the k-th target word fails to match the preset keyword; when it is detected that the k-th target word does not include other characters except the preset keyword, determining that the k-th target word matches the preset keyword successfully.

[0007] Optionally, after detecting that the k-th target word does not include other characters except the preset keyword and determining that the k-th target word matches the preset keyword successfully, the keyword matching method further includes: when detecting that there are S target words in the target text that match the preset keyword successfully, determining the S target words as S matching words, where S is an integer greater than 1; when the difference between the offsets of any two matching words among the S matching words is less than or equal to a preset value, determining the two first-type words as a keyword pair, where the offset is used to represent the distance between the matching word and the starting position of the target text; adding the S matching words and the keyword pair to the target set.

[0008] Optionally, when detecting that the k-th target word in the target text is a word in the target language and the k-th target word includes the preset keyword, before detecting whether the k-th target word includes other characters except the preset keyword, the method further includes: splitting the target text into N sub-texts, where N is an integer greater than 1; inputting the N sub-texts into a thread pool, where the thread pool includes multiple threads, and the multiple threads are used to collaboratively process the N sub-texts; controlling the multiple threads in the thread pool to extract the target words of each sub-text in the N sub-texts in a parallel processing manner, and detecting the language of each target word and whether each target word includes the preset keyword.

[0009] Optionally, splitting the target text into N sub-texts includes: performing N - 1 target operations on the target text until N sub-texts are obtained, where each target operation is used to split a sub-text with a text length equal to the preset length on the basis of the target text.

[0010] Optionally, the j-th target operation among the N - 1 target operations, where j is an integer greater than or equal to 1 and less than N - 1, includes the following steps: detecting whether the length of the target text is greater than the preset length; in the case where the length of the target text is greater than the preset length, performing a cutting process on the target text according to the preset length to obtain a first sub-text and a second sub-text, where the first sub-text is a text with the preset length, and the second sub-text is the remaining text of the target text except the first sub-text; using the second sub-text obtained after the j-th target operation as the target text for the (j + 1)-th target operation.

[0011] Optionally, adding the S matching words and the keyword pair to the target set includes: adding the position information of the S matching words in the target text and the S matching words to the target set; adding the position information of the keyword pair in the target text and the keyword pair to the target set.

[0012] Optionally, after adding the S matching words and keyword pairs to the target set, the method further includes: creating a target lock in the thread pool, where the target lock is used to manage the update of the target set; when the t-th thread in the thread pool needs to update the target set, requesting to obtain the target lock, where t is an integer greater than or equal to 1; when the target lock is occupied by other threads, prohibiting the t-th thread from updating the target set; when the target lock is not occupied, obtaining the target lock through the t-th thread and updating the target set; after the t-th thread updates the target set, releasing the target lock.

[0013] According to another aspect of the present application, there is also provided a keyword matching device, including: a detection unit, when detecting that the k-th target word in the target text is a word in the target language and the k-th target word includes a preset keyword, detecting whether the k-th target word includes other characters except the preset keyword, where k is an integer greater than or equal to 1; a first determination unit, when detecting that the k-th target word includes other characters except the preset keyword, determining that the k-th target word fails to match the preset keyword; a second determination unit, when detecting that the k-th target word does not include other characters except the preset keyword, determining that the k-th target word matches the preset keyword successfully.

[0014] According to another aspect of the present application, there is also provided a computer-readable storage medium, in which a computer program is stored. When the computer program runs, the device where the computer-readable storage medium is located executes the above-mentioned keyword matching method.

[0015] According to another aspect of the present application, there is also provided an electronic device, including one or more processors and a memory, where the memory is used to store one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors execute the above-mentioned keyword matching method.

[0016] In the present application, when detecting that the k-th target word in the target text is a word in the target language and the k-th target word includes a preset keyword, it is detected whether the k-th target word includes other characters except the preset keyword, where k is an integer greater than or equal to 1. Then, when detecting that the k-th target word includes other characters except the preset keyword, it is determined that the k-th target word fails to match the preset keyword. When detecting that the k-th target word does not include other characters except the preset keyword, it is determined that the k-th target word matches the preset keyword successfully. That is, through the sub-word filtering mechanism, the purpose of accurately identifying independent English keywords is achieved, thereby realizing the technical effect of improving the accuracy of English keyword matching, and further solving the technical problem that English keywords cannot be effectively identified when using the jieba word segmentation algorithm for keyword matching in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The drawings described herein are provided to further understand the present application and form a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:

[0018] Figure 1 is a flowchart of an optional keyword matching method according to an embodiment of the present application;

[0019] Figure 2 is the process of completing keyword matching according to the word segmentation algorithm provided by the embodiment of the present application Figure 1 ;

[0020] Figure 3 is the process of completing keyword matching according to the word segmentation algorithm provided by the embodiment of the present application Figure 2 ;

[0021] Figure 4 is a flowchart of an optional thread pool processing according to an embodiment of the present application;

[0022] Figure 5 is a schematic diagram of an optional keyword matching device according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0023] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0024] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order different from those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0025] It should be noted that the information collected in this application (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for display, data for analysis, etc.) are information and data that have been authorized by the user or fully authorized by all parties. Moreover, for the processing of relevant data such as collection, storage, use, processing, transmission, provision, disclosure, and application, all comply with relevant laws, regulations, and standards, necessary confidentiality measures are taken, it does not violate public order and good customs, and corresponding operation entrances are provided for users to choose to authorize or refuse.

[0026] According to an embodiment of the present application, a method embodiment of a keyword matching method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0027] It should be noted that an intelligent matching system can be the execution subject of the keyword matching method of the embodiment of the present application. It can be understood that the keyword matching method provided by the embodiment of the present application can also be executed by other systems or devices as the execution subject, and the embodiment of the present application does not make specific limitations on this.

[0028] Figure 1 is a flowchart of an optional keyword matching method according to an embodiment of the present application, as Figure 1 shown, the method includes the following steps:

[0029] Step S101, when it is detected that the k-th target word in the target text is a word in the target language and the k-th target word includes a preset keyword, detect whether the k-th target word includes other characters except the preset keyword.

[0030] In step S101, k is an integer greater than or equal to 1.

[0031] Optionally, the target text refers to the text to be processed, that is, the original text input to the word segmentation algorithm for keyword matching.

[0032] Optionally, the preset keyword refers to a keyword that is pre-configured by the user and needs to be searched in the text.

[0033] Optionally, the intelligent matching system first processes the target text to obtain multiple target words, then detects the multiple target words according to the preset keywords, and then checks whether the target word is the target language (such as English). When the target word is the target language, it detects whether the preset keywords are included in the target word. When the target word is a word in the target language and the k-th target word includes the preset keywords, it detects whether there are other characters in the k-th target word except the preset keywords.

[0034] It should be noted that the target language can be any other language in addition to English.

[0035] It should also be noted that the embodiments of the present application perform keyword matching based on the improvement of the traditional word segmentation algorithm. Specifically, it mainly adds the filtering of keyword sub-words in the English state, the matching of Chinese and English keyword pairs, and the use of batch processing plus thread pool processing technology in the traditional word segmentation algorithm. Among them, the traditional word segmentation algorithm, such as the jieba word segmentation algorithm, is a widely used Chinese word segmentation tool and algorithm. It combines the forward maximum matching algorithm based on the dictionary and the hidden Markov model based on statistics to segment Chinese text to achieve efficient word segmentation. In the jieba word segmentation algorithm, if the keyword is in English, when it is a sub-word of a longer English keyword, a matching result will also be obtained through character matching, which does not meet the requirements of keyword matching. For example, if the keyword is "other", if the English sentence contains "another", a matching result of the "other" keyword will also be obtained, which is an incorrect matching. Based on the improved jieba word segmentation algorithm in the embodiments of the present application, the system can accurately identify English keywords.

[0036] Step S102, when it is detected that the k-th target word includes other characters except the preset keywords, it is determined that the k-th target word fails to match the preset keywords.

[0037] Optionally, when the intelligent matching system detects that the k-th target word includes other characters in addition to the preset keywords, it is determined that the k-th target word fails to match the preset keywords.

[0038] Optionally, if the target word contains other characters except the preset keywords, then according to the improved word segmentation algorithm logic provided by the present application, this target word will not be considered an exact match of the preset keywords. For example, if "other" in "projectother" is detected, since "project" also follows "other", it will not be recognized as a match of the preset keyword "other". This step ensures that the algorithm can accurately distinguish independent keywords when processing English keywords, avoiding the possibility of incorrect matching.

[0039] In step S103, when it is detected that the k-th target word does not include any other characters except the preset keyword, it is determined that the k-th target word matches the preset keyword successfully.

[0040] Optionally, when the intelligent matching system detects that the k-th target word does not include any other characters except the preset keyword, that is, when the k-th target word is exactly the same as the preset keyword without any extra characters or substrings, it is determined that the k-th target word matches the preset keyword successfully.

[0041] Optionally, Figure 2 is the process of keyword matching completed according to the word segmentation algorithm provided by the embodiments of the present application Figure 1 , as Figure 2 shown, first, the keywords configured by the user and the text to be processed are input. The CutAll function (full-mode word segmentation function) of the jieba word segmentation algorithm is called, and the keywords configured by the user and the text to be processed are input into the CutAll function of the jieba word segmentation algorithm. Then, it enters the Cut function (word segmentation cutting function) that includes the text preprocessing process for text cutting. Then, the text is batch-cut and input into the thread pool for processing (the thread pool created in the constructor of the jieba word segmentation algorithm). All the batch-cut text is processed in the thread pool, including word segmentation processing and keyword matching processing. After that, the processing result of the thread pool, that is, the keyword matching result, is obtained, and finally the keyword matching result is output.

[0042] Optionally, the intelligent matching system stores the keywords configured by the user in the map structure of C++ and passes them layer by layer into the word segmentation function of the improved jieba word segmentation algorithm. In the word segmentation function, the words cut out are compared with the keywords to distinguish whether they are sub-words. If they are sub-words in the English state, no matching is performed, so as to achieve accurate matching of keywords in the English state, rather than simple character matching.

[0043] As can be seen from the content of steps S101 to S103, in this application, when it is detected that the k-th target word in the target text is a word in the target language and the k-th target word contains a preset keyword, it is detected whether the k-th target word contains other characters except the preset keyword, where k is an integer greater than or equal to 1. Then, when it is detected that the k-th target word contains other characters except the preset keyword, it is determined that the k-th target word fails to match the preset keyword. When it is detected that the k-th target word does not contain other characters except the preset keyword, it is determined that the k-th target word matches the preset keyword successfully. That is, through the sub-word filtering mechanism, the purpose of accurately identifying independent English keywords is achieved, thereby realizing the technical effect of improving the matching accuracy of English keywords, and further solving the technical problem that English keywords cannot be effectively identified when using the jieba word segmentation algorithm in the prior art.

[0044] In an optional embodiment, when the intelligent matching system detects that there are S target words in the target text that match the preset keyword successfully, it determines that these S target words are S matching words, where S is an integer greater than 1. When the difference between the offsets of any two matching words among the S matching words is less than or equal to a preset value, it is determined that these two first-type words are a keyword pair, where the offset is used to represent the distance between the matching word and the starting position of the target text. Finally, the S matching words and the keyword pairs are added to the target set.

[0045] Optionally, on the basis that the improved jieba word segmentation algorithm can accurately identify words in the target language, when the intelligent matching system processes the text using the improved jieba word segmentation algorithm, it can identify S target words that exactly match the preset keywords configured by the user. For each matching word, the system will calculate its offset in the target text, that is, the distance between the matching word and the starting position of the text. Then, it detects the offset difference between any two of the S matching words. If this difference is less than or equal to the preset value, then these any two matching words will be determined as a keyword pair. Among them, the setting of the preset value depends on the specific application scenario and the definition of the keyword pair. It can help the system identify those keyword pairs that are closely related in the text. Finally, all the S matching words that match successfully and the combinations identified as keyword pairs are added to a target set. This set will include all the target words that match successfully in the target text and their position information, as well as the details of the keyword pairs, providing comprehensive information for subsequent text analysis and processing.

[0046] From the above content, we can see that the intelligent matching system provides the ability of Jieba word segmentation algorithm in keyword word pair matching. By accurately calculating the offset of matching words and using preset values ​​to judge word pairs, the system can more intelligently identify keyword relationships in text.

[0047] In an optional embodiment, the intelligent matching system divides the target text into N sub-texts, where N is an integer greater than 1, and inputs the N sub-texts into a thread pool, wherein the thread pool includes multiple threads, and the multiple threads are used to collaboratively process the N sub-texts, and control the multiple threads in the thread pool to adopt a parallel processing method to extract the target words of each sub-text in the N sub-texts, and detect the language of each target word and whether each target word includes preset keywords.

[0048] Optionally, first, the intelligent matching system divides the original target text into N smaller sub-texts. The purpose of this process is to decompose a large text processing task into multiple smaller sub-tasks that can be processed in parallel. For example, if the target text is a document containing thousands of words, the system may divide it into N paragraphs or sentence groups for subsequent parallel processing. Then, these N sub-texts are input into the thread pool for processing, wherein the thread pool contains multiple threads, which will be used to collaboratively process the N sub-texts. Each thread is responsible for processing one or more sub-texts, so that word segmentation and keyword matching can be performed on multiple sub-texts at the same time, greatly improving the processing speed. For example, if there are 5 threads in the thread pool, then these 5 threads will process N sub-texts in parallel, and the number of sub-texts processed by each thread depends on the thread pool strategy and the size of the sub-texts.

[0049] Optionally, multiple threads in the thread pool use parallel processing to segment each subtext and detect keywords therein. This process includes identifying the language of each target word and determining whether the word includes preset keywords. For example, if a subtext contains "project", "data", and "password", the thread will identify "project" and "data" as English words, "password" as a Chinese word, and check whether they are in the preset keyword list. For English keywords, the system will further check whether there are subwords (i.e., words containing characters other than preset keywords) that are mismatched, ensuring that only fully matched words are identified. Finally, when all threads have completed the processing of their respective subtexts, the intelligent matching system will collect and summarize the processing results of these threads, including all extracted keywords and their corresponding language and text position information.

[0050] As can be seen from the above, the intelligent matching system adopts a multi-threaded batch processing method for text. It performs sentence segmentation and keyword search and matching within multiple threads, which can significantly reduce the text processing time, improve the efficiency of keyword extraction, and greatly enhance the running efficiency of the program. At the same time, the optimization of sub-word filtering and language identification for English keywords ensures the accuracy of keyword matching, avoiding false positives and false negatives, and enhancing the reliability of information retrieval and text analysis. Overall, this implementation method not only accelerates the text processing process but also improves the effectiveness of keyword matching.

[0051] In an alternative embodiment, the intelligent matching system performs N - 1 target operations on the target text until N sub-texts are obtained. Each target operation is used to split a sub-text with a text length equal to the preset length based on the target text.

[0052] Optionally, the intelligent matching system first receives the target text input by the user as the initial object for the entire text processing. The target text can be a document of any length and contain a mixture of Chinese and English content. Then, a preset length is set, which is determined according to the system processing capacity and text analysis requirements. For example, if the system processing capacity is strong, the preset length can be set to 500 characters to ensure that each sub-text contains sufficient information. If the system processing capacity is limited or the text data volume is extremely large, the preset length can be set smaller to improve the processing speed. Then, the system performs N - 1 target operations on the target text, and each operation splits a sub-text with a length equal to the preset length based on the target text. This process is achieved through a sliding window method, that is, starting from the starting position of the target text, each operation splits a fixed-length text, and then moves to the next position and repeats this operation until the entire target text is completely processed. Through the above N - 1 target operations, the system finally divides the target text into N sub-texts, and the length of each sub-text is equal to the preset length. This strategy enables subsequent multi-threaded processing to process these sub-texts in parallel, greatly improving the efficiency of text analysis.

[0053] Optionally, the preset lengths set for English target words and Chinese target words can be different.

[0054] As can be seen from the above, by implementing this strategy of N - 1 target operations, the intelligent matching system can efficiently split long texts, providing an optimized basis for multi-threaded processing, keyword matching, and word pair recognition. This strategy improves the efficiency of text processing and is particularly suitable for processing large amounts of long Chinese-English mixed texts. The optimization in the text preprocessing stage helps the entire system to be more efficient and accurate in the subsequent processing stage, significantly enhancing the overall performance of text analysis.

[0055] In an alternative embodiment, when the length of the target text is greater than a preset length, the intelligent matching system performs a cutting process on the target text according to the preset length to obtain a first sub-text and a second sub-text, where the first sub-text is a text of the preset length, and the second sub-text is the remaining text of the target text except the first sub-text. Then, the second sub-text obtained after completing the j-th target operation is used as the target text for the (j + 1)-th target operation.

[0056] Optionally, for the process of the target operation, the intelligent matching system first determines the preset length of the target text. If the text length is greater than the preset length, the target text is cut to obtain a first sub-text and a second sub-text, where the first sub-text is a text of the preset length, and the second sub-text is the remaining text of the target text except the first sub-text. Then, the second sub-text is used as the target text for the next target operation.

[0057] Optionally, Figure 3 is the process of completing keyword matching according to the word segmentation algorithm provided by the embodiments of the present application Figure 2 , as Figure 3 shown, first, the keywords configured by the user and the text to be processed (target text) are input. The CutAll function of the jieba word segmentation algorithm is called, and the keywords configured by the user and the text to be processed are input into the CutAll function of the jieba word segmentation algorithm. Then, it enters the Cut function including the text preprocessing process for text cutting. It is judged whether the starting offset of the preprocessed text is less than the text length (the length of the target text). If the starting offset of the preprocessed text is less than the length of the target text, cutting is performed to obtain a text segment (a text segment preprocessed). The processed text is assigned to the thread pool for further processing. The thread pool will cut the text into words, then perform keyword matching, and then add the matching keywords obtained by the thread pool processing to the keyword result set. On the contrary, if the starting offset of the preprocessed text is greater than the length of the target text, it directly waits for all the text in the thread pool to be processed, returns the keyword matching result, and outputs the keyword matching result.

[0058] For example, input a text (target text) with a length of 1000 characters and a preset keyword, call the CutAll function of the jieba tokenization algorithm. Through the CutAll function, input the target text into the Cut function for text cutting. For the first cut, at this time, the offset (distance from the starting position of the target text) is 0, which is less than the length of the target text, so cutting is performed. After cutting according to a certain preset length (such as 100), two sub-texts (the first sub-text and the second sub-text) are obtained. Allocate the first sub-text to the thread pool for tokenization processing and keyword matching processing, and add the matching results to the keyword result set. At the same time, add 1 to the end position of the first half of the sub-text (the first sub-text) as the starting offset for the next cut, that is, 101 (which is also equivalent to the offset of the second sub-text relative to the starting position of the target text); during the second cut, repeat the above steps for judgment and cutting, and also add 1 to the end position of the first half of the sub-text as the starting offset for the next cut, that is, 201; loop the above operations. After the last cut, the obtained offset is 1001, which is greater than the total length of 1000 at this time, so no next loop is performed, and the loop ends. Wait for all the text in the thread pool to be processed, return the keyword matching results, and output the keyword matching results.

[0059] As can be seen from the above content, the intelligent matching system loops to cut the target text according to the preset length, and finally obtains multiple sub-texts, that is, decomposes the long text into smaller and more easily processed blocks to meet the needs of multi-threaded processing. In this way, the system can process multiple text blocks in parallel, thereby significantly improving the processing speed.

[0060] In an alternative embodiment, the intelligent matching system first adds the position information of S matching words in the target text and the S matching words to the target set, and then adds the position information of the keyword pairs in the target text and the keyword pairs to the target set.

[0061] Optionally, the intelligent matching system adds S matching words and their position information in the target text to the target set. This process includes recording the start offset and end offset of each matching word to accurately locate the position of the keyword in the text. For example, if "password" is successfully matched in a text, its position information will be recorded, such as the start offset being 21 and the end offset being 29. Then, "password" and its position information will be added to the target set. At the same time, the intelligent matching system continues to process keyword pairs, integrating the identified keyword pairs and their position information into the same target set. This step is based on the previous matching word position information to calculate the offset difference between any two matching words. When this difference meets the definition of a keyword pair (i.e., less than or equal to a preset value), these two words are confirmed as a keyword pair, and their combination and position information in the text will also be recorded and added to the target set.

[0062] As can be seen from the above, by integrating the matching word information, the intelligent matching system can accurately track the position of each matching word in the text, which is crucial for identifying keyword pairs and subsequent text analysis. By recording the information of keyword pairs, including their combination and position, it helps to deeply understand the text structure and the relationship between keywords, improving the depth and accuracy of text analysis.

[0063] In an optional embodiment, the intelligent matching system creates a target lock within the thread pool. Here, the target lock is used to manage the update of the target set. When the t-thread in the thread pool needs to update the target set, it requests to acquire the target lock, where t is an integer greater than or equal to 1. When the target lock is occupied by other threads, the t-thread is prohibited from updating the target set. When the target lock is not occupied, the t-thread acquires the target lock and updates the target set. After the t-thread updates the target set, it releases the target lock.

[0064] Optionally, the intelligent matching system deploys a target lock inside the thread pool. This is a special control structure used for mutual exclusion control when multiple threads attempt to update the target set (the above keyword result set) simultaneously to prevent data conflicts. The creation of the target lock is part of the system initialization to ensure that the update of the target set can proceed orderly in a multi-threaded environment.

[0065] Optionally, when the t-th thread in the thread pool needs to update the target set, it first requests to acquire the target lock. If the target lock is currently occupied by another thread, the t-th thread will not be able to update the target set immediately and will enter a waiting state until the target lock is released. This mechanism avoids the situation where multiple threads modify the target set simultaneously and reduces the risk of data inconsistency. If the target lock is not occupied when the request is made, the t-th thread will acquire the target lock and start updating the target set. Once the update is complete, the t-th thread will immediately release the target lock, allowing other waiting threads to operate.

[0066] Optionally, Figure 4 is a flowchart of an optional thread pool processing according to an embodiment of the present application. As Figure 4 shown, in the thread pool, first, the preprocessed text (i.e., the sub-text obtained by cutting the original text) and keywords are input. According to the words in the dictionary (a set including words in multiple languages) and the preset keywords, the text is subjected to word segmentation and cutting processing to obtain multiple target words. The cut keywords are compared with the keywords to be matched. For keyword matching in the English state, reasonable sub-word filtering is performed to filter out character matching. Then, the positions of the matched keywords are added to the result set, and the corresponding keywords are extracted according to the positions of the keywords. At the same time, the keyword result set is locked to avoid data conflicts and inconsistencies. After that, the matched keywords and the offset of the keyword (i.e., the distance between the keyword and the starting position of the original text) are added to the result set, and finally the thread ends.

[0067] Optionally, the embodiment of the present application is applicable to the keyword extraction function in the C++ version and currently supports the matching of Chinese and English keywords and keyword pairs. After adding this function to the DLP product, it can perform security auditing and filtering on the outgoing files, and can determine whether the document is a sensitive file according to the keywords to protect the file from being leaked; after adding this function to the search engine, by matching the keywords in the user's search query, it can more effectively attract the target audience; in natural language processing tasks, it helps the system understand and parse the text, extract key information, and perform subsequent operations accordingly.

[0068] As can be seen from the above, the adoption of the target lock mechanism significantly improves the stability and reliability of the improved Jieba word segmentation algorithm in a multi-threaded environment. By ensuring that the update operation of the target set is mutually exclusive, it avoids the inconsistency problem caused by data competition, so that even in a high-concurrency scenario, the system can accurately perform keyword matching and result integration. The use of this target lock not only enhances the parallel processing ability of the algorithm but also ensures the correctness of the data.

[0069] An embodiment of the present application also provides a keyword matching device. It should be noted that the keyword matching device in the embodiment of the present application can be used to execute the keyword matching method provided in the embodiment of the present application. The following introduces the keyword matching device provided in the embodiment of the present application.

[0070] According to an embodiment of the present application, there is also provided a device for implementing the above keyword matching method. Figure 5 It is a schematic diagram of an optional keyword matching device according to an embodiment of the present application, as Figure 5 shown. The device includes: a detection unit 501, a first determination unit 502, and a second determination unit 503.

[0071] Optionally, the detection unit 501 is configured to detect whether there are other characters in the k-th target word in the target text except the preset keyword when it is detected that the k-th target word in the target text is a word in the target language and the k-th target word includes the preset keyword, where k is an integer greater than or equal to 1; the first determination unit 502 is configured to determine that the k-th target word fails to match the preset keyword when it is detected that there are other characters in the k-th target word except the preset keyword; the second determination unit 503 is configured to determine that the k-th target word matches the preset keyword successfully when it is detected that there are no other characters in the k-th target word except the preset keyword.

[0072] Optionally, the keyword matching device further includes: a third determination unit, a fourth determination unit, and a first processing unit. Among them, the third determination unit is configured to determine the S target words as S matching words when it is detected that there are S target words in the target text that match the preset keyword successfully, where S is an integer greater than 1; the fourth determination unit is configured to determine that the two first type words are a keyword pair when the difference between the offsets of any two matching words among the S matching words is less than or equal to a preset value, where the offset is used to represent the distance between the matching word and the starting position of the target text; the first processing unit is configured to add the S matching words and the keyword pair to the target set.

[0073] Optionally, the keyword matching device further includes: a first segmentation unit, a first input unit, and a first control unit. Among them, the first segmentation unit is configured to segment the target text into N sub-texts, where N is an integer greater than 1; the first input unit is configured to input the N sub-texts into a thread pool, where the thread pool includes multiple threads, and the multiple threads are used to collaboratively process the N sub-texts; the first control unit is configured to control the multiple threads in the thread pool to extract the target words of each sub-text in the N sub-texts and detect the language of each target word and whether each target word includes the preset keyword in a parallel processing manner.

[0074] Optionally, the first splitting unit includes: a first processing subunit, configured to perform N - 1 target operations on the target text until N sub - texts are obtained, where each target operation is used to split a sub - text with a text length equal to the preset length based on the target text.

[0075] Optionally, the first processing subunit includes: a first processing module and a first determination module. The first processing module is configured to, when the length of the target text is greater than the preset length, perform cutting processing on the target text according to the preset length to obtain a first sub - text and a second sub - text, where the first sub - text is a text with the preset length and the second sub - text is the remaining text of the target text except the first sub - text; the first determination module is configured to use the second sub - text obtained after the j - th target operation as the target text for the (j + 1) - th target operation.

[0076] Optionally, the first processing unit includes: a second processing subunit and a third processing subunit. The second processing subunit is configured to add the position information of S matching words in the target text and the S matching words to the target set; the third processing subunit is configured to add the position information of keyword pairs in the target text and the keyword pairs to the target set.

[0077] Optionally, the matching device for keywords further includes: a first creation unit, a first request unit, a first prohibition unit, a first acquisition unit, and a first release unit. The first creation unit is configured to create a target lock in the thread pool, where the target lock is used to manage the update of the target set; the first request unit is configured to request to acquire the target lock when the t - th thread in the thread pool needs to update the target set, where t is an integer greater than or equal to 1; the first prohibition unit is configured to prohibit the t - th thread from updating the target set when the target lock is occupied by other threads; the first acquisition unit is configured to acquire the target lock through the t - th thread and update the target set when the target lock is not occupied; the first release unit is configured to release the target lock after the t - th thread updates the target set.

[0078] According to another aspect of the present application, there is also provided a computer - readable storage medium. A computer program is stored in the computer - readable storage medium. When the computer program runs, it causes the device where the computer - readable storage medium is located to execute the above - mentioned keyword matching method.

[0079] According to another aspect of the present application, there is also provided an electronic device, including one or more processors and a memory. The memory is used to store one or more programs. When the one or more programs are executed by the one or more processors, it causes the one or more processors to execute the above - mentioned keyword matching method.

[0080] The serial numbers of the embodiments of the present application above are only for description and do not represent the superiority or inferiority of the embodiments.

[0081] In the above embodiments of the present application, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0082] The embodiments or examples of the present application are not exhaustive. They are only illustrations of some embodiments or examples and do not specifically limit the protection scope of the present application. Without conflict, each step in a certain embodiment or example can be implemented as an independent embodiment, and the steps can be combined arbitrarily. For example, the solution after removing some steps in a certain embodiment or example can also be implemented as an independent embodiment, and the order of the steps in a certain embodiment or example can be arbitrarily exchanged. In addition, the optional methods or optional examples in a certain embodiment or example can be combined arbitrarily; moreover, the various embodiments or examples can be combined arbitrarily. For example, some or all of the steps of different embodiments or examples can be combined arbitrarily, and a certain embodiment or example can be combined arbitrarily with the optional methods or optional examples of other embodiments or examples.

[0083] In the several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units can be a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of units or modules can be in an electrical or other form.

[0084] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0085] In addition, the functional units in the various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0086] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs.

[0087] The foregoing are only the preferred embodiments of this application. It should be noted that for those of ordinary skill in the art, without departing from the principle of this application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of this application.

Claims

1. A keyword matching method, characterized in that: include: When it is detected that the kth target word in the target text is a word of the target language and the kth target word includes a preset keyword, detecting whether the kth target word includes other characters except the preset keyword, wherein k is an integer greater than or equal to 1; When it is detected that the k-th target word includes other characters except the preset keyword, determining that the k-th target word fails to match the preset keyword; When it is detected that the k-th target word does not include other characters except the preset keyword, it is determined that the k-th target word matches the preset keyword successfully.

2. The keyword matching method according to claim 1, characterized in that: When it is detected that the k-th target word does not include any other characters except the preset keyword, after determining that the k-th target word successfully matches the preset keyword, the keyword matching method further includes: When it is detected that there are S target words in the target text that successfully match the preset keyword, the S target words are determined to be S matching words, where S is an integer greater than 1; When the difference between the offsets of any two matching words among the S matching words is less than or equal to a preset value, the two first-category words are determined to be a keyword pair, wherein the offset is used to represent the distance between the matching words and the starting position of the target text; The S matching words and the keyword word pairs are added to a target set.

3. The keyword matching method according to claim 1, characterized in that: When it is detected that the kth target word in the target text is a word of the target language and the kth target word includes a preset keyword, before detecting whether the kth target word includes other characters except the preset keyword, the method further includes: Divide the target text into N subtexts, where N is an integer greater than 1; Inputting the N subtexts into a thread pool, wherein the thread pool includes a plurality of threads, and the plurality of threads are used to collaboratively process the N subtexts; The multiple threads in the thread pool are controlled to adopt a parallel processing mode to extract the target word of each subtext in the N subtexts, and detect the language of each target word and whether each target word includes the preset keyword.

4. The keyword matching method according to claim 3, characterized in that: The target text is divided into N sub-texts, including: The target text is subjected to N-1 target operations until the N sub-texts are obtained, wherein each target operation is used to segment the target text to obtain a sub-text having a text length equal to a preset length.

5. The keyword matching method according to claim 4, characterized in that: The j-th target operation among the N-1 target operations, where j is an integer greater than or equal to 1 and less than N-1, comprises the following steps: Detecting whether the length of the target text is greater than the preset length; In the case where the length of the target text is greater than the preset length, the target text is segmented according to the preset length to obtain a first subtext and a second subtext, wherein the first subtext is the text of the preset length, and the second subtext is the remaining text of the target text except the first subtext; The second subtext obtained after completing the j-th target operation is used as the target text used for the j+1-th target operation.

6. The keyword matching method according to claim 2, characterized in that: Adding the S matching words and the keyword word pair to a target set includes: Adding the position information of the S matching words in the target text and the S matching words to the target set; The position information of the keyword pair in the target text and the keyword pair are added to the target set.

7. The keyword matching method according to claim 2, characterized in that: After adding the S matching words and the keyword word pair to the target set, the method further includes: Creating a target lock in the thread pool, wherein the target lock is used to manage updates of the target set; When the tth thread in the thread pool needs to update the target set, it requests to acquire the target lock, where t is an integer greater than or equal to 1; When the target lock is occupied by other threads, prohibiting the tth thread from updating the target set; When the target lock is not occupied, acquiring the target lock through the tth thread, and updating the target set; After the tth thread updates the target set, the target lock is released.

8. A keyword matching device, characterized in that: include: a detection unit, when detecting that a k-th target word in the target text is a word of the target language and the k-th target word includes a preset keyword, detecting whether the k-th target word includes other characters except the preset keyword, wherein k is an integer greater than or equal to 1; A first determining unit, when detecting that the k-th target word includes other characters except the preset keyword, determines that the k-th target word fails to match the preset keyword; The second determining unit determines that the kth target word successfully matches the preset keyword when it is detected that the kth target word does not include other characters except the preset keyword.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein when the computer program is executed, the device where the computer-readable storage medium is located executes the keyword matching method according to any one of claims 1 to 7.

10. An electronic device, characterized in that: It includes one or more processors and a memory, wherein the memory is used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors execute the keyword matching method described in any one of claims 1 to 7.