Regular expression generation method and device based on large language model

By adopting a regular expression generation method based on a large language model in intention recognition, and generating and storing regular expressions corresponding to the target intention label, the problems of low accuracy and inefficiency in the prior art are solved, and more efficient and accurate intention recognition is achieved.

CN120011524AActive Publication Date: 2025-05-16SHENZHEN WEIAIZHIYUN TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510487076.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-05-16
Estimated Expiration
2045-04-18

AI Technical Summary

Technical Problem

In the prior art, the accuracy of intention recognition is not high and the efficiency is low, making it difficult to quickly and accurately identify user intentions in actual business scenarios.

Method used

The regular expression generation method based on the large language model is adopted. By selecting candidate samples with the same intent tag in the sample library to form an intent document, extracting intent keywords and clustering, using the large language model to process the intent keyword set, generating regular expressions corresponding to the target intent tag, and storing them in the regular library.

Benefits of technology

The efficiency and accuracy of intention recognition are improved, so that user Q&A intentions can be quickly identified during the Q&A interactive stage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011524A_ABST
    Figure CN120011524A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a regular expression generation method and device based on a large language model, and the regular expression generation method based on the large language model comprises the steps: selecting candidate samples with the same intention tag from a sample library to form a plurality of intention documents, the candidate sample is composed of text data and an intention label corresponding to the text data; extracting an intention keyword corresponding to each intention document, and clustering the intention keywords to obtain an intention keyword set corresponding to the target intention tag; processing the intention keyword set by utilizing a large language model according to a target cue word to obtain a regular expression corresponding to the target intention tag; the regular expression is stored in a regular library, and the regular expression stored in the regular library is used for recognizing the question and answer intention of the user in the question and answer interaction stage.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this specification relate to the field of machine learning technology, and more particularly to a method and device for generating regular expressions based on a large language model. Background Art

[0002] With the development of computer and Internet technology, intent recognition has been applied in more and more scenarios, mainly to understand the user's intention. For example, in intelligent customer service scenarios, intent recognition is widely used in customer service robots. By identifying the intention of user input, the robot can quickly respond to user needs. In addition, the accuracy of intent recognition plays an important role in scenarios such as e-commerce intelligent customer service and virtual assistants. In the existing technology, most rule-based methods of intent recognition rely on dictionary matching, but lack flexibility; traditional machine learning classifies through statistical features, and deep learning (such as BERT, LSTM) improves accuracy by automatically capturing semantics. Although the above implementation can achieve the purpose of intent recognition, the accuracy of intent recognition is not high in actual business scenarios, and the efficiency is low. Therefore, an effective solution is urgently needed to solve the above problems. Summary of the invention

[0003] In view of this, the embodiments of this specification provide a method for generating a regular expression based on a large language model. One or more embodiments of this specification also relate to a regular expression generating device based on a large language model, a data processing method, a data processing device, a computing device, a computer-readable storage medium and a computer program product to solve the technical defects existing in the prior art.

[0004] According to a first aspect of an embodiment of this specification, a method for generating a regular expression based on a large language model is provided, comprising: Select candidate samples with the same intent label in the sample library to form multiple intent documents, wherein the candidate samples are formed based on text data and intent labels corresponding to the text data; Extract the intent keywords corresponding to each intent document, and obtain the intent keyword set corresponding to the target intent label by clustering the intent keywords; Using a large language model to process the intent keyword set according to the target prompt word, to obtain a regular expression corresponding to the target intent label; The regular expression is stored in a regular library, wherein the regular expression stored in the regular library is used to identify the user's question and answer intention in the question and answer interaction stage.

[0005] According to a second aspect of an embodiment of this specification, a data processing method is provided, including: Get the question text data submitted by the user; Performing regular matching on the question text data using the regular expressions stored in the regular library, wherein the regular expressions stored in the regular library are constructed according to the above method; Determine the target regular expression corresponding to the question text data according to the matching result, and use the intent label corresponding to the target regular expression as the intent label corresponding to the question text data; Answer text data is generated based on the intention label corresponding to the question text data, and is displayed to the user.

[0006] According to a third aspect of the embodiments of this specification, a regular expression generation device based on a large language model is provided, comprising: A selection module is configured to select candidate samples with the same intent labels from a sample library to form a plurality of intent documents, wherein the candidate samples are formed based on text data and intent labels corresponding to the text data; An extraction module is configured to extract intent keywords corresponding to each intent document, and obtain an intent keyword set corresponding to a target intent tag by clustering the intent keywords; A processing module is configured to process the intention keyword set according to the target prompt word using a large language model to obtain a regular expression corresponding to the target intention label; The storage module is configured to store the regular expression in a regular library, wherein the regular expression stored in the regular library is used to identify the user's question and answer intention in the question and answer interaction stage.

[0007] According to a fourth aspect of an embodiment of this specification, there is provided a data processing device, including: An acquisition module is configured to acquire question text data submitted by a user; A matching module is configured to perform regular matching on the question text data using a regular expression stored in a regular library, wherein the regular expression stored in the regular library is constructed according to the above method; A determination module is configured to determine a target regular expression corresponding to the question text data according to the matching result, and use the intent tag corresponding to the target regular expression as the intent tag corresponding to the question text data; The display module is configured to generate answer text data based on the intention label corresponding to the question text data and display it to the user.

[0008] According to a fifth aspect of an embodiment of this specification, a computing device is provided, including: Memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the above-mentioned regular expression generation method or data processing method based on a large language model are implemented.

[0009] According to the sixth aspect of the embodiments of this specification, a computer-readable storage medium is provided, which stores computer-executable instructions, and when the instructions are executed by a processor, the steps of the above-mentioned regular expression generation method or data processing method based on a large language model are implemented.

[0010] According to the seventh aspect of the embodiments of this specification, a computer program product is provided, including a computer program or instructions, which, when executed by a processor, implement the steps of the above-mentioned regular expression generation method or data processing method based on a large language model.

[0011] The regular expression generation method based on the large language model provided in this embodiment, in order to be able to construct a regular expression that can quickly and accurately identify the intent of text data, can first select candidate samples with the same intent label in the sample library to form multiple intent documents, wherein the candidate samples are composed of text data and the intent label corresponding to the text data; on this basis, the intent keywords corresponding to each intent document can be extracted, and the intent keywords can be clustered to obtain the intent keyword set corresponding to the target intent label, so that the keywords expressing the intent information corresponding to the target intent label can be integrated before the regular expression is constructed, so that the subsequent regular expression construction corresponding to the target intent label is more accurate. After that, the large language model can be used to process the intent keyword set according to the target prompt word, and the regular expression corresponding to the target intent label can be obtained according to the processing result; finally, the regular expression can be stored in the regular library, so that in the application stage, the regular expression stored in the regular library can be used to identify the user's question and answer intention in the question and answer interaction stage. Thereby effectively improving the efficiency and accuracy of intent recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 It is a flowchart of a method for generating a regular expression based on a large language model provided by an embodiment of this specification; Figure 2 is a flow chart of a data processing method provided by an embodiment of this specification; Figure 3 It is a structural diagram of a regular expression generation device based on a large language model provided by an embodiment of this specification; Figure 4 is a structural schematic diagram of a data processing device provided by an embodiment of this specification; Figure 5It is a structural block diagram of a computing device provided by an embodiment of this specification. DETAILED DESCRIPTION

[0013] Many specific details are described in the following description to facilitate a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the connotation of this specification, so this specification is not limited to the specific implementation disclosed below.

[0014] The terms used in one or more embodiments of this specification are only for the purpose of describing specific embodiments, and are not intended to limit one or more embodiments of this specification. The singular forms of "a", "said" and "the" used in one or more embodiments of this specification and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.

[0015] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of this specification, this information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".

[0016] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0017] First, the terms involved in one or more embodiments of this specification are explained.

[0018] Large language model technology is usually based on deep neural networks, especially the Transformer architecture, which uses massive amounts of data for training and can learn complex language patterns and semantic relationships. Large language models can capture the complex relationships between words through the self-attention mechanism and extract semantic information at multiple levels to understand the deep meaning of sentences. Large language model parameters are often very large, but this is also why large language models have context awareness and deep semantic expression capabilities.

[0019] Regular expression extraction tasks require the identification of specific patterns or structures from text, and usually require in-depth semantic understanding and pattern matching of the text. This is extremely consistent with the characteristics of large language models, because large language models can better capture the context and semantic relationships of text, providing higher flexibility and robustness than traditional regular expressions. Through the self-learning model, it can cope with complex and changing inputs, has strong generalization capabilities, and can automatically process large-scale data to reduce human intervention. Therefore, large language model technology is very suitable for regular expression extraction tasks.

[0020] In this specification, a method for generating a regular expression based on a large language model is provided. One or more embodiments of this specification also relate to a regular expression generating device based on a large language model, a data processing method, a data processing device, a computing device, a computer-readable storage medium, and a computer program product, which are described in detail one by one in the following embodiments.

[0021] See also Figure 1 , Figure 1 A flowchart of a method for generating a regular expression based on a large language model according to an embodiment of the present specification is shown, which specifically includes the following steps.

[0022] Step S102 , selecting candidate samples with the same intent labels in the sample library to form a plurality of intent documents, wherein the candidate samples are formed based on text data and intent labels corresponding to the text data.

[0023] The regular expression generation method based on a large language model provided in this embodiment can be applied to any scenario with intent recognition, such as intelligent customer service scenarios (by recognizing the intent of user input, the robot can quickly respond to user needs), e-commerce intelligent customer service scenarios (the intelligent customer service system automatically responds to questions by recognizing the user's intentions such as shopping and order inquiries, thereby reducing the cost of manual customer service and improving customer experience), virtual assistant scenarios (understanding user instructions through intent recognition to complete specific tasks. For example, when the user says "set an alarm" or "play music", the virtual assistant performs corresponding operations through intent recognition), medical and health assistant scenarios (through voice or text input, users can consult health issues or make appointments with doctors, and medical assistants provide users with corresponding help through intent recognition, such as diagnosing symptoms, prescribing, or making appointments), etc.

[0024] This embodiment describes the application of a regular expression generation method based on a large language model in a virtual assistant scenario as an example. The descriptions of other scenarios can be found in the description of this embodiment, and this embodiment will not be elaborated in detail here.

[0025] Specifically, the sample library refers to a database that stores candidate samples, wherein candidate samples refer to samples composed of text data and intent labels corresponding to the text data, text data refers to text that needs to be identified for intent, and intent labels refer to intent information corresponding to the text data. Correspondingly, the intent document refers to a document composed of candidate samples with the same intent labels selected from the sample library. Selecting candidate samples with the same intent labels to construct an intent document can ensure that the text data in the candidate samples contained in the intent document have the same intent, which facilitates the subsequent construction of regular expressions for the intent corresponding to each intent document.

[0026] Based on this, in order to build a regular expression that can quickly and accurately identify the intent of text data, we can first select candidate samples with the same intent label in the sample library to form multiple intent documents, where the candidate samples are based on text data and the intent label corresponding to the text data; on this basis, we can extract the intent keywords corresponding to each intent document, and cluster the intent keywords to obtain the intent keyword set corresponding to the target intent label, so that the keywords that express the intent information corresponding to the target intent label can be integrated before the regular expression is constructed, so that the subsequent regular expression construction corresponding to the target intent label is more accurate. After that, the large language model can be used to process the intent keyword set according to the target prompt word, and the regular expression corresponding to the target intent label can be obtained according to the processing result; finally, the regular expression can be stored in the regular library, so that in the application stage, the regular expression stored in the regular library can be used to identify the user's question and answer intention in the question and answer interaction stage. Thereby effectively improving the efficiency and accuracy of intent recognition.

[0027] In one or more implementations of this embodiment, before the step of selecting candidate samples with the same intent label in the sample library to form multiple intent documents is performed, the method further includes: Acquire text data, and perform regular matching on the text data using the regular expressions stored in the regular library; In case of successful matching, a target regular expression corresponding to the text data is determined; the intent tag corresponding to the target regular expression is used as the intent tag corresponding to the text data, and a question-answering interaction task is performed based on the intent tag corresponding to the text data.

[0028] In the event of a matching failure, the text data is input into the large language model for intent recognition to obtain an intent label corresponding to the text data; candidate samples are constructed based on the text data and the intent label corresponding to the text data, and the candidate samples are stored in the sample library.

[0029] Specifically, the target regular expression refers to the regular expression obtained after regular matching of text data, and its corresponding intent label is the label that represents the intent information corresponding to the regular expression. If the text data matches the target regular expression, it means that the intent label corresponding to the target regular expression can represent the intent of the text data. Correspondingly, the question-answering interactive task specifically refers to the task of answering text data based on the intent label of the text data, which can be achieved through a large language model.

[0030] Based on this, before constructing a regular expression, the text data needs to be processed to determine whether the text data can be used to construct a regular expression corresponding to a new intent. Therefore, the text data can be obtained and the regular expression stored in the regular library can be used to perform regular matching on the text data.

[0031] If the match is successful, it means that the regular expression corresponding to the intent of the text data has been stored in the regular library, so there is no need to use the text data to build a new regular expression. Therefore, the target regular expression corresponding to the text data can be determined; the intent label corresponding to the target regular expression is used as the intent label corresponding to the text data, and the question-and-answer interaction task is performed based on the intent label corresponding to the text data.

[0032] In the case of a match failure, it means that the regular expression corresponding to the intent of the text data has not been constructed, so a regular expression construction process is required, which can be implemented in conjunction with a large language model. Specifically, the text data can be input into a large language model for intent recognition, so that the intent label corresponding to the text data can be obtained based on the recognition result; thereafter, candidate samples are constructed based on the text data and the intent label corresponding to the text data, and the candidate samples are stored in a sample library, which can be used to subsequently select candidate samples with the same intent label from the sample library to form an intent document, which is used to construct a regular expression for the specified intent and is stored in the regular library for use in actual business scenarios.

[0033] For example, after obtaining the text data "Check today's weather", you can first use the regular expression stored in the regular library to perform regular matching on the text data; if the match is successful, you can extract the target regular expression corresponding to the text data from the regular library. For example, if the regular expression is {r"(query|view|understand)(.*?)(today|tomorrow|the day after tomorrow)?(.*?)weather"}, this regular expression can match keywords such as "query", "view", "weather", and "today". The intent corresponding to this regular expression is "query today's weather". Therefore, it can be determined that the intent of "check today's weather" is "query today's weather". Subsequently, the answer can be fed back for the text data based on the intent information.

[0034] If the match fails, it means that the intent of the text data cannot be identified through regular expressions. Therefore, the text data can be input into a large language model for intent recognition. The recognition result of the model determines that the intent corresponding to the text data is "query today's weather". Then, based on the text data "Check today's weather" and the intent "Query today's weather", a candidate sample {Check today's weather - Query today's weather} can be constructed and stored in the sample library.

[0035] Furthermore, when it is necessary to construct a regular expression, candidate samples with the same intent can be extracted from the sample library, and intent documents corresponding to different intents can be obtained by splicing the candidate samples. For example, by selecting candidate samples with the intent of "query today's weather" from the sample library, {How is today's weather - query today's weather}, {Is it sunny today - query today's weather}, {Will it rain today - query today's weather}..., at this time, the candidate samples can be spliced ​​to obtain the intent document corresponding to the intent of "query today's weather", that is, one document represents one intent, and so on. After obtaining multiple documents corresponding to multiple intents, regular expressions corresponding to different intents can be constructed and stored in the regular library, so that when the virtual assistant interacts with the user in question and answer, it can quickly identify the user's intent and perform question and answer operations.

[0036] In summary, regular expressions are first used to match text data to determine whether the intent corresponding to the text data is recognized, so as to select different links for subsequent processing according to the recognition results, thereby making data processing more efficient and reasonable.

[0037] Step S104, extracting the intent keywords corresponding to each intent document, and obtaining the intent keyword set corresponding to the target intent tag by clustering the intent keywords.

[0038] Specifically, after extracting candidate samples with the same intent label from the sample library to form intent documents, in order to accurately construct the regular expression corresponding to each intent, the intent keywords corresponding to each intent document can be extracted, and the intent keywords can be clustered to achieve deduplication and improve keyword accuracy, thereby obtaining a set of intent keywords corresponding to the target intent label for subsequent use.

[0039] Among them, the intent keywords specifically refer to the keywords extracted from each intent document, and the intent keywords corresponding to each intent document come from different candidate samples in the intent document. Correspondingly, the target intent label specifically refers to the intent label corresponding to each intent document, and the intent keyword set is the intent keywords obtained after clustering the intent keywords contained in the intent document. For example, intent document a contains intent keywords a1, a2, a3 and a4, and its corresponding intent label is A. By clustering intent keywords a1, a2, a3 and a4, we can get the intent keyword set {a1, a2, a4} corresponding to intent label A.

[0040] In one or more implementations of this embodiment, determining the intent keyword corresponding to any one of the multiple intent documents includes: Determine the intent document to be processed; use a keyword extraction algorithm to extract multiple initial intent keywords from the intent document to be processed; use a keyword deduplication algorithm to deduplicate the multiple initial intent keywords, and determine the intent keyword corresponding to the intent document to be processed based on the deduplication result.

[0041] Specifically, the intent document to be processed specifically refers to any one of the multiple intent documents. Correspondingly, the keyword extraction algorithm specifically refers to the TF-IDF algorithm, and the initial intent keyword specifically refers to the intent keyword extracted from the intent document that has not been deduplicated; correspondingly, the keyword deduplication algorithm specifically refers to the MMR algorithm.

[0042] Based on this, when performing keyword extraction for any intent document, the intent document to be processed can be determined first, and then the keyword extraction algorithm can be used to extract multiple initial intent keywords from the intent document to be processed. After obtaining multiple initial intent keywords, the keyword deduplication algorithm can be used to deduplicate the multiple initial intent keywords, and then the intent keywords corresponding to the intent document to be processed can be determined according to the deduplication results for subsequent use.

[0043] In summary, by using keyword extraction algorithm and keyword deduplication algorithm to determine the intended keywords, not only can the accuracy of keyword extraction be guaranteed, but also the subsequent calculation pressure caused by keyword redundancy can be avoided, thereby effectively improving the accuracy of regular expression construction.

[0044] In one or more implementations of this embodiment, the deduplication processing of the multiple initial intent keywords is performed using a keyword deduplication algorithm, and the intent keyword corresponding to the intent document to be processed is determined according to the deduplication processing result, including: A keyword deduplication algorithm is used to calculate the cosine similarity of the multiple initial intent keywords; the intent keywords to be deleted are determined based on the cosine similarity calculation results, and the intent keywords to be deleted are deleted from the multiple initial intent keywords; based on the deletion results, the remaining initial intent keywords are used as the intent keywords corresponding to the intent document to be processed.

[0045] Specifically, the intent keywords to be deleted specifically refer to the intent keywords that are repeated with other intent keywords among the multiple initial intent keywords.

[0046] Based on this, when using the keyword deduplication algorithm for deduplication processing, the keyword deduplication algorithm can be used to calculate the cosine similarity of multiple initial intent keywords; thereafter, the intent keywords to be deleted can be determined based on the cosine similarity calculation results, and the intent keywords to be deleted can be deleted from the multiple initial intent keywords, so that the remaining initial intent keywords can be used as the intent keywords corresponding to the intent document to be processed based on the deletion results, for subsequent use.

[0047] In the specific implementation, the TF-IDF algorithm is used to extract high-frequency keywords from the intent document, and keywords related to the intent tag corresponding to the intent document can be obtained. On this basis, the maximum marginal relevance (MMR) algorithm is used to filter the keywords to achieve the purpose of removing redundant and repeated keywords, thereby determining the intent keywords corresponding to each intent document and then performing subsequent regular expression construction.

[0048] It should be noted that the MMR algorithm can be implemented by the following formula (1) when performing deduplication processing: (1) Among them, S is a number of initial intent keywords extracted by the TF-IDF algorithm. is the set of keywords that have been selected. is one of the initial intent keywords. For the query (or target text, or textual representation of the intent category), For keywords With query The correlation between them is generally calculated using TF-IDF cosine similarity; For keywords With a selected keyword The similarity between them is usually calculated using cosine similarity. A weight parameter that controls relevance and diversity (usually set between 0.5 and 0.7).

[0049] It can be understood that the MMR algorithm is a greedy algorithm, which continuously finds the keyword that maximizes the formula among multiple initial intent keywords. ,Right now , by using the keyword as the intent keyword corresponding to the intent document, and then performing subsequent intent keyword extraction processing. Until the number of keywords reaches the set threshold, or the similarity between the newly added keywords and the extracted intent keywords is too high, the keyword extraction operation can be stopped, and then subsequent processing can be performed.

[0050] In summary, by calculating the cosine similarity to remove duplicate keywords, we can ensure that the remaining keywords are unique and associated with the target intent label, so that the subsequent regular expression construction is more accurate.

[0051] In one or more implementations of this embodiment, determining the set of intent keywords corresponding to the target intent tag includes: Determine multiple intent keywords to be clustered corresponding to the target intent label, wherein the multiple intent keywords to be clustered belong to the intent document corresponding to the target intent label; perform semantic recognition on the multiple intent keywords to be clustered respectively through a semantic recognition model to obtain a semantic vector corresponding to each intent keyword to be clustered; use a clustering algorithm to process the semantic vector corresponding to each intent keyword to be clustered, and construct an intent keyword set corresponding to the target intent label according to the processing result.

[0052] Specifically, the intent keywords to be clustered specifically refer to the intent keywords extracted from any intent document, and these intent keywords are all associated with the target intent label. Correspondingly, the semantic recognition model specifically refers to a model that constructs the semantic features corresponding to each intent keyword to be clustered, which can be implemented using the BERT model. Correspondingly, the semantic vector is the semantic feature expression corresponding to each intent keyword to be clustered. Correspondingly, the clustering algorithm can be implemented using K-means.

[0053] Based on this, when constructing an intent keyword set for the target intent tag corresponding to any intent document, we can first determine the multiple intent keywords to be clustered corresponding to the target intent tag, and the multiple intent keywords to be clustered belong to the intent document corresponding to the target intent tag; thereafter, in order to improve processing efficiency and accuracy, the semantic recognition model can be used to perform semantic recognition on the multiple intent keywords to be clustered respectively to obtain the semantic vector corresponding to each intent keyword to be clustered; at this time, the clustering algorithm is used to process the semantic vector corresponding to each intent keyword to be clustered, and the intent keyword set corresponding to the target intent tag can be constructed according to the processing result.

[0054] In specific implementation, when constructing a semantic vector for each keyword to be clustered through the semantic recognition model, it can be expressed by the following formula (2): (2) The clustering process can be expressed by the following formula (3): (3) in, Represents the total Intention, Indicates a high-frequency keyword extracted from a document under a certain intent (serial number j). Indicates the high-frequency keyword clustering center under this intent recognition, Represents the total set of keywords under intent j.

[0055] For example, after obtaining intent document a corresponding to intent label A and intent document b corresponding to intent label B, the TF-IDF algorithm can be used to extract high-frequency keywords from intent document a and intent document b. This can be understood as extracting high-frequency keywords that are exactly related to the intent label, rather than keywords that are high-frequency in all documents. According to the processing results, it is determined that the keywords contained in intent document a are {a1, a2, a3…an}, and the keywords contained in intent document b are {b1, b2, a3…bm}.

[0056] Furthermore, after obtaining the keywords corresponding to each intent document, the MMR algorithm can be used to deduplicate the keywords. The specific processing can be found in the above description. According to the processing results, the initial keyword set corresponding to intent label A is determined to be {a1, a2, a3, a5, a6, a7, a8}, and the initial keyword set corresponding to intent label B is determined to be {b1, b5, b6, b9, b11}.

[0057] Furthermore, by using the pre-trained model BERT to embed the keywords corresponding to each intent label into the semantic space, we can get the embedding vector corresponding to each keyword in the keyword set, and then use the clustering algorithm K-means to cluster the keywords with the same semantics to get the keywords corresponding to the intent label. Determine that the keyword set corresponding to intent label A is {a6, a7, a8}, and the keyword set corresponding to intent label B is {b1, b5, b6}. Then, we can construct the regular expression for each intent label according to the keyword set.

[0058] In summary, by adopting the method of semantic recognition when clustering intent keywords, the intent keywords finally determined can have the same intent, thereby effectively improving the accuracy of constructing regular expressions.

[0059] Step S106: Process the intention keyword set according to the target prompt word using a large language model to obtain a regular expression corresponding to the target intention label.

[0060] Step S108: storing the regular expression in a regular library, wherein the regular expression stored in the regular library is used to identify the user's question-answering intention during the question-answering interaction stage.

[0061] Specifically, after obtaining the intent keyword set corresponding to the target intent label as mentioned above, in order to improve the efficiency of constructing regular expressions, a large language model can be used to process the intent keyword set according to the target prompt word, and then obtain the regular expression corresponding to the target intent label. Thereafter, the regular expression can be stored in the regular library, so that the regular expressions stored in the regular library can be used to identify user question and answer intentions in the question and answer interaction stage.

[0062] The target prompt word specifically refers to the prompt word configured after optimization using the prompt word optimization project, which enables the large language model to construct a regular expression corresponding to each intent tag according to the prompt word. Correspondingly, the regular library specifically refers to a database storing regular expressions corresponding to different intent tags.

[0063] In actual applications, when using a large language model to extract the regular expression corresponding to each intent tag according to the target prompt word, white regular expressions and black regular expressions can be extracted separately to improve the pertinence and accuracy of regular expressions. Among them, white regular expressions can be used to match standard formats, such as email addresses, ID card numbers, etc., while black regular expressions can be used to process unstructured data, such as log cleaning.

[0064] Using the above example, after obtaining the keyword set {a6, a7, a8} corresponding to the intent label A and the keyword set {b1, b5, b6} corresponding to the intent label B, the large language model can be used to construct regular expressions for the keyword set {a6, a7, a8} and the keyword set {b1, b5, b6} according to the target prompt words, and then the regular expressions corresponding to the intent labels A and B can be obtained. For example, if the intent label A is "query today's weather", the regular expression obtained is {r"(query|view|understand)(.*?)(today|tomorrow|the day after tomorrow)?(.*?)weather"}, and if the intent label B is "open the calculator", the regular expression obtained is {r"(open|start|run)(.*?)(calculator|Calculator)"}. After storing it in the regular library, if the user mentions questions about the intent label A or B during the Q&A interaction with the virtual assistant, the user's intention can be directly determined through the corresponding regular expression, and the weather information or the operation of starting the calculator can be fed back to the user.

[0065] In one or more implementations of this embodiment, after the step of storing the regular expression in a regular library is performed, the following step is further included: Matching text data with unrecognized intent in the text database with regular expressions stored in the regular library; and determining and deleting associated text data in the text database according to the matching result.

[0066] In addition, considering that the construction of regular expressions needs to be combined with a large amount of text data contained in the text database in order to complete the construction of regular expressions with a wider coverage, after each regular expression is constructed, in order to avoid redundant construction operations, that is, repeated regular expression construction, the text data with unrecognized intent in the text database can also be matched with the regular expressions stored in the regular library; thereby, the associated text data in the text database is determined and deleted according to the matching results, so as to reduce the amount of data in the text database and improve the efficiency of regular expression construction.

[0067] In summary, in order to construct a regular expression that can quickly and accurately identify the intent of text data, we can first select candidate samples with the same intent label in the sample library to form multiple intent documents, where the candidate samples are composed of text data and the intent label corresponding to the text data; on this basis, we can extract the intent keywords corresponding to each intent document, and cluster the intent keywords to obtain the intent keyword set corresponding to the target intent label, so that the keywords expressing the intent information corresponding to the target intent label can be integrated before the regular expression is constructed, so that the subsequent regular expression construction corresponding to the target intent label is more accurate. After that, the large language model can be used to process the intent keyword set according to the target prompt word, and the regular expression corresponding to the target intent label can be obtained according to the processing result; finally, the regular expression can be stored in the regular library, so that in the application stage, the regular expression stored in the regular library can be used to identify the user's question and answer intention in the question and answer interaction stage. Thereby effectively improving the efficiency and accuracy of intent recognition.

[0068] See also Figure 2 , Figure 2 A flow chart of a data processing method provided according to an embodiment of the present specification is shown, which specifically includes the following steps.

[0069] Step S202, obtaining question text data submitted by the user.

[0070] Step S204, performing regular matching on the question text data using the regular expressions stored in the regular library, wherein the regular expressions stored in the regular library are constructed according to the above method.

[0071] Step S206, determining the target regular expression corresponding to the question text data according to the matching result, and using the intention label corresponding to the target regular expression as the intention label corresponding to the question text data.

[0072] Step S208, generating answer text data based on the intention label corresponding to the question text data, and displaying the answer text data to the user.

[0073] Specifically, question text data refers to the question text submitted by the user in the current business scenario, which includes but is not limited to e-commerce intelligent customer service scenarios (the intelligent customer service system automatically responds to questions by identifying the user's intentions such as shopping and order inquiries, thereby reducing the cost of manual customer service and improving customer experience), virtual assistant scenarios (understanding the user's instructions through intent recognition to complete specific tasks. For example, the user says "set an alarm" or "play music", and the virtual assistant performs corresponding operations through intent recognition), medical and health assistant scenarios (through voice or text input, users can consult health issues or make appointments with doctors, and medical assistants provide users with corresponding help through intent recognition, such as diagnosing symptoms, prescribing, or making appointments), etc.

[0074] Correspondingly, the answer text data specifically refers to the answer text displayed to the user according to the intent information after the intent is recognized on the question text data submitted by the user. In specific implementation, the answer text data can be generated by a large language model.

[0075] For example, in an intelligent customer service scenario, when receiving a question text submitted by a user, "When will my order arrive?", the regular expression stored in the regular library constructed using the same method as above can be used to identify the intent of the question text, thereby determining that the intent of the question text "When will my order arrive?" is "querying the order delivery status" based on the regular matching result; on this basis, the order information of the goods purchased by the user can be obtained, and the delivery status can be determined based on the order information, and the delivery status of the goods can be displayed to the user.

[0076] In summary, by using regular expressions stored in the regular library to quickly identify user intentions, the efficiency of handling user questions can be effectively improved, thereby improving the user experience.

[0077] Corresponding to the above method embodiment, this specification also provides an embodiment of a regular expression generation device based on a large language model, Figure 3 FIG. 1 shows a schematic diagram of a regular expression generation device based on a large language model provided by an embodiment of the present specification. Figure 3 As shown, the device comprises: A selection module 302 is configured to select candidate samples with the same intent label from the sample library to form a plurality of intent documents, wherein the candidate samples are formed based on text data and intent labels corresponding to the text data; An extraction module 304 is configured to extract intent keywords corresponding to each intent document, and obtain an intent keyword set corresponding to a target intent tag by clustering the intent keywords; The processing module 306 is configured to process the intention keyword set according to the target prompt word using a large language model to obtain a regular expression corresponding to the target intention label; The storage module 308 is configured to store the regular expression in a regular library, wherein the regular expression stored in the regular library is used to identify the user's question and answer intention in the question and answer interaction stage.

[0078] In an optional embodiment, before the step of selecting candidate samples with the same intent label in the sample library to form multiple intent documents is performed, the step further includes: Acquire text data, and use the regular expressions stored in the regular library to perform regular matching on the text data; if the match is successful, determine the target regular expression corresponding to the text data; use the intent label corresponding to the target regular expression as the intent label corresponding to the text data, and perform a question-answering interaction task based on the intent label corresponding to the text data.

[0079] In an optional embodiment, when the matching fails, the method further includes: The text data is input into the large language model for intent recognition to obtain an intent label corresponding to the text data; a candidate sample is constructed based on the text data and the intent label corresponding to the text data, and the candidate sample is stored in the sample library.

[0080] In an optional embodiment, determining the intent keyword corresponding to any one of the multiple intent documents includes: Determine the intent document to be processed; use a keyword extraction algorithm to extract multiple initial intent keywords from the intent document to be processed; use a keyword deduplication algorithm to deduplicate the multiple initial intent keywords, and determine the intent keyword corresponding to the intent document to be processed based on the deduplication result.

[0081] In an optional embodiment, determining the set of intent keywords corresponding to the target intent tag includes: Determine multiple intent keywords to be clustered corresponding to the target intent label, wherein the multiple intent keywords to be clustered belong to the intent document corresponding to the target intent label; perform semantic recognition on the multiple intent keywords to be clustered respectively through a semantic recognition model to obtain a semantic vector corresponding to each intent keyword to be clustered; use a clustering algorithm to process the semantic vector corresponding to each intent keyword to be clustered, and construct an intent keyword set corresponding to the target intent label according to the processing result.

[0082] In an optional embodiment, after the step of storing the regular expression in a regular library is executed, the method further includes: Matching text data with unrecognized intent in the text database with regular expressions stored in the regular library; and determining and deleting associated text data in the text database according to the matching result.

[0083] In an optional embodiment, the deduplication processing of the multiple initial intent keywords is performed using a keyword deduplication algorithm, and the intent keyword corresponding to the intent document to be processed is determined according to the deduplication processing result, including: A keyword deduplication algorithm is used to calculate the cosine similarity of the multiple initial intent keywords; the intent keywords to be deleted are determined based on the cosine similarity calculation results, and the intent keywords to be deleted are deleted from the multiple initial intent keywords; based on the deletion results, the remaining initial intent keywords are used as the intent keywords corresponding to the intent document to be processed.

[0084] In summary, in order to construct a regular expression that can quickly and accurately identify the intent of text data, we can first select candidate samples with the same intent label in the sample library to form multiple intent documents, where the candidate samples are composed of text data and the intent label corresponding to the text data; on this basis, we can extract the intent keywords corresponding to each intent document, and cluster the intent keywords to obtain the intent keyword set corresponding to the target intent label, so that the keywords expressing the intent information corresponding to the target intent label can be integrated before the regular expression is constructed, so that the subsequent regular expression construction corresponding to the target intent label is more accurate. After that, the large language model can be used to process the intent keyword set according to the target prompt word, and the regular expression corresponding to the target intent label can be obtained according to the processing result; finally, the regular expression can be stored in the regular library, so that in the application stage, the regular expression stored in the regular library can be used to identify the user's question and answer intention in the question and answer interaction stage. Thereby effectively improving the efficiency and accuracy of intent recognition.

[0085] The above is a schematic scheme of a regular expression generation device based on a large language model of this embodiment. It should be noted that the technical scheme of the regular expression generation device based on a large language model and the technical scheme of the regular expression generation method based on a large language model belong to the same concept, and the details not described in detail in the technical scheme of the regular expression generation device based on a large language model can be found in the description of the technical scheme of the regular expression generation method based on a large language model.

[0086] Corresponding to the above method embodiment, this specification also provides a data processing device embodiment, Figure 4 FIG. 1 is a schematic diagram showing the structure of a data processing device provided by an embodiment of the present specification. Figure 4 As shown, the device comprises: The acquisition module 402 is configured to acquire question text data submitted by the user; A matching module 404 is configured to perform regular matching on the question text data using a regular expression stored in a regular library, wherein the regular expression stored in the regular library is constructed according to the above method; The determination module 406 is configured to determine the target regular expression corresponding to the question text data according to the matching result, and use the intention label corresponding to the target regular expression as the intention label corresponding to the question text data; The display module 408 is configured to generate answer text data based on the intention tag corresponding to the question text data, and display it to the user.

[0087] In summary, by using regular expressions stored in the regular library to quickly identify user intentions, the efficiency of handling user questions can be effectively improved, thereby improving the user experience.

[0088] The above is a schematic scheme of a data processing device of this embodiment. It should be noted that the technical scheme of the data processing device and the technical scheme of the above data processing method belong to the same concept, and the details of the technical scheme of the data processing device that are not described in detail can be referred to the description of the technical scheme of the above data processing method.

[0089] Figure 5 The structure block diagram of a computing device 500 provided according to an embodiment of the present specification is shown. The components of the computing device 500 include but are not limited to a memory 510 and a processor 520. The processor 520 is connected to the memory 510 via a bus 530, and the database 550 is used to store data.

[0090] The computing device 500 also includes an access device 540 that enables the computing device 500 to communicate via one or more networks 560. Examples of these networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 540 may include one or more of any type of network interface (e.g., a network interface card (NIC)) that is wired or wireless, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a world-wide interoperability for microwave access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, and a near field communication (NFC).

[0091] In one embodiment of the present specification, the above components of the computing device 500 and Figure 5 Other components not shown in the figure may also be connected to each other, for example, via a bus. It should be understood that Figure 5 The computing device structure block diagram shown is only for the purpose of illustration, and is not intended to limit the scope of this specification. Those skilled in the art can add or replace other components as needed.

[0092] The computing device 500 may be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook, etc.), a mobile phone (e.g., a smart phone), a wearable computing device (e.g., a smart watch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or a personal computer (PC). The computing device 500 may also be a mobile or stationary server.

[0093] The processor 520 is used to execute the following computer executable instructions, which, when executed by the processor, implement the steps of the above-mentioned regular expression generation method or data processing method based on a large language model.

[0094] The above is a schematic scheme of a computing device of this embodiment. It should be noted that the technical scheme of the computing device and the technical scheme of the regular expression generation method or data processing method based on a large language model are of the same concept, and the details not described in detail in the technical scheme of the computing device can be found in the description of the technical scheme of the regular expression generation method or data processing method based on a large language model.

[0095] An embodiment of the present specification also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the steps of the above-mentioned regular expression generation method or data processing method based on a large language model.

[0096] The above is a schematic scheme of a computer-readable storage medium of this embodiment. It should be noted that the technical scheme of the storage medium and the technical scheme of the regular expression generation method or data processing method based on a large language model belong to the same concept, and the details not described in detail in the technical scheme of the storage medium can be found in the description of the technical scheme of the regular expression generation method or data processing method based on a large language model.

[0097] An embodiment of the present specification also provides a computer program product, including a computer program or instructions, which, when executed by a processor, implements the steps of the above-mentioned regular expression generation method or data processing method based on a large language model.

[0098] The above is a schematic scheme of a computer program product of this embodiment. It should be noted that the technical scheme of the computer program product and the technical scheme of the regular expression generation method or data processing method based on a large language model belong to the same concept, and the details not described in detail in the technical scheme of the computer program product can be found in the description of the technical scheme of the regular expression generation method or data processing method based on a large language model.

[0099] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0100] The computer instructions include computer program codes, which may be in source code form, object code form, executable files or some intermediate forms, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium, etc. It should be noted that the content contained in the computer-readable medium may be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.

[0101] It should be noted that, for the above-mentioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of this specification are not limited by the order of the actions described, because according to the embodiments of this specification, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the embodiments of this specification.

[0102] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0103] The preferred embodiments of this specification disclosed above are only used to help explain this specification. The optional embodiments do not describe all the details in detail, nor do they limit the invention to the specific implementation methods described. Obviously, many modifications and changes can be made according to the content of the embodiments of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of this specification, so that technicians in the relevant technical field can understand and use this specification well.

Claims

1. A regular expression generation method based on a large language model, characterized in that: include: Select candidate samples with the same intent label in the sample library to form multiple intent documents, wherein the candidate samples are formed based on text data and intent labels corresponding to the text data; Extract the intent keywords corresponding to each intent document, and obtain the intent keyword set corresponding to the target intent label by clustering the intent keywords; Using a large language model to process the intent keyword set according to the target prompt word, to obtain a regular expression corresponding to the target intent label; The regular expression is stored in a regular library, wherein the regular expression stored in the regular library is used to identify the user's question and answer intention in the question and answer interaction stage.

2. The method for generating a regular expression based on a large language model according to claim 1, characterized in that: Before the step of selecting candidate samples with the same intent label in the sample library to form multiple intent documents is performed, the method further includes: Acquire text data, and perform regular matching on the text data using the regular expressions stored in the regular library; In case of successful matching, determining a target regular expression corresponding to the text data; The intent label corresponding to the target regular expression is used as the intent label corresponding to the text data, and a question-answering interaction task is performed based on the intent label corresponding to the text data.

3. The method for generating a regular expression based on a large language model according to claim 2, characterized in that: In case of a match failure, this also includes: Inputting the text data into the large language model for intent recognition to obtain an intent label corresponding to the text data; A candidate sample is constructed based on the text data and the intent label corresponding to the text data, and the candidate sample is stored in the sample library.

4. The method for generating a regular expression based on a large language model according to claim 1, characterized in that: Determining the intent keyword corresponding to any one of the plurality of intent documents includes: Identify pending intent documents; Extracting a plurality of initial intent keywords from the intent document to be processed using a keyword extraction algorithm; The multiple initial intent keywords are deduplicated using a keyword deduplication algorithm, and the intent keywords corresponding to the intent document to be processed are determined based on the deduplication results.

5. The method for generating regular expressions based on a large language model according to claim 1, characterized in that: Determining the set of intent keywords corresponding to the target intent tag includes: Determine a plurality of to-be-clustered intent keywords corresponding to the target intent tag, wherein the plurality of to-be-clustered intent keywords belong to the intent document corresponding to the target intent tag; Performing semantic recognition on the multiple keywords to be clustered through a semantic recognition model to obtain a semantic vector corresponding to each keyword to be clustered; The semantic vector corresponding to each intent keyword to be clustered is processed using a clustering algorithm, and a set of intent keywords corresponding to the target intent label is constructed based on the processing results.

6. The method for generating regular expressions based on a large language model according to claim 1, characterized in that: After the step of storing the regular expression in the regular library is executed, the method further includes: Matching text data with unrecognized intent in the text database with regular expressions stored in the regular library; According to the matching results, the associated text data is determined in the text database and deleted.

7. The method for generating a regular expression based on a large language model according to claim 4, characterized in that: The method of performing deduplication processing on the multiple initial intent keywords by using a keyword deduplication algorithm, and determining the intent keyword corresponding to the intent document to be processed according to the deduplication processing result, includes: Calculating cosine similarity of the multiple initial intent keywords using a keyword deduplication algorithm; Determine the intended keyword to be deleted according to the cosine similarity calculation result, and delete the intended keyword to be deleted from the multiple initial intended keywords; According to the deletion result, the remaining initial intent keywords are used as the intent keywords corresponding to the intent document to be processed.

8. A data processing method, characterized in that: include: Get the question text data submitted by the user; Performing regular matching on the question text data using a regular expression stored in a regular library, wherein the regular expression stored in the regular library is constructed according to the method according to any one of claims 1 to 7; Determine the target regular expression corresponding to the question text data according to the matching result, and use the intent label corresponding to the target regular expression as the intent label corresponding to the question text data; Answer text data is generated based on the intention label corresponding to the question text data, and is displayed to the user.

9. A regular expression generation device based on a large language model, characterized in that: include: A selection module is configured to select candidate samples with the same intent labels from a sample library to form a plurality of intent documents, wherein the candidate samples are formed based on text data and intent labels corresponding to the text data; An extraction module is configured to extract intent keywords corresponding to each intent document, and obtain an intent keyword set corresponding to a target intent tag by clustering the intent keywords; A processing module is configured to process the intention keyword set according to the target prompt word using a large language model to obtain a regular expression corresponding to the target intention label; The storage module is configured to store the regular expression in a regular library, wherein the regular expression stored in the regular library is used to identify the user's question and answer intention in the question and answer interaction stage.

10. A data processing device, characterized in that: include: An acquisition module is configured to acquire question text data submitted by a user; A matching module, configured to perform regular matching on the question text data using a regular expression stored in a regular library, wherein the regular expression stored in the regular library is constructed according to the method according to any one of claims 1 to 7; A determination module is configured to determine a target regular expression corresponding to the question text data according to the matching result, and use the intent tag corresponding to the target regular expression as the intent tag corresponding to the question text data; The display module is configured to generate answer text data based on the intention label corresponding to the question text data and display it to the user.

11. A computing device, characterized in that: include: Memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the method according to any one of claims 1 to 8 are implemented.

12. A computer-readable storage medium, characterized in that: It stores computer executable instructions, which, when executed by a processor, implement the steps of the method described in any one of claims 1 to 8.

13. A computer program product, characterized in that The method comprises a computer program or an instruction, which implements the steps of the method according to any one of claims 1 to 8 when executed by a processor.

Citation Information

Patent Citations

  • Entity relationship extraction method and device

    CN111126067A

  • Text classification method of regular expression generated based on large language model

    CN117556049A

  • Regular expression generation method and device, equipment and storage medium

    CN118427406A

  • Variable extraction method, device and equipment based on large model and storage medium

    CN118782263A

  • Instruction intention recognition method and device, computing equipment and storage medium

    CN119167145A