Method and device for intention recognition and slot position extraction and medium

By filtering sensitive words, correcting errors and replacing synonyms on the text content input by users, combining intent classification model and text ordering technology, accurate intention recognition and slot extraction of fuzzy languages ​​are achieved, solving the problem that the existing technology is difficult to deal with fuzzy languages, and improving the accuracy and efficiency of interactions.

CN120163164APending Publication Date: 2025-06-17CHINA UNITED NETWORK COMM GRP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510237570.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

In the prior art, it is difficult to achieve accurate intention recognition and slot extraction when dealing with daily fuzzy language.

Method used

By obtaining the text content input by the user, determining whether there are sensitive words, performing text error correction and synonym replacement, using a pre-constructed intent classification model for intent recognition, and text ordering of the text according to the intent category, and finally performing slot extraction.

Benefits of technology

It realizes accurate intention recognition and efficient slot extraction of daily fuzzy language, improves the accuracy and efficiency of human-computer interaction, and is suitable for business indicator question-and-answer query and process action initiation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163164A_ABST
    Figure CN120163164A_ABST
Patent Text Reader

Abstract

The invention provides an intention recognition and slot position extraction method and device and a medium, and relates to the technical field of artificial intelligence. The method comprises the steps that text content input by a user is acquired; judging whether sensitive words exist in the text content or not; in response to the fact that no sensitive word exists in the text content, text error correction and synonym replacement are conducted on the text content, and the optimized text content is obtained; performing intention recognition on the optimized text content by using a pre-constructed intention classification model to obtain a corresponding intention category; performing text sequence adjustment on the optimized text content according to a standard format template corresponding to the intention category to obtain the text content in a standardized format; and performing slot extraction on the text content in the standardized format. According to the method, the device and the medium, the problem that accurate intention recognition and slot position extraction are difficult to realize when a daily fuzzy spoken language is processed in the prior art can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a method, device and medium for intent recognition and slot extraction. Background Art

[0002] With the rapid development of artificial intelligence technology and the explosive growth of information volume, people's demands for accurate information acquisition, quick action initiation, and intelligent human-computer interaction are becoming increasingly urgent. Specifically, the machine analyzes human spoken language, recognizes human intent, extracts key information, and then gives corresponding feedback according to the predefined actions corresponding to the intent and the parameter corresponding to the slot, completes the human-computer interaction, and achieves the purpose of information push or action initiation. Therefore, accurate intent recognition and slot extraction are the basis for the machine to give correct answers or quickly initiate actions.

[0003] However, in the prior art, it is difficult to achieve accurate intent recognition and slot extraction when dealing with daily fuzzy spoken language. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide a method, device and medium for intent recognition and slot extraction to solve the problem that it is difficult to achieve accurate intent recognition and slot extraction in the prior art when dealing with daily fuzzy spoken language.

[0005] In a first aspect, the present invention provides a method for intent recognition and slot extraction, the

[0006] method includes:

[0007] Obtain the text content input by the user;

[0008] Judge whether there are sensitive words in the text content;

[0009] In response to the absence of sensitive words in the text content, perform text correction and synonym replacement on the text content to obtain the optimized text content;

[0010] Use a pre-constructed intent classification model to perform intent recognition on the optimized text content to obtain the corresponding intent category;

[0011] According to the standard format template corresponding to the intent category, perform text reordering on the optimized text content to obtain the text content in a standardized format;

[0012] Perform slot extraction on the text content in the standardized format.

[0013] Further, the judgment of whether there are sensitive words in the text content specifically includes:

[0014] Use regular matching to preliminarily determine whether there are sensitive words in the text content;

[0015] In response to the preliminary judgment result that there are sensitive words, trigger an alarm;

[0016] In response to the preliminary judgment result that there are no sensitive words, input the text content and each sensitive word in the preset sensitive word library into the ELMo model based on the language model's embedding to obtain the text vectors of the text content and each sensitive word in the sensitive word library;

[0017] Calculate the cosine similarity between the text vector of the text content and the text vectors of each sensitive word in the sensitive word library;

[0018] If the cosine similarity between the text vector of the text content and the text vector of a certain sensitive word is greater than or equal to the preset first threshold, it is determined that there are sensitive words in the text content; otherwise, it is determined that there are no sensitive words in the text content and an alarm is triggered.

[0019] Further, the text error correction and synonym replacement of the text content to obtain the optimized text content specifically include:

[0020] Detect and correct the typos in the text content to obtain the corrected text content;

[0021] Based on the preset synonym table, perform synonym replacement on the corrected text content to obtain the optimized text content.

[0022] Further, the intent classification model includes a Bidirectional Encoder Representations from Transformers (BERT) model, a Deep Pyramid Convolutional Neural Network (DPCNN) model, and a fully connected layer. Using the pre-built intent classification model to perform intent recognition on the optimized text content to obtain the corresponding intent category specifically includes:

[0023] Use the pre-trained BERT model to encode the optimized text content and output the corresponding embedding sequence;

[0024] Input the embedding sequence into the DPCNN model, and through the DPCNN model, perform layer-by-layer processing on the embedding sequence to obtain the processed equal-length feature vectors;

[0025] Input the equal-length feature vectors into the fully connected layer and use the softmax function to output the probability distribution of each intent category;

[0026] According to the probability distribution, select the intent category with the highest probability as the intent category of the optimized text content.

[0027] Further, text reordering is performed on the optimized text content according to the standard format template corresponding to the intention category to obtain the text content in a standardized format, which specifically includes:

[0028] Obtain text reordering prompt words of the standard format template corresponding to the intention category, where the text reordering prompt words are used to sort the input text according to the requirements of the standard format template;

[0029] According to the text reordering prompt words, use the optimized text content as the input text, and perform text reordering using a preset large language model to obtain the text content in a standardized format.

[0030] Further, slot extraction is performed on the text content in a standardized format, which specifically includes:

[0031] Perform word segmentation on the text content in a standardized format, and vectorize each word after word segmentation;

[0032] Obtain the preset slot content of the intention category in the information table;

[0033] Use the cosine similarity algorithm to calculate the similarity between each vectorized word and the slot vector corresponding to the preset slot content. If the similarity exceeds a preset second threshold, determine that the word is the content of the corresponding slot, and extract the corresponding slot content;

[0034] Judge whether the extracted slots are consistent with the preset slots of the intention category. If not, continue to extract slots using the preset large language model or adopt any of the following filling methods: fill the slot information retained in the previous round of conversation into the missing slots, and fill the predefined default slot information into the missing slots.

[0035] Further, after slot extraction is performed on the text content in a standardized format, the method further includes:

[0036] Judge whether the extracted slot content meets the requirements according to the preset business rules;

[0037] In response to the slot content meeting the requirements, judge whether the user has the corresponding access permission;

[0038] In response to the user having the corresponding access permission, perform the action corresponding to the intention category according to the extracted slot content.

[0039] In a second aspect, the present invention provides a device for intention recognition and slot extraction, the device includes:

[0040] An input text acquisition module for acquiring the text content input by the user;

[0041] A sensitive word judgment module, connected to the input text acquisition module, for judging whether there are sensitive words in the text content;

[0042] An input text optimization module, connected to the sensitive word judgment module, for performing text error correction and synonym replacement on the text content in response to the absence of sensitive words in the text content to obtain the optimized text content;

[0043] An intention recognition module, connected to the input text optimization module, for using a pre-constructed intention classification model to perform intention recognition on the optimized text content to obtain the corresponding intention category;

[0044] A text reordering module, connected to the intention recognition module, for reordering the optimized text content according to the standard format template corresponding to the intention category to obtain the text content in a standardized format;

[0045] A slot extraction module, connected to the text reordering module, for performing slot extraction on the text content in the standardized format.

[0046] In a third aspect, the present invention provides a device for intention recognition and slot extraction, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to implement the method for intention recognition and slot extraction described in the first aspect above.

[0047] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method for intention recognition and slot extraction described in the first aspect above is implemented.

[0048] The method, device and medium for intention recognition and slot extraction provided by the present invention. First, obtain the text content input by the user; then determine whether there are sensitive words in the text content; and in response to the absence of sensitive words in the text content, perform text correction and synonym replacement on the text content to obtain the optimized text content; then use a pre-constructed intention classification model to perform intention recognition on the optimized text content to obtain the corresponding intention category; and perform text reordering on the optimized text content according to the standard format template corresponding to the intention category to obtain the text content in the standardized format; finally, perform slot extraction on the text content in the standardized format. By filtering sensitive words, performing text correction and synonym replacement on the text content input by the user, and then performing intention recognition and reordering based on the standard format template of the intention category on the optimized text content, the present invention finally realizes accurate intention recognition and efficient slot extraction for daily fuzzy spoken language. This method can be used as a pre-step for question-and-answer queries of business indicators and initiation of process actions, which helps to improve the intelligent operation level of the company and save labor costs. It solves the problem that it is difficult to achieve accurate intention recognition and slot extraction in the prior art when dealing with daily fuzzy spoken language. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 It is a flowchart of a method for intention recognition and slot extraction according to Embodiment 1 of the present invention;

[0050] Figure 2 It is a flowchart of another method for intention recognition and slot extraction according to Embodiment of the present invention;

[0051] Figure 3 It is a schematic structural diagram of the ELMo model according to an embodiment of the present invention;

[0052] Figure 4 It is a flowchart of the text optimization method according to an embodiment of the present invention;

[0053] Figure 5 It is an architecture diagram of the intention classification model according to an embodiment of the present invention;

[0054] Figure 6 It is a flowchart of slot extraction according to an embodiment of the present invention;

[0055] Figure 7 It is a schematic structural diagram of a device for intention recognition and slot extraction according to Embodiment 2 of the present invention;

[0056] Figure 8 It is a schematic structural diagram of a device for intention recognition and slot extraction according to Embodiment 3 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0057] To enable those skilled in the art to better understand the technical solution of the present invention, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.

[0058] It can be understood that the specific embodiments and drawings described herein are only used to explain the present invention, rather than limiting the present invention.

[0059] It can be understood that, without conflict, the various embodiments in the present invention and the features in the embodiments can be combined with each other.

[0060] It can be understood that, for the convenience of description, only the parts related to the present invention are shown in the drawings of the present invention, and the parts unrelated to the present invention are not shown in the drawings.

[0061] It can be understood that each unit and module involved in the embodiments of the present invention may correspond to only one entity structure, or may be composed of multiple entity structures, or multiple units and modules may also be integrated into one entity structure.

[0062] It can be understood that the terms "first", "second", etc. in the embodiments of the present invention are used to distinguish different objects, or to distinguish different processes for the same object, rather than to describe the specific order of the objects.

[0063] It can be understood that, without conflict, the functions and steps marked in the flowcharts and block diagrams of the present invention may occur in an order different from that marked in the drawings.

[0064] It can be understood that in the flowcharts and block diagrams of the present invention, the possible system architectures, functions, and operations of the systems, devices, equipment, and methods according to the embodiments of the present invention are shown. Among them, each block in the flowchart or block diagram may represent a unit, module, program segment, or code, which contains executable instructions for implementing the specified function. Moreover, each block or combination of blocks in the block diagram and flowchart can be implemented by a hardware-based system for implementing the specified function, or can be implemented by a combination of hardware and computer instructions.

[0065] It can be understood that the units and modules involved in the embodiments of the present invention can be implemented in software or in hardware. For example, the units and modules can be located in the processor.

[0066] Embodiment 1:

[0067] This embodiment provides a method for intention recognition and slot extraction, as Figure 1 shown, the method includes:

[0068] Step S101: Obtain the text content input by the user.

[0069] In this embodiment, taking the question-and-answer query of business indicators as an example, the text content input by the user can be very diverse and flexible. The user may use colloquial or vague expressions, such as "How was the development of the mobile network yesterday? For each city in ** province", "How was the performance of this province last month?", etc., to inquire about information on business indicators.

[0070] Step S102: Determine whether there are sensitive words in the text content.

[0071] In this embodiment, sensitive words often have destructiveness, illegality or inappropriateness. In order to ensure the health and safety of the network environment, the sensitive words in the text content are filtered first.

[0072] Optionally, determining whether there are sensitive words in the text content specifically includes:

[0073] Use regular matching to preliminarily determine whether there are sensitive words in the text content;

[0074] In response to the result of the preliminary judgment that there are sensitive words, trigger an alarm;

[0075] In response to the result of the preliminary judgment that there are no sensitive words, input the text content and each sensitive word in the preset sensitive word library into the ELMo (Embeddings from Language Models) model to obtain the text vectors of the text content and each sensitive word in the sensitive word library;

[0076] Calculate the cosine similarity between the text vector of the text content and the text vectors of each sensitive word in the sensitive word library;

[0077] If the cosine similarity between the text vector of the text content and the text vector of a certain sensitive word is greater than or equal to the preset first threshold, it is determined that there are sensitive words in the text content; otherwise, it is determined that there are no sensitive words in the text content and an alarm is triggered.

[0078] In this embodiment, in order to ensure the security and compliance of the input text, multiple detection methods are used to conduct a step-by-step review of the text. Specifically, regular matching can be used first to quickly determine whether there is obvious illegal content (i.e., sensitive words) in the input text. If so, an alarm is given to the user. Among them, regular matching is a technology that uses regular expressions to perform pattern matching on text. Regular expressions are composed of ordinary characters and special characters and can construct complex matching rules.

[0079] In this embodiment, since sensitive words cannot be exhaustively listed, it is possible to further determine whether there is any hidden or variant sensitive content in the text content based on the similarity calculation results between the text vector and each sensitive word in the sensitive word library, wherein the sensitive word library is used to identify and manage multiple types of sensitive words, including sensitive word ID (primary key), sensitive word, sensitive word type, update time and other fields. If it is determined that there is any hidden or variant sensitive content, an alarm is immediately triggered to prompt the user to modify it.

[0080] Step S103: In response to the absence of sensitive words in the text content, text error correction and synonym replacement are performed on the text content to obtain optimized text content.

[0081] In this embodiment, in order to improve the accuracy and readability of the text content, text error correction is performed on the text content. At the same time, in order to enrich the language expression or enhance the text diversity, synonym replacement is performed on the text content, that is, certain words in the text content are replaced with their synonyms.

[0082] Optionally, performing text error correction and synonym replacement on the text content to obtain the optimized text content specifically includes:

[0083] Detecting and correcting typos in the text content to obtain the corrected text content;

[0084] The corrected text content is replaced with synonyms based on a preset synonym table to obtain optimized text content.

[0085] In this embodiment, ERNIE SpellGCN, BERT and its variants, T5 (Text-to-Text Transfer Transformer) model, etc. can be used to detect and correct typos in user input text to prevent similar pronunciation and glyphs from confusing the text intent.

[0086] In this embodiment, the preset synonym table can be based on the pyltp word segmentation tool and the synonym table of Harbin Institute of Technology to perform synonym replacement.

[0087] Step S104: Use a pre-built intent classification model to perform intent recognition on the optimized text content to obtain a corresponding intent category.

[0088] In this embodiment, the intent classification model is pre-trained using a dataset containing question corpora and corresponding intent categories, and can perform intent recognition on the optimized text content. Among them, the intent categories can include asking for rankings, asking for details, asking for comparisons, asking for trends, asking for calibers, seal-using processes, leave-taking processes, etc. Each intent category can correspond to an intent ID and an intent name.

[0089] Optionally, the intent classification model includes a BERT (Bidirectional Encoder Representations from Transformers) model, a DPCNN (Deep Pyramid Convolutional Neural Network) model, and a fully connected layer. Using the pre-constructed intent classification model to perform intent recognition on the optimized text content to obtain the corresponding intent category specifically includes:

[0090] Using the pre-trained BERT model to encode the optimized text content and output the corresponding embedding sequence;

[0091] Inputting the embedding sequence into the DPCNN model, and processing the embedding sequence layer by layer through the DPCNN model to obtain processed equal-length feature vectors;

[0092] Inputting the equal-length feature vectors into the fully connected layer and using the softmax function to output the probability distribution of each intent category;

[0093] Selecting the intent category with the highest probability as the intent category of the optimized text content according to the probability distribution.

[0094] In this embodiment, the main body of the Bert model is stacked by 3 Transformer encoders. Each encoder consists of a multi-head self-attention mechanism, a feed-forward neural network, a residual connection, and a normalization layer. The DPCNN model contains 2 convolutional blocks, and each convolutional block consists of 2 convolutional layers and 1 pooling layer.

[0095] Specifically, first use the pre-trained BERT model to encode the optimized text content, then use the embedding sequence output by BERT as the input of DPCNN, input the equal-length feature vectors processed by DPCNN into the fully connected layer, and then use the softmax function to output the probability distribution of each category, and finally obtain the intent category of the optimized text content.

[0096] Step S105: Rearrange the optimized text content according to the standard format template corresponding to the intention category to obtain the text content in the standardized format;

[0097] In this embodiment, for the convenience of subsequent slot extraction, a corresponding standard format template is designed for each intention. These templates are designed according to specific intention categories and slot requirements. For example: [Time][Location][Indicator Name][Question Word]. Then, the optimized text content is rearranged according to the standard format template to obtain the text content in the standardized format.

[0098] Optionally, the step of rearranging the optimized text content according to the standard format template corresponding to the intention category to obtain the text content in the standardized format specifically includes:

[0099] Obtain the text rearrangement prompt words of the standard format template corresponding to the intention category, where the text rearrangement prompt words are used to sort the input text according to the requirements of the standard format template;

[0100] According to the text rearrangement prompt words, use the optimized text content as the input text and perform text rearrangement using a preset large language model to obtain the text content in the standardized format.

[0101] In this embodiment, text rearrangement prompt words are designed according to the standard format template, and then the general ability of a preset large language model (such as Qwen2.5 - 72B, GPT - 4, Llama - 3.1 - 405B, etc.) is used to perform text rearrangement to obtain the text content in the standardized format. For example, the text rearrangement prompt words for the text of the development indicator category can be as follows: "\n\nYou are an assistant for sentence order standardization. You need to adjust the word order of the input sentence according to the example without changing its meaning content, and answer in Chinese.\n\nNote: Do not return in Markdown format and do not mix other information.\n\n\nSentence order standardization example: Input: How about the mobile network development situation yesterday? In each city of ** Province\n\nOutput: Yesterday, in each city of ** Province, how about the mobile network development situation?\n\n". The preset large language model outputs the rearranged text content in the standardized format according to this text rearrangement prompt word and the optimized text content.

[0102] Step S106: Extract slots from the text content in the standardized format.

[0103] In this embodiment, user question corpora can be collected, the corpora and slot tags can be annotated, an information table can be constructed, and then slot extraction can be performed on the text content in the standardized format according to the information table. The information table is used to describe the mapping relationship between intent categories and slots, and the information table includes but is not limited to question corpora, corpus IDs, intent categories, intent IDs, intent slots, actions to be performed, etc.

[0104] Optionally, the slot extraction of the text content in the standardized format specifically includes:

[0105] Segment the text content in the standardized format and vectorize each segmented word.

[0106] Obtain the slot content preset for the intent category in the information table.

[0107] Use the cosine similarity algorithm to calculate the similarity between each vectorized word and the slot vector corresponding to the preset slot content. If the similarity exceeds a preset second threshold, it is determined that the word is the content of the corresponding slot, and the corresponding slot content is extracted.

[0108] Determine whether the extracted slot is consistent with the slot preset for the intent category. If not, continue to extract slots using the preset large language model or adopt any of the following filling methods: fill the missing slots with the slot information retained in the previous round of conversation, or fill the missing slots with the predefined default slot information.

[0109] In this embodiment, each intent category usually corresponds to multiple slots. For example, the information table can be as shown in Table 1, where the first intent category (asking about ranking) corresponds to three slots, and the corresponding slot contents are "** City", "November 2024", and "Income" respectively.

[0110] Table 1: Information Table

[0111]

[0112] In this embodiment, the cosine similarity algorithm is used to calculate the similarity between each word after vectorizing the text content and the corresponding slot vector. If it exceeds a certain threshold (i.e., the second threshold), it is determined as the content of that slot. Then, it is determined whether the extracted slot is consistent with the slot preset for the intent category. If it is consistent, it is determined that the integrity check is satisfied. If not, continue to extract slots using the preset large language model or adopt any of the following filling methods: fill the missing slots with the slot information retained in the previous round of conversation, or fill the missing slots with the predefined default slot information.

[0113] In an optional embodiment, if there is an inconsistency, the preset large language model (such as the Qwen2.5-72B model) can be used first to continue extracting slots. If the slot extraction by the large language model still cannot meet the slots preset for the intent category, it can be further determined whether this conversation is the first-round conversation. If it is not the first-round conversation, the slot information retained in the previous conversation is filled into the missing slots. If this conversation is the first-round conversation, the default information of the predefined slots is filled into the missing slots.

[0114] Optionally, after the slot extraction of the text content in the standardized format, the method further includes:

[0115] Judge whether the extracted slot content meets the requirements according to the preset business rules;

[0116] In response to the slot content meeting the requirements, judge whether the user has the corresponding access permission;

[0117] In response to the user having the corresponding access permission, perform the action corresponding to the intent category according to the extracted slot content.

[0118] In this embodiment, the preset business rules define that the slot content needs to meet specific formats, ranges, or logical conditions to ensure data accuracy and system security. For example, judge whether the extracted slot content meets the following requirements: "Whether the data accounting period is less than the current accounting period", "Whether the region exists", "Whether there is data for this accounting period in the system", etc. If the requirements are not met, the user is prompted with the reason for the error on the user interaction interface.

[0119] In this embodiment, if the slot content meets the requirements, it is further judged whether the user has the corresponding access permission. This access permission includes accessible intent categories, accessible regions, accessible business metrics, etc. Specifically, the system will compare the "accessible intents" and "accessible regions" or "accessible business metrics" information corresponding to the user ID in the user permission table according to the extracted slot content (such as the region specified by the user or the business metric to be queried, etc.). If they match, the system will determine that the user has the corresponding access permission and allow them to perform subsequent actions. Conversely, if the extracted slot content does not match their access permission, the system will reject their access request and may give corresponding error prompts or suggestions.

[0120] In this embodiment, the actions corresponding to the intent categories, for example, the action corresponding to the intent of asking for ranking is to query the ranking of a certain metric in the database; the action corresponding to the intent of asking for details is to query the detailed data of a certain metric in the database, etc.

[0121] It should be noted that in actual production, the method for intention recognition and slot extraction provided by the embodiments of the present invention can be used as a pre-step for question-and-answer queries of business indicators and initiation of process actions, which helps to improve the intelligent operation level of the company and save labor costs.

[0122] In a specific embodiment, to help the machine accurately understand the true purpose and intention expressed by the user and lay a foundation for subsequent data queries and action initiations, the present invention constructs a method for accurately recognizing intentions and extracting slots in fuzzy colloquial expressions based on the user's question content by using ELMo, ERNIESpellGCN, BERT, DPCNN, and Qwen2.5-72B technologies. The flowchart of this intention recognition and slot extraction method is as Figure 2 shown. The specific process is as follows: Obtain the text content of the user's question, store and process the data locally; use regular matching and the ELMo model to automatically identify sensitive content in the input text and perform corresponding processing; based on ERNIE SpellGCN and BERT technologies, perform text enhancement to identify and correct fuzzy semantics such as inverted word order, colloquial expressions, typos, synonyms, and abbreviations in the input text; based on the BERT+DPCNN text classification model, identify the user's intention; use BERT+DPCNN and Qwen2.5-72B to extract key slot information of the intention to which the input text belongs; and perform corresponding actions based on the user's intention and key slot information.

[0123] The following will introduce the specific implementation steps of intention recognition and slot extraction in detail:

[0124] I. Information collection

[0125] 1. Determine the types of intentions to be recognized, and set a unique identifier for the intention, that is, the intention ID.

[0126] 2. Determine the process actions corresponding to each intention and the slot information required for initiating actions, and set a unique identifier for the slot, that is, the slot ID. Use "intention ID_" as the prefix of the slot ID.

[0127] 3. Collect the user's question corpus, perform corpus and slot annotation, and construct an information table. For example, the information table can be as shown in Table 1.

[0128] 4. Construct a user permission table. For each intention, set its user table; for the slots of some intentions, such as: region, economic indicators, etc., set their user tables. For example, the user permission table can be as shown in Table 2.

[0129] Table 2: User permission table

[0130]

[0131] II. Sensitive Word Filtering

[0132] Construct a method for intelligently identifying sensitive words in text, and the specific steps are as follows.

[0133] 1. Construct a sensitive word library for identifying and managing various types of sensitive information (such content often has destructiveness, violation, or inappropriateness) to ensure the health and safety of the network environment. The sensitive word library is constructed and stored in the form of a sensitive word table, which contains fields such as sensitive word ID (primary key), sensitive word, sensitive word type, and update time.

[0134] 2. Use regular matching to quickly determine whether there is obvious illegal content in the input text. If so, give an alarm to the user.

[0135] 3. Input the text that still needs to be reviewed after screening (that is, the text where no violation was found in the previous step) and the sensitive words in the sensitive word library into Figure 3 the ELMo model shown to obtain corresponding text vectors, and calculate the cosine similarity between the text vector to be reviewed and the sensitive word vector. If the similarity between the input sentence and any sensitive word exceeds the preset threshold, it is considered that the sentence contains sensitive words.

[0136] The following details the process of using ELMo for word vectorization.

[0137] The corpus is vectorized as shown in the following formula:

[0138]

[0139] where, X i is the i-th text vector to be reviewed, is the m-th word vector of the i-th text; input the initial word vector of the text to be reviewed into the double-layer bidirectional LSTM (Long Short-Term Memory) layer. The forward language model believes that the probability p(t k ) of the word t k appearing is only affected by the previous k - 1 words.

[0140] p(t k ) = p(t k |t1,t2,…,t k-1 )

[0141] Therefore, the probability of the text appearing is:

[0142]

[0143] where, N represents the number of words in the text. Similarly, the probability of the text appearing calculated based on the backward model is:

[0144]

[0145] For the ELMo language model, the objective function is to take the forward and backward maximum likelihood functions:

[0146]

[0147] Where and are the parameters of the forward and backward language models.

[0148] Each LSTM module linearly processes and sigmoid activates the input vector, and inputs the processed result into the next LSTM module. After a two-layer bidirectional LSTM, each word has 2L + 1 representations The finally output ELMo word vector is:

[0149]

[0150] where α is the normalization coefficient, and the sum of all α is 1, is the input vector of word t k ; is obtained by splicing .

[0151] 5. Regularly update the sensitive word library to adapt to newly emerging social hotspots and rule changes.

[0152] It should be noted that in order to ensure the security and compliance of the input text, multiple detection methods are used to conduct a step-by-step review of the text. Specifically, after the input text undergoes multiple rounds of sensitive word detection, only the text that is not identified as containing sensitive words in each round of detection will be recognized as truly safe text. If sensitive words are identified in any round of detection, the system will immediately trigger an alarm and stop the subsequent detection of the text. Thus, ensuring that only the text that has been verified multiple times and found no sensitive words will enter the next stage, thereby effectively improving the accuracy and reliability of text review.

[0153] III. Text Error Correction

[0154] Construct an intelligent text optimization method for correcting incorrect expressions in text. Its flowchart is as Figure 4 shown, and the specific steps are as follows:

[0155] 1. Use ERNIE SpellGCN to detect and correct typos in the text input by the user, preventing the confusion of text intentions due to similar pronunciation and spelling.

[0156] (1) Install the PaddlePaddle library using pip install paddlepaddle, and install the PaddleHub, a pre-trained model management and application platform. Load the ERNIE SpellGCN model.

[0157] (2) Input the text that needs to be corrected to obtain the corrected text.

[0158] 2. Perform synonym replacement based on the pyltp word segmentation tool and the synonym list of Harbin Institute of Technology.

[0159] (1) Install the pyltp library, download the LTP model file, and download the synonym list of Harbin Institute of Technology.

[0160] (2) Initialize the LTP word segmenter and read the synonym list.

[0161] (3) Use the replace_synonyms function to perform synonym replacement.

[0162] (4) Recombine the replaced words into sentences.

[0163] IV. Intent Recognition

[0164] Based on BERT and DPCNN technologies, construct an intent classification model. Its architecture diagram is as Figure 5 shown, and the specific steps are as follows:

[0165] 1. Extract four fields (or two fields of the question corpus and the intent category), namely the corpus ID, the question corpus, the intent category, and the intent ID, from the information table as the intent recognition dataset.

[0166] 2. Randomly divide the intent recognition dataset into a training set, a validation set, and a test set according to a ratio of 6:2:2. The training set is used to train the Bert text classification model, the validation set is used for model tuning, and the test set is used to characterize the model performance.

[0167] 3. Construct a BERT+DPCNN text classification model. Use the pre-trained BERT model to encode the input text, use the embedded sequence output by BERT as the input of DPCNN, input the equal-length feature vectors processed by DPCNN into the fully connected layer, and then use the softmax function to output the probability distribution of each category.

[0168] (1) The main body of the Bert model is stacked by 3 Transformer encoders. Each encoder consists of a multi-head self-attention mechanism, a feed-forward neural network, a residual connection, and a normalization layer. The multi-head self-attention mechanism maps the input sequence to the Query, Key, and Value matrices respectively, as shown below:

[0169] Q=XWQ , K = XW K , V = XW V

[0170] Among them, W Q , W K , W V is the weight matrix.

[0171] Calculate the matching degree of each Query with all Keys using dot product similarity, and divide by the square root of the key dimension to prevent gradient vanishing or explosion

[0172]

[0173] Among them, d k is the dimension of Key.

[0174] Apply the softmax function to the attention scores to obtain the normalized attention weights

[0175]

[0176] Use the attention weights to perform a weighted sum on the value matrix, and the output is as follows:

[0177] Output = AttentionWeights · V

[0178] Concatenate the results of each head together and integrate through a linear transformation

[0179] Multi-Head(Q, K, V) = Concat(head1, head2,..., head h )W O

[0180] Among them, h is the number of heads, and W O is the linear transformation matrix.

[0181] The feed-forward neural network performs a non-linear transformation on the representation of each position, which usually includes two layers of linear transformation and an activation function in the middle

[0182] FFN(x) = max(0, xW1 + b1)W2 + b2

[0183] For the input x, the residual connection result is

[0184] Output = LayerNorm(x + Sublayer(x))

[0185] Among them, Sublayer(x) is the sublayer transformation

[0186] For x i,j perform normalization, and the result is as follows:

[0187]

[0188] Among them, μ j is the mean of the jth feature, is the variance of the jth feature, and ∈ is a very small constant.

[0189] (2) Input the feature vector output by BERT into the DPCNN model. The DPCNN model contains 2 convolution blocks, each of which consists of 2 convolution layers and 1 pooling layer. The word vector is input into the first convolution block, which undergoes two convolutions. The convolution kernel of each convolution layer is 250, the convolution kernel size is 3, the activation function is relu, and after 1 / 2 pooling; the word vector enters the second convolution block, which has the same structure as the first convolution block.

[0190] (3) The vector output by DPCNN is input into the fully connected layer, the activation function is softmax, the loss function is cross entropy, the optimizer is adaptive gradient descent, the probability of each intent category is output, and the text is classified into the category with the highest probability.

[0191] (4) Use the validation set data accuracy to measure the classification accuracy and output the average loss value and accuracy. After a certain number of iterations, the encoder, embedding layer, and convolutional layer parameters are determined based on the performance on the validation set.

[0192] 4. Train and tune the intent classification model; input the text of the intent to be identified into the training intent classification model to obtain the intent type and ID.

[0193] 5. Text reordering

[0194] Use Qwen2.5-72B+ prompt word project to adjust the text sequence and adjust the user input text to a standardized format to facilitate slot extraction.

[0195] (1) Install the CUDA toolchain, then create a new Python virtual environment and install vLLM using pip. Use the huggingface-cli tool to download the model. Start the vLLM service.

[0196] (2) For each intent, design a corresponding standard format template. An example is shown in Table 3.

[0197] Table 3: Standard format template

[0198]

[0199] (3) Design prompt words according to the standard format template and use the general capabilities of Qwen2.5-72B to perform text sorting. Examples of prompt words are shown in Table 4:

[0200] Table 4: Design of Text Reordering Prompt

[0201]

[0202] (4) Use the transformers library to load the Qwen2.5-72B model and the tokenizer. Input the prompt, process the input using the tokenizer, and generate the input for the model. Decode the IDs generated by the model to obtain the final text output.

[0203] (5) Compare the reordering results of the large model with the standard results of manual reordering, calculate the accuracy of the large model's reordering, and stop this step when the accuracy is higher than a certain threshold.

[0204] VI. Slot Extraction

[0205] The flow chart of slot extraction is as Figure 6 shown below, and the specific steps are as follows:

[0206] 1. Take the corpus ID, query corpus, intent category, intent ID, slot fields (or it can only include the query corpus and slot fields) from the information table as the slot extraction dataset.

[0207] 2. Use the BERT+DPCNN text classification model constructed in Section IV-2 to vectorize the slot content (i.e., the slot extraction dataset).

[0208] 3. Randomly divide the intent recognition dataset into a training set, a validation set, and a test set according to 6:2:2. The training set is used to train the Bert text classification model, the validation set is used for model tuning, and the test set is used to characterize the model performance.

[0209] 4. Use the cosine similarity algorithm to calculate the similarity between the text vector and the corresponding slot vector. If it exceeds a certain threshold, it is determined as the slot content. The formula for calculating the feature similarity between the text vector X and the corresponding slot vector Y is:

[0210]

[0211] 5. For the corpus X corresponding to intent N, check whether the slots extracted from corpus X are consistent with the slots preset for intent N. If they are consistent, the integrity check is satisfied, and the rationality check step is entered. If not, use the Qwen2.5-72B model to continue extracting slots. If the Qwen2.5-72B model still cannot extract slots that meet the slot content preset for intent N, fill the missing slots with the slot information retained in the previous conversation. If this is the first conversation in the dialogue, fill the missing slots with the default slot information predefined.

[0212] It should be noted that in Step 6-2, word segmentation is performed first, and in 6-3, each word is vectorized, and the word vector is compared with the template slot for similarity. If it is greater than a certain threshold, it is extracted. Finally, if the number of extracted ones is the same as the template slot, it is considered completed. Otherwise, the large model is used for extraction.

[0213] The following details the process of slot extraction using the open-source large model Qwen2.5-72B.

[0214] (1) Organize the format of the slot extraction dataset so that each record contains the corpus and slot labels. IOB (Inside, Outside, Beginning) is used to indicate whether each word belongs to a certain slot and its boundary. The specific example is as follows:

[0215] □□Jinan [B-LOC] last month [B-TIME] income [B-NAME] in [O] the whole province [O] of [O] ranking [O]

[0216] 1561531**** [B-NUM] last month [B-TIME] bill out [B-NAME] how much [O] money [O]

[0217] Among them, "O" indicates a non-slot word, "B-LOC" indicates the start of a place name, "B-TIME" indicates the start of a time, "B-NUM" indicates the start of a user number, and "B-NAME" indicates the start of an indicator name.

[0218] (2) Use BERT+DPCNN to convert the text into word vector form.

[0219] (3) Use the Transformers library of Hugging Face to load the Qwen2.5-72B model and specify it for sequence labeling tasks.

[0220] (4) Use the Trainer API provided by the Transformers library for pre-training. The learning rate is set to 1e-5, the batch size is set to 64, the number of training rounds is set to 10 rounds, the loss function is cross-entropy loss, and the optimizer is AdamW. Use the validation set to evaluate the slot extraction effect of the model.

[0221] 6. Judge whether the slot information conforms to the actual situation, including the following dimensions: "Whether the data accounting period is less than the current accounting period", "Whether the region exists", "Whether there is data for this accounting period in the system". If it does not meet the requirements, prompt the user with the error reason on the user interface.

[0222] 7. According to the intent, slot information and permission table, determine whether the user who asked the question has the permission to obtain or initiate the question content. If the user does not have the corresponding permission, the user authentication result and the corresponding administrator will be prompted on the user interaction interface.

[0223] 8. Based on the extracted slot key information, execute the predefined actions of the intent.

[0224] It should be noted that the method for intent recognition and slot extraction provided by the present invention has the following characteristics:

[0225] a) To avoid the risk of violations and protect the security of enterprise data, a method for intelligently identifying sensitive content in text and making corresponding treatments is constructed. The specific method is: construct a sensitive word library containing various types of sensitive information; use rule matching to quickly determine whether there is obvious illegal content in the input text; for content that is not determined as illegal by the rules, it is input into the ELMo model as text that still needs to be reviewed, and combined with the contextual semantics, it is determined whether there is implicit or variant sensitive content; regularly update the sensitive word library to adapt to emerging social hot spots and rule changes.

[0226] b) In order to increase the readability of the text and improve the accuracy of subsequent intent recognition, a text correction and text enhancement method is constructed to intelligently correct incorrect expressions in the text. The specific method is: using ERNIE SpellGCN to detect and correct typos in user input text to prevent similar pronunciation and glyphs from confusing the text intent; based on the pyltp word segmentation tool and the synonym table of Harbin Institute of Technology, synonym replacement is performed;

[0227] c) In order to improve the user experience and optimize the interactive experience, a method for identifying the intent of user colloquial language is constructed. The specific method is: define the user intent category and set the unique identifier of the intent, i.e., ID; collect corpus and manually annotate the intent corresponding to the corpus; randomly divide the annotated corpus into training set and test set, the training set is used to train the intent classification model, and the test set is used for model tuning; construct a BERT+DPCNN text classification model, use the pre-trained BERT model to encode the input text, use the embedded sequence output by BERT as the input of DPCNN, input the equal-length feature vector processed by DPCNN into the fully connected layer, and then use the softmax function to output the probability distribution of each category; train and tune the intent classification model; input the text of the intent to be identified into the training intent classification model to obtain the intent category and ID to which it belongs.

[0228] d) To build intelligent and efficient dialogue systems and services, a slot extraction method incorporating rationality verification and user intention analysis is constructed. The specific method is as follows: For the defined intention, set the corresponding slot extraction template, and use BERT+DPCNN to vectorize the text of the corresponding slot; obtain the intention ID corresponding to the text vector after being vectorized by BERT+DPCNN; use the cosine similarity algorithm to calculate the similarity between the text vector and the corresponding slot vector, and if it exceeds a certain threshold, it is determined as the content of that slot; for the text where the slots corresponding to the intention are not completely extracted, use the Qwen2.5-72B model to further extract slot information; judge whether the slot information conforms to reality; based on the intention and slot information, judge whether the user asking the question has the permission to obtain or initiate the content of their question; if the information conforms to reality and the user has the permission, execute the predefined action.

[0229] The method for intention recognition and slot extraction provided by the embodiments of the present invention first obtains the text content input by the user; then judges whether there are sensitive words in the text content; and in response to the absence of sensitive words in the text content, performs text correction and synonym replacement on the text content to obtain the optimized text content; then uses a pre-constructed intention classification model to perform intention recognition on the optimized text content to obtain the corresponding intention category; and reorders the optimized text content according to the standard format template corresponding to the intention category to obtain the text content in the standardized format; finally, performs slot extraction on the text content in the standardized format. The present invention filters sensitive words, corrects text, and replaces synonyms on the text content input by the user, and then performs intention recognition and reordering based on the standard format template of the intention category on the optimized text content, and finally realizes accurate intention recognition and efficient slot extraction of daily fuzzy colloquial language. This method can be used as a pre-step for question-and-answer queries of business indicators and initiation of process actions, which helps to improve the intelligent operation level of the company and save labor costs. It solves the problem that it is difficult to achieve accurate intention recognition and slot extraction in the prior art when dealing with daily fuzzy colloquial language.

[0230] Embodiment 2:

[0231] As Figure 7 shown, this embodiment provides a device for intention recognition and slot extraction for executing the above-mentioned method for intention recognition and slot extraction, including:

[0232] An input text acquisition module 11, configured to acquire the text content input by the user;

[0233] A sensitive word judgment module 12, connected to the input text acquisition module 11, and configured to judge whether there are sensitive words in the text content;

[0234] The input text optimization module 13, connected to the sensitive word judgment module 12, is configured to perform text error correction and synonym replacement on the text content in response to the absence of sensitive words in the text content, so as to obtain the optimized text content;

[0235] The intent recognition module 14, connected to the input text optimization module 13, is configured to perform intent recognition on the optimized text content using a pre-constructed intent classification model to obtain the corresponding intent category;

[0236] The text reordering module 15, connected to the intent recognition module 14, is configured to reorder the optimized text content according to the standard format template corresponding to the intent category to obtain the text content in a standardized format;

[0237] The slot extraction module 16, connected to the text reordering module 15, is configured to perform slot extraction on the text content in a standardized format.

[0238] Optionally, the sensitive word judgment module 12 specifically includes:

[0239] The regular matching unit is configured to preliminarily judge whether there are sensitive words in the text content using regular matching;

[0240] The first processing unit is configured to trigger an alarm in response to the result of the preliminary judgment that there are sensitive words;

[0241] The second processing unit is configured to input the text content and each sensitive word in the preset sensitive word library into the ELMo model based on the language model in response to the result of the preliminary judgment that there are no sensitive words, so as to obtain the text vectors of the text content and each sensitive word in the sensitive word library;

[0242] The first calculation unit is configured to calculate the cosine similarity between the text vector of the text content and the text vectors of each sensitive word in the sensitive word library;

[0243] The third processing unit is configured to judge that there are sensitive words in the text content if the cosine similarity between the text vector of the text content and the text vector of a certain sensitive word is greater than or equal to a preset first threshold, otherwise, judge that there are no sensitive words in the text content and trigger an alarm.

[0244] Optionally, the input text optimization module 13 specifically includes:

[0245] The typo correction unit is configured to detect and correct typos in the text content to obtain the corrected text content;

[0246] A synonym replacement unit for performing synonym replacement on the corrected text content based on a preset synonym table to obtain the optimized text content.

[0247] Optionally, the intent classification model includes a Bidirectional Encoder Representations from Transformers (BERT) model, a Deep Pyramid Convolutional Neural Network (DPCNN) model, and a fully connected layer. The intent recognition module 14 specifically includes:

[0248] An encoding unit for encoding the optimized text content using a pre-trained BERT model and outputting a corresponding embedding sequence;

[0249] A feature extraction unit for inputting the embedding sequence into the DPCNN model and performing layer-by-layer processing on the embedding sequence through the DPCNN model to obtain processed equal-length feature vectors;

[0250] A fully connected unit for inputting the equal-length feature vectors into the fully connected layer and using the softmax function to output the probability distribution of each intent category;

[0251] A category determination unit for selecting the intent category with the highest probability as the intent category of the optimized text content according to the probability distribution.

[0252] Optionally, the text reordering module 15 specifically includes:

[0253] A prompt word acquisition unit for acquiring text reordering prompt words of a standard format template corresponding to the intent category, where the text reordering prompt words are used to sort the input text according to the requirements of the standard format template;

[0254] A text reordering unit for using the preset large language model to reorder the optimized text content as the input text according to the text reordering prompt words to obtain the text content in a standardized format.

[0255] Optionally, the slot extraction module 16 specifically includes:

[0256] A vectorization unit for tokenizing the text content in the standardized format and vectorizing each tokenized word;

[0257] A first acquisition unit for acquiring the preset slot content of the intent category in the information table;

[0258] A second calculation unit for calculating the similarity between each vectorized word and the corresponding slot vector of the preset slot content using the cosine similarity algorithm, and determining that the word is the content of the corresponding slot and extracting the corresponding slot content if the similarity exceeds a preset second threshold.

[0259] A fourth processing unit, configured to determine whether the extracted slot is consistent with the slot preset for the intent category. If not, continue to extract slots using the preset large language model or adopt any of the following filling methods: filling the slot information retained in the previous round of conversation into the missing slots, and filling the predefined default slot information into the missing slots.

[0260] Optionally, the device further includes:

[0261] A rule judgment module, configured to judge whether the content of the extracted slot meets the requirements according to the preset service rules;

[0262] An authentication module, configured to judge whether the user has the corresponding access right in response to the content of the slot meeting the requirements;

[0263] An action execution module, configured to execute the action corresponding to the intent category according to the content of the extracted slot in response to the user having the corresponding access right.

[0264] Embodiment 3:

[0265] Reference Figure 8 , this embodiment provides a device for intent recognition and slot extraction, including a memory 21 and a processor 22. A computer program is stored in the memory 21, and the processor 22 is configured to run the computer program to execute the method for intent recognition and slot extraction in Embodiment 1.

[0266] Wherein, the memory 21 is connected to the processor 22. The memory 21 can adopt flash memory or read-only memory or other memories, and the processor 22 can adopt a central processing unit or a single-chip microcomputer.

[0267] Embodiment 4:

[0268] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the method for intent recognition and slot extraction in the above Embodiment 1.

[0269] The computer-readable storage medium includes volatile or non-volatile, removable or non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, computer program modules, or other data. Computer-readable storage media includes, but is not limited to, RAM (Random Access Memory), ROM (Read-Only Memory), EEPROM (Electrically Erasable Programmable Read Only Memory), flash memory or other memory technologies, CD-ROM (Compact Disc Read-Only Memory), digital versatile discs (DVDs) or other optical disc storage, magnetic cassettes, tapes, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer.

[0270] In summary, the method, apparatus, and medium for intent recognition and slot extraction provided by the embodiments of the present invention first obtain the text content input by the user; then determine whether there are sensitive words in the text content; and in response to the absence of sensitive words in the text content, perform text correction and synonym replacement on the text content to obtain the optimized text content; then use a pre-constructed intent classification model to perform intent recognition on the optimized text content to obtain the corresponding intent category; and reorder the optimized text content according to the standard format template corresponding to the intent category to obtain the text content in a standardized format; finally, perform slot extraction on the text content in the standardized format. The present invention filters sensitive words, corrects text, and replaces synonyms for the text content input by the user, then performs intent recognition and reorders based on the standard format template of the intent category on the optimized text content, and finally realizes accurate intent recognition and efficient slot extraction for daily fuzzy spoken language. This method can be used as a pre-step for question-and-answer queries of business indicators and initiation of process actions, which helps to improve the intelligent operation level of the company and save labor costs. It solves the problem that it is difficult to achieve accurate intent recognition and slot extraction in the prior art when dealing with daily fuzzy spoken language.

[0271] It can be understood that the above embodiments are merely exemplary embodiments adopted to illustrate the principles of the present invention, and the present invention is not limited thereto. For those of ordinary skill in the art, various modifications and improvements can be made without departing from the spirit and essence of the present invention, and these modifications and improvements are also considered within the protection scope of the present invention.

Claims

1. A method for intent recognition and slot extraction, characterized in that: The method comprises: Get the text content entered by the user; Determine whether there are sensitive words in the text content; In response to the absence of sensitive words in the text content, performing text error correction and synonym replacement on the text content to obtain optimized text content; Using a pre-built intent classification model to perform intent recognition on the optimized text content to obtain a corresponding intent category; Performing text sorting on the optimized text content according to the standard format template corresponding to the intention category to obtain the text content in a standardized format; Slot extraction is performed on the text content in a standardized format.

2. The method according to claim 1, characterized in that The determining whether there is a sensitive word in the text content specifically includes: Use regular matching to preliminarily determine whether there are sensitive words in the text content; In response to the preliminary judgment result that a sensitive word exists, an alarm is triggered; In response to the result of the preliminary judgment that there is no sensitive word, the text content and each sensitive word in the preset sensitive word library are input into an ELMo model embedded based on a language model to obtain a text vector of the text content and each sensitive word in the sensitive word library; Calculating the cosine similarity between the text vector of the text content and the text vector of each of the sensitive words in the sensitive word library; If the cosine similarity between the text vector of the text content and the text vector of a certain sensitive word is greater than or equal to a preset first threshold, it is determined that there is a sensitive word in the text content; otherwise, it is determined that there is no sensitive word in the text content and an alarm is triggered.

3. The method according to claim 1, characterized in that The performing text error correction and synonym replacement on the text content to obtain the optimized text content specifically includes: Detecting and correcting typos in the text content to obtain the corrected text content; The corrected text content is replaced with synonyms based on a preset synonym table to obtain optimized text content.

4. The method according to claim 1, characterized in that: The intent classification model includes a Transformer-based bidirectional encoder representation BERT model, a deep pyramid convolutional neural network DPCNN model and a fully connected layer. The pre-built intent classification model is used to perform intent recognition on the optimized text content to obtain the corresponding intent category, which specifically includes: Use the pre-trained BERT model to encode the optimized text content and output the corresponding embedding sequence; Inputting the embedded sequence into the DPCNN model, processing the embedded sequence layer by layer through the DPCNN model to obtain processed feature vectors of equal length; Input the equal-length feature vectors into the fully connected layer, and use the softmax function to output the probability distribution of each intent category; The intent category with the greatest probability is selected according to the probability distribution as the intent category of the optimized text content.

5. The method according to claim 1, characterized in that The step of performing text sorting on the optimized text content according to the standard format template corresponding to the intention category to obtain the text content in a standardized format specifically includes: Acquire a text reordering prompt word of a standard format template corresponding to the intention category, wherein the text reordering prompt word is used to sort the input text according to the requirements of the standard format template; According to the text reordering prompt words, the optimized text content is used as input text, and the text is reordered using a preset large language model to obtain the text content in a standardized format.

6. The method according to claim 5, characterized in that The extracting of slots from the text content in the standardized format specifically includes: Segmenting the text content in a standardized format, and vectorizing each segmented word; Get the preset slot content of the intent category described in the information table; Using a cosine similarity algorithm to calculate the similarity between each of the vectorized words and the slot vector corresponding to the preset slot content, if the similarity exceeds a preset second threshold, the word is determined to be the content of the corresponding slot, and the corresponding slot content is extracted; Determine whether the extracted slot is consistent with the slot preset in the intent category. If not, continue to extract slots using the preset large language model or adopt any of the following filling methods: fill the missing slots with the slot information retained in the previous round of dialogue, or fill the missing slots with the predefined slot default information.

7. The method according to claim 6, characterized in that After the slot extraction is performed on the text content in the standardized format, the method further comprises: Determine whether the extracted slot content meets the requirements according to the preset business rules; In response to the slot content meeting the requirement, determining whether the user has corresponding access rights; In response to the user having corresponding access rights, an action corresponding to the intent category is performed according to the extracted slot content.

8. A device for intent recognition and slot extraction, characterized in that: The device comprises: Input text acquisition module, used to obtain the text content input by the user; A sensitive word determination module, connected to the input text acquisition module, for determining whether there are sensitive words in the text content; An input text optimization module connected to the sensitive word judgment module is used to perform text error correction and synonym replacement on the text content in response to the absence of sensitive words in the text content, so as to obtain the optimized text content; An intention recognition module, connected to the input text optimization module, is used to use a pre-built intention classification model to perform intention recognition on the optimized text content to obtain a corresponding intention category; A text reordering module, connected to the intention recognition module, for performing text reordering on the optimized text content according to a standard format template corresponding to the intention category to obtain the text content in a standardized format; The slot extraction module is connected to the text reordering module and is used to extract slots from the text content in a standardized format.

9. An apparatus for intent recognition and slot extraction, characterized in that: It comprises a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to implement the method for intent recognition and slot extraction as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method for intent recognition and slot extraction as described in any one of claims 1 to 7 is implemented.