Data mining method and system of dialogue system
Through the data mining method of dialogue system, vector similarity clustering and pre-training language models are used to identify and update differences, solving the problem of insufficient recognition accuracy and generalization capabilities of existing systems in complex scenarios, realizing in-depth analysis of dialogue data and continuous optimization of the system.
Patent Information
- Application Number
- CN202510481123.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-08-08
AI Technical Summary
The existing task-based dialogue system lacks recognition accuracy and generalization capabilities when dealing with complex scenarios and diverse user needs, and the machine learning model performs poorly in the case of long-tail data, making it difficult to explore the potential value of subsequent data.
By identifying audio data, preprocessing and judging the recognition capability of the online NLP system, when not recognized, vector similarity clustering is performed, query clustering is matched, and the difference terms are classified using the pre-trained language model. When the confidence reaches the threshold, it is marked as an update term. Combining historical audio to build a corpus and Embedding model to train historical vectors.
It improves the recognition accuracy and generalization ability of the dialogue system, reduces the mining workload of long-tail queries, improves the accuracy of data mining, and supports continuous optimization and update of the system.
Smart Images

Figure CN120448484A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data mining technology, and in particular to a data mining method and system for a dialogue system. Background Art
[0002] With the development of artificial intelligence (AI) technology, task-based dialogue systems are increasingly being used in daily life. However, these systems often face challenges with low recognition accuracy and poor user experience when processing user queries. While improvements have been made to some extent through methods such as improving speech recognition models, optimizing corpora, and enhancing language model performance, the results remain limited.
[0003] To address the above issues, two solutions have been proposed in the existing technology: rule-based task-based dialogue system data mining and machine learning-based task-based dialogue system data mining. These solutions can improve user queries from the two perspectives of rule type classification and user dialogue training, thereby enhancing user experience.
[0004] However, of the two solutions mentioned above, the former is difficult to cope with complex conversation scenarios and diverse user needs, which limits the scalability and application scope of the system, resulting in a decrease in the system's recognition accuracy and generalization ability. The latter's model performance is easily affected by the quantity and quality of training data, performs poorly in the case of long-tail data, and it is difficult to tap the potential value of subsequent data. Summary of the Invention
[0005] In response to the problems existing in the prior art, an embodiment of the present invention provides a data mining method and system for a dialogue system.
[0006] An embodiment of the present invention provides a data mining method for a dialogue system, comprising:
[0007] Recognize audio data to obtain an original query, and preprocess the original query;
[0008] Determine whether the online NLP system can recognize the pre-processed original query. If the online NLP system cannot recognize it, perform vector similarity clustering on the original query to match query clusters.
[0009] Comparing the query cluster with the classification results of the online NLP system to determine the differences;
[0010] The difference item is classified by a pre-trained language model, and a classification confidence is checked to see whether it reaches a preset threshold. When the classification confidence reaches the preset threshold, the difference item is marked as an updated item.
[0011] In one embodiment, the method further comprises:
[0012] The online NLP system includes a rule-based dialogue system and a machine learning recognition system;
[0013] Determine whether the original query matches the rules of the rule-based dialogue system, and determine whether the original query belongs to the domain and intent in the machine learning recognition system.
[0014] In one embodiment, the method further comprises:
[0015] Obtain historical audio to build a corpus, and train the historical vector corresponding to the corpus through the Embedding model;
[0016] Grouping the historical vectors by a clustering algorithm to form a query cluster set;
[0017] A preset similarity metric is obtained, and several query clusters in the query cluster set that are closest in similarity to the original query are found.
[0018] In one embodiment, the method further comprises:
[0019] Selecting a pre-trained model and training the corpus based on the pre-trained model, wherein the corpus includes historical queries and annotation information, and the annotation information includes domain and intent;
[0020] Adjust the model structure and model parameters based on system requirements to obtain the trained pre-trained language model;
[0021] The difference items are input into a pre-trained language model, and a classification result is output.
[0022] In one embodiment, the method further comprises:
[0023] When the domain and intent of the difference item classification result are the same as the clustering domain and intent of any query cluster, it is determined whether the classification confidence of the difference item classification result is greater than a preset threshold.
[0024] In one embodiment, the method further comprises:
[0025] The difference items are transmitted to a manual regression verification system. When the difference items pass the screening of the manual regression verification system, the difference items are marked as update items, and the online NLP system is updated.
[0026] An embodiment of the present invention provides a data mining system for a dialogue system, comprising:
[0027] A recognition module, configured to recognize audio data, obtain an original query, and preprocess the original query;
[0028] A clustering module is used to determine whether the online NLP system can recognize the pre-processed original query. If the online NLP system cannot recognize it, the original query is clustered based on vector similarity to match query clusters.
[0029] A difference module, used to compare the query cluster with the classification results of the online NLP system to determine the difference items;
[0030] An updating module is used to classify the difference item through a pre-trained language model and check whether the classification confidence reaches a preset threshold. When the classification confidence reaches the preset threshold, the difference item is marked as an updated item.
[0031] In one embodiment, the system further comprises:
[0032] An acquisition module is used to acquire historical audio to construct a corpus and train the historical vector corresponding to the corpus through the Embedding model;
[0033] A grouping module, configured to group the historical vectors using a clustering algorithm to form a query cluster set;
[0034] The search module is configured to obtain a preset similarity metric and search for a number of query clusters in the query cluster set that are closest in similarity to the original query.
[0035] An embodiment of the present invention provides an electronic device, including a processor and a memory;
[0036] The processor is connected to the memory;
[0037] The memory is used to store executable program code;
[0038] The processor reads the executable program code stored in the memory to run a program corresponding to the executable program code, so as to execute the method described in one or more embodiments.
[0039] An embodiment of the present invention provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the data mining method for the above-mentioned dialogue system are implemented.
[0040] In view of the above, in one or more embodiments of the present specification, audio data is identified to obtain the original query, and the original query is preprocessed; it is determined whether the online NLP system can recognize the preprocessed original query. When the online NLP system cannot recognize it, the original query is clustered by vector similarity to match the query cluster; the query cluster is compared with the classification result of the online NLP system to determine the difference items; the difference items are classified by the pre-trained language model, and the classification confidence is checked to see whether it reaches a preset threshold. When the classification confidence reaches the preset threshold, the difference items are marked as updated items. In this way, valuable data can be mined by training methods such as domain and intent recognition, clustering and classification to achieve in-depth analysis of dialogue data, which can improve the recognition accuracy and generalization ability of the dialogue system, reduce the workload of long-tail query mining in the system, improve the accuracy of data mining, and the mined valuable data can be used for subsequent system updates and iterations. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0042] Figure 1 This is a flowchart of a data mining method for a dialogue system provided by one embodiment of this specification.
[0043] Figure 2 This is a flowchart of another data mining method for a dialogue system provided by an embodiment of this specification.
[0044] Figure 3 This is a structural diagram of a data mining system for a dialogue system provided by an embodiment of this specification.
[0045] Figure 4 This is a structural diagram of an electronic device provided by an embodiment of this specification. DETAILED DESCRIPTION
[0046] The subject matter described herein will now be discussed with reference to example embodiments. It should be understood that these embodiments are discussed only to enable those skilled in the art to better understand and implement the subject matter described herein, and are not intended to limit the scope of protection, applicability, or examples set forth in the claims. The functions and arrangements of the elements discussed may be changed without departing from the scope of protection of this specification. Various examples may omit, replace, or add various processes or components as needed. For example, the described method may be performed in an order different from the order described, and various steps may be added, omitted, or combined. In addition, features described relative to some examples may also be combined in other examples.
[0047] As used herein, the term "including" and its variations are open terms meaning "including but not limited to". The term "based on" means "based at least in part on". The terms "one embodiment" and "an embodiment" mean "at least one embodiment". The term "another embodiment" means "at least one other embodiment". The terms "first", "second", etc. may refer to different or the same objects. Other definitions may be included below, whether explicit or implicit. Unless the context clearly indicates otherwise, the definition of a term is consistent throughout the specification.
[0048] like Figure 1 As shown, an embodiment of the present invention provides a data mining method for a dialogue system, comprising:
[0049] Step S102: Identify audio data, obtain an original query, and pre-process the original query.
[0050] Specifically, during the operation of the dialogue system, audio data from user queries will be received, the audio data sent by the user will be identified, and the audio data collection can be triggered by voice keywords or by corresponding trigger buttons, without further limitation. After the audio data is collected, it can be transmitted to automatic speech recognition (ASR) for recognition. The ASR model can be selected based on the application field corresponding to the dialogue system. For example, for professional terms in a specific field, the recognition rate can be improved by adding relevant vocabulary to the ASR model's dictionary. The ASR model can convert audio data into original queries in text form (queries are information sent to find specific information in the database).
[0051] Furthermore, in order to improve the text quality of the original query converted by the ASR model, the recognized text can be preprocessed in combination with a typical corpus. The preprocessing can, for example, be done by using word segmentation tools and named entity recognition (NER) technology to annotate the text; it can also include filtering of stop words, such as removing common words that are not helpful for understanding semantics; it can also include synonym replacement and expansion, such as using a synonym library to replace certain key words with a more understandable form; it can also include anomaly detection and correction in the text.
[0052] Step S104 , determining whether the online NLP system can recognize the pre-processed original query. If the online NLP system cannot recognize the original query, vector similarity clustering is performed on the original query to match query clusters.
[0053] Specifically, after obtaining the original query from the user, the pre-processed original query is identified by the online NLP system. The online comprehensive NLP system consists of two parts: a traditional rule-based task-based dialogue system and a multi-level recognition system based on machine learning. The rule base in the traditional task-based dialogue system includes common query and answer patterns and response logic, enabling immediate responses to simple queries directly through rules. The multi-level recognition system based on machine learning can ensure coverage of the widest possible range of domains and intent types by collecting annotated dialogue data as a training set. The annotation process includes but is not limited to domain classification, intent identification, and the extraction of specific parameters (such as time and location). Model selection and training are then performed on the training set data. A classification algorithm (such as SVM, random forest, or deep learning model) is selected to divide the training set data into text domains. Further, more fine-grained models (such as CRF, LSTM, BERT, etc.) are applied to identify specific user intents.
[0054] Furthermore, online NLP systems can adopt a layered architecture, first using a rule-based system to process easily identifiable queries, then passing complex queries or those that don't conform to any rules to a machine learning module for in-depth analysis. If neither the rule-based system nor the machine learning module can identify the original query after preprocessing, vector similarity clustering is performed on the original query to match query clusters.
[0055] The process of performing vector similarity clustering on the original query may include:
[0056] 1. Collect audio data from historical queries, preprocess it based on the corresponding queries, and build a corpus. Based on the corpus, generate a context-sensitive vector representation for each query. For example, using an embedding model, you can train a historical vector corresponding to the query. This historical vector preserves the semantic information of the audio data in a high-dimensional space.
[0057] 2. Apply clustering algorithms (such as K-means, DBSCAN, hierarchical clustering, etc.) to group historical vectors to form multiple query clusters. The queries within each cluster have high similarity and usually correspond to specific fields and intentions.
[0058] 3. Similarly, the original query is vectorized and a preset similarity metric (such as cosine similarity, Euclidean distance, etc.) is obtained. The similarity between the vector corresponding to the original query and the center point of the existing query cluster is compared and calculated, and the closest clusters are found as the matched query clusters.
[0059] Step S106 : Compare the query cluster with the classification results of the online NLP system to determine the difference items.
[0060] Specifically, after using a clustering algorithm to identify several query clusters in the embedding model that are highly similar to the original query, the clusters are compared with the online NLP system to determine the differences in classification results. This comparison process can, for example, be performed using the machine learning module of the online NLP system to identify inconsistencies in the classification results. For each query cluster, the original query, the cluster assignment results, and the classification output of the online NLP system are recorded in detail to provide sufficient information for subsequent analysis.
[0061] Step S108 , classifying the difference item using a pre-trained language model, and checking whether the classification confidence reaches a preset threshold. When the classification confidence reaches the preset threshold, marking the difference item as an updated item.
[0062] Specifically, after determining the difference items, the difference items can be classified using a pre-trained language model. The language model can be trained using the Transformer architecture. The training process may include:
[0063] 1. Select a pre-trained model and prepare the query and annotation information in the corpus (such as domain classification, intent recognition, and slot filling).
[0064] 2. Fine-tune the pre-trained model to adapt the language representation to the task at hand. For example, add additional fully connected layers for multi-classification tasks. Properly set hyperparameters such as the learning rate, batch size, and number of training rounds to ensure effective model convergence. For example, grid search or Bayesian optimization can be used to find the optimal configuration. For the loss function, select an appropriate one and introduce regularization to prevent overfitting.
[0065] 3. After adjusting the model structure and parameters, train the corpus to obtain a pre-trained language model. Additionally, you can use k-fold cross-validation or other validation strategies to evaluate the model's performance and ensure its generalization capabilities. You can also score performance metrics such as precision and recall to determine the language model's training results.
[0066] The difference item is input into the pre-trained language model, and the classification confidence of the difference item is detected to see if it reaches a preset threshold, such as 0.7 or 0.8. The difference item is then considered a valuable difference item that can be updated. The detection of the classification confidence of the difference item can include, after the language model outputs the domain and intent of the difference item, comparing the output result with the training results in the corpus. When the output result is the same as the classification result and clustering domain intent of a query in the corpus, and the classification confidence reaches or exceeds 0.7, it can be considered a potential difference item. The specific data of the threshold can be adjusted at will, and no further restrictions are given here.
[0067] To calculate the numerical confidence level of a classification, we can use the final fully connected layer of the Transformer model, plus a softmax activation function, to convert the model output into a probability distribution for each category. The highest category in the probability distribution and its corresponding probability value are then found to determine the classification confidence level. The model output layer can also directly predict the probability of each category and determine the highest category in the probability distribution and its corresponding probability value.
[0068] Furthermore, when the classification confidence is greater than or equal to the preset threshold, it means that the difference item has high accuracy and representativeness. In order to promptly feed back the mined difference items to various parts of the system and promote continuous optimization of the system, the difference items can be marked as update items and updated to various parts of the dialogue system, including the online NLP system, corpus, Embedding model, and pre-trained language model of the Transformer architecture.
[0069] In addition, before marking the difference items as update items, the difference items can also be transmitted to the manual regression verification system for manual review to determine whether they pass. After the difference items pass the screening of the manual regression verification system, the difference items will be marked as update items, and the online NLP system will be updated for future system iterations and optimizations.
[0070] An embodiment of the present invention provides a data mining method for a dialogue system, which identifies audio data, obtains the original query, and preprocesses the original query; determines whether the online NLP system can recognize the preprocessed original query. When the online NLP system cannot recognize it, the original query is clustered based on vector similarity to match the query cluster; the query cluster is compared with the classification results of the online NLP system to determine the difference items; the difference items are classified using a pre-trained language model, and the classification confidence is checked to see if it reaches a preset threshold. When the classification confidence reaches the preset threshold, the difference items are marked as updated items. In this way, valuable data can be mined using training methods such as clustering and classification to achieve in-depth analysis of dialogue data, which can improve the recognition accuracy and generalization ability of the dialogue system, reduce the workload of the system's long-tail query mining, and improve the accuracy of data mining.
[0071] In another embodiment, the process of a data mining method for a dialogue system can be as follows: Figure 2 As shown, in this embodiment, a more specific data mining process of a dialogue system is shown, including:
[0072] 1. At the beginning of the process, a recording of the conversation between the user and the dialogue system is collected, such as a WAV audio file of the conversation.
[0073] 2. Use an automatic speech recognition (ASR) model to convert the WAV audio into the original query in text form.
[0074] 3. The original query text output from the ASR model provides basic text data for subsequent processing.
[0075] 4. For the original query, you can use an online NLP system to try to output the query domain and intent. You can also combine it with a corpus, perform embedding processing, generate a vector representation, convert the text into a numerical vector, perform vector similarity comparison and clustering, and determine the query domain and intent. You can also fine-tune the pre-trained language model of the Transformer architecture based on annotated data from a typical corpus and determine the query domain and intent after classification.
[0076] 5. Compare the query domain and intent differences between different methods, identify the different items, verify and screen them through manual regression, mark the qualified items as updated items, and update the online NLP system. This provides strong data support for the continuous optimization and upgrade of the dialogue system, ultimately providing users with a more intelligent and convenient dialogue experience. Continuously update and optimize the processing system to improve system performance and user experience.
[0077] In this embodiment, through the mining of difference items and subsequent system iteration, in-depth analysis and mining of conversation data are achieved.
[0078] See Figure 3 , Figure 3 This is a structural diagram of a data mining system for a dialogue system provided by an embodiment of the present application. Figure 3 As shown, the system includes:
[0079] Identification module S302, used to identify audio data, obtain original query, and pre-process the original query;
[0080] Clustering module S304, used to determine whether the online NLP system can recognize the pre-processed original query. If the online NLP system cannot recognize it, the original query is clustered based on vector similarity to match query clusters.
[0081] A difference module S306 is used to compare the query cluster with the classification results of the online NLP system to determine the difference items;
[0082] The updating module S308 is configured to classify the difference item using a pre-trained language model and check whether the classification confidence reaches a preset threshold. When the classification confidence reaches the preset threshold, the difference item is marked as an updated item.
[0083] In one embodiment, the system further comprises:
[0084] An acquisition module is used to acquire historical audio data to construct a corpus and train the historical vectors corresponding to the corpus through an embedding model;
[0085] A grouping module, configured to group the historical vectors using a clustering algorithm to form a query cluster set;
[0086] The search module is configured to obtain a preset similarity metric and search for a number of query clusters in the query cluster set that are closest in similarity to the original query.
[0087] Those skilled in the art will clearly understand that the technical solutions of the embodiments of the present application can be implemented with the help of software and / or hardware. "Unit" and "module" in this specification refer to software and / or hardware that can independently perform or cooperate with other components to perform specific functions, where the hardware can be, for example, a field programmable gate array (FPGA), an integrated circuit (IC), etc.
[0088] Each processing unit and / or module in the embodiments of the present application may be implemented by an analog circuit that implements the functions described in the embodiments of the present application, or may be implemented by software that executes the functions described in the embodiments of the present application.
[0089] See also Figure 4 , which shows a schematic diagram of the structure of an electronic device involved in an embodiment of the present application, the electronic device can be used to implement Figure 1 The method in the embodiment shown. Figure 4 As shown, the electronic device 400 may include: at least one processor 401 , at least one network interface 404 , a user interface 403 , a memory 405 , and at least one communication bus 402 .
[0090] The communication bus 402 is used to implement the connection and communication between these components.
[0091] The user interface 403 may include a display screen (Display) and a camera (Camera). Optionally, the user interface 403 may also include a standard wired interface and a wireless interface.
[0092] The network interface 404 may optionally include a standard wired interface or a wireless interface (such as a WI-FI interface).
[0093] The processor 401 may include one or more processing cores. The processor 401 utilizes various interfaces and circuits to connect the various components within the entire electronic device 400. By running or executing instructions, programs, code sets, or instruction sets stored in the memory 405, and calling data stored in the memory 405, the processor 401 executes various functions of the terminal 400 and processes data. Optionally, the processor 401 may be implemented in at least one hardware form of a digital signal processing (DSP), a field-programmable gate array (FPGA), or a programmable logic array (PLA). The processor 401 may integrate one or a combination of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. The CPU primarily processes the operating system, user interface, and application programs; the GPU is responsible for rendering and drawing the content to be displayed on the display; and the modem is used to handle wireless communications. It is understood that the modem may not be integrated into the processor 401 and may be implemented separately on a single chip.
[0094] Among them, the memory 405 may include a random access memory (RAM) or a read-only memory (Read-Only Memory). Optionally, the memory 405 includes a non-transitory computer-readable storage medium. The memory 405 can be used to store instructions, programs, codes, code sets or instruction sets. The memory 405 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area may store data involved in the above-mentioned various method embodiments, etc. The memory 405 may also be optionally at least one storage device located away from the aforementioned processor 401. As Figure 4 As shown, the memory 405 as a computer storage medium may include an operating system, a network communication module, a user interface module, and program instructions.
[0095] exist Figure 4In the electronic device 400 shown, the user interface 403 is mainly used to provide an input interface for the user and obtain the data input by the user; and the processor 401 can be used to call the image-generated interactive application stored in the memory 405, and specifically perform the following operations: recognize audio data, obtain the original query, and preprocess the original query; determine whether the online NLP system can recognize the preprocessed original query, when the online NLP system cannot recognize it, perform vector similarity clustering on the original query and match the query cluster; compare the query cluster with the classification result of the online NLP system to determine the difference items; classify the difference items through the pre-trained language model, and check whether the classification confidence reaches the preset threshold. When the classification confidence reaches the preset threshold, mark the difference item as an update item.
[0096] The present application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above method. The computer-readable storage medium may include, but is not limited to, any type of disk, including a floppy disk, an optical disk, a DVD, a CD-ROM, a microdrive, a magneto-optical disk, a ROM, a RAM, an EPROM, an EEPROM, a DRAM, a VRAM, a flash memory device, a magnetic card or an optical card, a nanosystem (including a molecular memory IC), or any type of medium or device suitable for storing instructions and / or data.
[0097] It should be noted that for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.
[0098] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0099] In the several embodiments provided in this application, it should be understood that the disclosed devices can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some service interface, and the indirect coupling or communication connection of the device or unit can be electrical or other forms.
[0100] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0101] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0102] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a memory, including a number of instructions for enabling a computer device (which can be a personal computer, server or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned memory includes: various media that can store program codes, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.
[0103] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable memory, which may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0104] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
Claims
1. A data mining method for a dialogue system, characterized in that: include: Recognize audio data to obtain an original query, and preprocess the original query; Determine whether the online NLP system can recognize the pre-processed original query. If the online NLP system cannot recognize it, perform vector similarity clustering on the original query to match query clusters. Comparing the query cluster with the classification results of the online NLP system to determine the differences; The difference item is classified by a pre-trained language model, and a classification confidence is checked to see whether it reaches a preset threshold. When the classification confidence reaches the preset threshold, the difference item is marked as an updated item.
2. The data mining method for a dialogue system according to claim 1, characterized in that: Determining whether the online NLP system can recognize the preprocessed original query includes: The online NLP system includes a rule-based dialogue system and a machine learning recognition system; Determine whether the original query matches the rules of the rule-based dialogue system, and determine whether the original query belongs to the domain and intent in the machine learning recognition system.
3. The data mining method for a dialogue system according to claim 1, characterized in that: The performing vector similarity clustering on the original query to match the query cluster includes: Obtain historical audio to build a corpus, and train the historical vector corresponding to the corpus through the Embedding model; Grouping the historical vectors by a clustering algorithm to form a query cluster set; A preset similarity metric is obtained, and several query clusters in the query cluster set that are closest in similarity to the original query are found.
4. The data mining method for a dialogue system according to claim 3, characterized in that: The classifying the difference items by using a pre-trained language model includes: Selecting a pre-trained model and training the corpus based on the pre-trained model, wherein the corpus includes historical queries and annotation information, and the annotation information includes domain and intent; Adjust the model structure and model parameters based on system requirements to obtain the trained pre-trained language model; The difference items are input into a pre-trained language model, and a classification result is output.
5. The data mining method for a dialogue system according to claim 4, characterized in that: The step of checking whether the classification confidence reaches a preset threshold includes: When the domain and intent of the difference item classification result are the same as the clustering domain and intent of any query cluster, it is determined whether the classification confidence of the difference item classification result is greater than a preset threshold.
6. The data mining method for a dialogue system according to claim 1, characterized in that: The method further comprises: The difference items are transmitted to a manual regression verification system. When the difference items pass the screening of the manual regression verification system, the difference items are marked as update items, and the online NLP system is updated.
7. A data mining system for a dialogue system, characterized in that: The system comprises: A recognition module, configured to recognize audio data, obtain an original query, and preprocess the original query; A clustering module is used to determine whether the online NLP system can recognize the pre-processed original query. If the online NLP system cannot recognize it, the original query is clustered based on vector similarity to match query clusters. A difference module, used to compare the query cluster with the classification results of the online NLP system to determine the difference items; An updating module is used to classify the difference item through a pre-trained language model and check whether the classification confidence reaches a preset threshold. When the classification confidence reaches the preset threshold, the difference item is marked as an updated item.
8. The data mining system for a dialogue system according to claim 7, characterized in that: The system further comprises: An acquisition module is used to acquire historical audio data to construct a corpus and train the historical vectors corresponding to the corpus through an embedding model; A grouping module, configured to group the historical vectors using a clustering algorithm to form a query cluster set; The search module is configured to obtain a preset similarity metric and search for a number of query clusters in the query cluster set that are closest in similarity to the original query.
9. An electronic device comprising a processor and a memory; The processor is connected to the memory; The memory is used to store executable program code; The processor reads the executable program code stored in the memory to run a program corresponding to the executable program code, so as to execute the method according to any one of claims 1 to 6.
10. A computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the method according to any one of claims 1 to 6 when executed by a processor.