An intelligent scenario dialogue analysis method and system based on model recognition
By collecting and processing historical dialogue data on the intelligent scene dialogue platform, selecting the best performance model, and analyzing real-time dialogue data, the problem of difficulty in handling noise and error information in multilingual dialogue text in existing technology is solved, and high-quality dialogue analysis and interaction efficiency are achieved.
Patent Information
- Application Number
- CN202510414754.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-04-03
AI Technical Summary
Existing intelligent scene dialogue analysis methods based on model recognition are difficult to deal with noise and error information in multilingual dialogue texts, and cannot accurately participle words and annotate part of speech.
By collecting historical dialogue data on the intelligent scene dialogue platform, performing text cleaning, word segmentation and part-of-speech annotation, high-quality historical dialogue standard data are obtained, and the model with the best performance is selected as the basic model through comparative experiments, and real-time dialogue data is analyzed.
It realizes high-quality processing of multilingual dialogue text, improves the accuracy of word segmentation and part-of-speech labeling, and improves the efficiency and quality of dialogue interaction.
Smart Images

Figure CN119940345B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of scenario dialogue analysis, and specifically relates to an intelligent scenario dialogue analysis method and system based on model recognition. Background Art
[0002] With the development of artificial intelligence technology, intelligent scenario dialogue applications are becoming increasingly widespread, such as intelligent customer service, smart home control, etc. However, existing dialogue analysis methods are difficult to accurately understand complex and diverse dialogue scenarios. The intelligent scenario dialogue analysis method and system based on model recognition can utilize advanced model recognition technology to accurately identify dialogue scenarios, understand semantics and intentions, and improve the efficiency and quality of dialogue interaction, which is of great significance in enhancing user experience, optimizing services, etc.
[0003] Existing intelligent scenario dialogue analysis methods and systems based on model recognition are difficult to fully process the noise and error information in dialogue data when facing dialogue texts containing multiple languages, and at the same time, they are unable to accurately segment words and label parts of speech. Therefore, it is necessary to provide an intelligent scenario dialogue analysis method and system based on model recognition to solve the above-mentioned problems. Summary of the Invention
[0004] To solve the above technical problems, an intelligent scenario dialogue analysis method and system based on model recognition are provided. This technical solution solves the problems that existing intelligent scenario dialogue analysis methods and systems based on model recognition are difficult to fully process the noise and error information in dialogue data when facing dialogue texts containing multiple languages, and at the same time, they are unable to accurately segment words and label parts of speech as mentioned in the above background art.
[0005] To achieve the above object, the technical solution adopted by the present invention is as follows:
[0006] An intelligent scenario dialogue analysis method based on model recognition, comprising:
[0007] Collect historical dialogue data on an intelligent scenario dialogue platform, perform text cleaning on the historical dialogue data to obtain historical dialogue text information;
[0008] Segment words and label parts of speech for the historical dialogue text information to obtain historical dialogue text part-of-speech information;
[0009] Perform data annotation on the historical dialogue data according to the historical dialogue text information and the historical dialogue text part-of-speech information to obtain historical dialogue standard data;
[0010] Obtain the length information of each historical dialogue text in the historical dialogue standard data, denoted as historical dialogue word length sub-information;
[0011] Divide the historical dialogue standard data according to the historical dialogue text length sub-information to obtain the historical dialogue standard sub-data;
[0012] Select the corresponding model to conduct a comparative experiment on the historical dialogue standard sub-data, and at the same time complete the construction of the initial model and obtain the accuracy corresponding to each initial model;
[0013] Select the initial model with the best performance as the basic model according to the accuracy corresponding to each initial model to obtain the basic model for intelligent scenario dialogue analysis;
[0014] Obtain the user's real-time dialogue data, preprocess the user's real-time dialogue data to obtain the dialogue data to be analyzed, and finally use the basic model for intelligent scenario dialogue analysis to analyze the dialogue data to be analyzed.
[0015] In an alternative embodiment, collecting historical dialogue data on the intelligent scenario dialogue platform and cleaning the historical dialogue data to obtain historical dialogue text information specifically includes:
[0016] Obtain the type information of the intelligent scenario dialogue platform, thereby determining the corresponding interface, and use the SDK provided by the official of the intelligent scenario dialogue platform to collect public dialogue data to obtain historical dialogue data;
[0017] Based on the intelligent scenario dialogue platform, obtain the preset scenario category information corresponding to the historical dialogue data, and preliminarily classify the historical dialogue data according to the preset scenario category information to obtain historical dialogue data of different preset scenario categories;
[0018] Store the historical dialogue data of different preset scenario categories in the same data set, and sequentially clean the historical dialogue data in the data set, including:
[0019] S1.1. Use regular expressions to remove the noise information in the dialogue text corresponding to the historical dialogue data;
[0020] S1.2. According to the preset scenario category information corresponding to the historical dialogue data, call the corresponding corpus, and use the historical dialogue data and the corresponding corpus to train an error correction model corresponding to the preset scenario category information. Then, use the error correction model to correct the dialogue text in the historical dialogue data, and use the corrected dialogue text in the historical dialogue data as the historical dialogue text information;
[0021] Among them, the intelligent scenario dialogue platform includes an online customer service platform, an intelligent voice assistant, and a social media platform.
[0022] In an alternative embodiment, the tokenization and part-of-speech tagging of the historical dialogue text information to obtain the historical dialogue text part-of-speech information specifically includes:
[0023] Obtain the language type information corresponding to the dialogue text in the historical dialogue data after error correction, and perform Chinese-English segmentation on the dialogue text in the historical dialogue data after error correction according to the language type information, while recording the segmentation position information;
[0024] Segment the Chinese and English in the historical dialogue text to obtain the historical dialogue Chinese text and the historical dialogue English text;
[0025] Select corresponding word segmentation tools to perform word segmentation on the historical dialogue Chinese text and the historical dialogue English text, including:
[0026] S2.1. Based on the intelligent scenario dialogue platform, obtain the preset scenario category information corresponding to the historical dialogue data again, so as to obtain the preset scenario domain information;
[0027] S2.2. According to the preset scenario domain information, load the corresponding custom Chinese dictionary, and use the Jieba word segmentation tool to segment the historical dialogue Chinese text, cut the Chinese text into individual Chinese words, and then store the Chinese words corresponding to each historical dialogue Chinese text in the same list to obtain the first historical dialogue text information;
[0028] S2.2. Load the corresponding custom English dictionary, use the word segmentation function in the NLTK library to perform word segmentation on the historical dialogue English text, cut the English text into individual English words, and then store the English words corresponding to each historical dialogue English text in the same list to obtain the second historical dialogue text information;
[0029] Perform part-of-speech tagging on the first historical dialogue text information and the second historical dialogue text information, including:
[0030] S3.1. Based on the first historical dialogue text information, obtain all the Chinese pinyin corresponding to the Chinese words in the first historical dialogue text information, and obtain all the homophonic Chinese words corresponding to the Chinese words based on the Chinese pinyin;
[0031] S3.2. Based on the second historical dialogue text information, traverse each English word in the second historical dialogue text information, and obtain all the recombined English words corresponding to each English word;
[0032] S3.3. Load the custom Chinese dictionary corresponding to Chinese words and the custom English dictionary corresponding to English texts respectively according to the preset scenario category information corresponding to the historical dialogue data. Then, use the preset scenario category information, the custom Chinese dictionary, and the custom English dictionary to screen all homophonic Chinese words and recombined English words, and extract the homophonic Chinese words and recombined English words included in the custom Chinese dictionary and the custom English dictionary to obtain the first historical dialogue standard text information and the second historical dialogue standard text information;
[0033] S3.4. Use the Jieba word segmentation tool to perform part-of-speech tagging on the Chinese words in the first historical dialogue standard text information, and use the part-of-speech tagger in the NLTK library to perform part-of-speech tagging on the English words in the second historical dialogue standard text information. At the same time, use the part-of-speech tagging results of the first historical dialogue standard text information and the second historical dialogue standard text information to record the historical dialogue text part-of-speech information;
[0034] S3.5. Recombine the first historical dialogue standard text information and the second historical dialogue standard text information after part-of-speech tagging through the segmentation position information to update the historical dialogue text information.
[0035] In an optional embodiment, the data annotation of the historical dialogue data according to the historical dialogue text information and the historical dialogue text part-of-speech information to obtain the historical dialogue standard data specifically includes:
[0036] Send the updated historical dialogue text information to professional annotators for annotation to obtain the historical dialogue text part-of-speech manual annotation information;
[0037] Compare the historical dialogue text part-of-speech information with the historical dialogue text part-of-speech manual annotation information, and remove the historical dialogue text part-of-speech information that does not match the historical dialogue manual annotation information to obtain the historical dialogue text part-of-speech standard information;
[0038] Use the historical dialogue text part-of-speech standard information to perform data annotation on the historical dialogue data to obtain the historical dialogue standard data.
[0039] In an optional embodiment, the obtaining of the length information of each historical dialogue text in the historical dialogue standard data, denoted as the historical dialogue word length sub-information, specifically includes:
[0040] Obtain the total number of Chinese characters of the Chinese words and the total number of letters of the English words in each historical dialogue text in the historical dialogue standard data;
[0041] Take the sum of the total number of Chinese characters of the Chinese words and the total number of letters of the English words in each historical dialogue text in the historical dialogue standard data as the length information of each historical dialogue text in the historical dialogue standard data to obtain the historical dialogue word length sub-information.
[0042] In an alternative embodiment, the dividing the historical dialogue standard data according to the historical dialogue text length sub-information to obtain historical dialogue standard sub-data specifically includes:
[0043] Traverse all the historical dialogue text length sub-information to obtain the median of all historical dialogue text lengths;
[0044] Divide the historical dialogue standard data with historical dialogue text lengths greater than or equal to the median of all historical dialogue text lengths into first historical dialogue standard data, and divide the historical dialogue standard data with historical dialogue text lengths less than the median of all historical dialogue text lengths into second historical dialogue standard data;
[0045] Store the first historical dialogue standard data and the second historical dialogue standard data in different data sets respectively to obtain historical dialogue standard sub-data.
[0046] In an alternative embodiment, the selecting corresponding models to conduct comparative experiments on the historical dialogue standard sub-data, simultaneously completing the construction of the initial models, and obtaining the accuracy corresponding to each initial model specifically includes:
[0047] Extract the historical dialogue standard data corresponding to the first historical dialogue standard data and the second historical dialogue standard data from the historical dialogue standard sub-data respectively, and divide them into a first training set, a first validation set, a first test set, a second training set, a second validation set, and a second test set according to 70% training set, 15% validation set, and 15% test set respectively;
[0048] Obtain a long short-term memory network model, a gated recurrent unit model, and a Transformer model respectively, and set initial hyperparameters for the long short-term memory network model, the gated recurrent unit model, and the Transformer model to complete the construction of the initial models;
[0049] Train the initial models using the first training set and the second training set respectively, and based on the first validation set, the first test set, the second validation set, and the second test set, through , obtain the accuracy corresponding to each initial model, where TP is the true positive, TN is the true negative, FP is the false positive, and FN is the false negative.
[0050] In an alternative embodiment, the selecting the initial model with the best performance as the basic model according to the accuracy corresponding to each initial model to obtain the intelligent scenario dialogue analysis basic model specifically includes:
[0051] Traverse all the initial models, and take the initial model with the highest accuracy as the initial model with the best performance, which is the basic model, so as to obtain the intelligent scenario dialogue analysis basic model;
[0052] Among them, the intelligent scenario dialogue analysis basic model includes a long dialogue analysis basic model and a short dialogue analysis basic model.
[0053] In an optional embodiment, the method of obtaining the user's real-time dialogue data, preprocessing the user's real-time dialogue data to obtain the dialogue data to be analyzed, and finally using the intelligent scenario dialogue analysis basic model to analyze the dialogue data to be analyzed specifically includes:
[0054] Obtain the length information of the dialogue text to be analyzed in the dialogue data to be analyzed, and obtain the median of the lengths of all historical dialogue texts;
[0055] If the length of the dialogue text to be analyzed is greater than or equal to the median of the lengths of all historical dialogue texts, use the long dialogue analysis basic model to analyze the dialogue data to be analyzed;
[0056] If the length of the dialogue text to be analyzed is less than the median of the lengths of all historical dialogue texts, use the short dialogue analysis basic model to analyze the dialogue data to be analyzed.
[0057] Furthermore, an intelligent scenario dialogue analysis system based on model recognition is proposed, which is used to implement the analysis method as described in any one of the above, including:
[0058] A collection module, which is used to collect historical dialogue data on the intelligent scenario dialogue platform;
[0059] A data processing module, which is used to clean the historical dialogue data to obtain historical dialogue text information, perform word segmentation and part-of-speech tagging on the historical dialogue text information to obtain historical dialogue text part-of-speech information, perform data annotation on the historical dialogue data according to the historical dialogue text information and the historical dialogue text part-of-speech information to obtain historical dialogue standard data, obtain the length information of each historical dialogue text in the historical dialogue standard data, record it as historical dialogue text length sub-information, and divide the historical dialogue standard data according to the historical dialogue text length sub-information to obtain historical dialogue standard sub-data;
[0060] A basic model construction module, which is used to select a corresponding model to conduct a comparative experiment on the historical dialogue standard sub-data, complete the construction of the initial model at the same time, obtain the accuracy rate corresponding to each initial model, and select the initial model with the best performance as the basic model according to the accuracy rate corresponding to each initial model to obtain the intelligent scenario dialogue analysis basic model;
[0061] An analysis module, which is used to obtain real-time user conversation data, preprocess the real-time user conversation data to obtain conversation data to be analyzed, and finally analyze the conversation data to be analyzed using an intelligent scenario conversation analysis basic model.
[0062] Compared with the prior art, the beneficial effects of the present invention are:
[0063] A method and system for intelligent scenario conversation analysis based on model recognition proposed by this solution removes noise information in the conversation text of historical conversation data through regular expressions, and then calls a corpus to train an error correction model according to preset scenario category information to correct the conversation text, so as to obtain high-quality historical conversation text information;
[0064] A method and system for intelligent scenario conversation analysis based on model recognition proposed by this solution performs Chinese-English segmentation according to language type information, loads a custom dictionary in combination with preset scenario domain information, and uses the Jieba word segmentation tool and the word segmentation function and part-of-speech tagger in the NLTK library for processing respectively. It also improves the accuracy of word segmentation and part-of-speech tagging through operations such as screening homophonic words and recombining words, and extracts text features more accurately. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] Figure 1 It is a flowchart of a method for intelligent scenario conversation analysis based on model recognition proposed by the present invention;
[0066] Figure 2 It is a flowchart of word segmentation for historical conversation text information in the present invention;
[0067] Figure 3 It is a flowchart of part-of-speech tagging for historical conversation text information in the present invention;
[0068] Figure 4 It is a system framework diagram of a system for intelligent scenario conversation analysis based on model recognition proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0069] The following description is used to disclose the present invention so that those skilled in the art can implement the present invention. The preferred embodiments described below are only examples, and those skilled in the art can think of other obvious variations.
[0070] Referring to Figure 1 - Figure 4 As shown, a method for intelligent scenario conversation analysis based on model recognition includes:
[0071] Collect historical conversation data on an intelligent scenario conversation platform, perform text cleaning on the historical conversation data to obtain historical conversation text information;
[0072] Tokenize and perform part-of-speech tagging on the historical dialogue text information to obtain the part-of-speech information of the historical dialogue text;
[0073] Based on the historical dialogue text information and the part-of-speech information of the historical dialogue text, perform data annotation on the historical dialogue data to obtain the historical dialogue standard data;
[0074] Obtain the length information of each historical dialogue text in the historical dialogue standard data, denoted as the sub-information of the historical dialogue text length;
[0075] Divide the historical dialogue standard data according to the sub-information of the historical dialogue text length to obtain the historical dialogue standard sub-data;
[0076] Select the corresponding model to conduct a comparative experiment on the historical dialogue standard sub-data, and at the same time complete the construction of the initial model, and obtain the accuracy rate corresponding to each initial model;
[0077] Select the initial model with the best performance as the basic model according to the accuracy rate corresponding to each initial model to obtain the intelligent scenario dialogue analysis basic model;
[0078] Obtain the user's real-time dialogue data, perform data preprocessing on the user's real-time dialogue data to obtain the dialogue data to be analyzed, and finally use the intelligent scenario dialogue analysis basic model to analyze the dialogue data to be analyzed.
[0079] Furthermore, collect historical dialogue data on the intelligent scenario dialogue platform, perform text cleaning on the historical dialogue data to obtain the historical dialogue text information, specifically including:
[0080] Obtain the type information of the intelligent scenario dialogue platform, so as to determine the corresponding interface to collect the public dialogue data using the SDK provided by the official of the intelligent scenario dialogue platform to obtain the historical dialogue data;
[0081] Based on the intelligent scenario dialogue platform, obtain the preset scenario category information corresponding to the historical dialogue data, and perform preliminary classification on the historical dialogue data according to the preset scenario category information to obtain the historical dialogue data of different preset scenario categories;
[0082] Store the historical dialogue data of different preset scenario categories in the same data set, and perform text cleaning on the historical dialogue data in the data set in turn, including:
[0083] S1.1. Use regular expressions to remove the noise information in the corresponding dialogue text of the historical dialogue data;
[0084] S1.2. According to the preset scenario category information corresponding to the historical conversation data, call the corresponding corpus, and use the historical conversation data and the corresponding corpus to train an error correction model corresponding to the preset scenario category information. Then, use the error correction model to correct the conversation text in the historical conversation data, and use the corrected conversation text in the historical conversation data as the historical conversation text information;
[0085] Among them, the intelligent scenario dialogue platform includes an online customer service platform, an intelligent voice assistant, and a social media platform.
[0086] Specifically, in this embodiment, first, it is necessary to determine the specific type of the intelligent scenario dialogue platform. For example, it is an online customer service platform (such as a customer service system used by an enterprise to communicate with customers), an intelligent voice assistant (such as a voice interaction assistant on a mobile phone), or a social media platform (such as WeChat, Weibo, etc.). After clarifying the platform type, find the corresponding interface according to this type. Use the software development kit (SDK) provided by the official of the intelligent scenario dialogue platform, and through this interface, collect the public dialogue data on the platform, so as to obtain the historical conversation data. Then, based on the obtained historical conversation data, obtain the preset scenario category information corresponding to these historical conversation data from the intelligent scenario dialogue platform. For example, in the online customer service platform, the preset scenario categories may include product consultation, after-sales service, complaint and suggestion, etc.; in the social media platform, there may be scenario categories such as daily communication, topic discussion, advertising, etc. Then, according to these preset scenario category information, conduct a preliminary classification of the historical conversation data, and classify the historical conversation data belonging to the same preset scenario category into one category, so as to obtain historical conversation data of different preset scenario categories. Then, store the historical conversation data of these different preset scenario categories in the same data set. Next, perform text cleaning operations on the historical conversation data in this data set in turn:
[0087] S1.1: Use regular expressions to process the corresponding conversation text in the historical conversation data. Regular expressions are a powerful text processing tool that can match and remove noise information in the conversation text according to specific rules, such as some special characters, garbled characters, irrelevant punctuation marks, etc., to make the conversation text cleaner and tidier.
[0088] S1.2: According to the preset scenario category information corresponding to the historical conversation data obtained previously, call the corresponding corpus. A large amount of language data related to this preset scenario is stored in the corpus. Then, use the historical conversation data and the corresponding corpus to train an error correction model, which is specifically for this preset scenario category information. After training, use this error correction model to correct the conversation texts in the historical conversation data, find and correct the errors in the conversation texts, such as typos and grammar errors. Finally, determine the conversation texts in the historical conversation data after error correction as the historical conversation text information.
[0089] Furthermore, perform word segmentation and part-of-speech tagging on the historical conversation text information to obtain the historical conversation text part-of-speech information, specifically including:
[0090] Obtain the language type information corresponding to the conversation text in the historical conversation data after error correction, and perform Chinese-English segmentation processing on the conversation text in the historical conversation data after error correction according to the language type information, and record the segmentation position information at the same time;
[0091] Segment the Chinese and English in the historical conversation text to obtain the historical conversation Chinese text and the historical conversation English text;
[0092] Select the corresponding word segmentation tool to perform word segmentation processing on the historical conversation Chinese text and the historical conversation English text, including:
[0093] S2.1: Based on the intelligent scenario dialogue platform, obtain the preset scenario category information corresponding to the historical conversation data again, so as to obtain the preset scenario domain information;
[0094] S2.2: According to the preset scenario domain information, load the corresponding custom Chinese dictionary, and use the Jieba word segmentation tool to segment the historical conversation Chinese text, cut the Chinese text into individual Chinese words and phrases, and then store the Chinese words and phrases corresponding to each historical conversation Chinese text into the same list to obtain the first historical conversation text information;
[0095] S2.2: Load the corresponding custom English dictionary, use the word segmentation function in the NLTK library to perform word segmentation processing on the historical conversation English text, cut the English text into individual English words, and then store the English words and phrases corresponding to each historical conversation English text into the same list to obtain the second historical conversation text information;
[0096] Perform part-of-speech tagging on the first historical conversation text information and the second historical conversation text information, including:
[0097] S3.1. Based on the first historical dialogue text information, obtain all the Chinese pinyin corresponding to the Chinese words and phrases in the first historical dialogue text information, and based on the Chinese pinyin, obtain all the Chinese words and phrases with the same pronunciation as the Chinese words and phrases;
[0098] S3.2. Based on the second historical dialogue text information, traverse each English word in the second historical dialogue text information, and obtain all the recombined English words corresponding to each English word;
[0099] S3.3. According to the preset scenario category information corresponding to the historical dialogue data, load the custom Chinese dictionary corresponding to the Chinese words and phrases and the custom English dictionary corresponding to the English text respectively. Then, use the preset scenario category information, the custom Chinese dictionary and the custom English dictionary to screen all the Chinese words and phrases with the same pronunciation and the recombined English words, and extract the Chinese words and phrases with the same pronunciation and the recombined English words included in the custom Chinese dictionary and the custom English dictionary to obtain the first historical dialogue standard text information and the second historical dialogue standard text information;
[0100] S3.4. Use the Jieba word segmentation tool to perform part-of-speech tagging on the Chinese words and phrases in the first historical dialogue standard text information, use the part-of-speech tagger in the NLTK library to perform part-of-speech tagging on the English words in the second historical dialogue standard text information, and at the same time, use the part-of-speech tagging results of the first historical dialogue standard text information and the second historical dialogue standard text information to record the historical dialogue text part-of-speech information;
[0101] S3.5. Through the segmentation position information, recombine the first historical dialogue standard text information and the second historical dialogue standard text information after part-of-speech tagging, and update the historical dialogue text information.
[0102] Specifically, in this embodiment, to obtain the language type information, the language type of the dialogue text can be judged by simple character features. For example, if most of the characters in the text are letters within the ASCII code range, it is judged as English; if it contains a large number of Chinese characters (judged by the Unicode range, such as \u4e00-\u9fa5), it is judged as Chinese. For a text mixed with Chinese and English, the main language type can be determined by further analyzing the character distribution ratio. Some mature language detection libraries can also be used, such as the langdetect library in Python, which can conveniently detect the language type of the text. For Chinese-English segmentation and recording the positions, traverse the dialogue text and distinguish Chinese and English according to the Unicode range of the characters. When encountering Chinese characters, mark them as the Chinese part; when encountering English letters, mark them as the English part. At the same time, record the start and end positions of the segmentation.
[0103] It can be understood that for obtaining preset scenario domain information in word segmentation, the preset scenario category information corresponding to historical dialogue data is obtained from the configuration file or database of the intelligent scenario dialogue platform. For example, in the e-commerce customer service scenario, the preset scenario categories may include product consultation, order query, after-sales processing, etc. According to these scenario categories, the corresponding preset scenario domain information is further determined, such as the types of products, the order process, etc. Chinese word segmentation requires installing the jieba library, using the jieba.load_userdict() method to load the custom Chinese dictionary, and then using the jieba.cut() method to segment the Chinese text of the historical dialogue. For English word segmentation, the NLTK library needs to be installed, and the nltk.tokenize.word_tokenize() method is used to segment the English text of the historical dialogue.
[0104] When performing part-of-speech tagging, to obtain homophonous Chinese words and recombined English words, some open-source pinyin libraries such as the pypinyin library can be used to obtain the pinyin of Chinese words, and then homophonous words are found through the pinyin to obtain Chinese homophonous words. For recombined English words, recombined English words are obtained by traversing the permutations and combinations of English letters. For screening standard text information, it is necessary to load the custom Chinese dictionary and English dictionary, traverse the homophonous Chinese words and recombined English words, and judge whether they are in the corresponding dictionary. If they are in the dictionary, they are extracted to obtain the first historical dialogue standard text information and the second historical dialogue standard text information.
[0105] For Chinese part-of-speech tagging, the jieba.posseg module can be used to perform part-of-speech tagging on Chinese words in the first historical dialogue standard text information. For English part-of-speech tagging, the nltk.pos_tag() method in the NLTK library is used to perform part-of-speech tagging on English words in the second historical dialogue standard text information. Finally, according to the previously recorded segmentation position information, the part-of-speech tagged first historical dialogue standard text information and the second historical dialogue standard text information are recombined to update the historical dialogue text information.
[0106] Furthermore, according to the historical dialogue text information and the historical dialogue text part-of-speech information, the historical dialogue data is data-labeled to obtain the historical dialogue standard data, specifically including:
[0107] The updated historical dialogue text information is sent to professional annotators for annotation to obtain the historical dialogue text part-of-speech manual annotation information;
[0108] The historical dialogue text part-of-speech information and the historical dialogue text part-of-speech manual annotation information are compared, and the historical dialogue text part-of-speech information that does not match the historical dialogue manual annotation information is removed to obtain the historical dialogue text part-of-speech standard information;
[0109] The historical conversation data is annotated using the standard part-of-speech information of the historical conversation text to obtain the standard data of the historical conversation.
[0110] Specifically, the part-of-speech manual tagging information of historical dialogue texts is obtained: first, the updated historical dialogue text information obtained after a series of previous processing (including text cleaning, Chinese-English segmentation, word segmentation, part-of-speech tagging, and text information update) is sent to professional taggers in an appropriate way (such as using a special data tagging platform, sending emails, or using specific project management tools). These professional taggers have relevant language knowledge and tagging experience. They will tag each word in the updated historical dialogue text information according to professional standards and understanding of text semantics, thereby obtaining the part-of-speech manual tagging information of historical dialogue texts. For example, in a dialogue text about e-commerce consultation, the tagger will clearly point out that the part of speech of the word "commodity" is a noun, and the part of speech of the word "purchase" is a verb, etc. Obtain the standard part-of-speech information of historical dialogue texts: Then, the part-of-speech information of historical dialogue texts obtained by automatic processing of the program (that is, the part-of-speech information obtained after Chinese-English segmentation, word segmentation, part-of-speech tagging, etc.) is carefully compared with the part-of-speech manual tagging information of historical dialogue texts given by professional taggers. During the comparison process, the word-by-word part-of-speech tagging results of the two are checked to see if they are consistent. For those parts of speech information of historical dialogue text that do not match the manual tagging information of historical dialogues, that is, the parts where the program's automatic tagging results are different from the manual tagging results, they are removed from the original part-of-speech information of historical dialogue texts. After such screening and processing, the remaining part-of-speech information of historical dialogue texts is determined as the standard part-of-speech information of historical dialogue texts. For example, if the program automatically tags the word "fast" as a noun, and the manual tagging is an adjective, since the two do not match, the result of "noun" automatically marked by the program is removed, and the tagging result of "adjective" that matches the manual tagging is finally retained. Obtaining the standard data of historical dialogues: Finally, the obtained standard part-of-speech information of historical dialogue texts is used to perform comprehensive data tagging on the original historical dialogue data. This means that the standard part-of-speech information of historical dialogue texts is accurately applied to each section of historical dialogue data, and each word in the historical dialogue data is given an accurate, manually reviewed and confirmed part-of-speech tag. After such labeling operations, standard data of historical conversations are obtained. These data can be used for subsequent model training, analysis and other tasks, providing a high-quality data foundation for intelligent scene conversation analysis.
[0111] Furthermore, the length information of each historical conversation text in the historical conversation standard data is obtained and recorded as the historical conversation text length sub-information, which specifically includes:
[0112] Obtain the total number of Chinese characters and the total number of letters of English words in each historical conversation text in the historical conversation standard data;
[0113] The sum of the total number of Chinese words and the total number of letters in English words in each historical dialogue text in the historical dialogue standard data is used as the length information of each historical dialogue text in the historical dialogue standard data to obtain the historical dialogue text length sub-information.
[0114] Specifically, first, for each historical dialogue text in the historical dialogue standard data, the total number of Chinese words and the total number of letters in English words need to be counted. Count the total number of Chinese words: traverse the Chinese words in each historical dialogue text. Since the text has been segmented in the previous processing, it is possible to check whether the words in the text are Chinese (by judging the Unicode encoding range of the characters, for example, the characters between \u4e00 and \u9fa5 are Chinese characters). Each time a Chinese word is recognized, the counter is incremented by 1. After all the words in the historical dialogue text are traversed, the value recorded by the counter is the total number of Chinese words in the text. For example, a text "I like apples", in which "I", "like" and "apple" are all Chinese words, the counter starts from 0 and increments by 1 in sequence, and finally the total number of Chinese words is 3. Count the total number of letters in English words: traverse the English words in each historical dialogue text. For each English word, calculate the number of letters it contains. This can be achieved by obtaining the character length of the word (in programming, the length function of a string can return the number of characters in a word). The number of letters in each English word is accumulated into a sum variable. After traversing all the English words in the historical dialogue text, the value recorded in the sum variable is the total number of letters in the English words in this text. For example, a text "I like apples", where the length of "I" is 1, the length of "like" is 4, and the length of "apples" is 6, the total number of letters in the English word after accumulation is 1 + 4 + 6 = 11. Then, the total number of Chinese words and the total number of letters in the English word obtained in each historical dialogue text are added. The result of this addition is used as the length information of the historical dialogue text. The above addition operation is repeated for each historical dialogue text, so as to obtain the length information corresponding to each historical dialogue text. This length information is called the historical dialogue text length sub-information, and the historical dialogue standard data can be further processed and analyzed based on this sub-information, such as classification by length.
[0115] Further, divide the historical dialogue standard data according to the historical dialogue text length sub-information to obtain historical dialogue standard sub-data, specifically including:
[0116] Traverse all the historical dialogue text length sub-information to obtain the median of all historical dialogue text lengths;
[0117] Divide the historical dialogue standard data with the historical dialogue text length greater than or equal to the median of all historical dialogue text lengths into the first historical dialogue standard data, and divide the historical dialogue standard data with the historical dialogue text length less than the median of all historical dialogue text lengths into the second historical dialogue standard data;
[0118] Store the first historical dialogue standard data and the second historical dialogue standard data in different data sets respectively to obtain historical dialogue standard sub-data.
[0119] Specifically, to obtain the median of all historical dialogue text lengths: First, we already have all the historical dialogue text length sub-information, which represents the length of each historical dialogue text. We need to organize these length data. Collect all the historical dialogue text length sub-information into a list or set, and this set contains the text length values corresponding to each historical dialogue. Sort the length values in this set in ascending order. The purpose of sorting is to facilitate the subsequent determination of the median.
[0120] Next, determine the median according to the number of sorted data. If the number of data is odd, then the median is the value in the middle position after sorting; if the number of data is even, the median is usually the average of the two middle values. Through such calculations, we obtain the median of all historical dialogue text lengths.
[0121] Dividing historical dialogue standard data: Prepare the historical dialogue standard data, which are the historical dialogue data with annotation information obtained after a series of previous processes (such as text cleaning, part-of-speech tagging, etc.). Check the historical dialogue text length sub-information corresponding to each piece of historical dialogue standard data in turn. Compare this length with the median we just calculated. If the text length of a certain historical dialogue is greater than or equal to the median, then classify this piece of historical dialogue standard data into the first historical dialogue standard data. This part of the data represents relatively long historical dialogues. If the text length of a certain historical dialogue is less than the median, classify it into the second historical dialogue standard data. This part of the data represents relatively short historical dialogues. Storing the divided data: Prepare two different data sets. For example, data structures such as lists, dictionaries, or database tables can be used to store the data. Store the first historical dialogue standard data in one of the data sets, which is specifically used to store the longer historical dialogue standard data. Store the second historical dialogue standard data in another data set, which is specifically used to store the shorter historical dialogue standard data. After such storage operations, we obtain the historical dialogue standard sub-data, which are stored in two different data sets respectively, facilitating subsequent different processing and analysis of long and short dialogue data, such as training with different models, etc.
[0122] Furthermore, select the corresponding models to conduct comparative experiments on the historical dialogue standard sub-data, and at the same time complete the construction of the initial models and obtain the accuracy corresponding to each initial model, specifically including:
[0123] Extract the historical dialogue standard data corresponding to the first historical dialogue standard data and the second historical dialogue standard data from the historical dialogue standard sub-data respectively, and divide them into a 70% training set, a 15% validation set, and a 15% test set respectively to obtain the first training set, the first validation set, the first test set, the second training set, the second validation set, and the second test set;
[0124] Obtain the long short-term memory network model, the gated recurrent unit model, and the Transformer model respectively, and set the initial hyperparameters for the long short-term memory network model, the gated recurrent unit model, and the Transformer model to complete the construction of the initial models;
[0125] Use the first training set and the second training set to train the initial models respectively, and based on the first validation set, the first test set, the second validation set, and the second test set, through to obtain the accuracy corresponding to each initial model, where TP is the true positive, TN is the true negative, FP is the false positive, and FN is the false negative.
[0126] Specifically, set initial hyperparameters for these three models (LSTM, GRU, Transformer) respectively. Hyperparameters are parameters that need to be set manually before model training, such as learning rate, number of hidden layers, number of neurons, number of iterations, etc. Different hyperparameter settings will affect the training process and performance of the model. By reasonably setting the initial hyperparameters, the initial model is constructed, enabling the model to meet the conditions for starting training. True positive example: The number of samples where the model predicts a positive example (such as predicting that a conversation belongs to a specific category) and the actual situation is indeed a positive example. For example, when judging whether a conversation is a customer complaint, if the model predicts "it is a complaint" and the actual conversation is indeed a customer complaint, such a situation is counted as a true positive example. True negative example: The number of samples where the model predicts a negative example (such as predicting that a conversation does not belong to a specific category) and the actual situation is indeed a negative example. For example, if the model predicts that a conversation "is not a complaint" and it is actually not a complaint, this is a true negative example. False positive example: The number of samples where the model predicts a positive example but the actual situation is a negative example. For example, if the model predicts that a conversation "is a complaint" but in fact the conversation is not a complaint, this situation is a false positive example. False negative example: The number of samples where the model predicts a negative example but the actual situation is a positive example. For example, if the model predicts that a conversation "is not a complaint" but in fact it is a customer complaint, this is a false negative example.
[0127] Furthermore, select the initial model with the best performance as the basic model according to the accuracy rate corresponding to each initial model to obtain the intelligent scenario dialogue analysis basic model, specifically including:
[0128] Traverse all initial models, and take the initial model with the highest accuracy rate as the initial model with the best performance, which is the basic model, so as to obtain the intelligent scenario dialogue analysis basic model;
[0129] Among them, the intelligent scenario dialogue analysis basic model includes a long conversation analysis basic model and a short conversation analysis basic model.
[0130] Furthermore, obtain the user's real-time conversation data, perform data preprocessing on the user's real-time conversation data to obtain the dialogue data to be analyzed, and finally use the intelligent scenario dialogue analysis basic model to analyze the dialogue data to be analyzed, specifically including:
[0131] Obtain the length information of the dialogue text to be analyzed in the dialogue data to be analyzed, and obtain the median of the lengths of all historical dialogue texts;
[0132] If the length of the dialogue text to be analyzed is greater than or equal to the median of the lengths of all historical dialogue texts, use the long conversation analysis basic model to analyze the dialogue data to be analyzed;
[0133] If the length of the dialogue text to be analyzed is less than the median of the lengths of all historical dialogue texts, the short dialogue analysis basic model is used to analyze the dialogue data to be analyzed.
[0134] Specifically, to determine the relationship between the length of the dialogue text to be analyzed and the median, it is necessary to compare the length information of the dialogue text to be analyzed just obtained with the median of the lengths of all historical dialogue texts. Analyze using the long dialogue analysis basic model: If the length of the dialogue text to be analyzed is greater than or equal to the median of the lengths of all historical dialogue texts, it means that this dialogue to be analyzed is relatively long. At this time, select the long dialogue analysis basic model. Input the dialogue data to be analyzed into the long dialogue analysis basic model, and this model will analyze the dialogue data to be analyzed according to the patterns and rules learned from previous training. The analysis process may include identifying the theme, sentiment tendency, intention, etc. of the dialogue, and finally output the analysis result. Analyze using the short dialogue analysis basic model: If the length of the dialogue text to be analyzed is less than the median of the lengths of all historical dialogue texts, it indicates that this dialogue to be analyzed is relatively short. At this time, select the short dialogue analysis basic model. Input the dialogue data to be analyzed into the short dialogue analysis basic model, and this model will analyze the dialogue data to be analyzed based on its own training results. It may also involve analyzing aspects such as theme, sentiment, and intention, and output the corresponding analysis result.
[0135] Furthermore, an intelligent scenario dialogue analysis system based on model recognition is proposed to implement the analysis method as described in any one of the above, including:
[0136] A collection module, which is used to collect historical dialogue data on the intelligent scenario dialogue platform;
[0137] A data processing module, which is used to clean the historical dialogue data to obtain historical dialogue text information, perform word segmentation and part-of-speech tagging on the historical dialogue text information to obtain historical dialogue text part-of-speech information, perform data annotation on the historical dialogue data according to the historical dialogue text information and historical dialogue text part-of-speech information to obtain historical dialogue standard data, obtain the length information of each historical dialogue text in the historical dialogue standard data, denoted as historical dialogue text length sub-information, and divide the historical dialogue standard data according to the historical dialogue text length sub-information to obtain historical dialogue standard sub-data;
[0138] A basic model construction module, which is used to select the corresponding model to conduct a comparative experiment on the historical dialogue standard sub-data, complete the construction of the initial model at the same time, obtain the accuracy corresponding to each initial model, and select the initial model with the best performance as the basic model according to the accuracy corresponding to each initial model to obtain the intelligent scenario dialogue analysis basic model;
[0139] Analysis module, which is used to obtain real-time user conversation data, preprocess the real-time user conversation data to obtain the conversation data to be analyzed, and finally analyze the conversation data to be analyzed using the intelligent scenario conversation analysis basic model.
[0140] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification is only the principle of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.
Claims
1. An intelligent scene dialogue analysis method based on model recognition, characterized in that: include: Collect historical conversation data on the intelligent scene conversation platform, perform text cleaning on the historical conversation data, and obtain historical conversation text information; Perform word segmentation and part-of-speech tagging on the historical dialogue text information to obtain the part-of-speech information of the historical dialogue text; According to the historical conversation text information and the part-of-speech information of the historical conversation text, the historical conversation data is annotated to obtain the historical conversation standard data; Obtaining the length information of each historical conversation text in the historical conversation standard data, and recording it as the historical conversation text length sub-information; Dividing the historical conversation standard data according to the historical conversation text length sub-information to obtain the historical conversation standard sub-data; Select the corresponding model to conduct a comparative experiment on the historical dialogue standard sub-data, complete the construction of the initial model, and obtain the accuracy of each initial model; According to the accuracy of each initial model, the initial model with the best performance is selected as the basic model to obtain the basic model of intelligent scene dialogue analysis; Acquire real-time user conversation data and perform data preprocessing on the real-time user conversation data to obtain conversation data to be analyzed. Finally, use the basic model of intelligent scenario conversation analysis to analyze the conversation data to be analyzed.
2. According to the method of intelligent scene dialogue analysis based on model recognition according to claim 1, it is characterized in that: The collecting of historical conversation data on the intelligent scene conversation platform and text cleaning of the historical conversation data to obtain historical conversation text information specifically includes: Obtain the type information of the smart scene dialogue platform to determine the corresponding interface. Use the SDK officially provided by the smart scene dialogue platform to collect public dialogue data and obtain historical dialogue data. Based on the intelligent scene dialogue platform, the preset scene category information corresponding to the historical dialogue data is obtained, and the historical dialogue data is preliminarily classified according to the preset scene category information to obtain the historical dialogue data of different preset scene categories; The historical conversation data of different preset scenario categories are stored in the same data set, and the text of the historical conversation data in the data set is cleaned in turn, including: S1.
1. Use regular expressions to remove noise information from the corresponding conversation text in the historical conversation data; S1.2, according to the preset scene category information corresponding to the historical dialogue data, calling the corresponding corpus, and using the historical dialogue data and the corresponding corpus to train an error correction model corresponding to the preset scene category information, and then correcting the dialogue text in the historical dialogue data through the error correction model, and using the dialogue text in the historical dialogue data after error correction as the historical dialogue text information; Among them, the intelligent scene dialogue platform includes an online customer service platform, an intelligent voice assistant and a social media platform.
3. The method for intelligent scene dialogue analysis based on model recognition according to claim 1, characterized in that: The word segmentation and part-of-speech tagging of the historical conversation text information to obtain the part-of-speech information of the historical conversation text specifically includes: Obtaining language type information corresponding to the dialogue text in the historical dialogue data after error correction, and performing Chinese-English segmentation processing on the dialogue text in the historical dialogue data after error correction according to the language type information, and recording segmentation position information at the same time; Segment the Chinese and English texts in the historical dialogue text to obtain the historical dialogue Chinese text and the historical dialogue English text; Select the corresponding word segmentation tool to perform word segmentation on the Chinese text of historical dialogues and the English text of historical dialogues, including: S2.
1. Based on the intelligent scene dialogue platform, the preset scene category information corresponding to the historical dialogue data is obtained again, thereby obtaining the preset scene field information; S2.2, according to the preset scene domain information, load the corresponding custom Chinese dictionary, and use the Jieba word segmentation tool to segment the historical conversation Chinese text, cut the Chinese text into individual Chinese words, and then store the Chinese words corresponding to each historical conversation Chinese text into the same list to obtain the first historical conversation text information; S2.2, load the corresponding custom English dictionary, use the word segmentation function in the NLTK library to perform word segmentation on the historical conversation English text, cut the English text into individual English words, and then store the English words corresponding to each historical conversation English text into the same list to obtain the second historical conversation text information; Part-of-speech tagging is performed on the first historical conversation text information and the second historical conversation text information, including: S3.
1. Based on the first historical conversation text information, obtain all Chinese pinyins corresponding to the Chinese words in the first historical conversation text information, and obtain all homophonic Chinese words corresponding to the Chinese words based on the Chinese pinyins; S3.2, based on the second historical conversation text information, traverse each English word in the second historical conversation text information, and obtain all recombined English words corresponding to each English word; S3.3, according to the preset scene category information corresponding to the historical conversation data, the custom Chinese dictionary corresponding to the Chinese words and the custom English dictionary corresponding to the English text are loaded respectively, and then all homophonic Chinese words and reorganized English words are screened by using the preset scene category information, the custom Chinese dictionary and the custom English dictionary, and the homophonic Chinese words and reorganized English words contained in the custom Chinese dictionary and the custom English dictionary are extracted to obtain the first historical conversation standard text information and the second historical conversation standard text information; S3.4, using the Jieba word segmentation tool to perform part-of-speech tagging on the Chinese words in the first historical dialogue standard text information, using the part-of-speech tagger in the NLTK library to perform part-of-speech tagging on the English words in the second historical dialogue standard text information, and using the part-of-speech tagging results of the first historical dialogue standard text information and the second historical dialogue standard text information to record the part-of-speech information of the historical dialogue text; S3.
5. By segmenting the position information, the first historical dialogue standard text information and the second historical dialogue standard text information after part-of-speech tagging are reorganized to update the historical dialogue text information.
4. The method for intelligent scene dialogue analysis based on model recognition according to claim 3 is characterized in that: The historical conversation data is annotated according to the historical conversation text information and the historical conversation text part-of-speech information to obtain the historical conversation standard data, specifically including: The updated historical conversation text information is sent to professional annotators for annotation to obtain manual part-of-speech annotation information of the historical conversation text; Compare the part-of-speech information of the historical conversation text with the part-of-speech manual annotation information of the historical conversation text, remove the part-of-speech information of the historical conversation text that does not match the manual annotation information of the historical conversation, and obtain the standard part-of-speech information of the historical conversation text; The historical conversation data is annotated using the standard part-of-speech information of the historical conversation text to obtain the standard data of the historical conversation.
5. The method for intelligent scene dialogue analysis based on model recognition according to claim 1, characterized in that: The length information of each historical conversation text in the historical conversation standard data is obtained, which is recorded as the historical conversation text length sub-information, and specifically includes: Obtain the total number of Chinese characters and the total number of letters of English words in each historical conversation text in the historical conversation standard data; The sum of the total number of Chinese words and the total number of letters in English words in each historical dialogue text in the historical dialogue standard data is used as the length information of each historical dialogue text in the historical dialogue standard data to obtain the historical dialogue text length sub-information.
6. The method for intelligent scene dialogue analysis based on model recognition according to claim 1, characterized in that: The historical conversation standard data is divided according to the historical conversation text length sub-information to obtain the historical conversation standard sub-data, which specifically includes: Traverse all the historical conversation text length sub-information to obtain the median value of all historical conversation text lengths; Classify the historical conversation standard data whose historical conversation text length is greater than or equal to the median of all historical conversation text lengths as the first historical conversation standard data, and classify the historical conversation standard data whose historical conversation text length is less than the median of all historical conversation text lengths as the second historical conversation standard data; The first historical conversation standard data and the second historical conversation standard data are respectively stored in different data sets to obtain historical conversation standard sub-data.
7. The method for intelligent scene dialogue analysis based on model recognition according to claim 1, characterized in that: The selection of the corresponding model conducts a comparative experiment on the historical dialogue standard sub-data, completes the construction of the initial model, and obtains the accuracy rate corresponding to each initial model, specifically including: Extracting historical conversation standard data corresponding to the first historical conversation standard data and the second historical conversation standard data from the historical conversation standard sub-data respectively, and dividing them into 70% training set, 15% validation set and 15% test set respectively, to obtain a first training set, a first validation set, a first test set, a second training set, a second validation set and a second test set; Obtain a long short-term memory network model, a gated recurrent unit model, and a Transformer model respectively, set initial hyperparameters for the long short-term memory network model, the gated recurrent unit model, and the Transformer model, and complete the construction of the initial model; The initial model is trained using the first training set and the second training set respectively, and the first validation set, the first test set, the second validation set and the second test set are used to train the initial model. , get the accuracy corresponding to each initial model, where TP is a true positive example, TN is a true negative example, FP is a false positive example, and FN is a false negative example.
8. The method for intelligent scene dialogue analysis based on model recognition according to claim 1, characterized in that: The method of selecting the best initial model as the basic model according to the accuracy rate corresponding to each initial model to obtain the basic model for intelligent scene dialogue analysis specifically includes: Traverse all initial models and take the initial model with the highest accuracy as the initial model with the best performance, that is, the basic model, so as to obtain the basic model of intelligent scene dialogue analysis; Among them, the intelligent scene dialogue analysis basic model includes a long dialogue analysis basic model and a short dialogue analysis basic model.
9. The method for intelligent scene dialogue analysis based on model recognition according to claim 8, characterized in that: The method of obtaining the user's real-time conversation data and preprocessing the user's real-time conversation data to obtain the conversation data to be analyzed, and finally analyzing the conversation data to be analyzed using the intelligent scene conversation analysis basic model, specifically includes: Obtaining the length information of the conversation text to be analyzed in the conversation data to be analyzed, and obtaining the median value of the length of all historical conversation texts; If the length of the conversation text to be analyzed is greater than or equal to the median length of all historical conversation texts, the long conversation analysis basic model is used to analyze the conversation data to be analyzed; If the length of the conversation text to be analyzed is less than the median length of all historical conversation texts, the short conversation analysis basic model is used to analyze the conversation data to be analyzed.
10. An intelligent scene dialogue analysis system based on model recognition, used to implement the analysis method according to any one of claims 1 to 9, characterized in that: include: A collection module, which is used to collect historical conversation data on the intelligent scene conversation platform; A data processing module, the data processing module is used to perform text cleaning on the historical conversation data to obtain historical conversation text information, perform word segmentation and part-of-speech tagging on the historical conversation text information to obtain the part-of-speech information of the historical conversation text, perform data tagging on the historical conversation data according to the historical conversation text information and the part-of-speech information of the historical conversation text to obtain historical conversation standard data, obtain the length information of each historical conversation text in the historical conversation standard data, record it as historical conversation text length sub-information, divide the historical conversation standard data according to the historical conversation text length sub-information, and obtain historical conversation standard sub-data; A basic model construction module, which is used to select a corresponding model to conduct a comparative experiment on the historical dialogue standard sub-data, complete the construction of the initial model, and obtain the accuracy rate corresponding to each initial model. According to the accuracy rate corresponding to each initial model, the initial model with the best performance is selected as the basic model to obtain the basic model of intelligent scene dialogue analysis; The analysis module is used to obtain the user's real-time conversation data, perform data preprocessing on the user's real-time conversation data, obtain the conversation data to be analyzed, and finally analyze the conversation data to be analyzed using the intelligent scene conversation analysis basic model.
Citation Information
Patent Citations
Telephone interruption recognition method based on semantic recognition and system thereof
CN113488024A
Intelligent dialogue platform based on large language model
CN118551772A