Training method and device of intention recognition model
By obtaining and correcting the initial user intent and training the intent recognition model based on the learning platform function identification text, we can solve the intent recognition errors caused by user accents and idioms, improve the generalization and fault tolerance of the model, and enhance the user experience of the learning platform.
Patent Information
- Application Number
- CN202510937341.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-08
- Publication Date
- 2025-09-30
AI Technical Summary
Existing intent recognition models have difficulty accurately identifying user intent when faced with diverse expressions of user accents and idioms, resulting in misrecognition or failure to recognize, affecting user experience.
By obtaining the initial text, determining the initial user intent, and modifying it based on the function identification text of the learning platform, standard user intent is generated, and a training dataset is constructed until an intent recognition model with strong generalization and fault tolerance is trained.
The recognition accuracy of the intent recognition model in the face of user accents and voice recognition errors has been improved, which has enhanced the practicality and user experience of the learning platform in voice interaction scenarios.
Smart Images

Figure CN120723915A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and more particularly to a method for training an intent recognition model. The present application also relates to a training apparatus for an intent recognition model, a computing device, a computer-readable storage medium, and a computer program product. Background Art
[0002] With the rapid development of artificial intelligence technology, human-computer interaction has also developed rapidly. Among them, human-computer interaction through voice is being applied in various fields. For example, in the field of education, learning platforms can develop different functions and provide users with relevant functions based on the user's intention corresponding to the user's voice command. For example, when a user issues the voice command "I want to practice speaking," the learning platform can use the intent recognition model to identify the user's intention as "speaking practice" and provide the user with functions related to "speaking practice."
[0003] However, in actual applications, due to the differences in users' accents and idioms, the intent recognition model may fail to recognize or misrecognize user intent.
[0004] Based on this, the present application provides a method and device for training an intent recognition model. Summary of the Invention
[0005] In view of this, embodiments of the present application provide a method for training an intent recognition model. This application also relates to a training apparatus for an intent recognition model, a computing device, a computer-readable storage medium, and a computer program product to address the above-mentioned problems in the prior art.
[0006] According to a first aspect of an embodiment of the present application, a method for training an intent recognition model is provided, comprising: Acquiring an initial text, where the initial text is determined by historical voice instructions received by the learning platform; updating the initial text according to the standard user intention to obtain a standard text; Determine an initial user intent corresponding to the initial text, determine a target identification text corresponding to the initial user intent based on the identification text of the function corresponding to the learning platform, and modify the initial user intent based on the target identification text to obtain a standard user intent; An initial intent recognition model is trained based on the standard text and the target identification text until a model training stop condition is reached, thereby obtaining a trained intent recognition model.
[0007] According to a second aspect of an embodiment of the present application, a training device for an intent recognition model is provided, comprising: A data acquisition module is configured to acquire an initial text, where the initial text is determined by historical voice instructions received by the learning platform; a text updating module configured to update the initial text according to the standard user intention to obtain a standard text; an intention modification module configured to determine an initial user intention corresponding to the initial text, determine a target identification text corresponding to the initial user intention based on the identification text of the function corresponding to the learning platform, and modify the initial user intention based on the target identification text to obtain a standard user intention; The model training module is configured to train the initial intent recognition model according to the standard text and the standard user intent until the model training stop condition is reached to obtain a trained intent recognition model.
[0008] According to a third aspect of an embodiment of the present application, a computing device is provided, including: memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer programs / instructions are executed by the processor, the steps of the training method of the above-mentioned intent recognition model are implemented.
[0009] According to a fourth aspect of an embodiment of the present application, a computer-readable storage medium is provided, which stores a computer program / instruction, which, when executed by a processor, implements the steps of the training method of the above-mentioned intent recognition model.
[0010] According to a fifth aspect of an embodiment of the present application, a computer program product is provided, comprising a computer program / instruction, which, when executed by a processor, implements the steps of the above-mentioned training method for the intent recognition model.
[0011] The training method for the intent recognition model provided in this application can obtain the initial text determined by the historical voice instructions received by the learning platform, and can determine the initial user intent corresponding to the initial text. Then, the target identification text can be determined based on the identification text of the function corresponding to the learning platform, and the initial user intent can be modified to obtain the standard user intent. The initial text can be updated based on the standard user intent to obtain the standard text. Finally, the initial intent recognition model can be trained based on the standard text and standard user intent until the model training stop condition is reached, thereby obtaining a trained intent recognition model.
[0012] The training method of the intent recognition model provided in this application can improve the generalization and fault tolerance of the trained intent recognition model. When the user has an accent or makes mistakes, and when there are errors in speech recognition, that is, when the text obtained during speech-to-text conversion is incorrect, the trained intent recognition model can still recognize the user's true intention, thereby improving the practicality of the learning platform in voice interaction scenarios and enhancing the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 This is a flowchart of a method for training an intent recognition model provided in one embodiment of the present application; Figure 2 This is a flowchart of a method for modifying initial user intent provided by an embodiment of the present application; Figure 3 An architectural diagram of a training system for an intent recognition model provided in one embodiment of the present application; Figure 4 This is a flowchart of a method for training an intent correction model provided in one embodiment of the present application; Figure 5 1 is a schematic structural diagram of a training device for an intent recognition model provided in one embodiment of the present application; Figure 6 This is a structural block diagram of a computing device provided in one embodiment of the present application. DETAILED DESCRIPTION
[0014] The following description sets forth many specific details to facilitate a thorough understanding of the present application. However, the present application can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the scope of the present application. Therefore, the present application is not limited to the specific implementations disclosed below.
[0015] The terms used in one or more embodiments of the present application are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of the present application. The singular forms "a", "the" and "the" used in one or more embodiments of the present application and the appended claims are also intended to include plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of the present application refers to and includes any or all possible combinations of one or more associated listed items.
[0016] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of the present application, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of the present application, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0017] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0018] First, the terms involved in one or more embodiments of the present application are explained.
[0019] Large Language Model: A Large Language Model (LLM), also known as a large language model, is an artificial intelligence model designed to understand and generate human language. LLMs are trained on large amounts of text data and can perform a wide range of tasks, including text summarization, translation, and sentiment analysis. LLMs are characterized by their massive scale, containing billions of parameters, which help them learn complex patterns in language data. Large language models are typically based on deep learning architectures, typically using the Transformer as their foundation.
[0020] SFT (Supervised Fine-Tuning of Large Models) refers to a fine-tuning strategy for large models. Large pre-trained models can be fine-tuned on supervised data for a specific task. This strategy improves the model's performance and makes it better suited to the specific task.
[0021] Intent Classification and Slot Filling: Intent classification and slot filling is a core task in semantic understanding, aiming to determine the purpose or intent of a user sentence. Slot extraction is a sequence labeling task aimed at identifying key entities or parameters in a user sentence.
[0022] With the rapid development of artificial intelligence technology, human-computer interaction has also developed rapidly. Among them, human-computer interaction through voice is being applied in various fields. For example, in the field of education, learning platforms can develop different functions and provide users with relevant functions based on the user's intention corresponding to the user's voice command. For example, when a user issues the voice command "I want to practice speaking," the learning platform can use the intent recognition model to identify the user's intention as "speaking practice" and provide the user with functions related to "speaking practice."
[0023] In practical applications, current intent recognition models primarily train text-intent pairs based on rules or templates. However, due to the diverse accents and idioms of users, these rule- or template-based text-intent pairs struggle to capture the diverse expressions of users. This can lead to the intent recognition model failing to recognize or misidentifying the user's true intent, resulting in a poor user experience.
[0024] Based on this, the present application specification provides a training method for an intent recognition model, which can improve the generalization and fault tolerance of the trained intent recognition model. When the user has an accent or makes mistakes, and when there are errors in speech recognition, that is, when the text obtained during speech-to-text conversion is incorrect, the trained intent recognition model can still recognize the user's true intention, thereby improving the practicality of the learning platform in voice interaction scenarios and enhancing the user experience.
[0025] In this application, a training method for an intent recognition model is provided. This application also involves a training device for an intent recognition model, a computing device, a computer-readable storage medium, and a computer program product, which are described in detail one by one in the following embodiments.
[0026] Figure 1 A flowchart of a method for training an intent recognition model according to an embodiment of the present application is shown, which specifically includes the following steps: Step 102: Acquire an initial text, where the initial text is determined by historical voice instructions received by the learning platform.
[0027] In this specification, the execution subject of the training method of the intent recognition model can be any computing device with computing capabilities, such as a server, a terminal, etc. For ease of description, the following description is made using a server as an example.
[0028] In practice, the learning platform is a terminal application that provides learning assistance services to users. It not only provides conventional learning resource management and teaching support functions, but also integrates a voice recognition module to support voice interaction with users. In other words, in this specification, the learning platform can be understood as an intelligent learning terminal that integrates learning tools and voice interaction devices.
[0029] In one or more embodiments of this specification, developers can flexibly develop various functional modules based on user needs to meet diverse learning scenarios. Examples of these functions include, but are not limited to, providing oral practice services, video course viewing services, automatic homework grading services, text translation services, and essay writing and analysis services.
[0030] In this manual, to improve the efficiency of searching for each function and facilitate intent matching, you can configure corresponding identifier text for each function of the learning platform. For example, "Speaking Practice" corresponds to the speaking practice function, "Video Class" corresponds to the video course viewing function, "Homework Correction" corresponds to the automatic homework correction function, "Lookup Translation" corresponds to the text translation function, "Composition Assistant" corresponds to the composition writing and analysis function, and so on.
[0031] Of course, in actual application, other functions can be added according to specific business needs, and their identification texts can be set accordingly.
[0032] It should be noted that the identification text can be generated by extracting key semantic words from the description text of the corresponding function, so as to be used in processes such as speech recognition, intent recognition and function calling.
[0033] In one or more embodiments of this specification, the server may obtain an initial text, which is determined by the learning platform based on historical voice commands received. Specifically, in actual application, the learning platform may store historical voice commands entered by users during the interaction process and convert these historical voice commands into corresponding text form, namely, the "initial text," through its internally integrated voice-to-text conversion module.
[0034] The speech-to-text conversion module may be a speech recognition model deployed locally on the learning platform, or a cloud-based speech recognition capability implemented by calling a remote service interface. This specification does not impose any specific restrictions on this.
[0035] The learning platform can then send the generated initial text to the server for subsequent processing. In subsequent steps, the server can construct a training dataset for training the initial intent recognition model based on the acquired initial text.
[0036] It should be understood that in actual application scenarios, the learning platform may send a large amount of initial text to the server, and the server may construct a training data set for training the initial intent recognition model based on the large amount of initial text.
[0037] In this specification, the server constructs a training data set based on the initial text corresponding to the historical voice commands in the learning platform. This not only enables in-depth analysis and mining of the behavior patterns of users using the learning platform, but also makes the constructed training data set more consistent with the expression methods of users using the learning platform, so that the trained intent recognition model can better recognize the true intentions of users using the learning platform.
[0038] Step 104: Determine the initial user intention corresponding to the initial text, determine the target identification text corresponding to the initial user intention based on the identification text of the function corresponding to the learning platform, and modify the initial user intention based on the target identification text to obtain the standard user intention.
[0039] After obtaining the initial text, the server can determine the initial user intent corresponding to the initial text. In one or more embodiments of the present specification, the server can input the initial text into the initial intent recognition model to obtain the initial user intent corresponding to the initial text. Since the initial user intent corresponding to the initial text needs to be modified based on the identification text of the function corresponding to the learning platform, that is, the initial user intent is an inaccurate user intent, the initial user intent has not undergone the semantic alignment processing of the learning platform function system, and it may have semantic deviations or classification errors, and still requires subsequent standardization processing to improve accuracy. Therefore, although the initial intent recognition model has not been trained, the initial user intent corresponding to the initial text can still be determined based on the initial intent recognition model.
[0040] Of course, the server can also determine the initial user intent corresponding to the initial text through other methods, such as using other pre-trained intent recognition models to determine the user intent corresponding to the initial text, thereby reducing the workload of subsequent modification of the initial user intent. Alternatively, the server can also determine the initial user intent corresponding to the initial text based on manual methods.
[0041] Then, the server can modify the initial user intent according to the identification text of the function corresponding to the learning platform to obtain the standard user intent.
[0042] In actual applications, the learning platform can pre-send the identification text of each supported function to the server. Specifically, the learning platform can pre-send the identification text of each function to the server. As previously mentioned, the function identification text can be generated by extracting key semantic words from the corresponding function description text. This identification text is used to represent the core semantic features of each function within the learning platform. The server can receive the identification text of the function corresponding to the learning platform and store it.
[0043] Furthermore, the server may calculate the text similarity between each identification text and the initial user intention, and determine the target identification text according to each text similarity, so as to modify the initial user intention according to the target identification text and obtain the standard user intention.
[0044] Among them, methods for calculating text similarity include but are not limited to cosine similarity based on word vectors (such as using BERT, Word2Vec, GloVe), scoring mechanisms based on semantic matching models, rule-based keyword matching methods, etc.
[0045] In one or more embodiments of the present specification, the identification text that is most similar to the initial user intent may be selected as the target identification text, so that the initial user intent may be modified based on the target identification text to obtain the standard user intent.
[0046] In one or more embodiments of the present specification, the server may replace the initial user intent with the target identification text to obtain the standard user intent, that is, the server may directly use the target identification text as the standard user intent.
[0047] See also Figure 2 , Figure 2 This is a flowchart of a method for modifying an initial user intention provided in one embodiment of this specification. Specifically, it includes the following steps: Step 202: Identify the intended characters of the initial user intention and identify the standard characters in the target identification text.
[0048] Step 204: Compare the intended character with the standard character to obtain a character comparison result.
[0049] Step 206: If there are abnormal characters in the character comparison result, adjust the intended characters according to the standard characters.
[0050] In one or more embodiments of the present specification, the server may further identify the intended characters of the initial user intent, identify standard characters in the target identification text, and compare the intended characters with the standard characters to obtain a character comparison result. Specifically, the server may compare the intended characters with the standard characters character by character to generate a character comparison result. Thus, if abnormal characters are present in the character comparison result, the intended characters may be adjusted based on the standard characters to obtain the standard user intent.
[0051] The intended character represents each independent Chinese character or word unit in the initial user intention, and the standard character represents each independent Chinese character or word unit in the target identification text.
[0052] It should be noted that how the server identifies the intended characters of the initial user's intention and identifies the standard characters in the target identification text is a relatively mature technology at present. It can be achieved by using Chinese word segmentation technology, character slicing algorithm, BERT-based tokenization, etc. This manual does not make specific restrictions on this.
[0053] In one or more embodiments of the present specification, when comparing intended characters with standard characters to obtain a character comparison result, the server may first determine the identical characters between the intended characters and the standard characters, and then align the initial user intent with the standard intent based on the identical characters. Alternatively, the server may use the identical characters as the starting character comparison point, thereby aligning the initial user intent with the standard intent. After the alignment is complete, the server may replace or correct the abnormal characters in the initial user intent based on the standard characters in the target identification text, thereby obtaining the final standard user intent.
[0054] For example, the initial text may be: I want to speak English for a while. Then the initial user intention may be: speak English for a while. Assuming that the target identification text is: English speaking training, the intention characters in the initial user intention may include 10, which are: speak English for a while. The standard characters in the target identification text may include 6, which are: speak English for a while. Therefore, it can be determined that the same characters include English, English, and spoken English. Then, you can randomly select an identical character as the alignment character to align the initial user intention and the identification user intention. You can also select the same character as the alignment character based on the preset rule to align the initial user intention and the identification user intention. The preset rule can be: After the initial user intention and the identification user intention are aligned using the alignment character, the initial user intention and the identification user intention are aligned. Figure 1 The number of corresponding characters is greater than a preset threshold to ensure semantic consistency of the modified intent. For example, "English / Spoken / Spoken" can be selected as the starting character for comparison.
[0055] In this specification, in order to preserve the user's expression and style, abnormal characters may include spelling errors, semantic inconsistencies, and missing characters, but not redundant characters.
[0056] Furthermore, based on the character alignment results, the server can replace abnormal characters with standard characters to generate the final standard user intent. Continuing with the above example, if the abnormal characters are due to lack of training, then "perform spoken English for a while" can be modified to: "Perform spoken English training for a while".
[0057] Based on this approach, rather than using the target identifier text directly as the standard user intent, character-level comparison and correction are used to standardize the intent while preserving the user's expressive characteristics. This allows for the construction of a training dataset, training the initial intent recognition model, and improving the generalization of the trained intent recognition model. Furthermore, using the identifier text of the learning platform's corresponding function as a benchmark for modifying the user's initial intent ensures that the user intent generated by the subsequently trained intent recognition model aligns with the learning platform's functions, avoiding the generation of user intent unrelated to the learning platform's functions. This allows the learning platform to provide users with relevant functional services.
[0058] Step 106: Update the initial text according to the standard user intention to obtain a standard text.
[0059] Step 108: Train an initial intent recognition model based on the standard text and the standard user intent until a first model training stop condition is reached to obtain a trained intent recognition model.
[0060] Therefore, in this specification, the server can update the initial text based on the standard user intent to obtain the standard text. Since the initial text corresponds to the user's voice command, that is, the initial text is generated based on the user's voice, it is not directly modified. Only updating the initial text based on the standard user intent can maximize the preservation of the user's expression style and habits. This allows the intent recognition model to learn different expression styles and habits, and can still recognize the user's true intention in the face of diverse user expressions.
[0061] In one or more embodiments of this specification, a target text corresponding to an initial user intent in the initial text can be determined, and the target text can be replaced with the standard user intent to obtain a standard text. Using the above example, if the initial text is "I want to practice speaking English for a while," the initial user intent can be "Practice speaking English for a while," and the target text is "Practice speaking English for a while," and the standard user intent is "Practice speaking English for a while," the standard text is "I want to practice speaking English for a while."
[0062] It should be understood that the target text corresponding to the initial user intention refers to the text in the initial text from which the initial user intention originates, which may be a partial text in the initial text or the initial text itself.
[0063] In this specification, standard text can be used as samples and target identification text as annotations to train the initial intent recognition model. Specifically, the standard text can be input into the initial intent recognition model to obtain a predicted intent. Based on the predicted intent and target identification text, a loss value is determined. Based on this loss value, the initial intent recognition model is trained until the first model training stop condition is met, resulting in a trained intent recognition model.
[0064] In one or more embodiments of this specification, there are many methods for calculating loss values, such as cross entropy loss function, maximum loss function, average loss function, etc. In this specification, the specific method of the loss function is not limited and is subject to actual application.
[0065] In one or more embodiments of the present specification, the model training stopping conditions include but are not limited to the determined loss being less than a preset loss threshold, the number of iterations reaching a preset number, the number of standard text-target identification text pairs used reaching a preset number, and the like.
[0066] The standard text-target identifier text pair constructed using this method preserves the user's expression habits and style, and can correct text errors caused by the user's accent, which can lead to intent recognition errors. Furthermore, when speech-to-text conversion errors occur, the text can be corrected, and the identifier text based on the function in the learning platform can be corrected. This improves the accuracy of the trained intent recognition model while preventing the intent recognition model from generating intents unrelated to the learning platform's functions.
[0067] The training method based on the above-mentioned intent recognition model can improve the generalization and fault tolerance of the trained intent recognition model. When the user has an accent or makes a slip of the tongue, or when there are errors in speech recognition, that is, when the text obtained during speech-to-text conversion is incorrect, the trained intent recognition model can still recognize the user's true intention, thereby improving the practicality of the learning platform in voice interaction scenarios and enhancing the user experience.
[0068] Furthermore, in one embodiment of the present specification, in actual applications, after the intent recognition model is trained, the trained intent recognition model can be deployed in a learning platform. When the learning platform receives a user's voice command, the voice command can be converted into text first, and then the text can be input into the trained intent recognition model to obtain the user intent output by the trained intent recognition model, so that the learning platform can enable the function corresponding to the user intent based on the user intent and the identification text of each function.
[0069] In one or more embodiments of the present specification, in order to further improve the accuracy of the user intent generated by the trained intent recognition model, prompt information may also be set.
[0070] Specifically, the learning platform can obtain text to be recognized, where the text to be recognized is determined by the voice command to be recognized currently received by the learning platform. It can also obtain prompt information, where the prompt information at least includes the identification text of the corresponding function of the learning platform. The text to be recognized and the prompt information can then be input into a trained intent recognition model to obtain the user intent corresponding to the text to be recognized.
[0071] Based on the above method, by including the identification text of the function corresponding to the learning platform as prompt information, it is possible to avoid the intent recognition model from generating intents that are unrelated to the services provided by the learning platform.
[0072] In addition, this manual also provides an architecture diagram of the training system for the intent recognition model, see Figure 3 , Figure 3 This is an architecture diagram of a training system for an intent recognition model provided in one embodiment of this specification. The training system for an intent recognition model may include a client 100 and a server 200; The client 100 is configured to send a training task for an intent recognition model to the server 200, wherein the training task includes an initial text and an identification text of a function corresponding to the learning platform, wherein the initial text is determined by historical voice commands received by the learning platform; The server 200 is configured to receive a training task, determine an initial user intent corresponding to the initial text, determine a target identification text corresponding to the initial user intent based on the identification text of the function corresponding to the learning platform, and modify the initial user intent based on the target identification text to obtain a standard user intent; update the initial text based on the standard user intent to obtain a standard text; train an initial intent recognition model based on the standard text and the target identification text until a first model training stop condition is reached, thereby obtaining a trained intent recognition model and model parameters of the trained intent recognition model; and send the model parameters to the client 100; The client 100 is also used to receive model parameters sent by the server 200.
[0073] Apply the solution of the embodiment of this specification to obtain an initial text, which is determined by historical voice instructions received by a learning platform; determine the initial user intention corresponding to the initial text, determine the target identification text corresponding to the initial user intention based on the identification text of the function corresponding to the learning platform, and modify the initial user intention based on the target identification text to obtain a standard user intention; update the initial text according to the standard user intention to obtain a standard text; train an initial intention recognition model based on the standard text and the target identification text until the first model training stop condition is reached, and obtain a trained intention recognition model.
[0074] This preserves the user's expression habits and style, and corrects text errors caused by the user's accent, which can lead to intent recognition errors. Furthermore, when speech-to-text errors occur, the text can be corrected, and the corrections can be made based on the text identifying the functions in the learning platform. This improves the accuracy of the trained intent recognition model while preventing the intent recognition model from generating intents unrelated to the learning platform's functions.
[0075] The training method based on the above-mentioned intent recognition model can improve the generalization and fault tolerance of the trained intent recognition model. When the user has an accent or makes a slip of the tongue, or when there are errors in speech recognition, that is, when the text obtained during speech-to-text conversion is incorrect, the trained intent recognition model can still recognize the user's true intention, thereby improving the practicality of the learning platform in voice interaction scenarios and enhancing the user experience.
[0076] The image processing system may include multiple clients 100 and a server 200. The clients 100 may be referred to as end-side devices, and the server 200 may be referred to as cloud-side devices. The multiple clients 100 may establish a communication connection through the server 200. In the training scenario of an intent recognition model, the server 200 is used to provide intent recognition model training task services between the multiple clients 100. The multiple clients 100 may act as either senders or receivers, communicating through the server 200.
[0077] Users can interact with the server 200 through the client 100 to receive data sent by other clients 100, or send data to other clients 100, etc. In the training scenario of the intent recognition model, the user can publish a data stream to the server 200 through the client 100. The server 200 generates detection results based on the data stream and pushes the model parameters to other clients that have established communication.
[0078] The client 100 and the server 200 are connected via a network. The network provides a medium for the communication link between the client 100 and the server 200. The network can include various connection types, such as wired or wireless communication links or fiber optic cables. The data transmitted by the client 100 may need to be encoded, transcoded, compressed, or other processing before being released to the server 200.
[0079] The client 100 can be a browser, an application (APP), a web application such as an H5 (HyperText Markup Language 5) application, a lightweight application (also known as a mini-program, a type of lightweight application), or a cloud application. The client 100 can be developed based on the software development kit (SDK) of the corresponding service provided by the server 200, such as a real-time communication (RTC) SDK. The client 100 can be deployed on a computing device and rely on the device or certain applications on the device to run. For example, the computing device can have a display screen and support information browsing, such as a personal mobile terminal such as a mobile phone, tablet computer, or personal computer. Various other types of applications can also be configured on the computing device, such as human-computer interaction applications, model training applications, text processing applications, web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.
[0080] The server 200 may include servers that provide various services, such as servers that provide communication services to multiple clients, servers that support backend training for models used on clients, and servers that process data sent by clients. It should be noted that the server 200 can be implemented as a distributed server cluster consisting of multiple servers or as a single server. The server can also be a server in a distributed system or a server integrated with blockchain. The server can also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), big data and artificial intelligence platforms, or intelligent cloud computing servers or intelligent cloud hosts equipped with artificial intelligence technology.
[0081] It is worth noting that the training method for the intent recognition model provided in the embodiments of this specification is generally performed by the server. However, in other embodiments of this specification, the client may also have similar functions to the server, thereby performing the training method for the intent recognition model provided in the embodiments of this specification. In other embodiments, the training method for the intent recognition model provided in the embodiments of this specification may also be performed jointly by the client and the server.
[0082] In addition, in one embodiment of this specification, a method for training an intention correction model is also provided. Figure 4 , Figure 4 This is a flowchart of a method for training an intent correction model provided in one embodiment of this specification, specifically including the following steps: Step 402: Obtain sample user intent and a label corresponding to the sample user intent, wherein the label corresponding to the sample user intent is determined based on an identification text of a function corresponding to the learning platform.
[0083] Step 404: Input the sample user intention into the initial intention correction model to obtain the predicted correction intention output by the initial intention correction model.
[0084] Step 406: Based on the predicted correction intent and the labels corresponding to the sample user intent, the initial intent correction model is trained until a model training stop condition is reached, thereby obtaining a trained intent correction model.
[0085] In this specification, the server may obtain a sample user intent and a label corresponding to the sample user intent, wherein the label corresponding to the sample user intent is determined based on the identification text of the function corresponding to the learning platform.
[0086] In one or more embodiments of this specification, sample text may be obtained, which may be text corresponding to a voice command issued by a user using the learning platform. Sample user intent corresponding to the sample text may also be determined. The specific method for determining the sample user intent from the sample text is not detailed here; the specific process may be consistent with the method used in the aforementioned intent recognition model training method.
[0087] Then, the label corresponding to the sample user intent can be determined. Specifically, the similarity between each identifier text and the sample user intent can be calculated, and the identifier text with the greatest similarity can be used as the label corresponding to the sample user intent. Of course, manual methods can also be used to determine the label of the sample user intent based on the identifier text corresponding to the function of the learning platform.
[0088] In one or more embodiments of this specification, the server may input a sample user intent into an initial intent correction model to obtain a predicted correction intent output by the initial intent correction model. The initial intent correction model is trained based on the labels corresponding to the predicted correction intent and the sample user intent until a training stop condition is met, thereby obtaining a trained intent correction model.
[0089] Specifically, the server can determine the loss value based on the labels corresponding to the predicted correction intent and the sample user intention, and then train the initial intent correction model based on the loss value until the model training stop condition is reached to obtain a trained intent correction model.
[0090] It should be noted that the loss value and model training stopping condition in the process of the training method of the intention correction model are the loss value and model training stopping condition corresponding to the initial intention correction model.
[0091] In one or more embodiments of this specification, there are many methods for calculating the loss value corresponding to the initial intent correction model, such as the cross entropy loss function, the maximum loss function, the average loss function, etc. In this specification, the specific method of the loss function is not limited and is subject to actual application.
[0092] In one or more embodiments of the present specification, the model training stopping conditions corresponding to the initial intent correction model include but are not limited to the determined loss being less than a preset loss threshold, the number of iterations reaching a preset number, the number of sample text-sample text corresponding annotation pairs used, that is, the number of training data reaching a preset number, etc.
[0093] In one or more embodiments of this specification, the trained intent modification model can be used to modify the initial user intent. Specifically, when modifying the initial user intent based on the target identifier text in step 104, the initial user intent can be modified based on the intent modification model trained based on the identifier text of the function corresponding to the learning platform.
[0094] Specifically, the initial user intent may be input into a trained intent correction model to obtain a predicted intent output by the trained intent correction model, which is the standard user intent.
[0095] The training method based on the above-mentioned intention correction model can train the user intention corresponding to the correction learning platform, thereby improving the modification efficiency and accuracy.
[0096] Corresponding to the above method embodiment, the present application also provides an embodiment of a training device for an intent recognition model. Figure 5 FIG. 1 shows a schematic diagram of a training device for an intention recognition model provided by an embodiment of the present application. Figure 5 As shown, the device includes: The data acquisition module 502 is configured to acquire an initial text, where the initial text is determined by historical voice instructions received by the learning platform; The intent modification module 504 is configured to determine an initial user intent corresponding to the initial text, determine a target identification text corresponding to the initial user intent based on the identification text of the function corresponding to the learning platform, and modify the initial user intent based on the target identification text to obtain a standard user intent; A text updating module 506 is configured to update the initial text according to the standard user intention to obtain a standard text; The model training module 508 is configured to train the initial intent recognition model according to the standard text and the target identification text until the model training stop condition is reached, thereby obtaining a trained intent recognition model.
[0097] Optionally, the intention modification module 504 is further configured to input the initial text into the initial intention recognition model to obtain an initial user intention corresponding to the initial text.
[0098] Optionally, the intention modification module 504 is further configured to calculate text similarities between each identification text and the initial user intention; and determine the target identification text according to each text similarity.
[0099] Optionally, the intention modification module 504 is further configured to replace the initial user intention with the target identification text.
[0100] Optionally, the intention modification module 504 is further configured to identify the intended characters of the initial user intention and identify the standard characters in the target identification text; compare the intended characters with the standard characters to obtain character comparison results; and if there are abnormal characters in the character comparison results, adjust the intended characters according to the standard characters.
[0101] Optionally, the model training module 508 is further configured to input the initial standard text into the initial intent recognition model to obtain a predicted intent; determine a loss value based on the predicted intent and the target identification text; and train the initial intent recognition model based on the loss value.
[0102] Optionally, the device also includes a model application module, which is configured to obtain text to be recognized, wherein the text to be recognized is determined by the voice instruction to be recognized currently received by the learning platform, and obtain prompt information, wherein the prompt information at least includes the identification text of the function corresponding to the learning platform, and input the text to be recognized and the prompt information into the trained intention recognition model to obtain the user intention corresponding to the text to be recognized.
[0103] The training device for the intent recognition model provided in this specification can improve the generalization and fault tolerance of the trained intent recognition model. When the user has an accent or makes mistakes, or when there are errors in speech recognition, that is, when the text obtained during speech-to-text conversion is incorrect, the trained intent recognition model can still recognize the user's true intention, thereby improving the practicality of the learning platform in voice interaction scenarios and enhancing the user experience.
[0104] The above is a schematic diagram of a training device for an intent recognition model according to this embodiment. It should be noted that the technical solution of this training device for an intent recognition model is based on the same concept as the technical solution of the training method for an intent recognition model described above. For details not described in detail in the technical solution for the training device for an intent recognition model, please refer to the description of the technical solution for the training method for an intent recognition model described above.
[0105] Figure 6 6 shows a block diagram of a computing device 600 according to an embodiment of the present application. Components of the computing device 600 include, but are not limited to, a memory 610 and a processor 620. The processor 620 is connected to the memory 610 via a bus 630, and a database 650 is used to store data.
[0106] The computing device 600 also includes an access device 640 that enables the computing device 600 to communicate via one or more networks 660. Examples of such networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 640 may include one or more of any type of network interface (e.g., a network interface controller (NIC)) whether wired or wireless, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, a near field communication (NFC) interface, and the like.
[0107] In one embodiment of the present application, the above components of the computing device 600 and Figure 6 Other components not shown in the figure may also be connected to each other, for example, via a bus. Figure 6 The computing device structure block diagram shown is for illustrative purposes only and is not intended to limit the scope of the present application. Those skilled in the art may add or replace other components as needed.
[0108] Computing device 600 can be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, personal digital assistant, laptop computer, notebook computer, netbook computer, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or personal computer (PC). Computing device 600 can also be a mobile or stationary server.
[0109] Among them, the processor 620 is used to execute the following computer program / instructions, which, when executed by the processor, implement the steps of the training method of the above-mentioned intent recognition model.
[0110] The above is a schematic diagram of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solution of the aforementioned intent recognition model training method are based on the same concept. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solution of the aforementioned intent recognition model training method.
[0111] An embodiment of this specification also provides a computer-readable storage medium storing a computer program / instruction, which, when executed by a processor, implements the steps of the above-mentioned training method for the intent recognition model.
[0112] The above is a schematic diagram of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium and the technical solution of the training method for the intent recognition model described above are based on the same concept. For details not described in detail in the technical solution of the storage medium, please refer to the description of the technical solution of the training method for the intent recognition model described above.
[0113] An embodiment of this specification also provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the above-mentioned training method for the intent recognition model.
[0114] The above is a schematic diagram of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product and the technical solution of the aforementioned intent recognition model training method are based on the same concept. For details not described in detail in the technical solution of the computer program product, please refer to the description of the technical solution of the aforementioned intent recognition model training method.
[0115] The foregoing description describes specific embodiments of the present application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0116] The computer instructions include computer program code, which may be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium may include any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal, and software distribution medium. It should be noted that the content of the computer-readable medium may be appropriately increased or decreased based on the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media does not include electric carrier signals and telecommunication signals.
[0117] It should be noted that for the aforementioned method embodiments, for ease of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.
[0118] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0119] The preferred embodiments of the present application disclosed above are intended only to help illustrate the present application. The optional embodiments do not describe all details in detail, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made based on the content of this application. This application selects and describes these embodiments in detail in order to better explain the principles and practical applications of this application, so that those skilled in the art can better understand and utilize this application. This application is limited only by the claims and their full scope and equivalents.
Claims
1. A training method for an intent recognition model, characterized in that: include: Acquiring an initial text, where the initial text is determined by historical voice instructions received by the learning platform; Determine an initial user intent corresponding to the initial text, determine a target identification text corresponding to the initial user intent based on the identification text of the function corresponding to the learning platform, and modify the initial user intent based on the target identification text to obtain a standard user intent; updating the initial text according to the standard user intention to obtain a standard text; An initial intent recognition model is trained based on the standard text and the target identification text until a model training stop condition is reached, thereby obtaining a trained intent recognition model.
2. The method according to claim 1, wherein Determining an initial user intent corresponding to the initial text includes: The initial text is input into the initial intent recognition model to obtain the initial user intent corresponding to the initial text.
3. The method according to claim 1, wherein Determining the target identification text corresponding to the initial user intention based on the identification text of the function corresponding to the learning platform includes: Calculating the text similarity between each identified text and the initial user intention; The target identification text is determined based on the similarity of each text.
4. The method according to claim 1, wherein Modifying the initial user intention based on the target identification text includes: The initial user intention is replaced with the target identification text.
5. The method according to claim 1, wherein Modifying the initial user intention based on the target identification text includes: Identifying the intended characters of the initial user intention and identifying the standard characters in the target identification text; Comparing the intended character with the standard character to obtain a character comparison result; In the case that there are abnormal characters in the character comparison result, the intended characters are adjusted according to the standard characters.
6. The method according to claim 1, wherein Training an initial intent recognition model based on the standard text and the target identification text includes: Inputting the standard text into the initial intent recognition model to obtain a predicted intent; Determining a loss value based on the predicted intent and the target identification text; The initial intent recognition model is trained according to the loss value.
7. The method according to claim 1, wherein The method further comprises: Acquire a text to be recognized, wherein the text to be recognized is determined by a voice instruction to be recognized currently received by the learning platform, and acquire prompt information, wherein the prompt information at least includes an identification text of a function corresponding to the learning platform, The text to be recognized and the prompt information are input into the trained intention recognition model to obtain the user intention corresponding to the text to be recognized.
8. A training device for an intent recognition model, characterized in that: include: A data acquisition module is configured to acquire an initial text, where the initial text is determined by historical voice instructions received by the learning platform; an intention modification module configured to determine an initial user intention corresponding to the initial text, determine a target identification text corresponding to the initial user intention based on the identification text of the function corresponding to the learning platform, and modify the initial user intention based on the target identification text to obtain a standard user intention; a text updating module configured to update the initial text according to the standard user intention to obtain a standard text; The model training module is configured to train the initial intent recognition model based on the standard text and the target identification text until the model training stopping condition is reached to obtain a trained intent recognition model.
9. A computing device, characterized in that include: memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer program / instructions are executed by the processor, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing a computer program / instruction, characterized in that: When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
11. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Voice conversation method and device, and computer readable storage medium
CN111986675A
Intention recognition method and device, electronic equipment and computer readable storage medium
CN112256845A
Voice intention recognition method and device, computer equipment and storage medium
CN112699213A