Intelligent dialogue method, system and device and storage medium
By introducing a generative model into the intelligent dialogue system, using the machine learning model trained by instruction learning to correct and expand the speech recognition text, and combining user historical data to personalize the answers, the shortcomings of the existing intelligent dialogue system in adapting to multi-field problems and personalized changes are solved, and higher adaptability and semantic understanding capabilities are achieved, and user experience and vehicle convenience and safety are improved.
Patent Information
- Application Number
- CN202311650108.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-04
- Publication Date
- 2025-06-06
AI Technical Summary
The existing intelligent dialogue system has shortcomings in adapting to various dialogue scenarios, solving multi-field problems, improving understanding and adapting to user personalized changes. Especially when dealing with problems that exceed the scope of inventory content, there are problems such as huge data and difficulty in managing and error accumulation.
By introducing a generative model, the machine learning model obtained by instruction learning training is used to correct and expand the processed speech recognition text, and personalize the answers based on the user's historical data.
The adaptability and semantic understanding capabilities of the intelligent dialogue system are improved, so that it can more accurately adapt to the on-board environment and users' personalized needs, improve users' driving and riding experience, and increase the convenience and safety of the vehicle.
Smart Images

Figure CN120108398A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of smart cars, and in particular to an intelligent dialogue method, system, device and storage medium. Background Art
[0002] In recent years, with the rapid development of artificial intelligence, intelligent services have been widely used in all walks of life in society. Among them, automotive intelligence includes intelligent dialogue systems. Intelligent dialogue systems can be used for navigation and route planning, entertainment control, vehicle function control, intelligent vehicle health management, and many other aspects. In the face of the continuous improvement of user needs, how intelligent dialogue systems can adapt to various dialogue scenarios, solve multi-field problems, improve understanding, and adapt to users' personalized changes has become an important issue in the fields of intelligent dialogue, intelligent vehicle-mounted, etc.
[0003] Therefore, it is necessary to provide an intelligent dialogue method, system, device and storage medium. Summary of the invention
[0004] One or more embodiments of the present specification provide an intelligent dialogue method, the method comprising: determining a processed text based on a user's speech recognition text; determining a generation result through a generation model based on the processed text, wherein the generation model is a machine learning model obtained through instruction learning training, and the generation model performs text correction on the processed text.
[0005] One or more embodiments of the present specification provide an intelligent dialogue system, the system comprising: a text determination module is configured to determine a processed text based on a user's speech recognition text; a result determination module is configured to determine a generated result through a generation model based on the processed text, wherein the generation model is a machine learning model obtained through instruction learning training, and the generation model performs text correction on the processed text.
[0006] One or more embodiments of the present specification provide an intelligent dialogue device, which includes at least one memory and at least one processor, wherein the at least one memory is used to store computer instructions, and the at least one processor executes the computer instructions or part of the instructions to implement an intelligent dialogue method.
[0007] One or more embodiments of the present specification provide a computer-readable storage medium, wherein the storage medium stores computer instructions. When a computer reads the computer instructions, the computer executes an intelligent dialogue method. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] This specification will be further described in the form of exemplary embodiments, which will be described in detail by the accompanying drawings. These embodiments are not restrictive, and in these embodiments, the same number represents the same structure, wherein:
[0009] Figure 1 is a schematic diagram of an application scenario of an intelligent dialogue system according to some embodiments of this specification;
[0010] Figure 2 is an exemplary flow chart of an intelligent dialogue method according to some embodiments of this specification;
[0011] Figure 3 is an exemplary schematic diagram of a generation model according to some embodiments of this specification;
[0012] Figure 4 is an exemplary schematic diagram of a training process of a generative model according to some embodiments of this specification;
[0013] Figure 5 It is an exemplary overall flow chart of the intelligent dialogue method shown in some embodiments of this specification. DETAILED DESCRIPTION
[0014] In order to more clearly illustrate the technical solutions of the embodiments of this specification, the following is a brief introduction to the drawings required for the description of the embodiments. Obviously, the drawings described below are only some examples or embodiments of this specification. For ordinary technicians in this field, this specification can also be applied to other similar scenarios based on these drawings without creative work. Unless it is obvious from the language environment or otherwise explained, the same reference numerals in the figures represent the same structure or operation.
[0015] It should be understood that the "system", "device", "unit" and / or "module" used herein are a method for distinguishing different components, elements, parts, portions or assemblies at different levels. However, if other words can achieve the same purpose, the words can be replaced by other expressions.
[0016] Unless the context clearly indicates an exception, the words "a", "an", "an" and / or "the" do not refer to the singular and may also include the plural. Generally speaking, the terms "include" and "comprise" only indicate the inclusion of the steps and elements that have been clearly identified, and these steps and elements do not constitute an exclusive list. The method or device may also include other steps or elements.
[0017] Flowcharts are used in this specification to illustrate the operations performed by the system according to the embodiments of this specification. It should be understood that the preceding or following operations are not necessarily performed precisely in order. Instead, the steps may be processed in reverse order or simultaneously. At the same time, other operations may also be added to these processes, or one or more operations may be removed from these processes.
[0018] In terms of intelligent dialogue system technology, there are mainly two methods: retrieval-based and guided. A retrieval-based dialogue system refers to a dialogue system that generates responses by matching user input with predefined stock answers. In application scenarios where specific questions and answers are known (for example, customer support, FAQs, reservation services, etc.), retrieval-based dialogue systems can provide efficient automated responses. A guided dialogue system refers to a dialogue system that can conduct free dialogue and provide responses based on the dialogue context. In application scenarios that require free dialogue, solving multi-domain problems, and adapting to users' changing needs (for example, virtual assistants, chatbots, customer support, and multi-domain Q&A, etc.), guided dialogue systems can interact with users more flexibly.
[0019] However, there are still some problems with the above methods. For the retrieval-based dialogue system, it is limited by the inventory answers and cannot handle questions beyond the scope of the inventory content, so there is a problem of limited flexibility. For the guided dialogue system, due to its complex system, it will lead to huge data that is difficult to manage and there will be problems such as error accumulation.
[0020] The above two methods do not have sufficient speech understanding capabilities. When the dialogue exceeds the scope of the specified template, the dialogue system will fail. And for the same questions from different users (such as car owners, etc.), the dialogue system can only provide the same answer. Furthermore, there is no fault tolerance for errors in text conversion caused by noise in the car, which affects the normal operation of the downstream modules of the dialogue system.
[0021] In view of this, some embodiments of the present invention provide an intelligent dialogue method, system, device and storage medium, which use the generative model in intelligent dialogue scenarios (such as in-vehicle dialogue scenarios), and enable the model to adapt to various scenarios and colloquial questions through instruction learning, thereby further improving the model's understanding ability. Introducing a user adaptive processing unit in the intelligent dialogue system can encode the user's historical data so that the generative model can answer questions according to the user's preferences. In the pre-training stage of the generative model, the task of correcting the speech recognition text (ASR) to the correct text is added, so that the generative model can understand the distribution of the speech recognition text well, inject speech recognition text knowledge, and enable the generative model to have better interaction capabilities on the speech recognition text. As a result, the intelligent dialogue system can accurately adapt to various intelligent dialogue scenarios (such as in-vehicle dialogue scenarios, etc.) and have higher semantic understanding capabilities, thereby meeting the personalized needs of users, improving the user's driving and riding experience, and increasing the convenience and safety of the vehicle.
[0022] Figure 1 1 is a schematic diagram of an application scenario of an intelligent dialogue system according to some embodiments of this specification. In some embodiments, an application scenario 100 of an intelligent dialogue system may include a processor 110, a network 120, an intelligent terminal 130, a storage device 140, and a vehicle 150.
[0023] The processor 110 may be used to process data related to the application scenario 100 of the intelligent dialogue system. For example, the processor 110 may process the user's speech recognition text by executing the intelligent dialogue method disclosed in this specification to give a personalized answer. Exemplarily, the processor 110 may determine the processed text based on the user's speech recognition text, and determine the generated result based on the processed text by generating a model.
[0024] In some embodiments, the processor 110 may be a single server or a server group. The server group may be centralized or distributed. In some embodiments, the processor 110 may be local or remote. In some embodiments, the processor 110 may be implemented on a cloud platform. As an example only, the cloud platform may include a private cloud, a public cloud, a hybrid cloud, a community cloud, a distributed cloud, an internal cloud, a multi-layer cloud, etc., or any combination thereof.
[0025] The network 120 may include any suitable network that can facilitate information and / or data exchange for the application scenario 100 of the intelligent dialogue system. The network 120 enables communication between the components and with other components outside the system to facilitate the exchange of data and / or information. For example, the speech recognition text stored in the storage device 140 can be transmitted to the processor 110 via the network 120 for processing. For another example, the processor 110 can transmit the generated results to the storage device 140 via the network 120.
[0026] In some embodiments, the network 120 may be any one or more of a wired network or a wireless network. For example, the network 120 may include a cable network, an optical fiber network, the Internet, a wireless local area network (WLAN), a Bluetooth network, a ZigBee network, a near field communication (NFC), a line within a device, a cable connection, etc., or any combination thereof. The network connection between the various parts may be in one of the above-mentioned ways, or in multiple ways.
[0027] The smart terminal 130 refers to the front-end device of the vehicle's monitoring, management and other systems. In some embodiments, the processor 110 may be a part of the smart terminal 130. For example, the smart terminal 130 may include a microcontroller unit (MCU), etc. In some embodiments, the smart terminal 130 may be a part of the vehicle 150. The vehicle 150 refers to a vehicle with an intelligent voice function. For example, a smart car, a smart bicycle, etc. The user can interact with the application scenario 100 of the intelligent dialogue system via the smart terminal 130. The user refers to an individual or a group related to the intelligent dialogue, etc. For example, the user is one or more of the passengers and / or drivers who have the need to use the intelligent dialogue.
[0028] The storage device 140 may store data, instructions, and / or any other information. In some embodiments, the storage device 140 may store data and / or information obtained from at least one component of the application scenario 100 of the intelligent dialogue system or an external data source. In some embodiments, the storage device may also store data and / or instructions related to the intelligent dialogue system. For example, the storage device 140 may store computer instructions for conducting intelligent dialogues. For another example, the storage device 140 may store speech recognition text, processed text, historical data, etc. In some embodiments, the storage device 140 may be connected to the network 120 to communicate with the processor 110.
[0029] In some embodiments, the storage device 140 may be a storage device inside or outside the intelligent dialogue system. In some embodiments, the storage device may be a part of the processor 110. In some embodiments, the storage device 140 may include a mass storage device, a removable storage device, a volatile read-write memory, a read-only memory (ROM), etc., or any combination thereof.
[0030] For more details about the speech recognition text, processed text, historical data, etc. mentioned above, see Figure 2-Figure 5 and its related description.
[0031] In some embodiments, the intelligent dialogue system may include a text determination module, a result determination module, and a training module. In some embodiments, the text determination module, the result determination module, and the training module may be partial components of the processor 110 .
[0032] In some embodiments, the determine text module may be configured to determine processed text based on the user's speech recognition text.
[0033] In some embodiments, the determination module can be configured to determine the generation result based on the processed text through a generation model, wherein the generation model is a machine learning model obtained through instruction learning training, and the generation model performs text correction on the processed text.
[0034] In some embodiments, the input for generating the model may also include historical data of the user.
[0035] In some embodiments, the generation model includes an adaptive layer, an encoding layer, and a classification layer, and the result determination module can be further configured to determine user feature data through the adaptive layer based on historical data; determine text feature data through the encoding layer based on processed text; and determine the generation result through the classification layer based on user feature data and text feature data.
[0036] In some embodiments, the result determination module can be further configured to determine the corrected text based on the processed text through the generation model, where the corrected text is the correct text of the processed text; determine the expanded text based on the corrected text through the generation model, where the expanded text is at least one similar text of the revised text; and determine the generated result through the generation model based on the expanded text and historical data.
[0037] In some embodiments, the training module can be configured to obtain at least one first training sample, wherein each of the at least one first training sample includes a sample processed text and a label of a sample corrected text representing the sample processed text; based on the at least one first training sample, fine-tune the initial generation model to determine a pre-training model; and determine the generation model based on the pre-training model.
[0038] In some embodiments, the training module can be further configured to determine at least one second training sample based on the target instruction set, wherein the at least one second training sample includes an instruction in the target instruction set; and fine-tune the pre-trained model to determine a generated model based on a label corresponding to an instruction in the target instruction set and the at least one second training sample.
[0039] For more information about the intelligent dialogue system (text module, result determination module, and training module), see Figure 2-Figure 5 and its related description.
[0040] It should be noted that the above description of the application scenarios and components of the intelligent dialogue system is only for the convenience of description and does not limit this specification to the scope of the embodiments. It is understandable that for those skilled in the art, after understanding the principle of the system, it is possible to arbitrarily combine the components or form a subsystem connected with other components without deviating from this principle.
[0041] Figure 2 is an exemplary flow chart of the intelligent dialogue method according to some embodiments of this specification. Figure 2 As shown, process 200 may include the following steps. In some embodiments, Figure 2 The process 200 shown may be called and / or executed by the processor 110 .
[0042] Step 210, based on the user's speech recognition text, determine the processed text. In some embodiments, step 210 can be performed by a text determination module.
[0043] The speech recognition text refers to the text content of the user's speech. In some embodiments, the processor 110 can obtain the speech recognition text in a variety of ways, for example, using Automatic Speech Recognition (ASR) technology or a machine learning model (such as an ASR model) to convert the user's speech into corresponding speech recognition text. The user's speech can be converted into a corresponding speech recognition text in the smart terminal 130 or in the intelligent dialogue system. The processor 110 can obtain the recorded user speech or the converted speech recognition text of the smart terminal 130 through the network 120. For more information about the user, see Figure 1 Related description.
[0044] The processed text refers to the processed speech recognition text. In some embodiments, the processor can obtain the processed text by preprocessing the speech recognition text. The preprocessing can include filtering out invalid data such as empty data, too short data, repeated data, other characters, etc. in the text data.
[0045] Step 220, based on the processed text, determine the generation result through the generation model. In some embodiments, step 220 can be performed by a determination result module.
[0046] The generated result refers to the answer result made to the user question. For example, if the user question is: "How is the weather today?", the corresponding generated result may be: "Today's weather is sunny and cloudy". In some embodiments, the processor may determine the generated result in a variety of ways based on the processed text. For example, the processor may generate the generated result based on the processed text by generating a model output.
[0047] The generation model refers to a model used to determine the generation result. In some embodiments, the generation model can be a machine learning model obtained through instruction learning training, and the generation model can perform text correction on the processed text. For more information about obtaining a model through instruction learning training, please refer to Figure 4 Related description.
[0048] Text correction refers to the process of correcting the processed text to obtain the corrected text. As shown in Table 2, text correction can refer to the process of correcting: "Navigate to Beijin Capital International Airport." to "Navigate to Beijing Capital International Airport." For more information about text correction, see Figure 4 and its related description.
[0049] In some embodiments, the generative model may be a language model, such as a large language model.
[0050] A large language model (LLM) refers to a machine learning model formed by training based on deep learning technology, large-scale data, and computing resources. Large language models are mainly aimed at natural language processing, but can also evolve to process other forms of data. For example, a large language model can include any one or combination of a large language model based on a pre-training and fine-tuning mode (e.g., BERT (Bidirectional Encoder Representation from Transformers)), a large language model based on a conversational mode (e.g., GPT (Generative Pre-Trained Transformer)), etc.
[0051] In some embodiments, the generative model may be a large language model based on a pre-training plus fine-tuning mode. For example, the generative model is a model obtained through training based on at least one of a BLOOM model (Big Science Large Open-science Open-access MultilingualLanguage Model), a model that interacts in a conversational manner (such as ChatGPT (Chat Generative Pre-Trained Transformer)), or other models.
[0052] In some embodiments of the present specification, a large language model with high semantic understanding ability can make the determination of the generated results more accurate, so that the intelligent dialogue system can better adapt to the in-vehicle environment.
[0053] In some embodiments, the input of the generation model 222 may include processed text 221 , and the output may be a generation result 223 .
[0054] In some embodiments, the generative model can be obtained by training a large language model (such as a BLOOM model, etc.) based on the initial training sample. The initial training sample may include the processed sample text, and the label corresponding to the initial training sample may be the actual generation result of the sample speech recognition text corresponding to the processed sample text. The initial training sample and the corresponding label can be obtained based on historical data. For more information about the processed sample text and the sample speech recognition text, please refer to Figure 4 and its related description.
[0055] In some embodiments, the processor can construct a loss function based on the output of the BLOOM model and the labels corresponding to the initial training samples; based on the iterative update of the loss function until the loss function is less than a threshold or converges, or the training cycle reaches a threshold and other conditions are met, the trained BLOOM model is obtained, that is, the generated model.
[0056] In some embodiments, the generative model can also be obtained based on the first training sample training. Each of the at least one first training sample can include a sample processed text and a label representing the sample corrected text of the sample processed text. In some embodiments, the processor can fine-tune the initial generative model based on the at least one first training sample to determine a pre-trained model; and determine the generative model based on the pre-trained model. For the model structure and training process of the generative model, please refer to Figure 3 and Figure 4 Related description.
[0057] In some embodiments, the input for generating the model may also include historical data of the user.
[0058] Historical data refers to historical data related to the current user. For example, historical data may include historical data related to the current user's behavior. The current user refers to an individual or group that is using the intelligent dialogue. For example, one or more of the passengers and / or drivers of a smart car. In some embodiments, historical data may include user preference data (for example, the user's favorite music type is rock music, etc.), user habit data (for example, the user's frequently navigated destination is community A, etc.), etc.
[0059] In some embodiments, the processor may obtain historical data through a historical data table in a storage device. In some embodiments, the historical data table may be constructed by recording the historical data of the user. For example, the processor constructs the historical data table by the relevant types and corresponding relevant contents of the user's historical data. Exemplarily, the historical data table may be as shown in Table 1 below: Table 1 type content Music Type Rock music, classical music Navigation location Airport, Community A …… …… Song Title Shame, Moonlight Sonata
[0060] The first column in Table 1 indicates the relevant type of historical data; the second column indicates the relevant content of the historical data.
[0061] In some embodiments, the historical data of the user input into the generation model may be in a concatenated format of different types and corresponding contents. For example, the historical data may be represented as “[CLS] Music type: rock [SEP] Location: airport [SEP] Song name: shameless [SEP]” and the like.
[0062] In some embodiments, when the input for generating the model includes historical data of a user, the initial training sample may also include sample historical data of a sample user.
[0063] In some embodiments of the present specification, adding the user's historical data to the generation model input is equivalent to introducing a user adaptive processing unit in the intelligent dialogue system, so that the generation model can provide the user with personalized answers that better meet the user's needs based on the user's preferences and habits.
[0064] In some embodiments, the generation model may include an adaptive layer, an encoding layer, and a classification layer. The processor may determine user feature data through the adaptive layer based on historical data; determine text feature data through the encoding layer based on processed text; and determine the generation result through the classification layer based on user feature data and text feature data. For more information about this section, please refer to Figure 3 and its related description.
[0065] In some embodiments, the processor can determine the revised text based on the processed text through the generation model; determine the expanded text based on the revised text through the generation model; and determine the generated result based on the expanded text and historical data through the generation model. For more information about this section, please refer to Figure 5 and its related description.
[0066] In some embodiments of the present specification, based on the processed speech recognition text, the machine learning model obtained through instruction learning training can accurately determine the answer to the user's question, so that the intelligent dialogue system can better adapt to the in-vehicle environment, have higher semantic understanding ability, and provide better services to users.
[0067] It should be noted that the above description of the process 200 is only for example and illustration, and does not limit the scope of application of this specification. For those skilled in the art, various modifications and changes can be made to the process 200 under the guidance of this specification. However, these modifications and changes are still within the scope of this specification.
[0068] Figure 3is an exemplary schematic diagram of a generation model according to some embodiments of the present specification.
[0069] In some embodiments, Figure 3 As shown, the generation model 222 may include an adaptive layer 222-1, an encoding layer 222-2, and a classification layer 222-3. The processor may determine user feature data 320 through the adaptive layer 222-1 based on the historical data 310; determine text feature data 330 through the encoding layer 222-2 based on the processed text 221; and determine the generation result 223 through the classification layer 222-3 based on the user feature data 320 and the text feature data 330.
[0070] The adaptive layer refers to a machine learning model used to determine user feature data, such as a recurrent neural network (RNN). The adaptive layer can extract features from historical data and generate feature vectors (user feature data).
[0071] In some embodiments, the input of the adaptive layer can be historical data, and the output can be user feature data. For more information about historical data, see Figure 2 and its related description.
[0072] User feature data refers to data reflecting user-related features. In some embodiments, the user feature data may be a feature vector and / or feature tensor of fixed dimensions that characterize user features. The fixed dimensions may be preset manually or set by system default. For example, the user feature data may be a two-dimensional feature vector (e.g., [a, b]), or a third-order feature tensor (e.g., (i, j, k)).
[0073] In some embodiments, the encoding layer may be a machine learning model. For example, the encoding layer may include a deep neural network, or at least one of other structures. Exemplarily, the encoding layer may be an encoder layer (Encoder Layer) of a transformer model (Transformer model). The encoding layer may encode the processed text to generate a feature vector and / or feature tensor (text feature data) with the same dimension as the user feature data.
[0074] In some embodiments, the input of the encoding layer may be processed text, and the output may be text feature data.
[0075] Text feature data refers to data reflecting features related to the user input question. In some embodiments, the text feature data may be a feature vector and / or feature tensor of fixed dimension that characterizes the user input question. In some embodiments, the text feature data and the user feature data have the same dimension. For example, if the user feature data is a two-dimensional feature vector (e.g., [a, b]), then the text feature data also needs to be a two-dimensional feature vector (e.g., [c, d]).
[0076] In some embodiments, the classification layer may be a machine learning model. For example, the classification layer may include a deep neural network, or at least one of other structures. Exemplarily, the classification layer may be a classification layer (Softmax Layer) of a Transformer model, etc.
[0077] In some embodiments, the input of the classification layer may be user feature data and text feature data, and the output may be a generation result.
[0078] In some embodiments, the processor may add the user feature data and the text feature data to fuse the information to obtain fused feature data, and input the fused feature data into the classification layer. The fused feature data may be characterized as a feature vector and / or feature tensor obtained by adding the user feature data and the text feature data. For example, the user feature data is a two-dimensional feature vector [a, b], the text feature data is a two-dimensional feature vector [c, d], and the fused feature data may be [a+c, b+d].
[0079] In some embodiments, the processor inputs the fused feature data into the classification layer, and the classification layer can convert the output answer to the question into a probability distribution representing the answer to the question, thereby determining the generated result.
[0080] In some embodiments, the generative model may be obtained in a variety of ways. For example, the processor may obtain the generative model by jointly training the adaptive layer, the encoding layer, and the classification layer.
[0081] In some embodiments, the joint training process may include: forming training samples of the classification layer based on the joint training samples through an adaptive layer and a coding layer, and performing training through supervised learning based on training labels of the training samples.
[0082] In some embodiments, the model parameters of the generated model can be obtained by training the initial adaptive layer, the initial coding layer and the initial classification layer using multiple joint training samples. Each joint training sample may include sample historical data and sample processed text. The sample historical data can be obtained based on a large amount of historical data. The sample processed text can be determined based on the sample speech recognition text. The label corresponding to the joint training sample can be the actual generation result of the sample speech recognition text corresponding to the sample processed text. In some embodiments, the label corresponding to the joint training sample can be obtained based on historical records. For more information about sample processed text and sample speech recognition text, please refer to Figure 4 and its related description.
[0083] The training of the initial adaptive layer, the initial coding layer, and the initial classification layer may include one or more iterations. As an example only, in the current iteration, for each joint training sample, the processor may input the sample history data into the intermediate adaptive layer, and input the sample processed text into the intermediate coding layer, and then input the user feature data output by the intermediate adaptive layer and the text feature data output by the intermediate coding layer into the intermediate classification layer, thereby determining the generation result of the sample speech recognition text. If the current iteration is the first iteration, the intermediate adaptive layer, the intermediate coding layer, and the intermediate classification layer may be the initial adaptive layer, the initial coding layer, and the initial classification layer, respectively. If the current iteration is another iteration, the intermediate layers may be the layers updated in the previous iteration. The processor may construct a loss function based on the generation result output by the intermediate classification layer and the label corresponding to the joint training sample, and synchronously update the parameters of the intermediate adaptive layer, the intermediate coding layer, and the intermediate classification layer based on the loss function. When the loss function meets the preset conditions (such as loss function convergence, loss function value is less than the preset value, etc.), the parameter update of each layer is completed, and the trained adaptive layer, coding layer, and classification layer are obtained, that is, the generation model training is completed.
[0084] In some embodiments, the processor may obtain at least one first training sample, wherein each of the at least one first training sample includes a sample processed text and a label representing a sample corrected text of the sample processed text; based on the at least one first training sample, fine-tune the initial generation model to determine a pre-training model; and based on the pre-training model, determine the generation model. For more information about this section, please refer to Figure 4 and its related description.
[0085] In some embodiments, the processor may also determine at least one second training sample based on the target instruction set, where the at least one second training sample includes an instruction in the target instruction set; and fine-tune the pre-trained model based on the label corresponding to the instruction in the target instruction set and the at least one second training sample to determine the generated model. Figure 4 and its related description.
[0086] In some embodiments of the present specification, historical data and processed text are converted into corresponding feature vectors and / or feature tensors of the same dimension through an adaptive layer and a coding layer, respectively, and then the probability distribution of the answers to the questions is mapped through a classification layer. This can improve the adaptive ability of the model, determine personalized generation results that are more in line with the user's historical preferences and habits, and help improve user satisfaction with the intelligent dialogue system.
[0087] Figure 4 It is an exemplary schematic diagram of the training process of the generative model shown in some embodiments of the present specification.
[0088] In some embodiments, the generative model can be obtained through multiple stages of training. For example, the training process of the generative model may include a pre-training process and a training process. In some embodiments, the pre-training process includes: based on the first training sample and the corresponding label, fine-tuning the initial generative model to determine the pre-training model. The specific process of the pre-training process includes the following:
[0089] In some embodiments, Figure 4 As shown, the training process of the generation model 222 includes: obtaining at least one first training sample 410, wherein each of the at least one first training sample 410 includes a sample processed text 410-1 and a label 410-2 representing a sample corrected text of the sample processed text; based on the at least one first training sample 410, fine-tuning the initial generation model 420 to determine the pre-trained model 430; based on the pre-trained model 430, determining the generation model 222.
[0090] The initial generation model refers to the basic model for obtaining the generation model. In some embodiments, the initial generation model can be a machine learning model with the following custom structure or other structures. For example, the initial generation model can be a large language model such as BLOOM, ChatGPT, etc.
[0091] The pre-trained model refers to a model obtained after fine-tuning the initial generation model. In some embodiments, the pre-trained model can be a machine learning model with a custom structure as described below, or other structures. For example, the pre-trained model can be a large language model trained from scratch, or a model that is pre-trained twice based on a large language model such as BLOOM.
[0092] In some embodiments, each of the at least one first training sample may include a sample processed text and a label representing a sample corrected text of the sample processed text.
[0093] The sample processed text refers to the processed text used to train the initial generation model. In some embodiments, the sample processed text can be obtained through sample speech recognition text. In some embodiments, the processor can convert a large amount of historical user input speech into sample speech recognition text by using Automatic Speech Recognition (ASR) technology or a machine learning model (such as an ASR model). The processor preprocesses the sample speech recognition text to obtain the sample processed text. For more information about the preprocessing process, please refer to Figure 2 and its related description.
[0094] The sample corrected text refers to the corrected text used to train the initial generation model. The corrected text refers to the correct text after the processed text with translation errors is corrected. In some embodiments, the label of the sample corrected text of the sample processed text can be obtained by screening the sample processed text with translation errors from a large number of historical processed texts, and manually correcting the sample processed text to obtain the corresponding correct text as the label of the sample corrected text of the sample processed text. Exemplarily, the processed text and the corrected text can be as shown in the following Table 2: Table 2
[0095] The first column in Table 2 represents the processed text; the second column represents the corrected text. The sample processed text and the corresponding sample corrected text can be any one or more groups in Table 2.
[0096] In some embodiments, a first training sample may include a sentence in the sample processed text and the first t-1 words of the sample corrected text of the sentence. The label corresponding to the first training sample may include the tth actual word of the sample corrected text of the sentence in the sample processed text.
[0097] The training method of the pre-trained model may include an autoregressive encoding method and the like.
[0098] Exemplarily, the autoregressive encoding method can be expressed by the following formula (1): Among them, x 1:T Represents the first to T words in the sample text after processing, y 1:T Represents the first to T words of the sample correction text, y 1 and t represents the first and tth words of the sample correction text, y 1:t-1Represents the 1st to t-1th words of the sample revised text. Formula (1) represents the probability of predicting the occurrence of the word at position t in the sample revised text by using a sentence in the sample processed text and the first t-1 words of the sample revised text of this sentence. The value of T does not exceed the total number of words in a sentence in the sample processed text, and the value of t can be any integer from 1 to T.
[0099] In some embodiments, the pre-trained model can be obtained by adjusting the initial generation model parameters so that the probability of predicting the occurrence of T position words is maximized, that is, the probability that the prediction result of the initial generation model is consistent with the actual result is maximized.
[0100] In some embodiments of the present specification, the initial generation model is fine-tuned through sample processed text and labels of sample corrected text representing the sample processed text, so that intelligent dialogue knowledge can be injected into the pre-trained model, thereby improving the pre-trained model's understanding and reasoning capabilities on speech recognition text, thereby enabling the subsequently determined generation model to better generate the best answer based on the user's habits, thereby improving the user experience.
[0101] In some embodiments, the processor can determine the generated model in a variety of ways based on the pre-trained model. The training process refers to secondary training based on the pre-training process. In some embodiments, the training process includes: fine-tuning the pre-trained model based on the second training sample and the corresponding label to determine the generated model. The specific process of the training process includes the following:
[0102] In some embodiments, Figure 4 As shown, the processor can determine at least one second training sample 450 based on the target instruction set 440, and the at least one second training sample 450 includes an instruction 450-1 in the target instruction set; based on an instruction 450-1 in the target instruction set and a label 460 corresponding to the at least one second training sample, the pre-trained model 430 is fine-tuned to determine the generated model 222.
[0103] The target instruction set refers to a set of instructions used to guide the generation model to determine the generation result. The target instruction set may include one or more instructions. The instruction is related to the specific needs of the user. The instruction can be expressed by a sentence, a prompt word, a paragraph, etc. For example, the instruction may be "Please plan a route to community A, do not use highways or toll roads, choose the shortest road, and notify me in advance before changing lanes." etc.
[0104] In some embodiments, the processor may construct an original instruction set; based on the original instruction set, determine the target instruction set through a similarity generation model, where the similarity generation model is a large language model.
[0105] The original instruction set refers to a set of basic instructions. For example, the original instruction set may include original instructions such as "Please navigate to community A, avoid highways and toll roads, choose the shortest route, and remind me to change lanes in advance."
[0106] In some embodiments, the processor can construct and obtain the original instruction set in a variety of ways. For example, the original instruction set can be obtained by manual construction. Exemplarily, the manager or staff of the intelligent dialogue system can manually input at least one original instruction. The processor can determine the at least one original instruction manually input as the original instruction set.
[0107] In some embodiments, the target instruction set may refer to a set of instructions that are similar and / or identical to the original instruction. In some embodiments, the target instruction set may include one or more sets of target instructions, wherein a set of target instructions may include an original instruction and one or more similar instructions corresponding thereto. For example, a set of target instructions may include an original instruction: "Please navigate to Community A, avoid highways and toll roads, choose the shortest route, and remind me in advance to change lanes." and multiple similar instructions similar thereto: "Please plan a route to Community A, do not use highways or toll roads, choose the shortest road, and notify me in advance before changing lanes.", "Navigate to Community A, avoid highways and toll roads, choose the shortest path, and remind me in advance before turning.", "For navigation to Community A, please avoid highways and toll roads, choose the shortest route, and inform me in advance when to change lanes.", etc.
[0108] Exemplarily, a set of target instructions may be shown in Table 3 below: Table 3
[0109] In Table 3, the first column represents the target instruction; the first row of the second column represents the original instruction; the second row of the second column represents the similar instruction similar to the original instruction. The target instruction includes the original instruction and the similar instruction.
[0110] In some embodiments, the processor may expand the original instruction set to obtain the target instruction set in a variety of ways (such as similar sentence generation technology, etc.). Similar sentence generation technology refers to a technology that can generate similar sentences based on the original sentence. For example, similar sentence generation technology may include multiple, such as one-key Chinese data enhancement tool (Natural Language Processing Chinese Data Augmentation, NLPCDA), machine learning models, etc.
[0111] In some embodiments, the processor may determine the target instruction set based on the original instruction set by using a similarity generation model.
[0112] In some embodiments, the similarity generation model refers to a model that can be used to generate similar sentences. The similarity generation model can be a machine learning model. For example, a language model, etc. The language model can be a large language model, etc.
[0113] In some embodiments, the input of the similarity generation model may include an original instruction set, and the output is a target instruction set. In some embodiments, one input of the similarity generation model may include an original instruction, and the output is one or more corresponding target instructions.
[0114] In some embodiments, the similarity generation model can be trained by using a plurality of generated training samples with labels.
[0115] In some embodiments, generating a training sample may include a sample original instruction set. Generating a training sample may be acquired based on historical data. A label corresponding to the generated training sample may include an actual target instruction set. In some embodiments, a generated training sample may include a sample original instruction, and a label corresponding to a generated training sample may be one or more actual target instructions.
[0116] In some embodiments, the labels corresponding to the training samples may be generated by manual annotation based on historical data. For example, the processor may obtain one or more target instructions corresponding to the original instructions in the historical data, and construct an actual target instruction set by counting the target instructions corresponding to multiple groups of original instructions.
[0117] In some embodiments, the processor can construct a loss function based on the output of the similarity generation model and the label corresponding to the generated training sample; iteratively update the loss function until the loss function is less than a threshold or converges, or the training cycle reaches a threshold and other conditions are met, thereby obtaining a trained similarity generation model.
[0118] In some embodiments of the present specification, by constructing an original instruction set, the generative model can fully understand the conversations occurring in various scenarios in the intelligent dialogue system. By using similar generative models, the original instruction set can be quickly expanded to improve the comprehensiveness and accuracy of the determined target instruction set, and problems such as overfitting of the generative model or low accuracy of the output generation results due to too few or too single or incomplete instructions can be avoided.
[0119] In some embodiments, at least one second training sample may include an instruction in the target instruction set. In some embodiments, the processor may determine at least one second training sample in a variety of ways based on the target instruction set. For example, during the training process, the second training sample may be randomly extracted from the target instruction set by the processor or extracted by a preset rule, so that in the case of the same sample during multiple training processes, different instructions can make the model focus more comprehensive and the training effect better. The preset rules can be set manually based on experience or by system default.
[0120] In some embodiments, the label corresponding to the at least one second training sample may be an actual generated result corresponding to the second training sample, or may be a manually annotated correct answer.
[0121] In some embodiments, the processor may fine-tune the pre-trained model to determine a generated model based on an instruction in the target instruction set and a label corresponding to at least one second training sample.
[0122] In some embodiments, the processor may pre-train and fine-tune the pre-trained model based on the pre-trained model by using the second training sample and the label corresponding to the second training sample to determine the generated model.
[0123] In some embodiments, the processor may randomly extract a sample instruction from the target instruction set as a second training sample, and use the actual generated result (or correct answer) corresponding to the sample instruction as the label corresponding to the second training sample, and use the sample instruction and the corresponding actual generated result (or correct answer) as an example of a set of training data. For example, an example of training data may include a sample instruction: "Please plan a route to Community A, do not use highways or toll roads, choose the shortest road, and notify me in advance before changing lanes.", and a label corresponding to the sample instruction: "The shortest path has been planned. Please go straight along Huaihai Road for about 1 km, turn left at the second intersection, and continue to go straight along Yinhe Road for about 2 km. The entire journey is about 3 km and is expected to take 8 minutes. Start navigating now: Go straight at the traffic light ahead, continue straight, and turn left soon. Please change lanes in time, turn left, and continue straight. You have arrived at your destination, and this navigation is over." The processor can determine multiple examples of multiple sets of training data through labels corresponding to multiple sample instructions and multiple sample instructions.
[0124] In some embodiments, the processor can input an example of the above-mentioned set of training data as a whole into the pre-training model. Through multiple examples of multiple sets of training data, the pre-training model can continuously learn and fine-tune, and then determine the trained generation model. The number of examples of the above-mentioned training data can be multiple, and the processor can be set according to actual needs.
[0125] The second training samples input into the pre-trained model may include multiple types, such as navigation, playing songs, setting reminders, etc. By learning from a variety of different training data, the pre-trained model can learn different intelligent dialogue scenarios, and then determine the trained generation model to improve the accuracy of the generation results output by the subsequent generation model.
[0126] In some embodiments, the generation model may include an adaptive layer, a coding layer, and a classification layer. The processor may perform a training process based on the adaptive layer, the coding layer, and the classification layer to obtain the generation model. For more information about the training process of the generation model, see Figure 3 Description of joint training.
[0127] In some embodiments of the present specification, a generative model is obtained by training with instructions in a target instruction set, so that the scope of the generative model can be more comprehensive, and the model can learn knowledge in intelligent dialogues, acquire more semantic understanding capabilities, and thus accurately determine the generation results for various scenarios.
[0128] Figure 5 is an exemplary overall flow chart of the intelligent dialogue method according to some embodiments of this specification. Figure 5 As shown, process 500 may include the following steps. In some embodiments, Figure 5 The process 500 shown may be called and / or executed by the processor 110 .
[0129] Step 510, based on the processed text, determine the revised text by generating a model.
[0130] Corrected text refers to the correct text of the processed text. For more information about processed text and corrected text, see Figure 2 and Figure 4 Related description.
[0131] In some embodiments, the processor may input the processed text into a generation model and output a corrected text. The model training method for implementing the function of outputting the corrected text may include an autoregressive encoding method, etc. For more information about the model training process, see Figure 4 The pre-training process will not be described here.
[0132] Step 520, based on the revised text, determine the expanded text by generating a model.
[0133] In some embodiments, the processor may input the revised text into a generation model and output the expanded text.
[0134] In some embodiments, the generative model that implements the function of outputting expanded text can be trained by multiple labeled expanded training samples.
[0135] In some embodiments, the extended training sample may include sample revised text. The label corresponding to the extended training sample may include the actual expanded text. The extended training sample may be obtained based on historical data. The label corresponding to the extended training sample may be manually annotated.
[0136] In some embodiments, the processor can construct a loss function based on the output of the generation model and the label corresponding to the extended training sample; iteratively update the loss function until the loss function is less than a threshold or converges, or the training cycle reaches a threshold and other conditions are met, thereby obtaining a trained generation model.
[0137] For more information on how to train a generative model to output expanded text, see Figure 4 The training method of similar generative models in is not described here.
[0138] The expanded text refers to at least one similar text of the revised text. For example, the revised text is: "Please navigate to community A, avoid highways and toll roads, choose the shortest route, and remind me in advance to change lanes." The expanded text can be: "Navigate to community A, avoid highways and toll roads, choose the shortest route, and remind me in advance before turning." and other similar texts.
[0139] Step 530, based on the expanded text and historical data, determine the generation result through the generation model.
[0140] In some embodiments, the processor may input the expanded text and historical data into a generation model and output a generation result.
[0141] In some embodiments, the generative model can be trained by using a plurality of third training samples with labels.
[0142] In some embodiments, the third training sample may include sample expanded text and sample historical data. The label corresponding to the third training sample may include an actual generated result. The third training sample may be obtained based on historical data. The label corresponding to the third training sample may be manually annotated.
[0143] In some embodiments, the processor can construct a loss function based on the output of the generation model and the label corresponding to the third training sample; iteratively update the loss function until the loss function is less than a threshold or converges, or the training cycle reaches a threshold and other conditions are met, thereby obtaining a trained generation model.
[0144] For more information on generating results and historical data, see Figure 2 and its related description.
[0145] In some embodiments, the generation model can be obtained by training multiple rounds based on the fourth training sample. The fourth training sample includes the processed sample text and the sample historical data. The fourth training sample can be obtained based on the historical data. The label corresponding to the fourth training sample can be the actual generation result, which can be manually labeled.
[0146] In some embodiments, the processor can input the processed sample text into the generation model and output the corrected sample; then input the output corrected sample into the generation model and output the expanded text; then input the output expanded text and sample history data into the generation model and output the generation result. A loss function is constructed based on the output generation result and the label corresponding to the fourth training sample; the loss function is used to continuously update the parameters of the generation model. Through parameter updating, a trained generation model is obtained.
[0147] In some embodiments of the present description, the model is enabled to correct the processed text through the pre-training process, and the intelligent dialogue model (generative model) is injected with intelligent dialogue knowledge, so that the generative model can obtain more semantic understanding capabilities. The model is enabled to expand the corrected text through the instruction learning process, and the generative model can accurately adapt to the complex scenarios of intelligent dialogue by combining instruction learning. By combining the expanded text and historical data, the generative model can give the best answer according to the user's habits, so as to meet the personalized needs of different users in various scenarios and improve the user experience.
[0148] One or more embodiments of the present specification provide an intelligent dialogue device, which includes a processor and a memory; the memory is used to store instructions, and when the instructions are executed by the processor, the device implements an intelligent dialogue method.
[0149] One or more embodiments of the present specification provide a computer-readable storage medium, wherein the storage medium stores computer instructions. When a computer reads the computer instructions in the storage medium, the computer executes an intelligent dialogue method based on speech recognition text.
[0150] The solutions described in this manual, if they involve the processing of personal information, will be processed on the premise of having a legal basis (such as obtaining the consent of the subject of personal information, or being necessary for the performance of a contract, etc.), and will only be processed within the scope of regulations or agreements. If a user refuses to process personal information other than the necessary information for basic functions, it will not affect the user's use of basic functions.
[0151] The basic concepts have been described above. Obviously, for those skilled in the art, the above detailed disclosure is only for example and does not constitute a limitation of this specification. Although not explicitly stated here, those skilled in the art may make various modifications, improvements and corrections to this specification. Such modifications, improvements and corrections are suggested in this specification, so such modifications, improvements and corrections still belong to the spirit and scope of the exemplary embodiments of this specification.
[0152] At the same time, this specification uses specific words to describe the embodiments of this specification. For example, "one embodiment", "an embodiment", and / or "some embodiments" refer to a certain feature, structure or characteristic related to at least one embodiment of this specification. Therefore, it should be emphasized and noted that "one embodiment" or "an embodiment" or "an alternative embodiment" mentioned twice or more in different positions in this specification does not necessarily refer to the same embodiment. In addition, certain features, structures or characteristics in one or more embodiments of this specification can be appropriately combined.
[0153] In addition, unless explicitly stated in the claims, the order of the processing elements and sequences described in this specification, the use of alphanumeric characters, or the use of other names are not intended to limit the order of the processes and methods of this specification. Although the above disclosure discusses some invention embodiments that are currently considered useful through various examples, it should be understood that such details are only for illustrative purposes, and the attached claims are not limited to the disclosed embodiments. On the contrary, the claims are intended to cover all modifications and equivalent combinations that are consistent with the essence and scope of the embodiments of this specification. For example, although the system components described above can be implemented by hardware devices, they can also be implemented only by software solutions, such as installing the described system on an existing server or mobile device.
[0154] Similarly, it should be noted that in order to simplify the description disclosed in this specification and thus help understand one or more embodiments of the invention, in the above description of the embodiments of this specification, multiple features are sometimes combined into one embodiment, figure or description thereof. However, this disclosure method does not mean that the features required by the subject matter of this specification are more than the features mentioned in the claims. In fact, the features of the embodiments are less than all the features of the single embodiment disclosed above.
[0155] In some embodiments, numbers describing the number of components and attributes are used. It should be understood that such numbers used in the description of the embodiments are modified by the modifiers "about", "approximately" or "substantially" in some examples. Unless otherwise specified, "about", "approximately" or "substantially" indicate that the numbers are allowed to vary by ±20%. Accordingly, in some embodiments, the numerical parameters used in the specification and claims are approximate values, which may change according to the required features of individual embodiments. In some embodiments, the numerical parameters should take into account the specified significant digits and adopt the general method of retaining digits. Although the numerical domains and parameters used to confirm the breadth of their range in some embodiments of this specification are approximate values, in specific embodiments, the setting of such numerical values is as accurate as possible within the feasible range.
[0156] Each patent, patent application, patent application publication, and other materials, such as articles, books, specifications, publications, documents, etc., cited in this specification are hereby incorporated by reference in their entirety. Except for application history documents that are inconsistent with or conflicting with the contents of this specification, documents that limit the broadest scope of the claims of this specification (currently or later attached to this specification) are also excluded. It should be noted that if the descriptions, definitions, and / or use of terms in the materials attached to this specification are inconsistent or conflicting with the contents described in this specification, the descriptions, definitions, and / or use of terms in this specification shall prevail.
[0157] Finally, it should be understood that the embodiments described in this specification are only used to illustrate the principles of the embodiments of this specification. Other variations may also fall within the scope of this specification. Therefore, as an example and not a limitation, alternative configurations of the embodiments of this specification may be considered consistent with the teachings of this specification. Accordingly, the embodiments of this specification are not limited to the embodiments explicitly introduced and described in this specification.
Claims
1. An intelligent dialogue method, It is characterized in that The method comprises: Determine processed text based on the user's speech recognition text; Based on the processed text, a generation result is determined by a generation model, wherein the generation model is a machine learning model obtained through instruction learning training, and the generation model performs text correction on the processed text.
2. The method according to claim 1, It is characterized in that The input of the generation model also includes the historical data of the user.
3. The method according to claim 2, It is characterized in that The generation model includes an adaptive layer, a coding layer and a classification layer, and the method further includes: Based on the historical data, determining user characteristic data through the adaptive layer; Based on the processed text, determining text feature data through the encoding layer; The generation result is determined by the classification layer based on the user feature data and the text feature data.
4. The method according to claim 1, It is characterized in that The training process of the generative model includes: Acquire at least one first training sample, wherein each of the at least one first training sample includes a sample processed text and a label representing a sample corrected text of the sample processed text; Based on the at least one first training sample, fine-tune the initial generation model to determine a pre-training model; Based on the pre-trained model, the generation model is determined.
5. The method according to claim 4, It is characterized in that Based on the pre-trained model, determining the generated model includes: Based on the target instruction set, determining at least one second training sample, wherein the at least one second training sample includes an instruction in the target instruction set; Based on an instruction in the target instruction set and a label corresponding to the at least one second training sample, the pre-trained model is fine-tuned to determine the generated model.
6. The method according to claim 1, It is characterized in that Determining the generation result by generating a model based on the processed text includes: Based on the processed text, determining a revised text through the generation model, wherein the revised text is a correct text of the processed text; Based on the revised text, determining an expanded text by the generation model, wherein the expanded text is at least one similar text to the revised text; Based on the expanded text and historical data, the generation result is determined by the generation model.
7. An intelligent dialogue system, It is characterized in that The system comprises: The text determination module is configured to determine a processed text based on the user's speech recognition text; The result determination module is configured to determine a generation result based on the processed text through a generation model, wherein the generation model is a machine learning model obtained through instruction learning training, and the generation model performs text correction on the processed text.
8. The system of claim 7, It is characterized in that The system further comprises a training module, wherein the training module is configured to: Acquire at least one first training sample, wherein each of the at least one first training sample includes a sample processed text and a label representing a sample corrected text of the sample processed text; Based on the at least one first training sample, fine-tune the initial generation model to determine a pre-training model; Based on the pre-trained model, the generation model is determined.
9. An intelligent dialogue device, It is characterized in that The device includes at least one memory and at least one processor, wherein the at least one memory is used to store computer instructions, and the at least one processor executes the computer instructions or part of the instructions to implement the method according to any one of claims 1 to 6.
10. A computer-readable storage medium, It is characterized in that The storage medium stores computer instructions. When a computer reads the computer instructions, the computer executes the method according to any one of claims 1 to 6.