Voice information rejection method and device and electronic equipment
By using small language models and large language models to identify and reject voice information in smart home devices, the problem of misidentification of voice information after extending the listening time is solved, and more accurate voice interaction is achieved.
Patent Information
- Application Number
- CN202510442949.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-06-13
AI Technical Summary
Existing smart home devices can easily cause misidentification of voice information after extending the listening time, resulting in incorrect or unnecessary replies.
By applying the voice information refusal method in the terminal device, the collected voice information is converted and recognized using a small language model and a large language model. The specific steps include: obtaining text input information, initially identifying the small language model, if the recognition result is to be confirmed, further identifying the large language model, and determining whether to perform rejection processing based on the comprehensive score.
It effectively avoids misidentification of information without semantic information or information without human-computer dialogue, and improves the accuracy and user experience of voice interaction.
Smart Images

Figure CN120148510A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of natural language processing, and particularly relates to a method, apparatus, and electronic device for rejecting speech information recognition. Background Art
[0002] Currently, electronic devices such as smart home devices usually support voice interaction with users. For example, after the user wakes up the device by voice, the user can have a human-machine conversation with the smart home device. In order to support the user to have multiple rounds of conversations with the smart home device during a single wake-up process, the current mainstream approach is to extend the listening duration of the smart home device, that is, after the device is woken up or after a conversation, as long as a human-machine conversation instruction from the user is received again within a preset duration, the relevant operation indicated by the instruction can continue to be executed.
[0003] However, extending the listening duration of the smart home device is likely to cause some misrecognition situations. For example, during the listening process, user input unrelated to human-machine interaction is received, such as misrecognizing the user's conversation with other objects as a human-machine interaction instruction, and at this time, incorrect or unnecessary responses will be generated. Summary of the Invention
[0004] In view of this, the present application provides a method, apparatus, and electronic device for rejecting speech information recognition to avoid misrecognizing non-human-machine interaction instructions included in the collected speech information.
[0005] The technical solutions provided by the present application are as follows:
[0006] According to an embodiment of the first aspect of the present application, a method for rejecting speech information recognition is provided. The method is applied to a terminal device, and the terminal device is used for voice interaction with a user. The method includes:
[0007] Obtain text input information, where the text input information is obtained by converting the collected speech input information;
[0008] Input the text input information into a trained small language model to obtain a first recognition result, where the first recognition result is used to indicate whether the text input information has semantics and whether it is a human-machine conversation;
[0009] If the first recognition result is a to-be-confirmed result, input the text input information into a trained large language model to obtain a comprehensive score corresponding to the text input information; the to-be-confirmed result means that the text input information has no semantics and / or is not a human-machine conversation;
[0010] Determine whether to perform rejection processing on the speech input information corresponding to the text input information according to the comprehensive score.
[0011] Optionally, before inputting the text input information into the trained small language model, the method further includes:
[0012] Determine whether the text input information conforms to a preset rule. The conformity to the preset rule includes: the text input information matches the reference information included in the stored text information library, and the reference information refers to the input information that needs to be responded to; or, the text input information satisfies the structure defined by a preset regular expression, and the structure defined by the regular expression refers to the structure of the text input information that needs to be responded to;
[0013] If so, determine to identify the text input information to respond to the corresponding voice input information of the text input information;
[0014] If not, continue to execute the step of inputting the text input information into the trained small language model.
[0015] Optionally, the small language model includes a prediction model, a first classifier, and a second classifier. The step of inputting the text input information into the trained small language model to obtain a first recognition result includes:
[0016] Input the text input information into the prediction model to obtain a text feature vector, and the text feature vector is used to represent the overall feature of the text input information;
[0017] Process the text feature vector through the first classifier to obtain a first classification result corresponding to the text input information, and the first classification result includes having semantics or no semantics;
[0018] Process the text feature vector through the second classifier to obtain a second classification result corresponding to the text input information, and the second classification result includes human-computer dialogue or not human-computer dialogue;
[0019] Use the first classification result and the second classification result corresponding to the text input information as the first recognition result.
[0020] Optionally, the large language model is configured with at least one classification category. The step of inputting the text input information into the trained large language model to obtain a comprehensive score corresponding to the text input information includes:
[0021] Process the text input information based on the large language model to obtain the score value corresponding to each classification category of the text input information; the classification categories include at least one of the following: semantic smoothness score, dialogue subject score, business instruction score, where the dialogue subject score is used to represent the probability that the text input information belongs to a human-machine dialogue, and the business instruction score is used to represent the probability that the text input information is a business instruction corresponding to the business supported by the electronic device that collects the text input information;
[0022] Determine the comprehensive score corresponding to the text input information according to the score value corresponding to each classification category.
[0023] Optionally, the classification categories further include at least one of the following: hot word score, previous context score, and high-frequency word score;
[0024] Among them, the hot word score refers to an additional score for the text input information when the text input information contains predefined hot words, and the hot words are words related to business instructions used to represent instruction objects; the previous context score is used to represent the degree of relevance between the text input information and the previous human-machine dialogue information; the high-frequency word score refers to an additional score for the text input information when the text input information contains predefined high-frequency words, and the high-frequency words are words used to represent actions with a frequency greater than a preset threshold during human-machine dialogue.
[0025] Optionally, the determining whether to perform rejection recognition processing on the voice input information corresponding to the text input information according to the comprehensive score includes:
[0026] If the comprehensive score is greater than or equal to a preset score threshold, determine to recognize the text input information to respond to the voice input information corresponding to the text input information;
[0027] If the comprehensive score is less than the preset score threshold, determine to perform rejection recognition processing to reject responding to the voice input information corresponding to the text input information.
[0028] Optionally, the training method of the small language model includes:
[0029] Collect sample data, where the sample data includes at least one of network data, business data, and generated data. The network data includes an open-source unsupervised data set collected from an online platform, the business data includes positive sample data of human-machine dialogue in a specified business, and the generated data includes negative sample data generated according to a large language model or negative sample data obtained by processing the business data according to specified rules;
[0030] Annotate each sample data to obtain the annotated sample data; the annotated sample data includes a first label and a second label, the first label is used to identify whether the sample data has semantics, and the second label is used to identify whether the sample data is a human-machine dialogue;
[0031] Input the annotated sample data into the prediction model to obtain a sample feature vector; the sample feature vector is used to represent the overall feature of the text input information;
[0032] Input the sample feature vector into the first classifier, and train the first classifier according to the sample feature vector and the first label corresponding to the sample data;
[0033] For each sample data, if it is determined that the first label corresponding to the sample data indicates that the sample data has semantics, input the sample feature vector corresponding to the sample data into the second classifier, and train the second classifier according to the sample feature vector and the second label corresponding to the sample data.
[0034] Optionally, the training method of the large language model includes:
[0035] Collect sample data, the sample data includes at least one of network data, business data, and generated data, the network data includes an open-source unsupervised data set collected from an online platform, the business data includes human-machine dialogue positive sample data in a specified business, and the generated data includes negative sample data generated according to the large language model or negative sample data obtained by processing the business data according to specified rules;
[0036] Annotate each sample data to obtain the annotated sample data; the annotated sample data includes a third label and a score value corresponding to each classification category obtained, and the third label is used to identify whether rejection processing is performed on the sample data;
[0037] Input the annotated sample data into the large language model, and train the large language model according to the sample data, the score value corresponding to each classification category to which the sample data is annotated, and the third label corresponding to the sample data.
[0038] According to an embodiment of the second aspect of the present application, there is provided a voice information rejection device, which is applied to a terminal device, and the terminal device is used for voice interaction with a user. The device includes:
[0039] An obtaining unit, configured to obtain text input information, where the text input information is converted based on the collected voice input information;
[0040] An input unit for inputting the text input information into a trained small language model to obtain a first recognition result, where the first recognition result is used to indicate whether the text input information has semantics and whether it is a human-machine dialogue;
[0041] A determination unit for, if the first recognition result is a to-be-confirmed result, inputting the text input information into a trained large language model to obtain a comprehensive score corresponding to the text input information; the to-be-confirmed result means that the text input information has no semantics and / or is not a human-machine dialogue;
[0042] A processing unit for determining whether to perform rejection processing on the voice input information corresponding to the text input information according to the comprehensive score.
[0043] Optionally, before inputting the text input information into the trained small language model, the input unit is further configured to:
[0044] Determine whether the text input information conforms to a preset rule, where conforming to the preset rule includes: the text input information matches the reference information included in the stored text information library, and the reference information is the input information that needs to be responded to; or, the text input information satisfies the structure defined by a preset regular expression, and the structure defined by the regular expression is the structure of the text input information that needs to be responded to;
[0045] If so, determine to recognize the text input information to respond to the voice input information corresponding to the text input information;
[0046] If not, continue to execute the step of inputting the text input information into the trained small language model.
[0047] Optionally, the small language model includes a prediction model, a first classifier, and a second classifier, and the input unit is specifically configured to:
[0048] Input the text input information into the prediction model to obtain a text feature vector, where the text feature vector is used to characterize the overall feature of the text input information;
[0049] Process the text feature vector through the first classifier to obtain a first classification result corresponding to the text input information, where the first classification result includes having semantics or having no semantics;
[0050] Process the text feature vector through the second classifier to obtain a second classification result corresponding to the text input information, where the second classification result includes human-machine dialogue or not being a human-machine dialogue;
[0051] Take the first classification result and the second classification result corresponding to the text input information as the first recognition result.
[0052] Optionally, the large language model is configured with at least one classification category, and the determining unit is specifically configured to:
[0053] Process the text input information based on the large language model to obtain a score value corresponding to each classification category of the text input information; the classification categories include at least one of the following: semantic smoothness score, dialogue subject score, business instruction score, where the dialogue subject score is used to represent the probability that the text input information belongs to a human-machine dialogue, and the business instruction score is used to represent the probability that the text input information is a business instruction corresponding to the business supported by the electronic device that collects the text input information;
[0054] Determine the comprehensive score corresponding to the text input information according to the score value corresponding to each classification category.
[0055] Optionally, the classification categories further include at least one of the following: hot word score, previous context score, and high-frequency word score;
[0056] Among them, the hot word score refers to an additional score for the text input information when the text input information contains a predefined hot word, and the hot word is a vocabulary related to a business instruction and used to represent an instruction object; the previous context score is used to represent the degree of relevance between the text input information and the previous human-machine dialogue information; the high-frequency word score refers to an additional score for the text input information when the text input information contains a predefined high-frequency word, and the high-frequency word is a vocabulary used to represent an action and having a frequency of occurrence greater than a preset threshold during human-machine dialogue.
[0057] Optionally, the processing unit is specifically configured to:
[0058] If the comprehensive score is greater than or equal to a preset score threshold, determine to recognize the text input information to respond to the voice input information corresponding to the text input information;
[0059] If the comprehensive score is less than the preset score threshold, determine to perform a rejection recognition process on the text input information to reject responding to the voice input information corresponding to the text input information.
[0060] Optionally, the training method of the small language model includes:
[0061] Collect sample data, where the sample data includes at least one of network data, business data, and generated data. The network data includes an open-source unsupervised dataset collected from an online platform. The business data includes positive sample data of human-machine conversations in a specified business. The generated data includes negative sample data generated according to a large language model or negative sample data obtained by processing the business data according to specified rules;
[0062] Annotate each sample data to obtain the annotated sample data. The annotated sample data includes a first label and a second label. The first label is used to identify whether the sample data has semantics, and the second label is used to identify whether the sample data is a human-machine conversation;
[0063] Input the annotated sample data into the prediction model to obtain a sample feature vector. The sample feature vector is used to represent the overall feature of the text input information;
[0064] Input the sample feature vector into the first classifier, and train the first classifier according to the sample feature vector and the first label corresponding to the sample data;
[0065] For each sample data, if it is determined that the first label corresponding to the sample data indicates that the sample data has semantics, input the sample feature vector corresponding to the sample data into the second classifier, and train the second classifier according to the sample feature vector and the second label corresponding to the sample data.
[0066] Optionally, the training method of the large language model includes:
[0067] Collect sample data, where the sample data includes at least one of network data, business data, and generated data. The network data includes an open-source unsupervised dataset collected from an online platform. The business data includes positive sample data of human-machine conversations in a specified business. The generated data includes negative sample data generated according to a large language model or negative sample data obtained by processing the business data according to specified rules;
[0068] Annotate each sample data to obtain the annotated sample data. The annotated sample data includes a third label and a score value corresponding to each classification category obtained. The third label is used to identify whether rejection processing is performed on the sample data;
[0069] Input the annotated sample data into the large language model, and train the large language model according to the sample data, the score value corresponding to each classification category to which the sample data is annotated, and the third label corresponding to the sample data.
[0070] According to an embodiment of the third aspect of the present application, an electronic device is provided, which includes: a processor and a machine-readable storage medium; the machine-readable storage medium stores machine-executable instructions that can be executed by the processor; the processor is configured to execute the machine-executable instructions to implement the method described in the first aspect.
[0071] As can be seen from the above technical solutions, the present application converts the collected voice input information into text input information, and preliminarily identifies the text input information through a trained small language model. When the recognition result of the small language model is information to be confirmed, that is, the text input information is meaningless information and / or the text input information is not a human-machine dialogue, based on the trained large language model, the text input information is further recognized to obtain a comprehensive score of the text input information. Finally, according to the comprehensive score, it is determined whether to recognize or reject the voice input information. Through two confirmations of two models, the situation of misrecognition caused by non-human-machine dialogue or meaningless information included in the collected voice input information is avoided. BRIEF DESCRIPTION OF THE DRAWINGS
[0072] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.
[0073] Figure 1A It is a wake-up scenario diagram of an electronic device provided by an embodiment of the present application;
[0074] Figure 1B It is another wake-up scenario diagram of an electronic device provided by an embodiment of the present application;
[0075] Figure 2 It is a flowchart of a method for rejecting recognition of voice information provided by an embodiment of the present application;
[0076] Figure 3 It is a training schematic diagram of a small language model provided by an embodiment of the present application;
[0077] Figure 4 It is a training schematic diagram of a large language model provided by an embodiment of the present application;
[0078] Figure 5 It is an overall flowchart of a method for rejecting recognition of voice information provided by an embodiment of the present application;
[0079] Figure 6 It is a structural diagram of a device for rejecting recognition of voice information provided by an embodiment of the present application;
[0080] Figure 7 It is a structural diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0081] To enable those skilled in the art to better understand the technical solutions provided by the embodiments of the present application and make the above-mentioned objects, features, and advantages of the embodiments of the present application more obvious and understandable, the technical solutions in the embodiments of the present application will be further described in detail below with reference to the accompanying drawings.
[0082] Current electronic devices, such as smart home devices, usually support voice interaction with users. For example, after the user wakes up the device by voice, they can have a human-machine conversation with the smart home device.
[0083] Please refer to Figure 1A , Figure 1A which is a wake-up scenario diagram of an electronic device provided by an embodiment of the present application.
[0084] As Figure 1A shown, the user can wake up the voice interaction system of the electronic device in the sleep state through a preset wake-up voice command, such as "Hello, xxx". After the wake-up is completed, the user can input voice input information, such as "Help me xxx", to the electronic device. After receiving the corresponding voice input information, the electronic device will execute the operation indicated by the voice input information. After the execution of the operation is completed, the electronic device will enter the sleep state again until it is woken up by the user's preset wake-up voice command again.
[0085] In order to support multi-round conversations between the user and the smart home device during a single wake-up, the listening duration of the smart home device can be extended, that is, after the device is woken up and after each operation indicated by the voice input information is executed, as long as the user's voice input information is received again within the preset duration, the relevant operations indicated by the voice input information can continue to be executed.
[0086] However, extending the listening duration of the smart home device is likely to cause some misrecognition situations. For example, during the listening process, voice information unrelated to human-computer interaction is received, such as misrecognizing the user's conversation with other objects as voice input information. At this time, an incorrect or unnecessary reply will be generated for this voice information.
[0087] Please refer to Figure 1B , Figure 1B which is another wake-up scenario diagram of an electronic device provided by an embodiment of the present application.
[0088] As Figure 1B shown, to solve the above problems, a rejection recognition module can be added to the electronic device with the extended listening function to perform rejection recognition judgment on the voice input information. In the case of determining not to reject, the electronic device executes the operation indicated by the voice input information. In the case of determining to reject, the operation indicated by the voice input information is refused to be executed to avoid the impact of the reply formed by misrecognition on the user.
[0089] Based on this, the present application proposes a method for rejecting speech information recognition to avoid misrecognition of information without semantic meaning or information that is not for human-machine conversation.
[0090] Please refer to Figure 2 , Figure 2 which is the flowchart of the method for rejecting speech information recognition provided by the embodiment of the present application.
[0091] In this embodiment, the method is applied to a terminal device, which can be a device for voice interaction with users, such as a smart home device, a smart wearable device, etc. Specifically, it can be applied to devices with a voice assistant function, such as a smartphone, a smart watch, a smart bracelet, a smart speaker, etc. The present application does not limit this.
[0092] As Figure 2 shown, the method may include the following steps:
[0093] Step 201, obtain text input information.
[0094] Among them, the text input information can be obtained by converting the collected voice input information.
[0095] Since the above-mentioned devices with a voice assistant function usually receive user instructions as voice input information from users, in this embodiment, the collected voice input information can be first converted into text input information through speech recognition technology. Specifically, the method of converting voice input information into text input information can be to input the voice input information into a trained deep neural network or convolutional neural network, etc. to convert it into text input information. The present application does not limit this.
[0096] So far, the description of step 201 ends, and the following is to execute step 202.
[0097] Step 202, input the text input information into a trained small language model to obtain a first recognition result.
[0098] In this embodiment, the first recognition result can be used to indicate whether the text input information has semantics and whether it is for human-machine conversation.
[0099] Among them, whether the text input information has semantics means whether the information is a meaningful information that can be normally understood. For example, some faulty sentences or incomplete sentences, such as "I go today" can be used as text input information without semantics.
[0100] And whether the text input information is for human-machine conversation means whether the information is for interaction with an electronic device, rather than for interaction between users and other objects, such as the conversation during daily chat between two people.
[0101] In this embodiment, the processed text input information can be obtained through a trained small language model to obtain a first recognition result for the text input information.
[0102] Among them, the trained small language model refers to the small language model obtained by fine-tuning the pre-trained small language model. Fine-tuning means continuing to train the pre-trained model through a dataset of a specific task and adjusting its parameters so that the fine-tuned model can better complete the specific task. As for the process of fine-tuning the pre-trained small language model to obtain the trained small language model, it will be described in detail below and will not be elaborated here.
[0103] As an embodiment, the small language model may include a prediction model, a first classifier, and a second classifier. The specific method for obtaining the first recognition result by inputting the text input information into the trained small language model may include:
[0104] Input the text input information into the prediction model to obtain a text feature vector, which is used to represent the overall feature of the text input information;
[0105] Process the text feature vector through the first classifier to obtain a first classification result corresponding to the text input information. The first classification result includes having semantics or no semantics;
[0106] Process the text feature vector through the second classifier to obtain a second classification result corresponding to the text input information. The second classification result includes human-machine dialogue or not human-machine dialogue;
[0107] Take the first classification result and the second classification result corresponding to the text input information as the first recognition result.
[0108] In this embodiment, the small language model may include a prediction model, which may be a model pre-trained based on a Transformer encoder, such as a BERT (Bidirectional Encoder Representations from Transformers) model. The present application does not limit this.
[0109] Taking the trained BERT model as an example, after inputting the text input information into the prediction model, a text feature vector, that is, a CLS (Classification Token) vector, and multiple Token vectors can be obtained. The CLS vector can be used to represent the overall feature of the text input information, and each Token vector can correspond to a character in the text input information.
[0110] In this embodiment, two classifiers are added at the position of outputting the CLS vector (i.e., the text feature vector), denoted as the first classifier and the second classifier respectively. These two classifiers can be binary classifiers. Among them, the first classifier is used to determine whether the text input information has semantics, and the second classifier is used to determine whether the text input information is a human-machine dialogue.
[0111] After the text input information is input into the prediction model to obtain the text feature vector, the text feature vector can be input into the first classifier and the second classifier respectively to obtain the first classification result and the second classification result.
[0112] Specifically, the process of determining the first classification result and the second classification result can be as follows:
[0113] The text feature vector is input into the first classifier to obtain the first probability corresponding to the text feature vector. The first probability is used to indicate the probability that the text input information has semantics. The text feature vector is input into the second classifier to obtain the second probability corresponding to the text feature vector. The second probability is used to indicate the probability that the text input information is a human-machine dialogue.
[0114] If the first probability is less than the first threshold, it is determined that the first classification result corresponding to the text input information is without semantics. If the first probability is not less than the first threshold, it is determined that the first classification result corresponding to the text input information is with semantics. If the second probability is less than the second threshold, it is determined that the second classification result corresponding to the text input information is not a human-machine dialogue. If the second probability is not less than the second threshold, it is determined that the second classification result corresponding to the text input information is a human-machine dialogue.
[0115] In this embodiment, after the first classification result and the second classification result are determined, the first recognition result can be determined according to the first classification result and the second classification result. For example, if the first classification result is "with semantics" and the second classification result is "not a human-machine dialogue", then the first recognition result is "with semantics and not a human-machine dialogue".
[0116] In this embodiment, if the first recognition result is with semantics and is a human-machine dialogue, it is determined that the text input information can be recognized to respond to the corresponding voice input information of the text input information.
[0117] Among them, the specific method for responding to the voice input information corresponding to the text information can be to generate a response information corresponding to the voice input information and reply to the voice input information. For example, if the voice input information is "What's the weather like tomorrow", the process of responding to the voice input information can be to generate a response information to the voice input information, such as "Tomorrow will be sunny, with a maximum temperature of 25 degrees and a minimum temperature of 18 degrees", and output the response information to reply to the voice input information. The specific method for responding to the voice input information corresponding to the text information can also be to execute the operation indicated by the voice input information. For example, if the voice input information is "Play the next song", then execute the operation of stopping the current music playback and starting to play the next song. The present application does not limit the specific method for responding to the voice input information corresponding to the text information.
[0118] If the first recognition result is meaningless and / or not a human-computer dialogue, it is necessary to further recognize the text input information through a large language model.
[0119] Among them, the first recognition result being meaningless and / or not a human-computer dialogue includes the following three situations: being meaningless and being a human-computer dialogue, being meaningless and not being a human-computer dialogue, and having meaning and not being a human-computer dialogue.
[0120] In this embodiment, the training process of the small language model will be described in detail below and will not be elaborated here.
[0121] So far, the description of step 202 ends, and step 203 is executed below.
[0122] Step 203, if the first recognition result is a result to be confirmed, input the text input information into the trained large language model to obtain the comprehensive score corresponding to the text input information.
[0123] Among them, the result to be confirmed refers to that the text input information is meaningless and / or not a human-computer dialogue.
[0124] When it is determined that the first recognition result is a result to be confirmed, that is, the text input information is meaningless and / or not a human-computer dialogue, the present application further inputs the text input information into the trained large language model for recognition to obtain the comprehensive score corresponding to the text input information, so as to confirm whether the voice input information corresponding to the text input information needs to be rejected through the comprehensive score.
[0125] Among them, the trained large language model refers to the large language model obtained by fine-tuning the pre-trained large language model. As for the process of fine-tuning the pre-trained large language model to obtain the trained large language model, it will be described in detail below and will not be elaborated here.
[0126] As an example, the large language model is configured with at least one classification category. The text input information is input into the trained large language model to obtain the comprehensive score corresponding to the text input information, including:
[0127] Based on the large language model's processing of the text input information, the score value corresponding to each classification category of the text input information is obtained; the classification categories include at least one of the following: semantic smoothness score, dialogue subject score, business instruction score. Among them, the dialogue subject score is used to represent the probability that the text input information belongs to a human-machine dialogue, and the business instruction score is used to represent the probability that the text input information is a business instruction corresponding to the business supported by the electronic device that collected the text input information;
[0128] According to the score value corresponding to each classification category, the comprehensive score corresponding to the text input information is determined.
[0129] In addition, the classification categories may further include at least one of: hot word score, previous context score, and high-frequency word score;
[0130] Among them, the hot word score refers to the additional score for the text input information when it contains predefined hot words. A hot word is a vocabulary related to a business instruction that represents an instruction object; the previous context score is used to represent the degree of relevance between the text input information and the previous human-machine dialogue information; the high-frequency word score refers to the additional score for the text input information when it contains predefined high-frequency words. A high-frequency word is a vocabulary that represents an action and has a frequency greater than a preset threshold during human-machine dialogue.
[0131] In this example, the large language model can be pre-configured with multiple classification categories, such as semantic smoothness score, dialogue subject score, business instruction, hot word score, previous context score, and high-frequency word score, etc.
[0132]
[0133] Table 1
[0134] Please refer to Table 1. Table 1 shows the multiple classification categories pre-configured in the above large language model, the optional value distribution of this category, and related descriptions. Among them, the optional value distribution is only exemplary, and can be adjusted according to actual needs.
[0135] The following briefly introduces the classification categories listed in Table 1:
[0136] The semantic smoothness score can represent the level of semantic smoothness and is used to measure whether the grammar of the text input information is smooth and whether the semantics are reasonable. For example, for the text input information "What's the weather like today", its semantic smoothness is relatively high, and the semantic smoothness score detected by the large language model can be relatively high, such as 8 - 10 points; for the text input information "I'm going today", its semantic smoothness is relatively low, and the semantic smoothness score detected by the large language model can be relatively low, such as 0 - 2 points.
[0137] The dialogue subject score is used to indicate the probability that the text input information is a human - machine dialogue. For example, for the text input information "Please help me check the weather for tomorrow", the probability that it is a human - machine dialogue is relatively high, and the dialogue subject score detected by the large language model can be relatively high, such as 8 - 10 points; for the text input information "What do you want to eat tonight", the probability that it is a human - machine dialogue is relatively low, and the dialogue subject score detected by the large language model can be relatively low, such as 0 - 2 points.
[0138] The business instruction score is used to indicate the probability that the text input information is a business instruction corresponding to the business supported by the electronic device that collects the text input information. For example, for some intelligent wearable devices such as smart earphones, for the text input information "Play the next song", since music playback is a business supported by smart earphones, the probability that "Play the next song" is a business instruction corresponding to the business supported by smart earphones is relatively high, and the business instruction score detected by the large language model can be relatively high, such as 8 - 10 points; for the text input information "Please turn on the bedroom light", since controlling smart home devices is not a business supported by smart earphones, the probability that "Please turn on the bedroom light" is a business instruction corresponding to the business supported by smart earphones is relatively low, and the business instruction score detected by the large language model can be relatively high, such as 0 - 2 points.
[0139] The hot - word score can be an additional score for the text input information when the text input information contains predefined hot words. A hot word refers to a vocabulary related to business instructions used to represent the instruction object. For example, the user sets a nickname for an object corresponding to a certain business supported by the electronic device. Exemplarily, the electronic device running the voice information rejection recognition method can control the turning on and off of the air conditioner. The user can set the nickname "XX" of the air conditioner as a hot word and set the corresponding hot - word score, such as 8 points. When the obtained text input information includes the hot word "XX", the hot - word score corresponding to the text input information is 8 points; when the obtained text input information does not include the hot word "XX", the hot - word score corresponding to the text input information is 0 points.
[0140] The above score is used to represent the relevance between the text input information and the previous human-machine dialogue information. The previous human-machine dialogue information refers to the response information of the electronic device to the input text information in the previous round of the human-machine dialogue that was not rejected during the current wake-up period of the electronic device. For example, if the previous round of the non-rejected human-machine dialogue includes: the input text information is "Please help me book a train ticket from Beijing to Shanghai tomorrow morning", and the response of the electronic device to the above text information is "The train tickets from Beijing to Shanghai tomorrow morning are sold out. Would you like me to book a ticket for tomorrow afternoon?", at this time, if the obtained input text information is "It is possible to book a train ticket for tomorrow afternoon", then its relevance to the previous human-machine dialogue information "The train tickets from Beijing to Shanghai tomorrow morning are sold out. Would you like me to book a ticket for tomorrow afternoon?" is relatively high, and the above score detected by the large language model can be relatively high, such as 8-10 points; if the obtained input text information is "Please turn off the bedroom lights", then its relevance to the previous human-machine dialogue information "The train tickets from Beijing to Shanghai tomorrow morning are sold out. Would you like me to book a ticket for tomorrow afternoon?" is relatively low, and the above score detected by the large language model can be relatively low, such as 0-2 points.
[0141] The high-frequency word score refers to the additional score for the text input information when the text input information contains pre-defined high-frequency words. High-frequency words refer to words representing actions that appear more frequently than a preset threshold during human-machine dialogue. For example, in the scenario of controlling smart home devices through an electronic device, action words such as "turn on" and "turn off" are often used. At this time, the above-mentioned commonly used action words can be configured as high-frequency words, and the corresponding high-frequency word scores are configured for them. For example, if the high-frequency word scores of 8 points are configured for "turn on" and "turn off", then when the obtained input text information includes the high-frequency word "turn on" or "turn off", the high-frequency word score corresponding to this text input information is 8 points, and when the obtained text input information does not include the high-frequency word "turn on" or "turn off", the high-frequency word score corresponding to this text input information is 0 points.
[0142] It should be noted that the examples of various types of scores and the specific score settings in the above text are all exemplary. The large language model proposed in the embodiments of the present application can at least obtain one of the semantic smoothness score, the dialogue subject score, and the business instruction, and can also obtain one or more of the hot word score, the above score, and the high-frequency word score. Further, the comprehensive score of the text input information is determined according to the obtained scores. The present application does not limit this.
[0143] After the large language model determines the scores corresponding to the text input data, the comprehensive score can be determined according to the scores corresponding to each score category.
[0144] As an embodiment, the method for determining the comprehensive score may include:
[0145] Add up the scores corresponding to each score category to obtain a comprehensive score;
[0146] Alternatively, according to the weights preset for each score category, perform a weighted sum on each score category to obtain a comprehensive score.
[0147] In this embodiment, the method for determining the comprehensive score based on the scores corresponding to each score category can be to add up the scores corresponding to each obtained score category to obtain a comprehensive score.
[0148] As a preferred embodiment, weights corresponding to each score category can be preset for each score category, and a weighted sum is performed on each score category to obtain a comprehensive score.
[0149] Exemplarily, for a score category that has a greater impact on determining whether the obtained text input information belongs to the text input information that needs to be rejected for recognition, a higher weight can be set.
[0150] For example, text input information with a higher score for the dialogue subject is usually human-machine dialogue information, and it is very likely that the text input information needs to be recognized. Then, a higher weight can be preset for the score of the dialogue subject. After determining the score of the dialogue subject of a certain text input information through a large language model, when calculating the comprehensive score, the product of the score of the dialogue subject and the corresponding weight can be used as the weighted score of the dialogue subject to calculate the comprehensive score. For example, if the large language model determines that the score of the dialogue subject of a certain text input information is 8 points, and the weight configured for the score of the dialogue subject is 2, then when calculating the comprehensive score by weighted calculation, the weighted score of the dialogue subject is 8 * 2 = 16 points.
[0151] As an embodiment, the method for setting the weight corresponding to each score category for each score category can also be to directly set the optional value range of the score. For example, for the score of the dialogue subject, a relatively large optional value range can be set, such as 0 - 20 points, which is equivalent to increasing the weight corresponding to this score category.
[0152] Thus, the description of step 203 ends, and step 204 is executed below.
[0153] Step 204, determine whether to perform rejection recognition on the voice input information corresponding to the text input information according to the comprehensive score.
[0154] In this embodiment, the specific method for determining whether to perform rejection recognition on the voice input information corresponding to the text input information according to the comprehensive score may include:
[0155] If the comprehensive score is greater than or equal to a preset score threshold, it is determined to recognize the text input information to respond to the voice input information corresponding to the text input information;
[0156] If the comprehensive score is less than the preset score threshold, it is determined to reject the recognition of the text input information, so as to reject the response to the voice input information corresponding to the text input information.
[0157] In this embodiment, a comprehensive score threshold can be preset. If the comprehensive score is greater than or equal to the score threshold, it indicates that the text input information is likely to need to be recognized and responded to; if the comprehensive score is less than the score threshold, it indicates that the text input information is likely to be information that needs to be rejected for recognition, and then the response to the voice input information corresponding to the text input information can be rejected.
[0158] So far, the description of step 204 ends.
[0159] As an embodiment, before inputting the text input information into the trained small language model, the method further includes:
[0160] Determine whether the text input information conforms to the preset rules. Conforming to the preset rules includes: the text input information matches the reference information included in the stored text information library, and the reference information refers to the input information that needs to be responded to; or, the text input information meets the structure defined by the preset regular expression, and the structure defined by the regular expression refers to the structure of the text input information that needs to be responded to;
[0161] If so, it is determined to recognize the text input information to respond to the voice input information corresponding to the text input information;
[0162] If not, continue to execute the step of inputting the text input information into the trained small language model.
[0163] In this embodiment, a text information library can be pre-configured. The text information library stores multiple pieces of reference information, that is, the input information that needs to be responded to. The reference information can be the common human-computer interaction service information in the services supported by the electronic device that collects the voice input information.
[0164] Before inputting the text input information into the small language model, if it is determined that the text input information matches any reference information in the text information library, it indicates that the text input information is likely to be the input information that needs to be responded to. At this time, there is no need to perform subsequent judgments of the small language model and the large language model, and the text input information can be recognized to respond to the voice input information corresponding to the text input information.
[0165] Similarly, the structure of the text input information that needs to be responded to can also be defined by a preset regular expression. For the text input information that meets the regular expression, there is no need to perform subsequent judgments of the small language model and the large language model, and the text input information can be recognized to respond to the voice input information corresponding to the text input information.
[0166] The training methods of the small language model and the large language model proposed in this application will be briefly described below. It should be noted that the training methods for the small language model and the large language model are the processes of fine-tuning the pre-trained small language model and the pre-trained large language model to obtain the trained small language model and the trained large language model.
[0167] Please refer to Figure 3 , Figure 3 which is a schematic diagram of small language model training provided by an embodiment of this application.
[0168] As Figure 3 shown, the training method of the small language model may include:
[0169] Collect sample data, where the sample data includes at least one of network data, business data, and generated data. The network data includes an open-source unsupervised data set collected from an online platform. The business data includes positive sample data of human-machine conversations in a specified business. The generated data includes negative sample data generated according to the large language model or negative sample data obtained by processing the business data according to specified rules;
[0170] Annotate each sample data to obtain the annotated sample data; the annotated sample data includes a first label and a second label. The first label is used to identify whether the sample data has semantics, and the second label is used to identify whether the sample data is a human-machine conversation;
[0171] Input the annotated sample data into a prediction model to obtain a sample feature vector; the sample feature vector is used to represent the overall feature of the text input information;
[0172] Input the sample feature vector into a first classifier, and train the first classifier according to the sample feature vector and the first label corresponding to the sample data;
[0173] For each sample data, if it is determined that the first label corresponding to the sample data indicates that the sample data has semantics, input the sample feature vector corresponding to the sample data into a second classifier, and train the second classifier according to the sample feature vector and the second label corresponding to the sample data.
[0174] In this embodiment, various sample data can be collected in advance. Among them, the data sources mainly include network collection, model generation, and business data. For network collection, open-source unsupervised datasets can be used as the basis, and the data sources include but are not limited to news websites, open-source platforms, etc., ensuring that the collected data is diverse and covers different fields, topics, and language styles. The business data part, as instruction data, can be the common human-computer interaction business information in the business supported by the electronic device that collects voice input information. The business data is the absolute positive sample of human-computer dialogue. Negative samples can be generated by models and rules. By setting negative sample-related prompt words in the large language model, a batch of negative samples can be obtained, or the positive samples can be processed in a specified manner (such as changing semantics, deleting some sentences, etc.) to obtain a batch of negative samples.
[0175] After completing the collection of the data, the data can be labeled. The labeling content includes: the first label and the second label. Among them, the first label is used to identify whether the sample data has semantics, and the second label is used to identify whether the sample data is a human-computer dialogue.
[0176] It should be noted that since the above-mentioned business data is a definite positive sample and the generated data is a definite negative sample, these two types of data do not require manual labeling and can be automatically and uniformly labeled according to the first label and the second label corresponding to the positive samples and negative samples.
[0177] In this embodiment, when training the small language model (i.e., fine-tuning the pre-trained small language model), first input the sample data. After word embedding by the completed training prediction model BERT, the text is converted into a vector representation with a text length of, for example, n * 768 dimensions. Take the first vector [CLS] as the vector representing the overall feature of the sample data (sample feature vector). Further, use the cross-entropy loss function and the Adam optimizer to train the classifier 1 (the first classifier). When the label1 (the first label) indicates that there is semantics, the classifier 2 (the second classifier) will continue to be trained. The classifier 2 is only used to judge the dialogue subject.
[0178] In this embodiment, decouple the training objectives of the two classifiers. Compared with training two models separately, the method of using one model with two classifiers can reduce the deployment difficulty and achieve the effect of parameter sharing, and the accuracies of the two tasks are more balanced.
[0179] In this embodiment, the cross-entropy loss function is as follows:
[0180]
[0181] Among them, p(x) is the calibration value, q(x) is the predicted value, and Loss(p, q) is the loss value.
[0182] The overall loss function of multi-task learning is as follows:
[0183] Loss total = Loss 分类器1 + Loss 分类器2
[0184] Among them, Loss total is the overall loss value of the small language model, Loss 分类器1 is the loss value of the first classifier, Loss 分类器2 is the loss value of the second classifier. When label1 (the first label) is 0, Loss 分类器2 is 0, and the loss values of the first classifier and the second classifier can be obtained through the above cross-entropy loss function.
[0185] So far, the description of the training method of the small language model ends.
[0186] Next, the training method of the large language model will be described.
[0187] Please refer to Figure 4 , Figure 4 which is the training schematic diagram of the large language model provided by the embodiment of the present application.
[0188] As Figure 4 shown, the training method of the large language model may include:
[0189] Collect sample data, where the sample data includes at least one of network data, business data, and generated data. The network data includes an open-source unsupervised data set collected from an online platform, the business data includes human-machine dialogue positive sample data in a specified business, and the generated data includes negative sample data generated according to the large language model or negative sample data obtained by processing the business data according to specified rules;
[0190] Annotate each sample data to obtain the annotated sample data; the annotated sample data includes a third label and the score value corresponding to each classification category obtained. The third label is used to identify whether to reject the recognition of the sample data;
[0191] Input the annotated sample data into the large language model, and train the large language model according to the sample data, the score value corresponding to each classification category to which the sample data is annotated, and the third label corresponding to the sample data.
[0192] In this embodiment, a suitable open-source large language model can be selected, such as Qwen2.5, etc., and the data annotated with the semantic smoothness degree, dialogue subject, and business instruction score is sorted into a format suitable for model input.
[0193] In this embodiment, when annotating sample data, for the annotation of the semantic fluency score, the corresponding score can be marked, such as 0 - 10 points; for the annotation of the dialogue subject score, it can be marked whether the dialogue subject is a human - machine dialogue or not. In application, a score will be obtained according to the trained model, such as 0 - 10 points. The higher the score, the greater the possibility of it being a human - machine dialogue; for the annotation of the business instruction score, it can be "yes" or "no". In application, a score will be obtained according to the trained model, such as 0 - 10 points. The higher the score, the greater the possibility of it being a business instruction supported by the electronic device.
[0194] As an example, for a piece of text "Please query tomorrow's weather", a possible annotation input is [Text: Please query tomorrow's weather, Semantic fluency: 8 points, Dialogue subject: Human - machine dialogue, Business instruction: Yes], and the output is [Not rejected].
[0195] The training process (i.e., the process of fine - tuning the pre - trained large - language model) can be trained using the cross - entropy loss function and the Adam optimizer. The cross - entropy loss function can efficiently measure the difference between the model's prediction result and the true annotation. By minimizing this loss value, the model is prompted to continuously optimize its prediction ability. The Adam optimizer, with its characteristic of adaptive learning rate adjustment, can dynamically allocate an appropriate learning rate for each parameter during the training process, not only accelerating the model's convergence speed but also effectively avoiding the training process falling into a local optimum.
[0196] It should be noted that the hot - word score, context score, and high - frequency word score can also be annotated according to actual needs. The descriptions of various scores have been detailed above and will not be elaborated here.
[0197] Finally, the comprehensive score can be determined based on the various scores obtained by the large - language model, and further, whether to reject the speech input information corresponding to the text input information can be determined according to the comprehensive score.
[0198] So far, the description of Figure 4 ends.
[0199] So far, the description of Figure 2 ends.
[0200] This application initially identifies the text input information converted from the voice input information based on a trained small language model. When the recognition result of the small language model indicates that the text input information has no semantic information and / or the text input information is not a human-machine conversation, the trained large language model is further used to identify the text input information to obtain the comprehensive score of the text input information. Finally, based on the comprehensive score, it is determined whether to recognize or reject the text input information. Through two confirmations by two models, the situation of misrecognition caused by non-human-machine conversations or no semantic information is avoided.
[0201] At the same time, as described above, in the method proposed in this application, different modules are decoupled. When optimization is required, different modules can be modified separately according to business requirements or model capabilities, greatly improving the application flexibility of this solution.
[0202] In addition, in the solution proposed in this application, the difficulty of obtaining sample data is relatively low. Only partial annotation needs to be completed manually (that is, for the annotation of network data, generated data can be automatically annotated as negative samples, and business data can be automatically annotated as positive samples), greatly reducing the workload of manual annotation.
[0203] Next, Figure 5 a rejection method proposed in this application will be described as a whole.
[0204] Please refer to Figure 5 , Figure 5 which is a schematic diagram of the overall process of the rejection method provided by the embodiments of this application.
[0205] As Figure 5 shown, the overall process of this rejection method includes:
[0206] 1) User input: Receive the voice input information input by the user.
[0207] 2) Format conversion: Convert the language input information input by the user into text input information.
[0208] 3) Rule judgment: First, perform rule judgment on the text input information. If the text input information hits a preset rule, then directly obtain the result of "not rejecting", that is, it is considered that this input needs to be responded to and there is no need to further judge through the model. At this time, the text input information can be directly recognized to respond to the voice input information corresponding to the text input information.
[0209] 4) Small language model processing: If the text input information does not match the rules, it is input into the small language model. The small language model classifies the input. When the detection results of classifier 1 (the first classifier) and classifier 2 (the second classifier) both exceed the preset threshold (i.e., the first recognition result indicates semantic meaning and is a human-machine conversation), it can be determined as "not rejected", indicating that the small language model believes that this input needs to be responded to; otherwise (i.e., the first recognition result indicates no semantic meaning and / or is not a human-machine conversation), the text input information is passed to the large language model for processing.
[0210] 4) Large language model processing: The large language model receives the text input information and comprehensively judges by combining various score information such as semantic smoothness score, dialogue subject score, instruction score, context score, high-frequency word score, and hot word score to obtain a comprehensive score. If the comprehensive score is higher than the preset score threshold, the system outputs "not rejected", indicating that the large language model believes that this text input information needs to be responded to; if the comprehensive score is lower than the preset score threshold, it outputs "rejected", that is, the system refuses to process this text input information.
[0211] Thus far, the description of Figure 5 ends.
[0212] Please refer to Figure 6 , Figure 6 which is the structure diagram of the rejection recognition device proposed in the embodiment of this application. As Figure 6 shown, this device is applied to a terminal device, and the terminal device is used for voice interaction with the user. This device may include an acquisition unit 601, an input unit 602, a determination unit 603, and a processing unit 604. Specifically, this device includes:
[0213] The acquisition unit 601 is used to acquire text input information, and the text input information is obtained by converting the collected voice input information;
[0214] The input unit 602 is used to input the text input information into the trained small language model to obtain a first recognition result, and the first recognition result is used to indicate whether the text input information has semantic meaning and whether it is a human-machine conversation;
[0215] The determination unit 603 is used to, if the first recognition result is a result to be confirmed, input the text input information into the trained large language model to obtain a comprehensive score corresponding to the text input information; the result to be confirmed means that the text input information has no semantic meaning and / or is not a human-machine conversation;
[0216] The processing unit 604 is used to determine whether to perform rejection recognition processing on the voice input information corresponding to the text input information according to the comprehensive score.
[0217] Optionally, before inputting the text input information into the trained small language model, the input unit 602 is further configured to:
[0218] Determine whether the text input information conforms to a preset rule. Conforming to the preset rule includes: the text input information matches the reference information included in the stored text information library, where the reference information refers to the input information that needs to be responded to; or, the text input information satisfies the structure defined by a preset regular expression, where the structure defined by the regular expression refers to the structure of the text input information that needs to be responded to;
[0219] If so, determine to recognize the text input information to respond to the corresponding voice input information of the text input information;
[0220] If not, continue to execute the step of inputting the text input information into the trained small language model.
[0221] Optionally, the small language model includes a prediction model, a first classifier, and a second classifier. The input unit 602 is specifically configured to:
[0222] Input the text input information into the prediction model to obtain a text feature vector, where the text feature vector is used to characterize the overall feature of the text input information;
[0223] Process the text feature vector through the first classifier to obtain a first classification result corresponding to the text input information, where the first classification result includes having semantics or no semantics;
[0224] Process the text feature vector through the second classifier to obtain a second classification result corresponding to the text input information, where the second classification result includes human-machine dialogue or not human-machine dialogue;
[0225] Use the first classification result and the second classification result corresponding to the text input information as the first recognition result.
[0226] Optionally, the large language model is configured with at least one classification category. The determination unit 603 is specifically configured to:
[0227] Process the text input information based on the large language model to obtain a score value corresponding to each classification category of the text input information; the classification categories include at least one of the following: semantic smoothness score, dialogue subject score, business instruction score, where the dialogue subject score is used to represent the probability that the text input information belongs to a human-machine dialogue, and the business instruction score is used to represent the probability that the text input information is a business instruction corresponding to the business supported by the electronic device that collects the text input information;
[0228] Determine the comprehensive score corresponding to the text input information according to the score value corresponding to each classification category.
[0229] Optionally, the classification categories further include at least one of the following: hot word score, previous context score, and high-frequency word score;
[0230] Among them, the hot word score refers to an additional score for the text input information when it contains predefined hot words. A hot word is a vocabulary related to the business instruction and used to represent the instruction object; the previous context score is used to indicate the degree of relevance between the text input information and the previous human-machine dialogue information; the high-frequency word score refers to an additional score for the text input information when it contains predefined high-frequency words. A high-frequency word is a vocabulary used to represent actions and with a frequency greater than a preset threshold during human-machine dialogue.
[0231] Optionally, the processing unit 604 is specifically configured to:
[0232] If the comprehensive score is greater than or equal to the preset score threshold, it is determined to recognize the text input information to respond to the corresponding voice input information of the text input information;
[0233] If the comprehensive score is less than the preset score threshold, it is determined to perform a rejection recognition process on the text input information to reject responding to the corresponding voice input information of the text input information.
[0234] Optionally, the training method of the small language model includes:
[0235] Collect sample data, where the sample data includes at least one of network data, business data, and generated data. The network data includes an open-source unsupervised data set collected from an online platform, the business data includes positive human-machine dialogue sample data in a specified business, and the generated data includes negative sample data generated according to a large language model or negative sample data obtained by processing business data according to specified rules;
[0236] Annotate each sample data to obtain the annotated sample data; the annotated sample data includes a first label and a second label. The first label is used to identify whether the sample data has semantics, and the second label is used to identify whether the sample data is a human-machine dialogue;
[0237] Input the annotated sample data into a prediction model to obtain a sample feature vector; the sample feature vector is used to characterize the overall feature of the text input information;
[0238] Input the sample feature vector into a first classifier and train the first classifier according to the sample feature vector and the first label corresponding to the sample data;
[0239] For each sample data, if it is determined that the first label corresponding to the sample data indicates that the sample data has semantics, the sample feature vector corresponding to the sample data is input into the second classifier, and the second classifier is trained according to the sample feature vector and the second label corresponding to the sample data.
[0240] Optionally, the training method of the large language model includes:
[0241] Collect sample data, where the sample data includes at least one of network data, business data, and generated data. The network data includes an open-source unsupervised data set collected from an online platform, the business data includes human-machine dialogue positive sample data in a specified business, and the generated data includes negative sample data generated according to the large language model or negative sample data obtained by processing the business data according to specified rules;
[0242] Annotate each sample data to obtain the annotated sample data; the annotated sample data includes a third label and a score value corresponding to each classification category obtained, and the third label is used to indicate whether to reject the recognition of the sample data;
[0243] Input the annotated sample data into the large language model, and train the large language model according to the sample data, the score value corresponding to each classification category obtained by which the sample data is annotated, and the third label corresponding to the sample data.
[0244] So far, the description of the voice information rejection recognition device in Figure 6 is completed.
[0245] This application embodiment also provides Figure 6 a description of the hardware structure of the device shown. The hardware structure is the structure in the electronic device shown in Figure 7 Please refer to Figure 7 , Figure 7 which is the electronic device structure diagram provided by this application embodiment. As shown in Figure 7 , the hardware structure may include: a processor and a machine-readable storage medium, and the machine-readable storage medium stores machine-executable instructions that can be executed by the processor; the processor is used to execute the machine-executable instructions to implement the method disclosed in the above examples of this application.
[0246] Based on the same application concept as the above method, this application embodiment also provides a machine-readable storage medium, on which several computer instructions are stored. When the computer instructions are executed by the processor, the method disclosed in the above examples of this application can be implemented.
[0247] Exemplarily, the above machine-readable storage medium can be any electronic, magnetic, optical or other physical storage device that can contain or store information such as executable instructions, data, and the like. For example, the machine-readable storage medium can be: RAM (Random Access Memory), volatile memory, non-volatile memory, flash memory, storage drives (such as hard disk drives), solid-state drives, any type of storage disk (such as optical discs, DVDs, etc.), or similar storage media, or a combination thereof.
[0248] The above are only the preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of protection of the present application.
Claims
1. A method for rejecting voice information, characterized in that: The method is applied to a terminal device, the terminal device is used to perform voice interaction with a user, and the method includes: Obtaining text input information, wherein the text input information is converted based on the collected voice input information; Inputting the text input information into the trained small language model to obtain a first recognition result, where the first recognition result is used to indicate whether the text input information has semantic meaning and whether it is a human-computer dialogue; If the first recognition result is a pending confirmation result, the text input information is input into the trained large language model to obtain a comprehensive score corresponding to the text input information; the pending confirmation result means that the text input information has no semantics and / or is not a human-computer dialogue; Determine whether to reject the speech input information corresponding to the text input information according to the comprehensive score.
2. The method according to claim 1, characterized in that Before inputting the text input information into the trained small language model, the method further includes: Determining whether the text input information complies with a preset rule, where the compliance with the preset rule includes: the text input information matches the reference information included in the stored text information library, where the reference information refers to the input information that needs to be responded to; or the text input information satisfies the structure defined by a preset regular expression, where the structure defined by the regular expression refers to the structure of the text input information that needs to be responded to; If yes, determining to recognize the text input information in response to the voice input information corresponding to the text input information; If not, continue to execute the step of inputting the text input information into the trained small language model.
3. The method according to claim 1, characterized in that The small language model includes a prediction model, a first classifier and a second classifier, and the text input information is input into the trained small language model to obtain a first recognition result, including: Inputting the text input information into the prediction model to obtain a text feature vector, wherein the text feature vector is used to characterize the overall features of the text input information; Processing the text feature vector by the first classifier to obtain a first classification result corresponding to the text input information, wherein the first classification result includes semantics or no semantics; Processing the text feature vector by the second classifier to obtain a second classification result corresponding to the text input information, wherein the second classification result includes human-computer dialogue or non-human-computer dialogue; The first classification result and the second classification result corresponding to the text input information are used as the first recognition result.
4. The method according to claim 1, characterized in that The large language model is configured with at least one score category, and the inputting the text input information into the trained large language model to obtain a comprehensive score corresponding to the text input information includes: The text input information is processed based on the large language model to obtain a score value corresponding to each score category corresponding to the text input information; the score category includes at least one of the following: a semantic fluency score, a dialogue subject score, and a business instruction score, wherein the dialogue subject score is used to indicate the probability that the text input information belongs to a human-computer dialogue, and the business instruction score is used to indicate the probability that the text input information is a business instruction corresponding to a business supported by the electronic device that collects the text input information; A comprehensive score corresponding to the text input information is determined according to the score value corresponding to each score category.
5. The method according to claim 4, characterized in that The score category also includes at least one of the following: a hot word score, a previous context score, and a high-frequency word score; Among them, the hot word score refers to the additional score of the text input information when the text input information contains a pre-defined hot word, and the hot word refers to a word related to the business instruction and used to represent the instruction object; the above score is used to indicate the degree of relevance between the text input information and the previous human-computer dialogue information; the high-frequency word score refers to the additional score of the text input information when the text input information contains a pre-defined high-frequency word, and the high-frequency word refers to a word used to represent an action whose frequency of appearance in human-computer dialogue is greater than a preset threshold.
6. The method according to claim 1, characterized in that The determining, according to the comprehensive score, whether to reject the speech input information corresponding to the text input information comprises: If the comprehensive score is greater than or equal to a preset score threshold, determining to recognize the text input information in response to the voice input information corresponding to the text input information; If the comprehensive score is less than a preset score threshold, it is determined to reject the text input information so as to refuse to respond to the voice input information corresponding to the text input information.
7. The method according to claim 3, characterized in that The training method of the small language model includes: Collecting sample data, the sample data including at least one of network data, business data, and generated data, the network data including an open source unsupervised data set collected from an online platform, the business data including positive sample data of human-computer dialogue in a specified business, and the generated data including negative sample data generated according to a large language model or negative sample data obtained by processing the business data according to specified rules; Labeling each sample data to obtain labeled sample data; the labeled sample data includes a first label and a second label, the first label is used to identify whether the sample data has semantics, and the second label is used to identify whether the sample data is a human-computer dialogue; Inputting the labeled sample data into the prediction model to obtain a sample feature vector; the sample feature vector is used to characterize the overall features of the text input information; Inputting the sample feature vector into the first classifier, and training the first classifier according to the sample feature vector and the first label corresponding to the sample data; For each sample data, if it is determined that the first label corresponding to the sample data indicates that the sample data has semantics, the sample feature vector corresponding to the sample data is input into the second classifier, and the second classifier is trained according to the sample feature vector and the second label corresponding to the sample data.
8. The method according to claim 5, characterized in that The training method of the large language model includes: Collecting sample data, the sample data including at least one of network data, business data, and generated data, the network data including an open source unsupervised data set collected from an online platform, the business data including positive sample data of human-computer dialogue in a specified business, and the generated data including negative sample data generated according to a large language model or negative sample data obtained by processing the business data according to specified rules; Labeling each sample data to obtain labeled sample data; the labeled sample data includes a third label and a score value corresponding to each score category, and the third label is used to identify whether to perform rejection processing on the sample data; The labeled sample data is input into the large language model, and the large language model is trained according to the sample data, the score value corresponding to each score category labeled with the sample data, and the third label corresponding to the sample data.
9. A voice information rejection device, characterized in that: The device is applied to a terminal device, and the terminal device is used to perform voice interaction with a user. The device includes: An obtaining unit, configured to obtain text input information, wherein the text input information is converted based on the collected voice input information; An input unit, used for inputting the text input information into the trained small language model to obtain a first recognition result, wherein the first recognition result is used to indicate whether the text input information has semantics and whether it is a human-computer dialogue; a determination unit, configured to input the text input information into a large language model that has been trained to obtain a comprehensive score corresponding to the text input information if the first recognition result is a pending confirmation result; the pending confirmation result means that the text input information has no semantics and / or is not a human-computer dialogue; A processing unit is used to determine whether to reject the voice input information corresponding to the text input information according to the comprehensive score.
10. An electronic device, characterized in that: The electronic device comprises: a processor and a machine-readable storage medium; the machine-readable storage medium stores machine-executable instructions that can be executed by the processor; the processor is used to execute the machine-executable instructions to implement the method described in any one of claims 1 to 8.