Voice interaction method, server and computer readable storage medium
Through natural language processing and similarity calculation, the similarity between the voice request and the target text is determined, which solves the problem of ignoring semantic information in the prior art, and achieves a more accurate and reliable ignoring result, which improves the robustness and naturalness of voice interaction.
Patent Information
- Application Number
- CN202510332245.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art may ignore the semantic information contained in the audio mode when making a refusal judgment based on audio, resulting in errors in refusal judgments and affecting the voice interaction effect between humans and machines.
By obtaining the current voice request, natural language processing is performed to determine the target text, calculate the similarity between the current voice request and the target text, and determine the rejection result based on the similarity degree to achieve multi-dimensional rejection result determination.
This method can ensure the accuracy and reliability of the recognised results to a certain extent, ensure the stability and nature of voice interaction, thereby improving the voice interaction between users and machines.
Smart Images

Figure CN120015018A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of voice interaction technology, and in particular to a voice interaction method, a server, and a computer-readable storage medium. Background Art
[0002] In related technologies, rejection judgment is usually performed on audio, such as determining whether there is a human voice in the audio through voice activity detection to determine whether to reject the audio. However, when performing rejection judgment based on audio, semantic information contained in the audio modality may be ignored, which may lead to errors in rejection judgment, affecting the voice interaction effect between humans and machines to a certain extent. Summary of the invention
[0003] The present application provides a voice interaction method, a server and a computer-readable storage medium.
[0004] A voice interaction method provided in an embodiment of the present application includes:
[0005] Get the current voice request;
[0006] Performing natural language processing on the current voice request to determine a target text corresponding to the current voice request;
[0007] Determining a first similarity between the current voice request and the target text according to the current voice request and the target text;
[0008] Determining a rejection result of the current voice request according to the current voice request, the target text, and the first similarity level;
[0009] The voice interaction is performed according to the rejection result.
[0010] Thus, in the implementation of the present application, the rejection result of the current voice request can be determined based on the current voice request itself and the target text corresponding to the current voice request, as well as the first similarity between the current voice request and the corresponding target text, thereby achieving multi-dimensional rejection result determination, and thus the accuracy and reliability of the rejection result of the current voice request can be guaranteed to a certain extent, and therefore, the robustness and naturalness of the voice interaction based on the rejection result can be guaranteed, thereby guaranteeing the voice interaction effect between the user and the vehicle and other machines. And, compared to the method of determining the rejection result of the current voice request by the current voice request itself or the text corresponding to the current voice request, the implementation of the present application is based on the first similarity between the current voice request and the target text corresponding to the current voice request, so that in the process of determining the rejection result of the current voice request, the mapping relationship between the current voice request and the corresponding target text can be taken into account, so the effectiveness and reliability of the rejection result can be further guaranteed.
[0011] In certain embodiments of the present application, determining the first similarity between the current voice request and the target text according to the current voice request and the target text includes:
[0012] Encoding the current voice request and the target text respectively to determine first encoding information of the current voice request and second encoding information of the target text;
[0013] The first similarity level is determined according to the first encoding information and the second encoding information.
[0014] Thus, in an embodiment of the present application, the current voice request and the target text may be encoded respectively to determine the first encoding information of the current voice request and the second encoding information of the target text, and the similarity between the current voice request and the corresponding target text may be determined based on the first encoding information and the second encoding information, so that the first similarity degree may be determined based on the first encoding information determined by encoding the current voice request and the second encoding information determined by encoding the target text. Compared with directly calculating the similarity between the current voice request and the target text to determine the first similarity degree, the determination of the first similarity degree can be completed relatively quickly, thereby ensuring efficient determination of the rejection result, and further ensuring timely and natural voice interaction.
[0015] In certain embodiments of the present application, determining the rejection result of the current voice request according to the current voice request, the target text and the first similarity level includes:
[0016] When the first similarity level is less than or equal to a first preset similarity threshold, the rejection result is determined to be a rejection.
[0017] Thus, in an embodiment of the present application, when the first degree of similarity between the current voice request and the corresponding target text is less than or equal to the first preset similarity threshold, the rejection result of the current voice request can be determined as rejection, thereby achieving the determination of the rejection result of the current voice request.
[0018] In certain embodiments of the present application, determining the rejection result of the current voice request according to the current voice request, the target text and the first similarity level includes:
[0019] When the first similarity level is greater than a second preset similarity threshold, the rejection result is determined according to the current voice request and / or the target text.
[0020] Thus, in an embodiment of the present application, when the first degree of similarity between the current voice request and the corresponding target text is greater than a second preset similarity threshold, the rejection result of the current voice request can be determined based on the current voice request and / or the target text corresponding to the current voice request, thereby ensuring efficient determination of the rejection result of the current voice request.
[0021] In certain embodiments of the present application, when the first similarity is greater than a second preset similarity threshold, determining the rejection result according to the current voice request and / or the target text includes:
[0022] When the first similarity level is greater than a second preset similarity threshold, the rejection result is determined according to the audio features of the current voice request and / or the semantics of the target text.
[0023] Thus, in an embodiment of the present application, when the first degree of similarity between the current voice request and the corresponding target text is greater than the second preset similarity threshold, the rejection result of the current voice request can be determined based on the audio features of the current voice request and / or the semantics of the target text, thereby ensuring the validity and reliability of the rejection result of the current voice request.
[0024] In certain embodiments of the present application, determining the rejection result of the current voice request according to the current voice request, the target text and the first similarity level includes:
[0025] Based on a pre-trained natural language processing model, the rejection result is determined according to the current voice request, the target text and the first similarity level.
[0026] In this way, in the implementation mode of the present application, the rejection result of the current voice request can be determined based on the pre-trained natural language processing model according to the current voice request, the target text, and the current voice request and the corresponding target text, so that the rejection result of the current voice request can be determined based on the natural language processing model, thereby ensuring the effectiveness and reliability of the rejection result of the current voice request.
[0027] In certain embodiments of the present application, the training step of the natural language processing model includes:
[0028] Obtaining a voice request sample and a text label of the voice request sample;
[0029] Determining a second similarity between the voice request sample and the text label according to the voice request sample and the text label;
[0030] Determining a rejection prediction result according to the voice request sample, the text label and the second similarity level;
[0031] Model training is performed according to the rejection prediction result to determine the natural language processing model.
[0032] In this way, in the implementation mode of the present application, a voice request sample and a text label of the voice request sample can be obtained, and based on the voice request sample and the text label, a second similarity degree between the voice request sample and the text label can be determined, and based on the voice request sample, the text label and the second similarity degree, a rejection prediction result can be determined, and model training can be performed based on the rejection prediction result to determine the natural language processing model, thereby realizing the training of the natural language processing model.
[0033] In some embodiments of the present application, the voice request sample includes a plurality of samples, and determining the rejection prediction result according to the voice request sample, the text label and the second similarity level includes:
[0034] Determining a plurality of the rejection prediction results of the voice request sample according to the second similarity between the voice request sample and the text label of each of the voice request samples;
[0035] The performing model training according to the rejection prediction result to determine the natural language processing model includes:
[0036] Model training is performed according to the multiple rejection prediction results of each of the voice request samples to determine the natural language processing model.
[0037] Thus, in the implementation mode of the present application, multiple rejection prediction results of a voice request sample can be determined according to the second similarity between a voice request sample and the text label of each voice request sample, and model training can be performed according to the multiple rejection prediction results of each voice request sample to determine the natural language processing model, so that the difficulty of model training is increased, thereby ensuring sufficient training of the model and ensuring model performance.
[0038] An embodiment of the present application provides a server, including a memory and a processor, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the above-mentioned voice interaction method is implemented.
[0039] An embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by one or more processors, the above-mentioned voice interaction method is implemented.
[0040] The server and computer-readable storage medium provided by the embodiments of the present application can determine the rejection result of the current voice request based on the current voice request itself and the target text corresponding to the current voice request, as well as the first similarity between the current voice request and the corresponding target text, thereby realizing multi-dimensional rejection result determination, and thus can guarantee the accuracy and reliability of the rejection result of the current voice request to a certain extent, and therefore can guarantee the robustness and naturalness of the voice interaction based on the rejection result, thereby guaranteeing the voice interaction effect between the user and the vehicle and other machines. And, compared with the method of determining the rejection result of the current voice request by the current voice request itself or the text corresponding to the current voice request, the embodiments of the present application are based on the first similarity between the current voice request and the target text corresponding to the current voice request, so that in the process of determining the rejection result of the current voice request, the mapping relationship between the current voice request and the corresponding target text can be taken into account, so the effectiveness and reliability of the rejection result can be further guaranteed.
[0041] Additional aspects and advantages of the embodiments of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through the practice of the embodiments of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the description of the embodiments in conjunction with the following drawings, in which:
[0043] Figure 1 A flowchart of a voice interaction method in certain implementation modes of the present application;
[0044] Figure 2 A flowchart of a voice interaction method in certain implementation modes of the present application;
[0045] Figure 3 A schematic diagram of application scenarios in certain embodiments of the present application;
[0046] Figure 4 A flowchart of a voice interaction method in certain implementation modes of the present application;
[0047] Figure 5 A schematic diagram of an application scenario in certain embodiments of the present application. DETAILED DESCRIPTION
[0048] The embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals represent the same or similar elements or elements having the same or similar functions from beginning to end. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the embodiments of the present application, and cannot be understood as limiting the embodiments of the present application.
[0049] In the relevant technology, Rejection may refer to a situation where the dialogue system cannot understand or refuses to respond to the user's voice input. For example, rejection may be triggered when the user's voice content exceeds the system's preset capabilities (including but not limited to the user asking unsupported commands, questions beyond the knowledge base, etc.). For another example, rejection may be triggered when the user's pronunciation is unclear, the sentence is incomplete, or contains dialects / accents, resulting in a failure of voice recognition. For another example, rejection may be triggered when the background noise interference is severe and the effective voice signal cannot be extracted.
[0050] In order to ensure the correct handling of rejection, the rejection recognition process in related technologies is usually "first perform rejection judgment on audio, and then perform rejection judgment on text". However, in this process, only audio modal information is extracted when performing rejection judgment on audio, and only text modal information is extracted when performing rejection judgment on text, but the audio modality contains acoustic, semantic and other information, and the text modality contains semantic and other information. Also, if rejection has been determined when performing rejection judgment on audio, rejection judgment will no longer be performed on text. In other words, if the rejection judgment result on audio is incorrect, the audio cannot be retrieved through subsequent rejection judgment on text. Also, before performing rejection judgment on text, the upstream ASR (Automatic Speech Recognition) module is usually required to recognize the audio signal to obtain the text to be rejected. In other words, the recognition accuracy of the ASR module may affect the rejection judgment result on text.
[0051] Based on the above problems you may encounter, please refer to Figure 1 , the present application embodiment provides a voice interaction method, including:
[0052] 01: Get the current voice request;
[0053] 02: Perform natural language processing on the current voice request to determine the target text corresponding to the current voice request;
[0054] 03: Determine a first similarity between the current voice request and the target text according to the current voice request and the target text;
[0055] 04: Determine a rejection result of the current voice request according to the current voice request, the target text and the first similarity level;
[0056] 05: Perform voice interaction based on the rejection result.
[0057] The embodiment of the present application provides a voice interaction device. The voice interaction method of the embodiment of the present application can be implemented by the voice interaction device of the embodiment of the present application. Specifically, the position determination device includes an acquisition module, a processing module, a first similarity determination module, a rejection result determination module and an interaction module. Among them, the acquisition module is used to obtain the current voice request. The processing module is used to perform natural language processing on the current voice request to determine the target text corresponding to the current voice request. The first similarity determination module is used to determine the first similarity between the current voice request and the target text based on the current voice request and the target text. The rejection result determination module is used to determine the rejection result of the current voice request based on the current voice request, the target text and the first similarity. The interaction module is used to perform voice interaction based on the rejection result.
[0058] The embodiment of the present application also provides a server, and the server includes a memory and a processor. The voice interaction method of the embodiment of the present application can be implemented by the server of the embodiment of the present application. Specifically, a computer program is stored in the memory, and the processor is used to obtain the current voice request, and to perform natural language processing on the current voice request, determine the target text corresponding to the current voice request, and to determine the first similarity between the current voice request and the target text based on the current voice request and the target text, and to determine the rejection result of the current voice request based on the current voice request, the target text and the first similarity, and to perform voice interaction based on the rejection result.
[0059] Specifically, in the related art, the rejection judgment for audio and text is optimized independently using single modal information, and only audio modal information is extracted when rejecting audio, and only text modal information is extracted when rejecting text, but the audio modality contains acoustic, semantic and other information, and the text modality contains semantic and other information. Therefore, the implementation of the present application provides a scheme for combining audio and text modal information to achieve audio-text modality joint modeling rejection, thereby jointly optimizing the rejection effect and reducing the occurrence of error accumulation and the like.
[0060] Specifically, in an embodiment of the present application, when a user in a vehicle cabin wants to use voice control to perform actions such as turning off the air conditioner or opening the window at the current moment, and therefore says sentences such as "turn off the air conditioner" or "open the window", the vehicle can capture the audio signal in the vehicle cabin through audio receiving components such as a microphone, thereby obtaining the current voice request, and forwarding the current voice request to a server that is connected to the vehicle in communication.
[0061] Next, when the server receives the current voice request forwarded by the vehicle, it may perform natural language processing such as ASR (Automatic Speech Recognition) on the current voice request, thereby determining the target text corresponding to the current voice request.
[0062] Next, the server may determine a first similarity between the current voice request and the target text according to the current voice request and the target text.
[0063] Then, the server may determine whether to reject the current voice request, or determine the rejection result corresponding to the current voice request, based on the current voice request, the target text corresponding to the current voice request, and the first similarity between the current voice request and the target text.
[0064] Finally, the server can perform subsequent voice interaction processes based on the rejection result corresponding to the current voice request.
[0065] For example, when the rejection result corresponding to the current voice request is not rejected, the server can perform natural language processing such as slot recognition on the target text of the current voice request, so as to determine the vehicle control instruction corresponding to the current voice request and which can be used to implement the current voice request, and then send the vehicle control instruction to the vehicle so that the vehicle can perform actions such as turning on the air conditioner and closing the windows according to the received vehicle control instruction, thereby realizing voice interaction with the user.
[0066] For example, when the rejection result corresponding to the current voice request is rejection, the server can send a preset rejection processing instruction to the vehicle. When the vehicle receives the rejection processing instruction, it can ignore the current voice request, or provide feedback to the user through voice broadcast, text display, etc., such as "Sorry, I don't understand what you mean, please say it again.", thereby realizing voice interaction with the user.
[0067] Thus, in the implementation of the present application, the rejection result of the current voice request can be determined based on the current voice request itself and the target text corresponding to the current voice request, as well as the first similarity between the current voice request and the corresponding target text, thereby achieving multi-dimensional rejection result determination, and thus the accuracy and reliability of the rejection result of the current voice request can be guaranteed to a certain extent, and therefore, the robustness and naturalness of the voice interaction based on the rejection result can be guaranteed, thereby guaranteeing the voice interaction effect between the user and the vehicle and other machines. And, compared to the method of determining the rejection result of the current voice request by the current voice request itself or the text corresponding to the current voice request, the implementation of the present application is based on the first similarity between the current voice request and the target text corresponding to the current voice request, so that in the process of determining the rejection result of the current voice request, the mapping relationship between the current voice request and the corresponding target text can be taken into account, so the effectiveness and reliability of the rejection result can be further guaranteed.
[0068] In one example, the natural language processing performed on the current voice request may be ASR (Automatic Speech Recognition) processing.
[0069] In one example, the server may call a pre-trained ASR model to perform speech recognition on the current voice request, thereby determining the target text corresponding to the current voice request.
[0070] In one example, if the first similarity between the current voice request and the corresponding target text is too low, it means that the natural language processing of the current voice request is incorrect, resulting in the current voice request and the corresponding target text not being mapped to a spatially close or identical position. In this case, the rejection result of the current voice request can be determined as rejection.
[0071] In one example, since there may be noise in the in-vehicle voice interaction scene, including but not limited to road noise, tire noise, wind noise, etc., when the current voice request sent by the vehicle contains the user's voice and road noise, tire noise, wind noise, etc., the user's voice part may be vague, but when the current voice request is processed by natural language, the target text corresponding to the current voice request is also determined, and the target text in this case is difficult to match the user's voice part, which leads to the first similarity between the current voice request and the corresponding target text being too low, and then the rejection result of the current voice request can be determined as rejection.
[0072] In one example, the current voice request may be first subjected to voice activity detection (VAD). If the voice activity detection fails to detect the presence of human voice in the current voice request, the rejection result of the current voice request is determined to be rejection. Conversely, if the voice activity detection detects the presence of human voice in the current voice request, the target text is detected, such as detecting whether the target text can generate a corresponding vehicle control instruction.
[0073] If a corresponding vehicle control instruction cannot be generated according to the target text, the rejection result of the current voice request is determined to be rejection. On the contrary, if a corresponding vehicle control instruction can be generated according to the target text, the rejection result according to the current voice request is determined according to the first similarity between the current voice request and the corresponding target text to perform a subsequent voice interaction process.
[0074] See also Figure 2 In certain embodiments of the present application, step 03 includes:
[0075] 030: Encode the current voice request and the target text respectively to determine first encoding information of the current voice request and second encoding information of the target text;
[0076] 031: Determine a first similarity level according to the first coding information and the second coding information.
[0077] The first similarity determination module of the implementation mode of the present application is also used to encode the current voice request and the target text respectively, determine the first encoding information of the current voice request and the second encoding information of the target text, and determine the first similarity based on the first encoding information and the second encoding information.
[0078] The processor of the embodiment of the present application is also used to encode the current voice request and the target text respectively, determine the first encoding information of the current voice request and the second encoding information of the target text, and determine the first similarity level based on the first encoding information and the second encoding information.
[0079] Specifically, in order to efficiently determine the first degree of similarity between the current voice request and the corresponding target text, in an embodiment of the present application, the current voice request and the corresponding target text can be encoded separately to determine the first encoding information of the current voice request and the second encoding information of the target text, and then, the first degree of similarity between the first encoding information and the second encoding information can be calculated to determine the first degree of similarity between the current voice request and the corresponding target text.
[0080] It can be understood that, compared to the method of directly calculating the first similarity between the current voice request and the corresponding target text, by respectively encoding the current voice request and the corresponding target text to determine the first encoding information and the second encoding information, and then calculating the first similarity between the first encoding information and the second encoding information to determine the "first similarity between the current voice request and the corresponding target text", the "first similarity between the current voice request and the corresponding target text" can be determined by the encoded current voice request (i.e., the first encoding information) and the encoded target text (i.e., the second encoding information), thereby reducing the difficulty of calculating the first similarity between the voice text to a certain extent, thereby improving the efficiency of determining the first similarity between the voice text.
[0081] In one example, the current voice request may be sequentially subjected to spectrogram conversion and pitch serialization processing, and the processed current voice request may be input into a pre-trained bidirectional transformer structure model for encoding, thereby determining first encoding information.
[0082] In one example, the target text corresponding to the current voice request can be encoded using a pre-trained BERT model based on a bidirectional transformer structure to determine the second encoding information.
[0083] Thus, in an embodiment of the present application, the current voice request and the target text may be encoded respectively to determine the first encoding information of the current voice request and the second encoding information of the target text, and the similarity between the current voice request and the corresponding target text may be determined based on the first encoding information and the second encoding information, so that the first similarity degree may be determined based on the first encoding information determined by encoding the current voice request and the second encoding information determined by encoding the target text. Compared with directly calculating the similarity between the current voice request and the target text to determine the first similarity degree, the determination of the first similarity degree can be completed relatively quickly, thereby ensuring efficient determination of the rejection result, and further ensuring timely and natural voice interaction.
[0084] In certain embodiments of the present application, step 05 includes:
[0085] When the first similarity level is less than or equal to the first preset similarity threshold, the rejection result is determined to be rejection.
[0086] The rejection result determination module of the implementation manner of the present application is further configured to determine that the rejection result is a rejection when the first similarity level is less than or equal to a first preset similarity threshold.
[0087] The processor of the embodiment of the present application is further configured to determine that the rejection result is a rejection when the first similarity level is less than or equal to a first preset similarity threshold.
[0088] Specifically, in an embodiment of the present application, the server may determine the rejection result of the current voice request as rejection when a first similarity between the current voice request and the corresponding target text is less than or equal to a first preset similarity threshold.
[0089] In one example, the first similarity between the current voice request and the corresponding target text is determined based on the cosine similarity calculation formula, and the first preset similarity threshold is 0.2. In other words, when the first similarity between the current voice request and the corresponding target text is less than or equal to 0.2, the rejection result of the current voice request is determined as rejection.
[0090] Thus, in an embodiment of the present application, when the first degree of similarity between the current voice request and the corresponding target text is less than or equal to the first preset similarity threshold, the rejection result of the current voice request can be determined as rejection, thereby achieving the determination of the rejection result of the current voice request.
[0091] In certain embodiments of the present application, step 05 includes:
[0092] When the first similarity level is greater than a second preset similarity threshold, a rejection result is determined according to the current voice request and / or the target text.
[0093] The rejection result determination module of the implementation manner of the present application is also used to determine the rejection result according to the current voice request and / or the target text when the first similarity level is greater than a second preset similarity threshold.
[0094] The processor of the embodiment of the present application is also used to determine a rejection result according to the current voice request and / or the target text when the first similarity level is greater than a second preset similarity threshold.
[0095] Specifically, in an embodiment of the present application, a rejection judgment may be first made using a first degree of similarity between the current voice request and the corresponding target text, and then a rejection judgment may be made based on the current voice request itself and / or the target text corresponding to the current voice request. In other words, a rejection judgment may be first made using a first degree of similarity between the current voice request and the corresponding target text, and when the similarity between the first degree of similarity between the current voice request and the corresponding target text is higher than a second preset similarity threshold, a rejection judgment may be made based on the current voice request itself and / or the target text corresponding to the current voice request.
[0096] In one example, the first degree of similarity between the current voice request and the corresponding target text is determined based on a cosine similarity calculation formula, and further, the second preset similarity threshold is 0.2. In other words, when the first degree of similarity between the current voice request and the corresponding target text is greater than 0.2, the server performs a rejection judgment based on the current voice request itself and / or the target text corresponding to the current voice request.
[0097] Further, in one example, when the first similarity between the current voice request and the corresponding target text is less than or equal to 0.2, the server may determine the rejection result of the current voice request as rejection.
[0098] Thus, in an embodiment of the present application, when the first degree of similarity between the current voice request and the corresponding target text is greater than a second preset similarity threshold, the rejection result of the current voice request can be determined based on the current voice request and / or the target text corresponding to the current voice request, thereby ensuring efficient determination of the rejection result of the current voice request.
[0099] In certain embodiments of the present application, when the first degree of similarity is greater than the second preset similarity threshold, the step of determining a rejection result according to the current voice request and / or the target text includes:
[0100] When the first similarity level is greater than the second preset similarity threshold, a rejection result is determined according to the audio features of the current voice request and / or the semantics of the target text.
[0101] The rejection result determination module of the implementation mode of the present application is also used to determine the rejection result according to the audio features of the current voice request and / or the semantics of the target text when the first similarity is greater than a second preset similarity threshold.
[0102] The processor of the embodiment of the present application is also used to determine a rejection result based on the audio features of the current voice request and / or the semantics of the target text when the first similarity level is greater than a second preset similarity threshold.
[0103] Specifically, in an embodiment of the present application, when the first degree of similarity between the current voice request and the corresponding target text is higher than a second preset similarity threshold, the rejection recognition result of the current voice request can be determined based on the audio features of the current voice request and / or the semantics of the target text.
[0104] In one example, the server may perform detection such as speech clarity (Speech Clarity) detection, signal-to-noise ratio detection, voice activity detection (VAD, Voice Activity Detection), signal-to-noise ratio (Signal-to-Noise Ratio, SNR) detection, etc. on the current voice request, so as to determine the audio characteristics of the current voice request.
[0105] More specifically, taking the signal-to-noise ratio detection as an example, after performing a signal-to-noise ratio detection on the current voice request, if the signal-to-noise ratio of the current voice request is lower than the preset value, it means that the current voice request contains more noise such as road noise, tire noise, etc., and the server can determine that the rejection result of the current voice request is rejection.
[0106] In one example, the server may call a pre-trained natural language processing model to perform natural language processing on the target text corresponding to the current voice request to determine whether the current voice request belongs to fields such as in-vehicle voice interaction, thereby determining a rejection result. For example, if the target text corresponding to the current voice request is used to control the vehicle to perform specific actions such as turning on the air conditioner, closing the windows, etc., then the current voice request belongs to the field of in-vehicle voice interaction, and the rejection result of the current voice request is not rejection. Conversely, if the target text corresponding to the current voice request is not used to control the vehicle to perform specific actions, the server may determine that the rejection result of the current voice request is rejection.
[0107] In one example, the server may determine the final rejection result by combining the audio features of the current voice request and the semantics of the target text of the current voice request. For example, the server may determine that the rejection result of the current voice request is rejection based on the audio features of the current voice request, and that the rejection result of the current voice request is rejection based on the semantics of the target text, and determine that the rejection result of the current voice request is rejection. Conversely, if the server determines that the rejection result of the current voice request is not rejection based on the audio features of the current voice request, or determines that the rejection result of the current voice request is not rejection based on the semantics of the target text, the server determines that the rejection result of the final output is not rejection.
[0108] Thus, in an embodiment of the present application, when the first degree of similarity between the current voice request and the corresponding target text is greater than the second preset similarity threshold, the rejection result of the current voice request can be determined based on the audio features of the current voice request and / or the semantics of the target text, thereby ensuring the validity and reliability of the rejection result of the current voice request.
[0109] To more clearly illustrate the implementation of this application, please refer to Figure 3 , Figure 3 Schematic diagram of application scenarios in certain embodiments of the present application. Specifically, Figure 3As shown, in the embodiment of the present application, for the current voice request and the target text of the current voice request, the current voice request can be encoded by a text encoder to extract audio information, thereby obtaining an audio embedding (i.e., the first encoding information), and the target text can be encoded by a text encoder to extract text information, thereby obtaining a text embedding (i.e., the second encoding information).
[0110] Then, a first similarity calculation is performed on the audio embedding and the text embedding, such as calculating the cosine similarity.
[0111] Then, if the first similarity between the audio embedding and the text embedding is less than or equal to the preset threshold, it is directly rejected. On the contrary, if the first similarity between the audio embedding and the text embedding is greater than the preset threshold, a rejection judgment can be made through the audio rejection module and the current voice request, such as performing VAD detection on the current voice request through the audio rejection module, and judging whether to reject it according to the VAD detection result.
[0112] Furthermore, if the first similarity between the audio embedding and the text embedding is greater than a preset threshold, a rejection judgment can also be made through the text rejection module and the target text of the current voice request, such as judging whether the target text does not fall into the field of in-vehicle voice interaction based on the semantics of the target text to determine whether to reject it.
[0113] In certain embodiments of the present application, step 04 includes:
[0114] Based on the pre-trained natural language processing model, the rejection result is determined according to the current voice request, the target text and the first similarity level.
[0115] The rejection result determination module of the implementation mode of the present application is also used to determine the rejection result based on the pre-trained natural language processing model according to the current voice request, the target text and the first similarity level.
[0116] The processor of the implementation mode of the present application is also used to determine a rejection result based on a pre-trained natural language processing model, according to the current voice request, the target text and the first similarity level.
[0117] Specifically, in order to further ensure the validity of the rejection result of the current voice request, in an implementation manner of the present application, the rejection result of the current voice request can be determined based on a pre-trained natural language processing model.
[0118] Specifically, after obtaining the current voice request, the target text corresponding to the current voice request, and the first similarity between the current voice request and the corresponding target text, the server may input the current voice request, the target text corresponding to the current voice request, and the first similarity between the current voice request and the corresponding target text into a pre-trained natural language processing model. Furthermore, the natural language processing model may perform reasoning based on the current voice request, the target text corresponding to the current voice request, and the first similarity between the current voice request and the corresponding target text, thereby outputting a rejection result of the current voice request.
[0119] In this way, in the implementation mode of the present application, the rejection result of the current voice request can be determined based on the pre-trained natural language processing model according to the current voice request, the target text, and the current voice request and the corresponding target text, so that the rejection result of the current voice request can be determined based on the natural language processing model, thereby ensuring the effectiveness and reliability of the rejection result of the current voice request.
[0120] See also Figure 4 In certain embodiments of the present application, the training step of the natural language processing model includes:
[0121] 06: Obtain voice request samples and text labels of voice request samples;
[0122] 07: Determine a second similarity between the voice request sample and the text label according to the voice request sample and the text label;
[0123] 08: Determine a rejection prediction result based on the voice request sample, the text label and the second similarity level;
[0124] 09: Conduct model training based on the rejection prediction results and determine the natural language processing model.
[0125] The voice interaction device of the embodiment of the present application also includes a sample acquisition module, a second similarity determination module, a prediction module and a training module. Among them, the sample acquisition module is used to obtain a voice request sample and a text label of the voice request sample. The second similarity determination module is used to determine the second similarity between the voice request sample and the text label based on the voice request sample and the text label. The prediction module is used to determine the rejection prediction result based on the voice request sample, the text label and the second similarity. The training module is used to perform model training based on the rejection prediction result and determine the natural language processing model.
[0126] The processor of the embodiment of the present application is also used to obtain a voice request sample and a text label of the voice request sample, and determine a second similarity between the voice request sample and the text label based on the voice request sample and the text label, and determine a rejection prediction result based on the voice request sample, the text label and the second similarity, and perform model training based on the rejection prediction result to determine a natural language processing model.
[0127] Specifically, in the embodiments of the present application, the training of the above-mentioned natural language processing model can be performed using pre-collected voice request samples and text labels of the voice request samples.
[0128] Specifically, the server may first obtain a voice request sample and a text label of the voice request sample. In one example, the server may call a pre-trained ASR model to perform voice recognition on each voice request sample, and determine the text label of each voice request sample based on the voice recognition result of each voice request sample.
[0129] Next, the server may input the voice request sample and the text label of the voice request sample into the model to be trained. Then, the model to be trained determines the second similarity between the voice request sample and the text label of the voice request sample according to the voice request sample and the text label of the voice request sample.
[0130] Next, the model to be trained can make a rejection judgment on the voice request sample based on the voice request sample, the text label of the voice request sample, and the second degree of similarity between the voice request sample and the text label of the voice request sample, thereby determining the rejection prediction result of the voice request sample, that is, whether to reject the voice request sample.
[0131] Finally, the server can update the parameters such as weights and biases in the natural language processing model according to the rejection prediction results output by the model, thereby training the natural language processing model. In one example, the server also obtains the rejection label of the voice request sample, and then the server can calculate the loss function based on the difference between the rejection label of the voice request sample and the rejection prediction result to update the parameters such as weights and biases in the natural language processing model, thereby training the natural language processing model.
[0132] In this way, in the implementation mode of the present application, a voice request sample and a text label of the voice request sample can be obtained, and based on the voice request sample and the text label, a second similarity degree between the voice request sample and the text label can be determined, and based on the voice request sample, the text label and the second similarity degree, a rejection prediction result can be determined, and model training can be performed based on the rejection prediction result to determine the natural language processing model, thereby realizing the training of the natural language processing model.
[0133] In some implementations of the present application, the voice request sample includes multiple ones, and then, step 08 includes:
[0134] Determining a plurality of rejection prediction results for a voice request sample according to a second similarity between a voice request sample and a text label of each voice request sample;
[0135] And, step 09 includes:
[0136] Model training is performed based on multiple rejection prediction results of each voice request sample to determine the natural language processing model.
[0137] The prediction module of the embodiment of the present application is also used to determine multiple rejection prediction results of a voice request sample according to the second similarity between the text label of a voice request sample and each voice request sample. The training module is also used to perform model training according to the multiple rejection prediction results of each voice request sample to determine the natural language processing model.
[0138] The processor of the embodiment of the present application is also used to determine multiple rejection prediction results of a voice request sample based on a second similarity between a voice request sample and a text label of each voice request sample, and to perform model training based on the multiple rejection prediction results of each voice request sample to determine a natural language processing model.
[0139] Specifically, in order to ensure the prediction accuracy of the natural language processing model, in the implementation mode of the present application, the server may introduce a rejection prediction result corresponding to a voice request sample and the text label of each other voice request sample during the training of the natural language processing model. Then, for any voice request sample, the natural language processing model may predict a rejection prediction result based on the voice request sample and the text label of the voice request sample during training, and may also predict multiple rejection prediction results based on the voice request sample and the text label of each other voice request sample. Then, the server may train the natural language processing model based on all the rejection prediction results of the voice request sample.
[0140] To more clearly illustrate the implementation of this application, please refer to Figure 5 , Figure 5 Schematic diagram of application scenarios in certain embodiments of the present application. Specifically, Figure 5 As shown, in the embodiment of the present application, the natural language processing model can encode each voice request sample when obtaining N voice request samples (i.e., audio) and the text label of each voice request sample, thereby obtaining N voice request sample encoding results A, i.e., A1, A2, A3, ..., A NFurthermore, the natural language processing model can also encode the text label of each voice request sample, thereby obtaining N text label encoding results T, namely T1, T2, T3, ..., T N .
[0141] Next, the natural language processing model can make corresponding rejection predictions based on the combination A·T of any speech request sample encoding result A and any text label encoding result T, such as the combination A1·T1 formed by the speech request sample encoding result A1 of the first speech request sample and the text label encoding result T1 of the first speech request sample, and the combination A3·T1 formed by the speech request sample encoding result A3 of the third speech request sample and the text label encoding result T1 of the first speech request sample, that is, the second similarity between A and T in each combination A·T, and output the corresponding rejection prediction results based on the second similarity. Furthermore, for any speech request sample, the model can output N rejection prediction results.
[0142] Finally, the server can train the model based on the N rejection prediction results of each voice request sample.
[0143] In one example, the loss function may be calculated based on N rejection prediction results for each voice request sample.
[0144] Furthermore, in one example, when the model is trained based on the loss function calculated based on the N rejection prediction results of each voice request sample, it can be ensured that the misaligned combination A i ·T j The corresponding rejection prediction result can be close to rejection. i ·T i The corresponding rejection prediction result can be close to the rejection label of the i-th voice request sample. Wherein, i and j are both positive integers, i≠j, i∈[1,n], j∈[1,n].
[0145] Thus, in the implementation mode of the present application, multiple rejection prediction results of a voice request sample can be determined according to the second similarity between a voice request sample and the text label of each voice request sample, and model training can be performed according to the multiple rejection prediction results of each voice request sample to determine the natural language processing model, so that the difficulty of model training is increased, thereby ensuring sufficient training of the model and ensuring model performance.
[0146] The embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by one or more processors, the above-mentioned voice interaction method is implemented.
[0147] In the description of this specification, the descriptions with reference to the terms "specifically", "further", "particularly", "understandably", etc. are intended to mean that the specific features, structures, materials or characteristics described in conjunction with the embodiments or examples are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms are not intended to refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, unless they are contradictory.
[0148] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code that includes one or more executable instructions for implementing the steps of a specific logical function or process, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may not be performed in the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by technicians in the technical field to which the embodiments of the present application belong.
[0149] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limitations on the present application. Ordinary technicians in this field can change, modify, replace and modify the above embodiments within the scope of the present application.
Claims
1. A voice interaction method, characterized in that: include: Get the current voice request; Performing natural language processing on the current voice request to determine a target text corresponding to the current voice request; Determining a first similarity between the current voice request and the target text according to the current voice request and the target text; Determining a rejection result of the current voice request according to the current voice request, the target text, and the first similarity level; The voice interaction is performed according to the rejection result.
2. The method according to claim 1, characterized in that The determining, according to the current voice request and the target text, a first similarity between the current voice request and the target text comprises: Encoding the current voice request and the target text respectively to determine first encoding information of the current voice request and second encoding information of the target text; The first similarity level is determined according to the first encoding information and the second encoding information.
3. The method according to claim 1, characterized in that The step of determining a rejection result of the current voice request according to the current voice request, the target text, and the first similarity level includes: When the first similarity level is less than or equal to a first preset similarity threshold, the rejection result is determined to be a rejection.
4. The method according to claim 2, characterized in that: The step of determining a rejection result of the current voice request according to the current voice request, the target text, and the first similarity level includes: When the first similarity level is greater than a second preset similarity threshold, the rejection result is determined according to the current voice request and / or the target text.
5. The method according to claim 4, characterized in that The step of determining the rejection result according to the current voice request and / or the target text when the first similarity level is greater than a second preset similarity threshold includes: When the first similarity level is greater than a second preset similarity threshold, the rejection result is determined according to the audio features of the current voice request and / or the semantics of the target text.
6. The method according to claim 1, characterized in that The step of determining a rejection result of the current voice request according to the current voice request, the target text, and the first similarity level includes: Based on a pre-trained natural language processing model, the rejection result is determined according to the current voice request, the target text and the first similarity level.
7. The method according to claim 6, characterized in that The training steps of the natural language processing model include: Obtaining a voice request sample and a text label of the voice request sample; Determining a second similarity between the voice request sample and the text label according to the voice request sample and the text label; Determining a rejection prediction result according to the voice request sample, the text label and the second similarity level; Model training is performed according to the rejection prediction result to determine the natural language processing model.
8. The method according to claim 7, characterized in that The voice request samples include a plurality of samples, and determining the rejection prediction result according to the voice request samples, the text labels and the second similarity level includes: Determining a plurality of the rejection prediction results of the voice request sample according to the second similarity between the voice request sample and the text label of each of the voice request samples; The performing model training according to the rejection prediction result to determine the natural language processing model includes: Model training is performed according to the multiple rejection prediction results of each of the voice request samples to determine the natural language processing model.
9. A server, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the method according to any one of claims 1 to 8 is implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by one or more processors, the method according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Information processing method and device and storage medium
CN111580773A
Semantic rejection method, semantic rejection device, vehicle and medium
CN113221580A
Voice interaction method, vehicle and computer readable storage medium
CN114049884A
Intention recognition method and device, computer equipment and computer readable storage medium
CN114678014A
Method and device for processing audio data, audio data processing equipment and medium
CN116959421A