Voice intention recognition method, electronic equipment and storage medium

By modifying and retraining the Whisper model, adding an intent adaptation layer and a dual-stream decoder enables it to handle text transcription and intent recognition at the same time, solving the problems of resource waste and accuracy loss in the prior art, and achieving an efficient voice interaction architecture.

CN120279912AActive Publication Date: 2025-07-08HANGZHOU QIUGUOJIHUA TECHNOLOGY CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510758291.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-07-08
Estimated Expiration
2045-06-09

AI Technical Summary

Technical Problem

In the prior art, voice interaction devices need to deploy ASR and NLU models separately, resulting in waste of computing resources and storage space and loss of accumulated accuracy. How to build a lightweight end-side voice interaction architecture while ensuring voice interaction accuracy.

Method used

By modifying and retraining the Whisper model, an intent adaptation layer and a dual-stream decoder are added, allowing it to handle text transcription and intent recognition tasks simultaneously, including mapping speech embedding vectors at the input layer to the intent recognition feature space, and building text transcription and intent recognition flows at the output layer, dynamic selection of operation processes.

Benefits of technology

It realizes that without reducing text transcription accuracy, intent recognition capabilities are enhanced, accuracy losses caused by model compression are avoided, and the end-side deployment cost is optimized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279912A_ABST
    Figure CN120279912A_ABST
Patent Text Reader

Abstract

The invention provides a voice intention recognition method, electronic equipment and a storage medium, and belongs to the technical field of data processing, and the method comprises the steps: receiving voice information, and determining a target task type corresponding to the voice information according to the triggering information of the voice information; according to the target task type, a corresponding target input prompt is set for the voice information, and the target input prompt comprises a text transcription prompt and / or an intention recognition prompt; the voice information provided with the target input prompt is input to a pre-trained Whisper model, the pre-trained Whisper model is called to execute operation corresponding to the target input prompt on the voice information, an execution result is output, and the execution result comprises a text transcription result and / or an intention recognition result; and obtaining an execution result output by the pre-trained Whisper model, and executing the execution result. According to the invention, the problems of resource waste and accumulated precision loss caused by independent deployment of double models in the prior art can be solved, and the end-side deployment cost is optimized at the same time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and in particular, to a voice intent recognition method, an electronic device, and a storage medium. Background Art

[0002] In recent years, the demand for voice interaction in end-side devices such as mobile terminals and Internet of Things devices has grown rapidly, and its technical implementation highly depends on the collaborative processing of an Automatic Speech Recognition (ASR) model and a Natural Language Understanding (NLU) model. Existing technologies generally adopt a dual-model independent deployment architecture: First, the ASR model converts the voice signal into a text sequence, then the NLU model parses the text semantics and extracts the user's intent, and finally drives the device to execute an operation.

[0003] However, the ASR and NLU models usually need to be separately deployed as independent inference units, and both of them require a large amount of computing resources and storage space. To reduce the deployment cost of the dual models, existing technologies have proposed model lightweighting techniques (such as knowledge distillation and parameter quantization), but such methods need to compress the ASR and NLU separately, resulting in cumulative accuracy loss.

[0004] Therefore, how to build a lightweight end-side voice interaction architecture while ensuring the accuracy of voice interaction is a technical problem that urgently needs to be solved at present. Summary of the Invention

[0005] The purpose of this application is to provide a voice intent recognition method, an electronic device, and a storage medium to solve the above problems.

[0006] To achieve the above purpose, in a first aspect, this application proposes a voice intent recognition method, which includes: Receiving voice information, and determining the target task type corresponding to the voice information according to the trigger information of the voice information; According to the target task type, setting a corresponding target input prompt for the voice information, where the target input prompt includes a text transcription prompt and / or an intent recognition prompt; Inputting the voice information with the target input prompt set into a pre-trained Whisper model, calling the pre-trained Whisper model to perform an operation corresponding to the target input prompt on the voice information, and outputting an execution result, where the execution result includes a text transcription result and / or an intent recognition result; Obtaining the execution result output by the pre-trained Whisper model and executing the execution result.

[0007] In some embodiments, the pre-trained Whisper model includes an input layer and an output layer. The input layer includes an intent adaptation layer, and the output layer includes a two-stream decoder. The two-stream decoder includes a text transcription stream and an intent recognition stream, and is further configured to select and execute the text transcription stream and / or the intent recognition stream according to an input prompt.

[0008] In some embodiments, when the target input prompt includes an intent recognition prompt, calling the pre-trained Whisper model to perform an operation corresponding to the target input prompt on the speech information and output an execution result includes: Performing feature extraction on the speech information to generate an original speech embedding vector; Based on the intent recognition prompt, calling the intent adaptation layer to map the original speech embedding vector to the feature space of the intent recognition stream to generate a new embedding vector; Concatenating the new embedding vector and the original speech embedding vector to generate a reconstructed embedding vector; Inputting the reconstructed embedding vector into the intent recognition stream, and calling the intent recognition stream to generate an intent recognition result based on the reconstructed embedding vector.

[0009] In some embodiments, the intent recognition stream includes a multi-layer perceptron and a hierarchical intent classifier. Calling the intent recognition stream to generate an intent recognition result based on the reconstructed embedding vector includes: Performing feature extraction on the reconstructed embedding vector to generate an intent feature vector; Enhancing the features of the intent feature vector through the multi-layer perceptron to generate an enhanced intent feature vector; Inputting the enhanced intent feature vector into the hierarchical intent classifier, and forming intent prediction results at several levels through the hierarchical intent classifier; Performing structured processing on the intent prediction results at several levels to generate an intent recognition result.

[0010] In some embodiments, the intent recognition stream further includes a termination judge. Inputting the enhanced intent feature vector into the hierarchical intent classifier and forming intent prediction results at several levels through the hierarchical intent classifier includes: Inputting the enhanced intent feature vector into the first level of the hierarchical intent classifier to obtain an intent prediction result at the first level; Inputting the intent prediction result at the current level and the enhanced intent feature vector into the termination judge to determine whether to perform intent prediction at the next level; If so, input the intention prediction result of the current level and the enhanced intention feature vector into the next level of the hierarchical intention classifier, and repeat the step of inputting the intention prediction result of the current level and the enhanced intention feature vector into the termination judge to determine whether to perform intention prediction at the next level; If not, terminate the prediction.

[0011] In some embodiments, when the target input prompt includes a text transcription prompt, the pre-trained Whisper model is called to perform an operation corresponding to the target input prompt on the speech information, and an execution result is output, where the execution result includes a text transcription result and / or an intention recognition result, including: Feature extraction is performed on the speech information through a multi-head attention mechanism and a feed-forward neural network to generate an initial speech feature vector, where the initial speech feature vector includes long-term dependencies and context information in the speech information; Based on the text transcription prompt, the text transcription stream is called to convert the initial speech feature vector into a text sequence, and the text sequence is output as the text transcription result.

[0012] In some embodiments, before inputting the speech information with the target input prompt into the pre-trained Whisper model, it further includes: Obtain an initial Whisper model, and replace the translation prompt of the initial Whisper model with an intention recognition prompt; Add an intention recognition layer to the input layer of the initial Whisper model; and construct a two-stream decoder at the output layer of the initial Whisper model to form a modified Whisper model; The two-stream decoder includes a text transcription stream and an intention recognition stream, and the two-stream decoder is used to select and execute the text transcription stream and / or the intention recognition stream according to the input prompt; The modified Whisper model is trained through a training sample set, a text transcription sample set, and an intention recognition sample set including a number of quadruple data to obtain a pre-trained Whisper model.

[0013] In some embodiments, the training of the initial Whisper model through a training sample set, a text transcription sample set, and an intention recognition sample set including a number of quadruple data to obtain a pre-trained Whisper model includes: Obtain a training sample set including a number of quadruple data, where the quadruple data includes speech, text, input prompt, and intention recognition result; Obtain a text transcription sample set and an intent recognition sample set, and set corresponding input prompts for each sample in the text transcription sample set and the intent recognition sample set respectively; Input the text transcription sample set and the intent recognition sample set into the initial Whisper model to obtain a text transcription result set and an intent recognition result set respectively output by the initial Whisper model; Calculate a text transcription loss based on the text transcription result set and the training sample set, and calculate an intent classification loss based on the intent recognition result set and the training sample set; Perform feedback update on the initial Whisper model based on the text transcription loss and the intent classification loss until the training is completed.

[0014] In a second aspect, the present application proposes an electronic device, including: One or more processors; A memory for storing one or more programs, wherein, when the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the voice intent recognition method as described above.

[0015] In a third aspect, the present application proposes a storage medium, which stores executable instructions, and when the instructions are executed by a processor, the processor is caused to execute the voice intent recognition method as described above.

[0016] Compared with the prior art, the beneficial effects of the present application include: By modifying and retraining the Whisper model, the pre-trained Whisper model can handle text transcription and intent recognition tasks simultaneously, solving the problems of resource waste and cumulative accuracy loss caused by the independent deployment of dual models in the prior art. Specifically as follows: First, by modifying the prompt mechanism of the Whisper model and replacing the translation task prompt with an intent recognition prompt, the Whisper model that originally did not have the intent recognition function can perform the intent recognition operation corresponding to the intent recognition prompt. Second, an intent adaptation layer is added to the input layer of the Whisper model. The original speech embedding vector is mapped to the feature space of intent recognition through linear projection and concatenated with the original embedding vector to generate a reconstructed embedding vector, enabling the model to capture the intent features in the speech without relying on an independent intent recognition model. Third, by constructing a two-stream decoder in the output layer, including a text transcription stream and an intent recognition stream, the model can dynamically select and execute the corresponding stream according to the input prompt, thus flexibly responding to different task requirements and improving the processing efficiency. The above-mentioned improvement points not only retain the high accuracy of the Whisper model in text transcription but also enhance its performance in intent recognition, avoid the accuracy loss caused by model compression, and optimize the end-side deployment cost. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required in the embodiments. It should be understood that the following drawings only show some embodiments of the present application and should not be regarded as limiting the scope of the present application.

[0018] Figure 1 It is a schematic flowchart of a voice intent recognition method in an embodiment; Figure 2 It is a schematic flowchart of performing intent recognition in an embodiment; Figure 3 It is a schematic flowchart of the detailed process of performing intent recognition in an embodiment; Figure 4 It is a schematic flowchart of the detailed process of hierarchical intent recognition in an embodiment; Figure 5 It is a schematic flowchart of performing text transcription in an embodiment; Figure 6 It is a schematic flowchart of modifying and training the initial Whisper model in an embodiment; Figure 7 It is a schematic structural diagram of an electronic device involved in the voice intent recognition method in the embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0020] All terms used in the present application (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification, and should not be interpreted in an idealized or overly rigid manner.

[0021] For example, terms such as "first" and "second" used in the present application may be used herein to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish the first element from another element.

[0022] For another example, terms such as "comprising" and "including" used in the present application indicate the presence of features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0023] As described above, in recent years, the demand for voice interaction in end-side devices such as mobile terminals and Internet of Things devices has grown rapidly, and its technical implementation highly depends on the collaborative processing of an Automatic Speech Recognition (ASR) model and a Natural Language Understanding (NLU) model. The existing technology generally adopts a dual-model independent deployment architecture: First, the ASR model converts the voice signal into a text sequence, then the NLU model parses the text semantics and extracts the user's intention, and finally drives the device to execute an operation. However, the ASR and NLU models usually need to be separately deployed as independent inference units, and both of them require a large amount of computing resources and storage space. To reduce the deployment cost of the dual models, the existing technology has proposed model lightweighting techniques (such as knowledge distillation and parameter quantization), but such methods need to compress the ASR and NLU separately, resulting in cumulative accuracy loss. Therefore, how to construct a lightweight end-side voice interaction architecture while ensuring the accuracy of voice interaction is a technical problem that urgently needs to be solved at present. For this reason, this application proposes a voice intention recognition method, an electronic device, and a storage medium. By modifying and retraining the Whisper model (Whisper is a general speech recognition model developed by OpenAI, focusing on accurately converting speech into text, especially outstanding in multilingual and complex environments), the pre-trained Whisper model not only retains high accuracy in text transcription but also enhances its performance in intention recognition, can simultaneously process text transcription and intention recognition tasks, solves the problems of resource waste and cumulative accuracy loss caused by the independent deployment of dual models in the existing technology, avoids the accuracy loss caused by model compression, and at the same time optimizes the end-side deployment cost.

[0024] In one embodiment, as Figure 1 shown, an embodiment of this application proposes a voice intention recognition method, and the method includes the following steps: Step S10, receive voice information, and determine the target task type corresponding to the voice information according to the trigger information of the voice information.

[0025] In this embodiment, the voice information refers to the acoustic wave signal collected by the user through a voice input device (such as a microphone), usually represented in the form of a digital audio waveform or a Mel spectrogram. The target task type refers to the type of task that the model needs to execute, including text transcription and / or intent recognition. The trigger information refers to the specific features in the voice information used to identify the task type, which can be the trigger method and keywords when the user triggers the voice input. For example, if the software that the user triggers the voice input is an input method software, the target task type can be determined as text transcription according to this trigger method. If the trigger information also includes keywords such as "send message", the target task type is determined as text transcription plus intent recognition. If the voice input is triggered by an instruction call, the target task type can be determined as intent recognition according to the trigger method of this instruction call. The trigger information can also be keywords, context features in the voice information, or multimodal features combined with other sensor data (such as visual information). For example, in the scenario of MR glasses, the trigger information can include keywords in the voice and visual information collected by the glasses (such as the information of the object the user is looking at or gestures), which is used to assist in judging the task type.

[0026] Step S20: Set a corresponding target input prompt for the voice information according to the target task type.

[0027] In this embodiment, the input prompt refers to the identifier used to guide the model to execute the target task, including a text transcription prompt and / or an intent recognition prompt. Among them, the text transcription prompt refers to the identifier used to instruct the model to execute the text transcription task, and the intent recognition prompt refers to the identifier used to instruct the model to execute the intent recognition task.

[0028] In some embodiments, the input prompt can be an additional dimension of the model input or a marker at a specific position, and can also be used as the input prefix of the voice information. For example, when the input prompt is used as the input prefix of the voice information, and the input prefix includes [decoder_start_token_id, language_id, task_id, notimestamps], then "task_id" can be set to "transcribe" (text transcription).

[0029] In some embodiments, the language prompt can also be set for the voice information according to the language information of the voice information. For example, when the language prompt is used as the input prefix of the voice information, and the input prefix includes [decoder_start_token_id, language_id, task_id, notimestamps], then "language_id" can be set to "CN" (Chinese) or "EN" (English), etc.

[0030] Step S30: Input the voice information with the target input prompt into the pre-trained Whisper model, and call the pre-trained Whisper model to perform an operation corresponding to the target input prompt on the voice information and output the execution result.

[0031] In this embodiment, the pre-trained Whisper model refers to a model that has been modified and retrained based on the initial Whisper model (Initial Whisper is a general speech recognition model developed by OpenAI, which focuses on accurately converting speech into text, especially outstanding in multi-language and complex environments). As a result, the pre-trained Whisper model not only retains high accuracy in text transcription but also enhances its performance in intent recognition, and can handle both text transcription and intent recognition tasks simultaneously. The execution result includes the text transcription result and / or the intent recognition result.

[0032] In some embodiments, the pre-trained Whisper model includes an input layer and an output layer. The input layer includes an intent adaptation layer, which is a sub-layer of the input layer and is used to map the embedding vector of the voice information to the feature space of the intent recognition task. The output layer includes a dual-stream decoder, which includes a text transcription stream and an intent recognition stream. The dual-stream decoder is also used to select and execute the text transcription stream and / or the intent recognition stream according to the input prompt. The text transcription stream is used to generate the text transcription result based on the multi-head attention mechanism and the feed-forward neural network, and the intent recognition stream is applicable to generate the intent recognition result based on the multi-layer perceptron and the hierarchical intent classifier.

[0033] In some embodiments, the pre-trained Whisper model can be called to perform an operation corresponding to the text transcription prompt on the voice information and output the text transcription result, and / or the pre-trained Whisper model can be called to perform an operation corresponding to the intent recognition prompt on the voice information and output the intent recognition result.

[0034] In some embodiments, when the target input prompt includes the intent recognition prompt, as Figure 2 shown, step S20 includes: Step A10: Extract features from the voice information to generate the original voice embedding vector.

[0035] In some embodiments, the voice information can be converted into Mel spectrogram, and by extracting features from the Mel spectrogram, the original voice embedding vector can be generated. The original voice embedding vector is the preliminary feature representation of the pre-trained Whisper model for the voice information, which can be used to carry the acoustic features of the voice information (such as pitch, phoneme, speech rate, etc.) and is represented as a vector with a fixed dimension.

[0036] In some embodiments, the process of feature extraction can directly use the encoder of the pre-trained Whisper model, or algorithms such as the Short-Time Fourier Transform (STFT) or wavelet transform, or a pre-trained feature extraction model (such as ResNet) can be embedded to generate a more advanced feature representation.

[0037] Step A20, based on the intent recognition prompt, call the intent adaptation layer to map the original speech embedding vector to the feature space of the intent recognition stream, generating a new embedding vector.

[0038] In some embodiments, based on the intent recognition prompt, call the intent adaptation layer to map the original speech embedding vector to the feature space of the intent recognition stream through linear projection, generating a new embedding vector. The new embedding vector has stronger feature expression ability and is more suitable for the intent recognition task.

[0039] Step A30, splice the new embedding vector and the original speech embedding vector to generate a reconstructed embedding vector.

[0040] In this embodiment, the reconstructed embedding vector refers to the new vector after splicing, which contains the acoustic features and intent features of the speech information. By splicing the new embedding vector and the original speech embedding vector, it is possible to introduce intent features while retaining the acoustic features of the speech information, further enhancing the feature expression ability of the model.

[0041] In some embodiments, the new embedding vector and the original speech embedding vector can be weighted and spliced to generate a reconstructed embedding vector, where the weights of the new embedding vector and the original speech embedding vector can be fixed or obtained through pre-training and learning.

[0042] The attention mechanism can also be used to splice the new embedding vector and the original speech embedding vector to generate a reconstructed embedding vector. Specifically, the new embedding vector can be regarded as the query vector, the key vector can be regarded as the original speech embedding vector, the similarity between the query vector (new embedding vector) and the key vector (original speech embedding vector) is calculated through dot product or scaled dot product to generate attention weights, and the attention weights are applied to the splicing of the new embedding vector and the original speech embedding to generate a reconstructed embedding vector.

[0043] Step A40, input the reconstructed embedding vector into the intent recognition stream, and call the intent recognition stream to generate an intent recognition result based on the reconstructed embedding vector.

[0044] In some embodiments, the intent recognition stream includes a multi-layer perceptron and a hierarchical intent classifier, such as Figure 3 As shown, step A40 includes: Step A41: Extract features from the reconstructed embedding vector to generate an intent feature vector.

[0045] Extract the feature vector related to intent recognition from the reconstructed embedding vector as the intent feature vector.

[0046] Step A42: Enhance the features of the intent feature vector through the multi-layer perceptron to generate an enhanced intent feature vector.

[0047] In this embodiment, the multi-layer perceptron (MLP) is a feed-forward neural network composed of multiple fully-connected layers, which is used for feature transformation and enhancement.

[0048] In some embodiments, the reconstructed embedding vector is input into the multi-layer perceptron, and the intent feature vector is non-linearly transformed through the fully-connected layers and activation functions (such as ReLU, Swish, or Mish) in the multi-layer perceptron to generate an enhanced intent feature vector.

[0049] Step A43: Input the enhanced intent feature vector into the hierarchical intent classifier to form intent prediction results at several levels through the hierarchical intent classifier.

[0050] In this embodiment, the hierarchical intent classifier refers to a classifier used to process multi-level intent recognition tasks and is composed of classifiers at multiple levels. When the enhanced intent feature vector is input into the hierarchical intent classifier, the classifiers at each level sequentially predict the intents at the corresponding levels and pass the results to the next level, thereby forming intent prediction results at several levels.

[0051] In some embodiments, the intent recognition flow further includes a termination judge, such as Figure 4 As shown, Step A43 includes: Step A431: Input the enhanced intent feature vector into the first level of the hierarchical intent classifier to obtain the intent prediction result of the first level.

[0052] In this embodiment, the enhanced intent feature vector is input into the first level of the hierarchical intent classifier to obtain the probability distribution corresponding to various possible first-level intents, and the first-level intent with the highest probability is used as the intent prediction result of the first level.

[0053] Exemplarily, the voice information is "Take a selfie for me", and the enhanced intent feature vector includes the embedded representations of the keywords "selfie" and "photo" in the voice information and context information. The first level of the hierarchical intent classifier outputs various first-level intents, such as "Open the camera app", "Open the gallery", "Turn on the front camera", "Find selfie photos", etc., based on the embedded representations in the enhanced intent feature vector, and takes the first-level intent with the highest probability, such as "Open the camera app", as the intent prediction result of the first level.

[0054] Step A432, input the intent prediction result of the current level and the enhanced intent feature vector into the termination judge.

[0055] Step A433, determine whether to perform intent prediction for the next level.

[0056] In this embodiment, the termination judge refers to a module for determining whether to continue predicting the intent of the next level, and can be a small multi-layer perceptron (MLP). The intent prediction result of the current level and the enhanced intent feature vector are concatenated into a high-dimensional fusion vector, and the high-dimensional fusion vector is input into the termination judge. The termination judge outputs a probability value indicating the possibility of continuing the prediction. By determining whether the probability value reaches a preset threshold, it is determined whether to perform intent prediction for the next level.

[0057] In some embodiments, the termination judge is pre-trained so that it has the ability to judge the number of levels of intent prediction. Specifically, a learning sample set containing multi-level intent labels is obtained, and each sample includes voice, text, multi-level intent labels, and a mark indicating whether to continue the prediction. The termination judge is trained using this training sample set to obtain a pre-trained termination judge.

[0058] Step A434, if so, input the intent prediction result of the current level and the enhanced intent feature vector into the next level of the hierarchical intent classifier, and repeat Step A432 and Step A433.

[0059] In this embodiment, if the probability value output by the termination judge is greater than the preset threshold, the intent prediction result of the current level and the enhanced intent feature vector are input into the classifier of the next level, and the judgment steps are repeated and executed.

[0060] Exemplarily, the enhanced intent feature vector is input into the first level of the hierarchical intent classifier to obtain the probability distribution corresponding to various possible first-level intents, and the first-level intent with the highest probability is used as the intent prediction result of the first level.

[0061] Exemplarily, an enhanced intent feature vector including the embedded representations of the keywords "selfie" and "photo" in the voice information and context information, as well as the intent prediction result at the first level (opening the camera app) are input into the second level of the hierarchical intent classifier to obtain the probability distributions of various secondary intents, such as "opening the front camera", "taking a photo", etc., and the secondary intent with the highest probability, such as "opening the front camera", is used as... Then, the intent prediction result "opening the front camera" at the second level and the enhanced intent feature vector are input into the termination discriminator to determine whether to continue predicting the tertiary intent. If so, the above process is repeated to obtain the intent prediction result "taking a photo" at the third level.

[0062] Step A435, if not, then terminate the prediction.

[0063] In this embodiment, if the probability value output by the termination discriminator is less than or equal to the preset threshold, further intent prediction is stopped.

[0064] Step A44, perform structured processing on the intent prediction results of the several levels to generate an intent recognition result.

[0065] In this embodiment, the intent prediction results of multiple levels are combined into a structured output format. For example, if the intent prediction result at the first level is "opening the camera app", the intent prediction result at the second level is "opening the front camera", and the intent prediction result at the third level is "taking a photo", the generated intent recognition result can be "opening the camera app\nopening the front camera\ntaking a photo".

[0066] In some embodiments, when the target input prompt includes a text transcription prompt, as Figure 5 shown, the step S20 includes: Step B10, perform feature extraction on the voice information through a multi-head attention mechanism and a feed-forward neural network to generate an initial voice feature vector.

[0067] In this embodiment, the multi-head attention mechanism is a mechanism for capturing long-term dependencies in sequence data, and captures the correlations between different positions through parallel calculations of multiple attention heads. The feed-forward neural network refers to a simple neural network structure that includes multiple fully connected layers for feature transformation. The initial voice feature vector extracted through the multi-head attention mechanism and the feed-forward neural network can include the long-term dependencies and context information in the voice information.

[0068] Step B20, based on the text transcription prompt, call the text transcription stream to convert the initial voice feature vector into a text sequence, and output the text sequence as the text transcription result.

[0069] In this embodiment, the text transcription stream is actually an end-to-end speech recognition decoder natively implemented by the Whisper model. This decoder consists of multiple Transformer layers and is responsible for gradually generating a text sequence from the initial speech feature vectors generated by the encoder. During the generation process of each step of the text sequence, the decoder predicts the next most likely character or word based on the current context and the previously generated text, generates the final text sequence, and interacts with the output of the encoder through the attention mechanism to obtain relevant context information.

[0070] Step S40: Obtain the execution result output by the pre-trained Whisper model and execute the execution result.

[0071] In this embodiment, the electronic device executing the speech intent recognition method obtains the execution result from the pre-trained Whisper model and performs corresponding operations according to the result type. If the execution result is a text transcription result, the text transcription result is displayed to the user. If the execution result is an intent recognition result, corresponding application operations are triggered based on the intent recognition result.

[0072] In the speech intent recognition method proposed in the embodiments of this application, by modifying and retraining the Whisper model (Whisper is a general speech recognition model developed by OpenAI, which focuses on accurately converting speech into text, especially performing outstandingly in multilingual and complex environments), the pre-trained Whisper model not only retains high accuracy in text transcription but also enhances its performance in intent recognition. It can handle both text transcription and intent recognition tasks simultaneously, solves the problems of resource waste and cumulative accuracy loss caused by the independent deployment of dual models in the prior art, avoids the accuracy loss caused by model compression, and at the same time optimizes the end-side deployment cost.

[0073] In one embodiment, as Figure 6 shown, before the step S30, it further includes: Step S00: Obtain the initial Whisper model and replace the translation prompt of the initial Whisper model with an intent recognition prompt.

[0074] In this embodiment, the initial Whisper model refers to a general speech recognition model developed by OpenAI, which focuses on accurately converting speech into text, especially performing outstandingly in multilingual and complex environments. By modifying the prompt mechanism of the initial Whisper model and replacing the translation prompt with an intent recognition prompt, the new task can be adapted while keeping the main structure of the model unchanged.

[0075] Step S01: Add an intent recognition layer to the input layer of the initial Whisper model.

[0076] In this embodiment, an intent adaptation layer is added to the input layer of the initial Whisper model to map the original speech embedding vector to the feature space of the intent recognition task.

[0077] Step S02, a two-stream decoder is constructed in the output layer of the initial Whisper model to form a modified Whisper model.

[0078] In this embodiment, the two-stream decoder includes a text transcription stream and an intent recognition stream, and the two-stream decoder is used to select and execute the text transcription stream and / or the intent recognition stream according to the input prompt. The initial Whisper model itself has a text transcription function. By expanding the initial Whisper model and adding an intent recognition stream, it can handle both text transcription and intent recognition tasks simultaneously. The multi-layer perceptron and hierarchical intent classifier in the intent recognition stream are components specifically designed for the intent recognition task, and they work together with the encoder and decoder of the initial Whisper model to achieve multi-task processing. Step S03, the modified Whisper model is trained with a training sample set, a text transcription sample set, and an intent recognition sample set each containing a number of quadruple data to obtain a pre-trained Whisper model.

[0079] In some embodiments, a training sample set containing a number of quadruple data is obtained, where the quadruple data includes speech, text, input prompt, and intent recognition result; a text transcription sample set and an intent recognition sample set are obtained, and corresponding input prompts are set for each sample in the text transcription sample set and the intent recognition sample set; the text transcription sample set and the intent recognition sample set are input into the initial Whisper model to obtain a text transcription result set and an intent recognition result set respectively output by the initial Whisper model; a text transcription loss is calculated based on the text transcription result set and the training sample set, and an intent classification loss is calculated based on the intent recognition result set and the training sample set; the initial Whisper model is feedback updated based on the text transcription loss and the intent classification loss until the training is completed.

[0080] In the speech intent recognition method proposed by the embodiments of this application, by modifying and retraining the Whisper model, the pre-trained Whisper model can handle text transcription and intent recognition tasks simultaneously, solving the problems of resource waste and cumulative accuracy loss caused by the independent deployment of dual models in the prior art. Specifically as follows: First, by modifying the prompt mechanism of the Whisper model and replacing the translation task prompt with an intent recognition prompt, the Whisper model that originally does not have the intent recognition function can perform the intent recognition operation corresponding to the intent recognition prompt. Second, an intent adaptation layer is added to the input layer of the Whisper model. The original speech embedding vector is mapped to the feature space of intent recognition through linear projection and concatenated with the original embedding vector to generate a reconstructed embedding vector, so that the model can capture the intent features in the speech without relying on an independent intent recognition model. Third, by constructing a two-stream decoder in the output layer, including a text transcription stream and an intent recognition stream, the model can dynamically select and execute the corresponding stream according to the input prompt, thus flexibly responding to different task requirements and improving the processing efficiency. The above-mentioned improvement points not only retain the high accuracy of the Whisper model in text transcription, but also enhance its performance in intent recognition, avoid the accuracy loss caused by model compression, and optimize the end-side deployment cost.

[0081] In one embodiment, a computer-readable storage medium is provided, on which executable instructions are stored. When the instructions are executed by a processor, the processor executes the steps in the above method embodiments.

[0082] In one embodiment, an electronic device is further provided, including one or more processors; a memory, and one or more programs are stored in the memory. Wherein, when the one or more programs are executed by the one or more processors, the one or more processors execute the steps in the above method embodiments.

[0083] In one embodiment, as Figure 7 shown, it shows a schematic structural diagram of an electronic device for implementing the embodiments of the present application. The electronic device 700 includes a central processing unit (CPU) 701, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 702 or the program loaded from the storage section 708 into the random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the electronic device 700 are also stored. The CPU 701, ROM 702, and RAM 703 are connected to each other through a bus 704. The input / output (I / O) interface 705 is also connected to the bus 704.

[0084] The following components are connected to the I / O interface 705: an input section 706 including a keyboard, a mouse, etc.; an output section 707 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. as well as a speaker, etc.; a storage section 708 including a hard disk, etc.; and a communication section 709 including a network interface card such as a LAN card, a modem, etc. The communication section 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the I / O interface 705 as needed. A removable medium 711 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is installed on the drive 710 as needed so that a computer program read from it can be installed into the storage section 708 as needed.

[0085] Specifically, according to an embodiment of the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present application includes a computer program product including a computer-readable medium carrying instructions. In such an embodiment, the instructions can be downloaded and installed from a network through the communication section 709, and / or installed from the removable medium 711. When the instructions are executed by a central processing unit (CPU) 701, the various method steps described in the present application are executed.

[0086] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.

[0087] In addition, those skilled in the art can understand that although some of the embodiments herein include certain features included in other embodiments rather than other features, the combination of the features of different embodiments means that it is within the scope of the present application and forms different embodiments. For example, any one of the above embodiments can be used in any combination. The information disclosed in this background art section is only intended to deepen the understanding of the overall background art of the present application, and should not be regarded as an admission or any form of implication that this information constitutes the prior art known to those skilled in the art.

Claims

1. A method for speech intent recognition, characterized in that, The method includes: Receiving voice information and determining the target task type corresponding to the voice information according to the trigger information of the voice information; Setting a corresponding target input prompt for the voice information according to the target task type, where the target input prompt includes a text transcription prompt and / or an intent recognition prompt; Inputting the voice information with the target input prompt into a pre-trained Whisper model, calling the pre-trained Whisper model to perform an operation corresponding to the target input prompt on the voice information, and outputting an execution result, where the execution result includes a text transcription result and / or an intent recognition result; Obtaining the execution result output by the pre-trained Whisper model and executing the execution result.

2. The voice intent recognition method according to claim 1, wherein The pre-trained Whisper model includes an input layer and an output layer. The input layer includes an intent adaptation layer, and the output layer includes a two-stream decoder. The two-stream decoder includes a text transcription stream and an intent recognition stream, and the two-stream decoder is further configured to select and execute the text transcription stream and / or the intent recognition stream according to the input prompt.

3. The voice intention recognition method according to claim 2, wherein When the target input prompt includes an intent recognition prompt, the step of calling the pre-trained Whisper model to perform an operation corresponding to the target input prompt on the voice information and outputting an execution result includes: Performing feature extraction on the voice information to generate an original voice embedding vector; Based on the intent recognition prompt, calling the intent adaptation layer to map the original voice embedding vector to the feature space of the intent recognition stream to generate a new embedding vector; Concatenating the new embedding vector and the original voice embedding vector to generate a reconstructed embedding vector; Inputting the reconstructed embedding vector into the intent recognition stream, and calling the intent recognition stream to generate an intent recognition result based on the reconstructed embedding vector.

4. The voice intention recognition method according to claim 3, characterized in that The intent recognition stream includes a multi-layer perceptron and a hierarchical intent classifier. The step of calling the intent recognition stream to generate an intent recognition result based on the reconstructed embedding vector includes: Performing feature extraction on the reconstructed embedding vector to generate an intent feature vector; Performing feature enhancement on the intent feature vector through the multi-layer perceptron to generate an enhanced intent feature vector; Inputting the enhanced intent feature vector into the hierarchical intent classifier, and forming intent prediction results at several levels through the hierarchical intent classifier; Performing structured processing on the intent prediction results at the several levels to generate an intent recognition result.

5. The voice intention recognition method according to claim 4, wherein The intent recognition stream further includes a termination judge. The step of inputting the enhanced intent feature vector into the hierarchical intent classifier and forming intent prediction results at several levels through the hierarchical intent classifier includes: Inputting the enhanced intent feature vector into the first level of the hierarchical intent classifier to obtain an intent prediction result at the first level; Inputting the intent prediction result at the current level and the enhanced intent feature vector into the termination judge to determine whether to perform intent prediction at the next level; If so, input the intention prediction result of the current level and the enhanced intention feature vector into the next level of the hierarchical intention classifier, and repeat the step of inputting the intention prediction result of the current level and the enhanced intention feature vector into the termination judge to determine whether to perform intention prediction at the next level; If not, terminate the prediction.

6. The voice intention recognition method according to claim 2, characterized in that, When the target input prompt includes a text transcription prompt, call the pre-trained Whisper model to perform an operation corresponding to the target input prompt on the voice information, and output an execution result, where the execution result includes a text transcription result and / or an intention recognition result, including: Extract features from the voice information through a multi-head attention mechanism and a feed-forward neural network to generate an initial voice feature vector, where the initial voice feature vector includes long-term dependencies and context information in the voice information; Based on the text transcription prompt, call the text transcription stream to convert the initial voice feature vector into a text sequence, and output the text sequence as the text transcription result.

7. The voice intent recognition method according to claim 1, characterized in that Before inputting the voice information with the target input prompt into the pre-trained Whisper model, it further includes: Obtain an initial Whisper model, and replace the translation prompt of the initial Whisper model with an intention recognition prompt; Add an intention recognition layer to the input layer of the initial Whisper model; and Construct a two-stream decoder in the output layer of the initial Whisper model to form a modified Whisper model; The two-stream decoder includes a text transcription stream and an intention recognition stream, and the two-stream decoder is used to select and execute the text transcription stream and / or the intention recognition stream according to the input prompt; Train the modified Whisper model through a training sample set, a text transcription sample set, and an intention recognition sample set including a number of quadruple data to obtain a pre-trained Whisper model.

8. The voice intention recognition method according to claim 7, wherein The training of the initial Whisper model through a training sample set, a text transcription sample set, and an intention recognition sample set including a number of quadruple data to obtain a pre-trained Whisper model includes: Obtain a training sample set including a number of quadruple data, where the quadruple data includes voice, text, input prompt, and intention recognition result; Obtain a text transcription sample set and an intention recognition sample set, and set corresponding input prompts for each sample in the text transcription sample set and the intention recognition sample set respectively; Input the text transcription sample set and the intention recognition sample set into the initial Whisper model to obtain a text transcription result set and an intention recognition result set respectively output by the initial Whisper model; Calculate a text transcription loss based on the text transcription result set and the training sample set, and calculate an intention classification loss based on the intention recognition result set and the training sample set; Based on the text transcription loss and the intent classification loss, perform feedback updates on the initial Whisper model until the training is completed.

9. An electronic device, characterized in that, Comprising: One or more processors; A memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the speech intent recognition method according to any one of claims 1 to 8.

10. A storage medium, characterized in that, The storage medium stores executable instructions, and when the instructions are executed by a processor, the processor is caused to execute the speech intent recognition method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Intelligent voice service management system and method driven by natural language understanding

    CN118194875A

  • Multi-modal voice model rejection training method, system and device, and storage medium

    CN118262724A

  • Intention recognition method and device based on voice interaction, equipment and medium

    CN118711585A

  • Classroom knowledge point coverage analysis method based on keyword and semantic matching

    CN119047481A

  • Multilingual subtitle generation method based on Whisper and FunASR double-channel speech recognition large model

    CN119517038A