Intention recognition method, model training method and electronic equipment
By determining the matching between the target input speech and the registered speech in the intent recognition model, the problem of inaccurate intent recognition in voice interaction in noisy environments is solved, achieving higher recognition accuracy and an optimized user experience.
Patent Information
- Application Number
- CN202511087470.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-04
- Publication Date
- 2025-11-11
AI Technical Summary
In noisy environments, the intent recognition of voice interaction is easily interfered with by background noise, leading to inaccurate recognition and affecting the user experience.
By acquiring the pre-set target registered speech, a pre-trained intent recognition model is used to determine whether the target input speech matches the registered speech. Intent recognition is only performed when a match is found. The model includes voiceprint feature extraction and a large language model to ensure speaker consistency.
It improves the accuracy of intent recognition, reduces background noise interference, and optimizes the user experience.
Smart Images

Figure CN120932634A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of natural language processing technology, and in particular to an intent recognition method, a model training method, and an electronic device. Background Technology
[0002] Currently, intent recognition in voice interaction involves key technologies such as speech recognition and natural language processing, with the goal of understanding the intent behind the user's voice input and executing the corresponding operation. However, due to the noisy environment in which the user is located, there may be interference from other sounds when the user outputs their voice, which can lead to inaccurate intent recognition and affect the user experience. Summary of the Invention
[0003] In view of this, the purpose of this disclosure is to propose an intent recognition method, a model training method, and an electronic device to solve or partially solve the above-mentioned problems.
[0004] For the purposes described above, a first aspect of this disclosure provides an intent recognition method, comprising:
[0005] In response to receiving the target input voice, obtain the preset target registration voice;
[0006] The target input speech and the target registered speech are input into a pre-trained intent recognition model and processed by the intent recognition model.
[0007] The process, via the intent recognition model, further includes:
[0008] Determine whether the target input speech matches the target registered speech;
[0009] In response to the target input speech matching the target registered speech, intent recognition is performed based on the target input speech and the intent recognition result is output.
[0010] Based on the same inventive concept, a second aspect of this disclosure provides a method for training an intent recognition model, comprising:
[0011] Obtain the training dataset, wherein each piece of training data in the training dataset includes training input speech data, training registered speech data, and text annotation data;
[0012] For each piece of training data, determine whether the training input speech data matches the training registered speech data, and determine the target format corresponding to the training data based on the determination result;
[0013] Obtain an initial intent recognition model, and train the initial intent recognition model using the training data in the target format until a preset training termination condition is met, thereby obtaining the intent recognition model.
[0014] Based on the same inventive concept, a third aspect of this disclosure proposes an electronic device including a memory, a processor, and a computer program stored in the memory and executable by the processor, wherein the processor implements the method described above when executing the computer program.
[0015] As can be seen from the above, the intent recognition method, model training method, and electronic device provided in this disclosure, when receiving target input speech, indicate that intent recognition needs to be performed on the user's target input speech. A preset target registration speech is obtained; the target registration speech is pre-determined speech data containing speaker features. Subsequently, it can be determined whether intent recognition of the target input speech is necessary based on the target registration speech. The target input speech and the target registration speech are input into a pre-trained intent recognition model to determine whether the target input speech and the target registration speech match, i.e., whether the user who issued the target input speech and the user who issued the target registration speech are the same user. If the target input speech and the target registration speech match, it indicates that the user who issued the target input speech and the user who issued the target registration speech are the same user. At this time, intent recognition is performed based on the target input speech, and the intent recognition result is output. Verification by the speaker of the target input speech avoids intent recognition deviations caused by the speaker of the received target input speech not being the speaker of the pre-selected target registration speech, reduces interference from background noise on intent recognition, improves recognition accuracy, and optimizes user experience. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in this disclosure or related technologies, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 This is a flowchart of the intent recognition method according to an embodiment of the present disclosure;
[0018] Figure 2 This is a first schematic diagram of an intent recognition model according to an embodiment of the present disclosure;
[0019] Figure 3 This is a second schematic diagram of the intent recognition model according to an embodiment of the present disclosure;
[0020] Figure 4 This is a schematic diagram of the adapter structure in an embodiment of this disclosure;
[0021] Figure 5This is a third schematic diagram of an intent recognition model according to an embodiment of the present disclosure;
[0022] Figure 6 This is a schematic diagram of the feature compression integration module in an embodiment of this disclosure;
[0023] Figure 7 This is a fourth schematic diagram of the intent recognition model according to an embodiment of the present disclosure;
[0024] Figure 8 This is a fifth schematic diagram of the intent recognition model according to an embodiment of the present disclosure;
[0025] Figure 9 This is a flowchart of the intent recognition model training method according to an embodiment of the present disclosure;
[0026] Figure 10 This is a schematic diagram of an intent recognition device according to an embodiment of the present disclosure;
[0027] Figure 11 This is a schematic diagram of an intent recognition model training device according to an embodiment of the present disclosure;
[0028] Figure 12 This is a schematic diagram of an electronic device according to an embodiment of the present disclosure. Detailed Implementation
[0029] To make the objectives, technical solutions, and advantages of this disclosure clearer, the following detailed description is provided in conjunction with specific embodiments and the accompanying drawings.
[0030] It should be noted that, unless otherwise defined, the technical or scientific terms used in the embodiments of this disclosure should have the ordinary meaning understood by one of ordinary skill in the art to which this disclosure pertains. The terms "first," "second," and similar terms used in the embodiments of this disclosure do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0031] The following are definitions of terms used in this disclosure:
[0032] ResNet: Deep residual network, primarily addresses the degradation problem of deep networks.
[0033] ECAPA-tdnn: ECAPA-tdnn is a TDNN-based neural network for speaker recognition and is one of the best single-model approaches currently available. It is based on TDNN and incorporates Res2Net and SENet modules to enhance the capture of contextual information.
[0034] CAM++: CAM++ is a speaker recognition model based on densely connected time-delay neural networks.
[0035] Large Language Models (LLMs) are deep learning models trained on massive amounts of text data, enabling them to generate natural language text or understand the meaning of language text. These models can provide in-depth knowledge and language production on a wide range of topics through training on large datasets. The core idea is to learn patterns and structures of natural language through large-scale unsupervised training, mimicking human language cognition and generation processes to some extent.
[0036] Currently, intent recognition in voice interaction involves key technologies such as speech recognition and natural language processing, aiming to understand the intent behind the user's voice input and execute corresponding operations. However, due to the noisy environment in which the user is located, there may be interference from other sounds when the user outputs speech. That is, during the speech recognition process, there will be a certain loss of recognition accuracy due to background noise, especially in the case of background human voices. This introduces prior error information into the subsequent natural language understanding task, resulting in the accumulation of recognition errors.
[0037] For example, in a home setting, when users interact with devices via voice, they may be affected by background noise, such as the sound of a TV drama playing or conversations of other family members. This may introduce redundant information during voice recognition, potentially making it difficult to accurately identify the user's intent, especially when multiple people are present.
[0038] Meanwhile, current intent recognition for voice interaction typically involves user input, followed by speech recognition to convert the speech signal into text, and then using natural language understanding technology to extract intent and key information from the recognized text. This process requires first converting speech information into text, and then performing intent recognition based on the text. This involves a cascading module problem, leading to the accumulation of errors between modules and ultimately resulting in poor intent recognition performance.
[0039] Based on the above description, one embodiment of this disclosure proposes an intent recognition method, such as... Figure 1 As shown, the method includes:
[0040] Step 101: In response to receiving the target input voice, obtain the preset target registration voice.
[0041] In practice, when a target input voice is received, it indicates that the user has a need for intent recognition, meaning that intent recognition needs to be performed on the target input voice. A preset target registered voice is obtained, wherein the target registered voice is pre-determined voice data containing speaker characteristics. Subsequently, it can be determined whether intent recognition of the target input voice is necessary based on the target registered voice.
[0042] Specifically, if it is subsequently determined that the speaker corresponding to the target input speech and the target registered speech is the same speaker, then the target input speech should be recognized. If it is determined that the speaker corresponding to the target input speech and the target registered speech is a different speaker, then the target input speech can be determined to be received background noise, and no further intent recognition operation is required.
[0043] In this embodiment, the target registration voice can be voice data selected by the user. For example, taking an in-vehicle scenario, the user can record their own voice information in the vehicle's infotainment system, i.e., record a registration voice; that is, the vehicle's infotainment system can store multiple registration voices. Before using the vehicle, the driver can select the voice information they recorded from the stored multiple registration voices; the voice information recorded by the driver is the target registration voice.
[0044] Another example is its application on smart devices, such as mobile phones. Users can record their own voice information on their phones, that is, record registration voice, and then use the recorded registration voice as the target registration voice.
[0045] Understandably, to reduce the number of registered voice recordings stored, and thus the memory occupied by them, when multiple recordings are entered, the speaker corresponding to each recording can be identified. For multiple recordings by the same speaker, only one should be retained.
[0046] Step 102: Input the target input speech and the target registered speech into the pre-trained intent recognition model, and process them through the intent recognition model;
[0047] In practice, after acquiring the target registered speech, the target input speech and the target registered speech are input into a pre-trained intent recognition model. The intent recognition model processes the target input speech and the target registered speech to obtain the intent recognition result. Optionally, the intent recognition model can be a model trained using deep learning technology.
[0048] Specifically, the process of processing via the intent recognition model in step 102 includes:
[0049] Step 1021: Determine whether the target input speech matches the target registered speech.
[0050] Step 1022: In response to the matching of the target input speech with the target registered speech, perform intent recognition based on the target input speech and output the intent recognition result.
[0051] In practice, the intent recognition model is used to determine whether the target input speech and the target registered speech match. In this embodiment, it can be understood as determining whether the user who outputs the target input speech is the same user as the user who outputs the target registered speech, that is, determining whether the speaker corresponding to the target input speech and the target registered speech is the same speaker.
[0052] If the target input speech matches the target registered speech, it means that the speakers corresponding to the target input speech and the target registered speech are the same person. The intent recognition model performs intent recognition on the received target input speech, obtains the intent recognition result, and outputs it. Subsequently, based on the intent recognition result, the functional components can be controlled to execute the corresponding functions.
[0053] For example, taking an in-car scenario, if the intention is recognized based on the target input voice and the intention recognition result is to play music, then the in-car multimedia function will be turned on and the corresponding music will be played.
[0054] If the target input speech does not match the target registered speech, it means that the speakers corresponding to the target input speech and the target registered speech are different speakers. That is, the target input speech may be received background noise. No further intent recognition operation is needed, and intent recognition is stopped.
[0055] The above scheme, upon receiving a target input voice, indicates the need for intent recognition. A preset target registration voice is obtained, which is pre-determined voice data containing speaker characteristics. Subsequent determinations based on the target registration voice indicate whether intent recognition of the target input voice is necessary. The target input voice and the target registration voice are input into a pre-trained intent recognition model to determine if they match, i.e., whether the user sending the target input voice and the user sending the target registration voice are the same user. If they match, it indicates that the user sending the target input voice and the user sending the target registration voice are the same user. In this case, intent recognition is performed based on the target input voice, and the intent recognition result is output. Verification using the speaker of the target input voice avoids errors caused by the speaker of the received target input voice not being the pre-selected speaker of the target registration voice, reduces interference from background noise, improves recognition accuracy, and optimizes the user experience.
[0056] In some embodiments, the intent recognition model includes a voiceprint feature extraction model and a large language model. Figure 2 A schematic diagram of the intent recognition model is shown. For example... Figure 2 As shown, step 1021, determining whether the target input speech matches the target registered speech, specifically includes:
[0057] Step 10211: Input the target input speech and the target registered speech into the voiceprint feature extraction model respectively. After processing by the speech feature extraction model, obtain the speaker features corresponding to the target input speech and the registered speaker features corresponding to the target registered speech.
[0058] Step 10212: Input the speaker features and the registered speaker features into the large language model, and use the large language model to determine whether the target input speech matches the target registered speech.
[0059] In specific implementation, the intent recognition model includes a voiceprint feature extraction model and a large language model. The voiceprint feature extraction model is used to extract the voiceprint features corresponding to the target input speech and the target registered speech, so that the large language model can determine whether the speakers of the target input speech and the target registered speech are the same speaker based on the voiceprint features.
[0060] Specifically, the target input speech is input into the voiceprint feature extraction model, and the voiceprint feature extraction model is used to extract the target input speech to obtain the speaker features corresponding to the target input speech. The speaker features can be used to determine the speaker of the target input speech.
[0061] The target registered speech is input into the voiceprint feature extraction model, and the voiceprint feature extraction model is used to extract the target registered speech to obtain the registered speaker features corresponding to the target registered speech. The registered speaker features can be used to determine the speaker of the target registered speech.
[0062] In this embodiment, the speaker features and the registered speaker features are preferably voiceprint features, wherein voiceprint features are biometric features composed of more than a hundred feature dimensions such as wavelength, frequency and intensity, and have the characteristics of stability, measurability and uniqueness.
[0063] After obtaining the speaker features and registered speaker features, the speaker features, registered speaker features and speech features are input into a large language model, and the large language model is used to determine whether the target input speech matches the target registered speech.
[0064] Specifically, a large language model can be used to determine the consistency between speaker features and registered speaker features. That is, to determine the consistency between the first voiceprint feature corresponding to the target input speech and the second voiceprint feature corresponding to the target registered speech, and to determine that the consistency value between the two is greater than a preset threshold, thus determining that the target input speech and the target registered speech are the speech output by the same speaker, that is, the target input speech and the target registered speech match.
[0065] The above scheme utilizes the uniqueness of speaker features, as different users have different voiceprint features. Therefore, speaker features are extracted through a voiceprint feature extraction model, which yields voiceprint features. Based on these voiceprint features, it is possible to quickly determine whether the target input speech and the target registered speech are output by the same user, thereby determining whether the target input speech and the target registered speech match. This improves judgment efficiency, shortens the time for intent recognition, and enhances user experience.
[0066] In some embodiments, the voiceprint feature extraction model includes a voiceprint feature extraction module and a first feature mapping module. Figure 3 A schematic diagram of the intent recognition model is shown, as follows: Figure 3 As shown, in step 10121, the target input speech and the target registered speech are respectively input into the voiceprint feature extraction model. After processing by the voiceprint feature extraction model, the speaker features corresponding to the target input speech and the registered speaker features corresponding to the target registered speech are obtained, specifically including:
[0067] Step 10a: Input the target input speech and the target registered speech into the voiceprint feature extraction module, and use the voiceprint feature extraction module to extract the target input speech and the target registered speech to obtain the initial speaker features corresponding to the target input speech and the initial registered speaker features corresponding to the target registered speech.
[0068] Step 10b: Input the initial speaker features and the initial registered speaker features into the first feature mapping module, and use the first feature mapping module to map the initial speaker features and the initial registered speaker features to obtain the speaker features corresponding to the target input speech and the registered speaker features corresponding to the target registered speech.
[0069] In specific implementation, the voiceprint feature extraction model includes a voiceprint feature extraction module and a first feature mapping module. The voiceprint feature extraction module is used to extract voiceprint features of speech, and the first feature mapping module is used to map high-dimensional voiceprint feature data to the low-dimensional space where the large language model is located, so that the speaker features output by the first feature mapping module and the registered speaker features can be input into the large language model for processing.
[0070] The target input speech is fed into the voiceprint feature extraction module, which extracts the initial speaker features. These initial speaker features are then fed into a first feature mapping module, which maps the initial speaker features to obtain the speaker features corresponding to the target input speech.
[0071] The target registered speech is input to the voiceprint feature extraction module, which extracts and processes the voiceprint feature to obtain the initial registered speaker features corresponding to the target registered speech. The initial registered speaker features are then input to the first feature mapping module, which maps the initial registered speaker features to obtain the registered speaker features corresponding to the target registered speech.
[0072] In this embodiment, the first feature mapping module is used to map the high-dimensional initial speaker features and the initial registered speaker features to the low-dimensional space where the large language model resides, respectively, to obtain speaker features and registered speaker features. Exemplarily, the first feature mapping module can be an adapter structure, wherein the adapter structure is a conformer structure based on a hybrid expert layer (MoE), such as... Figure 4 As shown, Figure 4A schematic diagram of the adapter in this embodiment is shown. The following explanation uses the mapping process of the first feature mapping module on the initial registered speaker features to obtain the registered speaker features corresponding to the target registered speech sound as an example, specifically including:
[0073] The adapter includes a feedforward network, a multi-head self-attention module, convolutional layers, and a hybrid expert layer. Voiceprint feature data is fed into the feedforward network and then into the multi-head self-attention module. This module improves the model's expressive and learning capabilities by segmenting the input features into multiple "heads" and processing each head independently. Key feature data from the initial registered speaker features is extracted through convolutional layers and then fed into the hybrid expert layer. This layer includes a gating network and multiple expert networks. The gating network determines the participation level of each expert network based on a set of gating values, guiding the key feature data from the initial registered speaker features to the appropriate expert network. Finally, the outputs of all expert networks are aggregated to obtain the registered speaker features.
[0074] In this embodiment, the voiceprint feature extraction module can be a speaker tokenizer. The speaker tokenizer uses residual vector quantization to convert continuous voiceprint features into discrete tokens. The residual quantization can be adjusted according to task requirements, achieving a balance between computational complexity and reconstruction quality. Simultaneously, using a tokenizer can compress the voiceprint signal, reducing storage and transmission costs.
[0075] In some embodiments, the voiceprint feature extraction module includes a voiceprint encoder and a first feature compression integration module, such as... Figure 5 As shown, Figure 5 A schematic diagram of the intent recognition model is shown. In step 10a, the target input speech and the target registered speech are respectively input to the voiceprint feature extraction module. The voiceprint feature extraction module is used to extract and process the target input speech and the target registered speech to obtain the initial speaker features corresponding to the target input speech and the initial registered speaker features corresponding to the target registered speech. Specifically, this includes:
[0076] Step a: Input the target input speech and the target registered speech into the voiceprint encoder respectively, and extract them using the voiceprint encoder to obtain at least one first speaker feature corresponding to the target input speech and at least one first registered speaker feature corresponding to the target registered speech;
[0077] Step b: Input the at least one first speaker feature into the first feature integration and compression module, and use the first feature integration and compression module to perform fusion processing to obtain the initial speaker feature;
[0078] Step c: Input the at least one first registered speaker feature into the first feature integration and compression module, and use the first feature integration and compression module to perform fusion processing to obtain the initial registered speaker feature.
[0079] In specific implementation, the voiceprint feature extraction module includes a voiceprint encoder and a first feature compression integration module. The voiceprint encoder is a pre-trained neural network model, which can be ResNet, ECAPA-tdnn, CAM++, etc., and is not specifically limited.
[0080] Since the voiceprint encoder may have a multi-layer structure, each layer outputs a corresponding feature. The first feature compression module then integrates and compresses all the features output by the voiceprint encoder to obtain a one-dimensional feature. Specifically, the first feature compression and integration module is used to compress and fuse the features output by the voiceprint encoder to obtain a one-dimensional feature, which is the initial speaker feature.
[0081] Specifically, the target input speech is input to a voiceprint encoder, and extracted using the voiceprint encoder to obtain at least one first speaker feature corresponding to the target input speech. The obtained at least one first speaker feature is input to a first feature integration and compression module, and fused using the first feature integration and compression module to obtain initial speaker features.
[0082] The target registered speech is input into a voiceprint encoder, which extracts at least one first registered speaker feature corresponding to the target registered speech. The obtained at least one first registered speaker feature is then input into a second feature integration and compression module, which performs fusion processing to obtain the initial registered speaker feature.
[0083] In this embodiment, the first feature compression and integration module is used to compress and fuse the features output by the voiceprint encoder, such as... Figure 6 As shown, Figure 6 A schematic diagram of the feature compression and integration module described in this application is shown. Taking the processing of all first speaker features using the feature compression and integration module to obtain initial speaker features as an example, the specific processing steps include:
[0084] All first speaker features are combined to form a feature map U, wherein the feature map U has dimensions H×W×C, and global average pooling of the channels is used, corresponding to F in the figure. sq (·), directly compressing the H×W×C feature map U containing global information into a 1×1×C feature vector Z, where the channel features of C feature maps are compressed into a single value, that is, compressing multiple spatial information into one channel. The corresponding formula is:
[0085]
[0086] Among them, z c Let x be the c-th element of Z. c (i,j) represents the pixel value at position (i,j) of the c-th channel in feature map U. The global average value z of this channel can be obtained by summing the values at all spatial positions (i,j) and dividing by the sum of the spatial dimensions H×W. c This allows the spatial information of each channel to be compressed into a single value, thereby capturing the global information of the channel.
[0087] Adaptive recalibration is employed, corresponding to F in the figure. ex The (·,W) architecture employs a two-layer fully connected gate mechanism. The first fully connected layer compresses the C channels into C / r channels to reduce computation, followed by a ReLU non-linear activation layer. The second fully connected layer restores the number of channels to C, and then obtains weights s through Sigmoid activation. The final s has a dimension of 1×1×C, which is used to characterize the weights of the C feature maps in feature map U, where r refers to the compression ratio.
[0088] Reweighting is performed, corresponding to F in the diagram. scale The (·,·) operator applies the attention weights obtained earlier to the features of each channel. This involves using an activation operation to learn the non-linear relationships between channels, generating a weight for each channel, calculating the weights using two fully connected layers and a non-linear activation function, and then multiplying the weights by the original features to perform a weighted summation across the channel space.
[0089] Finally, the output feature map is converted from multi-channel to single-channel, i.e., using F... red The · operator transforms the dimension from C to 1, meaning the feature map shape changes from H×W×C to T×D×1. Specific transformation methods can include global average pooling, 1×1 convolution, and channel-wise weighted summation. The formula is as follows:
[0090] s=σ(W2δ(W1z))
[0091]
[0092] Where y represents the initial speaker features, W1 is the weight matrix of the first fully connected layer, used to map the global average z of each channel of the input feature map to a low-dimensional space, i.e., to perform dimensionality reduction, thereby reducing the number of parameters and improving the model's generalization ability. W2 is the weight matrix of the second fully connected layer, used to map the dimensionality-reduced features back to the space of the original number of channels C, and s is the output weight vector. cLet be the weight vector corresponding to the c-th channel, δ be the ReLU activation function, and σ be the Sigmoid activation function.
[0093] In some embodiments, the intent recognition model further includes a speech feature extraction model, such as... Figure 2 As shown, Figure 2 A schematic diagram of the intent recognition model is shown. In step 1022, in response to the matching of the target input speech and the target registered speech, intent recognition is performed based on the target input speech and the intent recognition result is output. Specifically, this includes:
[0094] Step 10221: In response to the matching of the target input speech with the target registered speech, the target input speech is input to the speech feature extraction model, and the speech feature extraction model is used to extract the speech features corresponding to the target input speech;
[0095] Step 10222: Input the speech features into the large language model, and use the large language model to perform recognition processing on the speech features to obtain the target intent and target semantic slots;
[0096] Step 10223: Output the target intent and the target semantic slot as the intent recognition result.
[0097] In practice, the intent recognition model also includes a speech feature extraction model, which is used to extract feature representations from speech.
[0098] After determining that the target input speech and the target registered speech belong to the same speaker, intent recognition should be performed on the target input speech. The target input speech is then fed into a speech feature extraction model, which extracts speech features corresponding to the target input speech. These speech features include the speech content of the target input speech.
[0099] The extracted speech features are input into a large language model, which is then used for recognition to obtain the target intent and target semantic slots. The target intent and target semantic slots are then output as the intent recognition result. The large language model can be a classic and widely used model. Furthermore, large models with different parameter sizes can be selected based on computing power and memory requirements; no specific limitation is imposed.
[0100] In this embodiment, intent represents the user's purpose in the target input speech. For example, the target input speech is "Check the weather in Beijing today," and the intent of this target input speech is "Check the weather." Another example is the target input speech "Play a song," and the intent of this target input speech is "Recommend music."
[0101] Each intent can be configured with one or more semantic slots. Semantic slots are closely related to the intent and represent the specific information and parameters required during implementation. Specifically, a semantic slot includes a slot and its corresponding slot value. A slot is a container for an entity or slot value, and a slot value is a keyword in the target input speech. A slot refers to the key information that needs to be collected from the target input speech. For example, for the intent to query the weather, the configured slots can include location slots and time slots. The location slot determines which location's weather needs to be queried, and the time slot determines when the weather needs to be queried.
[0102] A slot value refers to the specific parameters of a slot, also known as the entity of a slot. For example, if the target input voice is "Check the weather in Beijing today", a location slot and a time slot can be extracted from this voice. The entity of the location slot is "Beijing", and the entity of the time slot is the current system date.
[0103] In this embodiment, the corresponding intents differ depending on the application scenario of the method described in this disclosure. For example, if the application scenario involves in-vehicle voice recognition of a user's voice commands, the corresponding intents include navigation, playing music, and making phone calls. If the application scenario involves controlling smart home devices based on voice commands, the corresponding intents include controlling lights, air conditioning, and television. If the application scenario involves recognizing a user's health needs, the corresponding intents include making an appointment for an appointment and finding available appointments. If the application scenario involves an electronic device recognizing a user's voice intent, the corresponding intents include playing music, setting an alarm clock, and checking the weather. If the application scenario involves recognizing a user's work needs, the corresponding intents include scheduling meetings, sending emails, and finding files.
[0104] Furthermore, the corresponding semantic slots differ depending on the intent. For example, if the intent is to query the weather, the corresponding semantic slots would be location, time, etc. If the intent is to play music, the corresponding semantic slots would be song, artist, album, etc. If the intent is to set a reminder, the corresponding semantic slots would be time, event, etc. If the intent is to query knowledge, the corresponding semantic slot would be the query topic. If the intent is to open an application, the corresponding semantic slot would be the application name, etc.
[0105] The above scheme, when determining the match between the target input speech and the target registered speech, extracts the target input speech through a speech feature extraction model, filters out the key features required for subsequent intent recognition, avoids irrelevant features from affecting intent recognition, and thus improves the accuracy of intent recognition.
[0106] In some embodiments, the speech feature extraction model includes a speech feature extraction module and a second feature mapping module, such as... Figure 7 As shown, Figure 7 A schematic diagram of the intent recognition model is shown. In step 10221, the target input speech is input to the speech feature extraction model, and the speech feature extraction model is used to extract the speech features corresponding to the target input speech. Specifically, this includes:
[0107] Step 10A: Input the target input speech into the speech feature extraction module, and use the speech feature extraction module to extract and process the target input speech to obtain initial speech features;
[0108] Step 10B: Input the initial speech features into the second feature mapping module, and use the second feature mapping module to map the initial speech features to obtain speech features.
[0109] In specific implementation, the speech feature extraction model includes a speech feature extraction module and a second feature mapping module. The speech feature extraction module is used to extract initial speech features, and the second feature mapping module is used to map high-dimensional speech feature data to the low-dimensional space where the large language model is located, so that the speech features output by the second feature mapping module can be input into the large language model for processing.
[0110] The target input speech is input to the speech feature extraction module, which extracts and processes the speech to obtain initial speech features. The initial speech features are then input to the second feature mapping module, which maps the initial speech features to obtain further speech features. In this embodiment, the second feature mapping module can also be an adapter, and its specific structure is the same as that of the first feature mapping module, which will not be described again here.
[0111] In this embodiment, the speech feature extraction module can be a speech tokenizer. The speech tokenizer uses residual vector quantization to convert continuous voiceprint features into discrete tokens. The residual quantization can be adjusted according to task requirements, achieving a balance between computational complexity and reconstruction quality. Furthermore, using a tokenizer can compress the voiceprint signal, reducing storage and transmission costs.
[0112] In some embodiments, the speech feature extraction module includes a speech encoder and a second feature compression integration module, such as... Figure 8 As shown, Figure 8 A schematic diagram of the intent recognition model is shown. In step 10A, the target input speech is input to the speech feature extraction module, and the speech feature extraction module is used to extract and process the target input speech to obtain initial speech features, specifically including:
[0113] Step A: Input the target input speech into the speech encoder, and use the speech encoder to extract the target input speech to obtain at least one first speech feature;
[0114] Step B involves inputting at least one first speech feature into the second feature integration and compression module, and then using the second feature integration and compression module to perform fusion processing to obtain initial speech features.
[0115] In specific implementation, the speech feature extraction module includes a speech encoder and a second feature compression integration module. The main purpose of the speech encoder is to extract feature representations from speech. It can be a speech model trained on large-scale data, such as the Wshiper encoder or the HuBERT self-supervised learning model, or other encoders trained on large-scale data models, such as the Paraformer encoder. There is no specific limitation.
[0116] Since a speech encoder may have a multi-layered structure, each layer outputs a corresponding feature. A second feature compression module then integrates and compresses all the features output by the speech encoder to obtain a one-dimensional feature. Specifically, the second feature compression and integration module is used to compress and fuse the features output by the speech encoder to obtain a one-dimensional feature, which is the initial speech feature.
[0117] Specifically, the target input speech is input to a speech encoder, which extracts at least one first speech feature. The obtained first speech feature is then input to a second feature integration and compression module, which performs fusion processing on the at least one first speech feature to obtain initial speech features.
[0118] In some embodiments, after obtaining the intent recognition result, a corresponding function can be executed based on the intent recognition result to meet the user's needs. That is, after outputting the target intent and the target semantic slot as the intent recognition result, the method further includes:
[0119] Step 103: Determine the control command based on the target intent and the target semantic slot;
[0120] Step 104: According to the control instruction, control the target function execution component to execute the corresponding function.
[0121] In practice, the target semantic slots are filled according to the target intent to obtain control instructions, and then the target function execution component corresponding to the control instructions is determined. The target function execution component is then controlled to execute the corresponding function according to the control instructions.
[0122] For example, the target intent is determined to be "play Zhang San's song A", and the target semantic slot is the song and artist. The target semantic slot is filled according to the target intent, resulting in the control instruction to play Zhang San's song A using the multimedia function. The target function execution component is determined to be the vehicle multimedia component, and the vehicle multimedia component is controlled to play Zhang San's song A.
[0123] The above scheme determines control commands based on the target intent and the target semantic slot, and executes the corresponding functions. This achieves accurate recognition of the corresponding text commands based on the user's voice, and executes the corresponding operations based on the text commands, thereby achieving efficient and accurate voice interaction and improving the user experience.
[0124] Based on the same inventive concept, another embodiment of this disclosure provides an intent recognition model training method for training the intent recognition model described in the above embodiment, characterized in that, as Figure 9 As shown, the method specifically includes:
[0125] Step 201: Obtain the training dataset, wherein each piece of training data in the training dataset includes training input speech data, training registered speech data, and text annotation data;
[0126] Step 202: For each piece of training data, determine whether the training input speech data matches the training registered speech data, and determine the target format corresponding to the training data based on the determination result;
[0127] Step 203: Obtain the initial intent recognition model, and train the initial intent recognition model using the training data in the target format until the preset training termination condition is met, thereby obtaining the intent recognition model.
[0128] In practice, a training dataset is obtained, wherein each piece of training data in the training dataset includes training input speech data, training registered speech data, and text annotation data.
[0129] In this embodiment, the training input speech data and the training registration speech data are preferably in WAV format, which is a digital audio format for storing sound waveforms. If other formats are used, they are converted to WAV format using a format conversion tool. Furthermore, in this embodiment, the sampling rate of the training input speech data and the training registration speech data is 16000, and 16-bit PCM encoding is used.
[0130] The text annotation data includes the storage address of the training input speech data, the speaker information of the training input speech data, the text content of the training input speech data, the actual training intent corresponding to the training input speech data, and the actual semantic slot information corresponding to the actual training intent.
[0131] In this embodiment, the text annotation data is formatted and stored in JSON format. For example, the text annotation data is:
[0132] {
[0133] "wav_path":"path / to / audio.wav",
[0134] “speaker”:“S00001”,
[0135] “text”:“Play Zhang San’s song”,
[0136] "intent":"<|play_music|>",
[0137] “slot”:{
[0138] Singer: "Zhang San"
[0139] }
[0140] }
[0141] Because intent recognition models can determine whether the target input speech matches the target registered speech, they need to be specifically trained to address this capability. Specifically:
[0142] For each piece of training data, it is determined whether the training input speech data matches the training registered speech data, and the target format corresponding to the training data is determined based on the determination result. The target format is the format in which the training data is input into the initial intent recognition model.
[0143] Specifically, if the training input speech data does not match the training registration speech data, it means that the user who outputs the training input speech data and the user who outputs the training registration speech data are not the same person, and the first label is used for labeling.
[0144] For example, if the first marker is <|speaker_mismatch|>, then when the training input speech data does not match the training registered speech data, the target format corresponding to the training data is determined to be "".<speaker_in><speaker_ref> Verify if the input voice matches the registered voice.<speech_in> If a match is successful, the intent is identified and information <|speaker_mismatch|> is extracted.
[0145] If the training input speech data matches the training registration speech data, it means that the user who outputs the training input speech data and the user who outputs the training registration speech data are the same person, and the second label is used for labeling.
[0146] For example, if the second tag is <|speaker_match|>, the content of the speech is playing Zhang San's song, the intent is <|play_music|>, and the corresponding semantic slot is the singer Zhang San, then the corresponding training data format is as follows:<speaker_in><speaker_ref> Verify if the input voice matches the registered voice.<speech_in> If a match is successful, the intent is identified and information is extracted. The song "<|speaker_match|>" by Zhang San ("<|singer_eos|>") is played.
[0147] in<speaker_in> Speaker features are used to embed training input speech data.<speaker_ref> Registered speaker features used to embed training registered speech data<speech_in> This is used to embed input speech features, and semantic slots are marked with <|singer_sos|> and <|singer_eos|>. This constructs prompts and training data for the large model, allowing the large model to learn and imitate the reasoning steps of the data, thereby forming a chain of thought and effectively solving the problem.
[0148] In some embodiments, the initial intent recognition model includes an initial speech feature extraction model, an initial voiceprint feature extraction model, and an initial large language model. Step 203 involves training the initial intent recognition model using the training data in the target format until a preset training termination condition is met, thereby obtaining the intent recognition model. Specifically, this includes:
[0149] Step 2031: Train the initial speech feature extraction model and the initial voiceprint feature extraction model using the training data in the target format until the first preset training termination condition is met, and obtain the speech feature extraction model and the voiceprint feature extraction model.
[0150] Step 2032: Freeze the model parameters of the speech feature extraction model and the voiceprint feature extraction model, and train the initial large language model using the training data in the target format until the second preset training termination condition is met to obtain the large language model.
[0151] In practice, a multi-stage training approach is adopted when training the initial intent recognition. First, the initial speech feature extraction model and the initial voiceprint feature extraction model are trained using training data in the target format until the first preset training termination condition is met, thus obtaining the speech feature extraction model and the voiceprint feature extraction model.
[0152] The first preset training termination condition includes at least one of the following: determining that all training data in all target formats are input into the initial speech feature extraction model and the initial voiceprint feature extraction model for training; determining that the loss functions of the initial speech feature extraction model and the initial voiceprint feature extraction model converge to the first convergence threshold; or determining that the initial speech feature extraction model and the initial voiceprint feature extraction model are iteratively trained to the first preset number of iterations.
[0153] For example, the first preset training termination condition is that all training data in the target format are input into the initial speech feature extraction model and the initial speaker feature extraction model for training:
[0154] There are fifty sets of training data in the target format. The first preset training termination condition is that all the training data in the target format has been input into the initial speech feature extraction model and the initial voiceprint feature extraction model for training. That is, when all fifty sets of data have been input into the initial speech feature extraction model and the initial voiceprint feature extraction model, there is no target format training data that has not yet been input into the initial speech feature extraction model and the initial voiceprint feature extraction model. At this time, the training of the initial speech feature extraction model and the initial voiceprint feature extraction model is determined to be completed, and the speech feature extraction model and the voiceprint feature extraction model are obtained.
[0155] In another example, the first preset training termination condition is that the loss functions of both the initial speech feature extraction model and the initial speaker feature extraction model converge to a first convergence threshold:
[0156] Training data in the target format is input into the initial speech feature extraction model and the initial speaker fingerprint feature extraction model for training, and the training results are output. Based on the training results and the training intent recognition results, the loss function includes at least one of the following: mean squared error loss function, cross-entropy loss function, logarithmic loss function, exponential loss function, squared loss function, or absolute value loss function, etc. When the loss functions of the initial speech feature extraction model and the initial speaker fingerprint feature extraction model both converge to a first convergence threshold, it is determined that the first preset training termination condition is met, and the speech feature extraction model and the speaker fingerprint feature extraction model are obtained.
[0157] In another example, the first preset training termination condition is to determine the initial speech feature extraction model and the initial voiceprint feature extraction model and iterate them for a first preset number of iterations.
[0158] The training data in the target format is input into the initial speech feature extraction model and the initial voiceprint feature extraction model for iterative training. The number of iterations is recorded. When the number of iterations is equal to the first preset number of iterations, the first preset training termination condition is met, and the speech feature extraction model and the voiceprint feature extraction model are obtained.
[0159] After obtaining the speech feature extraction model and the voiceprint feature extraction model, the model parameters of the speech feature extraction model and the voiceprint feature extraction model are frozen. The initial large language model is trained using training data in the target format until the second preset training termination condition is met, and the large language model is obtained.
[0160] In this embodiment, the second preset training termination condition includes at least one of the following: determining that all training data in the target format have been input into the initial large language model for training, determining that the loss function of the initial large language model has converged to the second convergence threshold, or determining that the initial large language model has been iteratively trained to the second preset number of iterations.
[0161] For example, during model training, x = x0x1…x m The sequence of inputs to the model, such as training input speech data, training registration speech data, and text annotation data, y = y0y1…y n The corresponding target output sequence refers to the actual training intent corresponding to the training input speech data and the actual semantic slot information corresponding to the actual training intent. In supervised training fine-tuning, the training objective is to maximize the probability of outputting the target output sequence given the input sequence, as shown in the following formula:
[0162]
[0163] Where θ is the model parameter, D is the training sample dataset, i.e., the training data in the target format, P(y|x,θ) is the conditional probability, which represents the probability that the model outputs y given the input x and the model parameter θ, and L is the probability, which represents the log probability that the model predicts the true label y given the input x and the parameter θ.
[0164] It should be noted that the method of this disclosure embodiment can be executed by a single device, such as a computer or server. The method of this embodiment can also be applied to a distributed scenario, where multiple devices cooperate to complete the task. In such a distributed scenario, one of these devices may execute only one or more steps of the method of this disclosure embodiment, and the multiple devices will interact with each other to complete the method described.
[0165] It should be noted that the above description describes some embodiments of this disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in a different order than that shown in the above embodiments and still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0166] Based on the same inventive concept, corresponding to any of the above-described embodiments, this disclosure also provides an intent recognition device.
[0167] refer to Figure 10 The intent recognition device includes:
[0168] The receiving module 301 is configured to receive the target input voice and obtain the preset target registration voice;
[0169] Model processing module 302 is configured to input the target input speech and the target registered speech into a pre-trained intent recognition model, and process them through the intent recognition model;
[0170] The model processing module further includes:
[0171] The judgment unit is configured to determine whether the target input speech matches the target registered speech;
[0172] The recognition unit is configured to perform intent recognition based on the target input speech and output the intent recognition result in response to the matching of the target input speech with the target registered speech.
[0173] Specifically, the intent recognition model includes a voiceprint feature extraction model and a large language model; the judgment unit is specifically configured as follows:
[0174] The target input speech and the target registered speech are respectively input into the voiceprint feature extraction model. After processing by the voiceprint feature extraction model, the speaker features corresponding to the target input speech and the registered speaker features corresponding to the target registered speech are obtained.
[0175] The speaker features and the registered speaker features are input into the large language model, and the large language model is used to determine whether the target input speech matches the target registered speech.
[0176] Specifically, the voiceprint feature extraction model includes a voiceprint feature extraction module and a first feature mapping module, and the judgment unit is specifically configured as follows:
[0177] The target input speech and the target registered speech are respectively input into the voiceprint feature extraction module. The voiceprint feature extraction module is used to extract and process the target input speech and the target registered speech to obtain the initial speaker features corresponding to the target input speech and the initial registered speaker features corresponding to the target registered speech.
[0178] The initial speaker features and the initial registered speaker features are respectively input into the first feature mapping module. The first feature mapping module is used to map the initial speaker features and the initial registered speaker features to obtain the speaker features corresponding to the target input speech and the registered speaker features corresponding to the target registered speech.
[0179] Specifically, the voiceprint feature extraction module includes a voiceprint encoder and a first feature compression integration module, and the judgment unit is specifically configured as follows:
[0180] The target input speech and the target registered speech are respectively input into the voiceprint encoder, and extracted using the voiceprint encoder to obtain at least one first speaker feature corresponding to the target input speech and at least one first registered speaker feature corresponding to the target registered speech.
[0181] The at least one first speaker feature is input into the first feature integration and compression module, and the first feature integration and compression module is used to perform fusion processing to obtain the initial speaker feature;
[0182] The at least one first registered speaker feature is input into the first feature integration and compression module, and the first feature integration and compression module is used to perform fusion processing to obtain the initial registered speaker feature.
[0183] Specifically, the intent recognition model further includes a speech feature extraction model, and the recognition unit is specifically configured as follows:
[0184] In response to the target input speech matching the target registered speech, the target input speech is input to the speech feature extraction model, and the speech feature extraction model is used to extract the speech features corresponding to the target input speech;
[0185] The speech features are input into the large language model, and the speech features are processed by the large language model to obtain the target intent and target semantic slots;
[0186] The target intent and the target semantic slot are output as intent recognition results.
[0187] Specifically, the speech feature extraction model includes a speech feature extraction module and a second feature mapping module, and the recognition unit is specifically configured as follows:
[0188] The target input speech is input to the speech feature extraction module, and the speech feature extraction module is used to extract and process the target input speech to obtain initial speech features;
[0189] The initial speech features are input into the second feature mapping module, and the initial speech features are mapped using the second feature mapping module to obtain speech features.
[0190] Specifically, the speech feature extraction module includes a speech encoder and a second feature compression integration module, and the recognition unit is specifically configured as follows:
[0191] The target input speech is input into the speech encoder, and the speech encoder is used to extract the target input speech to obtain at least one first speech feature;
[0192] The at least one first speech feature is input into the second feature integration and compression module, and the second feature integration and compression module is used to perform fusion processing to obtain the initial speech features.
[0193] Specifically, the device further includes a first control module, which is specifically configured to:
[0194] Determine control instructions based on the target intent and the target semantic slot;
[0195] According to the control instructions, the target function execution component is controlled to perform the corresponding function.
[0196] Specifically, the device further includes a second control module, which is specifically configured to:
[0197] In response to a mismatch between the target input speech and the target registered speech, intent recognition is stopped.
[0198] Based on the same inventive concept, corresponding to the methods of any of the above embodiments, this disclosure also provides an intent recognition device for training an intent recognition model in any of the above embodiments.
[0199] refer to Figure 11 The intent recognition model training device includes:
[0200] The acquisition module 401 is configured to acquire a training dataset, wherein each piece of training data in the training dataset includes training input speech data, training registered speech data, and text annotation data.
[0201] The judgment module 402 is configured to determine whether the training input speech data matches the training registered speech data for each piece of training data, and determine the target format corresponding to the training data based on the judgment result.
[0202] The training module 403 is configured to acquire an initial intent recognition model, train the initial intent recognition model using the training data in the target format until a preset training termination condition is met, and obtain the intent recognition model.
[0203] Specifically, the initial intent recognition model includes an initial speech feature extraction model, an initial voiceprint feature extraction model, and an initial large language model, and the training module 403 is specifically configured as follows:
[0204] The initial speech feature extraction model and the initial voiceprint feature extraction model are trained using the training data in the target format until the first preset training termination condition is met, thereby obtaining the speech feature extraction model and the voiceprint feature extraction model.
[0205] Freeze the model parameters of the speech feature extraction model and the voiceprint feature extraction model, and train the initial large language model using the training data in the target format until the second preset training termination condition is met, thereby obtaining the large language model.
[0206] For ease of description, the above apparatus is described in terms of its functions, divided into various modules. Of course, in implementing this disclosure, the functions of each module can be implemented in one or more software and / or hardware.
[0207] The system described above is used to implement the corresponding intent recognition method in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0208] Based on the same inventive concept, corresponding to the methods of any of the above embodiments, this disclosure also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the intent recognition method described in any of the above embodiments.
[0209] Figure 12 This embodiment illustrates a more specific hardware structure of an electronic device, which may include a processor 1010, a memory 1020, an input / output interface 1020, a communication interface 1040, and a bus 1050. The processor 1010, memory 1020, input / output interface 1020, and communication interface 1040 are interconnected internally via the bus 1050.
[0210] The processor 1010 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.
[0211] The memory 1020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 1020 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program code is stored in the memory 1020 and is called and executed by the processor 1010.
[0212] The input / output interface 1020 is used to connect input / output modules to realize information input and output. Input / output modules can be configured as components within the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Input devices may include keyboards, mice, touchscreens, microphones, various sensors, etc., while output devices may include displays, speakers, vibrators, indicator lights, etc.
[0213] The communication interface 1040 is used to connect a communication module (not shown in the figure) to enable communication between this device and other devices. The communication module can communicate via wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, Wi-Fi, Bluetooth, etc.).
[0214] Bus 1050 includes a pathway for transmitting information between various components of the device (e.g., processor 1010, memory 1020, input / output interface 1020, and communication interface 1040).
[0215] It should be noted that although the above-described device only shows the processor 1010, memory 1020, input / output interface 1020, communication interface 1040, and bus 1050, in specific implementations, the device may also include other components necessary for normal operation. Furthermore, those skilled in the art will understand that the above-described device may only include the components necessary for implementing the embodiments of this specification, and not necessarily all the components shown in the figures.
[0216] The electronic devices described above are used to implement the corresponding intent recognition methods in any of the foregoing embodiments and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0217] Based on the same inventive concept, corresponding to the methods of any of the above embodiments, this disclosure also provides a non-transitory computer-readable storage medium that stores computer instructions for causing the computer to execute the intent recognition method as described in any of the above embodiments.
[0218] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.
[0219] The computer instructions stored in the storage medium of the above embodiments are used to cause the computer to execute the intent recognition method as described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0220] Based on the same concept, corresponding to any of the above embodiments, this application also provides a computer program product, including computer program instructions, which, when run on a computer, cause the computer to perform the method described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0221] It is understood that before using the technical solutions of the various embodiments in this disclosure, users will be informed of the type, scope of use, and usage scenarios of the personal information involved in an appropriate manner, and user authorization will be obtained.
[0222] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose, based on the prompt message, whether to provide personal information to the software or hardware such as electronic devices, applications, servers, or storage media performing the operations of this disclosed technical solution.
[0223] As an optional but not limited implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0224] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0225] Those skilled in the art should understand that the discussion of any of the above embodiments is merely exemplary and is not intended to imply that the scope of this disclosure (including the claims) is limited to these examples; within the framework of this disclosure, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations of different aspects of the embodiments of this disclosure as described above, which are not provided in detail for the sake of brevity.
[0226] Additionally, to simplify the description and discussion, and to avoid obscuring the embodiments of this disclosure, the provided drawings may or may not show well-known power / ground connections to integrated circuit (IC) chips and other components. Furthermore, the apparatus may be shown in block diagram form to avoid obscuring the embodiments of this disclosure, and this also takes into account the fact that the details of implementation of these block diagram apparatuses are highly dependent on the platform on which the embodiments of this disclosure will be implemented (i.e., these details should be fully understood by those skilled in the art). While specific details (e.g., circuitry) have been set forth to describe exemplary embodiments of this disclosure, it will be apparent to those skilled in the art that the embodiments of this disclosure may be implemented without these specific details or with variations thereof. Therefore, these descriptions should be considered illustrative rather than restrictive.
[0227] Although this disclosure has been described in conjunction with specific embodiments thereof, many substitutions, modifications, and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may be used with the embodiments discussed.
[0228] This disclosure is intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. An intent recognition method, characterized in that, include: In response to receiving the target input voice, obtain the preset target registration voice; The target input speech and the target registered speech are input into a pre-trained intent recognition model and processed by the intent recognition model. The process, via the intent recognition model, further includes: Determine whether the target input speech matches the target registered speech; In response to the target input speech matching the target registered speech, intent recognition is performed based on the target input speech and the intent recognition result is output.
2. The method according to claim 1, characterized in that, The intent recognition model includes a voiceprint feature extraction model and a large language model; Determining whether the target input speech matches the target registered speech includes: The target input speech and the target registered speech are respectively input into the voiceprint feature extraction model. After processing by the voiceprint feature extraction model, the speaker features corresponding to the target input speech and the registered speaker features corresponding to the target registered speech are obtained. The speaker features and the registered speaker features are input into the large language model, and the large language model is used to determine whether the target input speech matches the target registered speech.
3. The method according to claim 2, characterized in that, The voiceprint feature extraction model includes a voiceprint feature extraction module and a first feature mapping module; The step of inputting the target input speech and the target registered speech into the voiceprint feature extraction model, and processing them through the voiceprint feature extraction model to obtain the speaker features corresponding to the target input speech and the registered speaker features corresponding to the target registered speech, includes: The target input speech and the target registered speech are respectively input into the voiceprint feature extraction module. The voiceprint feature extraction module is used to extract and process the target input speech and the target registered speech to obtain the initial speaker features corresponding to the target input speech and the initial registered speaker features corresponding to the target registered speech. The initial speaker features and the initial registered speaker features are respectively input into the first feature mapping module. The first feature mapping module is used to map the initial speaker features and the initial registered speaker features to obtain the speaker features corresponding to the target input speech and the registered speaker features corresponding to the target registered speech.
4. The method according to claim 3, characterized in that, The voiceprint feature extraction module includes a voiceprint encoder and a first feature compression integration module; The step of inputting the target input speech and the target registered speech into the voiceprint feature extraction module, and using the voiceprint feature extraction module to extract and process the target input speech and the target registered speech respectively, to obtain the initial speaker features corresponding to the target input speech and the initial registered speaker features corresponding to the target registered speech, includes: The target input speech and the target registered speech are respectively input into the voiceprint encoder, and extracted using the voiceprint encoder to obtain at least one first speaker feature corresponding to the target input speech and at least one first registered speaker feature corresponding to the target registered speech. The at least one first speaker feature is input into the first feature integration and compression module, and the first feature integration and compression module is used to perform fusion processing to obtain the initial speaker feature; The at least one first registered speaker feature is input into the first feature integration and compression module, and the first feature integration and compression module is used to perform fusion processing to obtain the initial registered speaker feature.
5. The method according to claim 2, characterized in that, The intent recognition model also includes a speech feature extraction model; The step of responding to the matching of the target input speech with the target registered speech, performing intent recognition based on the target input speech, and outputting the intent recognition result includes: In response to the target input speech matching the target registered speech, the target input speech is input to the speech feature extraction model, and the speech feature extraction model is used to extract the speech features corresponding to the target input speech; The speech features are input into the large language model, and the speech features are processed by the large language model to obtain the target intent and target semantic slots; The target intent and the target semantic slot are output as intent recognition results.
6. The method according to claim 5, characterized in that, The speech feature extraction model includes a speech feature extraction module and a second feature mapping module; The step of inputting the target input speech into the speech feature extraction model and using the speech feature extraction model to extract the speech features corresponding to the target input speech includes: The target input speech is input to the speech feature extraction module, and the speech feature extraction module is used to extract and process the target input speech to obtain initial speech features; The initial speech features are input into the second feature mapping module, and the initial speech features are mapped using the second feature mapping module to obtain speech features.
7. The method according to claim 6, characterized in that, The speech feature extraction module includes a speech encoder and a second feature compression integration module; The step of inputting the target input speech to the speech feature extraction module and using the speech feature extraction module to extract and process the target input speech to obtain initial speech features includes: The target input speech is input into the speech encoder, and the speech encoder is used to extract the target input speech to obtain at least one first speech feature; The at least one first speech feature is input into the second feature integration and compression module, and the second feature integration and compression module is used to perform fusion processing to obtain the initial speech features.
8. A method for training an intent recognition model, used to train the intent recognition model according to any one of claims 1-7, characterized in that, include: Obtain the training dataset, wherein each piece of training data in the training dataset includes training input speech data, training registered speech data, and text annotation data; For each piece of training data, determine whether the training input speech data matches the training registered speech data, and determine the target format corresponding to the training data based on the determination result; An initial intent recognition model is obtained, and the initial intent recognition model is trained using the training data in the target format until a preset training termination condition is met, thereby obtaining the intent recognition model.
9. The method according to claim 8, characterized in that, The initial intent recognition model includes an initial speech feature extraction model, an initial voiceprint feature extraction model, and an initial large language model. The step of training the initial intent recognition model using the training data in the target format until a preset training termination condition is met to obtain the intent recognition model includes: The initial speech feature extraction model and the initial voiceprint feature extraction model are trained using the training data in the target format until the first preset training termination condition is met, thereby obtaining the speech feature extraction model and the voiceprint feature extraction model. Freeze the model parameters of the speech feature extraction model and the voiceprint feature extraction model, and train the initial large language model using the training data in the target format until the second preset training termination condition is met, thereby obtaining the large language model.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the method as claimed in any one of claims 1 to 9.