Wake-up word processing method, wake-up word processing device and computer storage medium
By cutting the end-to-end acoustic models on the cloud side, generating an acoustic encoder model and issuing it to the end-side equipment, the detection efficiency and accuracy problems of traditional models under personalized needs and limited computing power are solved, and efficient and accurate wake-up word recognition is achieved.
Patent Information
- Application Number
- CN202510583093.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-06-03
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional fixed keyword models are difficult to cope with personalized needs in terms of scale, adaptability and robustness, especially due to the limited computing power of the end-side equipment, the wake-up word detection effect and efficiency are not high.
A wake-up word processing method is proposed, which obtains and maps the registered word sequence through the end-to-end acoustic model on the cloud side, locates the relevant core networks and generates an acoustic encoder model, reduces the amount of model calculation, and sends the registered embedding vector and acoustic encoder model to the end-side device for wake-up word recognition.
The acoustic model is cut through the cloud side and sent to the end-side device, which improves the wake-up word detection efficiency and recognition accuracy of the end-side device, meets personalized needs and takes into account recognition accuracy and low power consumption.
Smart Images

Figure CN120089130A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of human-computer interaction technology, and particularly to a wake word processing method, a wake word processing device, and a computer storage medium. Background Art
[0002] With the large-scale popularization of voice interaction products such as smart speakers and mobile phone voice assistants, in order to meet the diverse usage habits and flexibility of users, an open solution that allows users to define wake words by themselves has become the mainstream trend. However, traditional fixed keyword models are difficult to cope with the growing personalized needs in terms of scale, adaptability, and robustness. To solve this problem, researchers have proposed various multi-modal fusion technologies that combine voice and text information to improve the accuracy and customization ability of keyword detection. At the same time, the high-speed development of cloud computing power and the limited computing power of end-side devices form a sharp contrast in the keyword recognition system.
[0003] The limited computing power of end-side devices results in low detection effect and efficiency of wake words when users use the devices for human-computer interaction. Summary of the Invention
[0004] To solve the above technical problems, this application proposes a wake word processing method, a wake word processing device, and a computer storage medium.
[0005] To solve the above technical problems, this application proposes a wake word processing method. The wake word processing method is applied to the cloud side, and the wake word processing method includes: In response to a user registration instruction, input the user registration information into an end-to-end acoustic model to obtain a registration word element sequence; Map the registration word element sequence to a registration embedding vector; Use the registration word element sequence to locate the relevant core network of the end-to-end acoustic model; Use the relevant core network to generate an acoustic encoder model, where the model calculation amount of the acoustic encoder model is less than that of the end-to-end acoustic model; Send the registration embedding vector and the acoustic encoder model to the end-side device for wake word recognition.
[0006] Wherein, the user registration information is voice registration information or text registration information; The step of inputting the user registration information into the end-to-end acoustic model to obtain a registration word element sequence includes: Input the text registration information into the text modeling unit of the end-to-end acoustic model, obtain the text word elements of each text in the text registration information, and combine them into the registration word element sequence; Alternatively, input the voice registration information into the end-to-end acoustic model for recognition to obtain the registered token sequence.
[0007] Among them, the method of using the registered token sequence to locate the relevant core network of the end-to-end acoustic model includes: Obtain the acoustic unit activation paths of each token in the end-to-end acoustic model; Extract the relevant acoustic unit activation paths based on the registered token sequence, prune the acoustic unit activation paths of the remaining tokens or set their weights to zero to obtain the relevant core network corresponding to the registered token sequence.
[0008] Among them, after using the registered token sequence to locate the relevant core network of the end-to-end acoustic model, the wake word processing method further includes: Reduce the network layer depth and / or parameter dimension of the relevant core network.
[0009] To solve the above technical problems, the present application also proposes another wake word processing method. The wake word processing method is applied to a wake word processing system. Among them, the wake word processing system includes a cloud side and an end-side device; the wake word processing method includes: The end-side device uploads the user registration instruction and user registration information to the cloud side; The cloud side inputs the user registration information into the end-to-end acoustic model to obtain the registered token sequence; The cloud side maps the registered token sequence to a registered embedding vector; The cloud side uses the registered token sequence to locate the relevant core network of the end-to-end acoustic model; The cloud side uses the relevant core network to generate an acoustic encoder model, where the model calculation amount of the acoustic encoder model is less than that of the end-to-end acoustic model; The cloud side sends the registered embedding vector and the acoustic encoder model to the end-side device; The end-side device encodes the user real-time input using the acoustic encoder model to obtain a real-time embedding vector, and compares the registered embedding vector with the real-time embedding vector; When the vector comparison is successful, the end-side device wakes up the corresponding device with the wake word corresponding to the user real-time input.
[0010] Among them, the user real-time input is real-time voice input; The end-side device encodes the user real-time input using the acoustic encoder model to obtain a real-time embedding vector, including: The end-side device obtains a number of audio slices based on the real-time voice input; The edge device inputs the several audio slices into the acoustic encoder model to extract an audio feature matrix; The edge device inputs the audio feature matrix into the acoustic pronunciation classification branch to obtain the posterior probability of the pronunciation of the text token; The edge device segments the real-time speech input into several pronunciation paragraphs and the pronunciation duration of each text token according to the maximum posterior probability of the pronunciation of each audio slice; The edge device embeds the real-time speech input according to the several pronunciation paragraphs and the pronunciation duration to extract the audio-text embedding vector of the real-time speech input.
[0011] Wherein, the edge device embeds the real-time speech input according to the several pronunciation paragraphs and the pronunciation duration to extract the audio-text embedding vector of the real-time speech input, including: The edge device obtains the frame sequence covered by each text token according to the several pronunciation paragraphs and the pronunciation duration; The edge device aggregates the speech frames of the real-time speech input in the frame sequence to obtain the concentrated speech feature corresponding to each text token; The edge device associates the text token with the corresponding concentrated speech feature to generate the audio-text embedding vector.
[0012] Wherein, after the cloud edge device sends the registration embedding vector and the acoustic encoder model to the edge device, the wake-up word processing method further includes: The edge device reduces the network layer depth and / or parameter dimension of the acoustic encoder model according to the device performance.
[0013] To solve the above technical problems, the present application also proposes a wake-up word processing device, which includes a memory and a processor coupled to the memory; wherein, the memory is used to store program data, and the processor is used to execute the program data to implement the wake-up word processing method as described above.
[0014] To solve the above technical problems, the present application also proposes a computer storage medium, which is used to store program data, and when the program data is executed by a computer, it is used to implement the wake-up word processing method as described above.
[0015] Compared with the prior art, the beneficial effects of the present application are as follows: In response to a user registration instruction, the cloud side inputs the user registration information into the end-to-end acoustic model to obtain a registration token sequence; maps the registration token sequence to a registration embedding vector; uses the registration token sequence to locate the relevant core network of the end-to-end acoustic model; generates an acoustic encoder model by using the relevant core network, wherein the model calculation amount of the acoustic encoder model is less than that of the end-to-end acoustic model; and sends the registration embedding vector and the acoustic encoder model to the end-side device for wake-word recognition. Through the above wake-word processing method, the cloud side trims the acoustic model according to the user registration information and then sends it to the end-side device, improving the wake-word detection efficiency and recognition accuracy of the end-side device. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. Among them: Figure 1 is a schematic flowchart of an embodiment of the wake-word processing method provided by the present application; Figure 2 is a schematic flowchart of the acquisition and cloud processing solution of the registration information provided by the present application; Figure 3 is Figure 1 a specific schematic flowchart of step S13 of the wake-word processing method shown; Figure 4 is a schematic flowchart of another embodiment of the wake-word processing method provided by the present application; Figure 5 is a schematic flowchart of the real-time input processing of the end-side model provided by the present application; Figure 6 is Figure 4 a specific schematic flowchart of step S27 of the wake-word processing method shown; Figure 7 is a schematic structural diagram of an embodiment of the wake-word processing device provided by the present application; Figure 8 is a schematic structural diagram of an embodiment of the computer storage medium provided by the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0017] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts belong to the scope of protection of the present application.
[0018] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present application and the above accompanying drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described here, for example, can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products or devices.
[0019] Based on the idea of end-cloud collaboration, the wake-up word processing method of the present application introduces the dual-modal fusion of voice registration and text registration, combines the development and deployment of multi-modal custom wake-up words, and provides users with flexible and efficient keyword detection capabilities.
[0020] Aiming at the pain points of the prior art, the wake-up word processing method of the present application proposes a "multi-modal custom wake-up word system with end-cloud collaboration". A large-scale acoustic model and a text encoder are deployed in the cloud, supporting text and voice dual-modal registration. At the same time, the cloud computing power is fully utilized to train and prune the model. By only sending the sub-network and embedding vector related to the target wake-up word to the end side, accurate real-time detection can be completed under limited computing power conditions, taking into account both recognition accuracy and low-power consumption requirements. The wake-up word processing method of the present application effectively solves the problem that it is difficult to maintain high accuracy in the prior art when the single-modal adaptability is poor and the end-side computing power is limited, enabling users to flexibly change the wake-up word at any time according to personal preferences or business scenarios and ensuring the detection performance.
[0021] In actual deployment, the wake word processing method of the present application supports both text registration and voice registration modes, and completes the initialization and detection of custom wake words through the cooperation of the terminal and the cloud. The main purpose of this terminal-cloud collaborative architecture is to upload the registration information input by the user to the cloud, where the large-scale acoustic network is trimmed and compressed as necessary, and then sent to the lightweight model on the terminal side. The large-scale text encoder model is used to generate registration embedding vectors to ensure that the detection task can be completed efficiently and accurately on local devices with limited computing power, memory, and power consumption.
[0022] For details, please refer to Figure 1 and Figure 2 , Figure 1 which is a schematic flowchart of an embodiment of the wake word processing method provided by the present application. Figure 2 which is a schematic flowchart of the solution for obtaining and cloud processing of registration information provided by the present application.
[0023] The wake word processing method of the present application is applied to a wake word processing device. Among them, the wake word processing device of the present application can be a server, a terminal device, or a system in which the server and the terminal device cooperate with each other. Correspondingly, each part included in the wake word processing device, such as each unit, sub-unit, module, and sub-module, can be all set in the server, all set in the terminal device, or respectively set in the server and the terminal device.
[0024] Furthermore, the above-mentioned server can be hardware or software. When the server is hardware, it can be implemented as a distributed server cluster composed of multiple servers or as a single server. When the server is software, it can be implemented as multiple software or software modules, such as software or software modules used to provide a distributed server, or as a single software or software module, which is not specifically limited here.
[0025] It should be noted that the wake word processing device of the present application can be a cloud platform, cloud device, or cloud server on the cloud side, etc.
[0026] As Figure 1 shown, the specific steps are as follows: Step S11: In response to the user registration instruction, input the user registration information into the end-to-end acoustic model to obtain the registration token sequence.
[0027] In the embodiment of the present application, as Figure 2 shown, the user clicks or starts the wake word registration function on the terminal device, such as registering the wake word by voice, and the user inputs the voice registration information on the terminal device; or registering the wake word by text, and the user inputs the text registration information on the terminal device.
[0028] Specifically, when the user selects text registration, the text information to be used as the wake-up word can be directly input in the graphical interface or other input interface of the terminal device.
[0029] When the user chooses voice registration, he needs to speak the target wake-up word to the terminal device or its attached microphone. The voice stream is then sent to the cloud after preliminary recording processing.
[0030] After the cloud side receives the user registration information uploaded by the end-side device, it inputs the user registration information into the end-to-end acoustic model to extract the relevant registration word sequence, that is, the token sequence.
[0031] Specifically, in one implementation, when the user registration information is text registration information, the cloud side processes the text registration information through the following process: The cloud side formats the text registration information, that is, removes unnecessary spaces, punctuation or other special characters that may interfere with system recognition.
[0032] The cloud side converts the registered text into a corresponding token sequence based on the system's built-in modeling units, such as pinyin or other text segmentation methods, and conversion dictionaries (such as a comparison table of pinyin to Chinese characters, or characters to subwords). For example, if the user input "你好小美" is used as the modeling unit, it can be converted into several tokens such as "nǐ"-"hǎo"-"xiǎo"-"měi".
[0033] Finally, the cloud side will perform a simple validity check on the recognized token sequence, such as checking whether it is empty or too short. If it passes the check, it will be temporarily stored in the cloud to prepare for the next step of text encoding.
[0034] In another implementation, when the user registration information is voice registration information, the cloud side processes the voice registration information through the following process: A large-scale noise reduction network is deployed on the cloud side to remove environmental noise, reverberation and other interference in the voice registration information, ensuring higher voice quality input to the acoustic model.
[0035] The cloud side inputs the noise-reduced speech into a large-scale acoustic model. Because the model is not limited by the computing power of the client side, it can use high capacity and non-causal structure, and has better recognition accuracy. Finally, the text token sequence corresponding to the speech is output, just like the consistent modeling unit during text registration, recognizing the speech "你好小美" as "nǐ"-"hǎo"-"xiǎo"-"měi".
[0036] Finally, the cloud side will perform a simple validity check on the recognized token sequence, such as checking whether it is empty or too short in length. If the verification passes, it will be temporarily stored in the cloud to prepare for the next text encoding.
[0037] Step S12: Map the registered token sequence to a registered embedding vector.
[0038] In the embodiment of the present application, for the token sequences obtained through the two approaches shown in the above step S11 (whether directly transformed from text or obtained through cloud speech recognition), the cloud side will call a pre-trained text encoder to generate an embedding vector, so as to map the token sequence to a registered embedding vector.
[0039] Specifically, the cloud side inputs the token sequence into the cloud text encoder, maps each token in the sequence to a high-dimensional vector space, and performs context interaction or other specific feature extraction. Finally, the cloud side fuses this token sequence into a registered embedding vector or a sequence of vectors.
[0040] The cloud side saves the generated registered embedding vector in the cloud temporary storage area for subsequent linkage with operations related to the acoustic model. At the same time, the cloud side will also send this registered embedding vector to the terminal device later for initializing or updating the embedding matching module of the terminal device.
[0041] Step S13: Use the registered token sequence to locate the relevant core network of the end-to-end acoustic model.
[0042] In the embodiment of the present application, after completing text or voice registration, the cloud has obtained the text token information corresponding to the target wake-up word. The cloud side has pre-trained a large-scale end-to-end acoustic model, but this large model is not entirely sent to the terminal side. Instead, the relevant core network is obtained by cropping according to the actually registered wake-up word information and sent to the terminal device.
[0043] Specifically, the cropping strategy of the present application can ensure that the terminal-side model only focuses on the acoustic feature paths closely related to the target wake-up word, greatly reducing the model scale and computing power requirements. For the cropping process, please continue to refer to Figure 3 , Figure 3 is Figure 1 the specific process schematic diagram of step S13 of the wake-up word processing method shown.
[0044] As Figure 3 shown, the specific steps are as follows: Step S131: Obtain the acoustic unit activation paths of each token in the end-to-end acoustic model.
[0045] Step S132: Extract the relevant acoustic unit activation paths based on the registered token sequence, prune the acoustic unit activation paths of the remaining tokens or set their weights to zero, to obtain the relevant core network corresponding to the registered token sequence.
[0046] In the embodiments of the present application, the cloud side locates the relevant sub-network according to the token sequence: when training or deploying the model at the cloud side, the key paths of each token or its pronunciation variant in the acoustic model will be recorded. For example, for pinyin such as "nǐ" - "hǎo", it will correspond to specific acoustic unit activation paths. The cloud side extracts the acoustic unit activation paths corresponding to the token sequence extracted in step S11, and prunes or sets the weights of a large number of other unused pronunciation units to zero.
[0047] Step S133: Reduce the network layer depth and / or parameter dimension of the relevant core network.
[0048] In the embodiments of the present application, after the cloud side locates the core network necessary for the target token, it will further reduce the network layer depth and parameter dimension. For example, reducing redundant convolutional kernels, sparsifying some fully connected layers, etc., to achieve a double reduction in computational complexity and power consumption.
[0049] Step S14: Generate an acoustic encoder model using the relevant core network, where the model computational complexity of the acoustic encoder model is less than that of the end-to-end acoustic model.
[0050] In the embodiments of the present application, the cloud side takes the core network obtained by pruning in step S13 as a relatively complete and independently operable small acoustic encoder module, and packages the weights and structure definitions in the cloud.
[0051] Step S15: Send the registered embedding vector and the acoustic encoder model to the end-side device for wake-word recognition.
[0052] In the embodiments of the present application, the cloud side transmits the registered embedding vector calculated by the text encoder in step S12 and the acoustic encoder model obtained by pruning in step S14 to the end-side device for storage through the network, for subsequent invocation by the similarity determination module of the end-side device.
[0053] In this application, on the cloud side, in response to a user registration instruction, the user registration information is input into the end-to-end acoustic model to obtain a registration token sequence; the registration token sequence is mapped to a registration embedding vector; the relevant core network of the end-to-end acoustic model is located using the registration token sequence; an acoustic encoder model is generated using the relevant core network, where the computational complexity of the acoustic encoder model is less than that of the end-to-end acoustic model; the registration embedding vector and the acoustic encoder model are sent to the device side for wake word recognition. Through the above wake word processing method, the cloud side trims the acoustic model according to the user registration information and then sends it to the device side, improving the wake word detection efficiency and recognition accuracy of the device side.
[0054] Please continue to refer to Figure 4 and Figure 5 , Figure 4 which is a schematic flowchart of another embodiment of the wake word processing method provided by this application, Figure 5 and is a schematic flowchart of real-time input processing of the device-side model provided by this application.
[0055] As Figure 4 shown, the specific steps are as follows: Step S21: The device side uploads the user registration instruction and the user registration information to the cloud side.
[0056] Step S22: The cloud side inputs the user registration information into the end-to-end acoustic model to obtain a registration token sequence.
[0057] Step S23: The cloud side maps the registration token sequence to a registration embedding vector.
[0058] Step S24: The cloud side locates the relevant core network of the end-to-end acoustic model using the registration token sequence.
[0059] Step S25: The cloud side generates an acoustic encoder model using the relevant core network, where the computational complexity of the acoustic encoder model is less than that of the end-to-end acoustic model.
[0060] Step S26: The cloud side sends the registration embedding vector and the acoustic encoder model to the device side.
[0061] In the embodiment of this application, steps S21 to S26 are basically the same as Figure 1 the content of steps S11 to S15 in the wake word processing method shown, and will not be elaborated here.
[0062] Furthermore, after receiving the acoustic encoder model sent by the cloud side, the device side can continue to reduce the network layer depth and / or parameter dimension of the acoustic encoder model according to the device performance, specifically manifested as: reducing redundant convolutional kernels, sparsifying some fully connected layers, etc., to achieve a dual reduction in computational complexity and power consumption.
[0063] Step S27: The edge device encodes the user's real-time input using the acoustic encoder model to obtain a real-time embedding vector, and compares the registered embedding vector with the real-time embedding vector.
[0064] In the embodiment of the present application, after the edge device completes initialization after receiving the cropped acoustic model and the registered embedding vector returned from the cloud side, it will enter the real-time processing state, that is, determine the similarity of the wake-up word for the user's real-time input.
[0065] Specifically, the edge device encodes the user's real-time input according to the acoustic encoder model returned from the cloud side, and compares the obtained real-time embedding vector with the registered embedding vector returned from the cloud side. For the specific real-time input processing process, please continue to refer to Figure 5 Refer to Figure 6 , Figure 6 which Figure 4 is the specific flowchart of step S27 of the wake-up word processing method shown.
[0066] As Figure 6 shown, the specific steps are as follows: Step S271: The edge device obtains a number of audio slices based on the real-time voice input.
[0067] In the embodiment of the present application, when the edge device is in the listening or wake-up detection state, the real-time collected voice stream will first go through basic signal processing steps: usually, with a frame length of 25 ms and a frame shift of 10 ms, slice the input voice stream to obtain a number of audio slices.
[0068] Step S272: The edge device inputs the number of audio slices into the acoustic encoder model to extract the audio feature matrix.
[0069] In the embodiment of the present application, the edge device calculates the Mel-Frequency Cepstral Coefficients (MFCC) for each frame of audio slice, and combines the first-order or second-order difference information to capture the dynamic changes of pronunciation. The calculated audio feature matrix is the input of the acoustic encoder model.
[0070] Step S273: The edge device inputs the audio feature matrix into the acoustic pronunciation classification branch to obtain the posterior probability of the pronunciation of the text token.
[0071] In the embodiment of the present application, as Figure 5As shown in the figure, the edge device inputs the audio feature matrix into the acoustic pronunciation classification branch and the audio embedding branch respectively. Among them, the acoustic pronunciation classification branch outputs the posterior probability of pronunciation after aggregation of the current frame or several frames based on softmax classification or other forms, and then infers the text token index most likely corresponding to this moment. The audio embedding branch outputs a hidden layer representation h_frame(t) at each time step t for subsequent generation of the joint embedding of the audio-text token.
[0072] Step S274: The edge device segments the real-time voice input into several pronunciation paragraphs and the pronunciation duration of each text token according to the maximum value of the posterior probability of pronunciation of each audio slice.
[0073] In the embodiment of the present application, the edge device in the acoustic pronunciation classification branch determines the pronunciation token detected at the current moment based on the maximum value of the posterior probability. For example, if the corresponding token index is "hǎo", if the token classification result is still maintained in the subsequent several frames, it is regarded as the same continuous pronunciation paragraph. The start and end frames of this paragraph can be determined according to the switching moment between this token and the next token, so as to obtain its duration range.
[0074] Step S275: The edge device embeds the real-time voice input according to several pronunciation paragraphs and the pronunciation duration to extract the audio-text embedding vector of the real-time voice input.
[0075] In the embodiment of the present application, for the frame sequence covered by each token (j), the edge device will aggregate the hidden layer representations output by it in the audio embedding branch. The edge device calculates the mean of these vectors to obtain a time aggregation vector of this token (i), which represents the mapping relationship between the audio and this token, that is, in the current speech paragraph, the concentrated speech features corresponding to the token.
[0076] When all tokens (i) are calculated, the edge device obtains a set of joint embeddings {token(0), token(1),...token(n)}, corresponding to each text token and its audio representation in the current speech, so as to more finely represent whether the speech and the text match precisely.
[0077] Step S28: When the vector comparison is successful, the edge device wakes up the corresponding device with the wake-up word corresponding to the user's real-time input.
[0078] In the embodiment of the present application, after the edge device obtains the audio-text token joint vector sequence, it will call the registered embedding vector saved locally for similarity calculation, structure determination and response.
[0079] Specifically, the client device uses a similarity metric (such as cosine similarity or other distance functions) to calculate the degree of proximity between the token embedding of the current input speech and the target registered embedding in the vector space. If the overall matching degree exceeds a certain threshold, it is preliminarily determined that the speech contains the target wake-up word.
[0080] Once the registered target wake-up word is confirmed to appear in the audio, the end-side device will directly trigger the corresponding wake-up or application logic locally. For example, if this wake-up word is specifically used to activate a certain intelligent assistant function, the end-side device will schedule device resources, start the corresponding functional module, and give the user a prompt. In addition, if it is determined that the wake-up condition is not met, the end-side device will continue to monitor the subsequent voice stream in real time and repeat the aforementioned frame processing, acoustic encoding and similarity judgment process.
[0081] The present application provides a multi-mode custom wake-up word system that combines large-scale cloud-based training with lightweight deployment on the end side, significantly improving the accuracy and real-time performance of keyword recognition.
[0082] Innovations include but are not limited to the following aspects: First, on the cloud side, a convolutional and multi-head attention fusion acoustic encoder is used in conjunction with a multi-layer Transformer text encoder to achieve multimodal fusion of text and speech registration methods.
[0083] Second, the cloud side cuts the target keywords and generates corresponding sub-networks and embeddings, which are then sent to the end-side devices, greatly reducing the computing and storage overhead of the end-side devices.
[0084] Third, an evaluation method combining the acoustic pronunciation posterior probability and cosine similarity is adopted to perform precise pronunciation alignment and keyword matching on real-time speech, effectively reducing false triggering and enhancing robustness to noise and accents.
[0085] The overall process can not only make full use of cloud computing resources to complete large-scale data augmentation and training, but also perform rapid detection on the terminal side with low power consumption, which is suitable for personalized wake-up needs in multiple scenarios.
[0086] Specifically, the model training and model structure of the acoustic model used in this application are as follows: Regarding the collection and preprocessing of training data: This application builds a large-scale training data set of multi-modal custom wake-up words in the cloud, covering two modes: voice and text. The voice part contains about 10,000 hours of effective recordings, and the recording sources cover studio scenes and daily environment sampling, involving smartphones, smart speakers, and various microphone arrays. All voices go through a unified preprocessing process: First, apply the noise reduction algorithm based on spectral subtraction to eliminate background noise, set the audio sampling rate to 16,000 times per second, and the quantization precision to 16 bits. Subsequently, eliminate obvious distortion or severely silent segments to ensure the stable proportion of effective audio segments. The text part contains the transcription results aligned frame by frame with the above speech, and also includes plain text from web crawling and internal dialogue corpora, with a scale of approximately 20 million sentences. All texts are uniformly removed of non-standard characters and converted into a specified Token sequence to keep the subsequent training of the text encoder consistent.
[0087] To enhance the adaptability to environmental and accent differences, this application adopts various augmentation strategies for the training data. Specifically, it includes: introducing random time-domain stretching, superimposing noise scripts, and mild reverberation at the speech level; constructing synonymous phrase replacement and partial word order perturbation at the text level to avoid the model overfitting to a specific speaking style or text format. Through this series of operations, a diverse and high-quality training dataset is finally formed, providing a robust foundation for the model to recognize wake words in multiple scenarios and under multiple pronunciation conditions.
[0088] Regarding the network structure and training of the acoustic encoder: Each audio segment is framed in a way with a fixed frame length of 25 milliseconds and a frame shift of 10 milliseconds. Calculate 40-dimensional Mel-frequency cepstral coefficients for each frame, and additionally add the first-order difference and second-order difference to obtain a 120-dimensional acoustic feature vector for each frame. The entire speech is composed of this sequence as the input and fed into the main network of the acoustic encoder.
[0089] Specifically, the acoustic encoder consists of 3 two-dimensional convolutional modules and 2 temporal attention modules. The size of the convolutional kernel in the first layer is fixed at 3 rows and 3 columns, the number of channels is fixed at 32, and the stride is 1; the size of the convolutional kernel in the second layer is also 3 rows and 3 columns, the number of channels is fixed at 64, and the stride is 1; the size of the convolutional kernel in the third layer is still 3 rows and 3 columns, the number of channels is fixed at 128, and the stride is 2, which is used to halve the frame rate in the time domain direction. After each layer of convolution, a max-pooling window size of 2 and a stride of 2 are used to achieve double downsampling of the feature map. The convolution output is batch-normalized and added with an activation function based on the rectified linear unit. This 3-layer convolutional structure extracts low-level and mid-high-level local patterns of the speech signal and significantly reduces the subsequent computational amount.
[0090] After the convolutional block, 2 multi-head self-attention modules are additionally introduced for deeper temporal modeling. Each self-attention layer contains 8 attention heads, the internal tensor dimension is 256, and the output results of each head are concatenated and mapped back to 256 dimensions through a feed-forward network. Through residual connection and layer normalization, the training stability of the network is enhanced. At this point, a frame-level feature sequence with a significantly shortened length but stronger representation ability can be obtained.
[0091] After passing through the aforementioned convolutional and self-attention layers, two branches are separated at the end of the network: Acoustic Pronunciation Classification Branch: A 1-layer linear mapping projects the 256-dimensional hidden vector to 41 dimensions, including 40 tone-marked Chinese phonetic alphabets and 1 blank symbol. This branch uses the Connectionist Temporal Classification loss function to ensure the temporal alignment of pronunciation and text tokens.
[0092] Audio Embedding Branch: A 1-layer fully connected layer reduces the 256-dimensional hidden vector to 128 dimensions, followed by a hyperbolic tangent activation. The output is a lower-dimensional audio embedding containing time series information, which is used for subsequent comparison or similarity calculation with the text embedding in the vector space.
[0093] Regarding the network structure and training of the text encoder: All texts involved (the keyword text input by the user during registration or the Chinese phonetic alphabet sequence recognized by the cloud) need to be converted into a Token sequence consistent with the acoustic modeling unit. In this case, the tone-marked Chinese phonetic alphabets are used as the only Token set, with a fixed number of 40. Each Token is first mapped to a 128-dimensional trainable word vector to represent the morphological information.
[0094] Specifically, the text encoder consists of 3 layers of Transformer encoding blocks. Each layer of the Transformer contains multi-head self-attention and 2 layers of feed-forward networks. The number of heads in the multi-head self-attention is 8, and the internal vector dimension of each head is 256. The output of the multi-head is concatenated and then pressed back to 256 dimensions through a 1-layer feed-forward network. Residual paths and layer normalization are used to prevent gradient vanishing or explosion. After 3 layers of stacking, a deep context representation can be constructed for the Chinese phonetic alphabet sequence, extracting rich semantic and positional information.
[0095] After Transformer encoding, the hidden output corresponding to each Chinese phonetic alphabet Token is 256 dimensions. To be consistent with the 128-dimensional vector generated by the audio embedding branch, the text encoder finally uses a 1-layer linear mapping to reduce 256 dimensions to 128 dimensions. This 128-dimensional vector is used as the text embedding and is subsequently compared with the audio embedding in the audio-text similarity calculation. The training of the text encoder at this stage is completed through joint training with the acoustic model or separate pre-training with text data, ensuring that the final embedding space can reflect the pronunciation and semantic features of the text.
[0096] To enable the system to accurately detect the target keyword in the custom wake-up word scenario, it is necessary to combine the posterior probability of the CTC branch and the vector similarity between audio and text. The implementation steps are as follows: During training, the CTC branch maps the frame sequence to a sequence of tone-marked Chinese phonetic alphabets, and through modeling the blank symbol, it realizes elastic alignment with variable durations. After training, given any audio segment, the temporal sequence of the Chinese phonetic alphabet tokens output by the CTC branch determines the start and end ranges of each pronunciation on the time axis.
[0097] The audio embedding branch generates a 128-dimensional vector in each frame or small window. When the entire audio corresponding to a pinyin token needs to be evaluated, the system performs arithmetic averaging on the embedding vectors of all frames within the time domain of the token to obtain a stable audio feature vector.
[0098] During registration or after speech recognition is completed, a text encoder is used to map each pinyin of the target keyword to a 128-dimensional embedding. This forms a vector sequence that is aligned with the audio end.
[0099] The cosine similarity between the above audio vector and the text vector is calculated. If the similarity is higher than the set threshold and the posterior probability of the CTC branch is sufficient, it is determined that the audio segment contains the corresponding pinyin token, and then whether the keyword appears is determined.
[0100] When performing large-scale joint training in the cloud, this application defines the following loss function to simultaneously optimize the pronunciation alignment capability and the audio-text fusion effect.
[0101] The CTC loss directly measures the degree of fit of the CTC branch to the correct pinyin sequence at the frame level. Given the time sequence length T and the target pinyin sequence length U, the CTC algorithm constructs all possible alignment paths on blanks and repeated paths and calculates the total probability log(p) of the true label path. The larger this value is, the more accurate the alignment between the frame sequence and the true label sequence is. The loss function takes the negative logarithm, and the lower the value, the more perfect the model alignment is.
[0102] make, L_CTC=-log(p(pinyin sequence|acoustic feature)) If each frame output by the CTC branch can correctly correspond to the target pinyin, this loss will be significantly reduced.
[0103] Embedding contrast loss: Audio embeddings are compared with text embeddings in a 128-dimensional space. This application uses a contrastive learning method based on cosine similarity: Positive sample: audio frame sequence and corresponding text token.
[0104] Negative samples: audio frame sequences with irrelevant text tokens.
[0105] In order to bring positive samples closer and negative samples further away, a fixed threshold m is introduced during training. If the cosine similarity of the positive sample is lower than m, it is penalized. If the cosine similarity of the negative sample is higher than m, it is also penalized. The loss function is recorded as L_embed = Σ_max(0, m-cos(audio vector, text vector)) + Σ_max(0, cos(audio vector, text vector)-m) This structure continuously forces the model to learn the discrimination ability between audio and text during joint training, thus supporting more accurate cross-modal matching.
[0106] The final objective function consists of two parts: L = L_CTC + α × L_embed where α is the balance coefficient. At the beginning of training, the proportion of the CTC loss is increased to ensure that the model first learns accurate pronunciation alignment; later, the proportion of the embedding contrast loss is gradually increased to improve the discrimination of the audio-text joint embedding when the system performs keyword matching. Through distributed training on a large-scale dataset and iterating for dozens of rounds, the model gradually converges to an optimal state that takes both accurate alignment and cross-modal discrimination into account.
[0107] Those skilled in the art can understand that in the above method of the specific implementation manner, the writing order of each step does not mean a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.
[0108] To implement the above wake word processing method, the present application also proposes a wake word processing device. For details, please refer to Figure 7 , Figure 7 which is a schematic structural diagram of an embodiment of the wake word processing device provided by the present application.
[0109] The wake word processing device 400 of this embodiment includes a processor 41, a memory 42, an input / output device 43, and a bus 44.
[0110] The processor 41, the memory 42, and the input / output device 43 are respectively connected to the bus 44. Program data is stored in the memory 42, and the processor 41 is configured to execute the program data to implement the wake word processing method described in the above embodiment.
[0111] In the embodiments of the present application, the processor 41 may also be referred to as a CPU (Central Processing Unit). The processor 41 may be an integrated circuit chip with signal processing capabilities. The processor 41 may also be a general-purpose processor, a digital signal processor (DSP, Digital Signal Process), an application-specific integrated circuit (ASIC, Application Specific Integrated Circuit), a field-programmable gate array (FPGA, Field Programmable Gate Array), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor 41 may also be any conventional processor, etc.
[0112] The present application also provides a computer storage medium. Please continue to refer to Figure 8 , Figure 8 FIG. is a schematic structural diagram of an embodiment of the computer storage medium provided by the present application. A computer program 61 is stored in the computer storage medium 600. When the computer program 61 is executed by a processor, it is used to implement the wake-up word processing method of the above embodiment.
[0113] When the embodiments of the present application are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0114] The above are only the embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structural or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.
Claims
1. A wake-up word processing method, characterized in that: The wake-up word processing method is applied to the cloud side, and the wake-up word processing method includes: In response to a user registration instruction, inputting user registration information into an end-to-end acoustic model to obtain a registration word unit sequence; Mapping the registered word-meta sequence into a registered embedding vector; Locating a relevant core network of the end-to-end acoustic model using the registered word sequence; Generating an acoustic encoder model using the relevant core network, wherein a model calculation amount of the acoustic encoder model is less than a model calculation amount of the end-to-end acoustic model; The registered embedding vector and the acoustic encoder model are sent to the end-side device for wake-up word recognition.
2. The wake-up word processing method according to claim 1, characterized in that: The user registration information is voice registration information or text registration information; The step of inputting user registration information into the end-to-end acoustic model to obtain a registration word sequence includes: Inputting the text registration information into the text modeling unit of the end-to-end acoustic model, obtaining the text word units of each text in the text registration information, and combining them into the registration word unit sequence; Alternatively, the speech registration information is input into the end-to-end acoustic model for recognition to obtain the registration word sequence.
3. The wake-up word processing method according to claim 1, characterized in that: The method of using the registered word sequence to locate the relevant core network of the end-to-end acoustic model includes: Obtaining an acoustic unit activation path of each word in the end-to-end acoustic model; Relevant acoustic unit activation paths are extracted based on the registered word-unit sequence, and the acoustic unit activation paths of the remaining words are pruned or the weights are reset to zero to obtain the relevant core network corresponding to the registered word-unit sequence.
4. The wake-up word processing method according to claim 3, characterized in that: After locating the relevant core network of the end-to-end acoustic model by using the registered word-unit sequence, the wake-up word processing method further includes: Reduce the network layer depth and / or parameter dimension of the relevant core network.
5. A wake-up word processing method, characterized in that: The wake-up word processing method is applied to a wake-up word processing system, wherein the wake-up word processing system includes cloud-side and terminal-side devices; the wake-up word processing method includes: The terminal device uploads the user registration instruction and user registration information to the cloud side; The cloud side inputs the user registration information into the end-to-end acoustic model to obtain a registration word sequence; The cloud side maps the registered word-unit sequence into a registered embedding vector; The cloud side locates a relevant core network of the end-to-end acoustic model using the registered word sequence; The cloud side generates an acoustic encoder model using the relevant core network, wherein the model calculation amount of the acoustic encoder model is less than the model calculation amount of the end-to-end acoustic model; The cloud side sends the registration embedding vector and the acoustic encoder model to the end-side device; The end-side device encodes the user's real-time input by using the acoustic encoder model to obtain a real-time embedding vector, and compares the registered embedding vector with the real-time embedding vector; When the vector comparison is successful, the terminal side device wakes up the corresponding device using the wake-up word input by the user in real time.
6. The wake-up word processing method according to claim 5, characterized in that: The real-time user input is real-time voice input; The end-side device encodes the user's real-time input by using the acoustic encoder model to obtain a real-time embedding vector, including: The terminal side device acquires a plurality of audio slices based on the real-time voice input; The client-side device inputs the plurality of audio slices into the acoustic encoder model to extract an audio feature matrix; The terminal device inputs the audio feature matrix into the acoustic pronunciation classification branch to obtain the pronunciation posterior probability of the text word; The terminal device divides the real-time voice input into a plurality of pronunciation segments and the pronunciation duration of each text word according to the maximum value of the posterior probability of pronunciation of each audio slice; The terminal-side device embeds the real-time voice input according to the plurality of pronunciation segments and the pronunciation duration to extract an audio-text embedding vector of the real-time voice input.
7. The wake-up word processing method according to claim 6, characterized in that: The end-side device embeds the real-time voice input according to the plurality of pronunciation paragraphs and the pronunciation duration to extract an audio-text embedding vector of the real-time voice input, including: The terminal device acquires a frame sequence covered by each text word according to the plurality of pronunciation paragraphs and the pronunciation duration; The terminal device performs vector aggregation on the speech frames of the frame sequence for the real-time speech input to obtain concentrated speech features corresponding to each text word; The terminal device associates the text word with the corresponding concentrated speech feature to generate the audio text embedding vector.
8. The wake-up word processing method according to claim 5, characterized in that: After the cloud side sends the registration embedding vector and the acoustic encoder model to the client side device, the wake-up word processing method further includes: The end-side device reduces the network layer depth and / or parameter dimension of the acoustic encoder model according to the device performance.
9. A wake-up word processing device, characterized in that: The wake-up word processing device includes a memory and a processor coupled to the memory; The memory is used to store program data, and the processor is used to execute the program data to implement the wake-up word processing method as described in any one of claims 1 to 8.
10. A computer storage medium, characterized in that: The computer storage medium is used to store program data, and when the program data is executed by a computer, it is used to implement the wake-up word processing method as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Speech recognition method, device and equipment, and storage medium
CN110176230A
Voice wake-up template acquisition method and device, electronic equipment and computer readable storage medium
CN111326146A
Mmethod, device and equipment for determining network model pruning strategy and storage medium
CN112149829A
Speech recognition method and device, server and computer readable storage medium
CN112802461A
Voice emotion recognition method and device, equipment and storage medium
CN113129927A