Multi-mode user-defined word voice wake-up system

Through the multimodal custom word voice wake-up system, combined with text and voice templates for registration, the problems of high computational overhead and limited recognition accuracy in existing technologies are solved, and high-precision custom word detection is achieved, especially on difficult example datasets.

CN120833799APending Publication Date: 2025-10-24SHANGHAI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410491440.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-04-23
Publication Date
2025-10-24

AI Technical Summary

Technical Problem

Existing wake-up word detection systems are difficult to meet users' personalized customized word needs. They have high computational overhead and limited recognition accuracy. Single-modal methods cannot meet the detection performance requirements of predefined word systems.

Method used

A multimodal custom word voice wake-up system is adopted, which combines text and voice templates for registration. The stability of text as template input and the speaker pronunciation information of voice templates are utilized. Cross-modal matching and identification are performed through feature extraction, pattern extraction and pattern identification modules. A lightweight network is used for feature extraction and training. Joint negative example mining and speech synthesis are introduced to generate confusing samples to enhance robustness.

Benefits of technology

It improves the recognition accuracy of target wake-up words, reduces additional parameter overhead, outperforms single-modal methods, especially on difficult datasets, and achieves high-precision custom word detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120833799A_ABST
    Figure CN120833799A_ABST
Patent Text Reader

Abstract

A multi-mode user-defined word voice wake-up system comprises a feature extraction module, a mode extraction module and a mode identification module, the feature extraction module carries out feature extraction processing according to voice signals and text information to obtain corresponding feature representation results of voice and text, and the mode identification module carries out mode identification according to the feature representation results. The mode extraction module performs cross-modal feature matching processing according to input query voice and registration information to obtain a joint feature result of the query voice and the registration information, and the mode identification module performs multi-scale identification processing according to the input joint feature information to obtain a posterior probability result of the query voice to a target wake-up word. According to the method, the text and the voice template are simultaneously utilized to perform multi-modal registration, and the stability of the text serving as the template input and the conventional pronunciation information of the target speaker contained in the voice template are fully exerted, so that the accuracy of target wake-up word recognition is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, and particularly relates to a multi-modal self-defined word voice wake-up system using text and voice templates for registration. BACKGROUND

[0002] As a core link of intelligent voice interaction, the wake-up word detection system aims to continuously detect the interaction behavior of the user and the intelligent device. The existing wake-up word detection system usually identifies predefined keywords. In order to realize high-accuracy keyword detection, a large amount of predefined word data needs to be collected to train the model, however, this method is difficult to meet the personalized custom word demand of the user. At present, the self-defined wake-up word detection system mainly adopts a sample matching method. This method mainly has two input modes: one is to input the voice sample of the custom word as a template for voice template matching; the other is to input the text of the custom word as a template for detection by cross-modal matching. Although the two systems show very flexible characteristics for new defined words, they are far from meeting the detection performance requirements of the pre-defined word system. SUMMARY

[0003] The present application proposes a multi-modal self-defined word voice wake-up system, which simultaneously uses text and voice templates for multi-modal registration, fully utilizes the stability of text as a template input and the target speaker's regular pronunciation information contained in the voice template, so as to improve the accuracy of target wake-up word recognition.

[0004] The present application is realized by the following technical solutions:

[0005] The present application relates to a multi-modal self-defined word voice wake-up system, comprising a feature extraction module, a mode extraction module and a mode discrimination module, wherein: the feature extraction module performs feature extraction processing according to the voice signal and the text information, to obtain corresponding feature representation results of the voice and the text; the mode extraction module performs cross-modal feature matching processing according to the input query voice and the registration information, to obtain joint feature results of the query voice and the registration information; and the mode discrimination module performs multi-scale discrimination processing according to the input joint feature information, to obtain a posterior probability result of the query voice to the target wake-up word.

[0006] The feature extraction module comprises a query branch feature extraction unit and a registration branch feature extraction unit, wherein the query branch feature extraction unit encodes the input voice signal into frame-level wake-up word features through a lightweight acoustic encoder composed of multiple layers of attention networks; the registration branch feature extraction unit encodes the registered wake-up word text into phoneme sequence features and sub-word sequence features through a pre-trained text encoder, and then encodes the registered wake-up word voice input into frame-level wake-up word features through a pre-trained voice encoder.

[0007] The mode extraction module comprises a text query voice attention unit and a voice query voice attention unit, wherein the text query voice attention unit receives the query voice wake-up word features extracted from the query branch feature extraction unit, and the phoneme sequence features of the registered text and the sub-word sequence features of the registered text extracted from the pre-trained text encoder of the registration branch feature extraction unit, performs cross-modal matching between the text and audio modalities to determine whether the query voice contains the target wake-up, and outputs text-voice joint features; the voice query voice attention unit receives the query voice wake-up word features extracted from the query branch feature extraction unit and the registered voice wake-up word features extracted from the pre-trained voice encoder of the registration branch feature extraction unit, verifies whether the query voice matches the voice content in the template audio, and outputs voice-voice joint features.

[0008] The text query voice attention unit and the voice query voice attention unit are both composed of a single-layer attention module.

[0009] The mode discrimination module comprises a speech-level text fusion unit, a speech-level voice fusion unit, a speech-level discrimination unit, a phoneme sequence discrimination unit and a sub-word sequence discrimination unit, wherein the speech-level text fusion unit extracts one-dimensional speech-level text features from the text-voice joint features, the speech-level voice fusion unit extracts one-dimensional speech-level voice features from the voice-voice joint features, the speech-level discrimination unit performs two-classification according to the one-dimensional speech-level text features and the one-dimensional speech-level voice features, and calculates the binary cross-entropy loss with the true value; the phoneme sequence discrimination unit extracts the phoneme feature sequence corresponding to the original input from the text-voice joint features and performs two-classification, and calculates the binary cross-entropy loss of the phoneme sequence with the true value; the sub-word sequence discrimination unit performs two-classification on the sub-word feature sequence, and calculates the binary cross-entropy loss of the sub-word sequence with the true value.

[0010] The speech-level text fusion unit and the speech-level voice fusion unit are single-layer recurrent neural networks.

[0011] The speech-level discrimination unit, the phoneme sequence discrimination unit and the text sequence discrimination unit are all composed of a single-layer fully connected classification network structure.

[0012] The application relates to a multi-modal self-defined word voice wake-up method based on the above system, wherein in an offline stage, a text and an acoustic encoder are used for offline feature extraction of target text and template voice, a light feature extraction network is constructed and trained, in an online stage, online feature extraction of input voice is performed through the trained light feature extraction network, text-voice attention prediction is performed on input online voice features and offline target text features respectively, text posterior probability is obtained, voice-voice attention prediction is performed on input previous voice features and offline template voice features, voice posterior probability is obtained, and wake-up scores are obtained through fusion of the text posterior probability and the voice posterior probability and classification of a wake-up word classification module.

[0013] The training of the light feature extraction network is performed in an end-to-end mode, real data sets and synthesized wake-up word negative example data sets are used as training data, and a total loss function wherein: is a speech level loss, is a phoneme sequence detection loss, and is a sub-word sequence detection loss.

[0014] The negative example data set is a confusion sample generated through data augmentation, so as to enhance the robustness of the self-defined wake-up word system to easily confused words, and the specific method is as follows: negative example mining is performed on a target wake-up word to obtain wake-up word negative examples, and voice synthesis is performed on the wake-up word negative examples to generate confusion samples.

[0015] The negative example mining includes rule-based negative example mining and negative example mining based on large language model interaction, wherein: the rule-based scheme mainly aims to obtain synonyms, homophones and sorting words of the input target wake-up word. The rule-based synonym method is to calculate the text embedding distance between the input wake-up word and each word in the common word dictionary, and take the 25 words with the closest distance as the easily confused synonyms of the input wake-up word. The rule-based homophone method first needs to split the input target wake-up word into sub-words, then calculate the phoneme edit distance between each sub-word and each character in the common word dictionary, take the nearest 10 characters as the homophones of the sub-word, and then backfill to the original wake-up word, so as to constitute the homophone wake-up word negative example of the wake-up word. The sorting word is relatively simple, and only needs to exhaust all permutations and combinations of the target wake-up word, and remove the self, so as to constitute the sorting word wake-up word negative example of the target wake-up word. The negative example mining based on large language model interaction is a supplement to the wake-up word negative example generation, and the method is to interact with the large language model to generate some words or phrases with similar meanings or similar pronunciations to the target wake-up word, aiming to generate some easily confused negative examples commonly used in life.

[0016] The voice synthesis adopts a multi-lingual voice synthesis model of TSCM-TTS, generates voice as a confusion sample through four hyperparameters of language, target speaker, wake-up word, and cloning times.

[0017] The utterance-level detection loss Used to evaluate the similarity between the query branch and the support branch. If the query audio is the target wake-up word, the label is 1, otherwise 0, and the binary cross-entropy is used to calculate the loss.

[0018] The phoneme-level detection loss Aim to enhance the ability of the model to distinguish similar pronunciations (e.g. "young people" and "youth group"). The phoneme sequence discriminative unit converts the phoneme features into phoneme-level posterior probability. The phoneme true value sequence depends on the alignment information of the voice audio and the text phoneme. If the phoneme sequence of the voice label matches the phoneme sequence of the wake-up word text label, it is 1, otherwise 0.

[0019] The subword-level detection loss Aim to take advantage of the semantic differences between similar words and enhance the discriminative ability of the model. This method is very effective on near-phonetic words but not on near-synonymous words, such as the English words "waiter" (waiter) and "water" (water). The subword sequence discriminative module converts the text features into subword-level posterior probability. The subword true value sequence depends on the alignment information of the voice audio and the text. If the subword sequence of the voice label matches the subword sequence of the wake-up word text label, it is 1, otherwise 0. Technical effects

[0020] Compared with the prior art, the performance of the present application exceeds other single-modal methods, exceeds the method of speech recognition, and performs excellently on difficult example datasets. The present application effectively applies a large model with a large number of parameters to the wake-up word system on the wake-up word task, but does not introduce additional parameter overhead. BRIEF DESCRIPTION OF DRAWINGS

[0021] Figure 1 The present application is a system schematic diagram;

[0022] Figure 2 The present application is a double-branch text encoder schematic diagram;

[0023] Figure 3 The present application is a mode discriminative unit schematic diagram;

[0024] Figure 4 The present application is a data augmentation method schematic diagram;

[0025] Figure 5Schematic diagram of the negative example text mining subroutine in the data augmentation method;

[0026] Figure 6 This is a schematic diagram of the speech synthesis subroutine in the data augmentation method;

[0027] Figure 7 A schematic diagram of a web page illustrating a speech synthesis subroutine in a data augmentation method;

[0028] Figure 8 Schematic diagram of the sliding window reasoning process. DETAILED DESCRIPTION

[0029] like Figure 1 As shown, a multimodal custom word voice wake-up system involved in this embodiment includes: a feature extraction module, a pattern extraction module and a pattern identification module, wherein: the feature extraction module performs feature extraction processing based on the voice signal and text information to obtain the corresponding feature representation results of the voice and text; the pattern extraction module performs cross-modal feature matching processing based on the input query voice and registration information to obtain the joint feature results of the query voice and registration information; the pattern identification module performs multi-scale identification processing based on the input joint feature information to obtain the posterior probability result of the query voice to the target wake-up word.

[0030] like Figure 1 As shown, the feature extraction module includes: a query branch feature extraction unit and a registration branch feature extraction unit, wherein: the query branch feature extraction unit encodes the input speech signal into a frame-level wake-up word feature through a lightweight acoustic encoder composed of a multi-layer attention network; the registration branch feature extraction unit encodes the registered wake-up word text input into phoneme sequence features and subword sequence features respectively through a pre-trained text encoder, and then encodes the registered wake-up word voice input into a frame-level wake-up word feature through a pre-trained voice encoder.

[0031] The query branch is implemented using but not limited to the Conformer architecture described in "Conformer: Convolution-augmented transformer for speech recognition" by Gulati A, Qin J, Chiu CC, et al. ([J].arXiv preprintarXiv:2005.08100, 2020).

[0032] The registration branch adopts but is not limited to a pre-training text encoder DistilBERT (Sanh V, Debut L, Chaumond J, et al. DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter [J]. arXiv preprint arXiv:1910.01108, 2019.) and a pre-training speech encoder XLR-53 (Sharma M. Multi-lingual multi-task speech emotion recognition using wav2vec2.0 [C] / / ICASSP 2022-2022 IEEE international conference on acoustics, speech and signal processing (ICASSP). IEEE, 2022: 6907-6911.) to implement.

[0033] The mode extraction module includes a text query speech attention unit and a speech query speech attention unit, wherein: the text query speech attention unit receives query speech wake-up word features extracted from the query branch feature extraction unit and the phoneme sequence features of the registered text extracted from the pre-training text encoder of the registration branch feature extraction unit and the subword sequence features of the registered text As shown in Figure 2 , cross-modal matching is performed between the text and audio modalities to determine whether the query speech contains the target wake-up and output text-speech joint features; the speech query speech attention unit receives query speech wake-up word features extracted from the query branch feature extraction unit and the registered speech wake-up word features extracted from the pre-training speech encoder of the registration branch feature extraction unit Verify whether the query speech matches the speech content in the template audio and output speech-speech joint features, wherein: denotes the frame length of the query audio, and denote the number of registered text phonemes and registered text subwords, respectively, denotes the frame length of the registered audio, and d denotes the frame dimension.

[0034] The cross-modal matching between text and audio modalities refers to: the text query voice attention unit recognizes the cross-modal correlation between the registered text and the query voice, and the converted features and Connect along the time dimension in: and All by and Transformed, by For example, you need And the corresponding type code e pos and position code e type .Then use self-attention to calculate the text-speech joint features Specifically: in: Indicates the frame length of the query audio. and They represent the number of registered text phonemes and registered text subwords respectively, and d represents the frame dimension.

[0035] The query voice matches the voice content in the template audio. The voice query voice attention unit recognizes the voice modality correlation between the registration voice and the query voice, and the converted features and Connect along the time dimension in: and All by and Transformed, the specific form is by For example, you need And the corresponding type code e pos and position code e type , and then use self-attention to calculate the speech-speech joint feature Specifically: in: Indicates the frame length of the query audio. represents the frame length of the registered audio, and d represents the frame dimension.

[0036] like Figure 3 As shown, the pattern discrimination module includes: the pattern identification module includes: a discourse level text fusion unit, a discourse level speech fusion unit, a discourse level identification unit, a phoneme sequence identification unit and a subword sequence identification unit, wherein: the discourse level text fusion unit extracts the text-speech joint features extracted from the text query speech attention unit into a one-dimensional discourse level text feature The speech-level speech fusion unit extracts the speech-speech joint features from the speech query speech attention unit to one-dimensional speech-level speech features Specifically, The utterance-level discrimination unit receives the one-dimensional utterance-level text features and the one-dimensional utterance-level speech features extracted from the utterance-level text fusion unit and the utterance-level speech fusion unit, performs binary classification, and calculates the binary cross-entropy loss with the true value, specifically: Here, W, b and σ represent trainable weights, biases and Sigmoid functions respectively; the phoneme sequence discrimination unit receives the phoneme feature sequence corresponding to the original input in the text-speech joint features extracted from the text query speech attention unit performs binary classification, and calculates the binary cross-entropy loss of the phoneme sequence with the true value, specifically: where: phon represents the frame index in the range w, b and σ represent trainable weights, biases and Sigmoid functions respectively; the subword sequence discrimination unit receives the subword feature sequence corresponding to the original input in the text-speech joint features extracted from the text query speech attention unit performs binary classification, and calculates the binary cross-entropy loss of the subword sequence with the true value, specifically: The subscript "text" represents the frame index in the range Here, W, b and σ represent trainable weights, biases and Sigmoid functions respectively.

[0037] Through specific actual experiments, on the test platform (CPU: Intel(R) Xeon(R) Gold 6226R, GPU: NVIDIA GeForce RTX 3090), using the deep learning framework Pytorch, in the offline stage, the target text and the template speech are offline feature extraction through the text and acoustic encoder, the lightweight feature extraction network is constructed and trained, in the online stage, the input speech is online feature extraction through the trained lightweight feature extraction network, the text-speech attention prediction is performed on the input online speech features and offline target text features respectively, the text posterior probability is obtained, the speech-speech attention prediction is performed on the input previous speech features and offline template speech features, the speech posterior probability is obtained, the text posterior probability and the speech posterior probability are fused, and the wake-up score is obtained through the wake-up word classification module.

[0038] This example utilizes the large-scale English speech dataset LibriSpeech to construct a custom wake-word speech dataset LibriPhrase. The training set is generated from the train-clean-100 / 360 subsets of the LibriSpeech dataset, while the test dataset is generated from the train-others-500 subset. The LibriPhrase test dataset includes two subsets: LibriPhrase Easy (LE) and LibriPhrase Hard (LH). The duration of the speech segments is between 0.5 to 2 seconds, containing 1 to 4 English words or phrases. The LE test set consists of random negatives, with low phonetic similarity between the enrollment wake-word and the query wake-word. The LH test set consists of hard negatives, with very high phonetic similarity between the enrollment wake-word and the query wake-word, thus prone to errors, and is mainly used to simulate hard-case scenarios. The English custom wake-word detection model is trained on this dataset.

[0039] This example introduces the large-scale Chinese speech dataset WenetPhrase, which is used to evaluate custom wake-word detection in Mandarin. This dataset utilizes approximately 1000 hours of WenetSpeech M / S data, which is segmented by a forced alignment algorithm. On this basis, the Jieba segmentation tool is used to obtain a target word list, and speech segments with a duration of 0.5 to 2 seconds, containing 2 to 6 characters, are selected. This process results in approximately 122K training classes and 54K test classes, with a total of 2.9M samples. Two subsets, WenetPhrase Easy (WE) and WenetPhrase Hard (WH), are constructed. The WE test set consists of random negatives, with low phonetic similarity between the enrollment wake-word and the query wake-word. The WH test set consists of hard negatives, with very high phonetic similarity between the enrollment wake-word and the query wake-word, thus prone to errors, and is mainly used to simulate hard-case scenarios. The Chinese custom wake-word detection model is trained on this dataset.

[0040] This example uses the Speech Commands dataset to evaluate the generalization ability of the custom wake-word algorithm on out-of-domain data in an English environment. This dataset is a common corpus containing 30 wake-words. A multi-class task evaluation is performed. Specifically, 10 target wake-words are randomly selected, while the remaining 20 wake-words are designated as an unknown class.

[0041] Qualcomm Keywords. This dataset contains 50 speakers saying four wake words 4,270 speech samples (Hey Snapdragon, Hey Android, Hey Lumina, and Hey Alexa), which is very suitable for evaluating the performance of custom wake-word detection algorithms and speaker recognition algorithms.

[0042] Hey Snips dataset is a speaker-independent set containing around 11K wake words and 86.5K non-wake-word utterances. To evaluate the performance of Hey Snips one-word detection or wake-word detection, the error rate is measured within a specified duration.

[0043] The present embodiment uses in Figure 4 , Figure 5 , Figure 6 The data augmentation pipeline detailed above is used to augment the existing Libriphrase and proposed WenetPhrase datasets, and then a custom wake-word detection system is trained end-to-end. The query audio is pre-processed to obtain 80-dimensional mel-spectrogram features, and then a lightweight Conformer is used to extract 128-dimensional frame-level audio embeddings. In the enrollment branch, a multi-lingual G2P module is used to extract 128-dimensional phoneme embeddings, a multi-lingual DistilBERT text encoder is used to encode the text into 768-dimensional text embeddings, and an XLR-53 multi-lingual self-supervised speech base model is used to extract 1024-dimensional frame-level template audio embeddings. It is worth noting that the parameters of all the enrollment branches are fixed, and then three lightweight mappers are used to convert all the inputs into a unified 128-dimensional space. The training process uses the Adam optimizer for about 50k steps of training. The specific model parameters are shown in Table 1.

[0044] Table 1

[0045] Custom wake-word detection few-shot fine-tuning details: In order to adapt the proposed custom wake-word detection system to the wake dataset for specific wake-word customization, few-shot fine-tuning is performed. This involves using the data augmentation process in Figure 3 , Figure 4 , Figure 5 to build a large number of challenging negative examples for each target wake word, as well as a small number of real-world wake-word utterances (e.g., 5, 10, and 50), and an Adam optimizer is used to fine-tune for 5k steps with a lower learning rate. The inference process of the personalized wake-word detection system proposed in this embodiment includes two main stages. In the registration stage, the user provides the text and audio templates of the target wake-word, the system infers the phonemes of the text, i.e., the text features, and the audio features of the voice, and stores them in the memory without repeated extraction. In the subsequent test stage, the wake-word system searches for the wake-word. The detection output is a sentence-level probability score P utt , and if the score is greater than a threshold value, it is considered to be a wake-word.

[0046] The experimental data obtained are shown in Tables 2-4:

[0047] Table 2: Word detection accuracy test on LibriPhrase dataset

[0048] Table 3: Word detection accuracy test on WenetPhrase dataset

[0049] Table 3: Multi-wake-word accuracy test

[0050] Table 4: Target wake-word optimization accuracy test

[0051] Compared with the prior art, the present application has high word detection accuracy, achieves the most advanced effect on the LibriPhrase and WenetPhrase public datasets, has good custom word effect, and achieves a recognition accuracy of 95.9% in the closed set and 90.8% in the open set in any multi-wake-word scenario (10 categories). It has good customization ability and customization effect. Using only 5 template audios, it can achieve a low false alarm threshold on the target wake-word. On the HeySnapdragon public dataset, under the condition of 0.05 false alarms per hour, the FRR is only 0.64%, achieving the performance of a production pre-defined word.

[0052] The above specific embodiments can be adjusted in different ways by those skilled in the art without departing from the principles and purposes of the present application. The protection scope of the present application is subject to the claims and is not limited by the above specific embodiments. Each implementation within the scope is subject to the constraints of the present application.

Claims

1. A custom word voice wake-up system based on text and voice prompts, characterized by, The application relates to a wake-up word recognition method and device. The feature extraction module, the pattern extraction module and the pattern identification module, wherein: the feature extraction module performs feature extraction processing according to the voice signal and the text information, and obtains corresponding feature representation results of the voice and the text; the pattern extraction module performs cross-modal feature matching processing according to the input query voice and the registration information, and obtains joint feature results of the query voice and the registration information; and the pattern identification module performs multi-scale identification processing according to the input joint feature information, and obtains posterior probability results of the query voice on the target wake-up word. The feature extraction module comprises a query branch feature extraction unit and a registration branch feature extraction unit, wherein: the query branch feature extraction unit encodes the input voice signal into frame-level wake-up word features through a lightweight acoustic encoder composed of multiple layers of attention networks; and the registration branch feature extraction unit encodes the registered wake-up word text into phoneme sequence features and subword sequence features through a pre-trained text encoder, and then encodes the registered wake-up word voice into frame-level wake-up word features through a pre-trained voice encoder.

2. The text and voice prompt based custom word voice wake-up system of claim 1, wherein, The pattern extraction module comprises a text query voice attention unit and a voice query voice attention unit, wherein: the text query voice attention unit receives the query voice wake-up word features extracted from the query branch feature extraction unit, and the phoneme sequence features and the subword sequence features of the registered text extracted from the pre-trained text encoder of the registration branch feature extraction unit, performs cross-modal matching between the text and the audio modalities to determine whether the query voice contains the target wake-up, and outputs text-voice joint features; and the voice query voice attention unit receives the query voice wake-up word features extracted from the query branch feature extraction unit and the registered voice wake-up word features extracted from the pre-trained voice encoder of the registration branch feature extraction unit, verifies whether the voice content in the query voice and the template audio matches, and outputs voice-voice joint features.

3. The text and voice prompt based custom word voice wake-up system of claim 1, wherein, The pattern identification module comprises a speech-level text fusion unit, a speech-level voice fusion unit, a speech-level identification unit, a phoneme sequence identification unit and a subword sequence identification unit, wherein: the speech-level text fusion unit extracts one-dimensional speech-level text features from the text-voice joint features; the speech-level voice fusion unit extracts one-dimensional speech-level voice features from the voice-voice joint features; the speech-level identification unit performs two-classification according to the one-dimensional speech-level text features and the one-dimensional speech-level voice features respectively, and calculates a binary cross-entropy loss with a true value; the phoneme sequence identification unit extracts phoneme feature sequences corresponding to the original input from the text-voice joint features, performs two-classification, and calculates a binary cross-entropy loss of the phoneme sequence with a true value; and the subword sequence identification unit performs two-classification on the subword feature sequence, and calculates a binary cross-entropy loss of the subword sequence with a true value.

4. The text and voice prompt based custom word voice wake-up system of claim 1, wherein, ​ 5. A multimodal self-defined word voice wake-up method based on the system of any one of claims 1-4, characterized in that, In the offline stage, the target text and the template voice are offline feature extracted by the text and acoustic encoder, a lightweight feature extraction network is constructed and trained, in the online stage, the input voice is online feature extracted by the trained lightweight feature extraction network, the input online voice feature and the offline target text feature are respectively text-voice attention predicted to obtain the text posterior probability, the input prior voice feature and the offline template voice feature are voice-voice attention predicted to obtain the voice posterior probability, the text posterior probability and the voice posterior probability are fused and classified by the wake-up word classification module to obtain the wake-up score.

6. The multimodal self-defining word voice wake-up method of claim 5, wherein, The training lightweight feature extraction network refers to: training the model in an end-to-end manner, using a real data set and a synthesized wake-up word negative example data set as training data, and using a total loss function wherein: is a word-level loss for evaluating the similarity between query branches and support branches, is a phoneme sequence detection loss for enhancing the model to distinguish similar pronunciations, is a subword sequence detection loss for utilizing the semantic difference between similar words and enhancing the distinguishing ability of the model, The negative example dataset refers to: the confusion samples generated by data augmentation to enhance the robustness of the custom wake-up word system to easily confused words, specifically: generating wake-up word negative examples of the target wake-up word through negative example mining, and then performing voice synthesis on the wake-up word negative examples to generate confusion samples.

7. The multimodal self-defining word voice wake-up method of claim 6, wherein, The negative example mining includes rule-based negative example mining and large language model interaction-based negative example mining, wherein: the rule-based synonym method calculates the text embedding distance between the input wake-up word and each word in the commonly used word dictionary, and takes the 25 words with the closest distance as the easily confused synonyms of the input wake-up word. The voice synthesis adopts a multilingual voice synthesis model of TSCM-TTS, and generates voice as confusion samples through four hyperparameters of language, target speaker, wake-up word and cloning times.