Multi-mode-based multi-language self-defined instruction identification method and multi-mode-based multi-language self-defined instruction identification system
Through the multimodal fusion and metric learning framework, we build a supporting input processing unit and a metric discriminator, which solves the problems of insufficient computational efficiency, scalability and robustness of existing voice command recognition technology, and realizes efficient and low-latency multi-language custom command recognition, adapts to complex environments and supports personalized applications.
Patent Information
- Application Number
- CN202510505944.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-09-23
AI Technical Summary
Existing voice command recognition technology has shortcomings in computational efficiency, scalability, and adaptability to complex environments, making it difficult to meet diverse needs. In particular, its robustness is limited in real-time, multilingual, and personalized scenarios.
Adopting a multimodal fusion and metric learning framework, by constructing a support input processing unit, a query input processing unit and a metric discriminator, combining a multimodal encoder, a self-attention mechanism and a cross-attention mechanism, using a spike CTC strategy and a metric learning algorithm, efficient and scalable multi-language custom command recognition is achieved.
It improves the system's computing efficiency and instruction word expansion capabilities, enhances robustness in complex environments, supports multi-language and personalized applications, reduces latency and simplifies the training process, and improves recognition accuracy and adaptability.
Smart Images

Figure CN120690187A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of voice command recognition, and in particular to a multi-modal and multi-language custom command recognition method and system. Background Art
[0002] Keyword spotting (KWS) is a key technology for converting human voice commands into executable computer instructions. Its core task is to detect specific activation words or keywords from a continuous speech stream without fully transcribing the entire speech content. This technology is a crucial component of modern human-computer interaction and is widely used in smart speakers, smartphones, smart homes, and in-vehicle systems. Its basic process involves voice signal acquisition, feature extraction, acoustic model training, and command word recognition. These steps convert user voice input into executable instructions, enabling intelligent operation.
[0003] Traditional voice command recognition relies on the Gaussian Mixture Model-Hidden Markov Model (GMM-HMM) framework. This approach maps speech signals to words and performs recognition by modeling the relationship between acoustic and pronunciation features and language models. However, GMM-HMM suffers from limited accuracy in complex environments (such as noise and multiple accents) and is less adaptable to large-scale data sets. With the advancement of deep learning, technologies based on deep neural networks (DNNs) have gradually become mainstream. DNNs, using architectures such as multilayer perceptrons (MLPs), convolutional neural networks (CNNs), recurrent neural networks (RNNs), and long short-term memory (LSTMs), have significantly improved recognition accuracy and robustness. In recent years, Transformer-based models have further advanced the technology. These models abandon traditional recursive and convolutional structures and instead rely on a self-attention mechanism to capture long-term and short-term dependencies in speech sequences, demonstrating superior performance. Building on this foundation, the Conformer (Convolution-augmented Transformer) architecture combines the local feature extraction capabilities of convolutional neural networks with the global modeling capabilities of transformers to efficiently process contextual information in audio sequences, further improving recognition performance. In addition, the spike CTC (Connectionist Temporal Classification) algorithm was also introduced, which determines the key frame features by calculating the phoneme probability distribution, achieves efficient sequence decoding, reduces noise interference, and improves voice wake-up accuracy.
[0004] The application scenarios of voice command recognition technology are becoming increasingly diverse, requiring systems to possess low latency, high accuracy, and strong robustness. For example, smart home devices must respond quickly to user commands, in-vehicle systems must accurately recognize commands in noisy environments, and multilingual environments must adapt to different accents. However, existing technologies still lack scalability of command words, computational efficiency, and adaptability to complex scenarios. Traditional methods struggle to support the addition of new command words through simple training, and deep learning models require high computing resources and a complex training process. These limitations restrict the technology's application in small devices and personalized scenarios.
[0005] Therefore, a new technical solution is urgently needed to improve the command word expansion capability and environmental adaptability while ensuring real-time performance and accuracy, so as to meet diverse needs and promote the development of voice command recognition technology.
[0006] To achieve the above goals, existing technical solutions mainly rely on deep neural networks and self-attention mechanisms for voice command recognition. For example, patent CN116844534A proposes a voice recognition method that combines convolution and Transformer architectures. It uses the self-attention mechanism to capture long-term and short-term dependencies and improve the recognition accuracy of complex speech patterns. In addition, patent CN115376498A describes a voice recognition method that uses self-supervised learning to extract features from unlabeled audio, optimize context information modeling, and enhance the ability to process diverse speech. Regarding multimodal methods, patent CN118553235A introduces a Transformer network that integrates voice and video data, improving the robustness and accuracy of voice recognition. In addition, patent CN116013309A proposes a voice recognition system based on a lightweight Transformer network, aiming to improve efficiency and accuracy. These solutions further improve the performance of voice command recognition by combining convolutional neural networks (CNNs) with multi-layer perceptrons (MLPs) to extract audio features, or by jointly optimizing acoustic and language models.
[0007] As can be seen from the above, current voice command recognition technology has made progress with the support of deep neural networks and self-attention mechanisms, but it still has significant shortcomings. Existing technologies have at least the following defects:
[0008] On the one hand, existing technologies perform poorly in balancing computational efficiency and recognition performance. Taking patent CN116844534A as an example, it adopts a Convolution-Transformer hybrid architecture to capture long-term and short-term dependencies through a self-attention mechanism to improve the recognition accuracy of complex speech patterns. However, the FLOPs (floating-point operations) of this model grows quadratically with the length of the input sequence, resulting in high latency when processing long voice commands, making it difficult to meet real-time requirements. Patent CN116013309A proposes a speech recognition system based on a lightweight Transformer network, which aims to improve efficiency but may sacrifice recognition accuracy. This performance degradation stems from over-compression of the model capacity, resulting in the loss of high-frequency details of the acoustic features, especially in noisy environments or dialect scenarios, where robustness is significantly reduced. These problems limit the application of the system in scenarios that require fast response.
[0009] On the other hand, the lack of scalability of instruction words has become a major bottleneck of the existing technology. Patent CN118553235A introduces a Transformer network that integrates voice and video data to improve the robustness and accuracy of speech recognition, but its design is mainly aimed at a fixed instruction set. Adding new custom instructions requires retraining the model, which is inefficient and seriously restricts personalized applications. Patent CN202311049621 uses generative adversarial networks (GANs) to enhance accented speech data and improve the model's tolerance for accents, but data enhancement relies on specific accent datasets. When extended to other accents or languages such as Japanese, Korean and other agglutinative languages, the phoneme error rate is high, and data needs to be collected again and the model needs to be trained, which is costly. Although traditional classifier algorithms are lightweight, adding new instructions still requires redesigning the model, which makes it difficult to meet the needs of personalized and multilingual scenarios.
[0010] Furthermore, existing methods lack the ability to decouple features in scenarios with multiple interference coupling, such as noise and accents, resulting in a sharp decline in model generalization performance and limited robustness in complex environments. Patent CN115376498A describes a speech recognition method that uses self-supervised learning to extract features from unlabeled audio, optimize contextual information modeling, and enhance the ability to handle diverse speech. However, in practical applications, uneven data distribution or noise interference can lead to feature extraction bias, affecting the model's generalization ability. While the CTC algorithm can optimize phoneme decoding, it fails to fully utilize multimodal data when combined with traditional classifiers, resulting in insufficient adaptability.
[0011] Therefore, in view of the defects of the existing technology, it is necessary to propose a technical solution to solve the technical problems existing in the existing technology. Summary of the Invention
[0012] In view of this, it is indeed necessary to provide a multi-modal multi-language custom command recognition method and system, based on the multi-modal fusion and metric learning framework, to achieve efficient and scalable multi-language custom command recognition.
[0013] In order to solve the technical problems existing in the prior art, the technical solutions of the present invention are as follows:
[0014] A multi-language custom instruction recognition method based on multimodality includes the following steps:
[0015] Step S1: construct a large multimodal model and train the model, wherein the constructed large multimodal model includes:
[0016] Construct a support input processing unit to extract multimodal features using a multimodal encoder and output a support vector after feature fusion;
[0017] Constructing a query input processing unit for processing the query voice data and outputting a query vector;
[0018] Constructing a metric discriminator for calculating the matching degree between the support vector output by the support input processing unit and the query vector output by the query input processing unit to determine whether to activate the voice command;
[0019] Step S2: Based on step S1, register the user-defined voice command; wherein, the support feature vector of the user-registered command is obtained and stored by the support input processing unit trained in step S1;
[0020] Step S3: Acquire user voice and use the model generated in step S2 for inference and command recognition; wherein, the user voice is processed by the above-mentioned trained query input processing unit to output a query vector, and the matching degree with the pre-stored support feature vector is calculated to recognize the user voice command.
[0021] As a further improvement, the support input processing unit performs the following steps:
[0022] Step S11: extracting features of the registration data through a multimodal encoder, wherein an audio encoder, a phoneme encoder, and a text encoder are used to perform feature extraction on the multimodal dataset;
[0023] Step S12: Using the peak CTC strategy, the features of the frame with the highest probability of the phoneme are obtained to optimize the phoneme alignment.
[0024] Step S13: Use the self-attention mechanism and the cross-attention mechanism to fuse the multimodal features to obtain the support vector.
[0025] As a further improvement, in step S11, different pre-trained large models are used to extract features for the user registration data of three different modalities; wherein the user registered voice data is The text data is Then the phoneme data is obtained through the text data:
[0026] For speech data, Wav2Vec is used for feature encoding;
[0027] For text data, use pre-trained BERT to extract user registration text data features;
[0028] For phoneme data, a G2P encoder is used to extract features of user-registered phoneme data.
[0029] As a further improvement, in step S12, the possibility of the phoneme appearing at each moment is identified, and the frame features with the highest probability value are selected on the time axis, and then these frame features are fused.
[0030] As a further improvement, in step S1, the model is trained and optimized by combining the overall loss, the peak CTC loss, and the metric learning loss. The overall loss function is constructed as follows:
[0031] L=λ1L utt +λ2L CTC +λ3L metric
[0032] Among them, λ1, λ2 and λ3 are balance coefficients, L utt is the sentence-level loss, L metric To measure the learning loss, L CTC is the peak CTC loss.
[0033] As a further improvement, the query input processing unit performs the following steps:
[0034] receiving query voice data;
[0035] and using the audio encoder with the same weight in step S11 to perform feature encoding on the query speech to obtain deep features;
[0036] The encoded feature vector is further filtered for core features using the spike CTC strategy mentioned in step S12;
[0037] Then, feature interaction is performed through the adaptive instance normalization module to generate a query vector with semantic discriminative power.
[0038] As a further improvement, the cosine similarity measurement method is used to construct the metric discriminator, which measures the similarity of two vectors by calculating the angle between them.
[0039] As a further improvement, step S2 includes the following steps:
[0040] Step S21: User command input and analysis: The user submits command information via voice or text through a terminal device. The system pre-processes the input voice signal, segmenting and encoding the text data, and uses the natural language processing module to perform semantic analysis on the command content, extracting keywords, command categories, and language information, and constructing a preliminary command vector representation.
[0041] Step S22: instruction registration and mapping; wherein, based on the analysis results obtained in the S21 stage, the user instructions are mapped to preset semantic categories, and the multimodal features extracted by the supporting input processing unit are used to fuse the voice, text and phoneme data to form a stable instruction feature vector.
[0042] As a further improvement, step S3 includes the following steps:
[0043] Step S31: Real-time capture of user voice signals, and framing and pre-processing of the captured voice signals;
[0044] Step S32: In model reasoning and command recognition, the deep feature vector of the query voice is obtained after Wav2Vec model encoding and peak CTC feature screening, and then interactively fused with the features stored in the support vector library through adaptive instance normalization and cross-attention modules. The matching degree is calculated using cosine similarity, and judged based on the preset dynamic threshold: when the similarity exceeds the threshold, the system determines it as a valid command and triggers the corresponding device control or information query function, otherwise it feedbacks the "command not recognized" message or prompts the user to confirm.
[0045] The present invention also discloses a multi-modal multi-language custom instruction recognition system, comprising:
[0046] A support input processing unit is used to extract multimodal features using a multimodal encoder and output a support vector after performing feature fusion;
[0047] A query input processing unit, configured to process the query voice data and output a query vector;
[0048] The metric discriminator is used to calculate the matching degree between the support vector output by the support input processing unit and the query vector output by the query input processing unit to determine whether to activate the voice command.
[0049] Compared with the existing technology, the technical solution of the present invention can overcome the shortcomings of the existing technology in computing efficiency, scalability and robustness, and improve system performance and practicality. First, in response to the defect of high latency of the existing model Conformer, a metric learning algorithm is used to simplify the recognition process, reduce computational complexity, ensure low latency, high precision, and adapt to the needs of edge devices. Secondly, the instruction word expansion capability is improved. In response to the problem of insufficient flexibility of the existing solution MM-KWS, by combining metric learning with spike CTC decoding, users can dynamically add new instructions through a small number of samples without retraining the model, supporting multi-language and personalized applications. Finally, the robustness and training efficiency are enhanced. In response to the limitation of high recognition difficulty in complex environments, a multimodal architecture is designed to fuse voice and text information, and spike CTC is used to reduce noise and accent interference and improve adaptability. At the same time, the training process is simplified, and users are supported to self-register samples to adapt to individual characteristics, reduce resource requirements, and improve deployment efficiency and experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 This is a flowchart of a multi-language custom instruction recognition method based on multimodality of the present invention.
[0051] Figure 2 This is a flowchart of the support input processing unit in the present invention.
[0052] Figure 3 This is a flowchart of the peak CTC process in the present invention.
[0053] Figure 4 This is a flowchart of the feature fusion module in the present invention.
[0054] Figure 5 This is a block diagram of the discrimination process of the metric discriminator in the present invention.
[0055] Figure 6 This is a specific flow chart of step S2 in the present invention.
[0056] Figure 7 This is a specific flow chart of step S3 in the present invention.
[0057] Figure 8 This is a schematic diagram of the core algorithm framework of a multi-language custom instruction recognition system based on multimodality of the present invention.
[0058] The following specific embodiments will further illustrate the present invention in conjunction with the above-mentioned drawings. DETAILED DESCRIPTION
[0059] The technical solution provided by the present invention will be further described below with reference to the accompanying drawings.
[0060] See also Figure 1, shown is a flowchart of a multi-language custom instruction recognition method based on multimodality provided by the present invention, comprising the following steps:
[0061] Step S1: Construct a large multimodal model and train the model. The core algorithm framework of the constructed large multimodal model is shown in the following figure: Figure 8 As shown, including:
[0062] Construct a support input processing unit to extract multimodal features using a multimodal encoder and output a support vector after feature fusion;
[0063] Constructing a query input processing unit for processing the query voice data and outputting a query vector;
[0064] Constructing a metric discriminator for calculating the matching degree between the support vector output by the support input processing unit and the query vector output by the query input processing unit to determine whether to activate the voice command;
[0065] Step S2: Based on step S1, register the user-defined voice command; wherein, the support feature vector of the user-registered command is obtained and stored by the support input processing unit trained in step S1;
[0066] Step S3: Acquire user voice and use the model generated in step S2 for inference and command recognition; wherein, the user voice is processed by the above-mentioned trained query input processing unit to output a query vector, and the matching degree with the pre-stored support feature vector is calculated to recognize the user voice command.
[0067] In step S1 above, the construction and preliminary training of the underlying multimodal large model are mainly realized, including data preparation, model architecture design, training process management, etc. It mainly includes steps such as corpus data collection and preprocessing, model construction, model training and optimization.
[0068] In step S2 above, the user customizes the voice command. This step enables the user to customize the personalized command through natural language, and the system automatically recognizes, parses and registers the command. It mainly includes step S21 of user command input and parsing, and step S22 of command registration and mapping.
[0069] In step S3 above, model inference and command recognition ensure that the system can accurately identify whether the voice input is a user command and respond accordingly. This mainly includes step S31 user voice streaming input and step S32 model inference and command recognition.
[0070] Furthermore, during the corpus data collection and preprocessing steps, we used open-source speech datasets and text-to-speech technology to collect and generate anticipated data, building a large-scale training database with 2-4 words per category. Furthermore, we preprocessed the dataset, including but not limited to converting the speech sampling rate to 16kHz and enhancing the speech data through time warping. Furthermore, we collected phoneme information from the training database to generate the word metadata files and phoneme labels corresponding to each category for subsequent use.
[0071] Furthermore, in the multimodal large model constructed in step S1, a support branch, that is, a support input processing unit, is constructed to realize feature extraction of different modes of user registration instructions; a query branch, that is, a query input processing unit, is constructed to realize feature extraction of the query voice input by the user during actual use, and perform feature interaction on the support branch; a metric discriminator is constructed to calculate the confidence score of the feature vectors output by the support branch and the query branch using a metric algorithm, and perform discrimination.
[0072] See also Figure 2 , which is a flowchart of the support input processing unit in the present invention, performs the following steps:
[0073] Step S11: extracting features of the registration data through a multimodal encoder, wherein an audio encoder, a phoneme encoder, and a text encoder are used to perform feature extraction on the multimodal dataset;
[0074] Step S12: Using the peak CTC strategy, the features of the frame with the highest probability of the phoneme are obtained to optimize the phoneme alignment.
[0075] Step S13: Use the self-attention mechanism and the cross-attention mechanism to fuse the multimodal features to obtain the support vector.
[0076] In a preferred embodiment, in step S11, different pre-trained large models are used to perform feature extraction for user registration data of three different modalities;
[0077] For speech data, Wav2Vec is used for feature encoding;
[0078] For text data, use pre-trained BERT to extract user registration text data features;
[0079] For phoneme data, a G2P encoder is used to extract features of user-registered phoneme data.
[0080] In a preferred embodiment, in step S12, the possibility of the phoneme appearing at each moment is identified, and the frame features with the highest probability value are selected on the time axis, and then these frame features are fused.
[0081] Specifically, let the voice data registered by the user be The text data is Then the phoneme data is obtained through the text data: Features of the registered data are extracted by constructing a multimodal encoder.
[0082] For user registration data of three different modalities, different pre-trained large models are used for feature extraction. First, for speech data, Wav2Vec is used for feature encoding. Wav2Vec is a deep learning model used for automatic speech recognition (ASR). It performs end-to-end speech-to-text conversion by directly mapping speech signals to text representations. The self-supervised learning method is adopted. First, the potential audio features are learned from the original speech data, and then fine-tuned to adapt to the specific speech recognition task through Fine-tuning. However, in the present invention, no fine-tuning is performed, but the open source pre-trained weights are directly used for feature extraction. When the application language is actually determined, the model weights fine-tuned for different languages can be used. For example, in a Chinese environment, the Wav2Vec model fine-tuned on Chinese data is used for feature encoding. Specifically, the encoded feature vector can be expressed as It is defined as the following formula:
[0083]
[0084] Among them, ω a0 Represents the pre-trained weights of the Wav2Vec model;
[0085] Similarly, for text data, use pre-trained BERT (Bidirectional Encoder Representations from Transformers) to extract user registration text data The BERT model is a pre-trained language model based on the Transformer architecture that extracts features from text information through a bidirectional context approach. In the BERT model, text input is first processed through word segmentation, converted into subword units, and then processed by a multi-layer Transformer encoder. Unlike traditional unidirectional language models, BERT's bidirectionality makes it more comprehensive in understanding context and can simultaneously utilize contextual information for representation. The encoding process can be defined as the following formula:
[0086]
[0087] Among them, ω t0 Represents the pre-trained weights of the BERT model;
[0088] Similarly, for phoneme data Use G2P (Grapheme-to-Phoneme) encoder to extract user registered phoneme data G2P is a grapheme-to-phoneme model based on a recurrent neural network (RNN) that converts written text (letters) into phonemes (pronunciation units). The basic principle is to use a machine learning model to map letters or glyphs in the text to corresponding phonemes based on the language's spelling rules, vocabulary, and contextual information, and output phoneme-level features. G2P phoneme encoders usually rely on neural networks or other deep learning methods to process complex pronunciation rules, especially in the case of polyphones and irregular spellings, to make accurate conversions. The encoding process can be defined as the following formula:
[0089]
[0090] Among them, ω p0 Represents the pre-trained weights of the G2P model;
[0091] Furthermore, in step S12, to address the redundant frame problem of speech sequences, a dynamic feature selection mechanism is proposed based on the optimization strategy of spike CTC (Connectionist Temporal Classification), which optimizes phoneme alignment by extracting the frame features with the highest probability of phonemes on the time axis. Its structural block diagram is shown in the figure. Figure 3 shown.
[0092] Specifically, the system identifies the possibility of the occurrence of phonemes at each moment, selects the frame features with the highest probability value on the time axis, and then fuses these frame features. This method can effectively reduce the interference of background noise and speech misreading on the recognition results, especially in more complex noise environments. Compared with traditional CTC training methods, the spike CTC strategy avoids interference caused by redundant frames or pronunciation instability by accurately capturing the core information of phonemes in audio signals. Traditional CTC training methods often rely on simpler alignment methods, which may cause some phoneme information to be lost or misplaced, thereby affecting recognition results. The spike CTC strategy optimizes the selection of frame features, which not only improves the model's sensitivity to pronunciation details, but also improves its adaptability in changing speech environments.
[0093] The probability formula for each single aligned time step of CTC is:
[0094]
[0095] Where π is a redundant text character sequence or phoneme sequence, and can also be a pinyin sequence in Chinese speech recognition. x represents the input of the speech, and y represents the output of the neural network. Spike CTC selects the frame features corresponding to the probability peak and splices them into the key frame feature matrix F key ∈R K×d , where K is the number of phonemes. Because many consecutive frames can correspond to the same word, or the output is empty, a many-to-one function is defined to merge the repeated characters in the neural network output sequence to obtain a unique output sequence
[0096] Furthermore, in step S13, the self-attention mechanism and the cross-attention mechanism are used to fuse the multimodal features to obtain the support vector v s The feature fusion module consists of two cross-attention modules and one self-attention module. Its structural block diagram is as follows Figure 4 shown.
[0097] The self-attention mechanism allows the model to dynamically adjust its representation of an input sequence based on the relationships between elements within the sequence. Specifically, self-attention calculates the similarity between each element in the sequence and all other elements, generating a weighted representation so that the output of each element depends not only on its own information but also on information from other related elements in the sequence. This mechanism can effectively capture long-range dependencies in sequences and excels in many natural language processing tasks, such as machine translation and text generation. The advantage of self-attention is that it can process all elements in a sequence in parallel while focusing on information at different positions in the input sequence.
[0098] The self-attention mechanism formula is as follows:
[0099]
[0100] Among them, Q, K and V are the matrices of query, key and value, T represents the transpose of the matrix, and d2 is the dimension of the key vector, which is used to scale the dot product to have a more stable gradient when performing the softmax operation. For the audio feature encoding result, the self-attention mechanism is used to further focus on its own potential dependencies, which is defined as:
[0101]
[0102] Cross-Attention is a commonly used mechanism in neural networks, especially in natural language processing and computer vision tasks. It allows the model to focus on relevant information from different sources while processing the input. In cross-attention, one input sequence (text or phoneme features) interacts and pairs with another sequence (such as the query or target phoneme features), and the information is weighted and combined by calculating the attention weight between them. This enables the model to flexibly fuse features from different modalities or different levels, improving task performance. It splits the input tensor into two parts. and Then one of the parts is used as the query set and the other part is used as the key-value set. Its output is a tensor of size n×d2. For each row vector, its attention weight for all row vectors is given. Specifically, the cross attention mechanism calculation formula is as follows:
[0103]
[0104] Where Q = X1W Q ,K=V=X2W K ,T represents the transpose of the matrix, d k The dimension of the key-value collection;
[0105] Therefore, the semantic features and phoneme features of the text are aligned through the cross-attention mechanism to guide the model to focus on cross-modal features, which is defined as:
[0106]
[0107] In summary, the support feature vector v finally output by the support branch is s It can be defined by the following formula:
[0108]
[0109] Furthermore, step S1 constructs a query branch, that is, a query input processing unit, which first receives the query voice data And use the Wav2Vec model with the same weight in step S1211 to encode the query speech to obtain deep features Defined as:
[0110]
[0111] For the encoded feature vector First, the core features are further screened using the peak CTC strategy mentioned in step S1212. Then, the adaptive instance normalization module (AdaIN) is used to perform feature interaction and generate a query vector v with semantic discriminative power. q . The main function of the adaptive instance normalization module is to dynamically adjust the normalization parameters of the audio representation according to the input keywords, so that the model can adapt to different keywords. Compared with traditional static normalization, this design enables the audio feature distribution corresponding to each keyword to be independently regulated, generating different normalization parameters, thereby achieving adaptive processing of different keywords. The ability to share information between different keywords while maintaining specific processing capabilities for each keyword can significantly improve the performance of the model in open vocabulary keyword detection tasks. The adaptive instance normalization module is a normalization layer in the following form:
[0112]
[0113] Among them, z and v are the characteristic directions of the same shape, μ and σ represent the characteristic mean and characteristic standard deviation respectively;
[0114] Therefore, in the present invention, by s and Through the adaptive instance normalization module, feature interaction is performed and the final query feature vector v is obtained q , expressed as:
[0115]
[0116] Furthermore, in step S1, the design of the metric discriminator is constructed based on the query vector v q and support vector v s First, the model calculates the similarity between the query vector and the support vector through forward propagation. These vectors represent the relationship between the input data and the relevant content in the existing knowledge base. To evaluate the similarity between the two, this project uses the cosine similarity metric. Cosine similarity measures the similarity between two vectors by calculating the angle between them. In this model, the query vector and the support vector are regarded as two vectors in a high-dimensional space. The formula for cosine similarity is:
[0117]
[0118] Among them, · represents the vector dot product operation, ||v q || and ||v s || are the norms of the query vector and support vector (i.e., their modulo lengths), respectively. Cosine similarity measures the similarity between the query vector and the support vector by calculating the angle between them. A value closer to 1 indicates a higher similarity.
[0119] A dynamic threshold τ is preset (adaptively adjusted according to the environmental noise). If the calculated similarity is greater than τ, it is determined to be a valid instruction; otherwise, it is rejected.
[0120] The specific discrimination flow chart of the metric discriminator is as follows Figure 5 shown.
[0121] Furthermore, in step S1, during model training and optimization, the present invention adopts a combination of overall loss, peak CTC loss and metric learning loss to construct an overall loss function as follows:
[0122] L=λ1L utt +λ2L CTC +λ3L metric
[0123] Among them, λ1, λ2 and λ3 are balance coefficients, which can be adjusted according to experimental results and application requirements;
[0124] Specifically, L utt is the sentence-level loss, that is, the loss caused by whether the query sample is judged correctly, L metric To measure the learning loss, that is, the cosine similarity of the final support vector, it is defined as:
[0125]
[0126] Among them, y c represents the true label, Represents the probability predicted by the model;
[0127] Furthermore, in the pre-training stage, large-scale open source speech, text and phoneme data are used to build a training database, and the pre-trained models of each modality are fully learned. At the same time, the weights of the pre-trained models are partially frozen to ensure the stability of subsequent migration. In the fine-tuning stage, small-batch online fine-tuning is adopted for user-defined instruction samples, and data enhancement and cross-validation techniques are combined to prevent overfitting. At the same time, a dynamic threshold adjustment mechanism based on environmental noise estimation is introduced, and the cosine similarity judgment criterion is updated in real time, so as to maintain a high recognition accuracy in different noise environments.
[0128] Furthermore, the optimization strategy uses the Adam optimizer, combined with the learning rate decay strategy, to achieve alternating updates of the weights of each module, ensuring the coordinated evolution of the support branch, query branch, and cross-attention module. Finally, through character error rate indicators such as validation set evaluation, the hyperparameters are further optimized to achieve low-latency, high-accuracy deployment requirements.
[0129] Further, see Figure 6In step S21, during user command input and analysis, the user submits command information in voice or text form through the terminal device. The system preprocesses the input voice signal, including noise reduction, normalization, and frame processing, and simultaneously performs word segmentation and encoding on the text data. Then, the natural language processing module performs semantic analysis on the command content, extracts keywords, command categories, and language information, and constructs a preliminary command vector representation.
[0130] Furthermore, in step S22, during instruction registration and mapping, the user instruction is mapped to a preset semantic category based on the parsing result obtained in step S21, and the multimodal features extracted by the supporting branches are used to fuse the speech, text and phoneme data to form a stable instruction feature vector. The process of generating instruction vectors in steps S21 and S22 is as follows: Figure 6 shown.
[0131] Furthermore, in step S31, during the user voice streaming input, the system deploys a lightweight pre-processing module on the edge device to achieve real-time capture of the user voice signal, and performs framing and pre-processing on the captured voice, dividing the long voice stream into multiple short time-series segments, which are gradually input into the model for inference to reduce system response delay;
[0132] Furthermore, in step S32 model reasoning and command recognition, the deep feature vector of the query voice is obtained after Wav2Vec model encoding and peak CTC feature screening, and then interactively fused with the features stored in the support vector library through adaptive instance normalization and cross-attention modules, and the matching degree is calculated using cosine similarity, and judged according to the preset dynamic threshold: when the similarity exceeds the threshold, the system determines it as a valid command and triggers the corresponding device control or information query function, otherwise it feedbacks the "command not recognized" information or prompts the user to confirm to ensure that the risk of misidentification is minimized. The real-time processing flow of steps S31 and S32 is as follows: Figure 7 shown.
[0133] Furthermore, the system allows users to confirm or correct the recognition results. The feedback information is used to adaptively update the model online. It combines dynamic sampling and incremental learning methods to continuously expand the multimodal database, further improve feature mapping and similarity calculation strategies, and thus achieve long-term stable operation in multi-language and multi-accent scenarios.
[0134] In summary, the overall technical solution realizes a closed-loop process from data collection, model building, online fine-tuning to real-time command recognition through the organic connection of the above steps. It fully overcomes the shortcomings of existing technologies in real-time, command scalability and robustness, and has the significant advantages of high precision, low latency and flexible expansion, ensuring broad application prospects and market value in the field of multimodal, multi-language custom voice command recognition.
[0135] See also Figure 8 The present invention also discloses a multi-language custom instruction recognition system based on multi-modality, comprising:
[0136] A support input processing unit is used to extract multimodal features using a multimodal encoder and output a support vector after performing feature fusion;
[0137] A query input processing unit, configured to process the query voice data and output a query vector;
[0138] The metric discriminator is used to calculate the matching degree between the support vector output by the support input processing unit and the query vector output by the query input processing unit to determine whether to activate the voice command.
[0139] The system executes the above method to perform model training and optimization, user-defined voice command registration, and reasoning and command recognition.
[0140] By adopting the above technical solution, the present invention has the following technical effects:
[0141] 1. Through multimodal model design, the model has rich scalability and customization features, and uses a spike CTC strategy and metric learning algorithm to improve model performance. This is the key point and protection point of this invention. Through training in three modalities: audio, text, and phonemes, the model can more comprehensively understand the input data, improve robustness, better cope with the situation where a single modal data is missing or noisy, enhance generalization ability, and better generalize to new and unknown scenarios by learning the correlation between different modal data.
[0142] 2. Use metric learning algorithms to replace traditional classifiers. Metric learning learns the similarity measure between data points, so that similar data points are close to each other in the high-dimensional embedding space, while dissimilar data points are kept farther away. This enables the system to more accurately cluster or classify different categories of data, thereby achieving stronger overall recognition and prediction performance than traditional classifiers.
[0143] 3. By learning phonemes, users can register their own keywords and voice samples for recognition to more accurately identify and adapt to the user's voice characteristics, such as accent, speaking speed, pronunciation habits, etc., thereby providing a more personalized and smooth voice recognition experience, improving system flexibility, and avoiding security risks and privacy leaks.
[0144] 4. Through the spike CTC decoding algorithm, the features of the frame with the highest probability of the phoneme are obtained and fused, thereby reducing the interference of noise and misreading, fully ensuring the scalability of command words, and improving the accuracy of recognizing different accents.
[0145] The above embodiments are only intended to help understand the method and core concept of the present invention. It should be noted that, without departing from the principles of the present invention, a number of improvements and modifications may be made to the present invention by those skilled in the art, and such improvements and modifications also fall within the scope of protection of the claims of the present invention.
[0146] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A multi-language custom instruction recognition method based on multimodality, characterized in that: The following steps are involved: Step S1: construct a large multimodal model and train the model, wherein the constructed large multimodal model includes: Construct a support input processing unit to extract multimodal features using a multimodal encoder and output a support vector after feature fusion; Constructing a query input processing unit for processing the query voice data and outputting a query vector; Constructing a metric discriminator for calculating the matching degree between the support vector output by the support input processing unit and the query vector output by the query input processing unit to determine whether to activate the voice command; Step S2: Based on step S1, register the user-defined voice command; wherein, the support feature vector of the user-registered command is obtained and stored by the support input processing unit trained in step S1; Step S3: Acquire user voice and use the model generated in step S2 for inference and command recognition; wherein, the user voice is processed by the above-mentioned trained query input processing unit to output a query vector, and the matching degree with the pre-stored support feature vector is calculated to recognize the user voice command.
2. The multi-language custom instruction recognition method based on multimodality according to claim 1, characterized in that: The support input processing unit performs the following steps: Step S11: extracting features of the registration data through a multimodal encoder, wherein an audio encoder, a phoneme encoder, and a text encoder are used to perform feature extraction on the multimodal dataset; Step S12: Using the peak CTC strategy, the features of the frame with the highest probability of the phoneme are obtained to optimize the phoneme alignment. Step S13: Use the self-attention mechanism and the cross-attention mechanism to fuse the multimodal features to obtain the support vector.
3. The multi-language custom instruction recognition method based on multimodality according to claim 2, characterized in that: In step S11, different pre-trained large models are used to extract features for user registration data of three different modalities; wherein the user registered voice data is The text data is Then the phoneme data is obtained through text data: For speech data, Wav2Vec is used for feature encoding; For text data, use pre-trained BERT to extract user registration text data features; For phoneme data, a G2P encoder is used to extract features of user-registered phoneme data.
4. The multi-language custom instruction recognition method based on multimodality according to claim 2, characterized in that: In step S12, the possibility of the phoneme appearing at each moment is identified, and the frame features with the highest probability value are selected on the time axis, and then these frame features are fused.
5. The multi-language custom instruction recognition method based on multimodality according to claim 1, characterized in that: In step S1, the model is trained and optimized by combining the overall loss, the peak CTC loss and the metric learning loss, and the overall loss function is constructed as follows: L=λ1L utt +λ2L CTC +λ3L metric Among them, λ1, λ2 and λ3 are balance coefficients, L utt is the sentence-level loss, L metric To measure the learning loss, L CTC is the peak CTC loss.
6. The multi-language custom instruction recognition method based on multimodality according to claim 2, characterized in that: The query input processing unit performs the following steps: receiving query voice data; and using the audio encoder with the same weight in step S11 to perform feature encoding on the query speech to obtain deep features; The encoded feature vector is further filtered for core features using the spike CTC strategy mentioned in step S12; Then, feature interaction is performed through the adaptive instance normalization module to generate a query vector with semantic discriminative power.
7. The multi-language custom instruction recognition method based on multimodality according to claim 6, characterized in that: The cosine similarity metric is used to construct the metric discriminator, which measures the similarity between two vectors by calculating the angle between them.
8. The multi-language custom instruction recognition method based on multimodality according to claim 1, characterized in that: Step S2 includes the following steps: Step S21: User command input and analysis: The user submits command information via voice or text through a terminal device. The system pre-processes the input voice signal, segmenting and encoding the text data, and uses the natural language processing module to perform semantic analysis on the command content, extracting keywords, command categories, and language information, and constructing a preliminary command vector representation. Step S22: instruction registration and mapping; wherein, based on the analysis results obtained in the S21 stage, the user instructions are mapped to preset semantic categories, and the multimodal features extracted by the supporting input processing unit are used to fuse the voice, text and phoneme data to form a stable instruction feature vector.
9. The multi-language custom instruction recognition method based on multimodality according to claim 1, characterized in that: Step S3 includes the following steps: Step S31: Real-time capture of user voice signals, and framing and pre-processing of the captured voice signals; Step S32: In model reasoning and command recognition, the deep feature vector of the query speech is obtained after Wav2Vec model encoding and peak CTC feature screening, and then interactively fused with the features stored in the support vector library through adaptive instance normalization and cross-attention modules. The matching degree is calculated using cosine similarity, and the judgment is made based on the preset dynamic threshold: when the similarity exceeds the threshold, the system determines it as a valid command and triggers the corresponding device control or information query function. Otherwise, it feedbacks the "command not recognized" message or prompts the user to confirm.
10. The multi-modal, multi-language custom instruction recognition system according to any one of the methods in claims 1 to 9, characterized in that: include: A support input processing unit is used to extract multimodal features using a multimodal encoder and output a support vector after performing feature fusion; A query input processing unit, configured to process the query voice data and output a query vector; The metric discriminator is used to calculate the matching degree between the support vector output by the support input processing unit and the query vector output by the query input processing unit to determine whether to activate the voice command.
Citation Information
Patent Citations
Speech recognition method and device, model training method and device, medium and electronic equipment
CN115376498A
Speech recognition system and method based on lightweight Transform network
CN116013309A
Speech recognition method and device
CN116844534A
Technology for enhancing accent speech recognition based on generative adversarial network data
CN116863923A
Speech recognition method and system for multi-mode intelligent terminal
CN118553235A
Cited By
Custom voice instruction recognition method and device based on twin network, and electronic equipment
CN121662043A
Voice large model adaptation method and device for low-resource language
CN122116885A
A low-resource language-oriented speech large model adaptation method and device
CN122116885B