Command Word Confirmation Method, Device, Equipment and Storage Medium

By training voice recognition models on general and command-specific data and using self-attention mechanisms, the method addresses the challenge of recognizing user-defined commands, enhancing accuracy and adaptability, and improving user experience.

CN119811375BActive Publication Date: 2025-07-15深圳市友杰智新科技有限公司
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510296615.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-07-15
Estimated Expiration
2045-03-13

AI Technical Summary

Technical Problem

The existing voice recognition system lacks support for custom vocabulary, resulting in low accuracy in personalized command recognition and affecting user experience.

Method used

Training the recognition model on a common corpus, optimizing the model in combination with the command word corpus, and capturing the context information of audio embedding and text embedding through a self-attention mechanism, and identifying custom vocabulary using similarity comparison.

Benefits of technology

It significantly improves the recognition accuracy and user experience of custom vocabulary, enhances the flexibility and adaptability of the system, and can efficiently recognize user-defined command words.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119811375B_ABST
    Figure CN119811375B_ABST
Patent Text Reader

Abstract

This application relates to the field of speech recognition technology, and particularly to a command word confirmation method, device, equipment, and storage medium. The method includes: training an identification model on general corpus, then optimizing the identification model based on command word corpus, and fixing the model weights to obtain a fixed identification model; extracting audio embeddings of the input audio based on the fixed identification model, mapping the audio embeddings and the text embeddings corresponding to the input text into the same dimensional space and aligning them; processing the aligned audio embeddings through a self-attention mechanism to capture the context information in the audio and optimize the matching effect between the audio embeddings and the text embeddings; comparing the similarity between the audio embeddings and the text embeddings of the input text of custom vocabulary, and if the similarity exceeds the first preset threshold, identifying the command word as the corresponding custom vocabulary. This application significantly improves the identification ability of custom vocabulary and greatly improves the identification accuracy of personalized instructions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech recognition technology, and particularly to a method, device, equipment and storage medium for confirming command words. Background Art

[0002] In existing speech recognition systems, especially in smart home and smart terminal scenarios, command word recognition relies on a preset fixed vocabulary list, making it difficult to handle users' personalized needs or special instructions in specific application scenarios, such as "open my music playlist" or "set reading mode". Although the CTC (Connectionist Temporal Classification) algorithm is fast and memory-saving when processing sequence tasks, it has insufficient support for non-fixed and custom command words, resulting in frequent problems of misrecognition or non-recognition. In addition, introducing custom vocabulary faces the challenge of ensuring accurate recognition without significantly increasing the consumption of computing resources. Traditional methods require retraining the model or significantly increasing the number of parameters, which is time-consuming, laborious and increases the system complexity.

[0003] Therefore, there is a technical problem that the existing speech recognition system has insufficient support for custom vocabulary, resulting in low accuracy of personalized instruction recognition and affecting the user experience. Summary of the Invention

[0004] The main purpose of this application is to provide a method, device, equipment and storage medium for confirming command words, aiming to solve the technical problem that the existing speech recognition system has insufficient support for custom vocabulary, resulting in low accuracy of personalized instruction recognition and affecting the user experience.

[0005] To achieve the above invention purpose, this application proposes a method for confirming command words, the method includes:

[0006] Train a recognition model on a general corpus, then optimize the recognition model based on the command word corpus, and fix the model weights to obtain a fixed recognition model;

[0007] Extract the audio embedding of the input audio based on the fixed recognition model, map the audio embedding and the text embedding corresponding to the input text into the same dimensional space, and align them;

[0008] Process the aligned audio embedding through a self-attention mechanism to capture the context information in the audio and optimize the matching effect between the audio embedding and the text embedding;

[0009] Compare the similarity between the audio embedding and the text embedding of the input text of the custom vocabulary. If the similarity exceeds the first preset threshold, then recognize the command word as the corresponding custom vocabulary.

[0010] Further, after the step of processing the aligned audio embedding through the self-attention mechanism to capture the context information in the audio and optimize the matching effect between the audio embedding and the text embedding, the following steps are included:

[0011] Calculate the similarity at the frame level between the audio embedding and the text embedding;

[0012] If the similarity of a certain number of frames is lower than the second preset threshold, it is determined as misrecognition, and the current input audio is rejected for recognition.

[0013] Further, after the step of processing the aligned audio embedding through the self-attention mechanism to capture the context information in the audio and optimize the matching effect between the audio embedding and the text embedding, the following steps are also included:

[0014] Classify based on the command word list to generate a confusion recognition list;

[0015] When the recognized command word is in the confusion recognition list, traverse each command word in the confusion recognition list;

[0016] Calculate the similarity between the audio embedding and the text embedding of each command word respectively;

[0017] Select the command word with the highest similarity score as the final recognition result and update it to the corresponding command word.

[0018] Further, the step of training the recognition model on the general corpus, then optimizing the recognition model based on the command word corpus, and fixing the model weights to obtain the fixed recognition model includes:

[0019] Conduct preliminary training on the recognition model on the general corpus;

[0020] Optimize the already preliminarily trained model based on the command word corpus containing preset command words;

[0021] After the optimization is completed, fix the weight parameters of the model to form a fixed recognition model.

[0022] Further, the step of extracting the audio embedding of the input audio based on the fixed recognition model, mapping the audio embedding and the text embedding corresponding to the input text into the same dimensional space, and performing alignment includes:

[0023] Extract features from the input audio based on the fixed recognition model and generate an audio embedding;

[0024] Convert the input text into an embedding vector to generate a text embedding;

[0025] Map the audio embedding to the same dimensional space as the text embedding;

[0026] Using the Viterbi algorithm, align the audio embedding and the text embedding on the same time axis.

[0027] Further, the step of processing the aligned audio embedding through the self-attention mechanism to capture the context information in the audio and optimize the matching effect between the audio embedding and the text embedding includes:

[0028] Apply the self-attention mechanism to the aligned audio embedding to capture the context information in the audio sequence;

[0029] Based on the audio embedding and the text embedding after being processed by the self-attention mechanism, calculate the similarity frame by frame in the same dimensional space;

[0030] Use the contrast loss function to adjust the distance between the audio embedding and the text embedding for optimizing the matching effect between the audio embedding and the text embedding.

[0031] Further, the step of comparing the similarity between the audio embedding and the text embedding of the input text of the custom vocabulary, and if the similarity exceeds the first preset threshold, identifying the command word as the corresponding custom vocabulary includes:

[0032] Compare the similarity between the audio embedding and the text embedding frame by frame;

[0033] Compare the calculated similarity with the first preset threshold;

[0034] If the similarity exceeds the first preset threshold, identify the command word as the corresponding custom vocabulary and output the recognition result.

[0035] The second aspect of this application proposes a command word confirmation device, including:

[0036] A training module for training a recognition model on a general corpus, then optimizing the recognition model based on a command word corpus, and fixing the model weights to obtain a fixed recognition model;

[0037] A mapping module for extracting the audio embedding of the input audio based on the fixed recognition model, mapping the audio embedding and the text embedding corresponding to the input text into the same dimensional space, and aligning them;

[0038] An optimization module for processing the aligned audio embedding through the self-attention mechanism to capture the context information in the audio and optimize the matching effect between the audio embedding and the text embedding;

[0039] A comparison module for comparing the similarity between the audio embedding and the text embedding of the input text of the custom vocabulary, and if the similarity exceeds the first preset threshold, identifying the command word as the corresponding custom vocabulary.

[0040] The third aspect of this application also includes a computer device, comprising a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the steps of any one of the above-mentioned methods are implemented.

[0041] The fourth aspect of this application also includes a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of any one of the above-mentioned methods are implemented.

[0042] Beneficial effects:

[0043] This solution significantly improves the support ability of the command word recognition system for custom vocabulary, thereby greatly improving the recognition accuracy of personalized instructions and the user experience. First, the model is optimized by combining specific command word corpora on the basis of the general corpus and fixing the key weights, ensuring the stability and accuracy of the model when processing command words. Then, by mapping the audio embedding and the text embedding into the same dimensional space and aligning them, the effective fusion of cross-modal information is achieved, enhancing the understanding ability of the system. Introducing the self-attention mechanism further optimizes the audio embedding, enabling it to capture richer context information and solving the misrecognition problem caused by the lack of context support in traditional methods. Finally, by comparing the similarity between the audio embedding and the custom vocabulary text embedding, this solution can efficiently recognize the user-defined command words under low-threshold conditions, greatly improving the flexibility and precision of personalized services, effectively solving the problem of insufficient support for custom vocabulary in the prior art, and providing a more intelligent and personalized interaction experience for users. Description of the Drawings

[0044] Figure 1 It is a schematic flowchart of the command word confirmation method according to an embodiment of this application;

[0045] Figure 2 It is a schematic block diagram of the structure of the command word confirmation device according to an embodiment of this application;

[0046] Figure 3 It is a schematic block diagram of the structure of the computer device according to an embodiment of this application.

[0047] The realization of the purpose of this application, functional features and advantages will be further described with reference to the embodiments and the drawings. Detailed Embodiments

[0048] In order to make the purpose, technical solution and advantages of this application clearer, the following further details this application with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.

[0049] Those skilled in the art can understand that, unless specifically stated otherwise, the singular forms "a", "an", "above-mentioned" and "the" used herein may also include the plural forms. It should be further understood that the term "including" used in the specification of the present invention means the presence of features, integers, steps, operations, elements, modules and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, modules, components and / or their groups. It should be understood that when we say that an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any one of the one or more related listed items and all combinations.

[0050] Those skilled in the art can understand that, unless otherwise defined, all terms used herein (including technical terms and scientific terms) have the same meaning as the general understanding of those of ordinary skill in the field to which the present invention belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted in an idealized or overly formal sense unless specifically defined as here.

[0051] Referring to Figure 1 , an embodiment of the present invention provides a command word confirmation method, including steps S1-S4, specifically:

[0052] S1. Train an identification model on a general corpus, then optimize the identification model based on the command word corpus, and fix the model weights to obtain a fixed identification model;

[0053] S2. Extract the audio embedding of the input audio based on the fixed identification model, map the audio embedding and the text embedding corresponding to the input text into the same dimensional space, and align them;

[0054] S3. Process the aligned audio embedding through a self-attention mechanism to capture the context information in the audio and optimize the matching effect between the audio embedding and the text embedding;

[0055] S4. Compare the similarity between the audio embedding and the text embedding of the input text of the custom vocabulary. If the similarity exceeds a first preset threshold, the command word is recognized as the corresponding custom vocabulary.

[0056] The above-mentioned step S1 aims to construct an efficient and adaptable command word recognition model. The specific implementation process is divided into two main stages: first is the general corpus training stage, and then is the optimization stage based on specific command word corpus. In the general corpus training stage, widely applicable language datasets are selected as the basic training materials. These datasets usually contain daily conversations, news broadcasts, etc. in various language environments, with rich vocabulary and grammar structure variations. By using these general corpora to train the initial recognition model, it can be ensured that the model has good generalization ability and can handle various complex speech input scenarios. In this stage, the CTC (Connectionist Temporal Classification) algorithm is used for decoding. Because of its advantages such as no need for alignment for sequence tasks, fast calculation speed, and low memory occupancy, it is particularly suitable for embedded devices. As the training progresses, the model will continuously adjust its parameters on the validation set of the general corpus until the loss value or WER (Word Error Rate) no longer decreases significantly, indicating that the model has fully learned the patterns and features in the general corpus.

[0057] After completing the general corpus training, it enters the second stage, that is, the optimization based on the command word corpus. The goal of this stage is to enhance the model's recognition ability for command words in a specific domain. For this purpose, in addition to the original general corpus, a specially designed command word corpus needs to be introduced. These corpora cover the command words commonly used in smart home devices, such as "turn on the fan", "turn off the light", etc. Continuing to use the model obtained in the first stage as the starting point, the general corpus and the command word corpus are combined to further fine-tune the model. In this process, special attention is paid to the performance on the command word corpus. By continuously adjusting the model parameters until the loss value or WER on the validation set of the command word corpus no longer decreases. The key to this step is to ensure that the model can retain the original generalization ability while being specifically optimized for the command word recognition task. Once this goal is achieved, these part of weights will be fixed to prevent the impact on these key parameters during the subsequent training process, thus ensuring the accuracy and stability of the model when processing command words.

[0058] In addition, this step not only improves the model's recognition ability for predefined command words, but also lays a solid foundation for subsequent steps, including the generation and alignment of audio and text embeddings, the application of self-attention mechanism, and the final similarity comparison. The whole process emphasizes the step-by-step refinement strategy from general to specific domain, ensuring that the model has both broad adaptability and can provide accurate services in specific application scenarios.

[0059] In the above step S2, the main task is to extract the audio embeddings of the input audio from the fixed recognition model and map these audio embeddings and the corresponding text embeddings to the same dimensional space for alignment. First, after obtaining the fixed recognition model trained and optimized in stage S1, this model can effectively process general voice commands and domain-specific command words. In practical applications, when receiving a new audio input, the fixed recognition model is used to extract the feature representation from the audio data, that is, the audio embeddings. Specifically, this process involves the output of the last layer or several layers of the deep neural network, and the features before these layers often contain rich semantic information. For example, an audio embedding layer can be introduced by extracting the features before the audio linear layer to generate audio embeddings, which can capture the phoneme-level details in the audio signal.

[0060] Next, corresponding text embeddings need to be constructed for the input text. This step first involves converting the text into a digital representation form. For example, the original text is segmented into a sequence of words or characters by a tokenizer, and then a pre-trained word vector model such as Word2Vec, GloVe, or a more complex context-aware model such as lightweight BERT is used to convert each word into its corresponding embedding vector. To enable the audio embeddings and text embeddings to be compared in the same dimensional space, it is necessary to ensure that they have the same dimensional size. Therefore, methods such as a projection matrix can be used to map the text embeddings to the same dimensional space as the audio embeddings. In this process, some technical means such as a multi-layer perceptron (MLP) or a simple linear transformation may be used to ensure that the two types of embeddings can be operated in the same framework.

[0061] After completing the above steps, it enters the alignment stage. The goal of alignment is to ensure that the correspondence between the audio embeddings and text embeddings can reflect their true associations as accurately as possible. Here, the Viterbi algorithm can be used to perform alignment processing on the audio embeddings, so that each audio segment can find the most matching text part. For example, for a command word like "turn on the fan", through the alignment process, we can determine which parts of the audio sequence correspond to "turn", "on", "wind", and "fan". This precise alignment not only helps to improve the accuracy of command word recognition but also enhances the system's ability to understand complex sentence structures. Finally, through the aligned audio embeddings and text embeddings, the system can better understand the user's intention, laying a foundation for further optimizing the matching effect through the self-attention mechanism and achieving efficient recognition of custom vocabulary in subsequent steps. This process plays a bridging role in the entire command word confirmation process, connecting the front-end audio feature extraction and the back-end high-level semantic processing.

[0062] In step S3 above, the main task is to process the audio embeddings after alignment through the self-attention mechanism to capture the context information in the audio and optimize the matching effect between the audio embeddings and the text embeddings. First, after obtaining the aligned audio embeddings processed in step S2, these embeddings have initially reflected the relationship between the input audio and the corresponding text, but have not fully considered the influence of the context. The self-attention mechanism can effectively solve this problem. It allows the model to dynamically consider the information of other time steps when processing the audio features at each time step, thereby enhancing the ability to understand complex contexts.

[0063] In the specific implementation process, first, the aligned audio embeddings are passed as input to the self-attention layer. The core of the self-attention layer is to calculate the query (Q), key (K), and value (V). These three vectors are obtained by applying different linear transformations to the original audio embeddings. To further reduce the computational complexity and the number of parameters, a shared weight approach can be adopted, that is, Q, K, and V are transformed using the same weight matrix. Next, the similarity scores between each pair of time steps are calculated, usually through the dot product or the scaled dot product. The similarity scores are then transformed into a probability distribution through the softmax function, representing the importance of each time step relative to other time steps. Based on this probability distribution, a weighted sum of the value vectors is calculated to generate a new audio embedding representation, which contains the relevant information of the current time step and all other time steps, thus effectively capturing the context.

[0064] This processing method not only enables the audio embeddings to better reflect the long-range dependencies in the speech signal but also enhances the system's ability to understand complex sentence structures. For example, when processing a command word like "turn on the air conditioner", the self-attention mechanism can help distinguish between the two parts "turn on" and "air conditioner" and accurately identify the user's intention based on the context information. In addition, by introducing context information, the self-attention mechanism can effectively alleviate the misrecognition problem caused by the lack of background knowledge. For instance, when the user issues the instruction "turn on the fan", the system will not misidentify it as "turn on the light".

[0065] After self-attention processing, the new audio embeddings are used for more precise matching with text embeddings. In this process, by comparing the feature differences between audio embeddings and text embeddings for each frame, techniques such as Contrastive Loss and Asymmetric Proxy Loss (AsyP Loss) can be used to ensure that the distances between positive sample pairs are as close as possible, while the distances between negative sample pairs are as far apart as possible. This method not only improves the consistency between audio and text embeddings but also significantly enhances the overall performance of the command word recognition system, especially when dealing with challenging custom vocabulary. In this way, step S3 plays a crucial role in the entire command word confirmation process, greatly enhancing the robustness and adaptability of the system.

[0066] In the above step S5, the custom command words issued by the user are recognized by comparing the similarity between the audio embeddings and the text embeddings of the custom vocabulary. The core of this step lies in accurately determining whether the input audio matches the predefined or user-defined command words and identifying it as the corresponding custom vocabulary when certain conditions are met. In the specific implementation process, first, the audio embeddings processed by step S3 need to be prepared. These embeddings have captured rich context information through the self-attention mechanism and can better reflect the semantic details in the speech signal.

[0067] Next, for each custom vocabulary, its corresponding text embeddings are calculated in advance. This step usually uses a trained language model or embedding layer (such as BLSTM, BERT, etc.) to convert the custom vocabulary into a fixed-dimensional vector representation. To ensure that the text embeddings and audio embeddings can be compared in the same feature space, appropriate normalization or dimensional adjustment operations may be required for both. For example, the text embeddings can be mapped to the same dimensional space as the audio embeddings through linear transformation or other mapping methods.

[0068] After completing the above preparations, the similarity comparison stage is entered. For each input audio segment and its corresponding audio embeddings, the system will calculate the similarity with the text embeddings of all custom vocabularies in turn. Commonly used similarity measurement methods include cosine similarity, Euclidean distance, etc. Taking cosine similarity as an example, it measures the similarity between two vectors by calculating the cosine value of the angle between them, and the closer the value is to 1, the higher the similarity. In this process, the system will set a preset threshold, and only when the similarity between a pair of audio embeddings and text embeddings exceeds this threshold will it be considered that the audio segment has successfully matched the corresponding custom vocabulary.

[0069] It should be noted that considering that custom vocabulary may not have undergone specialized intensive training, a relatively low threshold can be set in the initial stage to improve recall, that is, to identify as many eligible custom vocabulary as possible. As the system is used for a longer time and data accumulates, this threshold can be gradually adjusted according to the actual situation to achieve a higher precision. In addition, in real-time application scenarios, a streaming comparison method can also be adopted, that is, similarity calculation is performed immediately whenever a new phoneme or command word is detected, so as to achieve fast response.

[0070] Through this detailed and flexible similarity comparison strategy, step S4 not only solves the problem of insufficient support for custom vocabulary in traditional speech recognition systems, but also significantly improves the accuracy of personalized command recognition and the user experience. Especially in scenarios such as smart homes and smart terminals, users can easily define their own command words, making the interaction more natural and fluent. Finally, through a series of carefully designed steps, from basic model training to advanced feature extraction and matching, this method constructs an efficient and robust command word recognition framework, providing users with an intelligent and personalized speech interaction solution.

[0071] In one embodiment, after the step of processing the aligned audio embedding through the self-attention mechanism to capture the context information in the audio and optimize the matching effect between the audio embedding and the text embedding, it includes:

[0072] S10. Calculate the frame-level similarity between the audio embedding and the text embedding;

[0073] S11. If the similarity of a certain number of frames is lower than the second preset threshold, it is determined as a misrecognition, and the current input audio is rejected for recognition.

[0074] In this embodiment, after completing the processing of the aligned audio embedding through the self-attention mechanism to capture the context information in the audio and optimize the matching effect between the audio embedding and the text embedding, the next key step is to calculate the frame-level similarity between the audio embedding and the text embedding, and decide whether to reject the recognition of the current input audio according to the similarity result. This process first requires converting the audio embedding and text embedding processed by the self-attention mechanism into the same feature space for direct comparison. Usually, these embeddings have been adjusted to the same dimension through mapping or transformation, so that the similarity between them can be directly calculated.

[0075] In specific implementation, first, a suitable similarity measurement method is selected to calculate the similarity between each pair of audio embeddings and text embeddings for each frame. Commonly used measurement methods include cosine similarity, Euclidean distance, etc. Taking cosine similarity as an example, it measures the similarity between two vectors by calculating the cosine value of the angle between them. The closer the value is to 1, the higher the similarity. For each frame of audio embedding and the corresponding text embedding, the system calculates the similarity score between them and records it. This step ensures that even in complex contexts, the subtle differences between audio and text can be accurately captured.

[0076] Subsequently, the system analyzes the similarity scores of all frames to determine whether there are certain frames with similarity scores lower than a preset second threshold. This threshold is preset based on the actual application scenario and dataset characteristics, aiming to filter out those audio segments that may have a risk of misidentification. If it is found that a certain number of frames (such as several consecutive frames or a specific proportion of frames) have similarity scores lower than this threshold, the system will consider that there is a high degree of uncertainty or the possibility of error in the current input audio, and thus reject the recognition of this audio input. This strategy effectively improves the robustness of the system and reduces the misidentification cases caused by background noise, unclear pronunciation, or other interference factors.

[0077] In addition, in practical applications, to further improve accuracy, other technical means such as Contrastive Loss and Asymmetric Proxy Loss (AsyP Loss) can also be combined to optimize the process of frame-level similarity calculation. For example, by comparing the distances between positive sample pairs (i.e., correctly matched audio embeddings and text embeddings) and negative sample pairs (i.e., incorrectly matched audio embeddings and text embeddings), it can be ensured that the distance between positive sample pairs is as small as possible, while the distance between negative sample pairs is as large as possible, thereby enhancing the model's ability to distinguish between correct and incorrect matches.

[0078] Generally speaking, by calculating the frame-level similarity between audio embeddings and text embeddings and setting a reasonable threshold to determine whether to reject the recognition of the current input audio, this method not only improves the accuracy and reliability of the command word recognition system, but also significantly enhances the system's ability to handle complex environments and user diversity. This provides a more intelligent, flexible, and efficient solution for voice interaction in scenarios such as smart homes and smart terminals. Through a series of carefully designed steps, from basic model training to final fine verification, the entire framework ensures high accuracy in speech recognition and good consistency in user experience.

[0079] In one embodiment, after the step of processing the aligned audio embedding through the self-attention mechanism to capture the context information in the audio and optimize the matching effect between the audio embedding and the text embedding, the following steps are further included:

[0080] S20. Classify based on the command word list to generate a list of easily confused identifications;

[0081] S21. When the recognized command word is in the list of easily confused identifications, traverse each command word in the list of easily confused identifications;

[0082] S22. Calculate the similarity between the audio embedding and the text embedding of each command word respectively;

[0083] S23. Select the command word with the highest similarity score as the final recognition result and update it to the corresponding command word.

[0084] In this embodiment, the key steps are to classify based on the command word list to generate a list of easily confused identifications, and further refine the command word recognition process on this basis. First, the system will generate a list of easily confused identifications according to the predefined command word list, especially those command words that are easily confused. For example, in smart home devices, command words such as "turn on the fan" and "turn off the fan", or "previous song" and "next song" are easily misrecognized due to similar pronunciations or close semantics. Therefore, these command words are classified into the list of easily confused identifications for more detailed processing later.

[0085] When the command word recognized by the system belongs to the list of easily confused identifications, a more refined recognition process will be entered. Specifically, for each recognized command word, the system will traverse each command word in the list of easily confused identifications and calculate the similarity between the input audio embedding and the text embedding corresponding to each command word respectively. In this process, the cosine similarity or other appropriate similarity measurement methods are usually used to quantify the matching degree between the audio embedding and the text embedding. For example, for the two command words "turn on the fan" and "turn off the fan", the system will calculate the similarity scores between the current input audio embedding and the text embeddings of these two command words respectively.

[0086] Subsequently, the system selects the command word with the highest similarity score as the final recognition result and updates the recognition to this command word. This method not only improves the accuracy of recognition but also effectively solves the problem of misrecognition caused by fuzzy speech signals or environmental noise. In addition, in this way, the system can better handle complex command words defined by users. Especially in the case of multiple approximate command words, it can still accurately understand the user's intention. This refined processing strategy significantly enhances the robustness and adaptability of the system, making the speech recognition system more intelligent and flexible in application scenarios such as smart homes and smart terminals. The whole process ensures high-accuracy speech recognition services from basic model training to final precise verification through a series of carefully designed steps, thereby improving the user experience.

[0087] In one embodiment, the step of training the recognition model on the general corpus, then optimizing the recognition model based on the command word corpus, and fixing the model weights to obtain a fixed recognition model includes:

[0088] S30. Conduct preliminary training on the recognition model on the general corpus;

[0089] S31. Optimize the model that has been preliminarily trained based on the command word corpus containing preset command words;

[0090] S32. After the optimization is completed, fix the weight parameters of the model to form a fixed recognition model.

[0091] In this embodiment, during the process of constructing the command word recognition model, it is first necessary to conduct preliminary training on the recognition model on the general corpus. The goal of this stage is to establish a basic model with broad adaptability that can handle daily conversations, news broadcasts, etc. in various language environments, ensuring that the model has good generalization ability. Specifically, when implementing, a large-scale general corpus containing rich vocabulary and grammatical structure variations is selected as the training data set, and deep learning frameworks such as TensorFlow or PyTorch are used to implement the model construction and training. The CTC (Connectionist Temporal Classification) algorithm is used for decoding, which is particularly suitable for embedded devices because it does not require alignment for sequence tasks, has a fast calculation speed, and low memory occupancy. Through multiple rounds of iterative training, the model parameters are continuously adjusted until the loss value or WER (Word Error Rate) on the validation set of the general corpus no longer decreases significantly, indicating that the model has fully learned the patterns and features in the general corpus.

[0092] After completing the initial training, enter the optimization phase based on the command word corpus. At this time, in addition to the original general corpus, a specially designed command word corpus needs to be introduced. These corpora cover common command words in smart home devices, such as "turn on the fan", "turn off the light", etc. Continue to use the model obtained in the first stage as the starting point, combine the general corpus with the command word corpus, and further fine-tune the model. During this process, pay special attention to the performance on the command word corpus. By continuously adjusting the model parameters until the loss value or WER on the validation set of the command word corpus no longer decreases. The key to this step is to ensure that the model can be specifically optimized for the command word recognition task while retaining its original generalization ability. For example, for a command word like "turn on the fan", the model needs to be able to accurately recognize the specific content of this command instead of misidentifying it as other similar commands.

[0093] Finally, after the optimization is completed, fix the weight parameters of the model to form a fixed recognition model. This process aims to prevent the impact on these key parameters during subsequent training, thereby ensuring the accuracy and stability of the model when processing command words. The fixed model can be directly applied to the actual scenario to process the user's voice input in real time and execute the corresponding commands. This step-by-step refinement strategy from general to specific domains not only improves the model's recognition ability for predefined command words but also lays a solid foundation for subsequent steps, including the generation and alignment of audio and text embeddings, the application of self-attention mechanisms, and the final similarity comparison. The entire process emphasizes that the model should not only have broad adaptability but also provide precise services in specific application scenarios, thereby providing a more intelligent and personalized interaction experience for users.

[0094] In one embodiment, the step of extracting the audio embedding of the input audio based on the fixed recognition model, mapping the audio embedding and the text embedding corresponding to the input text into the same dimensional space, and performing alignment includes:

[0095] S40. Extract features from the input audio based on the fixed recognition model and generate an audio embedding;

[0096] S41. Convert the input text into an embedding vector to generate a text embedding;

[0097] S42. Map the audio embedding to the same dimensional space as the text embedding;

[0098] S43. Use the Viterbi algorithm to align the audio embedding and the text embedding on the same time axis.

[0099] In this embodiment, in the process of extracting the audio embedding of the input audio based on a fixed recognition model, and mapping the audio embedding and the text embedding corresponding to the input text into the same dimensional space for alignment, it is first necessary to extract features from the input audio and generate the audio embedding. Specifically, a pre-trained fixed recognition model is utilized. This model has good generalization ability and optimization ability for specific command words, and can extract feature representations containing rich semantic information from the input audio. These features usually come from the output of the last layer or several layers of a deep neural network. For example, the audio embedding is generated by extracting the features before the audio linear layer. This process ensures that the audio embedding can capture the phoneme-level details in the speech signal, providing a solid foundation for subsequent steps.

[0100] Next, it is necessary to convert the input text into an embedding vector to generate the text embedding. This step first involves converting the original text into a digital representation form, such as splitting the text into a sequence of words or characters by a tokenizer. Then, a pre-trained word vector model such as Word2Vec, GloVe, or a more complex context-aware model such as lightweight BERT is used to convert each vocabulary into its corresponding embedding vector. To enable the audio embedding and the text embedding to be compared in the same dimensional space, it is necessary to ensure that they have the same dimensional size. Therefore, methods such as a projection matrix can be adopted to map the text embedding into the same dimensional space as the audio embedding. In this process, some technical means such as a multi-layer perceptron (MLP) or a simple linear transformation may be used to ensure that the two types of embeddings can be operated within the same framework.

[0101] After completing the above steps, it enters the crucial mapping and alignment stage. The audio embedding is mapped into the same dimensional space as the text embedding to ensure that the two can be compared within the same feature space. This process is usually achieved through a linear transformation or other mapping methods, so that the audio embedding and the text embedding are not only in the same dimensional space, but also can accurately reflect the semantic relationship between them. Subsequently, the Viterbi algorithm is used to align the audio embedding and the text embedding on the same time axis. The Viterbi algorithm is a dynamic programming algorithm commonly used in sequence alignment tasks, and determines the most likely correspondence between the audio and the text by calculating the optimal path. For example, for a command word such as "turn on the fan", through the alignment process, we can determine which parts of the audio sequence correspond to "turn", "on", "fan", and "wind". This precise alignment not only helps to improve the accuracy of command word recognition, but also enhances the system's ability to understand complex sentence structures.

[0102] Finally, through the aligned audio embeddings and text embeddings, the system can better understand the user's intention, laying a foundation for further optimizing the matching effect through the self-attention mechanism and achieving efficient recognition of custom vocabulary in the subsequent steps. This entire process serves as a bridge in the entire command word confirmation process, connecting the front-end audio feature extraction and the back-end advanced semantic processing, ensuring that the system not only has broad adaptability but also can provide precise services in specific application scenarios. In this way, the robustness of the speech recognition system and the user experience are significantly improved.

[0103] In one embodiment, the step of processing the aligned audio embeddings through the self-attention mechanism to capture the context information in the audio and optimize the matching effect between the audio embeddings and the text embeddings includes:

[0104] S50. Apply the self-attention mechanism to the aligned audio embeddings to capture the context information in the audio sequence;

[0105] S51. Based on the audio embeddings and text embeddings processed by the self-attention mechanism, calculate the frame-by-frame similarity in the same dimensional space;

[0106] S52. Use the contrastive loss function to adjust the distance between the audio embeddings and the text embeddings for optimizing the matching effect between the audio embeddings and the text embeddings.

[0107] In this embodiment, in the process of processing the aligned audio embeddings through the self-attention mechanism to capture the context information in the audio and optimize the matching effect between the audio embeddings and the text embeddings, it is first necessary to apply the self-attention mechanism to the aligned audio embeddings. Specifically, the goal of this stage is to enhance the ability of the audio embeddings to understand the context of each time step in the sequence, so as to better reflect the semantic details in the speech signal. In the implementation process, the aligned audio embeddings are first passed as input to the self-attention layer. The core of the self-attention layer is to calculate the query (Q), key (K), and value (V), and these three vectors are obtained by applying different linear transformations to the original audio embeddings. Next, calculate the similarity scores between each pair of time steps, usually implemented through dot product or scaled dot product. The similarity scores are then transformed into a probability distribution through the softmax function, indicating the importance of each time step relative to other time steps. Based on this probability distribution, a weighted sum of the value vectors is performed to generate a new audio embedding representation, which contains the relevant information of the current time step and all other time steps, thus effectively capturing the context.

[0108] After completing the above steps, enter the stage of calculating the similarity frame by frame between the audio embedding and the text embedding processed by the self-attention mechanism in the same dimensional space. To ensure that the audio embedding and the text embedding can be compared in the same feature space, it is necessary to map both of them to the same dimensional space first. This step can be achieved through linear transformation or other mapping methods, so that the audio embedding and the text embedding are not only in the same dimensional space, but also can accurately reflect the semantic relationship between them. Then, for each frame of audio embedding and the corresponding text embedding, the system will calculate the similarity score between them. Commonly used similarity measurement methods include cosine similarity, Euclidean distance, etc. For example, use cosine similarity to calculate the cosine value of the angle between each pair of audio embedding and text embedding. The closer the value is to 1, the higher the similarity. This frame-by-frame similarity calculation helps to accurately evaluate the matching degree between audio and text and provides a basis for further optimization.

[0109] Finally, based on the frame-by-frame similarity calculation, use the contrastive loss function to adjust the distance between the audio embedding and the text embedding to optimize the matching effect between the audio embedding and the text embedding. The contrastive loss function is a technique widely used in embedding learning, aiming to make the distance between positive sample pairs (that is, correctly matched audio embedding and text embedding) as small as possible, while making the distance between negative sample pairs (that is, incorrectly matched audio embedding and text embedding) as large as possible. In specific implementation, methods such as contrastive loss and asymmetric proxy loss (AsyPLoss) can be adopted. These loss functions update the model parameters through the backpropagation algorithm to make the matching between the audio embedding and the text embedding more accurate. For example, in a smart home device, for a command word like "turn on the fan", the system can more accurately distinguish the correct command word instead of misidentifying it as "turn off the fan". This method not only improves the accuracy of the command word recognition system, but also significantly enhances the robustness and adaptability of the system in complex environments, providing a more intelligent and personalized interaction experience for users. The entire process from basic model training to final fine verification ensures high accuracy of speech recognition and good consistency of user experience.

[0110] In one embodiment, the step of comparing the similarity between the audio embedding and the text embedding of the input text of the custom vocabulary and, if the similarity exceeds the first preset threshold, identifying the command word as the corresponding custom vocabulary includes:

[0111] S60. Compare the similarity between the audio embedding and the text embedding frame by frame;

[0112] S61. Compare the calculated similarity with the first preset threshold;

[0113] S62. If the similarity exceeds the first preset threshold, the command word is recognized as the corresponding custom vocabulary, and the recognition result is output.

[0114] In this embodiment, first, for the custom situation, the text customized by the user is pre-calculated as the embedding sequence of the text. This means that before actual use, the system has pre-processed and stored the text embeddings of all possible custom vocabularies. These text embeddings are converted into vector representations of fixed dimensions for custom vocabularies through a trained language model or embedding layer (such as BLSTM, BERT, etc.).

[0115] Next, the system compares the similarity between the audio embedding corresponding to the voice input and the text embedding frame by frame. This step is achieved by calculating the similarity score between each pair of audio embeddings and the corresponding text embeddings. Common similarity measurement methods include cosine similarity and Euclidean distance, etc. For example, the cosine value of the angle between each pair of audio embeddings and text embeddings is calculated using cosine similarity. The closer the value is to 1, the higher the similarity.

[0116] Then, the calculated similarity is compared with the first preset threshold. This preset threshold is pre-set based on the actual application scenario and dataset characteristics, and is used to determine whether the current input audio matches a certain custom vocabulary. It should be noted that since custom vocabularies may not have undergone specialized reinforcement training, a relatively low threshold can be set in the initial stage to improve the recall rate, that is, to identify as many eligible custom vocabularies as possible. Specifically, when the system recognizes a command word at the phoneme level, the threshold can be set very low at this time to ensure that as many potential matches as possible can be captured.

[0117] Finally, if the similarity exceeds the first preset threshold, the command word is recognized as the corresponding custom vocabulary, and the recognition result is output. Specifically, the system will select the custom vocabulary with the highest similarity score as the final recognition result and use it as the basis for command execution. This method not only solves the problem of insufficient support for custom vocabularies in traditional speech recognition systems, but also significantly improves the accuracy of personalized instruction recognition and the user experience. Especially in scenarios such as smart homes and smart terminals, users can easily define their own command words, making the interaction more natural and smooth.

[0118] Referring to Figure 2 , it is a structural block diagram of a command word confirmation device in an embodiment of the present application. The device includes:

[0119] A training module 100, configured to train an identification model on a general corpus, then optimize the identification model based on a command word corpus, and fix the model weights to obtain a fixed identification model;

[0120] A mapping module 200, configured to extract an audio embedding of the input audio based on a fixed recognition model, map the audio embedding and the text embedding corresponding to the input text into the same dimensional space, and perform alignment;

[0121] An optimization module 300, configured to process the aligned audio embedding through a self-attention mechanism, capture the context information in the audio, and optimize the matching effect between the audio embedding and the text embedding;

[0122] A comparison module 400, configured to compare the similarity between the audio embedding and the text embedding of the input text of a custom vocabulary, and if the similarity exceeds a first preset threshold, identify the command word as the corresponding custom vocabulary.

[0123] In one embodiment, the above device further includes a misrecognition module, including a misrecognition processing unit, configured to:

[0124] Calculate the similarity at the frame level between the audio embedding and the text embedding;

[0125] If there are a certain number of frames with similarity lower than a second preset threshold, it is determined as a misrecognition, and the current input audio is rejected for recognition.

[0126] In one embodiment, the above device further includes a confusion recognition module, including a confusion recognition processing unit, configured to:

[0127] Classify based on a command word list to generate a confusion recognition list;

[0128] When the recognized command word is in the confusion recognition list, traverse each command word in the confusion recognition list;

[0129] Calculate the similarity between the audio embedding and the text embedding of each command word respectively;

[0130] Select the command word with the highest similarity score as the final recognition result, and update the recognized command word accordingly.

[0131] In one embodiment, the above training module 100 includes a stage training unit, configured to:

[0132] Conduct preliminary training on the recognition model on a general corpus;

[0133] Optimize the already preliminarily trained model based on a command word corpus containing preset command words;

[0134] After the optimization is completed, fix the weight parameters of the model to form a fixed recognition model.

[0135] In one embodiment, the above mapping module 200 includes an embedding alignment unit, configured to:

[0136] Extract features from the input audio based on a fixed recognition model and generate an audio embedding;

[0137] Convert the input text into an embedding vector to generate a text embedding;

[0138] Map the audio embedding to the same dimensional space as the text embedding;

[0139] Use the Viterbi algorithm to align the audio embedding and the text embedding on the same time axis.

[0140] In one embodiment, the above optimization module 300 includes a similarity calculation unit for:

[0141] Apply a self-attention mechanism to the aligned audio embedding to capture context information in the audio sequence;

[0142] Based on the audio embedding and the text embedding processed by the self-attention mechanism, perform frame-by-frame similarity calculation in the same dimensional space;

[0143] Use a contrast loss function to adjust the distance between the audio embedding and the text embedding for optimizing the matching effect between the audio embedding and the text embedding.

[0144] In one embodiment, the above comparison module 400 includes a result output unit for:

[0145] Compare the similarity of the audio embedding and the text embedding frame by frame;

[0146] Compare the calculated similarity with a first preset threshold;

[0147] If the similarity exceeds the first preset threshold, recognize the command word as the corresponding custom vocabulary and output the recognition result.

[0148] Refer to Figure 3 , In the embodiments of the present application, a computer device is further provided. The computer device may be a server, and its internal structure may be as shown in Figure 3As shown in the figure. The computer device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the computer design is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store usage data and the like during the process of the command word confirmation method. The network interface of the computer device is used to communicate with an external terminal through a network connection. Further, the above computer device may also be provided with an input device, a display screen, and the like. When the above computer program is executed by the processor, it realizes a command word confirmation method, including the following steps: training an identification model on a general corpus, then optimizing the identification model based on command word corpus, and fixing the model weights to obtain a fixed identification model; extracting the audio embedding of the input audio based on the fixed identification model, mapping the audio embedding and the text embedding corresponding to the input text into the same dimensional space, and aligning them; processing the aligned audio embedding through a self-attention mechanism to capture the context information in the audio and optimize the matching effect between the audio embedding and the text embedding; comparing the similarity between the audio embedding and the text embedding of the input text of the custom vocabulary, and if the similarity exceeds a first preset threshold, identifying the command word as the corresponding custom vocabulary.

[0149] Those skilled in the art can understand that Figure 3 the structure shown in the figure is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied.

[0150] An embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it realizes a command word confirmation method, including the following steps: training an identification model on a general corpus, then optimizing the identification model based on command word corpus, and fixing the model weights to obtain a fixed identification model; extracting the audio embedding of the input audio based on the fixed identification model, mapping the audio embedding and the text embedding corresponding to the input text into the same dimensional space, and aligning them; processing the aligned audio embedding through a self-attention mechanism to capture the context information in the audio and optimize the matching effect between the audio embedding and the text embedding; comparing the similarity between the audio embedding and the text embedding of the input text of the custom vocabulary, and if the similarity exceeds a first preset threshold, identifying the command word as the corresponding custom vocabulary. It can be understood that the computer-readable storage medium in this embodiment may be a volatile readable storage medium or a non-volatile readable storage medium.

[0151] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium provided in this application and used in the embodiments can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0152] It should be noted that, in this article, the term "comprising", "including" or any other variation thereof is intended to cover a non-exclusive inclusion, such that a process, apparatus, article or method comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, apparatus, article or method. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, apparatus, article or method comprising the element.

[0153] The above are only the preferred embodiments of this application, and do not limit the patent scope of this application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of this application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of this application.

Claims

1. A command word confirmation method, characterized in that, The method includes: Training a recognition model on a general corpus, then optimizing the recognition model based on a command word corpus, and fixing the model weights to obtain a fixed recognition model; Extracting an audio embedding of the input audio based on the fixed recognition model, mapping the audio embedding and the text embedding corresponding to the input text into the same dimensional space, and performing alignment; Processing the aligned audio embedding through a self-attention mechanism to capture context information in the audio and optimize the matching effect between the audio embedding and the text embedding; Comparing the similarity between the audio embedding and the text embedding of the input text of the custom vocabulary. If the similarity exceeds a first preset threshold, the command word is recognized as the corresponding custom vocabulary; The step of processing the aligned audio embedding through a self-attention mechanism to capture context information in the audio and optimize the matching effect between the audio embedding and the text embedding includes: Applying a self-attention mechanism to the aligned audio embedding to capture context information in the audio sequence; Performing frame-by-frame similarity calculation in the same dimensional space based on the audio embedding and the text embedding processed by the self-attention mechanism; Using a contrastive loss function to adjust the distance between the audio embedding and the text embedding for optimizing the matching effect between the audio embedding and the text embedding.

2. The command word confirmation method according to claim 1, characterized in that After the step of processing the aligned audio embedding through a self-attention mechanism to capture context information in the audio and optimize the matching effect between the audio embedding and the text embedding, it includes: Performing frame-level similarity calculation on the audio embedding and the text embedding; If the similarity of a certain number of frames is lower than a second preset threshold, it is determined as a misrecognition, and the current input audio is rejected for recognition.

3. The command word confirmation method according to claim 1, wherein After the step of processing the aligned audio embedding through a self-attention mechanism to capture context information in the audio and optimize the matching effect between the audio embedding and the text embedding, it further includes: Classifying based on a command word list to generate a confusion recognition list; When the recognized command word is in the confusion recognition list, traversing each command word in the confusion recognition list; Calculating the similarity between the audio embedding and the text embedding of each command word respectively; Selecting the command word with the highest similarity score as the final recognition result and updating it to the corresponding command word.

4. The command word confirmation method according to claim 1, characterized in that The step of training a recognition model on a general corpus, then optimizing the recognition model based on a command word corpus, and fixing the model weights to obtain a fixed recognition model includes: Performing preliminary training on the recognition model on a general corpus; Optimizing the model that has been preliminarily trained based on a command word corpus containing preset command words; After the optimization is completed, fixing the weight parameters of the model to form a fixed recognition model.

5. The command word confirmation method according to claim 1, characterized in that The step of extracting an audio embedding of the input audio based on the fixed recognition model, mapping the audio embedding and the text embedding corresponding to the input text into the same dimensional space, and performing alignment includes: Extracting features from the input audio based on the fixed recognition model and generating an audio embedding; Converting the input text into an embedding vector to generate a text embedding; Mapping the audio embedding to the same dimensional space as the text embedding; Using the Viterbi algorithm to align the audio embedding and the text embedding on the same time axis.

6. The command word confirmation method according to claim 1, wherein The step of comparing the similarity between the audio embedding and the text embedding of the input text with custom vocabulary, and if the similarity exceeds a first preset threshold, identifying the command word as the corresponding custom vocabulary includes: Comparing the similarity between the audio embedding and the text embedding frame by frame; Comparing the calculated similarity with the first preset threshold; If the similarity exceeds the first preset threshold, identifying the command word as the corresponding custom vocabulary and outputting the recognition result.

7. A command word confirmation device for performing the method according to any one of claims 1-6, characterized in that, Including: A training module for training an identification model on a general corpus, then optimizing the identification model based on the command word corpus, and fixing the model weights to obtain a fixed identification model; A mapping module for extracting the audio embedding of the input audio based on the fixed identification model, mapping the audio embedding and the text embedding corresponding to the input text into the same dimensional space, and aligning them; An optimization module for processing the aligned audio embedding through a self-attention mechanism to capture the context information in the audio and optimize the matching effect between the audio embedding and the text embedding; A comparison module for comparing the similarity between the audio embedding and the text embedding of the input text with custom vocabulary, and if the similarity exceeds the first preset threshold, identifying the command word as the corresponding custom vocabulary.

8. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Voice synthesis method and device based on gated attention mechanism, equipment and medium

    CN119314463A

  • Mixing identification processing method and device, equipment and medium

    CN119600997A

  • Similar time series detection method and apparatus, program and recording medium

    US20040098225A1