Speech recognition method of intelligent glasses, intelligent glasses and storage medium
By setting up multi-microphone collection and multi-scale feature extraction weighted enhancement speech recognition methods on smart glasses, the limitations of speech acquisition with fixed wearing position of smart glasses are solved, and the accuracy and adaptability of speech recognition are improved.
Patent Information
- Application Number
- CN202510347905.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-29
AI Technical Summary
The position of the microphone is fixed when worn, resulting in great limitations in voice collection and difficulty in flexibly adjusting, especially when it is necessary to collect voices from the user, it is difficult to improve the accuracy of voice recognition in the prior art.
Smart glasses set up microphones at different locations to collect voice signals in real time, and perform multi-scale feature extraction through pre-trained artificial intelligence models, introduce weighted enhancement of attention mechanisms, and perform nonlinear activation and sequence modeling after deep fusion, which is finally converted into speech recognition results.
It enhances the comprehensive collection ability of smart glasses to collect speech in a fixed wearing position, improves the accuracy of speech recognition, and adapts to unique usage scenarios such as speech translation and command recognition.
Smart Images

Figure CN120388562A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of speech recognition, and particularly to a speech recognition method, a control device, a smart glasses and a computer-readable storage medium for a smart glasses. Background Art
[0002] Different from traditional portable voice devices (such as mobile phones, smart speakers, etc.), the usage position of smart glasses is relatively fixed and is usually worn on the head. This fixed wearing method makes the smart glasses have great limitations in sound collection. When a user uses the smart glasses for voice interaction, the position and orientation of its microphone are difficult to be flexibly adjusted according to the position and direction of the speaker.
[0003] Moreover, in actual usage scenarios, in addition to collecting the user's own voice, the smart glasses may also need to collect the voice of the person opposite the user. For example, in scenarios such as social communication and business negotiation, it is necessary to translate the voice of the other party in real time. However, since the smart glasses are worn on the user's head, the microphone is far from the mouth of the person opposite, which makes the voice of the person opposite be greatly attenuated and interfered during the transmission process.
[0004] This requires that the smart glasses need to have more accurate speech recognition ability than general voice devices. Therefore, how to optimize the speech recognition algorithm to adapt to the unique usage scenario and sound collection limitation of the smart glasses has become an urgent problem to be solved in the current smart glasses field.
[0005] The above content is only used to assist in understanding the technical solution of the present application, and does not represent an admission that the above content is prior art. Summary of the Invention
[0006] The main purpose of the present application is to provide a speech recognition method, a control device, a smart glasses and a computer-readable storage medium for a smart glasses, aiming to improve the accuracy of speech recognition of the smart glasses.
[0007] To achieve the above purpose, the present application provides a speech recognition method for a smart glasses, including the following steps: During the operation of the smart glasses, collect voice signals in real time; wherein, the smart glasses are provided with a plurality of microphones at different positions for collecting voice signals; Based on the input layer of the pre-trained artificial intelligence model, perform multi-scale feature extraction on the voice signals, and introduce an attention mechanism during the feature extraction process to learn the importance weights of different features, so as to weight and enhance the key features; Based on the middle layer of the artificial intelligence model, deeply fuse the weighted and enhanced multi-scale features, and perform non-linear activation and sequence modeling after the feature fusion; Based on the output layer of the artificial intelligence model, convert the feature vector output by the middle layer into a speech recognition result; Based on the current execution task, the smart glasses perform corresponding operations according to the speech recognition result; wherein, if the execution task is speech translation, the speech recognition result is converted into the corresponding translation text and output for display.
[0008] To achieve the above object, the present application further provides a control device, including: A collection module, configured to collect voice signals in real time during the operation of the smart glasses; wherein, the smart glasses are provided with a plurality of microphones at different positions for collecting voice signals; An extraction module, configured to perform multi-scale feature extraction on the voice signals based on the input layer of the pre-trained artificial intelligence model, and introduce an attention mechanism during the feature extraction process to learn the importance weights of different features, so as to weight and enhance the key features; A processing module, configured to deeply fuse the weighted and enhanced multi-scale features based on the middle layer of the artificial intelligence model, and perform non-linear activation and sequence modeling after the feature fusion; A conversion module, configured to convert the feature vector output by the middle layer into a speech recognition result based on the output layer of the artificial intelligence model; An execution module, configured to perform corresponding operations according to the speech recognition result based on the current execution task of the smart glasses; wherein, if the execution task is speech translation, the speech recognition result is converted into the corresponding translation text and output for display.
[0009] To achieve the above object, the present application further provides a smart glasses, the smart glasses include: a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the computer program is executed by the processor, the steps of the speech recognition method of the smart glasses as described above are implemented.
[0010] To achieve the above object, the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the speech recognition method of the smart glasses as described above are implemented.
[0011] The voice recognition method, control device, smart glasses and computer-readable storage medium provided by this application. The smart glasses are provided with microphones at multiple different positions to collect voice signals in real time, which can to a certain extent make up for the problems of fixed wearing positions and large limitations in sound collection, and collect the voices of users and the people opposite them more comprehensively and accurately. And when extracting features, an attention mechanism is introduced to weight and enhance key features, which can improve the ability to capture effective voice features. Finally, the feature vector is converted into an accurate recognition result, enabling the smart glasses to perform operations according to tasks (such as voice translation), which overall enhances the accuracy of voice recognition and adapts to the unique usage scenarios of smart glasses. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 It is a schematic diagram of the steps of the voice recognition method for smart glasses in an embodiment of this application; Figure 2 It is a schematic diagram of the control device in an embodiment of this application; Figure 3 It is a schematic diagram of the internal architecture of the smart glasses in an embodiment of this application.
[0013] The implementation, functional features and advantages of the purpose of this application will be further described in conjunction with the embodiments with reference to the drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0014] The embodiments of this application will be described in detail below. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals denote the same or similar elements or elements with the same or similar functions from beginning to end. The embodiments described below with reference to the drawings are exemplary and are intended to explain this application and should not be construed as a limitation of this application. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of this application.
[0015] In addition, if the descriptions in this application involve "first", "second", etc., they are only for descriptive purposes (such as for distinguishing the same or similar features), and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of these features. In addition, the technical solutions between the various embodiments can be combined with each other, but it must be based on the fact that those of ordinary skill in the art can implement them. When the combination of technical solutions appears to be contradictory or unable to be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by this application.
[0016] Referring to Figure 1 , in an embodiment, the voice recognition method for smart glasses includes: Step S10: During the operation of the smart glasses, collect voice signals in real time. Among them, the smart glasses are equipped with multiple microphones at different positions for collecting voice signals. Step S20: Based on the input layer of the pre-trained artificial intelligence model, perform multi-scale feature extraction on the voice signals, and introduce an attention mechanism during the feature extraction process to learn the importance weights of different features, so as to enhance the key features by weighting. Step S30: Based on the middle layer of the artificial intelligence model, deeply fuse the multi-scale features after weighted enhancement, and perform non-linear activation and sequence modeling after feature fusion. Step S40: Based on the output layer of the artificial intelligence model, convert the feature vector output by the middle layer into a speech recognition result. Step S50: The smart glasses perform corresponding operations according to the speech recognition result based on the current execution task. Among them, if the execution task is speech translation, the speech recognition result is converted into the corresponding translation text and output for display.
[0017] In this embodiment, the execution terminal of the embodiment can be the smart glasses, or other devices or apparatuses (such as a control device) that control the smart glasses.
[0018] As described in step S10, after the smart glasses have completed operations such as power-on and system initialization, each component and functional module are in a normal working state, ready to receive and process external voice signals at any time. Whether the user actively wakes up the smart glasses to activate a specific function, or the smart glasses are in the default listening state (such as the real-time translation mode), as long as it is running, it will continuously execute the voice signal collection task to ensure that no possible voice commands or voice information are missed.
[0019] Multiple microphones are set on the smart glasses at different positions. Microphones at different positions can collect sounds from all directions to achieve omnidirectional voice signal collection. For example, when the smart glasses are worn, the microphones in the front and on the sides can work simultaneously to ensure that no matter from which direction the sound comes, it can be effectively captured, avoiding dead corners in sound collection.
[0020] There are differences between the sound signals collected by multiple microphones. By analyzing and processing these differences, the interference of environmental noise can be effectively suppressed. At the same time, it is also possible to locate the sound source based on information such as the time difference and intensity difference of the sound received by different microphones, determine the direction of the sound source, which helps to process the voice signals more accurately later.
[0021] As described in step S20, the pre-trained artificial intelligence model is trained based on a large amount of speech data. This data covers speech in different accents, speaking speeds, and language scenarios, enabling the model to learn the general features and patterns of speech signals. Optional pre-trained models for speech processing include deep learning architectures such as convolutional neural networks, recurrent neural networks and their variants (such as LSTM, GRU), Transformer architectures, etc.
[0022] The input layer is the part of the model that receives the original speech signal. It is responsible for performing preliminary processing on the collected speech signal to make it suitable for subsequent feature extraction operations. The input layer will perform operations such as normalizing and dimension adjustment on the speech signal to ensure the consistency and effectiveness of the input data.
[0023] Speech signals contain rich information, and features at different scales can reflect different aspects of speech. Multi-scale feature extraction aims to analyze speech signals from multiple perspectives and granularities to capture more comprehensive feature information. For example, in the time domain, short-term features of speech (such as phoneme-level features) and long-term features (such as sentence-level prosodic features) can be extracted; in the frequency domain, the energy distribution and spectral features of different frequency bands can be analyzed.
[0024] Optionally, by performing sliding convolution on the speech signal with different-sized convolutional kernels, local features at different scales can be extracted. Small convolutional kernels can capture the detailed features of speech, such as the subtle changes in phonemes; large convolutional kernels can extract more macroscopic features, such as the rhythm and prosody of speech.
[0025] Optionally, the speech signal can be processed at different resolutions. For example, the speech signal can be downsampled or upsampled, and then feature extraction is performed separately. This can obtain features at different resolutions and enrich the diversity of features.
[0026] In speech signals, the importance of different features for speech recognition is different. The introduction of the attention mechanism is to enable the model to automatically learn the importance weights of different features, so as to pay more attention to key features and improve the accuracy of speech recognition.
[0027] First, the model calculates the attention score for each feature based on the extracted features. This score represents the importance of the feature in the current speech recognition task. The calculation of the attention score can be based on the correlation between features or other metrics.
[0028] Optionally, the attention scores are normalized through a non-linear function to obtain the importance weights of each feature. The sum of these weights is 1, ensuring that they can be reasonably distributed to each feature.
[0029] Multiply the calculated weights with the corresponding features to enhance the key features by weighting. In this way, in subsequent processing, the model will pay more attention to those key features with larger weights and relatively ignore those secondary features with smaller weights.
[0030] During the feature extraction process, the attention mechanism will evaluate and weight the extracted features in real time. For example, after each convolution operation or feature calculation, the attention scores will be immediately calculated and weighted, so that the feature extraction and enhancement processes are integrated with each other, continuously highlighting the key features.
[0031] The attention mechanism is dynamic and it will adjust the importance weights of features in real time according to different input speech signals. In different speech segments or different language environments, the key features may change, and the attention mechanism can adaptively capture these changes to ensure that the model always focuses on the most important features.
[0032] After the multi-scale feature extraction and the processing of the attention mechanism, the obtained are the multi-scale features enhanced by weighting. These features contain richer and more critical information of the speech signal, providing a better input for subsequent feature fusion and speech recognition. They will be passed to the middle layer of the artificial intelligence model for further processing.
[0033] As described in step S30, the multi-scale features describe the speech signal from different angles and granularities, but these features are scattered. The purpose of deep fusion is to integrate these scattered features into a more representative and comprehensive feature representation to capture more complex patterns and information of the speech signal and improve the model's understanding ability of speech.
[0034] Optionally, directly concatenate the features of different scales together on the feature dimension. This method is simple and direct and can retain the original information of each scale feature.
[0035] Or, assign different weights to the features of different scales, and then multiply their corresponding elements and sum them. The determination of the weights can be obtained through learning. For example, use a fully connected layer to automatically learn the importance weights of each scale feature. This method can fuse according to the importance of the features and highlight the role of key features.
[0036] Or, adopt a multi-level fusion strategy. First, perform preliminary fusion on some features of similar scales, and then further fuse the results of the preliminary fusion. For example, first fuse the short-term features and long-term features in the time domain respectively, and then fuse these two fusion results with the frequency domain features to gradually construct a more complex feature representation.
[0037] The features obtained through feature fusion are still the result of linear combination, and the expressive power of the model is limited. The introduction of non-linear activation functions can add non-linear factors to the model, enabling the model to learn more complex speech patterns and feature relationships, and enhancing the generalization ability of the model and its fitting ability for speech signals.
[0038] Optional activation functions include: ReLU (Rectified Linear Unit), Leaky ReLU, and Sigmoid function.
[0039] Speech signals are a type of time-series signal with obvious sequential features. For example, syllables and words in speech are arranged in sequence, and there is a context-dependent relationship between consecutive speech segments. The purpose of sequence modeling is to capture this temporal information and context relationship, so as to more accurately understand the semantics of speech.
[0040] Optionally, a recurrent neural network or a long short-term memory network can be used for sequence modeling.
[0041] Among them, a recurrent neural network is a neural network specifically designed to process sequence data. It captures information in the sequence through the cyclic transfer of hidden states. When processing a sequence of speech features, the hidden state of the recurrent neural network is updated based on the current input and the hidden state at the previous moment, thus retaining the historical information of the sequence.
[0042] Among them, the long short-term memory network is a variant of the recurrent neural network. By introducing a gating mechanism (input gate, forget gate, and output gate), it solves the problem of vanishing gradients. The forget gate controls how much information from the previous hidden state needs to be forgotten, the input gate controls how much information from the current input needs to be added to the cell state, and the output gate controls how much information from the cell state needs to be output as the hidden state at the current moment. The long short-term memory network can better capture the dependencies in long sequences.
[0043] After deep fusion, non-linear activation, and sequence modeling, the intermediate layer will output a feature vector. This feature vector integrates multi-scale information, non-linear features, and temporal context information of the speech signal, providing a rich and effective data basis for the subsequent output layer to convert it into a speech recognition result.
[0044] As described in step S40, the output layer is the connection bridge between the artificial intelligence model and the final speech recognition result. It receives the feature vector transmitted from the intermediate layer, which contains rich information after the speech signal has undergone multi-scale feature extraction, deep fusion, non-linear activation, and sequence modeling. The main function of the output layer is to further process and map this feature vector, converting it into a text form understandable by humans, that is, the speech recognition result.
[0045] When converting the feature vector into a speech recognition result, a language model can be combined to evaluate the reasonableness and likelihood of the word sequence in natural language. For example, in English, a sentence like "The cat is on the mat" is reasonable, while some random combinations of words do not conform to language habits. By combining the language model, possible word sequences can be screened and sorted during the decoding process, improving the accuracy and naturalness of the recognition result.
[0046] Optionally, for a model adopting the Transformer architecture, the output layer decodes based on attention, dynamically focusing on different parts of the input feature vector during the decoding process, thereby better capturing the correspondence between the speech signal and the text.
[0047] Optionally, for a model adopting a non-Transformer architecture, the output layer uses beam search decoding. Beam search is an improvement over greedy search. It retains multiple possible candidate paths (beam width) at each time step instead of only selecting one optimal path. In subsequent time steps, these candidate paths are continued to be expanded, and these paths are evaluated and sorted according to certain criteria (such as language model score, acoustic model score, etc.), and finally the path with the highest score is selected as the recognition result. Beam search balances computational efficiency and recognition accuracy to a certain extent.
[0048] Optionally, there may be some errors in the recognition result, such as typos, missing words, misidentifications, etc. Error correction and modification operations are performed in the post-processing stage. A spell checker can be used to correct spelling mistakes, and a grammar checker can be used to check and correct grammar errors. In addition, more intelligent error correction can be performed by combining context information and domain knowledge. For example, in speech recognition in the medical field, if some words that do not conform to medical terms appear in the recognition result, they can be corrected according to the domain dictionary.
[0049] Optionally, according to the requirements of the actual application, the recognition result is formatted. For example, after converting speech to text, punctuation marks, capital letters, etc. may need to be added to improve the readability of the text. The recognition result can also be typeset and organized according to specific output format requirements.
[0050] After the decoding process and post-processing steps, the finally obtained is a speech recognition result that conforms to natural language expression and application requirements. This result can be displayed in text form on the display screen of the smart glasses, or converted into speech through speech synthesis technology and fed back to the user, or transmitted to other devices or systems for further processing and application, such as being input as an instruction into a smart device to perform corresponding operations.
[0051] As described in step S50, the smart glasses first need to clarify the current execution task.
[0052] Optionally, before using the smart glasses, the user can set up the tasks to be executed in advance, such as selecting modes like "voice translation" and "voice command control". The smart glasses will store these preset information locally and, when receiving the speech recognition result, perform subsequent operations according to the preset tasks.
[0053] Optionally, the smart glasses can determine the current task based on keywords or specific expressions in the speech recognition result. For example, if the recognition result contains words such as "translation" or "translate", the smart glasses will automatically determine the current task as voice translation.
[0054] Once the execution task is determined, the smart glasses will match the speech recognition result with the corresponding operation rules according to the pre-written program logic. For example, if the task is to control smart home appliances and the speech recognition result is "turn on the air conditioner", the smart glasses will send an opening instruction to the air conditioner device through the built-in communication module (such as Bluetooth, Wi-Fi, etc.).
[0055] Optionally, if the current execution task is voice translation, the smart glasses will process it according to the following steps: The smart glasses usually integrate multiple translation engines, and the system will select a suitable translation engine according to factors such as user settings or network conditions, and pass the speech recognition result as input to the selected translation engine. The translation engine will analyze and process the input text according to its own algorithms and language models, and convert it from the source language to the target language. For example, if the speech recognition result is in Chinese and the target language is set to English, the translation engine will translate the Chinese text into the corresponding English sentence.
[0056] Optionally, the translated text is displayed on the display screen of the smart glasses. To improve readability, the format of the displayed text may be adjusted, such as setting appropriate font sizes, colors, and layout methods. At the same time, the source language text and the translated text can also be displayed on the display screen simultaneously for the user to view and compare conveniently.
[0057] Optionally, in addition to text display, the smart glasses can also convert the translated text into speech through text-to-speech technology for playback. The user can choose whether to turn on the speech playback function according to their own needs. For example, when it is inconvenient to view the screen, speech playback can provide a more convenient interaction method.
[0058] In one embodiment, the smart glasses are provided with microphones at multiple different positions to collect voice signals in real time, which can, to a certain extent, make up for the problems of fixed wearing positions and large limitations in sound collection, and collect the voices of the user and the person opposite more comprehensively and accurately. Moreover, an attention mechanism is introduced during feature extraction to weight and enhance key features, which can improve the ability to capture effective voice features. Finally, the feature vector is converted into an accurate recognition result, enabling the smart glasses to perform operations according to tasks (such as voice translation), overall enhancing the accuracy of voice recognition and adapting to the unique usage scenarios of smart glasses.
[0059] In one embodiment, based on the above embodiment, the voice recognition method of the smart glasses further includes: If the task to be executed is voice command recognition, the smart glasses query and execute the control instructions associated with the voice recognition result. Among them, the output layer uses a fully connected layer. After mapping the feature vector output by the intermediate layer to the category space, the Softmax function is applied to convert the output vector of the fully connected layer into a probability distribution, obtaining the prediction probability of each category to generate the voice recognition result.
[0060] In this embodiment, for the voice command recognition task, the output layer uses a fully connected layer mapping to map the feature vector output by the intermediate layer to the category space. Each neuron in the fully connected layer is connected to all neurons in the previous layer (intermediate layer). Suppose the dimension of the feature vector output by the intermediate layer is n, and there are m categories in the category space. The fully connected layer will contain m neurons. Through a series of weight parameters and bias parameters, the fully connected layer performs a linear transformation on the input feature vector to obtain an m-dimensional output vector. This process can be expressed by the mathematical formula: z = Wx + b; where x is the feature vector output by the intermediate layer, W is the weight matrix of the fully connected layer, b is the bias vector, and z is the output vector of the fully connected layer.
[0061] Furthermore, for the output vector z of the fully connected layer, the Softmax function is used for processing to convert it into a probability distribution. The role of the Softmax function is to convert the output values of the fully connected layer into probability values, and the sum of the probabilities of all categories is 1. In this way, the prediction probability of each category can be obtained.
[0062] According to the prediction probability of each category obtained by the Softmax function, the category with the highest probability is selected as the final voice recognition result. For example, if the prediction probability of category 1 is 0.8, the prediction probability of category 2 is 0.1, and the prediction probability of category 3 is 0.1, then the smart glasses will take the voice expression corresponding to category 1 as the voice recognition result and execute subsequent tasks based on this result, such as querying and executing the associated control instructions in the voice command recognition task.
[0063] The smart glasses first need to determine whether the currently executing task is voice command recognition. This can be based on the task recognition and matching methods mentioned above, such as the modes preset by the user or the keywords in the voice recognition results. For example, when the user says related expressions such as "execute command" or "issue instruction", the smart glasses determine the task as voice command recognition.
[0064] Once it is determined that the task is voice command recognition, the smart glasses will compare and query the voice recognition results with the pre-stored control instruction library. This instruction library contains a series of voice expressions and their corresponding control instructions. For example: The voice expression "turn on the light" corresponds to the control instruction of sending an on signal to the smart bulb.
[0065] The voice expression "wake up the voice assistant" corresponds to the control instruction of activating the built-in voice assistant of the smart glasses, and the user can further ask questions, query information or perform other operations to the voice assistant.
[0066] When the control instruction associated with the voice recognition result is found, the smart glasses can execute the corresponding control operation.
[0067] In one implementation, for voice command recognition, it can accurately query and execute the associated control instructions, and efficiently complete the task; the output layer uses a fully connected layer combined with the Softmax function to map the feature vector to the category space and convert it into a probability distribution, generating an accurate voice recognition result, improving the accuracy and reliability of the recognition.
[0068] In one embodiment, on the basis of the above embodiment, during the process of generating the voice recognition result, the output layer further introduces a context awareness mechanism, analyzes the previous voice content and interaction records, predicts the possible meaning of the current voice, and combines natural language processing technology to perform semantic understanding and reasoning on the voice recognition result.
[0069] In this embodiment, during the operation of the smart glasses, it will continuously record the user's voice content and related interaction operations. These data will be classified and stored in a local or cloud database, forming a detailed interaction log in chronological order. For example, record information such as each voice instruction issued by the user, the corresponding operation result, and the operation time. For the voice content, simple preprocessing will be performed, such as removing noise and performing preliminary voice feature extraction, etc., for subsequent analysis.
[0070] When the system needs to recognize the current speech, it will extract the previous speech content and interaction records from the database. According to the time point of the current speech, relevant information within a nearby time period is extracted to form a context window. For example, if the current speech is issued after a series of music playback operations by the user with the smart glasses, then the system will extract information such as the recent voice commands regarding music playback, the selected music genres, the played singers, etc. as the context.
[0071] The system will analyze the extracted context information using methods of natural language processing and machine learning. Through techniques such as keyword extraction and topic modeling, the main themes and intentions of the previous interactions are identified. Based on the analysis results of the context information, the system will predict the possible meaning of the current speech. It will generate a list of possible meanings according to the previous interaction patterns and themes, combined with some preliminary features of the current speech (such as the volume and intonation of the speech).
[0072] After predicting the possible meaning of the current speech, the system will combine natural language processing techniques to conduct in-depth semantic understanding and reasoning on the speech recognition result.
[0073] First, lexical analysis is performed on the speech recognition result, breaking it down into individual words or phrases and determining the part of speech and semantic category of each word. Then syntactic analysis is carried out to determine the grammatical relationships and sentence structures between the words. By analyzing the subject, predicate, object, etc. of the sentence, the basic semantic framework of the sentence is understood.
[0074] Combined with the context information and the predefined semantic knowledge base, semantic understanding is carried out on the results of lexical and syntactic analysis. The semantic knowledge base contains information such as semantic interpretations of various words and relationships between concepts. Based on semantic understanding, the system will conduct reasoning to solve possible ambiguity problems. At the same time, the system will also conduct further reasoning according to the user's historical preferences and usage habits.
[0075] Through the combination of the context awareness mechanism and natural language processing techniques, the system can optimize the preliminary speech recognition result. When generating the final speech recognition result, the context information, semantic understanding, and reasoning results are comprehensively considered, and the meaning that best conforms to the actual situation is selected as the final result. This can greatly improve the accuracy and intelligence of speech recognition, enabling the smart glasses to better understand the user's intentions and provide more personalized and accurate services. For example, when the user issues a vague voice command, the smart glasses can accurately execute corresponding operations according to the context and semantic analysis, such as playing specific music, querying relevant information, etc.
[0076] In one embodiment, based on the above embodiment, the middle layer also uses a convolutional layer to extract local features from the features after sequence modeling, and applies a pooling operation to reduce the dimension of the extracted features.
[0077] In this embodiment, in the middle layer, convolutional layers and pooling operations are usually used alternately. First, the convolutional layer performs local feature extraction on the features after sequence modeling to generate a set of feature maps. Then, the pooling operation reduces the dimensionality of these feature maps to obtain feature maps with a smaller size. This process can be repeated multiple times to form a stacked structure of multiple convolutional-pooling layers. Each time a convolutional-pooling operation is performed, the model can extract more advanced and abstract features while gradually reducing the dimensionality of the features, providing a more concise and effective feature representation for the subsequent processing of the output layer.
[0078] For example, after the first convolutional-pooling operation, the model may extract the basic phoneme features in the speech; after the second processing, it further extracts the features of syllables or words. Through this process of gradually deeper feature extraction and dimensionality reduction, the speech recognition model of the smart glasses can better understand and process speech signals, improving the accuracy and efficiency of recognition.
[0079] In one embodiment, based on the above embodiment, the speech recognition method of the smart glasses further includes: Before deploying the trained artificial intelligence model to the smart glasses, the model compression technology is used to convert the model parameters from high-precision floating-point numbers to low-precision integers, and then the compressed model is deployed to the smart glasses.
[0080] In this embodiment, an artificial intelligence model, especially a deep neural network model, usually has a large number of parameters. These parameters are generally stored in the form of high-precision floating-point numbers (such as 32-bit floating-point numbers), which will occupy a large amount of storage space and computing resources. The hardware resources of smart glasses (such as memory, storage capacity, and computing power) are relatively limited. Directly deploying an uncompressed model will cause many problems, such as slow device operation and excessive battery consumption. Therefore, it is of great significance to use the model compression technology to convert the model parameters into low-precision integers, which can not only reduce the storage space requirements of the model but also reduce the computational complexity and improve the running efficiency of the model on the smart glasses.
[0081] Before performing the parameter conversion, it is necessary to perform data statistics on the parameters of the model. By analyzing the distribution of all parameters in the trained model, the maximum and minimum values of the parameters are determined, so as to obtain the dynamic range of the parameters. For example, the mean and standard deviation of the parameters can be calculated to estimate the approximate range of the parameters. This range will be used as the basis for subsequent quantization to ensure that the information of the parameters can be retained as accurately as possible during the quantization process.
[0082] Assume that the distribution of the parameters is symmetric about zero. By determining a scaling factor, the high-precision floating-point parameters are mapped to the low-precision integer range. For example, for 8-bit integer quantization, the integer range is -128 to 127. The scaling factor is calculated as: scale = (the maximum value of the parameters - the minimum value of the parameters) / (the maximum value of the quantization range - the minimum value of the quantization range). Then, the floating-point parameter is divided by the scaling factor and rounded to the nearest integer to obtain the quantized integer parameter.
[0083] Considering that the parameter distribution may not be symmetric about zero, in addition to the scaling factor, a zero-point offset can be introduced. The zero-point offset is used to adjust the mapping relationship between the quantized integer and the original floating-point number, so that the quantization can more accurately reflect the actual distribution of the parameters. Similarly, for 8-bit integer quantization, the scaling factor and zero-point offset are calculated through a suitable algorithm to convert the floating-point parameter into an integer.
[0084] According to the selected quantization method, all parameters of the model (including weights and biases) are traversed and converted one by one. Each high-precision floating-point parameter is converted into the corresponding low-precision integer according to the scaling factor and zero-point offset. For example, a 32-bit floating-point weight parameter becomes an 8-bit integer after quantization. In this process, it is necessary to pay attention to controlling the quantization error to avoid a significant decrease in the model performance due to quantization.
[0085] Deploy the compressed and adjusted model to the smart glasses. This involves transferring the model file to the storage device of the smart glasses and configuring the corresponding running environment in the operating system of the smart glasses. When the voice recognition system of the smart glasses starts, it will load the compressed model and use low-precision integer operations to perform the voice recognition task. During operation, since the model parameters are stored and calculated in low-precision integer form, the required storage space and computing resources are greatly reduced, thus improving the voice recognition performance and battery life of the smart glasses.
[0086] In one embodiment, based on the above embodiment, the voice recognition method of the smart glasses further includes: When voice recognition is not required, control the voice recognition module of the smart glasses to enter the low-power sleep state; When a wake-up word is detected, wake up the voice recognition module to perform voice recognition.
[0087] In this embodiment, the system of the smart glasses will monitor in real time whether there is a voice recognition requirement.
[0088] Optionally, if the user explicitly tells the smart glasses to stop the voice recognition function through interaction methods such as touching or pressing a button, the system will record this instruction and prepare to put the voice recognition module into the sleep state. For example, when the user long-presses a specific button on the smart glasses, after the system receives this operation signal, it will parse and confirm this instruction.
[0089] Optionally, the system will build-in a time rule. When no voice input is received within a period of time (such as 5 minutes) and no related voice recognition tasks are triggered, it is determined that voice recognition is not required. The system will run a timer in the background, starting from the completion of the last voice recognition operation. When the timing reaches the preset time threshold, the sleep determination is triggered.
[0090] Optionally, if the application related to voice recognition is closed or switched to a function interface that does not rely on voice recognition, the system will also consider that voice recognition is not required currently. For example, when the user switches from a voice navigation application to a pure text reading application, the system will perceive the change in the application state.
[0091] Once it is determined that it is necessary to enter the low-power sleep state, the system will first release most of the resources occupied by the voice recognition module. This includes stopping the real-time data acquisition of the microphone, closing the relevant data transmission channels, and pausing operations such as preprocessing and feature extraction of voice signals. For example, stop sending a power supply signal to the microphone and interrupt the link for transmitting audio data from the microphone to the processing chip.
[0092] Optionally, reduce the working frequency of the voice recognition module to reduce the consumption of its computing resources. The voice recognition module usually consists of multiple processor cores and related circuits. In the sleep state, the system will adjust the clock frequencies of these processors so that they maintain the basic standby function at the lowest operating speed. For example, reduce the working frequency of the processor from the normal 1 GHz to 100 MHz.
[0093] The system will record the status information of the voice recognition module before entering the sleep state, such as the current configuration parameters, the progress of unfinished tasks, etc., and save this information to the non-volatile memory. In this way, it can quickly resume the previous state during subsequent wake-up. For example, save the current version of the voice recognition model, acoustic model parameters, etc.
[0094] Optionally, when the voice recognition module of the smart glasses is in the low-power sleep state, a lightweight wake-word detection sub-module will still be kept in the working state. This sub-module will continuously monitor the sound signals in the environment and perform simple feature extraction on them. For example, extract the frequency features, energy features, etc. of the sound to determine whether it may contain a wake word.
[0095] Match the extracted features with the pre-set wake word feature templates. These wake word feature templates are learned from a large amount of voice data during the training phase and can accurately represent the feature patterns of wake words. For example, the wake word of the smart glasses is "Hello, Xiaojing", and the system will compare the real-time extracted voice features with the feature template of "Hello, Xiaojing".
[0096] To avoid false wake-up, the system will set a confidence threshold. When the confidence of the matching result exceeds this threshold, it is determined that a wake word is detected. The calculation of confidence will comprehensively consider factors such as the similarity of feature matching and the clarity of the voice. For example, when the matching similarity reaches more than 80% and the signal-to-noise ratio of the voice meets certain requirements, it is determined that the confidence is high enough.
[0097] Once a wake word is detected, the system will immediately send a wake-up signal to the speech recognition module to start the restart process of the module. First, restore the power supply and data acquisition function of the microphone, re-establish the data transmission channel, so that the voice signal can be normally transmitted to the processing chip.
[0098] Read the previously saved status information from the non-volatile memory, and restore the configuration parameters and operating status of the speech recognition module to the state before entering the sleep state. For example, load the previously used speech recognition model and acoustic model parameters to ensure that speech recognition can be performed with the same recognition criteria.
[0099] After completing the module restart and status restoration, the speech recognition module officially starts the speech recognition process. Comprehensively process the subsequent received voice signals, including steps such as noise reduction, feature extraction, acoustic model matching, and language model decoding, and finally output accurate speech recognition results.
[0100] Through this low-power sleep and wake-up mechanism, the smart glasses can significantly reduce power consumption when not performing speech recognition, extend the battery life, and at the same time quickly respond when needed to provide timely speech recognition services to users.
[0101] In one embodiment, based on the above embodiment, before the step of performing multi-scale feature extraction on the voice signal by the input layer of the pre-trained artificial intelligence model and introducing an attention mechanism during the feature extraction process to learn the importance weights of different features to weighted enhance key features, it further includes: Adopt an adaptive filtering algorithm to dynamically adjust the filtering parameters, and remove the environmental noise in the voice signal based on the adjusted filtering parameters.
[0102] In this embodiment, in the actual application scenario, the voice signals collected by the smart glasses are often interfered by various environmental noises, such as noisy background voices, machine roars, wind noises, etc. These noises will seriously affect the quality of the voice signals and reduce the accuracy of subsequent voice recognition.
[0103] The core idea of the adaptive filtering algorithm is to continuously adjust the coefficients of the filter to minimize the error between the output of the filter and the desired signal. Optional adaptive filtering algorithms include the Least Mean Square (LMS) algorithm, the Recursive Least Squares (RLS) algorithm, etc. Here, the LMS algorithm is preferably used.
[0104] Before starting to process the voice signals, it is necessary to initialize the coefficients of the adaptive filter, usually initializing them to a zero vector. At the same time, set appropriate step size factors and filter orders. The selection of the filter order needs to be adjusted according to the actual situation. Generally speaking, the higher the order, the better the performance of the filter, but the computational complexity will also increase accordingly.
[0105] During the processing of the voice signals, the adaptive filter will receive the input signal and the reference signal in real time and continuously adjust the coefficients of the filter according to the error calculation and coefficient update formulas. As time goes by, the coefficients of the filter will gradually converge to the optimal value, enabling the filter to better adapt to the changes in environmental noises.
[0106] To ensure the effectiveness of the adaptive filtering algorithm, it is necessary to judge the convergence situation of the filter. The convergence of the filter can be judged by monitoring the mean square value of the error signal. When the mean square value of the error signal stabilizes within a small range, it is considered that the filter has converged.
[0107] After the coefficients of the adaptive filter converge to stable values, the adjusted filtering parameters can be used to filter the voice signals to remove the environmental noises therein. The specific operation is to pass the input voice signals through the filter with the adjusted coefficients to obtain the filtered voice signals. In the filtered voice signals, the components of the environmental noises are effectively suppressed, while the characteristic information of the voice signals is better retained, providing a cleaner and more accurate input for subsequent multi-scale feature extraction and attention mechanism learning.
[0108] By adopting the adaptive filtering algorithm to dynamically adjust the filtering parameters and remove the environmental noises, the quality of the voice signals can be significantly improved, thereby enhancing the voice recognition performance based on the pre-trained artificial intelligence model.
[0109] In addition, referring to Figure 2 , this application embodiment also provides a control device Z10, including: The acquisition module Z11 is used to collect voice signals in real time during the operation of the smart glasses. Among them, the smart glasses are provided with multiple microphones at different positions for collecting voice signals. The extraction module Z12 is used to perform multi-scale feature extraction on the voice signals based on the input layer of a pre-trained artificial intelligence model, and introduce an attention mechanism during the feature extraction process to learn the importance weights of different features, so as to enhance the key features by weighting. The processing module Z13 is used to deeply fuse the weighted multi-scale features based on the middle layer of the artificial intelligence model, and perform non-linear activation and sequence modeling after the feature fusion. The conversion module Z14 is used to convert the feature vector output by the middle layer into a speech recognition result based on the output layer of the artificial intelligence model. The execution module Z15 is used for the smart glasses to perform corresponding operations according to the speech recognition result based on the current execution task. Among them, if the execution task is speech translation, the speech recognition result is converted into the corresponding translation text and output for display.
[0110] Optionally, the control device Z10 can be a virtual control device (such as a virtual machine), or a physical device (such as a physical device other than the smart glasses that can execute the corresponding method).
[0111] In addition, an embodiment of the present application also provides a smart glasses, and the internal architecture of the smart glasses can be as Figure 3 shown, including a processor, a memory, a communication interface, and an input interface connected through a system bus. Among them, the processor is used to provide computing and control capabilities. The memory includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database is used to store data called by the computer program. The communication interface is used to communicate with an external terminal. The input interface is used to receive signals input by an external device. When the computer program is executed by the processor, it is used to implement a speech recognition method of a smart glasses as described in the above embodiments.
[0112] Those skilled in the art can understand that Figure 3 the structure shown in
[0113] In addition, the present application also provides a computer-readable storage medium, which includes a computer program. When the computer program is executed by a processor, the steps of the voice recognition method of the smart glasses as described in the above embodiments are implemented. It can be understood that the computer-readable storage medium in this embodiment can be a volatile readable storage medium or a non-volatile readable storage medium.
[0114] In summary, for the voice recognition method, control device, smart glasses, and computer-readable storage medium provided in the embodiments of the present application, the smart glasses are provided with microphones at multiple different positions to collect voice signals in real time, which can, to a certain extent, make up for the problems of fixed wearing positions and large limitations in sound collection, and collect the voices of users and the people opposite more comprehensively and accurately; and an attention mechanism is introduced during feature extraction to weight and enhance key features, which can improve the ability to capture effective voice features. Finally, the feature vector is converted into an accurate recognition result, enabling the smart glasses to perform operations according to tasks (such as voice translation). Overall, the voice recognition accuracy is enhanced, adapting to the unique usage scenarios of smart glasses.
[0115] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium provided in the present application and the embodiments can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0116] It should be noted that in this text, the term "comprise", "include" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, device, article or method comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, device, article or method. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, device, article or method comprising such element.
[0117] The above are only the preferred embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.
Claims
1. A voice recognition method for a smart glasses, characterized in that, Including: During the operation of the smart glasses, real-time collection of voice signals is performed; among them, the smart glasses are equipped with microphones at multiple different positions for collecting voice signals; Based on the input layer of the pre-trained artificial intelligence model, multi-scale feature extraction is performed on the voice signals, and an attention mechanism is introduced during the feature extraction process to learn the importance weights of different features, so as to weight and enhance the key features; Based on the middle layer of the artificial intelligence model, the multi-scale features after weighted enhancement are deeply fused, and non-linear activation and sequence modeling are performed after the feature fusion; Based on the output layer of the artificial intelligence model, the feature vector output by the middle layer is converted into a speech recognition result; The smart glasses perform corresponding operations according to the speech recognition result based on the current execution task; among them, if the execution task is speech translation, the speech recognition result is converted into the corresponding translation text and output for display.
2. The voice recognition method of the smart glasses according to claim 1, wherein, The speech recognition method of the smart glasses further includes: If the execution task is speech command recognition, the smart glasses query and execute the control instruction associated with the speech recognition result; Among them, the output layer uses a fully connected layer to map the feature vector output by the middle layer to the category space, and then applies the Softmax function to convert the output vector of the fully connected layer into a probability distribution to obtain the prediction probability of each category, so as to generate the speech recognition result.
3. The voice recognition method of the smart glasses according to claim 1 or 2, characterized in that, During the process of generating the speech recognition result, the output layer also introduces a context awareness mechanism to analyze the previous speech content and interaction records, predict the possible meaning of the current speech, and combine natural language processing technology to perform semantic understanding and reasoning on the speech recognition result.
4. The voice recognition method of the smart glasses according to claim 1, characterized in that, The middle layer also uses a convolutional layer to perform local feature extraction on the features after sequence modeling, and applies a pooling operation to reduce the dimension of the extracted features.
5. The voice recognition method of the smart glasses according to claim 1, characterized in that, The speech recognition method of the smart glasses further includes: Before deploying the trained artificial intelligence model to the smart glasses, a model compression technology is adopted to convert the model parameters from high-precision floating-point numbers to low-precision integers, and then the compressed model is deployed to the smart glasses.
6. The voice recognition method of the smart glasses according to claim 1, characterized in that, The speech recognition method of the smart glasses further includes: When speech recognition is not required, control the speech recognition module of the smart glasses to enter a low-power sleep state; When a wake-up word is detected, the speech recognition module is awakened to perform speech recognition.
7. The voice recognition method of the smart glasses according to claim 1, characterized in that, Before the step of performing multi-scale feature extraction on the voice signals based on the input layer of the pre-trained artificial intelligence model and introducing an attention mechanism during the feature extraction process to learn the importance weights of different features so as to weight and enhance the key features, it further includes: An adaptive filtering algorithm is used to dynamically adjust the filtering parameters, and the environmental noise in the voice signals is removed based on the adjusted filtering parameters.
8. A control device, characterized in that, Including: A collection module for real-time collection of voice signals during the operation of the smart glasses; among them, the smart glasses are equipped with microphones at multiple different positions for collecting voice signals; An extraction module for performing multi-scale feature extraction on the voice signals based on the input layer of the pre-trained artificial intelligence model, and introducing an attention mechanism during the feature extraction process to learn the importance weights of different features so as to weight and enhance the key features; A processing module, which is used to perform deep fusion on the weighted and enhanced multi-scale features based on the middle layer of the artificial intelligence model, and perform non-linear activation and sequence modeling after feature fusion; A conversion module, which is used to convert the feature vector output by the middle layer into a speech recognition result based on the output layer of the artificial intelligence model; An execution module, which is used for the smart glasses to perform corresponding operations according to the speech recognition result based on the current execution task; wherein, if the execution task is speech translation, the speech recognition result is converted into corresponding translation text and output for display.
9. An intelligent glasses, characterized in that, The smart glasses include a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, the steps of the speech recognition method of the smart glasses according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium. When the computer program is executed by the processor, the steps of the speech recognition method of the smart glasses according to any one of claims 1 to 7 are implemented.