Intelligent glasses-based voice recognition method, intelligent glasses and storage medium

By combining audio and lip-reading technologies, smart glasses have improved speech recognition accuracy in complex environments, solving the problem of low speech recognition rates in noisy environments and improving human-computer interaction efficiency.

CN116013285BActive Publication Date: 2026-06-12GEER TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GEER TECH CO LTD
Filing Date
2022-11-30
Publication Date
2026-06-12

AI Technical Summary

Technical Problem

Smart glasses have low voice recognition accuracy in complex environments, resulting in low efficiency of human-computer interaction.

Method used

Semantic recognition is performed by acquiring user voice signals, and lip reading models are used to detect lip shape changes. Audio and lip reading semantic information are fused to determine target semantic information, and ultrasonic detection modules and multi-microphone noise reduction technology are used to process noise.

Benefits of technology

It improves the accuracy of voice recognition in complex environments, enhances human-computer interaction efficiency, and avoids the impact of noise interference on voice recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116013285B_ABST
    Figure CN116013285B_ABST
Patent Text Reader

Abstract

The application discloses a voice recognition method based on smart glasses, smart glasses and a storage medium. The voice recognition method based on smart glasses comprises the following steps: acquiring a user sound signal, performing semantic recognition on the user sound signal to obtain audio semantic information; detecting a user's lip shape change feature, inputting the lip shape change feature into a converged lip language recognition model, and recognizing to obtain the user's lip language semantic information; and determining target semantic information according to the audio semantic information and the lip language semantic information. The application can improve the accuracy of voice recognition of smart glasses in a complex environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of augmented reality technology, and in particular to a speech recognition method based on smart glasses, smart glasses, and storage medium. Background Technology

[0002] Currently, user interaction with smart glasses typically relies on touchscreens, keyboards / mouse, remote controls, and gestures. However, for smart glasses, especially in everyday wear, these methods require users to manually operate specific devices or touch specific components (such as touchscreens on the temples of the glasses). This presents various limitations and results in low efficiency. Furthermore, voice recognition-based interactions often suffer from interference from ambient noise and other people's voices, significantly reducing recognition accuracy. Summary of the Invention

[0003] The main objective of this application is to provide a speech recognition method based on smart glasses, smart glasses, and a storage medium, aiming to solve the technical problem of low speech recognition accuracy in complex environments.

[0004] To achieve the above objectives, this application provides a speech recognition method based on smart glasses, the method comprising:

[0005] Acquire user voice signals and perform semantic recognition on the user voice signals to obtain audio semantic information;

[0006] The user's lip shape change features are detected and input into a converged lip reading recognition model to identify the user's lip reading semantic information;

[0007] The target semantic information is determined based on the audio semantic information and the lip-reading semantic information.

[0008] Optionally, the step of determining the target semantic information based on the audio semantic information and the lip-reading semantic information includes:

[0009] Detect whether there is an audio recognition failure time domain segment in the time domain of the audio semantic information;

[0010] If there is no audio recognition failure time segment, then the audio semantic information is used as the target semantic information;

[0011] If there is a time domain segment where audio recognition fails, the audio recognition failure time domain segment is repaired according to the lip reading semantic information to obtain repaired audio semantic information, and the repaired audio semantic information is used as the target semantic information.

[0012] Optionally, before the step of repairing the audio recognition failure time segment based on the lip-reading semantic information, the method further includes:

[0013] Detect whether there is a time domain segment in the lip-reading semantic information where lip-reading recognition failed;

[0014] If there is no time domain segment where lip reading recognition fails, then the following step is performed: repairing the audio recognition failure time domain segment based on the lip reading semantic information;

[0015] If there is a time domain segment where lip reading fails, then it is detected whether there is any overlap in the time domain between the audio recognition failure time domain segment and the lip reading failure time domain segment;

[0016] If there are no overlapping time domains, then the following step is performed: repairing the audio recognition failure time domain segment based on the lip-reading semantic information;

[0017] If there are overlapping time domains, semantic prediction is performed on the overlapping time domain segments where lip reading recognition failed, based on the lip reading semantic information and the audio semantic information, to obtain predicted semantic information; and target semantic information is determined based on the predicted semantic information, the lip reading semantic information, and the audio semantic information.

[0018] Optionally, the step of determining the target semantic information based on the predicted semantic information, the lip-reading semantic information, and the audio semantic information includes:

[0019] The predicted semantic information is filled into the overlapping time domain segments where lip reading failed, thus completing the predicted semantic filling process.

[0020] After the predicted semantic filling process is completed, the audio recognition failure time domain segment is repaired according to the lip reading semantic information to obtain the repaired audio semantic information, and the repaired audio semantic information is used as the target semantic information.

[0021] Optionally, the method further includes:

[0022] Obtain user lip shape samples corresponding to multiple lip shape change features, and the actual vocal information of the user corresponding to the user lip shape samples;

[0023] The actual vocal information of the user corresponding to the user lip-sync sample is used as the user vocal sample associated with the user lip-sync sample.

[0024] Multiple user lip-reading samples and user voice samples associated with each user voice sample are used as training samples to train the model, resulting in a converged lip-reading recognition model.

[0025] Optionally, the smart glasses include a glasses body and an ultrasonic detection module. The glasses body includes a frame and temples. The ultrasonic detection module is disposed on the frame, and the detection orientation of the ultrasonic detection module is directed towards the user's lips. The step of detecting the user's lip shape change features includes:

[0026] The ultrasonic detection module detects the user's lip shape changes.

[0027] Optionally, the smart glasses may further include a first sound detection module and a second sound detection module. The first sound detection module is disposed on the bridge of the frame near the user's lips, and the second sound detection module is disposed on the temple of the frame away from the user's lips. Before the step of acquiring the user's voice signal, the method further includes:

[0028] The first external audio information is collected through the first sound detection module, and the second external audio information is collected through the second sound detection module.

[0029] The environmental noise signal is identified based on the second external audio information;

[0030] Based on the identified environmental noise signal, the environmental noise signal in the first external audio information is denoised to obtain the user's voice signal.

[0031] Optionally, the step of performing noise reduction processing on the environmental noise signal in the first external audio information based on the identified environmental noise information includes:

[0032] Determine the noise cancellation signal corresponding to the environmental noise information, wherein the noise cancellation signal has the same frequency and amplitude as the environmental noise information, but opposite phase.

[0033] The noise-canceling wave signal is superimposed on the first external audio information to perform noise reduction processing on the environmental noise signal in the first external audio information.

[0034] This application also provides a smart glasses, which is a physical device. The smart glasses include: a memory, a processor, and a program for the speech recognition method based on the smart glasses, which is stored in the memory and can run on the processor. When the program for the speech recognition method based on the smart glasses is executed by the processor, it can implement the steps of the speech recognition method based on the smart glasses as described above.

[0035] This application also provides a storage medium, which is a computer-readable storage medium, on which a program implementing a speech recognition method based on smart glasses is stored. The program implementing the speech recognition method based on smart glasses is executed by a processor to implement the steps of the speech recognition method based on smart glasses as described above.

[0036] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the speech recognition method based on smart glasses as described above.

[0037] Considering that speech recognition is easily affected by noisy and complex environments, thus reducing the efficiency of human-computer interaction, in order to prevent speech recognition from being affected by noisy environments, improve the accuracy of speech recognition in complex environments, and thus improve the human-computer interaction efficiency of smart glasses systems. The technical solution of this application is to acquire user voice signals, perform semantic recognition on the user voice signals to obtain audio semantic information; detect the user's lip shape change features, input the lip shape change features into a convergent lip reading recognition model, and identify the user's lip reading semantic information; based on the audio semantic information and lip reading semantic information, determine the target semantic information, thereby realizing the comprehensive identification of the user's target semantic information by fusing and analyzing the user's lip shape change features and vocal audio features. This avoids the phenomenon that the speech recognition rate is greatly reduced when the user's voice is mixed with ambient noise, other people's voices, etc., which is the case with semantic recognition based solely on vocal audio. This improves the accuracy of speech recognition in complex environments for smart glasses. By combining speech recognition and lip reading recognition, this application analyzes and integrates the two to obtain a semantic information that is more in line with the context as an effective input to the system, solving the technical problem of low speech recognition accuracy of smart glasses in complex environments and improving the human-computer interaction efficiency of smart glasses. Attached Figure Description

[0038] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0039] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0040] Figure 1 This is a flowchart illustrating the first embodiment of the speech recognition method based on smart glasses in this application;

[0041] Figure 2This is a detailed flowchart of step S30 in the second embodiment of the speech recognition method based on smart glasses in this application;

[0042] Figure 3 This is a flowchart illustrating the third embodiment of the speech recognition method based on smart glasses in this application;

[0043] Figure 4 This is a schematic diagram of the mechanical structure of smart glasses according to an embodiment of this application;

[0044] Figure 5 This is a schematic diagram of the hardware structure of the hardware operating environment involved in the smart glasses in this application embodiment.

[0045] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0046] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0047] In this embodiment, the smart glasses of this application may be, for example, Mixed Reality (MR) glasses, Augmented Reality (AR) glasses, Virtual Reality (VR) glasses, Extended Reality (XR) glasses, or some combination thereof.

[0048] Example 1

[0049] Please refer to Figure 1 This embodiment provides a voice recognition method based on smart glasses, the voice recognition method including:

[0050] Step S10: Acquire user voice signal, and perform semantic recognition on the user voice signal to obtain audio semantic information;

[0051] In this embodiment, the user's voice signal can be obtained by capturing the user's speech through a microphone. It is understood that this microphone can be integrated into the smart glasses or can communicate with the smart glasses independently of them.

[0052] Those skilled in the art will understand that a pre-defined speech recognition algorithm can be used to perform semantic recognition on the user's voice signal to obtain audio semantic information. This speech recognition algorithm has been extensively studied by those skilled in the art and will not be elaborated upon here.

[0053] Step S20: Detect the user's lip shape change features, input the lip shape change features into the converged lip reading recognition model, and identify the user's lip reading semantic information;

[0054] In one embodiment, a camera can capture multiple consecutive images of a user, and based on these consecutive images, the user's lip-shape changes can be identified. The multiple consecutive images include images of the user's lips. It is understood that the camera can be integrated into smart glasses or can communicate with smart glasses independently. Specifically, the step of the smart glasses identifying the user's lip-shape changes based on the multiple consecutive images may include:

[0055] Step A10: Extract features from the multiple consecutive images according to the lip feature extraction algorithm and the mouth shape contour feature extraction algorithm to identify the user's mouth shape change features. The lip feature extraction algorithm includes at least one of a template-based feature extraction algorithm or an image pixel-based feature extraction algorithm. The mouth shape contour feature extraction algorithm includes at least one of a deformable template algorithm or a Snakes algorithm.

[0056] As an example, the method also includes:

[0057] Step B10: Obtain user lip shape samples corresponding to multiple lip shape change features, and the actual vocal information of the user corresponding to the user lip shape samples;

[0058] Step B20: Use the actual vocal information of the user corresponding to the user lip-sync sample as the user vocal sample associated with the user lip-sync sample.

[0059] Step B30: Use multiple user lip shape samples and user voice samples associated with each user voice sample as training samples to train the model and obtain a converged lip reading recognition model.

[0060] In this embodiment, during model training, the trained lip-reading model can be used to identify user lip-shape samples. If the matching degree between the identified lip-reading semantic information and the semantic information corresponding to the associated user vocal samples reaches a preset matching degree threshold, the lip-reading model is considered converged. Otherwise, it indicates that the lip-reading model has not yet converged and needs to be trained with more training samples to improve the accuracy of the lip-reading model in recognizing lip shape changes. The preset matching degree threshold can be set by those skilled in the art according to actual conditions; this embodiment does not impose a specific limitation. For example, the preset matching degree threshold could be 95%.

[0061] This embodiment acquires multiple user lip shape samples corresponding to lip shape change features, as well as the actual vocal information of the user corresponding to the lip shape sample. The actual vocal information of the user corresponding to the lip shape sample is used as the user vocal sample associated with the lip shape sample. Then, multiple user lip shape samples and the user vocal samples associated with each user vocal sample are used as training samples to train the model and obtain a converged lip reading recognition model. This effectively ensures the accuracy of the converged lip reading recognition model in recognizing lip reading semantic information.

[0062] Step S30: Determine the target semantic information based on the audio semantic information and the lip-reading semantic information.

[0063] Considering that speech recognition is easily affected by noisy and complex environments, thus reducing the efficiency of human-computer interaction, in order to prevent speech recognition from being affected by noisy environments, improve the accuracy of speech recognition in complex environments, and thus improve the human-computer interaction efficiency of smart glasses systems. The technical solution of this embodiment is to acquire the user's voice signal, perform semantic recognition on the user's voice signal to obtain audio semantic information; detect the user's lip shape change features, input the lip shape change features into a convergent lip reading recognition model, and identify the user's lip reading semantic information; based on the audio semantic information and lip reading semantic information, determine the target semantic information, thereby realizing the comprehensive identification of the user's target semantic information by fusing and analyzing the user's lip shape change features and vocal audio features. This avoids the phenomenon that the speech recognition rate is greatly reduced when the user's voice is mixed with ambient noise, other people's voices, etc., which is the case with semantic recognition based solely on vocal audio. This improves the accuracy of speech recognition of smart glasses in complex environments. This embodiment solves the technical problem of low accuracy of speech recognition of smart glasses in complex environments by combining speech recognition and lip reading recognition and fusing the two to analyze a semantic information that is more in line with the context as an effective input to the system, thus improving the human-computer interaction efficiency of smart glasses.

[0064] In one possible implementation, please refer to Figure 2 The step of determining the target semantic information based on the audio semantic information and the lip-reading semantic information includes:

[0065] Step S31: Detect whether there is an audio recognition failure time domain segment in the time domain of the audio semantic information;

[0066] Step S32: If there is no audio recognition failure time segment, then the audio semantic information is used as the target semantic information;

[0067] Step S33: If there is a time domain segment where audio recognition fails, then the audio recognition failure time domain segment is repaired according to the lip reading semantic information to obtain the repaired audio semantic information, and the repaired audio semantic information is used as the target semantic information.

[0068] In this embodiment, since the audio semantic information obtained by speech recognition is easily affected by ambient noise and other people's voices in a noisy environment, there may be a time domain in which the audio semantics cannot be effectively recognized during the speech recognition process. That is, the audio recognition fails in a certain time domain. The time domain in which the audio recognition fails is called the audio recognition failure time domain segment.

[0069] To aid understanding, consider two examples. For instance, if a user says "I want to open the refrigerator to get a drink" between 12:30:12:23 and 12:30:14:12, the recognized audio semantic information in that time range is "I want to open it to get a drink" due to noise. This is because the vocal signal between 12:30:13:28 and 12:30:13:94 cannot be effectively recognized. Therefore, the time range from 12:30:13:28 to 12:30:13:94 is the audio recognition failure segment. This can be corrected by identifying the lip-reading semantic information "refrigerator" corresponding to the lip-reading changes in that time range, and then repairing the audio recognition failure segment to obtain the repaired audio semantic information: "I want to open the refrigerator to get a drink".

[0070] For example, if a user says "I want to open the refrigerator to get a drink" between 12:30:12:23 and 12:30:14:12, the audio semantic information recognized in that time domain is "I want to open the refrigerator to get an elephant" due to noise. However, the audio semantic information "elephant" recognized between 12:30:13:54 and 12:30:14:12 is very inconsistent with the context, therefore the recognition fails, and the failed "elephant" is removed from the audio... After removing the semantic information, we get "I want to open the refrigerator to get it". At this time, the time domain from 12:30:13:55 to 12:30:14:12 is the audio recognition failure time domain segment. We can identify the lip-reading semantic information "drink" corresponding to the lip shape change features in the time domain from 12:30:13:55 to 12:30:14:12, and then repair the audio recognition failure time domain segment to get the repaired audio semantic information: "I want to open the refrigerator to get the drink".

[0071] This embodiment detects whether there is a time domain segment in the audio semantic information where audio recognition fails. If no such segment exists, the audio semantic information is used as the target semantic information. If a time domain segment does exist, the segment is repaired based on lip-reading semantic information to obtain repaired audio semantic information. This repaired audio semantic information is then used as the target semantic information, and the audio semantic information is corrected based on lip-reading semantic information to obtain the target semantic information. In other words, lip-reading semantic information and audio semantic information can be matched. This allows lip-reading semantic information to compensate for the problem of audio segments not being recognized or being misrecognized due to noise interference, thereby more accurately and effectively recognizing the target semantic information that the user wants to express and improving the accuracy of speech recognition in complex environments.

[0072] In this embodiment, just as in noisy environments, the audio semantic information obtained from speech recognition is easily affected by noise, resulting in the inability to effectively recognize audio semantics in a certain time domain. Similarly, due to the limitations of lip-reading algorithms, or because the lip changes during user speech are not obvious enough, lip-reading semantics cannot be effectively recognized in a certain time domain.

[0073] Based on this, this embodiment proposes an implementable method. In this embodiment, before the step of repairing the audio recognition failure time segment based on the lip-reading semantic information, the method further includes:

[0074] Step C10: Detect whether there is a time domain segment in the temporal domain of the lip reading semantic information where lip reading recognition failed;

[0075] Step C20: If there is no lip-reading recognition failure time domain segment, then perform the step of repairing the audio recognition failure time domain segment based on the lip-reading semantic information;

[0076] Step C30: If there is a time domain segment where lip reading fails, then detect whether there is an overlap in the time domain between the audio recognition failure time domain segment and the lip reading failure time domain segment.

[0077] Step C40: If there are no overlapping time domains, then perform the step of repairing the audio recognition failure time domain segment based on the lip reading semantic information.

[0078] Step C50: If there are overlapping time domains, then based on the lip-reading semantic information and the audio semantic information, perform semantic prediction on the lip-reading recognition failure time domain segments in the overlapping time domains to obtain predicted semantic information; and determine the target semantic information based on the predicted semantic information, the lip-reading semantic information and the audio semantic information.

[0079] For example, in step A50, the step of determining the target semantic information based on the predicted semantic information, the lip-reading semantic information, and the audio semantic information includes:

[0080] Step D10: Fill the predicted semantic information into the overlapping time domain segments where lip reading failed, thus completing the predicted semantic filling process;

[0081] Step D20: After the predicted semantic filling process is completed, the audio recognition failure time domain segment is repaired according to the lip reading semantic information to obtain the repaired audio semantic information, and the repaired audio semantic information is used as the target semantic information.

[0082] To aid understanding, let's consider two examples. For instance, if a user says "I'm hungry and want to eat sweet potatoes because they smell good" between 12:30:12:23 and 12:30:14:86, the lip-reading algorithm might fail to recognize the lip-reading information "I'm hungry for sweet potatoes because they smell good" in that time period. This is because the lip-reading features corresponding to certain semantic sounds are not clearly visible due to limitations in the algorithm's accuracy or the lack of distinct lip changes. The lip-reading information between 12:30:12:68 and 12:30:13:32 cannot be identified. Therefore, the time period between 12:30:12:68 and 12:30:13:32 is considered the time period where lip-reading failed to recognize the information.

[0083] For example, if a user says "I'm hungry and want to eat sweet potatoes because they smell good" between 12:30:12:23 and 12:30:14:86, the lip-reading algorithm might fail to recognize the audio semantic information "I'm hungry and want to eat a tree because sweet potatoes smell good" because of limitations in the algorithm's accuracy, or because some semantic lip changes are not obvious or easily confused. Specifically, the lip-reading semantic information "tree" recognized between 12:30:12:32 and 12:30:13:85 is highly inappropriate for the context, thus the recognition fails. Therefore, the time range from 12:30:12:32 to 12:30:13:85 is considered the time range where lip-reading recognition failed.

[0084] It's easy to understand that if there's a time-domain segment where lip-reading fails, and this time-domain segment overlaps with the time-domain segment where audio recognition fails, then the lip-reading semantic information alone cannot effectively repair the audio recognition failure segment. For example, if a user says "I'm hungry and want to eat sweet potatoes because they smell good" between 12:30:12:23 and 12:30:14:86, and if the user's time-domain segment between 12:30:12:32 and 12:30:13:85 is both a lip-reading failure segment and an audio recognition failure segment, and the actual target semantics of the time-domain segment between 12:30:12:32 and 12:30:13:85 is "sweet potatoes," then it's understandable that neither audio recognition nor lip-reading recognition can at most identify "I'm hungry and want to eat sweet potatoes because they smell good." In other words, based on the lip-reading semantic information, it is impossible to fill the audio recognition failure time domain segment with the unrecognized word "sweet potato" to obtain the complete and accurate target semantic information "I am hungry and want to eat sweet potato because sweet potato is very fragrant".

[0085] Therefore, the technical solution of this embodiment is to detect whether there is an overlap between the audio recognition failure time domain segment and the lip-reading failure time domain segment if there is a lip-reading failure time domain segment. If there is no overlap, it is determined that the audio recognition failure time domain segment can be effectively repaired based on the lip-reading semantic information, and the step of repairing the audio recognition failure time domain segment based on the lip-reading semantic information is executed. If there is an overlap, it is determined that the audio recognition failure time domain segment cannot be effectively repaired based on the lip-reading semantic information. Therefore, semantic prediction is performed on the overlapping lip-reading failure time domain segment based on the lip-reading semantic information and the audio semantic information to obtain predicted semantic information. For example, based on both the lip-reading semantic information and the audio semantic information, only "I am hungry and want to eat XX because sweet potatoes smell good" can be identified. In the time domain from 12:12.32 to 12:13.86, the character “XX” could not be identified. At this time, based on lip-reading semantic information and audio semantic information, combined with the preceding and following context, semantic prediction was performed on the overlapping time domain segments where lip-reading recognition failed. The predicted semantic information “XX” for the time domain from 12:12.32 to 12:13.86 was “sweet potato”, thus obtaining complete and accurate target semantic information. This further improved the accuracy of speech recognition in complex environments by fusing and analyzing speech recognition and lip-reading recognition to obtain a more context-appropriate semantic information as the effective input of the system.

[0086] Example 2

[0087] Based on the above embodiments of this application, in another embodiment of this application, the same or similar content as the above embodiments can be referred to the above description, and will not be repeated hereafter. Based on this, the smart glasses include a glasses body and an ultrasonic detection module. The glasses body includes a frame and temples. The ultrasonic detection module is disposed on the frame, and the detection orientation angle of the ultrasonic detection module is directed towards the user's lips. The step of detecting the user's lip shape change features includes:

[0088] Step E10: The ultrasonic detection module detects the user's lip shape change characteristics.

[0089] In this embodiment, the ultrasonic detection module can be an ultrasonic detection array, which consists of multiple ultrasonic detection units arranged in an array. Each ultrasonic detection unit may include an ultrasonic transmitter and an ultrasonic receiver. Specifically, this embodiment can control all ultrasonic detection units of the ultrasonic detection array to continuously emit ultrasonic signals to the user's lip area and receive the echo signals corresponding to the continuously returned ultrasonic signals. Then, the user's lip shape change characteristics can be determined based on the continuously returned echo signals.

[0090] Those skilled in the art will understand that controlling each ultrasonic cell corresponding to the ultrasonic array to emit ultrasonic signals at multiple detection points towards the user's lip area, and receiving the echo signals corresponding to the ultrasonic signals, and determining the user's lip shape characteristics based on the echo signals. Among these, the reception time and signal strength of the echo signals can be used as the basis for judging lip shape characteristics. Different reception times mean differences in the distance between the detected point and the ultrasonic cell, while different signal strengths mean different materials at the reflection location. For example, the lips, teeth, and tongue reflect ultrasonic waves differently. Thus, the user's lip shape characteristics are determined based on the echo signals, and the different lip shape characteristics of the user at consecutive moments are determined based on the echo signals returned at consecutive moments. Finally, the user's lip shape change characteristics are determined based on the different lip shape characteristics of the user at consecutive moments.

[0091] Furthermore, the smart glasses will also include a first sound detection module and a second sound detection module. The first sound detection module is located on the bridge of the glasses frame near the user's lips. Please refer to [reference needed]. Figure 3 The second sound detection module is located at one end of the temple of the glasses, away from the user's lips. Prior to the step of acquiring the user's voice signal, the method further includes:

[0092] Step S40: Collect first external audio information through the first sound detection module and collect second external audio information through the second sound detection module;

[0093] Step S50: Identify the ambient noise signal based on the second external audio information;

[0094] Step S60: Based on the identified environmental noise signal, perform noise reduction processing on the environmental noise signal in the first external audio information to obtain the user's voice signal.

[0095] For example, the step of performing noise reduction processing on the environmental noise signal in the first external audio information based on the identified environmental noise information includes:

[0096] Step F10: Determine the noise cancellation signal corresponding to the environmental noise information, wherein the noise cancellation signal has the same frequency and amplitude as the environmental noise information, but opposite phase.

[0097] Step F20: The noise-canceling wave signal is superimposed on the first external audio information to perform noise reduction processing on the environmental noise signal in the first external audio information.

[0098] In this embodiment, the first sound detection module can be a miniature microphone. It should be noted that, since the first sound detection module is located on the beam of the mirror frame near the user's lips, the pickup direction of the first sound detection module is directed towards the user's lips. It is easy to understand that the first external audio information collected by the first sound detection module includes both the user's voice signal and the ambient noise signal. The ambient noise signal can be the voices of other people and / or surrounding environmental noise, etc.

[0099] In this embodiment, the second sound detection module can also be a miniature microphone. It should be noted that, since the second sound detection module is located on the temple of the glasses at the end furthest from the user's lips, its pickup direction is further opposite to the user's lips. It is easy to understand that the second external audio information collected by the second sound detection module is an ambient noise signal, excluding the user's voice signal, or the user's voice signal is so small as to be negligible. This ambient noise signal can be the voices of other people and / or surrounding environmental noise, etc.

[0100] To help understand the installation position of the first sound detection module on the beam of the mirror frame near the user's lips, an example is provided. Please refer to [link / reference]. Figure 4 , Figure 4 This is a schematic diagram of the mechanical structure of smart glasses according to an embodiment of this application. The bridge of the glasses frame near the user's lips is... Figure 4 At positions 1 and 2, the bridge of the glasses, located furthest from the user's lips, is... Figure 4 Positions 3 and 4 are located in the image. In this embodiment, the first sound detection module can be located at... Figure 4 Positions 1 and / or 2 are shown in the diagram. It should be noted that the structure of the smart glasses illustrated in this example is only intended to aid in understanding this application and does not constitute a limitation on the structure of the smart glasses described in this application. Any further modifications or alterations based on the technical concept of this application should be within the scope of protection of this application.

[0101] This embodiment collects first external audio information through a first sound detection module and second external audio information through a second sound detection module; it identifies environmental noise signals based on the second external audio information; and it performs noise reduction processing on the environmental noise signals in the first external audio information based on the identified environmental noise signals to obtain the user's voice signal, thereby eliminating ambient noise in the first external audio information and improving the quality of the user's actual pronunciation signal in the first external audio information. This allows for more accurate and effective identification of the target semantic information that the user wants to express, avoiding the phenomenon of greatly reduced speech recognition rate caused by the mixing of user's voice with ambient noise, other people's voices, etc., thus improving the accuracy of speech recognition of smart glasses in complex environments.

[0102] Example 3

[0103] This invention also provides a voice recognition device based on smart glasses, the voice recognition device based on smart glasses comprising:

[0104] The audio semantic analysis module is used to acquire user voice signals and perform semantic recognition on the user voice signals to obtain audio semantic information.

[0105] The lip reading semantic analysis module is used to detect the user's lip shape change features, input the lip shape change features into the converged lip reading recognition model, and identify the user's lip reading semantic information;

[0106] The target semantic analysis module is used to determine the target semantic information based on the audio semantic information and the lip-reading semantic information.

[0107] Optionally, the target semantic analysis module is further configured to:

[0108] Detect whether there is an audio recognition failure time domain segment in the time domain of the audio semantic information;

[0109] If there is no audio recognition failure time segment, then the audio semantic information is used as the target semantic information;

[0110] If there is a time domain segment where audio recognition fails, the audio recognition failure time domain segment is repaired according to the lip reading semantic information to obtain repaired audio semantic information, and the repaired audio semantic information is used as the target semantic information.

[0111] Optionally, the target semantic analysis module is further configured to:

[0112] Detect whether there is a time domain segment in the lip-reading semantic information where lip-reading recognition failed;

[0113] If there is no time domain segment where lip reading recognition fails, then the following step is performed: repairing the audio recognition failure time domain segment based on the lip reading semantic information;

[0114] If there is a time domain segment where lip reading fails, then it is detected whether there is any overlap in the time domain between the audio recognition failure time domain segment and the lip reading failure time domain segment;

[0115] If there are no overlapping time domains, then the following step is performed: repairing the audio recognition failure time domain segment based on the lip-reading semantic information;

[0116] If there are overlapping time domains, semantic prediction is performed on the overlapping time domain segments where lip reading recognition failed, based on the lip reading semantic information and the audio semantic information, to obtain predicted semantic information; and target semantic information is determined based on the predicted semantic information, the lip reading semantic information, and the audio semantic information.

[0117] Optionally, the target semantic analysis module is further configured to:

[0118] The predicted semantic information is filled into the overlapping time domain segments where lip reading failed, thus completing the predicted semantic filling process.

[0119] After the predicted semantic filling process is completed, the audio recognition failure time domain segment is repaired according to the lip reading semantic information to obtain the repaired audio semantic information, and the repaired audio semantic information is used as the target semantic information.

[0120] Optionally, the voice recognition device based on smart glasses further includes a lip-reading training module, which is used for:

[0121] Obtain user lip shape samples corresponding to multiple lip shape change features, and the actual vocal information of the user corresponding to the user lip shape samples;

[0122] The actual vocal information of the user corresponding to the user lip-sync sample is used as the user vocal sample associated with the user lip-sync sample.

[0123] Multiple user lip-reading samples and user voice samples associated with each user voice sample are used as training samples to train the model, resulting in a converged lip-reading recognition model.

[0124] Optionally, the smart glasses include a glasses body and an ultrasonic detection module. The glasses body includes a frame and temples. The ultrasonic detection module is disposed on the frame, and the detection orientation of the ultrasonic detection module is directed towards the user's lips. The lip-reading semantic analysis module is further used for:

[0125] The ultrasonic detection module detects the user's lip shape changes.

[0126] Optionally, the smart glasses may further include a first sound detection module and a second sound detection module. The first sound detection module is disposed on the bridge of the glasses frame near the user's lips, and the second sound detection module is disposed on the temple of the glasses at the end away from the user's lips. The audio semantic analysis module is further used for:

[0127] The first external audio information is collected through the first sound detection module, and the second external audio information is collected through the second sound detection module.

[0128] The environmental noise signal is identified based on the second external audio information;

[0129] Based on the identified environmental noise signal, the environmental noise signal in the first external audio information is denoised to obtain the user's voice signal.

[0130] Optionally, the audio semantic analysis module is further configured to:

[0131] Determine the noise cancellation signal corresponding to the environmental noise information, wherein the noise cancellation signal has the same frequency and amplitude as the environmental noise information, but opposite phase.

[0132] The noise-canceling wave signal is superimposed on the first external audio information to perform noise reduction processing on the environmental noise signal in the first external audio information.

[0133] The speech recognition device based on smart glasses provided in this invention employs the speech recognition method based on smart glasses described in Embodiment 1 or Embodiment 2, thus solving the technical problem of low speech recognition accuracy in complex environments. Compared with the prior art, the beneficial effects of the speech recognition device based on smart glasses provided in this invention are the same as those of the speech recognition method based on smart glasses provided in the above embodiments, and other technical features in the speech recognition device based on smart glasses are the same as those disclosed in the previous embodiment method, and will not be repeated here.

[0134] Example 4

[0135] This invention provides a smart glasses system, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the speech recognition method based on smart glasses described in Embodiment 1 above.

[0136] The smart glasses of this application may be, for example, Mixed Reality (MR) glasses, Augmented Reality (AR) glasses, Virtual Reality (VR) glasses, Extended Reality (XR) glasses, or some combination thereof.

[0137] The following is for reference. Figure 5 It shows a schematic diagram of the hardware structure suitable for implementing smart glasses according to embodiments of the present disclosure. Figure 5 The smart glasses shown are merely an example and should not be construed as limiting the functionality and scope of use of the embodiments disclosed herein.

[0138] like Figure 5 As shown, the smart glasses may include a processing device 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM 1002) or a program loaded from a storage device into a random access memory (RAM 1004). The RAM 1004 also stores various programs and data required for the operation of the smart glasses. The processing device 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface is also connected to the bus 1005.

[0139] Typically, the following systems can be connected to I / O interface 1006: input devices 1007 including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. Communication device 1009 allows the smart glasses to communicate wirelessly or wiredly with other devices to exchange data. Although smart glasses with various systems are shown in the figures, it should be understood that it is not required to implement or possess all the systems shown. More or fewer systems can be implemented alternatively.

[0140] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of embodiments of this disclosure.

[0141] The smart glasses provided by this invention employ the speech recognition method based on smart glasses described in Embodiment 1 or Embodiment 2 above, solving the technical problem of low speech recognition accuracy in complex environments. Compared with the prior art, the beneficial effects of the smart glasses provided by this invention are the same as those of the speech recognition method based on smart glasses provided in Embodiment 1 above, and other technical features of the smart glasses are the same as those disclosed in the previous embodiment method, and will not be repeated here.

[0142] It should be understood that various parts of this disclosure can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics may be combined in any suitable manner in one or more embodiments or examples.

[0143] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

[0144] Example 5

[0145] This invention provides a storage medium, which is a computer-readable storage medium having computer-readable program instructions stored thereon, which are used to execute the voice recognition method based on smart glasses in the first embodiment described above.

[0146] The computer-readable storage medium provided in this embodiment of the invention may be, for example, a USB flash drive, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.

[0147] The aforementioned computer-readable storage medium may be included in the smart glasses; or it may exist independently and not assembled into the smart glasses.

[0148] The aforementioned computer-readable storage medium carries one or more programs. When the aforementioned one or more programs are executed by the smart glasses, the smart glasses cause the following: to acquire a user's voice signal, to perform semantic recognition on the user's voice signal to obtain audio semantic information; to detect the user's lip shape change features, to input the lip shape change features into a converged lip reading recognition model, and to identify the user's lip reading semantic information; and to determine target semantic information based on the audio semantic information and the lip reading semantic information.

[0149] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0150] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0151] The modules described in the embodiments of this disclosure can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.

[0152] The computer-readable storage medium provided by this invention stores computer-readable program instructions for executing the above-described speech recognition method based on smart glasses, thus solving the technical problem of low accuracy in speech recognition by smart glasses in complex environments. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this invention are the same as those of the speech recognition method based on smart glasses provided in Embodiment 1 or Embodiment 2, and will not be repeated here.

[0153] Example 6

[0154] This invention also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the speech recognition method based on smart glasses as described above.

[0155] The computer program product provided in this application solves the technical problem of low accuracy in speech recognition of smart glasses in complex environments. Compared with the prior art, the beneficial effects of the computer program product provided in the embodiments of this invention are the same as those of the speech recognition method based on smart glasses provided in Embodiment 1 or Embodiment 2 above, and will not be repeated here.

[0156] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent scope of this application.

Claims

1. A voice recognition method based on smart glasses, the smart glasses comprising a glasses body, the glasses body comprising a frame and temples, the smart glasses further comprising a first voice detection module and a second voice detection module, the first voice detection module being disposed on the bridge of the frame near the user's lips, and the second voice detection module being disposed on one end of the temples away from the user's lips, characterized in that, The method includes: The first external audio information is collected through the first sound detection module, and the second external audio information is collected through the second sound detection module. The environmental noise signal is identified based on the second external audio information; Based on the identified environmental noise signal, the environmental noise signal in the first external audio information is denoised to obtain the user's voice signal. The user's voice signal is semantically recognized to obtain audio semantic information; The user's lip shape change features are detected and input into a converged lip reading recognition model to identify the user's lip reading semantic information; Based on the audio semantic information and the lip-reading semantic information, the target semantic information is determined; The step of determining the target semantic information based on the audio semantic information and the lip-reading semantic information includes: Detect whether there is an audio recognition failure time domain segment in the time domain of the audio semantic information; If there is no audio recognition failure time segment, then the audio semantic information is used as the target semantic information; If there is a time domain segment where audio recognition fails, the audio recognition failure time domain segment is repaired according to the lip reading semantic information to obtain repaired audio semantic information, and the repaired audio semantic information is used as the target semantic information. Prior to the step of repairing the audio recognition failure time segment based on the lip-reading semantic information, the method further includes: Detect whether there is a time domain segment in the lip-reading semantic information where lip-reading recognition failed; If there is no time domain segment where lip reading recognition fails, then the following step is performed: repairing the audio recognition failure time domain segment based on the lip reading semantic information; If there is a time domain segment where lip reading fails, then it is detected whether there is any overlap in the time domain between the audio recognition failure time domain segment and the lip reading failure time domain segment; If there are no overlapping time domains, then the following step is performed: repairing the audio recognition failure time domain segment based on the lip-reading semantic information; If there are overlapping time domains, semantic prediction is performed on the lip-reading semantic information and the audio semantic information for the lip-reading recognition failure time domain segments in the overlapping time domains to obtain predicted semantic information; and target semantic information is determined based on the predicted semantic information, the lip-reading semantic information and the audio semantic information. The step of determining the target semantic information based on the predicted semantic information, the lip-reading semantic information, and the audio semantic information includes: The predicted semantic information is filled into the overlapping time domain segments where lip reading failed, thus completing the predicted semantic filling process. After the predicted semantic filling process is completed, the audio recognition failure time domain segment is repaired according to the lip reading semantic information to obtain the repaired audio semantic information, and the repaired audio semantic information is used as the target semantic information.

2. The speech recognition method based on smart glasses as described in claim 1, characterized in that, The method further includes: Obtain user lip shape samples corresponding to multiple lip shape change features, and the actual vocal information of the user corresponding to the user lip shape samples; The actual vocal information of the user corresponding to the user lip-sync sample is used as the user vocal sample associated with the user lip-sync sample. Multiple user lip-reading samples and user voice samples associated with each user voice sample are used as training samples to train the model and obtain a converged lip-reading recognition model.

3. The speech recognition method based on smart glasses as described in claim 1, characterized in that, The smart glasses include a glasses body and an ultrasonic detection module. The ultrasonic detection module is mounted on the frame, and the detection angle of the ultrasonic detection module is directed towards the user's lips. The step of detecting the user's lip shape changes includes: The ultrasonic detection module detects the user's lip shape changes.

4. The speech recognition method based on smart glasses as described in claim 1, characterized in that, The step of performing noise reduction processing on the environmental noise signal in the first external audio information based on the identified environmental noise information includes: Determine the noise cancellation signal corresponding to the environmental noise information, wherein the noise cancellation signal has the same frequency and amplitude as the environmental noise information, but opposite phase. The noise-canceling wave signal is superimposed on the first external audio information to perform noise reduction processing on the environmental noise signal in the first external audio information.

5. A type of smart glasses, characterized in that, The smart glasses include: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the steps of the speech recognition method based on smart glasses according to any one of claims 1 to 4.

6. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and the computer-readable storage medium stores a program for implementing a speech recognition method based on smart glasses. The program for implementing the speech recognition method based on smart glasses is executed by a processor to implement the steps of the speech recognition method based on smart glasses as described in any one of claims 1 to 4.