An AI intelligent voice interaction eyecup system method
By installing a voice acquisition device on the goggles for location and emotion analysis, and combining this with feedback questionnaires for optimization, the problem of inaccurate alerts from smart goggles in noisy environments has been solved, improving user experience and safety.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ANHUI SHUOHUI ENVIRONMENTAL PROTECTION TECHNOLOGY CO LTD
- Filing Date
- 2026-03-08
- Publication Date
- 2026-07-10
AI Technical Summary
Existing smart eye masks cannot provide targeted reminders in noisy environments and lack intelligent upgrades, resulting in a poor user experience.
A voice acquisition device is installed on the goggles to perform voice positioning and area division, recognize the user's voice emotions, provide personalized reminders, and upgrade the system through feedback questionnaires.
It improves user security and voice interaction experience, enhances the accuracy of user emotion analysis, and optimizes system functions based on user feedback.
Smart Images

Figure CN122369437A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a method for an AI-powered intelligent voice-interactive eye mask system, belonging to the field of eye mask voice interaction technology. Background Technology
[0002] There are many eye masks on the market, but existing eye masks are generally designed to improve sleep quality by blocking out light. However, with rising living standards, people seek to enhance their experience during short breaks. Current eye masks are not well-suited for these situations and can sometimes lead to oversleeping. Because traditional eye masks suffer from limited functionality and low technological sophistication, smart eye masks have been developed for voice interaction and reminders. These smart eye masks process voice interaction information in real-time. Using a real-time voice interaction processing method, device, electronic equipment, and storage medium, the system receives voice conversation information and determines whether preset sensitive content exists. If no sensitive content is found, a matching voice response is returned. If sensitive content is found, a matching voice reminder is returned, indicating the presence of sensitive content in the voice conversation. It can respond with voice reminders when there is sensitive content in the voice conversation, which can promptly remind users to stop the sensitive topic and avoid repeated issues, thus improving the user experience of voice interaction.
[0003] However, the voice interaction processing method proposed in the aforementioned comparative documents is not specific to the use of smart goggles, cannot adapt to the usage environment of goggles, cannot provide reminders in noisy environments, and cannot be upgraded to be intelligent.
[0004] In view of this, the present invention is proposed. Summary of the Invention
[0005] The purpose of this invention is to provide an AI-powered intelligent voice interaction goggle system method to solve the aforementioned problems. By installing a voice acquisition terminal on the goggle, it collects the user's voice and performs voice localization, ensuring that only the closest voice is collected. Distant voice information is defined as non-user-initiated speech, and ambient noise information is also recorded. Based on location, the system divides the area into danger zones, slightly dangerous zones, and safe zones, providing alerts when the user moves between these zones. This ensures basic voice interaction functionality while improving the effectiveness of alerts and reminders, thereby enhancing user safety. Furthermore, it incorporates voice emotion analysis and integrates data from various steps in the feature extraction process with language type analysis. The text content output in the class recognition step is analyzed, and the user's current emotion is determined simultaneously. By combining content and emotion to determine the user's current speech purpose, the obtained content will be more accurate. After use, all voice Q&A data generated that day is collected, and repetitive questions from the user are filtered out and defined as user-unsatisfied interaction items. A feedback questionnaire is sent to the user's mobile app, including the user-unsatisfied interaction items and providing reasons for dissatisfaction, such as "did not hear clearly" and "irrelevant answer". If the user selects the "irrelevant answer" option, the user is asked to fill in the specific meaning of their speech at that time. Based on the feedback, the voice interaction system is upgraded to help improve the system.
[0006] This invention achieves the above objectives through the following technical solution: an AI voice interaction interface for a smart eye mask. The treatment method includes the following steps: S1, Comprehensive Data Input: A voice acquisition terminal is installed on the goggles to collect the user's voice and record the surrounding noise information, which is used to remind the user when changing locations between different areas. S2, data preprocessing, which involves pre-emphasizing, framing, and windowing the voice data sent by the user; S3, Feature Extraction: Through the pre-emphasis, framing, and windowing steps in the data preprocessing steps, the collected user's entire speech is segmented into speech segments, and various features of the speech are extracted. The decoder then uses acoustics, language models, and dictionaries to find the word string that can output the signal with the highest probability from the input signal. S4, Language Recognition, identifies the language type of the speech data and outputs the speech as the corresponding text after recognition; S5, Voice Emotion Analysis, integrates the data from the feature extraction step and the text output from the language type recognition step, analyzes the user's voice content while determining their current emotion, and combines content with emotion to determine the purpose of the user's current voice. S6, Intelligent Response, provides relevant intelligent responses based on the analysis results in the voice emotion analysis step; S7, User Feedback: After use, all voice Q&A data generated that day is collected, allowing customers to provide feedback on unsatisfactory interactions. Based on the feedback, the voice interaction system is upgraded.
[0007] Furthermore, in the comprehensive data entry step, the specific operation method is as follows: a voice acquisition terminal is installed on the goggles to collect the voice emitted by the user, and the voice is located to ensure that only the closest voice is collected. The voice information at a distance is defined as non-user-initiated voice, and the noisy information is recorded. The area is divided into danger zone, mild danger zone and safe zone according to the location, and a reminder is issued when the user changes location between different zones.
[0008] Furthermore, in the data preprocessing step, the pre-emphasis method is to keep the low-frequency part of the signal unchanged, boost the high-frequency part of the signal, and deemphasize the low-frequency part of the signal while keeping the high-frequency part unchanged. The purpose of pre-emphasis / deemphasis is to boost the energy of the high-frequency part of the signal to compensate for the excessive attenuation of the high-frequency part by the channel. Before analyzing the speech signal s(n), the invalid part is filtered out by the filter and the high-frequency part is boosted.
[0009] Furthermore, in the data preprocessing step, short-time analysis adopts a frame-segmentation method. Genes may change between adjacent frames. Overlapping frame extraction is used to segment the speech signal and ensure a certain repetition rate. The speech signal is a non-stationary signal. When producing voiced sounds, the vocal cords vibrate regularly, that is, the fundamental frequency is relatively fixed in a short time range. The speech signal has short-time stationary characteristics.
[0010] Furthermore, in the feature extraction step, the specific method is as follows: through the pre-emphasis, framing, and windowing processing steps in the data preprocessing step, the collected user's entire speech is segmented into speech segments. The speech intensity, intensity level, loudness, pitch, fundamental period, fundamental frequency, harmonic-to-noise ratio, frequency perturbation, amplitude perturbation, and normalized noise energy data are removed. The short-time energy, short-time average amplitude, formants, glottal waves, speech rate, and pauses of each segment are calculated. The decoder uses the input signal to find the word string that can output the signal with the highest probability based on acoustics, language models, and dictionaries.
[0011] Furthermore, in the language type recognition step, the specific method is as follows: the speech type of the speech data is identified, including Chinese, English, and Japanese. After determining the language of the user's speech, the actual content is output, and the speech is output as the corresponding text.
[0012] Furthermore, in the user feedback step, the specific method is as follows: After use, all voice Q&A data generated that day is collected, and repetitive user questions are filtered out and defined as user-unsatisfied interaction items. A feedback questionnaire is then sent to the user's mobile app, including the user-unsatisfied interaction items and providing the reasons for dissatisfaction. The options are "did not hear clearly" and "irrelevant answer". If the user selects the "irrelevant answer" option, the user is asked to fill in the specific meaning of the voice at that time. Based on the feedback, the voice interaction system is upgraded.
[0013] The technical effects and advantages of this invention are as follows: By installing a voice acquisition terminal on the goggles, this invention collects the user's voice and locates the voice, ensuring that only the closest voice is collected. Distant voice information is defined as non-user-initiated voice, and ambient noise information is also recorded. Based on the location, the system is divided into danger zones, slightly dangerous zones, and safe zones, which provide reminders when the user changes locations between different zones. This ensures basic voice interaction functionality while improving the effectiveness of warnings and reminders to the user, thereby enhancing user safety.
[0014] Compared to traditional voice interaction processing methods, this invention adds user voice emotion analysis, integrates various data from the feature extraction step and the text content output from the language type recognition step, analyzes the user's voice content while determining their current emotion, and combines content with emotion to determine the purpose of the user's current voice, resulting in more accurate content.
[0015] This invention also includes a feedback function. After use, all voice Q&A data generated that day is collected, and repetitive user questions are filtered out and defined as unsatisfactory interaction items. A feedback questionnaire is then sent to the user's mobile app, containing the unsatisfactory interaction items and providing reasons for dissatisfaction, such as "did not hear clearly" and "answered irrelevantly." If the user selects the "answered irrelevantly" option, they are asked to specify the exact meaning of their voice at the time. Based on the feedback, the voice interaction system is upgraded, contributing to its improvement. Attached Figure Description
[0016] Figure 1 This is a flowchart of the present invention. Detailed Implementation
[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0018] Please see Figure 1 As shown, a method for an AI-powered intelligent voice-interactive eye mask system.
[0019] The system integrates data entry and a voice acquisition device installed on the goggles to collect user voice messages. It also locates the voice messages, ensuring only the closest voice messages are captured. Distant voice messages are defined as non-user-initiated. Additionally, it records ambient noise and divides the area into danger zones, slightly dangerous zones, and safe zones based on location. This provides alerts when the user moves between these zones, ensuring basic voice interaction while enhancing the effectiveness of warnings and alerts, thereby improving user safety.
[0020] Data preprocessing involves pre-emphasis, framing, and windowing of the user-generated voice data. Pre-emphasis maintains the low-frequency components of the signal while boosting the high-frequency components, and de-emphasizes the low-frequency components while preserving the high-frequency components. Both pre-emphasis and de-emphasis aim to boost the energy of the high-frequency components to compensate for excessive attenuation of the high-frequency components by the channel. Before analyzing the voice signal s(n), invalid components are filtered out, and the high-frequency components are boosted. Voice signals are non-stationary; during voiced speech, the vocal cords vibrate regularly, meaning the fundamental frequency is relatively fixed within a short timeframe. Voice signals exhibit short-time stationarity. Short-time analysis employs framing, where the fundamental frequency may change between adjacent frames. Overlapping frames are used to segment the voice signal while maintaining a certain repetition rate.
[0021] Feature extraction, through pre-emphasis, framing, and windowing steps in the data preprocessing process, extracts the features used for... The user's entire speech was segmented into speech fragments. Data on sound intensity, sound intensity level, loudness, pitch, fundamental period, fundamental frequency, harmonic-to-noise ratio, frequency perturbation, amplitude perturbation, and normalized noise energy were removed. Short-time energy, short-time average amplitude, formants, glottal waves, speech rate, and pauses of each fragment were calculated. The decoder then analyzed the input signal and, based on acoustics, language models, and a dictionary, searched for word strings that could output the signal with the highest probability.
[0022] Language recognition identifies the language of the speech data, including Chinese, English, and Japanese. After determining the language of the user's speech, the actual content is output, converting the speech into the corresponding text.
[0023] Voice emotion analysis integrates data from the feature extraction step and text output from the language recognition step. It analyzes the user's voice content while simultaneously determining their current emotion. By combining content and emotion, it determines the user's current voice purpose. Compared to traditional voice interaction processing methods, this method adds voice emotion analysis, analyzes the user's voice content while simultaneously determining their current emotion, and combines content and emotion to determine the user's current voice purpose, resulting in more accurate information.
[0024] Intelligent responses provide relevant intelligent replies based on the analysis results from the voice emotion analysis step.
[0025] User feedback is collected after each use. All voice Q&A data generated that day is aggregated, and repetitive questions are filtered out and defined as unsatisfactory interactions. A feedback questionnaire is then sent to the user's mobile app, including the unsatisfactory interactions and providing reasons for dissatisfaction, such as "did not hear clearly" or "answered irrelevantly." If the user selects the "answered irrelevantly" option, they are asked to specify the exact meaning of their voice at the time. Based on the feedback, the voice interaction system is upgraded.
[0026] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered in all respects as exemplary and non-limiting, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the present invention. No reference numerals in the claims should be construed as limiting the scope of the claims.
[0027] Furthermore, it should be understood that although this specification describes embodiments, not every embodiment contains only one independent technical solution. This narrative style is merely for clarity. Those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. A method for an AI-powered intelligent voice-interactive eye mask system, characterized in that, Includes the following steps: S1, Comprehensive Data Input: A voice acquisition terminal is installed on the goggles to collect the user's voice and record specific information about the user's location, which is used to issue reminders when the user changes locations between different areas. S2, data preprocessing, which involves pre-emphasizing, framing, and windowing the voice data sent by the user; S3, Feature Extraction: Through the pre-emphasis, framing, and windowing steps in the data preprocessing steps, the collected user's entire speech is segmented into speech segments, and various features of the speech are extracted. The decoder then uses acoustics, language models, and dictionaries to find the word string that can output the signal with the highest probability from the input signal. S4, Language Recognition, identifies the language type of the speech data and outputs the speech as the corresponding text after recognition; S5, Voice Emotion Analysis, integrates the data from the feature extraction step and the text output from the language type recognition step, analyzes the user's voice content while determining their current emotion, and combines content with emotion to determine the purpose of the user's current voice. S6, Intelligent Response, provides relevant intelligent responses based on the analysis results in the voice emotion analysis step; S7, User Feedback: After use, all voice Q&A data generated that day is collected. The specific method is as follows: Repeated questions from users are filtered out and defined as unsatisfactory interaction items. A feedback questionnaire is sent to the user's mobile app, including the unsatisfactory interaction items and providing reasons for dissatisfaction, such as "did not hear clearly" and "answered irrelevantly." If the user selects the "answered irrelevantly" option, the user is asked to fill in the specific meaning of their voice at that time. Based on the feedback, the voice interaction system is upgraded.
2. The method for an AI-powered intelligent voice-interactive eye mask system according to claim 1, characterized in that, In the comprehensive data entry step, the specific operation method is as follows: a voice acquisition terminal is installed on the goggles to collect the voice emitted by the user. At the same time, the voice is located to ensure that only the closest voice is collected. Voice information from a distance is defined as non-user-initiated voice. At the same time, specific information about the location is recorded. The project location is divided into dangerous areas, slightly dangerous areas, and safe areas to provide reminders when the user changes locations between different areas.
3. The method for an AI-powered intelligent voice-interactive eye mask system according to claim 1, characterized in that, In the data preprocessing step, the pre-emphasis method is to keep the low-frequency part of the signal unchanged, boost the high-frequency part of the signal, and de-emphasize the low-frequency part of the signal while keeping the high-frequency part unchanged. The purpose of pre-emphasis / de-emphasis is to boost the energy of the high-frequency part of the signal to compensate for the excessive attenuation of the high-frequency part by the channel. Before analyzing the speech signal s(n), the invalid part is filtered out by the filter and the high-frequency part is boosted.
4. A method for an AI-powered intelligent voice-interactive eye mask system according to claim 3, characterized in that, In the data preprocessing step, short-time analysis adopts a frame-segmentation method. Genes may change between adjacent frames. Overlapping frame extraction is used to segment the speech signal while ensuring a certain repetition rate. The speech signal is a non-stationary signal. When producing voiced sounds, the vocal cords vibrate regularly, meaning that the fundamental frequency is relatively fixed within a short time range. The speech signal has short-time stationary characteristics.
5. A method for an AI-powered intelligent voice-interactive eye mask system according to claim 1, characterized in that, In the feature extraction step, the specific method is as follows: through the pre-emphasis, framing and windowing processing steps in the data preprocessing step, the collected user's whole speech is segmented into speech segments. The speech intensity, intensity level, loudness, pitch, fundamental period, fundamental frequency, harmonic-to-noise ratio, frequency perturbation, amplitude perturbation and normalized noise energy data are removed. The short-time energy, short-time average amplitude, formants, glottal waves, speech rate and pauses of each segment are calculated. The decoder uses the input signal to find the word string that can output the signal with the highest probability based on acoustics, language model and dictionary.
6. A method for an AI-powered intelligent voice-interactive eye mask system according to claim 1, characterized in that, In the language type recognition step, the specific method is as follows: the speech type of the speech data is identified, including Chinese, English, and Japanese. After determining the language of the user's speech, the actual content is output, and the speech is output as the corresponding text.