AI intelligent glasses and interaction method thereof

By integrating voice, vision, and head vibration data to generate a unified contextual understanding vector, and combining it with a large AI model and emotion classification, personalized and humanized interaction of AI smart glasses is achieved, solving the problem of inaccurate intent recognition in existing technologies and improving the user experience.

CN120929939APending Publication Date: 2025-11-11SHENZHEN HUAYUE YUNPENG TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511052626.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-30
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing AI smart glasses are not accurate enough in initially judging user intentions and recognizing emotions, resulting in an unsatisfactory human-computer interaction experience and failing to fully leverage the advantages of intelligence.

Method used

By receiving the user's original instructions, calling the AI ​​big model to calculate the intention confidence distribution, integrating voice stream, visual stream and head vibration stream data to generate a unified contextual understanding vector, combining the emotion classification model to output the user's emotion label, and generating personalized feedback according to the interaction pattern mapping rules.

Benefits of technology

It improves the accuracy of intent recognition and enhances the user-friendly interactive experience, enabling smart glasses to better meet user needs and situations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120929939A_ABST
    Figure CN120929939A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of AI intelligent glasses, and particularly discloses AI intelligent glasses and an interaction method thereof, and the AI intelligent glasses are characterized in that an intention initial recognition module receives a user original instruction and calls an AI large model to calculate user intention confidence degree distribution; the multi-modal perception fusion module fuses voice stream data, visual stream data and head vibration stream data which are perceived in real time to generate a unified situation understanding vector; the intention comprehensive identification module analyzes the real intention of the user based on the unified situation understanding vector and user intention confidence distribution; the user emotion recognition module inputs the head micro-vibration feature vector and a perceived emotion type sequence and a perceived emotion degree sequence which are perceived based on voice stream data into an emotion classification model to output a user emotion tag; an interaction feedback generation module generates an interaction feedback result aiming at the real intention of the user according to an interaction mode mapping rule under the user emotion label; and the interaction between the intelligent glasses and the user better meets the current demand of the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of AI smart glasses technology, and in particular to an AI smart glasses and its interaction method. Background Technology

[0002] In today's booming era of smart wearable devices, AI smart glasses, as an emerging wearable device, are gradually changing the way people interact with technology. With the continuous advancement of artificial intelligence technology, smart glasses are no longer limited to simple information display but are developing towards more intelligent and human-centered interaction. AI smart glasses have demonstrated enormous application potential in numerous fields. In daily life, they can act like personal assistants, helping users quickly obtain information and complete various tasks, such as checking routes and schedules, making life more convenient and efficient. In the industrial sector, workers wearing AI smart glasses can obtain real-time work instructions and operational procedures, improving work efficiency and accuracy and reducing the risks of errors. In medical settings, doctors can use smart glasses to access patient data and obtain remote expert guidance during surgery, providing strong support for surgical success. With the popularization of 5G technology and the continuous optimization of artificial intelligence algorithms, the functions of AI smart glasses will be further expanded and strengthened. In the future, they are expected to become an indispensable intelligent companion in people's lives and work, achieving more natural and seamless human-computer interaction and driving various industries towards intelligentization.

[0003] However, existing AI smart glasses technology has many shortcomings. Its initial judgment of user intent is inaccurate, and its understanding of the user's context is insufficient, making it difficult to analyze the user's true intentions. Furthermore, because it cannot accurately identify user emotion tags, it cannot generate appropriate interactive feedback results based on corresponding interaction pattern mapping rules, resulting in a less than ideal human-computer interaction experience and failing to fully leverage the intelligent advantages of AI smart glasses.

[0004] Therefore, this invention proposes an AI smart glasses and its interaction method. Summary of the Invention

[0005] This invention provides AI smart glasses and their interaction method. It receives original user commands and calls a large AI model to calculate the user's intent confidence distribution. Leveraging powerful AI capabilities, it initially judges the user's intent tendency, laying the foundation for subsequent accurate recognition. It fuses voice stream, visual stream, and head vibration stream data to generate a unified contextual understanding vector, comprehensively capturing the user's situation by integrating multiple perceptual information, enabling the smart glasses to understand the user more deeply. Based on the unified contextual understanding vector and the user intent confidence distribution, it analyzes the user's true intent, integrating multiple aspects of information to improve the accuracy of intent recognition, allowing the smart glasses to accurately grasp user needs. By inputting head micro-vibration feature vectors and emotional sequences from voice perception into an emotion classification model to output user emotion labels, it observes the user's emotional state, providing a key basis for personalized interaction. Based on the interaction pattern mapping rules under the user's emotion labels, it generates interactive feedback results targeting the true intent, achieving more personalized and humanized interaction, improving the user experience, and making the interaction between the smart glasses and the user more closely aligned with the user's current state and needs.

[0006] This invention provides AI smart glasses, comprising:

[0007] The initial intent recognition module is used to receive the user's original instructions and call the AI ​​large model to calculate the user's intent confidence distribution;

[0008] The multimodal perception fusion module is used to fuse real-time perceived speech stream data, visual stream data, and head vibration stream data to generate a unified contextual understanding vector;

[0009] The intent integration and recognition module is used to analyze the user's true intent based on a unified contextual understanding vector and user intent confidence distribution.

[0010] The user emotion recognition module is used to input the head micro-vibration feature vector, the perceived emotion type sequence based on speech stream data, and the perceived emotion degree sequence into the emotion classification model, and output the user emotion label.

[0011] The interaction feedback generation module is used to generate interaction feedback results that target the user's true intentions based on the interaction pattern mapping rules under the user's emotion tags.

[0012] Preferably, the initial intent recognition module includes:

[0013] The multi-form receiving submodule is used to receive user input commands in various input formats.

[0014] The intent recognition submodule is used to calculate the user intent confidence distribution based on all the latest received user input commands by calling the AI ​​big model.

[0015] Preferred options also include:

[0016] The priority arbitration submodule is used to determine the priority of all user input commands that conflict when there are conflicting original user commands in multiple input formats, based on the priority arbitration principle, and to call the AI ​​big model to calculate the user intent confidence distribution based on all the latest received user input commands and their corresponding priorities.

[0017] Preferably, the multimodal perception fusion module includes:

[0018] The speech perception and analysis submodule is used to extract text and perceive emotions from real-time perceived speech stream data to obtain intention-perceived text, perceived emotion type sequence and corresponding perceived emotion degree sequence.

[0019] The visual perception analysis submodule is used to identify objects of attention based on real-time perceived visual stream data, and to obtain all objects of attention and the attention-focusing sliding path;

[0020] The head vibration sensing and analysis submodule is used to identify the direction and amplitude of head tremors based on real-time sensed head vibration flow data.

[0021] The multimodal fusion generation submodule is used to generate a unified contextual understanding vector based on intent-aware text, a sequence of perceived emotion types, a corresponding sequence of perceived emotion levels, all attention-focused objects, attention-focused sliding paths, head tremor direction, and head tremor amplitude.

[0022] Preferably, the multimodal fusion generation submodule includes:

[0023] The first intention evolution analysis unit is used to determine the first intention evolution sequence based on the perceived emotion type, perceived emotion degree and emotion-intention mapping rule;

[0024] The second intention evolution analysis unit is used to determine the second intention evolution sequence based on the direction and amplitude of head tremor and the action-intention mapping rule.

[0025] The comprehensive intention evolution analysis unit is used to match and analyze the first intention evolution sequence and the second intention evolution sequence with the attention focusing sliding path to obtain the comprehensive intention evolution sequence of the attention focusing sliding path;

[0026] The intention bias analysis unit is used to determine the intention shift and degree of intention shift for all attention-focused objects based on the comprehensive intention evolution sequence of the attention-focusing sliding path.

[0027] The Unified Contextual Understanding Unit is used to generate a unified contextual understanding vector based on the intent-aware text and the intent shift and degree of intent shift of all attention-focused objects.

[0028] Preferably, the integrated intention evolution analysis unit includes:

[0029] The path segmentation sub-unit is used to divide the attention focusing sliding path based on the gaze region type to obtain multiple attention focusing sub-sliding paths;

[0030] The weight allocation sub-unit is used to determine the path allocation weight of each attention focus sub-sliding path based on the gaze region type of each attention focus sub-sliding path;

[0031] The weighted operation subunit is used to align the first intention evolution sequence and the second intention evolution sequence with the attention focusing sliding path using timestamps, and to perform alignment weighting operation on the first intention evolution sequence and the second intention evolution sequence based on the path allocation weights of all attention focusing sub-sliding paths to obtain the comprehensive intention evolution sequence of the attention focusing sliding path.

[0032] Preferably, the intention-biased analysis unit includes:

[0033] The intent-direction determination subunit is used to determine the intent direction of all attention-focused objects based on the comprehensive intent evolution sequence of the attention-focusing sliding path;

[0034] The same-direction and opposite-direction path division sub-unit is used to filter out attention-focused objects that have undergone obvious intention shifts from all attention-focused objects based on the comprehensive intention evolution sequence of attention-focused sliding paths as dominant intention shift objects, and to determine the current same-direction intention accumulation path and future opposite-direction accumulation path of each dominant intention shift object based on the comprehensive intention evolution sequence of attention-focused sliding paths.

[0035] The cumulative bias determination subunit is used to determine the current same-direction cumulative bias and future opposite-direction cumulative bias of each dominant intention turning object based on the comprehensive intention evolution sequence;

[0036] The turning degree calculation subunit is used to calculate the intention turning degree of each attention focus object based on the current same-direction cumulative bias, future opposite-direction cumulative bias, current same-direction cumulative path, and future opposite-direction cumulative path of each dominant intention turning object.

[0037] Preferably, the steering degree calculation subunit includes:

[0038] The first steering degree determination end is used to calculate the comprehensive steering degree of the corresponding current homing intention cumulative path and the corresponding future dissimilar cumulative path based on the current homing intention cumulative bias degree and the future dissimilar cumulative bias degree of each dominant intention steering object;

[0039] The second turning degree determination end is used to calculate the first turning degree of all attention-focused objects in the current same-direction intention accumulation path and the second turning degree of all attention-focused objects in the corresponding future opposite-direction accumulation path of each dominant intention turning object based on the comprehensive intention evolution sequence and the comprehensive turning degree of the current same-direction intention accumulation path and the corresponding future opposite-direction accumulation path of each dominant intention turning object.

[0040] The third turning degree determination end is used to calculate the intention turning degree of each attention-focused object based on all turning degrees of each attention-focused object.

[0041] Preferably, the interactive feedback generation module includes:

[0042] The mapping rule retrieval submodule is used to retrieve the interaction pattern mapping rules under the user's emotion tag from the interaction pattern library;

[0043] The feedback pattern retrieval submodule is used to retrieve the interaction feedback pattern under the user's true intention from the interaction pattern mapping rules under the user's emotion tag.

[0044] The feedback result generation submodule is used to generate interactive feedback results based on the user's true intent, the interactive feedback pattern, and the user's true intent.

[0045] This invention provides an interaction method for AI smart glasses, comprising:

[0046] Receive the user's original instructions and call the AI ​​large model to calculate the user's intent confidence distribution;

[0047] Real-time perceived speech stream data, visual stream data, and head vibration stream data are fused to generate a unified contextual understanding vector;

[0048] The user's true intent is analyzed based on the unified contextual understanding vector and the user intent confidence distribution.

[0049] The head micro-vibration feature vector, the perceived emotion type sequence based on speech stream data, and the perceived emotion degree sequence are input into the emotion classification model to output the user's emotion label.

[0050] Based on the interaction pattern mapping rules under the user's emotion tag, generate interactive feedback results that target the user's true intentions.

[0051] The beneficial effects of this invention compared to existing technologies are as follows: It receives the user's original commands and calls upon a large AI model to calculate the user's intent confidence distribution. Leveraging powerful AI capabilities, it initially judges the user's intent tendency, laying the foundation for subsequent accurate recognition. It fuses voice stream, visual stream, and head vibration stream data to generate a unified contextual understanding vector, comprehensively capturing the user's situation by integrating multiple perceptual information, enabling smart glasses to understand the user more deeply. Based on the unified contextual understanding vector and the user intent confidence distribution, it analyzes the user's true intent, integrating multiple aspects of information to improve the accuracy of intent recognition, allowing smart glasses to accurately grasp user needs. By inputting the head micro-vibration feature vector and the emotional sequence perceived by voice into an emotion classification model to output user emotion labels, it gains insight into the user's emotional state, providing a key basis for personalized interaction. Based on the interaction pattern mapping rules under the user's emotion labels, it generates interactive feedback results targeting the true intent, achieving more personalized and humanized interaction, improving the user experience, and making the interaction between smart glasses and users more closely aligned with the user's current state and needs.

[0052] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures particularly pointed out in this application.

[0053] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0054] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:

[0055] Figure 1 This is a schematic diagram of the built-in module of the AI ​​smart glasses in an embodiment of the present invention;

[0056] Figure 2 This is a schematic diagram of the initial intent recognition module in an embodiment of the present invention;

[0057] Figure 3 This is a schematic diagram of the multimodal perception fusion module in an embodiment of the present invention. Detailed Implementation

[0058] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.

[0059] like Figure 1 As shown, the present invention provides an implementation method for AI smart glasses, including:

[0060] The initial intent recognition module is used to receive the user's original instructions and call the AI ​​large model to calculate the user's intent confidence distribution;

[0061] The multimodal perception fusion module is used to fuse real-time perceived speech stream data, visual stream data, and head vibration stream data to generate a unified contextual understanding vector;

[0062] The intent integration and recognition module is used to analyze the user's true intent based on a unified contextual understanding vector and user intent confidence distribution.

[0063] The user emotion recognition module is used to input the head micro-vibration feature vector, the perceived emotion type sequence based on speech stream data, and the perceived emotion degree sequence into the emotion classification model, and output the user emotion label.

[0064] The interaction feedback generation module is used to generate interaction feedback results that target the user's true intentions based on the interaction pattern mapping rules under the user's emotion tags.

[0065] In this embodiment, the user's original instruction refers to the unprocessed instruction information input by the user to the AI ​​smart glasses, which can take various forms such as voice instructions, gesture instructions, and touch instructions. For example, if a user says to the AI ​​smart glasses, "Check tomorrow's weather for me," or expresses a similar need through specific gestures, these instructions are considered the user's original instructions when received, serving as the starting information for the AI ​​smart glasses to understand the user's intent.

[0066] In this embodiment, the AI ​​large-scale model is an artificial intelligence model with powerful data processing and analysis capabilities. Trained on a large amount of data, it is capable of understanding and analyzing various natural languages ​​or other input information. In this invention, the AI ​​large-scale model is invoked to calculate the user intent confidence distribution based on the user's original commands. For example, the common GPT series models, through learning from massive amounts of text data, possess the ability to understand complex language expressions and infer intent. The AI ​​large-scale model in this invention is similar, used for preliminary analysis of the intent tendencies behind the user's original commands. This model is built through deep learning on a large number of user original commands, or their priorities (model input) and the user intent confidence distribution (model output).

[0067] In this embodiment, the user intent confidence distribution refers to the distribution of confidence levels among various possible user intents calculated by the AI ​​model based on the user's original command. It represents the probability of a user intent being a specific content. For example, when a user says "I want to watch a movie," the AI ​​model, after analysis, might conclude that the user's intent to search for nearby cinemas has a confidence level of 0.6, the intent to play a local movie file has a confidence level of 0.3, and the sum of the confidence levels for other intents is 0.1. This series of confidence values ​​constitutes the user intent confidence distribution, helping to initially determine the user's intent tendency.

[0068] In this embodiment, speech stream data, visual stream data, and head vibration stream data are real-time data acquired by the AI ​​smart glasses from different perceptual modalities. Speech stream data is continuous audio data generated by the user's speech, containing language content and emotional information, such as tone of voice and speech rate reflecting emotions. Visual stream data is image sequence data captured in real-time by the smart glasses' visual sensors (such as cameras), used to identify information such as the user's surrounding environment and objects of interest. Head vibration stream data is continuous data related to the user's head vibrations collected by sensors; information such as the direction and amplitude of head tremors may contain the user's underlying intentions or emotional state. For example, a user may have a larger head shaking amplitude when excited.

[0069] In this embodiment, the unified contextual understanding vector is a comprehensive vector generated by processing speech stream data, visual stream data, and head vibration stream data through a multimodal perception fusion module. It integrates key information extracted from different data modalities, comprehensively reflecting the user's context and intent cues, providing a comprehensive basis for accurately analyzing the user's true intent. For example, this vector may integrate information such as the intent text extracted from speech, the object of attention identified in vision, and the potential intent reflected by head vibrations, represented in a unified numerical vector form, facilitating further analysis by subsequent modules.

[0070] In this embodiment, the user's true intent refers to the intent that the user truly wants to express and achieve, derived from the analysis of the intent comprehensive recognition module after comprehensively considering the user's intent confidence distribution and unified contextual understanding vector. Compared to the intent initially judged based solely on the original command, the user's true intent, after multi-faceted information fusion analysis, is more accurate and closely matches the user's actual needs. For example, combining the user's voice command "search for that" (the user's original command) and the smart glasses' visual recognition of the image the user is looking at on the computer screen (visual stream data), the comprehensive analysis concludes that the user's true intent is to search for the image being looked at on the computer, rather than other possible ambiguous understandings.

[0071] In this embodiment, the head micro-vibration feature vector is a vector data extracted from head vibration stream data that can characterize the features of head micro-vibrations. It contains information such as the direction and amplitude of head tremors, which are extracted and encoded into a numerical vector form. It is used to reflect the potential intentions or emotional cues contained in the user's head micro-vibrations, providing key input information for the emotion classification model.

[0072] In this embodiment, the perceived emotion type sequence and perceived emotion intensity sequence based on speech stream data are obtained by performing emotion perception on the speech stream data. The perceived emotion type sequence records the emotion types identified during the speech process, such as happiness, sadness, and anger, arranged in chronological order. The perceived emotion intensity sequence corresponds to each emotion type, indicating its intensity in the speech expression. For example, in a speech segment, if the initially identified emotion type is "happiness" with an intensity of "moderate," and then the emotion type changes to "excitement" with an intensity of "high," two corresponding sequences are formed to assist in judging the user's emotional state.

[0073] In this embodiment, the emotion classification model is a trained model used to output a user's emotion label based on the input head micro-vibration feature vector and the perceived emotion type sequence and perceived emotion intensity sequence based on speech stream data. It learns from a large number of data samples with emotion labels (such as those containing head micro-vibration feature vectors and perceived emotion type and intensity sequences based on speech stream data) to master the mapping relationship between different features and emotions, thereby enabling accurate emotion classification of new input data. For example, after training, the model can identify when head micro-vibration exhibits a certain pattern and the speech emotion is "angry" and at a high level, outputting the user's emotion label as "angry and intensely emotional."

[0074] In this embodiment, the user emotion label is the output of the emotion classification model, used to identify the user's current emotional state. It summarizes the user's emotion in a concise label format, such as "happy," "frustrated," or "calm," providing a basis for generating personalized interactive feedback. Different emotion labels correspond to different interaction pattern mapping rules, enabling the AI ​​smart glasses to provide more appropriate feedback based on the user's emotions.

[0075] In this embodiment, the interaction pattern mapping rules under user emotion tags are pre-defined rule sets stored in an interaction pattern library. These rules specify which interaction feedback mode should be used for different user emotion tags in response to the user's true intent. For example, when the user's emotion tag is "happy," for a user's true intent to query information, the interaction feedback mode might be to display the query results with cheerful voice and rich facial expressions; while when the emotion tag is "frustrated," the interaction feedback mode might be more inclined towards comforting language and providing a concise solution.

[0076] In this embodiment, the interactive feedback result based on the user's true intent is the final response content presented to the user by the AI ​​smart glasses. It is generated based on the interaction pattern mapping rules under the user's emotion tag, combined with the user's true intent. For example, if the user's true intent is to search for nearby restaurants, and the emotion tag is "urgent," according to the corresponding rules, the interactive feedback result may be to quickly and concisely display information about the nearest restaurants and announce it to the user in a rapid voice to meet the user's current state and needs.

[0077] like Figure 2 As shown, in order to receive user input commands in multiple forms and to calculate the user intent confidence distribution based on these commands using a large AI model, an initial intent recognition module is further proposed, including:

[0078] The multi-form receiving submodule is used to receive user input commands in various input formats.

[0079] The intent recognition submodule is used to calculate the user intent confidence distribution based on all the latest received user input commands by calling the AI ​​big model.

[0080] In this embodiment, multiple input methods allow users to convey instructions to the AI ​​smart glasses in various ways. This reflects the multimodal interaction features of the AI ​​smart glasses, aiming to meet diverse user interaction needs. For example, users can input commands via voice, such as saying "Book me a flight to Beijing tomorrow"; they can also use gestures, such as specific hand movements, to express the intention to open a certain application; or they can use touch operations, such as swiping or clicking on specific touch areas of the smart glasses to input commands, like switching displayed content by swiping. These multiple input methods increase the flexibility and convenience of user interaction with the smart glasses, adapting to different scenarios and user habits.

[0081] In this embodiment, the AI ​​big data model is invoked to calculate the user intent confidence distribution based on all the latest received user input commands. This means that the AI ​​smart glasses integrate the latest input command information from various forms, and then, leveraging the powerful data analysis and intent inference capabilities of the AI ​​big data model, calculate the confidence distribution of various possible user intents. For example, a user might first say "Find me a place" via voice, and then point to a nearby shopping mall with a gesture. The AI ​​smart glasses will then combine these two commands and invoke the AI ​​big data model for analysis. The AI ​​big data model will consider the correlation between the voice content and the gesture, comprehensively judging the user's possible intents, such as wanting to find shops in the mall or learn about the mall's opening hours, and calculating the confidence of each of these different intents to form a user intent confidence distribution. This approach allows for a more comprehensive and accurate preliminary judgment of user intent tendencies, providing a foundation for more precise identification of the user's true intent in the future.

[0082] To determine priority through a priority arbitration principle when user original commands in multiple input formats conflict, and thus accurately calculate the user intent confidence distribution, the following further steps are proposed:

[0083] The priority arbitration submodule is used to determine the priority of all user input commands that conflict when there are conflicting original user commands in multiple input formats, based on the priority arbitration principle, and to call the AI ​​big model to calculate the user intent confidence distribution based on all the latest received user input commands and their corresponding priorities.

[0084] In this embodiment, conflicting user commands arise from multiple input methods. This means that when a user inputs commands to the AI ​​smart glasses through various means (such as voice, gestures, and touch), the intentions expressed by these commands may contradict each other or cannot be executed simultaneously. For example, if a user says "Close all applications" via voice while simultaneously attempting to open a new application via gesture, these two commands conflict, making it impossible for the AI ​​smart glasses to directly determine the user's true intention. Such conflicts may frequently occur in multimodal interactions, requiring an effective resolution mechanism to clarify the user's true needs.

[0085] In this embodiment, the priority arbitration principle is a pre-defined set of rules used to determine the priority of each instruction when multiple user-generated commands in various input formats conflict. These principles are typically based on factors such as the urgency of the command, the type of operation, and user habits. For example, voice commands may have higher priority than gesture commands because voice commands are more accurate in expressing complex intentions; or safety-related commands may have higher priority than ordinary function commands, such as the "emergency call" voice command, which always has higher priority to ensure priority response in emergency situations. Through clear priority arbitration principles, the AI ​​smart glasses can make reasonable decisions when commands conflict.

[0086] In this embodiment, the priority of all conflicting user input commands is determined based on a priority arbitration principle. This means that when multiple user commands in different input formats conflict, the AI ​​smart glasses evaluate and sort each conflicting command according to a pre-set priority arbitration principle, determining their respective priority levels. For example, when a user simultaneously issues a voice command to "open navigation to the airport" and a gesture command to switch applications, according to the set priority arbitration principle, the voice command is determined to have a higher priority due to its clearer task orientation and potential urgency, while the gesture command to switch applications has a relatively lower priority. This clear priority division helps the AI ​​smart glasses to accurately handle conflicting commands subsequently.

[0087] In this embodiment, the AI ​​big data model is used to calculate the user intent confidence distribution based on all the latest received user input commands and their corresponding priorities. This means that after determining the priorities of conflicting user input commands, the AI ​​smart glasses provide all the latest received commands and their corresponding priority information to the AI ​​big data model. The AI ​​big data model considers these priority factors when calculating the user intent confidence distribution. For example, the model assigns more weight to high-priority commands, tending to infer the user's intent based on those commands. Suppose a high-priority voice command expresses the intent to search for nearby hospitals, while a low-priority gesture command seems related to adjusting display brightness. When calculating the intent confidence distribution, the AI ​​big data model will focus on analyzing the intent to search for nearby hospitals, obtaining results such as a confidence score of 0.8 for searching for nearby hospitals and a confidence score of 0.2 for adjusting display brightness. This more accurately reflects the user's primary intent tendency in the case of command conflicts, laying the foundation for accurately identifying the user's true intent in the future.

[0088] like Figure 3 As shown, in order to perform targeted analysis on speech stream, visual stream, and head vibration stream data respectively, and to fuse the analysis results to generate a unified contextual understanding vector, a multimodal perception fusion module is further proposed, including:

[0089] The speech perception and analysis submodule is used to extract text and perceive emotions from real-time perceived speech stream data to obtain intention-perceived text, perceived emotion type sequence and corresponding perceived emotion degree sequence.

[0090] The visual perception analysis submodule is used to identify objects of attention based on real-time perceived visual stream data, and to obtain all objects of attention and the attention-focusing sliding path;

[0091] The head vibration sensing and analysis submodule is used to identify the direction and amplitude of head tremors based on real-time sensed head vibration flow data.

[0092] The multimodal fusion generation submodule is used to generate a unified contextual understanding vector based on intent-aware text, a sequence of perceived emotion types, a corresponding sequence of perceived emotion levels, all attention-focused objects, attention-focused sliding paths, head tremor direction, and head tremor amplitude.

[0093] In this embodiment, text extraction and emotion perception are performed on real-time perceived speech stream data to obtain intent-perceived text, a sequence of perceived emotion types, and a corresponding sequence of perceived emotion levels. This refers to the AI ​​smart glasses utilizing their speech perception analysis submodule to deeply process the continuous speech data spoken by the user in real time. For text extraction, speech is converted into text using speech recognition technology, and then natural language processing methods are used to extract text that directly reflects the core content of the user's intent—that is, intent-perceived text. For example, if a user says, "I'm going to the gym this afternoon, can you check the route for me?", the extracted text would be "Check the route to the gym this afternoon." For emotion perception, speech sentiment analysis technology is used to identify the emotional information contained in the speech, forming a sequence of perceived emotion types. The changes in emotion types during the speech are recorded in chronological order, such as ["expectation", "eagerness"]. Simultaneously, the intensity of each emotion is determined, forming a sequence of perceived emotion levels, such as ["moderate", "high"]. This information lays the foundation for a comprehensive understanding of the user's intent and emotional state, providing more tailored interactive feedback to the user's needs.

[0094] In this embodiment, the intent-aware text is extracted from real-time perceived speech stream data by the speech perception analysis submodule using speech recognition and natural language processing technologies. It is used to reflect the user's core intent. It is extracted from the user's speech content. For example, if a user says, "I want to eat hot pot tonight, can you recommend some restaurants?", the intent-aware text might be "Looking for recommendations for hot pot tonight," providing key textual information for understanding the user's intent.

[0095] In this embodiment, the perceived emotion type sequence and the corresponding perceived emotion degree sequence are products of emotion perception from speech stream data. The perceived emotion type sequence records the emotion categories identified during the speech process in chronological order, and the perceived emotion degree sequence indicates the intensity of each emotion type. For example, if a user excitedly says, "Great, I passed the exam," the corresponding emotion type in the perceived emotion type sequence might be "excitement," and the corresponding emotion degree in the degree sequence might be "high," thus helping to determine the impact of emotions on the user's intentions.

[0096] In this embodiment, the recognition of the object of attention based on real-time perceived visual stream data is achieved by using the visual perception function of AI smart glasses to analyze the image sequence acquired by the camera, and using computer vision technology to identify the specific object that the user is paying attention to. At the same time, the trajectory of attention shifting between objects is recorded. For example, in a classroom scene, the user looks at the blackboard first and then at the textbook. The blackboard and the textbook are the objects of attention. The trajectory of the gaze shift from the blackboard to the textbook is the sliding path of attention, which helps to understand the user's focus and changes in thinking.

[0097] In this embodiment, the object of attention is the specific object, person, or area that the user is currently focusing their attention on when analyzing visual stream data. For example, at an auto show, the user's object of attention may be a new type of car. Identifying this object helps to infer the relationship between the user's interests and intentions.

[0098] In this embodiment, the attention focus sliding path is a trajectory that records the transfer of user attention between different attention focus objects. It is obtained by tracking the direction and order of user eye movement in visual flow data. For example, in a shopping mall, a user's eye movement route from looking at clothes to looking at accessories and then to looking at the cashier is the attention focus sliding path, which can reflect the user's thinking process and demand direction when making a purchase.

[0099] In this embodiment, the direction and amplitude of head tremors are identified based on real-time perceived head vibration flow data. This is achieved by collecting head vibration data using sensors in smart glasses, and then using algorithms to analyze the direction of head tremors (such as up and down, left and right) and the quantitative amplitude representing the intensity of the tremors. For example, when a user is excited, they may tremble significantly from side to side, while when thinking, they may tremble slightly from up and down, providing a basis for judging the user's emotions or potential intentions.

[0100] To determine the intention shift and degree of intention shift for all attention-focused objects based on the mapping rules of emotion, action, and intention, and to generate a unified contextual understanding vector, a multimodal fusion generation submodule is further proposed, including:

[0101] The first intention evolution analysis unit is used to determine the first intention evolution sequence based on the perceived emotion type, perceived emotion degree and emotion-intention mapping rule;

[0102] The second intention evolution analysis unit is used to determine the second intention evolution sequence based on the direction and amplitude of head tremor and the action-intention mapping rule.

[0103] The comprehensive intention evolution analysis unit is used to match and analyze the first intention evolution sequence and the second intention evolution sequence with the attention focusing sliding path to obtain the comprehensive intention evolution sequence of the attention focusing sliding path;

[0104] The intention bias analysis unit is used to determine the intention shift and degree of intention shift for all attention-focused objects based on the comprehensive intention evolution sequence of the attention-focusing sliding path.

[0105] The Unified Contextual Understanding Unit is used to generate a unified contextual understanding vector based on the intent-aware text and the intent shift and degree of intent shift of all attention-focused objects.

[0106] In this embodiment, the emotion-intention mapping rule is a pre-defined association rule that establishes a connection between different types of perceived emotions and their corresponding degrees of perceived emotion and the evolution of the user's potential intention. It is used to infer the development direction of the user's intention from the user's emotional state and to provide a basis for analyzing the emotional dimension of the user's intention.

[0107] In this embodiment, determining the first intention evolution sequence based on the perceived emotion type and perceived emotion degree and the emotion-intention mapping rule means that by analyzing the perceived emotion type and degree in the user's speech and according to the emotion-intention mapping rule, a sequence reflecting the possible evolution order of the user's intention due to changes in emotion is sorted out, thereby initially outlining the development path of the user's intention from an emotional perspective.

[0108] Suppose a user says while choosing a phone, "Wow, this phone looks so beautiful, and the specs seem pretty good too!" Based on the emotion perception of the voice stream data, the perceived emotion type is "like," and the perceived emotion level is "moderate." According to a pre-defined emotion-intention mapping rule, this emotional state of "like (moderate)" may map to a series of intention developments. First, they might want to learn more about the phone's details; then they might consider buying it; and later they might think about how to match phone accessories. The first intention evolution sequence determined based on these mapping relationships might be: [learn about phone details, consider buying, think about matching accessories], this sequence initially shows the user's possible intention evolution path from an emotional perspective, providing clues based on emotion analysis for a comprehensive understanding of user intentions.

[0109] In this embodiment, the action-intention mapping rule is a set of rules that associate head movement features with user intentions. It specifies the possible intention changes corresponding to different head tremor directions and amplitudes, providing guidance for inferring user intentions from an action perspective.

[0110] In this embodiment, the second intention evolution sequence is determined based on the direction and amplitude of head tremors and the action-intention mapping rule. This involves generating a sequence reflecting the order of intention changes triggered by head movements, based on the head tremor direction and amplitude information perceived by the smart glasses and combined with the action-intention mapping rule. This provides a behavioral reference for a comprehensive understanding of user intentions. Imagine a user attending an electronics exhibition, standing in front of a display of a new computer. The user's head begins to tremble rapidly with small, noticeable back-and-forth amplitudes. Assume the action-intention mapping rule indicates that small, rapid back-and-forth head tremors represent interest in the displayed item, with larger amplitudes indicating higher interest. In this case, the amplitude of the user's head tremors corresponds to "high interest." Based on this rule, the possible second intention evolution sequence is: [further understanding of computer specifications, considering the possibility of purchase, inquiring about purchase-related matters]. This sequence, based on the intention tendencies reflected in the head movements, demonstrates the user's development process from initial interest in the computer revealed by head movements to a series of related intentions that may follow.

[0111] In this embodiment, the first intention evolution sequence and the second intention evolution sequence are matched and analyzed with the attention focusing sliding path to obtain the comprehensive intention evolution sequence of the attention focusing sliding path. This combines the intention evolution sequence obtained from the perspective of emotion and action with the path of user attention shifting between different objects, and comprehensively considers the influence of multiple factors on intention, thereby deriving a more comprehensive sequence that better reflects the evolution of user intention in the process of focusing on different objects.

[0112] In this embodiment, the intention shift and degree of intention shift for all attention-focused objects are determined based on the comprehensive intention evolution sequence of the attention-focused sliding path. This is achieved by analyzing the comprehensive intention evolution sequence to determine the direction of change of the user's intention when focusing on each object (i.e., the shift of intention) and the relative intensity of this change (i.e., the degree of intention shift), so as to clarify the differences and magnitude of the user's intentions towards different objects of attention.

[0113] In this embodiment, the intention shift and the degree of intention shift represent the direction and intensity of the change in a user's intention relative to previous objects of attention or the overall trend of intention when focusing on a particular object. Intention shift reflects the direction in which the intention changes, while the degree of intention shift quantifies the magnitude of this change, helping to more accurately grasp the changes in a user's intention towards different objects.

[0114] In this embodiment, a unified contextual understanding vector is generated based on the intent-aware text and the intent shift and degree of intent shift of all attention-focused objects. This integrates the intent-aware text extracted from speech with the intent shift and degree of shift information obtained from analyzing the attention-focused objects, presenting a comprehensive and integrated vector representation of the user's context and intent state. This provides a unified and comprehensive basis for subsequent accurate analysis of the user's true intent. For example, suppose a user says "I want to buy a pair of running shoes" in a shopping mall; this is the intent-aware text. At this time, the user's attention is sequentially focused on running shoes from different brands, such as first looking at brand A running shoes, then looking at brand B running shoes. Furthermore, when switching from brand A to brand B, the intent shifts from focusing on price to focusing on comfort; this is an intent shift, and analysis shows a high degree of intent shift. When generating the unified contextual understanding vector based on this information, the intent-aware text "buy running shoes" is encoded, and the intent shift and degree of interest in brands A and B running shoes are also encoded numerically. For example, the intent-aware text can be transformed into a specific numerical sequence representing the intent to "buy running shoes." Another numerical value can represent the direction of intent shift from brand A to brand B (e.g., 1 represents a shift from focusing on price to focusing on comfort), and yet another numerical value can represent the degree of intent shift (e.g., 8 represents a higher degree). These numerical values ​​combined form a unified contextual understanding vector, comprehensively reflecting the user's changing intent and focus within the given scenario.

[0115] To obtain the comprehensive intent evolution sequence of the attention-focusing sliding path by dividing, assigning weights, and performing weighted calculations, a comprehensive intent evolution analysis unit is further proposed, including:

[0116] The path segmentation sub-unit is used to divide the attention focusing sliding path based on the gaze region type to obtain multiple attention focusing sub-sliding paths;

[0117] The weight allocation sub-unit is used to determine the path allocation weight of each attention focus sub-sliding path based on the gaze region type of each attention focus sub-sliding path;

[0118] The weighted operation subunit is used to align the first intention evolution sequence and the second intention evolution sequence with the attention focusing sliding path using timestamps, and to perform alignment weighting operation on the first intention evolution sequence and the second intention evolution sequence based on the path allocation weights of all attention focusing sub-sliding paths to obtain the comprehensive intention evolution sequence of the attention focusing sliding path.

[0119] In this embodiment, gaze region type refers to a category classified according to the characteristics, function, or attributes of the area where the user's visual focus is located. For example, in an indoor scene, gaze region type may include furniture area, appliance area, decorative area, etc.; in an outdoor scene, it may include natural landscape area, building area, road area, etc. This classification helps to understand the content that the user is focusing on and its potential intentions in more detail.

[0120] In this embodiment, the attention-focusing sliding path is divided into multiple attention-focusing sub-sliding paths based on the gaze region type. This means that the complete path of the user's attention shifting between different objects is divided into multiple smaller segments according to the different gaze region types. For example, if the user's attention-focusing sliding path moves from the television (appliance area) in the living room to the sofa (furniture area), and then to the painting on the wall (decorative area), this path can be divided into a sub-sliding path from the television to the sofa (appliance area to furniture area) and a sub-sliding path from the sofa to the painting (furniture area to decorative area) according to the gaze region type. This allows for a clearer analysis of the changes in user intent along each small segment of the path.

[0121] In this embodiment, the path allocation weight for each attention-focusing sub-slide path is determined based on the gaze region type. This is because different gaze region types may have varying importance or relevance to the user's intent expression, assigning a weight value to each sub-slide path. For example, suppose a user is in a home furnishing store, and the attention-focusing slide path involves moving from the kitchenware section to the bedroomware section. Based on gaze region type, the kitchenware section and the bedroomware section are different gaze region types. If the user is planning to renovate their kitchen, then the kitchenware section is highly relevant to the user's current intent. In this case, the portion of the attention-focusing sub-slide path from the kitchenware section to the bedroomware section will be assigned a higher path allocation weight, such as 0.7; while the portion involving the bedroomware section will have a relatively lower weight, such as 0.3. This is because the gaze region type of the kitchenware section is more closely related to the user's current primary intent (renovating the kitchen) and should be given greater consideration when analyzing user intent.

[0122] In this embodiment, the first and second intention evolution sequences are timestamped and aligned with the attention-focusing sliding path. Then, based on the path allocation weights of all attention-focusing sub-sliding paths, an alignment weighted calculation is performed on the first and second intention evolution sequences to obtain the comprehensive intention evolution sequence of the attention-focusing sliding path. Specifically, timestamping alignment ensures that the first intention evolution sequence from sentiment analysis and the second intention evolution sequence from action analysis correspond to the attention-focusing sliding path in the time dimension, guaranteeing that elements in each sequence match the corresponding moments in the attention-focusing process. Then, according to the previously determined weights of each attention-focusing sub-sliding path, the elements of the first and second intention evolution sequences on the corresponding sub-sliding paths are weighted and calculated. For example, on a certain attention-focusing sub-sliding path, a certain intention element in the first intention evolution sequence has a weight of 0.6, and the corresponding element in the second intention evolution sequence has a weight of 0.4. The weighted calculation yields the comprehensive intention elements on that sub-sliding path. Finally, the results from all sub-sliding paths are integrated to form the comprehensive intention evolution sequence of the attention-focusing sliding path, thus comprehensively reflecting the user's intention evolution throughout the entire attention shift process.

[0123] To determine the intent shift and degree of the attention focus object based on the comprehensive intent evolution sequence of the attention-focused sliding path, an intent bias analysis unit is further proposed, including:

[0124] The intent-direction determination subunit is used to determine the intent direction of all attention-focused objects based on the comprehensive intent evolution sequence of the attention-focusing sliding path;

[0125] The same-direction and opposite-direction path division sub-unit is used to filter out attention-focused objects that have undergone obvious intention shifts from all attention-focused objects based on the comprehensive intention evolution sequence of attention-focused sliding paths as dominant intention shift objects, and to determine the current same-direction intention accumulation path and future opposite-direction accumulation path of each dominant intention shift object based on the comprehensive intention evolution sequence of attention-focused sliding paths.

[0126] The cumulative bias determination subunit is used to determine the current same-direction cumulative bias and future opposite-direction cumulative bias of each dominant intention turning object based on the comprehensive intention evolution sequence;

[0127] The turning degree calculation subunit is used to calculate the intention turning degree of each attention focus object based on the current same-direction cumulative bias, future opposite-direction cumulative bias, current same-direction cumulative path, and future opposite-direction cumulative path of each dominant intention turning object.

[0128] In this embodiment, the intention shift for all attention-focused objects is determined based on the comprehensive intention evolution sequence of the attention-focused sliding path. This is achieved by analyzing this comprehensive sequence and observing whether the relevant intention for each attention-focused object changes direction compared to the previous object as the user's attention shifts from one object to another. For example, if the comprehensive intention evolution sequence shows that the user's intention changes from viewing information to preparing to charge when they move from focusing on the phone (the first attention-focused object) to focusing on the charger (the second attention-focused object), then it is determined that an intention shift has occurred for the charger as the attention-focused object.

[0129] In this embodiment, the comprehensive intent evolution sequence based on the attention focus sliding path filters out attention focus objects that show significant intent shifts from all attention focus objects as dominant intent shift objects. This involves selecting objects with more obvious intent shifts and greater impact on the overall intent direction after determining the intent shifts of all attention focus objects. For example, in a shopping scenario, a user's attention focuses sequentially on clothes, shoes, and bags. The comprehensive intent evolution sequence shows that the intent doesn't change much from clothes to shoes, but a significant shift occurs from shoes to bags, such as from focusing on comfort to focusing on aesthetics. In this case, the bag might be selected as the dominant intent shift object because its intent shift is more crucial for understanding the changes in the user's overall shopping intent.

[0130] In this embodiment, the comprehensive intent evolution sequence based on the attention-focused sliding path determines the current same-direction intent accumulation path and the future opposite-direction accumulation path for each dominant intent-turning object. This involves analyzing, within the comprehensive intent evolution sequence, the paths formed by intent accumulation in the same direction as the current intent before the object becomes the dominant intent-turning object (i.e., the current same-direction intent accumulation path); and the paths formed by intent accumulation in a different direction after the object becomes the dominant intent-turning object (i.e., the future opposite-direction accumulation path). For example, when a user is browsing a travel guide, they first focus on hotel information (the dominant intent-turning object), and then their attention shifts to attraction descriptions. From the start of browsing the guide to focusing on hotel information, all intent accumulations related to accommodation form the current same-direction intent accumulation path; while from focusing on hotel information to focusing on attraction descriptions, intent accumulations related to attractions (different from the accommodation focus) form the future opposite-direction accumulation path.

[0131] In this embodiment, the current cumulative bias in the same direction and the future cumulative bias in the opposite direction for each dominant intent-turning object are determined based on the comprehensive intent evolution sequence. This is achieved by analyzing the comprehensive intent evolution sequence to measure the degree to which the intent of each dominant intent-turning object deviates from or concentrates on the current cumulative path in the same direction and the future cumulative path in the opposite direction. Assume a user is selecting a mobile phone on an e-commerce platform. Initially, the user focuses on the phone's camera function, browsing several phones that emphasize photography. Later, their attention shifts to the phone's gaming performance.

[0132] Imagine a user browsing electronics on an e-commerce platform. Initially, they focus on selecting a laptop, primarily concerned with processor performance. Later, their attention shifts to the laptop's portability. This shift from focusing on processor performance to focusing on portability represents a change in dominant intent towards the desired device.

[0133] Calculation of current cumulative bias: Before the dominant intent shifted, i.e., during the processor performance focus phase, the user viewed 10 items related to laptops. Of these, 7 were about the number of processor cores, 2 were about processor frequency, and 1 was about the processor brand. We can determine the current cumulative bias by calculating the percentage of user attention focused on a specific aspect (such as the number of processor cores).

[0134] Let the total number of attentions be N (here N = 10), and the number of attentions to a specific aspect (number of processor cores) be n (here n = 7). The bias formula can be simply expressed as bias = n ÷ N. Therefore, the bias of attention regarding the number of processor cores on the current cumulative path of same-direction intent is 7 ÷ 10 = 0.7. If we comprehensively consider various aspects of the processor, we can calculate the overall current cumulative bias using methods such as weighted averaging. For example, if the weight of processor core count is set to 0.5, frequency to 0.3, and brand to 0.2, then the current cumulative bias is 0.7 × 0.5 + 0.2 × 0.3 + 0.1 × 0.2 = 0.43.

[0135] Calculation of future anomaly cumulative bias: After the dominant intent shifts, i.e., the focus shifts to portability, the user viewed 8 pieces of information, including 3 about the computer's weight, 2 about its size, and 3 about its battery life (battery life is also related to portability). Following the same method, let the total number of attentions be M (here, M=8). Taking computer weight as an example, the number of attentions is m (here, m=3). Then, the bias for the attention point of computer weight is 3÷8=0.375. Similarly, the overall future anomaly cumulative bias is calculated using a weighted average. Assuming a weight weight of 0.4, a size weight of 0.3, and a battery life weight of 0.3, the future anomaly cumulative bias is 0.375×0.4+0.25×0.3+0.375×0.3=0.35625.

[0136] This calculation method yields the current cumulative bias in the same direction and the future cumulative bias in the opposite direction for each dominant intent-turning object. These reflect the degree to which users focus on different aspects before and after the intent shift.

[0137] To accurately calculate the degree of intention shift for each attentional focus object based on the cumulative bias of the dominant intention shift object (both homing and dishoming), a sub-unit for calculating the degree of shift is further proposed, including:

[0138] The first steering degree determination end is used to calculate the comprehensive steering degree of the corresponding current homing intention cumulative path and the corresponding future dissimilar cumulative path based on the current homing intention cumulative bias degree and the future dissimilar cumulative bias degree of each dominant intention steering object;

[0139] The second turning degree determination end is used to calculate the first turning degree of all attention-focused objects in the current same-direction intention accumulation path and the second turning degree of all attention-focused objects in the corresponding future opposite-direction accumulation path of each dominant intention turning object based on the comprehensive intention evolution sequence and the comprehensive turning degree of the current same-direction intention accumulation path and the corresponding future opposite-direction accumulation path of each dominant intention turning object.

[0140] The third turning degree determination end is used to calculate the intention turning degree of each attention-focused object based on all turning degrees of each attention-focused object.

[0141] In this embodiment, the comprehensive turning degree of the corresponding current homing intention cumulative path and the corresponding future dissimilar cumulative path is calculated based on the current homing intention cumulative bias and the future dissimilar cumulative bias of each dominant intention turning object. This is achieved by weighting the absolute value (e.g., with a weight of 0.6) and the average value (e.g., with a weight of 0.4) of the difference between the current homing intention cumulative bias and the future dissimilar cumulative bias of each dominant intention turning object to obtain the comprehensive turning degree of the corresponding current homing intention cumulative path and the corresponding future dissimilar cumulative path.

[0142] In this embodiment, based on the comprehensive intention evolution sequence and the comprehensive turning degree of the current same-direction intention accumulation path and the corresponding future opposite-direction accumulation path of each dominant intention turning object, the first turning degree of all attention-focused objects in the current same-direction intention accumulation path and the second turning degree of all attention-focused objects in the corresponding future opposite-direction accumulation path of each dominant intention turning object are calculated. Specifically, using the intention evolution information contained in the comprehensive intention evolution sequence and combining it with the calculated comprehensive turning degree, for each attention-focused object in the current same-direction intention accumulation path, the specific degree of intention turning when transitioning from this path to the future opposite-direction accumulation path is evaluated to obtain the first turning degree; similarly, for each attention-focused object in the future opposite-direction accumulation path, its corresponding intention turning degree, i.e., the second turning degree, is calculated.

[0143] Calculate the first degree of steering:

[0144] In the current path of accumulating unidirectional intent (attention-focused scenic spot stage), there are multiple objects of attention, such as specific scenic spot names (e.g., Huangshan, Taishan) and scenic spot features (e.g., natural scenery, historical culture).

[0145] Taking "Huangshan" as an example of an object of attention, assuming that by analyzing user behavior when browsing information related to Huangshan, we find that the proportion of user attention to Huangshan information during the entire attraction attention phase is p1 = 0.3, then the first degree of shift in attention to "Huangshan" as an object of attention, D1, can be calculated using the following formula:

[0146] D1=p1×S=0.3×0.388=0.1164;

[0147] The first degree of turning is calculated for each attention focus object in the current homing intent accumulation path using this method.

[0148] Calculate the second degree of steering:

[0149] In the future anisotropic cumulative path (focusing on the route planning stage), there are also multiple objects of attention, such as the starting point, the ending point, and the cities along the route.

[0150] Taking the "route starting point" as the focus of attention as an example, assuming that the proportion of user attention to the route starting point information in the entire route planning attention phase is p2 = 0.2, then the second turning degree D2 of the "route starting point" as the focus of attention is calculated as follows:

[0151] D2=p2×S=0.2×0.388=0.0776;

[0152] The second turning degree is calculated using this method for each attention focus object in the future divergent cumulative path.

[0153] In this embodiment, the total degree of turning for each attention-focused object refers to the set of various degrees of turning calculated for an attention-focused object based on different stages and paths during the analysis of the entire attention-focusing sliding path. These degrees of turning include the first degree of turning in the current same-direction intention accumulation path and the second degree of turning in the future opposite-direction accumulation path.

[0154] In this embodiment, the intentional shift degree of each attention focus object is calculated based on all shift degrees of each attention focus object, which means taking the average of all shift degrees of each attention focus object as the intentional shift degree of each attention focus object.

[0155] To generate interactive feedback results tailored to the user's true intent by retrieving interaction pattern libraries and interactive feedback patterns under the user's true intent, an interactive feedback generation module is further proposed, including:

[0156] The mapping rule retrieval submodule is used to retrieve the interaction pattern mapping rules under the user's emotion tag from the interaction pattern library;

[0157] The feedback pattern retrieval submodule is used to retrieve the interaction feedback pattern under the user's true intention from the interaction pattern mapping rules under the user's emotion tag.

[0158] The feedback result generation submodule is used to generate interactive feedback results based on the user's true intent, the interactive feedback pattern, and the user's true intent.

[0159] In this embodiment, the interaction pattern library is a pre-built collection that stores information related to various interaction patterns. It includes different user emotion tags and corresponding mapping rules for various interaction patterns. These rules specify the interaction methods to be adopted in different emotional states and for different user intentions. It is an important basis for AI smart glasses to generate appropriate interactive feedback, similar to a database containing rich interaction strategies.

[0160] In this embodiment, retrieving interaction pattern mapping rules under the user's emotion tag from the interaction pattern library means that the AI ​​smart glasses search for matching rules in the interaction pattern library based on the user's emotion tag output by the emotion classification model. For example, if the emotion classification model determines that the user's emotion tag is "happy," the system will search the interaction pattern library for a series of rules corresponding to the emotion "happy." These rules will guide the smart glasses on how to provide interactive feedback to the user's various possible intentions under the emotion of happiness.

[0161] In this embodiment, retrieving the interaction feedback pattern corresponding to the user's true intent from the interaction pattern mapping rules under the user's emotion tag involves further filtering out specific interaction feedback patterns applicable to the analyzed user's true intent from these rules after finding the interaction pattern mapping rules corresponding to the user's emotion tag. For example, in the rules corresponding to the "happy" emotion tag, if the user's true intent is to search for information about tourist attractions, then the interaction feedback patterns specifically for this intent under the happy emotion are found, which may include introducing the attractions in a cheerful tone of voice, displaying brightly colored pictures of the attractions, etc.

[0162] In this embodiment, the interactive feedback mode based on the user's true intent clarifies the specific interaction methods and strategies that AI smart glasses should adopt under a combination of a particular user's emotions and true intent. It specifies aspects such as the form of feedback (voice, text, images, etc.), the style of the feedback content (formal, lively, concise, etc.), and the focus of the feedback. For example, for a user's true intent to quickly obtain information about nearby hospitals in an anxious state, the interactive feedback mode might prioritize highlighting the nearest hospital's address and contact number in concise and clear text, accompanied by short and urgent voice prompts.

[0163] In this embodiment, the interactive feedback pattern based on the user's true intention and the generation of interactive feedback results tailored to the user's true intention mean that the AI ​​smart glasses generate the final response presented to the user based on the determined interactive feedback pattern and the specific content of the user's true intention. For example, if the interactive feedback pattern introduces tourist attractions in a vivid and interesting way, and the user's true intention is to learn about the attractions in Huangshan, then the interactive feedback result might be a lively voice describing the famous attractions of Huangshan and displaying some beautiful pictures of Huangshan, while also labeling the pictures with the names of the attractions and brief descriptions, to meet the user's needs under specific emotions and intentions.

[0164] This invention provides an implementation method for an interaction method of AI smart glasses, comprising:

[0165] Receive the user's original instructions and call the AI ​​large model to calculate the user's intent confidence distribution;

[0166] Real-time perceived speech stream data, visual stream data, and head vibration stream data are fused to generate a unified contextual understanding vector;

[0167] The user's true intent is analyzed based on the unified contextual understanding vector and the user intent confidence distribution.

[0168] The head micro-vibration feature vector, the perceived emotion type sequence based on speech stream data, and the perceived emotion degree sequence are input into the emotion classification model to output the user's emotion label.

[0169] Based on the interaction pattern mapping rules under the user's emotion tag, generate interactive feedback results that target the user's true intentions.

[0170] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of this invention and its equivalents, this invention also intends to include these modifications and variations.

Claims

1. An AI smart glasses, characterized in that, include: The initial intent recognition module is used to receive the user's original instructions and call the AI ​​large model to calculate the user's intent confidence distribution; The multimodal perception fusion module is used to fuse real-time perceived speech stream data, visual stream data, and head vibration stream data to generate a unified contextual understanding vector; The intent integration and recognition module is used to analyze the user's true intent based on a unified contextual understanding vector and user intent confidence distribution. The user emotion recognition module is used to input the head micro-vibration feature vector, the perceived emotion type sequence based on speech stream data, and the perceived emotion degree sequence into the emotion classification model, and output the user emotion label. The interaction feedback generation module is used to generate interaction feedback results that target the user's true intentions based on the interaction pattern mapping rules under the user's emotion tags.

2. The AI ​​smart glasses according to claim 1, characterized in that, The initial intent recognition module includes: The multi-form receiving submodule is used to receive user input commands in various input formats. The intent recognition submodule is used to calculate the user intent confidence distribution based on all the latest received user input commands by calling the AI ​​big model.

3. The AI ​​smart glasses according to claim 2, characterized in that, Also includes: The priority arbitration submodule is used to determine the priority of all user input commands that conflict when there are conflicting original user commands in multiple input formats, based on the priority arbitration principle, and to call the AI ​​big model to calculate the user intent confidence distribution based on all the latest received user input commands and their corresponding priorities.

4. The AI ​​smart glasses according to claim 1, characterized in that, The multimodal perception fusion module includes: The speech perception and analysis submodule is used to extract text and perceive emotions from real-time perceived speech stream data to obtain intention-perceived text, perceived emotion type sequence and corresponding perceived emotion degree sequence. The visual perception analysis submodule is used to identify attention-focused objects based on real-time perceived visual stream data, and to obtain all attention-focused objects and attention-focusing sliding paths; The head vibration sensing and analysis submodule is used to identify the direction and amplitude of head tremors based on real-time sensed head vibration flow data. The multimodal fusion generation submodule is used to generate a unified contextual understanding vector based on intent-aware text, a sequence of perceived emotion types, a corresponding sequence of perceived emotion levels, all attention-focused objects, attention-focused sliding paths, head tremor direction, and head tremor amplitude.

5. The AI ​​smart glasses according to claim 4, characterized in that, The multimodal fusion generation submodule includes: The first intention evolution analysis unit is used to determine the first intention evolution sequence based on the perceived emotion type, perceived emotion degree and emotion-intention mapping rule; The second intention evolution analysis unit is used to determine the second intention evolution sequence based on the direction and amplitude of head tremor and the action-intention mapping rule. The comprehensive intention evolution analysis unit is used to match and analyze the first intention evolution sequence and the second intention evolution sequence with the attention focusing sliding path to obtain the comprehensive intention evolution sequence of the attention focusing sliding path; The intention bias analysis unit is used to determine the intention shift and degree of intention shift for all attention-focused objects based on the comprehensive intention evolution sequence of the attention-focusing sliding path; The Unified Contextual Understanding Unit is used to generate a unified contextual understanding vector based on the intent-aware text and the intent shift and degree of intent shift of all attention-focused objects.

6. The AI ​​smart glasses according to claim 5, characterized in that, The integrated intentional evolution analysis unit includes: The path segmentation sub-unit is used to divide the attention focusing sliding path based on the gaze region type to obtain multiple attention focusing sub-sliding paths; The weight allocation sub-unit is used to determine the path allocation weight of each attention focus sub-sliding path based on the gaze region type of each attention focus sub-sliding path; The weighted operation subunit is used to align the first intention evolution sequence and the second intention evolution sequence with the attention focusing sliding path using timestamps, and to perform alignment weighting operation on the first intention evolution sequence and the second intention evolution sequence based on the path allocation weights of all attention focusing sub-sliding paths to obtain the comprehensive intention evolution sequence of the attention focusing sliding path.

7. The AI ​​smart glasses according to claim 5, characterized in that, Intent bias analysis unit, including: The intent-direction determination subunit is used to determine the intent direction of all attention-focused objects based on the comprehensive intent evolution sequence of the attention-focusing sliding path; The same-direction and opposite-direction path division sub-unit is used to filter out attention-focused objects that have undergone obvious intention shifts from all attention-focused objects based on the comprehensive intention evolution sequence of attention-focused sliding paths as dominant intention shift objects, and to determine the current same-direction intention accumulation path and future opposite-direction accumulation path of each dominant intention shift object based on the comprehensive intention evolution sequence of attention-focused sliding paths. The cumulative bias determination subunit is used to determine the current same-direction cumulative bias and future opposite-direction cumulative bias of each dominant intention turning object based on the comprehensive intention evolution sequence; The turning degree calculation subunit is used to calculate the intention turning degree of each attention focus object based on the current same-direction cumulative bias, future opposite-direction cumulative bias, current same-direction cumulative path, and future opposite-direction cumulative path of each dominant intention turning object.

8. The AI ​​smart glasses according to claim 7, characterized in that, The steering degree calculation subunit includes: The first steering degree determination end is used to calculate the comprehensive steering degree of the corresponding current homing intention cumulative path and the corresponding future dissimilar cumulative path based on the current homing intention cumulative bias degree and the future dissimilar cumulative bias degree of each dominant intention steering object; The second turning degree determination end is used to calculate the first turning degree of all attention-focused objects in the current same-direction intention accumulation path and the second turning degree of all attention-focused objects in the corresponding future opposite-direction accumulation path of each dominant intention turning object based on the comprehensive intention evolution sequence and the comprehensive turning degree of the current same-direction intention accumulation path and the corresponding future opposite-direction accumulation path of each dominant intention turning object. The third turning degree determination end is used to calculate the intention turning degree of each attention-focused object based on all turning degrees of each attention-focused object.

9. The AI ​​smart glasses according to claim 1, characterized in that, The interactive feedback generation module includes: The mapping rule retrieval submodule is used to retrieve the interaction pattern mapping rules under the user's emotion tag from the interaction pattern library; The feedback pattern retrieval submodule is used to retrieve the interaction feedback pattern under the user's true intention from the interaction pattern mapping rules under the user's emotion tag. The feedback result generation submodule is used to generate interactive feedback results based on the user's true intent, the interactive feedback pattern, and the user's true intent.

10. An interaction method for AI smart glasses, characterized in that, include: Receive the user's original instructions and call the AI ​​large model to calculate the user's intent confidence distribution; Real-time perceived speech stream data, visual stream data, and head vibration stream data are fused to generate a unified contextual understanding vector; The user's true intent is analyzed based on the unified contextual understanding vector and the user intent confidence distribution. The head micro-vibration feature vector, the perceived emotion type sequence based on speech stream data, and the perceived emotion degree sequence are input into the emotion classification model to output the user's emotion label. Based on the interaction pattern mapping rules under the user's emotion tag, generate interactive feedback results that target the user's true intentions.