Interaction control method, system and equipment for intelligent glasses

Through multimodal data acquisition and co-occurrence frequency analysis, smart glasses identify interaction intentions and evaluate their credibility, solving the problem of false triggering in a single modal interaction mode, and achieving more efficient multimodal interaction control.

CN120335613AActive Publication Date: 2025-07-18DANYANG JINGTONG GLASSES TECH INNOVATION SERVICE CENT

Patent Information

Application Number
CN202510489886.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-07-18
Estimated Expiration
2045-04-18

AI Technical Summary

Technical Problem

The existing smart glasses interactive mode relies on a single mode, resulting in high false triggering rate, unstable user experience, and lack of in-depth analysis of the synergy between multimodal information.

Method used

The multimodal data acquisition device obtains eye movement trajectory, head posture, voice input and facial expression changes, performs interactive feature extraction, analyzes the co-occurrence frequency matrix between modal pairs, calculates the credibility score of the interaction intention, and triggers the control command if it is greater than the threshold.

Benefits of technology

It reduces the erroneous operation rate of smart glasses, improves the multi-modal interaction experience, and improves the accuracy and reliability of interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120335613A_ABST
    Figure CN120335613A_ABST
Patent Text Reader

Abstract

The invention discloses an interaction control method, system and equipment for intelligent glasses, and relates to the related field of intelligent glasses interaction, and the method comprises the steps: obtaining multi-modal interaction information through a multi-modal data collection device; performing interaction feature extraction on the multi-modal interaction information, and identifying an interaction intention trigger point; analyzing a co-occurrence frequency matrix between the modal pairs according to the trigger points to obtain a co-occurrence frequency vector; and inputting the co-occurrence frequency vector into a credibility evaluation module, calculating a credibility score of the interaction intention, if the score is greater than a preset threshold value, triggering an interaction control instruction, and executing a control program by the intelligent glasses. The technical problems that an existing intelligent glasses interaction mode depends on a single mode, the false triggering rate is high, and the user experience is unstable are solved, and the technical effects that through multi-mode cross validation and dynamic intention credibility evaluation, the false operation rate of the intelligent glasses is reduced, and the multi-mode interaction experience is improved are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of smart glasses interaction, and particularly to an interaction control method, system and device for smart glasses. Background Art

[0002] With the continuous development of smart glasses technology, more and more application scenarios require smart glasses to have efficient and accurate interaction control capabilities. Traditional smart glasses interaction methods usually rely on single-modal input, such as voice commands or touch operations. However, single-modal interaction methods are often affected by environmental noise, user movement restrictions or external interference, resulting in frequent occurrences of mis-triggering and interaction failures. To solve this problem, multi-modal interaction control solutions have emerged in recent years. It combines multiple input methods such as voice, eye movement, head pose, and facial expression, and improves the accuracy and reliability of interaction by comprehensively analyzing multiple pieces of information.

[0003] However, the existing technology still faces some challenges, especially in the fusion and analysis of multiple modal information. How to extract effective interaction features from multi-modal data, accurately identify the user's interaction intention, and prevent mis-triggering are important issues in the current smart glasses interaction technology. In traditional methods, there is often a lack of in-depth analysis of the synergistic effects of multi-modal information, resulting in the interaction control system being easily interfered by misoperations and environmental factors, affecting the user experience.

[0004] Therefore, there is an urgent need for a new type of interaction control method that can make full use of the complementarity of multi-modal data, improve the accuracy of smart glasses interaction control, enhance its anti-mis-triggering ability, and thus meet the increasingly complex and changeable usage requirements. Summary of the Invention

[0005] This application provides an interaction control method, system and device for smart glasses, solves the technical problem that the existing smart glasses interaction method relies on a single modality, resulting in a high mis-triggering rate and unstable user experience, and achieves the technical effect of reducing the misoperation rate of smart glasses and improving the multi-modal interaction experience through multi-modal cross-validation and dynamic intention credibility evaluation.

[0006] The present application provides an interactive control method for smart glasses. The smart glasses are embedded with a multimodal data acquisition device, including: obtaining multimodal interaction information by the multimodal data acquisition device, where the multimodal interaction information includes an eye movement trajectory, head pose data, voice input signals, and facial expression changes stored in a time series; extracting interaction features from the multimodal interaction information to identify multimodal interaction intention trigger points; analyzing the co-occurrence frequency matrix between modality pairs according to the multimodal interaction intention trigger points to obtain a modality co-occurrence frequency vector; inputting the modality co-occurrence frequency vector into an interaction intention credibility evaluation module to calculate the credibility score of the current interaction intention. If the credibility score is greater than a first preset credibility threshold, an interactive control instruction is triggered, and the smart glasses execute a control program according to the interactive control instruction.

[0007] The present application also provides an interactive control system for smart glasses. The smart glasses are embedded with a multimodal data acquisition device, including: a data acquisition unit: obtaining multimodal interaction information by the multimodal data acquisition device, where the multimodal interaction information includes an eye movement trajectory, head pose data, voice input signals, and facial expression changes stored in a time series; a feature extraction unit: extracting interaction features from the multimodal interaction information to identify multimodal interaction intention trigger points; a modality analysis unit: analyzing the co-occurrence frequency matrix between modality pairs according to the multimodal interaction intention trigger points to obtain a modality co-occurrence frequency vector; an interactive control unit: inputting the modality co-occurrence frequency vector into an interaction intention credibility evaluation module to calculate the credibility score of the current interaction intention. If the credibility score is greater than a first preset credibility threshold, an interactive control instruction is triggered, and the smart glasses execute a control program according to the interactive control instruction.

[0008] The present application also provides an electronic device, including: a memory for storing executable instructions; a processor for implementing an interactive control method for smart glasses when executing the executable instructions stored in the memory.

[0009] A method, system and device for interactive control of smart glasses proposed in this application first obtains multimodal interaction information by a multimodal data acquisition device. This multimodal interaction information includes eye movement trajectories, head pose data, voice input signals, and facial expression changes stored in time series. Subsequently, interactive feature extraction is performed on the multimodal interaction information to identify multimodal interaction intention trigger points. Then, according to the multimodal interaction intention trigger points, the co-occurrence frequency matrix between modality pairs is analyzed to obtain a modality co-occurrence frequency vector. Finally, the modality co-occurrence frequency vector is input into an interactive intention credibility evaluation module to calculate the credibility score of the current interactive intention. If the credibility score is greater than the first preset credibility threshold, an interactive control instruction is triggered, and the smart glasses execute a control program according to the interactive control instruction. The technical effect of reducing the misoperation rate of smart glasses and improving the multimodal interaction experience is achieved. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings of the embodiments of the present invention will be briefly introduced below. Flowcharts are used in this application to illustrate the operations performed by the systems according to the embodiments of the present application. It should be understood that the operations above or below do not necessarily need to be executed precisely in order. On the contrary, according to the need, they can be executed in reverse order or simultaneously. At the same time, other operations can also be added to these processes, or one or several operations can be removed from these processes.

[0011] Figure 1 It is a schematic flowchart of a method for interactive control of smart glasses provided by an embodiment of the present application.

[0012] Figure 2 It is a schematic structural diagram of an interactive control system for smart glasses provided by an embodiment of the present application.

[0013] Figure 3 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application.

[0014] Description of reference numerals: data acquisition unit 11, feature extraction unit 12, modality analysis unit 13, interactive control unit 14, processor 21, memory 22, input device 23, output device 24. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0015] The above description is only an overview of the technical solutions of this application. In order to be able to understand the technical means of this application more clearly, it can be implemented according to the content of the description. And in order to make the above and other purposes, features and advantages of this application more obvious and understandable, the following specifically describes the specific embodiments of this application.

[0016] To make the objectives, technical solutions, and advantages of this application clearer, the following will further describe this application in detail with reference to the accompanying drawings. The described embodiments should not be construed as limitations on this application. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of this application.

[0017] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments and can be combined with each other without conflict. The terms "first / second" involved are only used to distinguish similar objects and do not represent a specific order for the objects. The terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or server that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or modules that are not clearly listed or are inherent to these processes, methods, products, or devices. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application.

[0018] An embodiment of this application provides an interactive control method for smart glasses. The smart glasses are embedded with a multimodal data acquisition device, as Figure 1 shown. The method includes:

[0019] Obtaining multimodal interaction information by the multimodal data acquisition device, where the multimodal interaction information includes an eye movement trajectory, head pose data, voice input signals, and facial expression changes stored in a time series.

[0020] Specifically, a multimodal data acquisition device is integrated inside the smart glasses. This device is used to simultaneously collect various interaction signals of the user during use. These signals cover information inputs in multiple dimensions, including but not limited to eye movement trajectories, head pose data, voice input signals, and facial expression changes. Among them, the eye movement trajectory is the path of the user's eye movement recorded in real time by an eye movement sensor, reflecting the gaze direction, gaze time, and saccade pattern; the head pose data is the rotation angle and motion state of the user's head detected by inertial sensors such as gyroscopes and accelerometers, used to identify natural postures such as nodding and shaking the head; the voice input signal is the user's voice command collected by a microphone array, used to identify keywords or semantic information; the facial expression change is the information of the user's facial muscle change captured by an embedded camera, used to judge emotions or specific expression operation intentions (such as opening the mouth, blinking, smiling, etc.). The multimodal information collected by the multimodal data acquisition device is synchronously recorded in chronological order to form a set of continuous interaction data streams with time tags, that is, multimodal interaction information, providing a comprehensive and real-time input basis for subsequent interaction feature extraction and intention recognition. This method significantly improves the context integrity of the interaction and helps the system make more accurate interaction judgments.

[0021] Extract interaction features from the multimodal interaction information to identify multimodal interaction intention trigger points.

[0022] Specifically, after the multimodal interaction information is collected, the collected information is stored in blocks according to the information type. Each block is responsible for storing a type of data, such as eye movement trajectories, head pose data, etc. Subsequently, multiple interaction feature extraction blocks are called, and each interaction feature extraction block is matched with the information storage block, so that the interaction feature extraction block corresponds to the information storage block one by one. Then, the interaction feature extraction block is used to extract the key features related to the interaction from the modal interaction information stored in the information storage block. These features can reflect the user's interaction intention. For example, the fixation dwell feature of the eye movement trajectory can reveal the area where the user's attention is concentrated, and the direction feature of the head pose can reflect whether the user's intention is related to a certain direction or target. By extracting these modal features, the multimodal interaction intention trigger points of the user can be identified. These trigger points serve as the basis for subsequent interaction processing, helping the system better understand the user's needs and respond.

[0023] In a possible implementation manner, the method for extracting interaction features from the multimodal interaction information to identify multimodal interaction intention trigger points includes:

[0024] The multi-modal interaction information is stored in blocks to obtain multiple interaction information storage blocks; multiple interaction feature extraction blocks are set, and the multiple interaction feature extraction blocks are correspondingly connected to the multiple interaction information storage blocks. The multiple interaction feature extraction blocks include pre-stored eye movement interaction features, head pose interaction features, voice keyword interaction features, and facial expression interaction features. Each interaction feature extraction block is used to receive the modal data of the interaction information storage block for interaction feature extraction and output a multi-modal interaction intention trigger point.

[0025] Specifically, to improve the efficiency and accuracy of multi-modal interaction information processing, first, the continuously collected multi-modal interaction information is segmented according to content features and divided into multiple interaction information storage blocks. Each block stores any one type of data among eye movement trajectories, head pose data, voice input signals, facial expression changes, etc. Subsequently, multiple interaction feature extraction blocks are configured, and each interaction feature extraction block corresponds one-to-one with the interaction information storage block, and is used to process and analyze the modal data in their respective corresponding blocks. Among them, for each interaction feature extraction block, feature extraction logics for different modalities are preset. For the interaction information storage block storing eye movement trajectories, the interaction feature extraction block pre-storing eye movement interaction features is used to analyze this interaction information storage block. The eye movement interaction features include features such as gaze dwell and rapid saccade. Gaze dwell can be determined by accumulating the timestamps of consecutive identical line-of-sight coordinates in the eye movement trajectory. When the user's line of sight stays in a small area for a long time, such as more than 300 ms, then this area is marked as a candidate intention trigger point; rapid saccade can be determined by recording the time and number of times the line of sight enters a certain area each time. When within a set time window (such as 2 seconds), the user's line of sight enters the same area multiple times (such as 2 times or more), then this area is marked as a candidate intention trigger point. For the interaction information storage block storing head pose data, the interaction feature extraction block pre-storing head pose interaction features is used to analyze this interaction information storage block. The head pose interaction features include features such as rotation direction and rotation amplitude. When the rotation direction in the head pose data is up or down, the number of upward or downward rotations whose rotation amplitude meets the amplitude threshold is counted. If the counted number meets the nod threshold, it indicates that the user is relatively satisfied with the currently gazed area. At this time, this area is marked as a candidate intention trigger point. For the interaction information storage block storing voice input signals, the interaction feature extraction block pre-storing voice keyword interaction features is used to analyze this interaction information storage block. The voice keyword interaction features include multiple keywords, such as "confirm", "like", "hate", "open", etc. When the voice input signal shows that the user makes a keyword response to a certain area, this area is marked as a candidate intention trigger point. For the interaction information storage block storing facial expression changes, the interaction feature extraction block pre-storing facial expression interaction features is used to analyze this interaction information storage block. The facial expression interaction features include multiple expression information, such as raising eyebrows, smiling, opening the mouth. When the user smiles while gazing at a certain area, it may indicate that the user has an intention for this area. At this time, this area is marked as a candidate intention trigger point. After all the interaction feature extraction blocks complete the processing of the stored data, all the candidate intention trigger points are summarized to obtain multi-modal interaction intention trigger points, providing an input basis for subsequent modal co-occurrence analysis and credibility evaluation.

[0026] Analyze the co-occurrence frequency matrix between modality pairs according to the multi-modal interaction intention trigger points, and obtain the modality co-occurrence frequency vector.

[0027] Specifically, use the identified multi-modal interaction intention trigger points to analyze whether there is synchronous or collaborative trigger behavior between different modalities. The purpose is to evaluate whether these modalities frequently appear simultaneously in the same interaction behavior, so as to construct a more reliable basis for judging interaction intentions. Specifically, process the multi-modal interaction intention trigger points according to the timestamps in the multi-modal interaction intention trigger points, and retain the trigger point information within a preset time window (for example, 10 seconds). Subsequently, combine all the modality interaction intention trigger points in pairs to form a number of modality pairs, such as eye movement - speech, speech - facial expression, etc., for counting whether these modality combinations are often triggered simultaneously. Then, analyze these modality pairs through the set co-occurrence determination conditions, calculate their co-occurrence frequencies, and arrange the co-occurrence frequencies of all modality pairs in rows and columns to generate a symmetric co-occurrence frequency matrix. Each element represents the collaboration strength of the corresponding modality pair. To facilitate subsequent credibility evaluation, convert the co-occurrence frequency matrix into a modality co-occurrence frequency vector according to the multi-modal intention weights. This vector is a concise representation of the collaboration degree between multi-modalities. The higher the value, the more likely the modalities belong to the same interaction intention. Through this series of processes, not only the collaborative relationship between modalities is established, but also a data foundation is laid for the subsequent credibility evaluation of interaction intentions, thereby improving the accuracy and robustness of multi-modal fusion decision-making.

[0028] In a possible implementation manner, analyzing the co-occurrence frequency matrix between modality pairs according to the multi-modal interaction intention trigger points includes:

[0029] Obtain the multi-modal interaction intention trigger points, including the trigger points of each modality within a preset time window; form modality pairs, and the modality pairs are obtained by combining all modalities in pairs; set co-occurrence determination conditions, count the co-occurrence times of each modality pair according to the co-occurrence determination conditions, calculate the co-occurrence frequency according to the counted co-occurrence times, and generate a co-occurrence frequency matrix.

[0030] Specifically, after the processing of the obtained multi-modal interaction intention trigger points is completed, the trigger points of each modality included in the multi-modal interaction intention trigger points are within the same preset time window. Subsequently, all modalities are randomly combined in pairs to form multiple modality pairs. For example, an eye movement-head pose modality pair, an eye movement-voice modality pair, and so on. After that, co-occurrence determination conditions are set, and based on these co-occurrence determination conditions, it is judged whether each modality pair co-occurs within the same time window. Among them, the definition of co-occurrence is usually based on temporal proximity, that is, if the time difference between the trigger points of two modalities is less than a preset threshold, it is considered that the trigger points of these two modalities co-occur, which means that they have interacted within a relatively short period of time and may represent the same interaction intention. For each modality pair, the number of times they co-occur in all interaction processes is counted, that is, it is checked whether each pair of modalities meets the co-occurrence determination conditions within the time window, and the number of co-occurrence events is recorded. For example, in the preset time window, the eye movement-voice modality pair is jointly triggered 5 times, so their co-occurrence frequency is 5. Then, the co-occurrence frequency of each pair of modalities is calculated. The co-occurrence frequency refers to the ratio of the number of co-occurrences of a modality pair to the number of its trigger points. After obtaining the co-occurrence frequencies of each pair of modalities, a co-occurrence frequency matrix is generated based on these co-occurrence frequencies. Each element of this matrix represents the co-occurrence frequency between a pair of modalities, and this matrix provides data support for subsequent analysis to help evaluate the degree of cooperation of different modality combinations. In summary, through the above steps, the interaction patterns between different modalities can be effectively analyzed, potential trigger points of user intentions can be identified, and accurate basis can be provided for subsequent credibility evaluation and interaction control.

[0031] In a possible implementation manner, it is judged whether the interaction intention trigger points of each modality pair meet the co-occurrence determination conditions. If they meet the co-occurrence determination conditions, a co-occurrence event is recorded; wherein, the co-occurrence determination condition is that the time difference between the interaction intention trigger points of the modality pair is less than a preset time difference threshold.

[0032] Specifically, for each pair of modalities, first, the interaction intention trigger point data of this modality pair within the preset time window is obtained, including the timestamps of the trigger points. Subsequently, the time difference between the trigger points of the two modalities in each modality pair is calculated, and the calculated time difference is compared with the preset time difference threshold (such as 300 ms). When the time difference between the trigger points of the modality pair is less than the preset time difference threshold, it is considered that these two modalities meet the co-occurrence determination conditions in terms of time, that is, they are triggered synchronously. Conversely, it is considered that these two modalities are not triggered within the same time period and do not meet the co-occurrence conditions. For each modality pair that meets the co-occurrence conditions, a co-occurrence event is recorded, which means that these two modalities have a common interaction intention trigger within the same time period. Repeat this process, conduct the above inspections and statistics on all modality pairs, and record the co-occurrence events, so as to provide data support for subsequent calculation of modality co-occurrence frequencies and credibility evaluation of interaction intentions.

[0033] In a possible implementation, the co-occurrence frequency is calculated according to the statistically obtained co-occurrence times. The co-occurrence frequency is the ratio of the co-occurrence times of the modality pair to the total number of trigger points, and the total number of trigger points is min(|T i |,|T j |); where T i is the number of interaction intention trigger points corresponding to the i-th modality, and T j is the number of interaction intention trigger points corresponding to the j-th modality.

[0034] Specifically, after counting the number of times C i,j that each pair of modalities satisfies the co-occurrence condition during the entire interaction process, the total number of trigger points of the i-th modality and the j-th modality are also counted respectively, denoted as T i and T j . These two values respectively represent the total number of times that interaction intention trigger points occur independently for these two modalities during the entire acquisition time. Subsequently, to avoid frequency calculation imbalance caused by differences in trigger frequencies of different modalities, min(|T i |,|T j |) is used as the normalization base for the co-occurrence frequency of the modality pair. By calculating the ratio of C i,j to min(|T i |,|T j |), the co-occurrence frequency of modality i and modality j is obtained.

[0035] Exemplarily, if the eye movement modality is triggered 50 times, the voice modality is triggered 30 times, and the eye movement and voice modalities co-occur 15 times, then their co-occurrence frequency is 15÷min(50, 30) = 0.5.

[0036] In a possible implementation, to obtain the modality co-occurrence frequency vector, the method includes:

[0037] Construct a multi-modal intention weight matrix; perform matrix calculation on the co-occurrence frequency matrix according to the multi-modal intention weight matrix to obtain the modality co-occurrence frequency vector, which is used to represent the synchronization degree of interaction intention trigger points between modalities.

[0038] Specifically, a multi-modal intention weight matrix is constructed based on prior knowledge and domain experience. This multi-modal intention weight matrix is used to represent the relative importance between different modalities. Each element in the matrix represents the weight between modality i and modality j, and this weight reflects the contribution degree of each pair of modalities to the interaction intention. Subsequently, based on this multi-modal intention weight matrix, matrix calculations are performed on the co-occurrence frequency matrix, that is, the multi-modal intention weight matrix is applied to the co-occurrence frequency matrix using matrix multiplication to obtain a modality co-occurrence frequency vector. This modality co-occurrence frequency vector is obtained through weighted calculation, where each element represents the synchronization degree of the interaction intention between modality i and other modalities, which can reflect the comprehensive cooperation degree between modality pairs and provide data support for the accuracy of interaction control and the anti-mis-touch ability.

[0039] Input the modality co-occurrence frequency vector into the interaction intention credibility evaluation module to calculate the credibility score of the current interaction intention. If the credibility score is greater than the first preset credibility threshold, an interaction control instruction is triggered, and the smart glasses execute the control program according to the interaction control instruction.

[0040] Specifically, after obtaining the modality co-occurrence frequency vector, input it into the interaction intention credibility evaluation module. This module will calculate the data in the modality co-occurrence frequency vector according to the connected fuzzy rule engine, analyze the cooperation degree between different modalities, and obtain a credibility score. The credibility score reflects the degree of recognition of the current interaction intention judgment. The higher the score, the more consistent the trigger points of multiple modalities. Subsequently, compare the calculated credibility score with the preset first credibility threshold. If the calculated credibility score is higher than this threshold, it indicates that the recognition of the current interaction intention is accurate and reliable. At this time, a corresponding interaction control instruction will be triggered, and this instruction will direct the smart glasses to perform specific operations or tasks, such as adjusting the display content, starting a certain application, executing a voice command, etc. Finally, the smart glasses perform actual operations according to the triggered interaction control instruction according to the predetermined control program. In this way, it is possible to make an intelligent response based on the user's multi-modal interaction intention, thereby improving the accuracy and fluency of the interaction experience, while avoiding mis-triggering and improving the reliability and intelligence of the interaction.

[0041] In a possible implementation manner, input the modality co-occurrence frequency vector into the interaction intention credibility evaluation module to calculate the credibility score of the current interaction intention. The method includes:

[0042] Define a fuzzy interval, where the fuzzy interval includes a first co-occurrence frequency region, a second co-occurrence frequency region, and a third co-occurrence frequency region. The first co-occurrence frequency region is a high co-occurrence frequency, the second co-occurrence frequency region is a medium co-occurrence frequency, and the third co-occurrence frequency region is a low co-occurrence frequency. Set a credibility rule based on the fuzzy interval, encapsulate the credibility rule to obtain a fuzzy rule engine, connect it to the interactive intention credibility evaluation module, calculate the credibility score for the modal co-occurrence frequency vector, and output the credibility score result.

[0043] Specifically, according to the co-occurrence frequencies of different modalities, a fuzzy interval is defined. This fuzzy interval is used to divide different co-occurrence frequency values into multiple regions to evaluate the credibility of interaction intentions. The divided fuzzy interval includes a first co-occurrence frequency region, a second co-occurrence frequency region, and a third co-occurrence frequency region. Among them, the first co-occurrence frequency region (high co-occurrence frequency) indicates a strong synchronous collaboration relationship between modality pairs, usually corresponding to a high co-occurrence frequency. A high co-occurrence frequency indicates that the frequency of multiple modalities being triggered simultaneously is very high, and the determination of interaction intentions is relatively reliable. The second co-occurrence frequency region (medium co-occurrence frequency) indicates a medium collaboration relationship between modality pairs, usually corresponding to a medium co-occurrence frequency. A medium co-occurrence frequency means that the modalities are synchronously triggered during some interaction processes, but not as strongly as in the high co-occurrence frequency region. The third co-occurrence frequency region (low co-occurrence frequency) indicates a weak synchronous relationship between modality pairs, usually corresponding to a low co-occurrence frequency. A low co-occurrence frequency indicates that there is less collaborative interaction between modalities, and the recognition of interaction intentions is relatively uncertain. Through this division, different intensities of modality collaboration relationships can be distinguished, thus providing a basis for credibility calculation. Subsequently, credibility rules are set according to the above fuzzy interval. These rules define how to calculate the credibility score according to different co-occurrence frequency regions. Specifically, the credibility score of each modality pair is determined according to the region where its co-occurrence frequency is located. For the first co-occurrence frequency region, the credibility score is relatively high, possibly close to the full score value, indicating that the confirmation of the interaction intention is relatively accurate; for the second co-occurrence frequency region, the credibility score is medium, indicating that the degree of confirmation of the interaction intention is average; for the third co-occurrence frequency region, the credibility score is low, indicating that the degree of confirmation of the interaction intention is weak. These rules help evaluate the relationship between the collaboration degree between modalities and the credibility score. After that, the above-set credibility rules are encapsulated into fuzzy rules, and a fuzzy rule engine is constructed. The fuzzy rule engine is an intelligent decision-making component. It can output the credibility score of the interaction intention based on the input co-occurrence frequency vector according to the fuzzy rules. The engine classifies the co-occurrence frequency values of each modality pair, determines the co-occurrence frequency region it belongs to, and then matches according to the credibility rules of the corresponding region to obtain a credibility score. Then, the fuzzy rule engine is communicatively connected to the interaction intention credibility evaluation module. The interaction intention credibility evaluation module can receive the modality co-occurrence frequency vector from the system and can also encapsulate the credibility score calculated by the fuzzy rule engine into a credibility score result and send it back to the system to reflect the reliability of the current interaction intention, ensuring that the corresponding control instructions can only be triggered under high credibility, and reducing the probability of occurrence of error operation events.

[0044] In a possible implementation manner, after calculating the credibility score of the current interaction intention, the method further includes:

[0045] If the confidence score is less than or equal to the first preset confidence threshold, determine whether the confidence score is greater than a second preset confidence threshold, where the second preset confidence threshold is less than the first preset confidence threshold; if the confidence score is greater than the second preset confidence threshold, send a waiting confirmation instruction to the user wearing the smart glasses, for prompting the user whether to trigger an interaction control instruction.

[0046] Specifically, after the interaction intention confidence evaluation module outputs the confidence score of the current interaction intention, compare the confidence score with the first preset confidence threshold. If the confidence score is less than or equal to the first preset confidence threshold, it means that the system's judgment of the current interaction intention is uncertain and the interaction control instruction cannot be directly triggered. To avoid directly ignoring a potentially valid interaction intention due to slight uncertainty, a second preset confidence threshold is also set, and its value is lower than the first preset confidence threshold. At this time, it will be determined whether the current confidence score is higher than the second preset confidence threshold to further distinguish between potentially valid and obviously invalid interaction intentions. If the confidence score is less than or equal to the second preset confidence threshold, it means that the confidence of this interaction intention is very low and it is very likely to be a mis-triggered behavior, so no response is made and it is directly ignored. If the confidence score is higher than the second preset confidence threshold but lower than the first preset confidence threshold, it means that this interaction intention has a certain possibility. Therefore, the user confirmation mechanism will be entered, that is, a waiting confirmation instruction will be sent to the user wearing the smart glasses, usually presented through visual prompts (such as popping up a confirmation prompt in the lens HUD), voice announcements or vibration prompts, etc. The prompt content may include "A possible interaction intention is detected. Do you confirm to execute?" "Please nod / voice 'yes' / tap to confirm to continue the operation." The user can confirm or cancel the operation through the set methods, such as nodding, voice confirmation (such as saying "yes"), touching the confirmation area, etc. If the user confirms, the corresponding interaction control instruction will be triggered; if the user cancels, or there is no confirmation action within the preset time, the system defaults not to execute this interaction control operation. Through this mechanism, when the system faces interaction intentions with medium confidence, it can avoid mis-triggering, ensure the responsiveness and flexibility of the interaction, and thus achieve a safe, controllable and user-friendly smart glasses interaction experience.

[0047] In the above text, with reference to Figure 1 A method for interaction control of smart glasses according to an embodiment of the present invention has been described in detail. Next, with reference to Figure 2 An interaction control system for smart glasses according to an embodiment of the present invention will be described.

[0048] An interactive control system for smart glasses according to an embodiment of the present invention is used to solve the technical problem that the existing interactive methods of smart glasses rely on a single modality, resulting in a relatively high false trigger rate and unstable user experience, and achieve the technical effect of reducing the misoperation rate of smart glasses and improving the multi-modal interactive experience through multi-modal cross-verification and dynamic intention credibility evaluation. An interactive control system for smart glasses includes: a data acquisition unit 11, a feature extraction unit 12, a modality analysis unit 13, and an interactive control unit 14.

[0049] Data acquisition unit 11: Obtain multi-modal interaction information from the multi-modal data acquisition device, where the multi-modal interaction information includes eye movement trajectories, head pose data, voice input signals, and facial expression changes stored in time series; Feature extraction unit 12: Extract interaction features from the multi-modal interaction information to identify multi-modal interaction intention trigger points; Modality analysis unit 13: Analyze the co-occurrence frequency matrix between modality pairs according to the multi-modal interaction intention trigger points to obtain a modality co-occurrence frequency vector; Interactive control unit 14: Input the modality co-occurrence frequency vector into an interactive intention credibility evaluation module to calculate the credibility score of the current interaction intention. If the credibility score is greater than a first preset credibility threshold, trigger an interactive control instruction, and the smart glasses execute a control program according to the interactive control instruction.

[0050] Next, the specific configuration of the feature extraction unit 12 will be described in detail. As described above, extract interaction features from the multi-modal interaction information to identify multi-modal interaction intention trigger points. The feature extraction unit 12 may further include: storing the multi-modal interaction information in blocks to obtain a plurality of interaction information storage blocks; setting a plurality of interaction feature extraction blocks, where the plurality of interaction feature extraction blocks are correspondingly connected to the plurality of interaction information storage blocks, and the plurality of interaction feature extraction blocks include pre-stored eye movement interaction features, head pose interaction features, voice keyword interaction features, and facial expression interaction features; wherein each interaction feature extraction block is used to receive the modality data of the interaction information storage block for interaction feature extraction and output multi-modal interaction intention trigger points.

[0051] Next, the specific configuration of the modality analysis unit 13 will be described in detail. As described above, analyze the co-occurrence frequency matrix between modality pairs according to the multi-modal interaction intention trigger points. The modality analysis unit 13 may further include: obtaining multi-modal interaction intention trigger points, including trigger points of each modality within a preset time window; forming modality pairs, where the modality pairs are obtained by pairwise combination of all modalities; setting co-occurrence determination conditions, and counting the co-occurrence times of each modality pair according to the co-occurrence determination conditions, and calculating the co-occurrence frequency according to the counted co-occurrence times to generate a co-occurrence frequency matrix.

[0052] Among them, the modal analysis unit 13 may further include: determining whether the interaction intention trigger points of each modal pair satisfy the co-occurrence determination condition, and if so, recording a co-occurrence event; wherein, the co-occurrence determination condition is that the time difference between the interaction intention trigger points of the modal pair is less than a preset time difference threshold.

[0053] Among them, the modal analysis unit 13 may further include: calculating the co-occurrence frequency according to the statistically obtained co-occurrence times, where the co-occurrence frequency is the ratio of the co-occurrence times of the modal pair to the total number of trigger points, and the total number of trigger points is min(|T i |,|T j |); where T i is the number of interaction intention trigger points corresponding to the i-th modality, and T j is the number of interaction intention trigger points corresponding to the j-th modality.

[0054] Among them, to obtain the modal co-occurrence frequency vector, the modal analysis unit 13 may further include: constructing a multi-modal intention weight matrix; performing matrix calculation on the co-occurrence frequency matrix according to the multi-modal intention weight matrix to obtain the modal co-occurrence frequency vector, which is used to represent the synchronization degree of the interaction intention trigger points between modalities.

[0055] Next, the specific configuration of the interaction control unit 14 will be described in detail. As described above, inputting the modal co-occurrence frequency vector into the interaction intention credibility evaluation module to calculate the credibility score of the current interaction intention, the interaction control unit 14 may further include: defining a fuzzy interval, which includes a first co-occurrence frequency region, a second co-occurrence frequency region, and a third co-occurrence frequency region, where the first co-occurrence frequency region is a high co-occurrence frequency, the second co-occurrence frequency region is a medium co-occurrence frequency, and the third co-occurrence frequency region is a low co-occurrence frequency; setting credibility rules based on the fuzzy interval, encapsulating the credibility rules to obtain a fuzzy rule engine, connecting it to the interaction intention credibility evaluation module, calculating the credibility score for the modal co-occurrence frequency vector, and outputting the credibility score result.

[0056] Among them, after calculating the credibility score of the current interaction intention, the interaction control unit 14 may further include: if the credibility score is less than or equal to the first preset credibility threshold, determining whether the credibility score is greater than the second preset credibility threshold, where the second preset credibility threshold is less than the first preset credibility threshold; if the credibility score is greater than the second preset credibility threshold, sending a waiting confirmation instruction to the user wearing the smart glasses to prompt the user whether to trigger an interaction control instruction.

[0057] The interactive control system for smart glasses provided by the embodiments of the present invention can execute the interactive control method for smart glasses provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.

[0058] Although the present application makes various references to certain modules in the system according to the embodiments of the present application, however, any number of different modules can be used and run on the user terminal and / or the server. The included individual units and modules are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of the functional units are only for the convenience of mutual distinction and do not limit the protection scope of the present invention.

[0059] Based on the foregoing embodiments, the embodiments of the present application further provide an electronic device. Figure 3 It is a schematic structural diagram of the electronic device provided by the embodiments of the present invention, showing a block diagram of an exemplary electronic device suitable for implementing the embodiments of the present invention. Figure 3 The shown electronic device is only an example and should not bring any limitation to the functions and usage scope of the embodiments of the present invention. The electronic device is presented in the form of a general computing device, and its components may include but are not limited to a processor 21, a memory 22, an input device 23, and an output device 24. Among them, the processor 21 may be one or more; the memory 22 may include a computer-readable medium and at least one program product, and this program product has a set (at least one) of program modules, and these program modules are configured to execute the functions of the embodiments of the present application.

[0060] The memory 22 shown in the embodiments of the present invention can adopt any combination of one or more computer-readable media; the computer-readable storage medium may be but is not limited to infrared rays, semiconductor systems, devices or components, or any combination of the above, for storing software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to an interactive control method for smart glasses in the embodiments of the present invention. The processor 21 executes various functional applications and data processing of the computer device by running the software programs, instructions, and modules stored in the memory 22, that is, to implement the above-mentioned interactive control method for smart glasses.

[0061] The above specific embodiments do not constitute a limitation on the protection scope of this application. Those skilled in the art should understand that various modifications, combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application shall be included within the protection scope of this application. In some cases, the actions or steps recorded in this application can be executed in a sequence different from that in the embodiments and still achieve the desired results. Additionally, the processes depicted in the drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

Claims

1. An interactive control method for smart glasses, characterized in that, The smart glasses are embedded with a multi-modal data acquisition device, and the method includes: Obtaining multi-modal interaction information by the multi-modal data acquisition device, where the multi-modal interaction information includes eye movement trajectories, head pose data, voice input signals, and facial expression changes stored in time series; Performing interaction feature extraction on the multi-modal interaction information to identify multi-modal interaction intention trigger points; Analyzing the co-occurrence frequency matrix between modality pairs according to the multi-modal interaction intention trigger points to obtain a modality co-occurrence frequency vector; Inputting the modality co-occurrence frequency vector into an interaction intention credibility evaluation module to calculate the credibility score of the current interaction intention. If the credibility score is greater than a first preset credibility threshold, an interaction control instruction is triggered, and the smart glasses execute a control program according to the interaction control instruction.

2. The method according to claim 1, wherein Performing interaction feature extraction on the multi-modal interaction information to identify multi-modal interaction intention trigger points, and the method includes: Storing the multi-modal interaction information in blocks to obtain a plurality of interaction information storage blocks; Setting a plurality of interaction feature extraction blocks, where the plurality of interaction feature extraction blocks are correspondingly connected to the plurality of interaction information storage blocks, and the plurality of interaction feature extraction blocks include pre-stored eye movement interaction features, head pose interaction features, voice keyword interaction features, and facial expression interaction features; Wherein, each interaction feature extraction block is used to receive the modality data of the interaction information storage block for interaction feature extraction and output multi-modal interaction intention trigger points.

3. The method according to claim 1, characterized in that Analyzing the co-occurrence frequency matrix between modality pairs according to the multi-modal interaction intention trigger points, and the method includes: Obtaining multi-modal interaction intention trigger points, including trigger points of each modality within a preset time window; Forming modality pairs, where the modality pairs are obtained by pairwise combining all modalities; Setting co-occurrence determination conditions, counting the co-occurrence times of each modality pair according to the co-occurrence determination conditions, calculating the co-occurrence frequency according to the counted co-occurrence times, and generating a co-occurrence frequency matrix.

4. The method according to claim 3, wherein Judging whether the interaction intention trigger points of each modality pair meet the co-occurrence determination conditions. If they meet the co-occurrence determination conditions, record a co-occurrence event; Wherein, the co-occurrence determination condition is that the time difference between the interaction intention trigger points of the modality pair is less than a preset time difference threshold.

5. The method according to claim 3, wherein Calculate the co-occurrence frequency according to the statistically obtained co-occurrence times. The co-occurrence frequency is the ratio of the co-occurrence times of the modality pair to the total number of trigger points, and the total number of trigger points is min(|T i |,|T j |); Among them, T i is the number of interaction intention trigger points corresponding to the i-th modality, and T j is the number of interaction intention trigger points corresponding to the j-th modality.

6. The method according to claim 3, wherein Obtaining a modality co-occurrence frequency vector, and the method includes: Constructing a multi-modal intention weight matrix; Performing matrix calculation on the co-occurrence frequency matrix according to the multi-modal intention weight matrix to obtain the modality co-occurrence frequency vector, which is used to represent the synchronization degree of interaction intention trigger points between modalities.

7. The method according to claim 1, wherein Inputting the modality co-occurrence frequency vector into an interaction intention credibility evaluation module to calculate the credibility score of the current interaction intention, and the method includes: Defining a fuzzy interval, where the fuzzy interval includes a first co-occurrence frequency region, a second co-occurrence frequency region, and a third co-occurrence frequency region. The first co-occurrence frequency region is a high co-occurrence frequency, the second co-occurrence frequency region is a medium co-occurrence frequency, and the third co-occurrence frequency region is a low co-occurrence frequency; Based on the fuzzy interval, credibility rules are set, and a fuzzy rule engine is encapsulated with the credibility rules. The fuzzy rule engine is connected to the interactive intention credibility evaluation module to calculate the credibility score of the modal co-occurrence frequency vector and output the credibility score result.

8. The method according to claim 1, wherein After calculating the credibility score of the current interactive intention, the method further includes: If the credibility score is less than or equal to the first preset credibility threshold, it is determined whether the credibility score is greater than the second preset credibility threshold, where the second preset credibility threshold is less than the first preset credibility threshold; If the credibility score is greater than the second preset credibility threshold, a waiting confirmation instruction is sent to the user wearing the smart glasses to prompt the user whether to trigger an interactive control instruction.

9. An interactive control system for smart glasses, characterized in that, The smart glasses are embedded with a multi-modal data acquisition device. The system is used to implement an interactive control method for smart glasses according to any one of claims 1-8. The system includes: Data acquisition unit: Obtain multi-modal interaction information from the multi-modal data acquisition device. The multi-modal interaction information includes eye movement trajectories, head pose data, voice input signals, and facial expression changes stored in time series; Feature extraction unit: Extract interaction features from the multi-modal interaction information to identify the trigger points of multi-modal interaction intentions; Modal analysis unit: Analyze the co-occurrence frequency matrix between modal pairs according to the trigger points of the multi-modal interaction intention to obtain a modal co-occurrence frequency vector; Interactive control unit: Input the modal co-occurrence frequency vector into the interactive intention credibility evaluation module to calculate the credibility score of the current interactive intention. If the credibility score is greater than the first preset credibility threshold, an interactive control instruction is triggered, and the smart glasses execute the control program according to the interactive control instruction.

10. An electronic device, characterized in that, The electronic device includes: A memory for storing executable instructions; A processor for implementing an interactive control method for smart glasses according to any one of claims 1 to 8 when executing the executable instructions stored in the memory.

Citation Information

Patent Citations

  • Multi-modal intention recognition method and device, electronic equipment and storage medium

    CN115618270A

  • Interaction system and method based on intelligent glasses and intelligent glasses

    CN119045770A

  • Multi-modal modeling of temporal interaction sequences

    US20140212853A1

  • Data anlaysis of interaction information for interacting participants to determine a connection intention

    US20250118075A1

  • Smart glasses, and interaction method and apparatus thereof

    WO2022037355A1

Cited By

  • Intelligent dialogue wake-up control method, dialogue system, intelligent device and storage medium

    CN120748400A

  • AI intelligent glasses and interaction method thereof

    CN121478127A