Intelligent glasses, interaction method and device of intelligent glasses, electronic equipment and medium

By integrating perception modules, decision-making modules, execution modules and feedback modules in smart glasses, the perception and autonomous decision-making of multi-dimensional environmental data is achieved, and the problem of intelligent wearable devices being one-sided and passive in the environment is solved, improving the level of intelligence and user experience.

CN120508207APending Publication Date: 2025-08-19BEIJING INSTITUTE OF GRAPHIC COMMUNICATION
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510575651.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-08-19

AI Technical Summary

Technical Problem

Smart wearable devices have a relatively one-sided understanding of the environment, passive interaction, and insufficient intelligence level, which cannot provide convenience to users and affect normal use.

Method used

The combination of perception module, decision module, execution module and feedback module is adopted to realize independent decision-making and active interaction through the perception, prediction of behavioral intentions and generation of multi-dimensional environmental data and user behavior data.

Benefits of technology

It realizes the active perception and independent learning of the environment by smart wearable devices, can adapt to environmental changes like humans, provide a highly personalized intelligent interactive experience, and improve the level of intelligence.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120508207A_ABST
    Figure CN120508207A_ABST
Patent Text Reader

Abstract

The invention relates to intelligent glasses, an interaction method and device of the intelligent glasses, electronic equipment and a medium. The sensing module, the decision making module, the execution module and the feedback module are arranged on the glasses body so as to sense multi-dimensional environment data of a surrounding environment and current behavior data of a user in a preset intelligent interaction mode, predict a behavior intention of the user, track a first behavior change of the user and a first environment change of the surrounding environment and send the first behavior change to the feedback module; and determining an instruction chain meeting a preset execution condition, executing the instruction chain to track a second behavior change of the user based on the instruction chain, and optimizing a generation algorithm of the instruction chain based on the second behavior change. Therefore, the technical problems that in the related technology, the intelligent wearable device understands the environment in a one-sided mode, interaction is passive, the intelligent wearable device can only operate according to a set rule, the intelligent level is insufficient, convenience cannot be provided for a user, and normal use of the user is easily affected are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of smart wearable devices, and in particular to smart glasses, a method and device for interacting with smart glasses, an electronic device, and a medium. Background Art

[0002] In related technologies, smart wearable devices mostly rely on preset rules or single modal input, and have defects such as one-sided understanding of the environment, passive interaction, lack of active perception of the environment, autonomous learning of user needs, and the ability to make intelligent decisions based on dynamic changes in the environment. It is difficult to achieve closed-loop learning of "perception-action" and the interaction data between physical entities and the environment are not fully utilized, resulting in insufficient intelligence level of smart wearable devices. The intelligent interaction they generate does not meet user needs, which not only fails to provide convenience for users, but also easily affects users' normal use, and urgently needs improvement. Summary of the Invention

[0003] The present application provides smart glasses, an interaction method, an apparatus, an electronic device, and a medium for smart glasses to solve technical problems in related technologies, such as that smart wearable devices have a relatively one-sided understanding of the environment, interact passively, can only operate according to established rules, and have insufficient intelligence, which not only fails to provide convenience to users but also easily affects users' normal use.

[0004] A first aspect of the present application provides a pair of smart glasses, comprising: a glasses body; a perception module arranged on the glasses body, for perceiving multidimensional environmental data of a surrounding environment and current behavior data of a user in a preset intelligent interaction mode; a decision module arranged on the glasses body, for predicting the user's behavioral intention based on the multidimensional environmental data, the current behavior data and historical event data, so as to generate a multimodal decision path, and track a first behavioral change of the user and a first environmental change of the surrounding environment, so as to determine an instruction chain that meets preset execution conditions in combination with the first behavioral change, the first environmental change and the multimodal decision path; an execution module arranged on the glasses body, for executing the instruction chain; and a feedback module arranged on the glasses body, for tracking a second behavioral change of the user based on the instruction chain, and optimizing a generation algorithm of the instruction chain based on the second behavioral change.

[0005] Optionally, in one embodiment of the present application, the perception module includes: a camera for acquiring environmental image data of the surrounding environment and pupil tracking data of the user; a microphone for collecting sound data of the surrounding environment and voice command data of the user; an ambient temperature sensor for collecting ambient temperature data of the surrounding environment; a human body temperature sensor for collecting body temperature data of the user; and a positioning unit for receiving positioning data of the user.

[0006] Optionally, in one embodiment of the present application, the perception module includes: a visual processing unit, used to obtain the pupil positioning result of the user based on the environmental image data and the pupil tracking data, so as to use the pupil positioning result to determine the user's attention level to the environmental information of the surrounding environment, and use the attention level to construct the user's attention map; a semantic segmentation unit, used to fuse the multi-dimensional environmental data to semantically align the environmental image data, the sound data and the voice command data, and perform semantic segmentation based on the aligned environmental image data, the sound data and the voice command data to construct environmental semantics.

[0007] Optionally, in one embodiment of the present application, the decision module includes: a prediction unit for fusing the historical event data, the attention map and the environmental semantics to predict the user's behavioral intention to generate the multimodal decision path; a triggering event unit for tracking the user's first behavioral change and the first environmental change of the surrounding environment to combine the first behavioral change, the first environmental change and the multimodal decision path to determine an instruction chain that meets preset execution conditions.

[0008] Optionally, in one embodiment of the present application, the execution module includes: a speaker, used to convert the voice execution instructions in the instruction chain into voice information and play the voice information; a display unit, used to convert the display execution instructions in the instruction chain into display content and display the display content.

[0009] Optionally, in one embodiment of the present application, the feedback module includes: a receiving unit, used to receive the second behavior change and the second environmental change of the surrounding environment after the execution module executes any instruction in the instruction chain; an optimization unit, used to optimize the semantic segmentation accuracy of the perception module, the reward function weight of the decision module, and the information density threshold of the execution module based on the second behavior change, the second environmental change and the operation success rate of any instruction.

[0010] A second aspect of the present application provides an interaction method for smart glasses, comprising the following steps: in a preset smart interaction mode, perceiving multidimensional environmental data of the surrounding environment and current behavior data of the user; predicting the user's behavioral intention based on the multidimensional environmental data, the current behavior data, and historical event data to generate a multimodal decision path, and tracking the user's first behavioral change and the first environmental change of the surrounding environment, so as to determine an instruction chain that meets preset execution conditions in combination with the first behavioral change, the first environmental change, and the multimodal decision path; executing the instruction chain, tracking the user's second behavioral change based on the instruction chain, and optimizing the instruction chain generation algorithm based on the second behavioral change.

[0011] The third aspect of the present application provides an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the interaction method of the smart glasses as described in the above embodiment.

[0012] A fourth aspect of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable the computer to execute the interaction method of the smart glasses as described in the above embodiment.

[0013] The fifth aspect of the present application provides a computer program product, including a computer program, which, when executed, is used to implement the above-mentioned smart glasses interaction method.

[0014] The embodiments of the present application may include a pair of glasses and a perception module, a decision module, an execution module, and a feedback module disposed on the pair of glasses. In a preset intelligent interaction mode, the pair of glasses perceives multi-dimensional environmental data of the surrounding environment and the user's current behavior data, predicts the user's behavioral intention, and tracks the user's first behavioral changes and the first environmental changes of the surrounding environment to determine a command chain that meets the preset execution conditions, and executes the command chain to track the user's second behavioral changes based on the command chain, and optimizes the command chain generation algorithm based on the second behavioral changes. Intelligent growth is achieved through the interaction between the physical entity and the environment, enabling the wearable device to perceive, understand, and adapt to environmental changes like humans. It can not only actively provide interaction, but also autonomously learn based on the user's behavior to adapt to the usage habits of different users, with a higher level of intelligence. This solves the technical problem in the related art that smart wearable devices have a relatively one-sided understanding of the environment, passive interaction, and can only operate according to established rules. The insufficient intelligence level not only fails to provide convenience to users, but also easily affects the user's normal use.

[0015] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:

[0017] Figure 1 A schematic structural diagram of smart glasses provided according to an embodiment of the present application;

[0018] Figure 2 This is a schematic structural diagram of smart glasses provided according to one embodiment of the present application;

[0019] Figure 3 This is a schematic structural diagram of smart glasses provided according to another embodiment of the present application;

[0020] Figure 4 A schematic diagram of an optimization principle provided according to an embodiment of the present application;

[0021] Figure 5 A schematic diagram of the autonomous learning principle of smart glasses provided according to one embodiment of the present application;

[0022] Figure 6 A schematic diagram of the functional framework of smart glasses provided according to one embodiment of the present application;

[0023] Figure 7 This is a flowchart of an interaction method for smart glasses provided according to an embodiment of the present application;

[0024] Figure 8 A schematic diagram of the structure of an electronic device provided according to an embodiment of the present application.

[0025] Among them, 1000-smart glasses, 100-glasses body, 101-frame, 102-lens, 103-left temple, 104-right temple, 105-nose pad, 200-perception module, 201-environmental perception camera, 202-pupil perception camera, 203-microphone, 204-environmental temperature sensor, 205-human body temperature sensor, 206-positioning unit, 207-capacitive gesture sensor, 300-decision module, 400-execution module, 401-speaker, 402-adaptive optical adjustment unit, 500-feedback module, 600-piezoelectric nanogenerator; 801-memory, 802-processor, 803-communication interface. DETAILED DESCRIPTION

[0026] The following describes in detail embodiments of the present application, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.

[0027] The following describes smart glasses, interaction methods, devices, electronic devices, and media for smart glasses according to embodiments of the present application with reference to the accompanying drawings. In response to the technical issues mentioned in the above background technology, in which smart wearable devices have a relatively one-sided understanding of the environment and passive interaction, and can only operate according to established rules, the level of intelligence is insufficient, not only failing to provide convenience to users, but also easily affecting the normal use of users, the present application provides smart glasses, which may include a glasses body and a perception module, a decision module, an execution module, and a feedback module disposed on the glasses body, to perceive multi-dimensional environmental data of the surrounding environment and the user's current behavior data in a preset intelligent interaction mode, predict the user's behavioral intention, and track the user's first behavioral change and the first environmental change of the surrounding environment to determine an instruction chain that meets the preset execution conditions, and execute the instruction chain to track the user's second behavioral change based on the instruction chain, and optimize the instruction chain generation algorithm based on the second behavioral change, so as to achieve intelligent growth through the interaction between the physical entity and the environment, so that the wearable device can perceive, understand, and adapt to environmental changes like humans, and can not only actively provide interaction, but also autonomously learn based on the user's behavior to adapt to the usage habits of different users, with a higher level of intelligence. This solves the technical problem in related technologies that smart wearable devices have a relatively one-sided understanding of the environment, passive interaction, can only operate according to established rules, and have insufficient intelligence, which not only fails to provide convenience to users but also easily affects users' normal use.

[0028] Specifically, Figure 1 A schematic diagram of the structure of smart glasses provided in an embodiment of the present application.

[0029] like Figure 1 As shown, the smart glasses 1000 include: a glasses body 100, a perception module 200, a decision module 300, an execution module 400 and a feedback module 500.

[0030] Specifically, if Figure 2 and Figure 3 As shown, the glasses body 100 may include: a frame 101, a lens 102, a left temple 103, a right temple 104 and a nose pad 105, wherein the frame 101 is used to fix the lens 102, and the left temple 103 and the right temple 104 are respectively connected to both sides of the frame 101 for wearing on the user's head.

[0031] The perception module 200, decision module 300, execution module 400 and feedback module 500 are integrated on the glasses body 100 to realize the perception of the environment, the generation of decisions, the execution of decisions and the optimization of decision algorithms based on the environment and user behavior.

[0032] The perception module 200 provided on the eyeglass body 100 is used to perceive the multi-dimensional environmental data of the surrounding environment and the user's current behavior data in a preset intelligent interaction mode.

[0033] The perception module 200 provided on the glasses body 100 can realize the perception of multi-dimensional environmental data and user behavior data.

[0034] First, the smart glasses 1000 of the embodiment of the present application can determine whether it has entered the smart interaction mode, for example, by setting a hard button on the glasses body 100, and determining whether to enter the smart interaction mode based on whether the hard button is triggered; receiving the user's mode-on voice command through the perception module 200 to enter the smart interaction mode; receiving the user's gesture execution through the perception module 200 to enter the smart interaction mode.

[0035] Optionally, in one embodiment of the present application, the perception module 200 includes: a camera, a microphone, an ambient temperature sensor, a human body temperature sensor and a positioning unit.

[0036] Among them, the camera is used to obtain environmental image data of the surrounding environment and the user's pupil tracking data.

[0037] Microphone, used to collect sound data from the surrounding environment and user's voice command data.

[0038] Ambient temperature sensor, used to collect ambient temperature data of the surrounding environment.

[0039] Human body temperature sensor, used to collect user's body temperature data.

[0040] The positioning unit is used to receive the user's positioning data.

[0041] Combine Figure 2 and Figure 3 As shown, the perception module 200 may include: cameras (environmental perception camera 201 and pupil perception camera 202 ), a microphone 203 , an ambient temperature sensor 204 , a human body temperature sensor 205 and a positioning unit 206 .

[0042] Among them, the environmental perception cameras 201 are located at the left and right ends of the outside of the frame 101. They can use the VIMA multimodal model to deconstruct the functional semantics of objects in the scene in real time, such as "sittable chairs" and "avoidable obstacles", and identify objects, people and actions in the scene in real time, construct a three-dimensional scene model, and predict the impact of dynamic changes in the environment on actions based on the 3D-VLA model, so that the glasses can understand the surrounding environment like humans.

[0043] The pupil sensing camera 202 is located in the middle of the inner side of the frame 101. It adopts a neuromorphic vision chip and uses the EyeTrAES algorithm to achieve high-fidelity and low-latency pupil tracking. It uses a large brain-like model to analyze gaze patterns, maps pupil positioning results to screen coordinates, calculates the length of stay in a specific area, and accurately collects the user's attention to environmental information. Combined with the PaLM-E model, it constructs a "first-person" perspective attention map to provide key data of user interest points for the decision module 300.

[0044] Microphone 203, located on the inside of the left temple 103, utilizes the Group Relative Policy Optimization (GRPO) algorithm to optimize the built-in DSP chip through reinforcement learning, thereby continuously collecting target scene environmental information and user voice commands. Microphone 203 incorporates a multi-channel beamforming array and uses the SayCan framework to semantically align voice commands with scene context.

[0045] The ambient temperature sensor 204 is located on the right outer side of the frame 101 and adopts the MLX90614 miniature infrared non-contact sensor. It uses the adaptive filtering LMS algorithm to adjust the filter coefficient in real time to suppress transient noise caused by reasons such as the evaporation of user sweat, accurately collect ambient temperature information, and provide ambient temperature parameters for the decision module 300.

[0046] The human body temperature sensor 205 is located at the nose pad 105 and is implanted in the flexible electronic skin. A meta-learning algorithm is used to establish an individualized body temperature baseline model, thereby eliminating the influence of ambient temperature on skin surface temperature measurement, approaching core body temperature or stabilizing epidermal temperature, and providing user body temperature data to the decision module 300, so that the glasses can make corresponding decisions based on changes in user body temperature.

[0047] The positioning unit 206, such as a GPS receiver, is located at the end of the right temple 104. The metal part of the temple is used as an antenna extension to reduce signal obstruction and feed back the user's location information to the decision module 300 in real time. After processing and response by the decision module 300, the corresponding decision is finally autonomously executed by the execution module 400.

[0048] In addition, the perception module 200 may further include a hand sensing unit, such as a capacitive gesture sensor 207, to obtain user gesture information.

[0049] The perception module 200 may further include a communication unit to receive information from other smart terminals associated with the smart glasses 1000 to associate the information from other smart terminals with the user's current motion data and environmental data.

[0050] Optionally, in one embodiment of the present application, the perception module 200 includes: a visual processing unit and a semantic segmentation unit.

[0051] Among them, the visual processing unit is used to obtain the user's pupil positioning results based on environmental image data and pupil tracking data, so as to use the pupil positioning results to determine the user's attention level to the environmental information of the surrounding environment, and use the attention level to construct the user's attention map.

[0052] The semantic segmentation unit is used to fuse multi-dimensional environmental data to semantically align environmental image data, sound data and voice command data, and perform semantic segmentation based on the aligned environmental image data, sound data and voice command data to construct environmental semantics.

[0053] The visual processing unit can be used to perform data fusion and pupil positioning on environmental image data and pupil tracking data to determine the user's line of sight focus. For example, the visual processing unit can use the RGB-D information image of the surrounding environment scene captured in real time by the environmental perception camera 201 to deconstruct the functional semantics of objects in the scene and construct a three-dimensional scene model. It can then combine the gaze heat map of the user's pupils in the environment tracked by the pupil perception camera 202 to determine the user's attention level to specific information to obtain the user's attention map.

[0054] The semantic segmentation unit can align the timestamps of multi-source data (such as audio and video synchronization) and achieve spatial alignment through coordinate transformation (such as sound source positioning and visual target matching). It then performs semantic segmentation based on the aligned environmental image data, sound data, and voice command data to construct environmental semantics. Through cross-modal semantic understanding and precise segmentation, it provides the smart glasses 1000 in complex environments with the ability to build interpretable semantic maps.

[0055] The decision module 300 provided on the eyeglass body 100 is used to predict the user's behavioral intention based on multi-dimensional environmental data, current behavior data and historical event data to generate a multimodal decision path, and track the user's first behavioral change and the first environmental change of the surrounding environment, so as to combine the first behavioral change, the first environmental change and the multimodal decision path to determine an instruction chain that meets the preset execution conditions.

[0056] Furthermore, the decision module 300 of the embodiment of the present application can be based on the multi-dimensional environmental information obtained by the perception module 200, and through deep learning algorithms and reinforcement learning strategies, it can integrate user historical behavior data and real-time perception information to achieve a fully closed-loop autonomous decision-making from "perception-reasoning-action".

[0057] During the actual execution process, the decision module 300 can match the historical event data based on the multi-dimensional environmental data and the user's current behavior data. For example, in the waiting room scenario, the user's current behavior is to pay attention to the train rear car sign, and the historical event data is the event of triggering the attention sign under similar waiting room scenarios and historical behaviors. The user's behavioral intention can be predicted, such as amplifying the sign information, interpreting the sign information, extracting key words from the sign information, etc., to generate a multimodal decision path. Furthermore, the embodiment of the present application can track the user's behavioral changes, such as whether the user moves away from the sign in a short period of time, whether the display content of the sign switches to the next page, whether there is reserved ticket information in other smart terminals associated with the smart glasses 1000, etc., and then determine the executable instruction chain from the multimodal decision path according to the behavioral changes and environmental changes.

[0058] For example, when the decision-making scenario occurs in a park, and the historical event is that the user took a leisure run in the park, the user's behavior changes are tracked to confirm whether the user has entered the historical running area and / or whether the environment is suitable for running (such as whether it is raining or whether the ambient temperature is too high), etc., and then a corresponding instruction chain is generated.

[0059] For example, when the decision-making scenario occurs in the office, the historical event is that the user remotely controls the start-up of the home smart home during off-duty hours. At this time, the user's behavior changes are tracked to confirm whether the user enters the parking area or enters the off-duty scene, etc., and then the corresponding instruction chain is generated to realize the remote start-up of the smart home.

[0060] Among them, historical event data includes the user's historical behavior, the multi-dimensional environment corresponding to the historical behavior, and the event ultimately triggered by the historical behavior.

[0061] Optionally, in one embodiment of the present application, the decision module 300 includes: a prediction unit and a trigger event unit.

[0062] Among them, the prediction unit is used to fuse historical event data, attention maps and environmental semantics to predict the user's behavioral intention and generate a multimodal decision path.

[0063] The trigger event unit is used to track the user's first behavior change and the first environmental change of the surrounding environment, so as to determine an instruction chain that meets the preset execution conditions by combining the first behavior change, the first environmental change and the multimodal decision path.

[0064] The decision module 300 may include a prediction unit. Based on the multi-dimensional environmental information acquired by the perception module 200, the prediction unit utilizes deep learning algorithms and reinforcement learning strategies to integrate historical user behavior data with real-time perception information. This unit then uses a world model to predict user behavior and environmental changes in real time, generating a multimodal decision path. For example, by analyzing the user's pupil dwell time and historical behavior data, the prediction unit can predict information that the user may be interested in or an upcoming action.

[0065] The decision-making module 300 may include a trigger event unit. Based on preset rules and real-time perception data, the trigger event unit maps environmental semantics into an executable instruction chain, enabling a fully closed-loop autonomous decision-making process from "perception-reasoning-action." When a user enters a specific area or receives a specific voice command, the relevant function is automatically activated, enabling autonomous decision-making and proactive response, providing an intelligent interactive experience for the user.

[0066] It is understandable that the instruction chain is not just asking whether to execute event A. If not, it jumps to event B, and if not, it jumps to event C. It can also include some conditional nodes. When the user's behavior changes or environmental changes are obtained, and when the behavior changes or environmental changes reach a certain node, such as when the user's body temperature or ambient temperature is abnormal during exercise, a conditional node is inserted to generate an exercise termination suggestion.

[0067] The execution module 400 disposed on the eyeglass body 100 is used to execute the instruction chain.

[0068] The execution module 400 can be used to execute the instruction chain, for example, displaying detailed information of the items that the user is interested in, taking photos and storing the objects that the user is interested in, issuing instruction inquiries to the user, etc.

[0069] Optionally, in one embodiment of the present application, the execution module 400 includes: a speaker and a display unit.

[0070] The speaker is used to convert the voice execution instructions in the instruction chain into voice information and play the voice information.

[0071] The display unit is used to convert the display execution instruction in the instruction chain into display content and display the display content.

[0072] like Figure 2 and Figure 3As shown, the execution module 400 uses a multimodal interaction engine to convert the decision results into adaptive display content, which is then broadcasted via speaker 401 and displayed visually using lens 102 as a display unit, presenting the results to the user in an intuitive manner. For example, when the decision module 300 predicts that the user may need navigation information, lens 102 automatically displays navigation arrows, while speaker 401 broadcasts navigation instructions, allowing the user to obtain the required information without manual operation.

[0073] The feedback module 500 provided on the eyeglass body 100 is used to track the second behavior change of the user based on the instruction chain, and optimize the generation algorithm of the instruction chain based on the second behavior change.

[0074] The feedback module 500 can track the user's behavior changes when the execution module 400 executes the instruction chain to determine whether the user is satisfied with the current decision, thereby optimizing the instruction chain generation algorithm.

[0075] For example, when the execution module 400 displays an inquiry about whether detailed information of a vase needs to be provided, the embodiment of the present application can determine, based on the user's pupil tracking, whether the user has moved his eyes away from the vase or whether the user's perspective is focused on the "yes" or "no" of the inquiry, or, based on the user's voice reply, to obtain the user's response to the inquiry, or, based on the user's gesture information, to obtain the user's response to the inquiry, so that when the user makes an affirmative decision, the current behavior, environmental data, and triggered event (querying for detailed information of the vase) are stored in the historical event data, and when the user makes a negative decision, corresponding optimization is performed.

[0076] Optionally, in one embodiment of the present application, the feedback module 500 includes: a receiving unit and an optimization unit.

[0077] The receiving unit is configured to receive a second behavior change and a second environmental change of the surrounding environment after the execution module 400 executes any instruction in the instruction chain.

[0078] The optimization unit is used to optimize the semantic segmentation accuracy of the perception module 200, the reward function weight of the decision module 300, and the information density threshold of the execution module 400 based on the second behavior change, the second environment change, and the operation success rate of any instruction.

[0079] The receiving unit can communicate data with the perception module 200 so as to obtain the user's behavior changes and environmental changes through the perception module 200 after the execution module 400 executes any instruction in the instruction chain.

[0080] The optimization unit can optimize the semantic segmentation accuracy of the perception module, the reward function weight of the decision module, and the information density threshold of the execution module 400 in real time, achieving continuous intelligent growth from "experience feedback-strategy iteration-capability evolution".

[0081] To sum up, in the perception stage, the embodiment of the present application collects multi-dimensional environmental information through the perception module 200, including scene information collected by the environmental perception camera 201, user pupil stay information collected by the pupil perception camera 202, voice information collected by the microphone 203, ambient temperature information collected by the ambient temperature sensor 204, user body temperature information collected by the human body temperature sensor 205, and user location information collected by the positioning unit 206.

[0082] Decision-making stage: Based on the multi-dimensional environmental information acquired by the perception module 200, the decision-making module 300 makes autonomous decisions through the prediction unit and the triggering event unit. The prediction unit uses deep learning algorithms and reinforcement learning strategies to integrate historical user behavior data with real-time perception information to generate a multimodal decision path. The triggering event unit maps environmental semantics into an executable instruction chain based on preset rules and real-time perception data.

[0083] Execution phase: The execution module 400 converts the decision result of the decision module 300 into adaptive display content through the multimodal interaction engine, and performs voice broadcast through the speaker 401 and image display through the lens 102.

[0084] Feedback stage: The feedback module 500 receives feedback signals such as pupil gaze duration, voice interaction frequency, and environmental operation success rate, and performs real-time optimization of the semantic segmentation accuracy of the perception module 200, the reward function weight of the decision module 300, and the information density threshold of the execution module 400, to achieve continuous intelligent growth from "experience feedback-strategy iteration-capability evolution".

[0085] Combine Figures 2 to 6 As shown, the interaction between the smart glasses and the smart glasses in the embodiment of the present application is described in detail using an embodiment.

[0086] like Figure 4 As shown, in the relevant technologies, smart glasses have certain problems, such as integrating different AI algorithms, using customized algorithms to complete predetermined task decisions, intermittently executing tasks, issuing tasks in different ways, relying on preset rules or single modal input, and having defects such as one-sided understanding of the environment, passive interaction, lack of active perception of the environment, autonomous learning of user needs, and the ability to make intelligent decisions based on dynamic changes in the environment.

[0087] Based on the above-mentioned defects, the embodiments of the present application can make corresponding improvements, such as realizing unified processing of multimodal information through large models, simulating human thinking to complete complex task decisions, continuously and dynamically generating intelligent actions, and making autonomous decisions, so as to realize active perception of the environment, autonomous learning and intelligent decision-making, and provide users with a highly personalized and intelligent machine-active interaction experience, thereby promoting the innovative application of embodied intelligence technology in the field of wearable devices.

[0088] like Figure 2 and Figure 3 As shown, in this embodiment of the present application, the eyeglass body 100 includes: a frame 101, lenses 102, left temples 103, right temples 104, and nose pads 105. The frame 101 is used to fix the lenses 102. The left temples 103 and right temples 104 are respectively connected to the sides of the frame 101 for wearing on the user's head. The lenses 102 are used for normal visual observation and displaying environmental interaction information. The core of the system lies in the integration of four modules: the perception module 200, the decision module 300, the execution module 400, and the feedback module 500, to construct a complete embodied intelligent system. In addition, in this embodiment of the present application, a piezoelectric nanogenerator 600 is also provided on the eyeglass body 100 to provide power for the device.

[0089] Among them, the perception module 200 collects multi-dimensional environmental information through multiple sensors, providing a comprehensive perception basis for autonomous decision-making.

[0090] Decision module 300, the core functional area of intelligent decision-making, includes a prediction unit and a triggering event unit. Based on the multi-dimensional environmental information acquired by perception module 200, decision module 300 integrates historical user behavior data with real-time perception information through deep learning algorithms and reinforcement learning strategies, achieving a fully closed-loop autonomous decision-making process from "perception-reasoning-action."

[0091] The execution module 400 uses a multimodal interaction engine to convert the decision results into adaptive display content, which is then presented to the user in an intuitive manner through a voice broadcast via the speaker 401 and a visual display via the lens 102. For example, if the decision module predicts that the user may need navigation information, the lens 102 will automatically display a navigation arrow, while the speaker 401 will simultaneously broadcast navigation instructions, allowing the user to obtain the required information without manual operation.

[0092] The feedback module 500 builds an embodied intelligent closed-loop learning system, receives feedback signals such as pupil gaze duration, voice interaction frequency, and environmental operation success rate, and performs real-time optimization of the semantic segmentation accuracy of the perception module 200, the reward function weight of the decision module 300, and the information density threshold of the execution module 400, thereby achieving continuous intelligent growth from "experience feedback-strategy iteration-capability evolution" to improve adaptability to the environment and the level of intelligence.

[0093] like Figure 5 and Figure 6 As shown, the smart glasses 1000 of the embodiment of the present application can be used as follows in actual application:

[0094] First, the perception module 200 performs multimodal data collection, and the environmental perception camera 201 captures the RGB-D information image of the surrounding scene in real time, deconstructs the functional semantics of objects in the scene, and constructs a three-dimensional scene model.

[0095] At the same time, the pupil sensing camera 202 tracks the gaze heat map of the user's pupils in the environment and accurately monitors the user's attention level to specific information.

[0096] At the same time, the microphone 203 continuously collects target scene environment information and user voice commands, and realizes semantic alignment of voice commands and scene context through the SayCan framework.

[0097] At the same time, the ambient temperature sensor 204 collects ambient temperature information and uses the adaptive filtering LMS algorithm to suppress transient noise.

[0098] At the same time, the human body temperature sensor 205 collects the user's body temperature data and establishes an individualized body temperature baseline model.

[0099] At the same time, the positioning unit 206, such as a GPS receiver, obtains the user's location information in real time and determines the user's motion status in combination with the IMU data.

[0100] At the same time, the capacitive hand sensor 207 can be used to collect the user's gestures to determine the user's needs based on the gesture actions.

[0101] After completing the multimodal data collection, the collected multimodal data (vision, hearing, temperature, location, etc.) will be fused through deep learning algorithms to build a rich perception environment and provide a comprehensive perception basis for autonomous decision-making.

[0102] Next, the prediction unit uses deep learning algorithms and reinforcement learning strategies based on the multi-dimensional environmental information obtained by the perception module 200, integrates the user's historical behavior data with real-time perception information, and makes real-time predictions of user behavior and environmental changes through world model preview to generate a multimodal decision path.

[0103] At the same time, the triggering event unit maps environmental semantics into an executable instruction chain based on preset rules and real-time perception data, achieving a fully closed-loop autonomous decision-making process from "perception-reasoning-action." When a user enters a specific area or receives a specific voice command, the relevant function is automatically activated, achieving autonomous decision-making and proactive response.

[0104] Then, the execution module 400 converts the decision result of the decision module 300 into adaptive display content through the multimodal interaction engine. When the decision module 300 predicts that the user may need navigation information, the lens 102 automatically displays a navigation arrow and the speaker 401 broadcasts the navigation prompt.

[0105] At the same time, an adaptive optical adjustment unit 402 can also be set on the execution module 400 to combine pupil tracking data with ambient light intensity to dynamically adjust the transmittance and display contrast of the lens 102 to achieve zero-delay personalized visual adaptation.

[0106] Then, the feedback module 500 receives feedback signals such as pupil fixation duration, voice interaction frequency, and environmental operation success rate. Based on the feedback signals, the semantic segmentation accuracy of the perception module 200, the reward function weight of the decision module 300, and the information density threshold of the execution module 400 are optimized in real time, ultimately achieving continuous intelligent growth from "experience feedback-strategy iteration-capability evolution", improving the adaptability and intelligence level of the embodiment of the application to the environment.

[0107] At the same time, the feedback module 500 is connected to the Internet through the wireless circuit module to obtain external data resources required for translation, query or recording in real time, establish and continuously update personalized knowledge graphs, and optimize search efficiency and information presentation accuracy.

[0108] Finally, through the user behavior data and environmental feedback collected by the feedback module 500, the device continuously learns and optimizes its own perception, decision-making and display capabilities.

[0109] According to the embodiments of the present application, the smart glasses proposed may include a glasses body and a perception module, a decision module, an execution module, and a feedback module disposed on the glasses body. In a preset intelligent interaction mode, the glasses can perceive multi-dimensional environmental data of the surrounding environment and the user's current behavior data, predict the user's behavioral intention, and track the user's first behavioral changes and the first environmental changes of the surrounding environment to determine an instruction chain that meets the preset execution conditions, and execute the instruction chain to track the user's second behavioral changes based on the instruction chain, and optimize the instruction chain generation algorithm based on the second behavioral changes. Through the interaction between the physical entity and the environment, intelligent growth is achieved, enabling the wearable device to perceive, understand, and adapt to environmental changes like humans. It can not only actively provide interaction, but also autonomously learn based on the user's behavior to adapt to the usage habits of different users, with a higher level of intelligence. This solves the technical problem in the related art that smart wearable devices have a relatively one-sided understanding of the environment, passive interaction, and can only operate according to established rules. The insufficient intelligence level not only fails to provide convenience to users but also easily affects the user's normal use.

[0110] Next, the interaction method of the smart glasses proposed according to the embodiments of the present application will be described with reference to the accompanying drawings.

[0111] Figure 7 This is a flowchart of the interactive method of the smart glasses according to an embodiment of the present application.

[0112] like Figure 7 As shown, the interaction method of the smart glasses includes the following steps:

[0113] In step S701, in a preset intelligent interaction mode, multi-dimensional environmental data of the surrounding environment and current behavior data of the user are perceived.

[0114] In step S702, based on the multidimensional environmental data, current behavior data and historical event data, the user's behavioral intention is predicted to generate a multimodal decision path, and the user's first behavioral change and the first environmental change of the surrounding environment are tracked to combine the first behavioral change, the first environmental change and the multimodal decision path to determine an instruction chain that meets the preset execution conditions.

[0115] In step S703, the instruction chain is executed, and the second behavior change of the user based on the instruction chain is tracked, and the generation algorithm of the instruction chain is optimized based on the second behavior change.

[0116] It should be noted that the aforementioned explanation of the smart glasses embodiment is also applicable to the interaction method of the smart glasses in this embodiment, and will not be repeated here.

[0117] According to the interaction method of smart glasses proposed in the embodiment of the present application, in a preset intelligent interaction mode, it is possible to perceive the multi-dimensional environmental data of the surrounding environment and the current behavior data of the user, predict the user's behavioral intention, and track the user's first behavioral changes and the first environmental changes of the surrounding environment to determine an instruction chain that meets the preset execution conditions, and execute the instruction chain to track the user's second behavioral changes based on the instruction chain, and optimize the instruction chain generation algorithm based on the second behavioral changes. Through the interaction between physical entities and the environment, intelligent growth is achieved, enabling wearable devices to perceive, understand and adapt to environmental changes like humans. Not only can they actively provide interaction, but they can also learn autonomously based on the user's behavior to adapt to the usage habits of different users, with a higher level of intelligence. This solves the technical problem in the related art that smart wearable devices have a relatively one-sided understanding of the environment, passive interaction, and can only operate according to established rules. The insufficient intelligence level not only fails to provide convenience to users, but also easily affects the user's normal use.

[0118] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. The electronic device may include:

[0119] A memory 801 , a processor 802 , and a computer program stored in the memory 801 and executable on the processor 802 .

[0120] When the processor 802 executes the program, the interaction method of the smart glasses provided in the above embodiment is implemented.

[0121] Furthermore, the electronic device further includes:

[0122] The communication interface 803 is used for communication between the memory 801 and the processor 802 .

[0123] The memory 801 is used to store computer programs that can be run on the processor 802.

[0124] The memory 801 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.

[0125] If the memory 801, processor 802, and communication interface 803 are implemented independently, the communication interface 803, memory 801, and processor 802 can be connected to each other via a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 8 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.

[0126] Optionally, in a specific implementation, if the memory 801, the processor 802 and the communication interface 803 are integrated on a chip, the memory 801, the processor 802 and the communication interface 803 can communicate with each other through an internal interface.

[0127] The processor 802 may be a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.

[0128] This embodiment also provides a computer-readable storage medium having a computer program stored thereon, which implements the above-mentioned smart glasses interaction method when executed by a processor.

[0129] An embodiment of the present application further provides a computer program product, including a computer program, which, when executed by a processor, implements the interaction method of the smart glasses provided by an embodiment of the present invention.

[0130] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or N embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and features of different embodiments or examples without contradiction.

[0131] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of technical features indicated. Thus, a feature specified as "first" or "second" may explicitly or implicitly include at least one such feature. In the description of this application, "N" means at least two, for example, two, three, etc., unless otherwise specifically defined.

[0132] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, fragment or portion of code comprising one or N executable instructions for implementing a custom logical function or process step, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may be performed in a different order than shown or discussed, including performing functions in a substantially simultaneous manner or in a reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present application pertain.

[0133] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection with one or N wires (electronic devices), a portable computer disk cartridge (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and programmable read-only memory (EPROM or flash memory), fiber optic devices, and a portable compact disc read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program can be obtained electronically by optically scanning the paper or other medium and then editing, interpreting or processing it in other suitable ways as necessary, and then storing it in a computer memory.

[0134] It should be understood that various parts of the present application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiment, the N steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used to implement: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0135] Those skilled in the art will understand that all or part of the steps in the method of the above embodiment can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.

[0136] In addition, the functional units in the various embodiments of the present application may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into a module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.

[0137] The storage medium mentioned above may be a read-only memory, a magnetic disk, or an optical disk, etc. Although the embodiments of the present application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present application. Persons skilled in the art may make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.

Claims

1. A pair of smart glasses, characterized in that: include: The glasses body; A sensing module provided on the glasses body, configured to sense multi-dimensional environmental data of the surrounding environment and the user's current behavior data in a preset intelligent interaction mode; a decision module provided on the eyeglass body, configured to predict the user's behavioral intention based on the multi-dimensional environmental data, the current behavioral data, and the historical event data, so as to generate a multimodal decision path, and to track a first behavioral change of the user and a first environmental change of the surrounding environment, so as to determine an instruction chain that meets a preset execution condition by combining the first behavioral change, the first environmental change, and the multimodal decision path; An execution module provided on the glasses body, configured to execute the instruction chain; The feedback module provided on the glasses body is used to track the second behavior change of the user based on the instruction chain, and optimize the generation algorithm of the instruction chain based on the second behavior change.

2. The smart glasses according to claim 1, wherein: The perception module includes: A camera, configured to obtain environmental image data of the surrounding environment and pupil tracking data of the user; A microphone, configured to collect sound data of the surrounding environment and voice command data of the user; An ambient temperature sensor, used to collect ambient temperature data of the surrounding environment; A human body temperature sensor, used to collect the user's body temperature data; A positioning unit is configured to receive positioning data of the user.

3. The smart glasses according to claim 2, wherein: The perception module includes: a visual processing unit, configured to obtain a pupil positioning result of the user based on the environmental image data and the pupil tracking data, determine the user's attention level to environmental information of the surrounding environment using the pupil positioning result, and construct an attention map of the user using the attention level; A semantic segmentation unit is used to fuse the multi-dimensional environmental data to semantically align the environmental image data, the sound data and the voice command data, and perform semantic segmentation based on the aligned environmental image data, the sound data and the voice command data to construct environmental semantics.

4. The smart glasses according to claim 3, wherein: The decision module includes: a prediction unit, configured to fuse the historical event data, the attention map, and the environmental semantics to predict the user's behavioral intention and generate the multimodal decision path; A triggering event unit is used to track the first behavioral change of the user and the first environmental change of the surrounding environment, so as to determine an instruction chain that meets the preset execution conditions in combination with the first behavioral change, the first environmental change and the multimodal decision path.

5. The smart glasses according to claim 1, wherein: The execution module includes: A speaker, configured to convert the voice execution instructions in the instruction chain into voice information and play the voice information; The display unit is configured to convert the display execution instruction in the instruction chain into display content and display the display content.

6. The smart glasses according to claim 1, wherein: The feedback module includes: a receiving unit, configured to receive the second behavior change and the second environmental change of the surrounding environment after the execution module executes any instruction in the instruction chain; An optimization unit is used to optimize the semantic segmentation accuracy of the perception module, the reward function weight of the decision module, and the information density threshold of the execution module based on the second behavior change, the second environment change, and the operation success rate of any one of the instructions.

7. A method for interacting with smart glasses, characterized in that: Applied to the smart glasses according to any one of claims 1 to 6, wherein the method comprises the following steps: In the preset intelligent interaction mode, it perceives the multi-dimensional environmental data of the surrounding environment and the user's current behavior data; Based on the multi-dimensional environmental data, the current behavior data, and the historical event data, predict the user's behavioral intention to generate a multimodal decision path, and track a first behavioral change of the user and a first environmental change of the surrounding environment to determine an instruction chain that meets preset execution conditions in combination with the first behavioral change, the first environmental change, and the multimodal decision path; The instruction chain is executed, and a second behavior change of the user based on the instruction chain is tracked, and a generation algorithm of the instruction chain is optimized based on the second behavior change.

8. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method for interacting with smart glasses as claimed in claim 7.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: The program is executed by a processor to implement the interaction method of the smart glasses as claimed in claim 7.

10. A computer program product, comprising a computer program, wherein when the computer program is executed, the computer program is used to implement the interaction method of the smart glasses according to claim 7.

Citation Information

Cited By

  • AI glasses and smart phone data interaction system under intelligent scene perception

    CN121397131A

  • AI glasses and smart phone data interaction system under intelligent scene perception

    CN121397131B

  • AI intelligent glasses and interaction method thereof

    CN121478127A

  • Intelligent audio glasses adaptive interaction method and device, equipment and medium

    CN122086249A