Multimodal interactive smart glasses system and implementation method
By integrating multimodal data and prioritizing environmental objects, the smart glasses system can accurately generate high-priority object commands even when the user gives incomplete instructions. This solves the problem of incomplete multimodal data integration and improves the user experience.
Patent Information
- Application Number
- CN202511564162.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-30
- Publication Date
- 2026-02-13
- Estimated Expiration
- 2045-10-30
AI Technical Summary
When the user instructions obtained from multimodal data fusion are incomplete, they cannot accurately match the user's actual intent, resulting in a poor user experience.
By integrating multimodal data acquisition hardware modules, including electrode modules, motion sensors, sound sensors, and image acquisition devices, and combining physiological electrical signals, sound information, motion information, and gaze image information, data fusion and environmental object priority classification are performed to recognize gestures or voice commands and generate user commands.
When a user performs an incomplete operation, the system can quickly and accurately generate instructions for high-priority objects, reducing redundant information interference, improving user convenience and accuracy, and reducing operational complexity.
Smart Images

Figure CN121029009B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of smart glasses interaction, and particularly relates to a multi-modal interactive smart glasses system and an implementation method. BACKGROUND
[0002] With the maturity of mixed reality (MR) and smart sensing technology, MR smart glasses have penetrated into multiple core scenarios of consumers and industries. On the consumer side, it can present real-time weather, schedules and other daily information, realize immersive games through virtual-real fusion, real-time translation of foreign languages, or scan products to present price evaluations when shopping; on the industry side, it can superimpose device drawings and remote guidance on engineers in industrial operation, help doctors to retrieve patient image data during surgery in medical scenarios, and intuitively display human structure, molecular composition and other content through 3D models in the education field. In these scenarios, the core value of smart glasses is to perceive user behavior and environmental information and convert it into service instructions to realize hands-free intelligent interaction support.
[0003] Early smart glasses mostly rely on single-modal data to provide services, such as only recognizing the "open navigation" instruction through voice or only positioning the application icon that the user gazes at through eye tracking. Single-modal data processing logic is simple, directly converting collected voice, eye movement and other single signals into operation instructions, and realizing basic convenience in specific scenarios, but it requires users to accurately give clear and accurate operation instructions, which requires high user cooperation and increases the difficulty of user operation, resulting in low user experience. With the development of multi-modal technology, smart glasses start to collect image, voice information and even motion posture data, and compensate for the limitations of single-modal through multi-dimensional information fusion, for example, image information collects that the user is gazing at the TV, voice information collects that the user instructs to turn up the volume, and through the fusion of image information and voice information, the user's intention can be accurately inferred even in the case of low user cooperation, significantly improving the user experience.
[0004] However, in most use scenarios, users may simplify input instructions "lazily", for example, a user wants to query the inventory of a product, image information collects that the user is gazing at a range containing a large number of products, and after combining image information and voice information, in order to satisfy (cover) the user's actual intention as much as possible, the smart glasses will query the inventory information of all products in the image information and superimpose the queried information on the user's gaze range, which will invade the user's normal operation field of view and affect the user's normal working environment, resulting in low user experience. SUMMARY
[0005] The application provides a multi-modal interactive intelligent glasses system and an implementation method to solve the problem that when the user instruction obtained by fusing multi-modal data is incomplete, the service result matching the actual intention of the user cannot be provided to the user through the user instruction, thereby causing a low user experience.
[0006] To solve the above technical problems, the application provides the following technical solutions.
[0007] The multi-modal interactive intelligent glasses implementation method comprises the following steps:
[0008] S10: integrating a hardware module for collecting multi-modal data in the intelligent glasses, wherein the hardware module comprises an electrode module, a motion sensor, a sound sensor and an image collection device, the multi-modal data comprises physiological electrical signals, sound information, motion information, eye image information and gaze image information, and the multi-modal data is stored together with the corresponding collection time;
[0009] S20: aligning the physiological electrical signals with the motion information through the collection time by the intelligent glasses, and fusing and processing the physiological electrical signals and the motion information into motion states and motion trajectories according to the time sequence relationship of the collection time; combining the motion trajectories with the gaze image information to generate environment information according to the collection time, performing contour recognition on the gaze image information, comparing the gaze image information with stored object contour features, marking the objects in the gaze image information as environment objects, extracting the relative positions, pixel colors and text descriptions of the environment objects in the gaze image information, generating object elements, and generating an individual database according to the environment information, the environment objects and the object elements;
[0010] S30: recognizing the environment objects in the gaze image information, matching the current environment information according to the environment objects, and performing priority division on the environment objects in the current environment information in combination with the motion trajectories and the collection time;
[0011] S40: collecting gesture instructions or voice instructions in the sound information or the image information, processing the gesture instructions into range constraint features according to a preset database, and recognizing content features or environment objects in the voice instructions, wherein the content features include any one of the relative positions, the pixel colors and the text descriptions;
[0012] S50: taking the content features, the environment objects and the range constraint features as logical elements, using at least one of the content features, the environment objects and the range constraint features to complete the logical elements according to the logical relationship between the stored logical elements and the priority of the current environment objects, and generating a user instruction according to the completed logical elements.
[0013] In addition, the multi-modal interactive intelligent glasses system using the multi-modal interactive intelligent glasses implementation method is provided.
[0014] The basic scheme principle and benefits are as follows: when the scheme divides the priority of the environmental object, the user's motion trajectory (such as walking to a certain commodity area), the collection time (such as long time staring at a shelf), and the environmental information (such as the warehouse, the store scene) are combined to judge the user's potential attention object from the time and space dimensions, and then physiological electrical signals (such as more likely to focus on nearby goods when stationary), and the range constraint characteristics of gestures (such as locking specific goods when picking up labels) are further filtered to finally distinguish high and low priority objects; this division method can directly narrow the instruction matching range, avoid the occupation of the user's field of view by full environmental object information (such as all goods on the dense shelf), and let the user quickly focus on the target object without manually filtering redundant information.
[0015] At the same time, this way of screening environmental objects in the scheme is suitable for the user's "lazy" incomplete operation (such as only making a gaze or a simple gesture), and the scheme can preferentially complete the instruction based on the high-priority object, reducing the user's repeated adjustment operation and the step of clearly instructing, which not only reduces the use complexity, but also avoids operation errors caused by information interference, and finally significantly improves the convenience and accuracy of the user during use.
[0016] Furthermore, the scheme first takes the content features, environmental objects, and range constraint characteristics as basic logical elements, relies on the pre-stored logical relationship in the individual database, combines the environmental object priority, and completes the missing elements of the user's incomplete instruction; at the same time, object element features are further integrated, including the relative position, pixel color, and edge contour of the environmental object, to build a precise correspondence between the features, objects, and instructions. In actual application, when the user issues a simplified instruction, the intelligent glasses first complete the instruction framework based on the logical elements, and then lock the specific target through the object element features to avoid confusion between similar objects. The corresponding effect is significant, on the one hand, it solves the problem of incomplete instructions caused by the user's simplified operation, without the user's repeated information supplement, reduces the operation steps, and improves the interaction efficiency; on the other hand, with the cross verification of multi-dimensional object element features, the instruction directionality is greatly enhanced, even in complex environments such as dense goods, it can avoid general matching deviation, reduce irrelevant object information interference, balance the use convenience and instruction accuracy, and optimize the overall use experience.
[0017] In summary, the scheme divides the priority of the environmental object through multi-dimensional data, completes the user's instruction by combining logical elements and object element features, solves the problems of incomplete instructions and information redundancy, reduces the user's operation steps and errors, improves the instruction directionality and accuracy in complex environments, and finally significantly optimizes the convenience and experience of user use.
[0018] Further, in step S40, the gesture instruction is also processed into a parameter change feature according to the preset database, and the content feature, the parameter change feature or the environmental object in the voice instruction is identified; in step S50, the parameter change feature is also taken as a logical element when the user instruction is generated according to the completed logical element.
[0019] In a home scenario, when the user issues an incomplete instruction (such as making a “lower” gesture), the scheme can filter valid objects based on the user's stationary state, gaze range (gaze image information display range) and object real-time state, exclude mobile phones without triggering conditions, lock open lights, and automatically complete the brightness parameter to realize instruction generation under low coordination, meet the user's “lazy” needs, and reduce operation coordination. In addition, the scheme extracts the edge contour, color and other appearance features of the object in the gaze image information, establishes an environmental object feature library, combines the gaze image information and the user's motion trajectory (environmental objects in the gaze range during movement), and fuses to generate environmental information (such as a desk scenario, a kitchen scenario), forming individualized scenario data. Taking the kitchen scenario as an example, the user enters the microwave oven, refrigerator and other environmental objects through historical gaze images when moving, takes the color and size of the microwave oven as object elements, temperature as a content feature, and high and low adjustment as a parameter change feature, and establishes a logical association with the microwave oven for storage, avoiding temperature generalization matching air conditioners, realizing scenario-based individualized services, and improving instruction generation accuracy.
[0020] Further, the intelligent glasses generate a connection request according to the locally stored identity information, in step S30, the intelligent glasses broadcast the connection request to the environmental objects when matching the current environmental information; after receiving the connection request, the environmental objects combine the identity information in the connection request with the locally stored authorization conditions to determine whether to establish a communication connection with the intelligent glasses, if the communication connection is established, the locally running data is sent to the intelligent glasses as feedback information of the connection request, if the communication connection is not established, the connection request is ignored; after receiving the feedback information, the intelligent glasses integrate the environmental objects that send the feedback information into an object set, combine the motion trajectory and the collection time of the corresponding gaze image information, and divide the environmental object set into multiple object sub-sets of different priorities; when completing the logical element, the environmental objects are matched according to the priority of the object sub-set.
[0021] The intelligent glasses build a security barrier for device communication based on the connection request and authorization verification mechanism of the identity information of the environmental object, prevent malicious response of unauthorized objects or leakage of sensitive running data, automatically filter irrelevant objects, and avoid non-target device interference with subsequent instruction matching. The local running data feedback of the environmental object not only provides the core reference of the object real-time state (such as whether it is in running) for the intelligent glasses, but also injects the actual running dimension basis for the priority division of the object set, so that the sorting no longer depends on the space-time information and is more in line with the use demand. The matching mode according to the priority further improves the accuracy and response efficiency of the logical element completion. For example, when the user is working, the intelligent glasses broadcast the connection request to the desktop lamp, standby display, and closed printer, and only the lamp returns the running data of "adjustable brightness / power on", and the intelligent glasses divide the lamp into a high-priority sub-set and exclude the display and printer; then the user makes a "lower" gesture, and the intelligent glasses lock the high-priority lamp and generate a "lower lamp brightness" instruction in the completion link. This broadcast handshake mode not only prevents unauthorized devices from leaking state information, but also ensures that only truly available objects enter the candidate, significantly reducing false responses.
[0022] Further, in step S20, the physiological electrical signal and the motion information are hard-synchronized according to the collection time to form a physiological change sequence, and the gaze image information frame is windowed according to a preset segmentation number to form a scene change sequence, so that the physiological change sequence and the scene change sequence are aligned in time sequence to form a state object synchronous sequence; when no gesture instruction or voice instruction is collected, an intention prior model is pre-generated through a dynamic Bayesian network, and the current state object synchronous sequence is input to output an object candidate set; in step S50, the environmental object is preferentially matched in the object candidate set.
[0023] The time sequence alignment of the physiological electrical signal and the gaze image information forms the state object synchronous sequence, which not only solves the misalignment problem of multi-modal data caused by collection delay, but also provides a dual-dimension basis of user physiological state and environmental object change for intention reasoning, avoiding the judgment deviation caused by single data dimension; the dynamic Bayesian network pre-generates the object candidate set based on the synchronous sequence, so that the matching range of step S50 is contracted from the full environmental object to the high-correlation candidate set, which not only improves the matching accuracy of incomplete instructions, but also reduces the power consumption of the present scheme and accelerates the instruction response speed.
[0024] Further, the method further comprises: obtaining historical trajectories corresponding to each environment information and all environment objects from historical data; calculating a motion range of the historical trajectories according to position information in the historical trajectories, and calculating a coincidence degree of the historical trajectories according to the motion range; performing similarity calculation on the environment objects according to object elements of the environment objects, taking environment objects with a similarity greater than a preset similarity threshold as same environment objects, and performing similarity comparison on the environment information according to the same environment objects, and taking environment information with a similarity greater than the preset similarity threshold as same environment information;
[0025] According to a preset statistical time, the appearance frequency of the environment objects corresponding to the same parameter change feature in the user instructions under the same environment information is taken as a comparison frequency, and the appearance frequency of the combination of the different parameter change features and the content features under the same environment object in the user instructions is taken as a use frequency; in step S50, when the environment objects in the logical elements are completed, the environment objects with a higher comparison frequency are preferentially selected; when the content features or the parameter change features in the logical elements are completed, the combination of the parameter change features and the content features and the corresponding appearance frequency are filtered out in combination with the environment objects and the environment information, the two feature combinations are sorted from high to low according to the appearance frequency, and if there is a content feature or a parameter change during this time, the sorting is adjusted again according to the content feature or the parameter change, and the two feature combinations in the front are selected from the sorting result to complete the logical elements.
[0026] Before the broadcast connection request, the smart glasses match the same environment information in historical data within a preset time according to the current motion trajectory and the corresponding gaze image information, obtain all user instructions from the matching result, count the appearance frequency of the same environment objects in the user instructions, and divide all the environment objects in the user instructions into common objects and uncommon objects according to the appearance frequency.
[0027] When the broadcast connection request is received, for the common objects, the smart glasses directly carry the locally stored authorization token to complete the fast handshake; for the uncommon objects, the smart glasses embed temporary permission application information in the connection request, and send a return information to the smart glasses when the receiving party of the connection request disagrees with the connection, the return information including a permission level and a maximum concurrency; the smart glasses establish a non-communication queue according to the environment objects that send the return information, dynamically adjust the sorting of the environment objects in the queue according to the permission level and the maximum concurrency in the return information, and set the interval time for sending the connection request to the environment objects by combining the waiting number of the environment objects in the queue with the maximum concurrency.
[0028] Based on the frequency statistics of user historical operation habits, the instruction generation is converted from general adaptation to personalized fitting, that is, when multiple candidate objects appear in the same priority, the object with the most historical occurrence times is preferentially selected to avoid the blindness of random matching; when the user has specified the object but lacks parameters or content, the existing logical elements are combined with the high-frequency combination of the environment object to complete the completion of the content characteristics or parameter change characteristics, eliminating random blind selection and reducing repeated correction.
[0029] At the same time, the intelligent glasses count the frequency of the appearance of the environment object according to the environmental information in the historical data and the user instruction, and divide the environment object into frequently used objects and infrequently used objects; for frequently used objects, the mirror directly carries the local stored authorization token to complete the fast handshake, significantly improving the connection efficiency and response speed of the frequently used scene; for infrequently used objects, the mirror embeds temporary permission application information in the connection request, the receiver returns the permission level and the maximum number of concurrent connections, and the mirror dynamically adjusts the broadcast interval and the queue length accordingly, avoiding invalid connection from occupying bandwidth and computing power, and realizing accurate allocation of resources; the same frequency data is used as the basis for classification of environment objects and supports dynamic optimization of connection strategy, achieving the functions of classification, authorization and queuing at one time, and simultaneously bringing the effects of connection acceleration, resource saving and seamless scene migration without adding new hardware.
[0030] Further, before completing the logical elements in step S50, the real-time running state of the environment object is obtained from the environment object that has established a communication connection, and the environment object in the closed state or without adjustable parameters is immediately excluded, and only the object set that remains open and operable is matched with the environment object; when the voice instruction and the gesture instruction point to the same object but the parameters are opposite, the current state of the environment object is used as the arbitration condition, the parameter consistent with the state change direction is preferentially selected, and the conflicting parameter is marked as a negative sample in reverse, and the intention prior model is immediately updated.
[0031] Before completing the logical elements, the intelligent glasses first use the real-time running state returned by the environment object that has established a communication connection to immediately exclude the devices in the closed state or without adjustable parameters, reduce the candidate set, and reduce the subsequent calculation amount; when the voice instruction and the gesture instruction issue opposite parameters to the same object, the current state change direction is used as the arbitration condition, the parameter that can be immediately executed is selected, and the eliminated parameter is automatically marked as a negative sample and written back to the intention prior model, so that the same state information simultaneously completes the three functions of range compression of the environment object, conflict arbitration of the user instruction and self-evolution of the intention prior model, realizing the collaborative value of more efficient instruction, less error and continuous optimization of the model.
[0032] Further, the intelligent glasses eliminate the environmental objects, acquire the environmental objects under the current environment information as initial environmental objects, acquire the object elements of the initial environmental objects, compare the initial environmental objects with the stored environmental objects according to the object elements, when the similarity obtained by comparison is greater than a preset similarity threshold, the stored environmental objects are taken as similar environmental objects, all similar environmental objects of the initial environmental objects under the current environment information are acquired, the coincidence degree between the current environment information and each environment information in the individual database is calculated according to all similar environmental objects, the environmental objects existing in the current environment information and a certain environment information in the individual database are taken as connection environmental objects, the proportion of the connection environmental objects in all environmental objects in the current environment information is taken as a first coincidence degree, the proportion of the connection environmental objects in all environmental objects in a certain environment information is taken as a second coincidence degree, and the ratio of the first coincidence degree to the second coincidence degree is taken as the coincidence degree proportion between the current environmental objects and the certain environmental objects.
[0033] According to the coincidence degree proportion, the environment information closest to the current environment information is matched from the stored environment information, if the coincidence degree proportion between the closest environment information is within a preset new environment proportion range, the current environment information is taken as similar environment information, and the connection environmental objects and the intention prior model are migrated to the similar environment information; if the proportion is lower than the lower limit of the new environment proportion range, a new environment information is established according to the current gaze image information and the motion trajectory, and is stored in the individual database, and then the new environment information is supplemented through the first demonstration operation of the user; if the coincidence proportion exceeds the upper limit of the new environment proportion range, the matched environment information is taken as the current environment information.
[0034] After the intelligent glasses eliminate the environmental objects, the current environmental objects and the object set under the historical environment information are subjected to coincidence proportion calculation, the proportion falling within a preset interval is regarded as a similar scene, and the object set and the intention prior model of the old scene are immediately migrated; the proportion lower than the interval is a new scene and is supplemented through the first demonstration, and the proportion higher than the interval is completely reused. The same proportion comparison simultaneously completes the scene similarity judgment and the intention prior model cold start, so that the commonly used environment is quickly adapted, the new environment is quickly formed, and the repeated modeling power is reduced.
[0035] Furthermore, when the current environment information is not the stored environment information, and no voice or gesture command is received within the preset waiting time, the environment objects that have been communicated with are prioritized based on their frequency of occurrence in historical user commands within the preset statistical time. A number of alternative commands are set according to the priority, and the most frequently occurring user operation command is selected from the historical user command data based on the environment object and its corresponding number of alternative commands as alternative commands. These alternative commands are then stored in the cache. When a user inputs a voice or gesture command, alternative commands are first selected based on the priority of the environment objects, or further selected based on their frequency of occurrence. The selected alternative commands are then used as the user command.
[0036] When smart glasses determine that a new scene is being created (i.e., not an environment already stored) and no instruction is received within the waiting time, the connected objects are sorted by historical usage frequency. The most frequently used operation instructions corresponding to high-frequency objects are extracted to generate a candidate instruction cache. Once the user speaks or makes a move, the cache is used as the first match. The same frequency sorting simultaneously completes the preloading during the silent period and the acceleration of the first instruction. The first interaction in a new scene does not need to be searched from scratch, the response time is shortened, and the user's cooperation is reduced.
[0037] Furthermore, if the overlap ratio between the current environmental information and the closest environmental information is lower than the lower limit of the new environmental ratio range, then the communication connection with the environmental object whose frequency of occurrence does not exceed the preset usage lower limit will be disconnected, and only the environmental objects that have a logical relationship with the preset content features will be retained among the environmental objects that have been connected.
[0038] If the proportion falls below the lower limit of the new environment's proportion range, the smart glasses immediately disconnect from communication with low-frequency objects, retaining only objects that have a logical relationship with preset content characteristics, forming a minimum usable set. This simultaneous disconnection action releases the communication load and safely prunes the environment object candidate pool, ensuring that in unfamiliar environments, commands are not mistakenly sent to unrelated devices while reliable instructions are quickly provided with the minimum set, balancing energy consumption, safety, and accuracy. Attached Figure Description
[0039] Figure 1 This is a flowchart of the method for implementing multimodal interactive smart glasses in Embodiment 1 of this solution. Detailed Implementation
[0040] The following will describe the concept and technical effects of the present invention clearly and completely with reference to embodiments, so as to fully understand the purpose, features and effects of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are all within the scope of protection of the present invention. Example
[0041] As Figure 1 shown, the multi-modal interactive intelligent glasses implementation method comprises the following steps:
[0042] S10: integrating a hardware module for collecting multi-modal data in the intelligent glasses, the hardware module comprising an electrode module, a motion sensor, a sound sensor and an image acquisition device, the multi-modal data comprising physiological electrical signals, sound information, motion information, eye image information and gaze image information; storing the multi-modal data together with the corresponding collection time;
[0043] S20: the intelligent glasses align the physiological electrical signals and the motion information through the collection time based on the historical data collected, and fuse and process them into motion states and motion trajectories according to the time sequence relationship of the collection time; combine the motion trajectories and the gaze image information to generate environmental information according to the collection time, perform contour recognition on the gaze image information, compare it with the stored object contour features, mark the objects in the gaze image information as environmental objects, and extract the relative position, pixel color (also including the image content of the product surface) and text description (for example, the text content on the product label and the text content on the product surface) of the environmental objects in the gaze image information to generate object elements, and generate an individual database according to the environmental information, the environmental objects and the object elements;
[0044] S30: identify the environmental objects in the gaze image information, match the current environmental information according to the environmental objects, and prioritize the environmental objects in the current environmental information according to the motion trajectories and the collection time;
[0045] S40: collect gesture instructions or voice instructions in the sound information or the image information, process the gesture instructions into range constraint features according to a preset database, and identify the content features or environmental objects in the voice instructions, wherein the content features include any one of the relative position, pixel color and text description;
[0046] S50: use at least one of the content features, the environmental objects and the range constraint features as logical elements, complete the logical elements according to the priority of the current environmental objects according to the logical relationship between the stored logical elements, and generate user instructions according to the completed logical elements.
[0047] The dry electrode module is embedded at the position where the left and right temples of the intelligent glasses contact the back of the ears and the position where the nose pad contacts the nose bridge to collect physiological electrical signals (including heart rate, skin electrical signals, etc.) by adhering to the skin. A motion sensor (such as a three-axis acceleration sensor) is integrated in the middle section of the temple to collect three-axis acceleration during user motion. A linear array of microphones is arranged at the lower end of the temple to form a beamforming channel for collecting environmental speech and bone conduction sound when the user speaks. A binocular camera is centrally placed at the front end of the frame to collect image information within the user's field of view as gaze image information. An infrared eye movement camera is arranged on the inner side of the frame to capture eye image information of the user. All sensors share a 50MHz master clock, which generates a 64-bit UTC timestamp after being divided by an FPGA, and writes the timestamp into the ring buffer of the mirror body security enclave with the data packet to achieve hard-synchronous storage of multi-modal data and collection time, providing a unified time basis for subsequent alignment, fusion, and priority division.
[0048] When the three-axis acceleration is fused and processed into a motion state and a motion trajectory according to the time sequence relationship of the collection time, the original data is first obtained from the three-axis accelerometer, the gravity component is separated (combined with the attitude information to remove) and filtered and denoised to retain pure motion acceleration. Based on the initial speed, the motion acceleration is integrated and accumulated according to the sampling time interval to obtain the velocity in each axis. Then, based on the initial position, the velocity is also integrated and accumulated, and the three-axis displacement is combined to obtain a three-dimensional motion trajectory. On this basis, a gyroscope is set inside the intelligent glasses to correct the integral accumulation deviation, and then the trajectory drift is corrected in reverse by fixing the objects in the gaze scene.
[0049] The administrator or user can pre-set the priority gradient according to individual habits, and define the division range of each priority gradient in the gaze image information for each scene information (for example, when inventorying in a warehouse, taking the center point of the gaze image information as the origin, and taking the circular ring to divide the gaze image information into 30% as high priority and 30%-60% as medium priority). When the environmental objects in the current environmental information are prioritized, the system will first analyze the current environment information in which the user is located (such as a store sales scene or a warehouse inventory scene), then call the pre-set gaze image range and priority gradient corresponding rules in this environment, and finally assign the priority bound to all environmental objects in the range to complete the priority division.
[0050] The administrator or user can set density constraints and position distribution constraints for different priority gradients according to individual habits. After the intelligent glasses divide the range of the gaze image information, the density and position distribution of the environmental objects in the division range are calculated based on the division result, and the environmental objects are further constrained by the density constraints and position constraints to exclude some environmental objects.
[0051] In a specific implementation, when the warehouse staff wears the smart glasses to carry out warehouse inventory, the device hardware module immediately starts to work cooperatively: the dry electrode module at the tail of the temple and the nose pad closely adheres to the skin to collect physiological electrical signals such as heart rate and skin electrical signals in real time; the three-axis acceleration sensor in the middle segment of the temple accurately captures the three-axis acceleration data of the warehouse staff during the movement between the shelves; the linear array of microphones at the lower end of the temple forms a beamforming channel to synchronously collect the inventory instruction voice and environmental sound of the warehouse staff, while preferentially recording the bone conduction sound when the user speaks to improve the recognition degree in a noisy environment; the binocular camera in the middle of the front end of the frame continuously takes images of the shelves and goods as the gaze image information; the infrared eye movement camera on the inner side of the frame captures the eye movement in real time to accurately track the gaze focus of the warehouse staff. Assuming that all sensors share a 50MHz master clock, a 64-bit UTC timestamp is generated after the FPGA frequency division, and the timestamp is written into the ring buffer of the mirror body security enclave together with the data packets collected by each module to realize the hard-synchronous storage of multi-modal data and collection time, and provide a unified and accurate time reference for subsequent data alignment, fusion and environmental object priority division.
[0052] Subsequently, the smart glasses process the motion data. First, the original data is obtained from the three-axis accelerometer, the gravity component is separated and removed in combination with the attitude information, and then the pure motion acceleration is retained by filtering algorithm denoising; the initial speed is taken as the reference, and the pure motion acceleration is accumulated and integrated according to the sampling time interval to obtain the axial velocity; then the initial position is taken as the reference, and the axial velocity is integrated and accumulated again to generate the three-dimensional motion trajectory of the warehouse staff. In the process, the gyroscope built-in the smart glasses corrects the integral accumulation deviation in real time, and the images of fixed objects such as warehouse pillars and goods location marks collected by the binocular camera will correct the trajectory drift in the opposite direction, ensuring the accuracy of the trajectory of the warehouse staff moving between different shelves, and avoiding the influence of trajectory deviation on subsequent environmental judgment.
[0053] The smart glasses further call the individual database to align the physiological electrical signals and motion information through the collection time, determine the motion state of the warehouse staff moving and stationary in combination with the timing relationship. When the warehouse staff stops in front of a row of beverage shelves, the physiological electrical signals tend to be stable, and the motion information has no obvious change, which is determined as "stationary state"; then the motion trajectory and gaze image information are associated to generate the "warehouse beverage shelf" environmental information. The gaze image is subjected to contour recognition, the recognition result is compared with the stored commodity contour features, objects such as bottled water and canned cola are marked as environmental objects, the relative positions of the commodities (such as the right side of the second layer of the shelf), the pixel colors (such as the blue label of bottled water and the red label of cola can), and the text descriptions (such as "550ml mineral water" and "330ml cola" on the label) are extracted to generate object elements, and the individual database is updated.
[0054] The priority classification stage is that the intelligent glasses first analyze the current "warehouse beverage shelf" environment information, call the rules set by the administrator in advance, and assume that the center point of the gaze image information is the origin, the high priority is the range of 30% of the circular ring, and the medium priority is the range of 30%-60%. The canned cola in the range of 30% is classified as high priority, and the 550ml mineral water in the range of 30%-60% is classified as medium priority. When the warehouse manager says "check the quantity", the microphone collects the voice instruction, recognizes the content feature "quantity", takes the content feature and the high-priority canned cola as logical elements, and combines the logical relationship between the canned cola and the quantity statistics in the individual database to complete the user instruction. The binocular camera distinguishes the category through the pixel color and text description of the goods and automatically counts, and efficiently completes the accurate inventory operation.
[0055] The warehouse manager sets the density and position distribution of the environment object as follows: only 5 evenly distributed environment objects in the range of 30% of the center of the gaze image information. At a certain moment, the warehouse manager stops in front of a certain shelf, and the gaze image information collected by the intelligent glasses shows that a large number of bottled water are placed on the shelf; according to the range division and calculation of the intelligent glasses, the 30% range of the center of the gaze image information contains 23 bottles of bottled water of different brands. At this time, the intelligent glasses start the screening according to the preset constraints: first, based on the position distribution constraint, the spatial positions of the 23 bottles of bottled water are calculated through image coordinates, and the bottled water gathered in the same small area (such as a certain grid of the shelf) is excluded, and the candidate objects with dispersed positions are preliminarily screened out; then, based on the density constraint, the candidate objects are further screened, and finally 5 bottled water evenly distributed in the range of 30% are determined as the target environment objects.
[0056] In this process, since the same kind of bottled water is placed as much as possible in the same area (such as the second layer left side of the shelf), although the 23 bottles of bottled water in this embodiment only cover 6 different brands, through the preferential screening of the dispersed area based on the position distribution constraint and the accurate control of the quantity based on the density constraint, the 5 bottles of bottled water finally selected are respectively from 5 different brands of concentrated areas, which realizes the uniform coverage of the product categories and better meets the actual needs of the warehouse manager to quickly inventory multiple categories. Embodiment
[0057] The difference between this embodiment and embodiment 1 is that in step S40, the gesture instruction is also processed into a parameter change feature according to the preset database, and the content feature, the parameter change feature, or the environment object in the voice instruction is recognized. In step S50, the parameter change feature is also taken as a logical element when the user instruction is generated according to the completed logical element.
[0058] The smart glasses generate a connection request according to the locally stored identity information, and in step S30, the smart glasses broadcast the connection request to the environment object when matching the current environment information; the environment object combines the identity information in the connection request with the locally stored authorization conditions to determine whether to establish a communication connection with the smart glasses, and if so, sends the locally running data as feedback information of the connection request to the smart glasses, and if not, ignores the connection request; after receiving the feedback information, the smart glasses integrate the environment object sending the feedback information into an object set, and combine the motion trajectory and the collection time of the corresponding gaze image information to divide the environment object set into multiple object sub-sets of different priorities; when completing the logical elements, the environment objects are matched in turn according to the priority of the object sub-set.
[0059] In specific implementation, after the user wears the smart glasses, the hardware modules of the smart glasses start working immediately: the dry electrode module at the back of the ear and the nose pad is attached to the skin to continuously collect physiological electrical signals such as heart rate and skin electrical signals; the three-axis acceleration sensor in the middle segment of the glasses leg starts to capture motion data; the microphone linear array at the lower end of the glasses leg forms a beamforming channel to synchronously collect environmental sound and user bone conduction voice; the binocular camera at the front end of the glasses frame takes images within the field of view as gaze image information, and the infrared eye movement camera on the inside captures eye image information. All sensors rely on a 50MHz master clock to generate a unified 64-bit UTC timestamp through FPGA frequency division, and the data packet is written into the ring buffer of the glasses security enclave along with the timestamp, realizing hard-synchronous storage of multi-modal data and laying a unified time basis for subsequent processing.
[0060] When the user walks from the bedroom to the study, the motion sensor collects real-time three-axis acceleration raw data. The smart glasses first separate the gravity component combined with the attitude information obtained by the gyroscope, retain the pure motion acceleration after filtering and noise reduction, then take the initial speed as the reference, accumulate the integral according to the sampling time interval to obtain the velocity in each axis, and further integrate and accumulate the three-axis displacement to generate a three-dimensional motion trajectory. In the process, the gyroscope continuously corrects the integral accumulation deviation, and the images of fixed objects such as the bedroom bed and the corridor picture collected by the binocular camera are also used to correct the trajectory drift in the opposite direction, accurately recording the complete motion trajectory in the bedroom scene, the corridor scene and the study scene. At this time, the motion state is determined as uniform motion.
[0061] After the user enters the study and sits down at the desk, the motion sensor detects that the three-axis acceleration is continuously zero (assuming that the user's head does not move or that the value detected by a small range of shaking is zero), and in combination with the stable heart rate and stable skin electrical signals collected by the dry electrode, the physiological electrical signals are determined to switch the motion state to the stationary state. The binocular camera focuses on the desk area at this time, collects the images of objects such as books, pens, table lamps, and computers, compares the object contour features with the stored object contour features through contour recognition, marks these objects as environmental objects, and extracts the relative positions (such as the table lamp on the left side of the desk) and pixel colors (such as the black pen and the white table lamp cover) of the objects to generate object elements. At the same time, the environmental information of the study posture scene, the marked environmental objects, and the object elements are associated and stored, and are updated to the individual database.
[0062] The smart glasses generate a connection request according to the user's biological characteristics (such as the user's eye feature information collected by the eye image information, including iris appearance feature, eyelid shape feature, etc.) stored locally, and the identity information such as device unique identifier, and broadcast the request to all environmental objects marked in the field of view (in the gaze image information or in the study scene) when matching the environmental information of the current study posture scene. The smart table lamp, smart phone and other devices with communication capability in the study receive the connection request, combine the identity information in the request with the authorization conditions such as the locally stored identity white list and scene permission matrix, and determine whether to establish a communication connection; objects such as books and pens without communication function do not respond to the request, and the closed computer (which cannot receive the connection request due to being in the closed state) also does not respond to the request. Among them, the smart table lamp and the smart phone are both in the authorized white list and the current scene meets the permission requirements, so a communication connection is established with the smart glasses, and the local running data (such as the current brightness and running mode of the smart table lamp, and the boot state and CPU occupancy of the smart phone) is sent to the smart glasses as feedback information; if there is an unauthorized foreign smart device (the smart phone next door also receives the signal), its verification fails (including the case that the smart phone next door does not answer or refuses the connection request), and no communication connection is established. After receiving the feedback information, the smart glasses integrate the smart table lamp and the smart phone to form an environmental object set, and combine the motion trajectory of the user after entering the study (finally stationary at the desk) and the collection time of the gaze image information collected by the binocular camera (the gaze duration of the smart table lamp is obviously longer than that of the smart phone which can be seen only by occasionally shaking the head), and divide the object set into two priority sub-sets: the smart table lamp is the high-priority sub-set, and the smart phone is the low-priority sub-set.
[0063] Based on the motion state, the collection time and the priority level of the object sub-set, the moving stage from the bedroom to the study and the stationary stage in the study are divided into two basic priority levels. The priority of the environmental objects such as the bed, the wardrobe and the hanging picture on the moving track is low, and the priority of the smart phone, the smart table lamp and the smart computer in the study is high. In the object sub-set (assuming that it only includes the smart phone and the smart table lamp) in the stationary stage in the study, the priority of the smart table lamp is the highest, and the priority of the smart phone is the second.
[0064] If the user makes a "lower" gesture (the binocular camera captures and processes it into a parameter change feature through a preset database) at this time, the smart glasses preferentially match in the high-priority objects in the study environment. In combination with the relationship between the logical elements, only the table lamp is associated with the "brightness adjustment" logic, and the real-time state is turned on. Therefore, the logical elements are completed, and the user instruction to lower the brightness of the table lamp is generated. The whole process does not require the user to provide a complete instruction, which fully meets the low-compliance use demand. Embodiment
[0065] The difference between this embodiment and Embodiments 1-2 is only that the gaze image information and the eye image information are combined, and the user's attention range in the gaze image information is inferred through the gaze image information and the relative position of the pupil in the eye image information, so as to further divide the priority of the object sub-set.
[0066] In specific implementation, after the smart glasses integrate the smart table lamp and the smart phone into the object set, the eye image information and the gaze image information after synchronization are called, the relative position of the pupil in the eye image (such as the vertical coordinate of the pupil center from the upper edge of the eyelid and the horizontal coordinate from the inner side of the white of the eye) is combined with the stored eye movement visual field mapping model, and the user's attention range in the gaze image is accurately inferred.
[0067] When the user is in a stationary posture, the infrared eye movement camera detects that the pupil center is long-term in the lower part of the eye image, which corresponds to the pen and the book on the desk in the gaze image. Occasionally, the user looks at the right table lamp through the peripheral light, and the mobile phone is not appeared in the attention range. Therefore, the priority of the table lamp is set to be higher than that of the mobile phone, so that the user instruction generation result is more consistent with the instantaneous gaze focus. Embodiment
[0068] The only difference between this embodiment and embodiments 1-3 is that, in step S20, physiological electrical signals and motion information are hard-synchronized according to the acquisition time to form a physiological change sequence. At the same time, the gaze image information frame is segmented by a window according to a preset number of segments (determined by the administrator based on computing power and calculation accuracy) to form a scene change sequence, so that the physiological change sequence and the scene change sequence are aligned in time to form a state object synchronization sequence. When no gesture command or voice command is acquired, an intention prior model is pre-generated through a dynamic Bayesian network, and the current state object synchronization sequence is used as input to output an object candidate set. In step S50, environmental objects are matched first from the object candidate set.
[0069] In practice, the physiological electrical signals and motion information after hard synchronization are extracted and divided into time windows according to a preset number of segments (i.e., the sampling period, assumed to be 100ms). Within each window, the mean heart rate and the variance of motion acceleration are calculated to generate a physiological change sequence containing heart rate and motion state (e.g., when walking in the living room, the heart rate is 75 beats / min + the variance of acceleration is 0.8; after sitting down in the study, the heart rate is 68 beats / min + the variance of acceleration is 0.1).
[0070] The gaze image information frame is segmented into time windows (consistent with the time granularity of physiological sequences). The object outline and relative position are identified within each window to generate a scene information and scene object scene change sequence (e.g., the scene change sequence of the living room window is the living room and sofa; the scene change sequence of the study window is the study and desk lamp and computer).
[0071] Using UTC timestamps as indexes, physiological change subsequences and scene change subsequences within the same time window are bound together to form a state object synchronization sequence (e.g., in the 16:05:00 window, heart rate 68, still, study, and desk lamp), thus solving the problem of misalignment between physiological states and environmental objects caused by multimodal data acquisition delays.
[0072] Leveraging the temporal dependencies and probabilistic reasoning capabilities of Dynamic Bayesian Networks (DBNs), conditional probabilistic relationships are constructed between physiological states, scene changes, and user intentions. For example, DBNs capture the temporal correlations of sequence data (e.g., walking followed by stillness often accompanies the intention to prepare for work, and the intention to prepare for work is highly correlated with a smart desk lamp), and can filter highly correlated objects through probability calculations. Historical state object synchronization sequences (e.g., multiple records of stillness in a study, a desk lamp, and brightness adjustments) are retrieved from the individual database. Using these sequences as input and the user's final operation object as the label, the transition probability matrix of the DBN is trained (e.g., in the combination of a still state and a study scene, the correlation probability of a smart desk lamp is higher than that of a smartphone). When no instruction is collected, the intention prior model generates a candidate set of objects each time. If no corresponding instruction is triggered subsequently, the rejected objects are marked as negative samples (e.g., a smart air conditioner mistakenly included as a candidate), and the database is immediately updated to update the probability matrix, reducing noise interference.
[0073] For example, when the user sits down in the study without immediately issuing an instruction, the priori model has output a candidate set centered on the smart desk lamp; at this time, the user makes a "lower" gesture, and step S50 does not need to traverse all objects in the study, but directly matches the logic of the desk lamp associated brightness adjustment in the candidate set, and quickly generates a user instruction to lower the brightness of the desk lamp. Embodiments
[0074] The difference between this embodiment and embodiments 1-4 is that the historical trajectory corresponding to each environment information and all environment objects are further obtained from the historical data; the motion range of the historical trajectory is calculated according to the position information in the historical trajectory, and the coincidence degree of the historical trajectory is calculated according to the motion range; the similarity of the environment objects is calculated according to the object elements of the environment objects, the environment objects with a similarity greater than a preset similarity threshold (set by an administrator or a user) are taken as the same environment objects, and the similarity of the environment information is compared according to the same environment objects, and the environment information with a similarity greater than the preset similarity threshold is taken as the same environment information.
[0075] According to a preset statistical time (set by a user according to personal habits), the appearance frequency of the environment objects corresponding to the same parameter change characteristics in the user instructions under the same environment information is taken as the comparison frequency, and the appearance frequency of the combination of different parameter change characteristics and content characteristics under the same environment objects in the user instructions is taken as the use frequency.
[0076] In step S50, when the environment objects in the logical elements are completed, the environment objects with a higher comparison frequency are preferentially selected; when the content characteristics or parameter change characteristics in the logical elements are completed, the combination of the parameter change characteristics and the content characteristics and the corresponding appearance frequency are first filtered out in combination with the environment objects and the environment information, the two characteristic combinations are sorted from high to low according to the appearance frequency, and if there is a content characteristic or a parameter change during this time, the sorting is adjusted again according to the content characteristic or the parameter change, and the two characteristic combinations in the front are selected from the sorting result to complete the logical elements.
[0077] Before the broadcast connection request, the smart glasses match the same environment information in the historical data within a preset statistical time according to the current motion trajectory and the corresponding gaze image information, obtain all user instructions from the matching results, count the appearance frequency of the same environment objects in the user instructions, and divide all environment objects in the user instructions into frequently used objects and infrequently used objects according to the appearance frequency.
[0078] When broadcasting the connection request, for the frequently used objects, the local stored authorization token is directly carried to complete the fast handshake; for the infrequently used objects, the temporary permission application information is embedded in the connection request, and the receiver of the connection request sends back the return information to the smart glasses when it disagrees with the connection, the return information including the permission level and the maximum concurrency; the smart glasses establish the un-communication queue according to the environment objects sending the return information, dynamically adjust the sorting of the environment objects in the queue according to the permission level and the maximum concurrency in the return information, and set the interval time for sending the connection request to the environment objects by combining the waiting number of the environment objects in the queue with the maximum concurrency.
[0079] Before the completion logic element in step S50, the real-time running state of the environment object is acquired from the environment object having established the communication connection, the environment object in the closed state or having no adjustable parameter is immediately excluded, and the environment object is matched only in the remaining object set which is in the open state and is operable; when the voice instruction and the gesture instruction point to the same object but the parameters are opposite, the current state of the environment object is taken as the arbitration condition, the parameter consistent with the state change direction is preferentially selected, and the conflict parameter is reversely marked as the negative sample, and the intention prior model is immediately updated.
[0080] In the implementation, the user wears the smart glasses and walks into the study, sits down and starts work. The smart glasses retrieve the historical data, count the occurrence of the environment objects corresponding to the same parameter change characteristics in the user instructions in the study environment, and the occurrence of the different parameter and content feature combinations under the same environment object, and accordingly classify the smart desk lamp and the smart computer as the frequently used objects and the smart bookshelf as the infrequently used object.
[0081] When broadcasting the connection request, the smart glasses directly use the local stored authorization token to quickly establish the connection for the smart desk lamp and the smart computer, and add the temporary permission application in the request for the smart bookshelf. After the smart bookshelf disagrees with the connection, the permission related information (the permission level and the maximum concurrency) is returned, and the smart glasses adjust the connection request sending interval for the smart bookshelf and the queuing sequence when not connected, so that the printer which has not been connected completes the communication connection first.
[0082] Before the completion logic element, the smart glasses acquire the real-time state of the connected object, find that the computer is in the closed state and has no adjustable parameter, and immediately exclude it, and only leave the desk lamp in the open state and operable as the candidate.
[0083] At a moment, the user says "increase" and makes the "decrease" gesture, both of which point to the desk lamp but the parameters are opposite. The smart glasses take the current running state of the desk lamp as the judgment basis, if the desk lamp is changing in the direction of increasing the brightness at this moment, the increasing parameter is preferentially selected, and simultaneously the desk lamp brightness in the scene state is marked as the inappropriate sample, and the intention prior model is immediately updated. EMBODIMENT
[0084] The difference between the present embodiment and embodiments 1-5 is that, after the smart glasses eliminate the environmental objects, the initial environmental objects under the current environmental information are obtained, the object elements of the initial environmental objects are obtained, the initial environmental objects are compared with the stored environmental objects according to the object elements, when the similarity obtained by comparison is greater than a preset similarity threshold, the stored environmental objects are taken as similar environmental objects, the similar environmental objects of all initial environmental objects under the current environmental information are obtained, and the coincidence degree between the current environmental information and each environmental information in the individual database is calculated according to all similar environmental objects. The environmental objects existing in the current environmental information and a certain environmental information in the individual database at the same time are taken as connection environmental objects, the proportion of the connection environmental objects in all environmental objects in the current environmental information is taken as the first coincidence degree, the proportion of the connection environmental objects in all environmental objects in a certain environmental information is taken as the second coincidence degree, and the ratio of the first coincidence degree to the second coincidence degree is taken as the coincidence degree proportion between the current environmental objects and a certain environmental object.
[0085] The environmental information closest to the current environmental information is matched from the stored environmental information according to the coincidence degree proportion, if the coincidence degree proportion between the closest environmental information is within a preset new environment proportion range (generally 40%-80%, which is set by the administrator according to the actual use of the user), the current environmental information is taken as similar environmental information, and the connection environmental objects and their intention prior models are migrated to the similar environmental information; if the proportion is lower than the lower limit of the new environment proportion range, a new environmental information is established according to the current gaze image information and the motion trajectory, and is stored in the individual database, and then the new environmental information is supplemented through the first demonstration operation of the user; if the coincidence proportion exceeds the upper limit of the new environment proportion range, the matched environmental information is taken as the current environmental information.
[0086] When the current environment information is not the stored environment information, and no voice instruction or gesture instruction is received within the preset waiting time (set by the user according to his own use experience, for example, 5 minutes), the environment objects that have been communicatively connected are prioritized according to the occurrence frequency of the historical user instructions of the environment objects that have been communicatively connected within the preset statistical time (the time for calculating the frequency, i.e., the preset statistical time, is set by the user according to his own personal use, for example, one month), and the number of alternatives is set according to the priority (specifically set by the user according to the frequency range, for example, 200 times / month or more, the number of alternatives is 5, 200 times / month-150 times / month, the number of alternatives is 4), and the user operation instruction with the highest occurrence frequency is selected as the alternative instruction from the historical user instruction data according to the environment object and the corresponding number of alternatives (for example, the user reads a book in the library, the coincidence degree of the environment object in the library with the stored environment object in the study is 70%, and the number of alternatives corresponding to the occurrence frequency of the table lamp is 5, and the 5 instructions with the highest occurrence frequency are selected from the historical user instructions including the table lamp, assuming that the statistical standard of the occurrence frequency is the last one month), and the alternative instruction is stored in the cache; when the user enters a voice instruction or a gesture instruction, the alternative instruction is first screened according to the priority of the environment object, or the alternative instruction is further screened according to the occurrence frequency of the alternative instruction, and the screened alternative instruction is used as the user instruction. For example, the table lamp is first matched according to the brightness, and then the occurrence frequencies of "table lamp brightness up" and "table lamp brightness down" are compared (assuming that the statistical standard of the occurrence frequency is the last one month), and the table lamp brightness up with the higher occurrence frequency is selected as the user instruction.
[0087] If the coincidence degree between the current environment information and the closest environment information is less than the lower limit of the new environment proportion range, the communication connection with the environment object whose occurrence frequency is less than the preset use lower limit (specifically set by the user according to the frequency range, for example, 100 times / month) is disconnected, and only the environment object that has a logical relationship with the preset content feature is retained in the environment object that has been communicatively connected.
[0088] In a specific implementation, when the user enters the library, the smart glasses worn by the user first eliminate the environment objects that have not been used for a long time (such as a phone in sleep mode), and then calculate the coincidence proportion of the objects such as the smart table lamp and the smart bookshelf in the current environment and the stored environment objects in the study. Assuming that the calculated coincidence proportion is within the preset new environment proportion range, the smart glasses determine that the library is a similar environment, and the coincident table lamp and its priori model are migrated to the current environment.
[0089] Because the library is a new environment not stored, and no user instruction is received within the preset waiting time, the smart glasses prioritize the smart table lamp, the smart bookshelf and the like according to the occurrence frequency of the connected environment objects in the historical instructions, assume that the priority of the smart table lamp is 5, set the number of alternatives according to the priority, assume that the number of alternatives of the smart table lamp is 5, filter the high-frequency operation instructions related to these objects from the historical data and store them in the cache, and assume that the high-frequency operation instructions of the table lamp from high to low are: increasing the brightness of the table lamp, decreasing the brightness of the table lamp, turning off the table lamp, and turning on the table lamp. When the user makes a "decrease" gesture, the "table lamp" is matched according to the priority first, and then the "table lamp" related instructions in the cache are generated according to the frequency to generate a decrease in the brightness of the table lamp as the user instruction.
[0090] After the user enters the bookstore, the smart glasses calculate the coincidence ratio of the current environment objects and the stored environment, assume that the calculated coincidence ratio is lower than the lower limit of the new environment ratio range, and determine that it is a completely new environment. The smart glasses disconnect the communication connection with the objects with low historical use frequency, assume that it is a watch, and only keep the environment objects related to the preset content features of "brightness" and "sound", such as smart earphones, smart phones and smart table lamps. Then, the new bookstore environment information is established according to the user's gaze image information and movement trajectory, and the environment information is supplemented and improved through the user's first demonstration operation (such as increasing the smart chair).
[0091] In this embodiment, a multi-modal interactive smart glasses system using a multi-modal interactive smart glasses implementation method is also included.
[0092] The above is only an embodiment of the application, and the application is not limited to the field involved in this embodiment. Well-known specific structures and characteristics in the scheme are not described in detail here. The ordinary skilled person in the art knows all the ordinary technical knowledge in the application field before the application date or the priority date, can know all the prior art in the field, and has the ability to apply conventional experimental means before that date. The ordinary skilled person in the art can improve and implement the scheme under the guidance of this application, and some typical well-known structures or well-known methods should not be an obstacle for the ordinary skilled person in the art to implement the application. It should be noted that, for those skilled in the art, without departing from the structure of the application, a number of modifications and improvements can be made, which should also be considered as the protection scope of the application, and these will not affect the effect and practicality of the application. The protection scope of this application should be subject to the content of its claims, and the specific implementation mode and the like in the specification can be used to explain the content of the claims.
Claims
1. A method for implementing multi-modal interactive smart glasses, characterized in that, The method comprises the following steps: S10: integrating a hardware module for collecting multi-modal data in the smart glasses, the hardware module comprising an electrode module, a motion sensor, a sound sensor, and an image acquisition device, and the multi-modal data comprising physiological electrical signals, sound information, motion information, eye image information, and gaze image information; The multi-modal data is stored together with the corresponding collection time; S20: the smart glasses align the physiological electrical signals and the motion information by the collection time, and fuse and process them into motion states and motion trajectories according to the time sequence relationship of the collection time; the motion trajectories and the gaze image information are combined to generate environmental information, the gaze image information is subjected to contour recognition, and the objects in the gaze image information are marked as environmental objects by comparing with the stored object contour features, the relative positions, pixel colors, and textual descriptions of the environmental objects in the gaze image information are extracted to generate object elements, and an individual database is generated according to the environmental information, the environmental objects, and the object elements; S30: the environmental objects in the gaze image information are recognized, the current environmental information is matched according to the environmental objects, and the environmental objects in the current environmental information are prioritized according to the motion trajectories and the collection time; S40: gesture instructions or voice instructions in the sound information or the image information are collected, the gesture instructions are processed into range constraint features according to a preset database, and content features or environmental objects in the voice instructions are recognized, the content features comprising any one of the relative positions, the pixel colors, and the textual descriptions; S50: the content features, the environmental objects, and the range constraint features are used as logical elements, at least one of the content features, the environmental objects, and the range constraint features is used to complete the logical elements according to the logical relationship between the stored logical elements and the priority of the current environmental objects, and a user instruction is generated according to the completed logical elements.
2. The multi-modal interactive smart glasses implementation method of claim 1, wherein: In step S40, the gesture instructions are processed into parameter change features according to the preset database, and the content features, the parameter change features, or the environmental objects in the voice instructions are recognized; in step S50, the parameter change features are also used as the logical elements when the user instruction is generated according to the completed logical elements.
3. The method of claim 2, wherein: The smart glasses generate a connection request according to the locally stored identity information, the smart glasses broadcast the connection request to the environmental objects when matching the current environmental information in step S30; the environmental objects receive the connection request, combine the identity information in the connection request with the locally stored authorization conditions, determine whether to establish a communication connection with the smart glasses, send the locally running data as feedback information of the connection request to the smart glasses if the communication connection is established, and ignore the connection request if the communication connection is not established; The smart glasses receive the feedback information, integrate the environmental objects that send the feedback information into an object set, divide the object set into a plurality of object sub-sets of different priorities according to the motion trajectories and the collection time of the corresponding gaze image information, and match the environmental objects according to the priorities of the object sub-sets when completing the logical elements.
4. The method as claimed in claim 2, wherein the multi-modal interactive smart glasses implementation method is characterized by: In step S20, the physiological signal and the motion information are hard-synchronized according to the collection time to form a physiological change sequence, and the gaze image information frames are windowed according to a preset segmentation number to form a scene change sequence, so that the physiological change sequence and the scene change sequence are aligned in time sequence to form a state object synchronous sequence; when no gesture instruction or voice instruction is collected, an intention prior model is pre-generated through a dynamic Bayesian network, the current state object synchronous sequence is taken as input, and an object candidate set is output; in step S50, the environmental objects in the object candidate set are preferentially matched.
5. The multi-modal interactive smart glasses implementation method of claim 3, wherein: Further comprising: obtaining historical trajectories corresponding to each environmental information and all environmental objects from historical data; calculating a motion range of the historical trajectories according to position information in the historical trajectories, and calculating a coincidence degree of the historical trajectories according to the motion range; performing similarity calculation on the environmental objects according to object elements of the environmental objects, taking environmental objects with a similarity greater than a preset similarity threshold as same environmental objects, and performing similarity comparison on the environmental information according to the same environmental objects, and taking environmental information with a similarity greater than the preset similarity threshold as same environmental information; According to a preset statistical time, the occurrence frequency of the environmental objects corresponding to the same parameter change feature in the user instructions under the same environmental information is taken as a comparison frequency, and the occurrence frequency of the combination of the different parameter change features and the content features under the same environmental object in the user instructions is taken as a use frequency; in step S50, when the environmental objects in the logical elements are completed, the environmental objects with a higher comparison frequency are preferentially selected; when the content features or the parameter change features in the logical elements are completed, the combination of the parameter change features and the content features and the corresponding occurrence frequency are filtered out in combination with the environmental objects and the environmental information, the two feature combinations are sorted from high to low according to the occurrence frequency, if there is a content feature or a parameter change in this time of completion, the sorting is adjusted again according to the content feature or the parameter change, and the two feature combinations in the front are selected from the sorting result to complete the logical elements; Before the broadcast connection request, the intelligent glasses match the same environmental information in the historical data of the current motion trajectory and the corresponding gaze image information within a preset time, obtain all user instructions from the matching result, count the occurrence frequency of the same environmental objects in the user instructions, and divide all the environmental objects in the user instructions into frequently-used objects and infrequently-used objects according to the occurrence frequency; When the broadcast connection request is received, for the frequently-used objects, the intelligent glasses directly complete the fast handshake by carrying the locally-stored authorization token; for the infrequently-used objects, the intelligent glasses embed temporary permission application information in the connection request, and send a return information to the intelligent glasses when the receiving party of the connection request disagrees with the connection, the return information including a permission level and a maximum concurrency number; The intelligent glasses establish a non-communication queue according to the environmental objects that send the return information, dynamically adjust the sorting of the environmental objects in the queue according to the permission level and the maximum concurrency number in the return information, and set the interval time for sending the connection request to the environmental objects in combination with the waiting number of the environmental objects in the queue and the maximum concurrency number.
6. The method of claim 4 or 5, wherein: Before the completion logic element in step S50, the real-time running state of the environment object is obtained for the environment object that has established a communication connection, the environment object in the closed state or without adjustable parameters is immediately excluded, and the environment object is matched in the remaining open and operable object set; when the voice instruction and the gesture instruction point to the same object but the parameters are opposite, the current state of the environment object is used as an arbitration condition, the parameter consistent with the state change direction is preferentially selected, and the conflict parameter is marked as a negative sample in reverse, and the intention prior model is immediately updated.
7. The multi-modal interactive smart glasses implementation method of claim 6, wherein: After the smart glasses exclude the environment object, the environment object under the current environment information is obtained as the initial environment object, the object element of the initial environment object is obtained, the initial environment object is compared with the stored environment object according to the object element, when the similarity obtained by the comparison is greater than a preset similarity threshold, the stored environment object is used as a similar environment object, all similar environment objects of the initial environment object under the current environment information are obtained, the coincidence degree between the current environment information and each environment information in the individual database is calculated according to all similar environment objects, the environment object existing in the current environment information and a certain environment information in the individual database at the same time is used as a connection environment object, the proportion of all environment objects of the connection environment object in the current environment information is used as a first coincidence degree, the proportion of all environment objects of the connection environment object in a certain environment information is used as a second coincidence degree, and the ratio of the first coincidence degree to the second coincidence degree is used as the coincidence degree proportion between the current environment object and a certain environment object. The environment information closest to the current environment information is matched from the stored environment information according to the coincidence degree proportion, if the coincidence degree proportion between the closest environment information is within a preset new environment proportion range, the current environment information is used as a similar environment information, and the connection environment object and its intention prior model are migrated to the similar environment information; if the proportion is lower than the lower limit of the new environment proportion range, a new environment information is established according to the current gaze image information and the motion trajectory, and is stored in the individual database, and the new environment information is supplemented through the first demonstration operation of the user; if the coincidence proportion exceeds the upper limit of the new environment proportion range, the matched environment information is used as the current environment information.
8. The multi-modal interactive smart glasses implementation method of claim 7, wherein: When the current environment information is not the stored environment information, and no voice instruction or gesture instruction is received within a preset waiting time, the priority of the environment object that has established a communication connection is sorted according to the appearance frequency of the historical user instructions of the environment object within a preset statistical time, the number of candidates is set according to the priority, the user operation instruction with the highest appearance frequency is selected as a candidate instruction from the historical user instruction data according to the environment object and the corresponding candidate number, and the candidate instruction is stored in the cache. When the user inputs a voice instruction or a gesture instruction, the candidate instruction is first screened according to the priority of the environment object, or the candidate instruction is further screened according to the appearance frequency of the candidate instruction, and the screened candidate instruction is used as the user instruction.
9. The multi-modal interactive smart glasses implementation method of claim 8, wherein: If the coincidence ratio between the current environment information and the closest environment information is lower than the new environment ratio range lower limit, the communication connection with the environment object whose occurrence frequency is not more than the preset use lower limit is disconnected, and only the environment object having a logical relationship with the preset content feature is reserved among the environment objects having the communication connection.
10. A multi-modal interactive smart glasses system, characterized in that, The multi-modal interactive smart glasses are implemented by using the method of any one of claims 1-9.
Citation Information
Patent Citations
Multi-modal fusion AR glasses intelligent control method and system
CN120469576A
Control Method, Terminal, and System
US20200034729A1