Intelligent glasses control method, electronic equipment and storage medium

By integrating multimodal data of smart glasses and finger rings, using large models to make interactive intention decisions, the problem of inaccurate control of smart glasses in noisy environments is solved, and flexible and accurate user interaction in different scenarios is achieved.

CN120559862APending Publication Date: 2025-08-29湖北星纪魅族集团有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510693477.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-08-29

AI Technical Summary

Technical Problem

In the environment where noise and inconvenient manual operation are not convenient, existing smart glasses control methods cannot effectively achieve normal user interaction with the device, especially when voice control is inaccurate or hand operation is limited in riding and noisy environments.

Method used

By integrating multimodal data of smart glasses and smart rings, pre-trained large models are used to splice and infer data, and user interaction commands are generated, including voice data, inertial measurement unit perception data and ring key data, and appropriate control methods are automatically selected to reduce the impact of the external environment on user operations.

Benefits of technology

In different application scenarios, smart glasses can automatically generate user interaction commands to ensure normal interaction with users, reduce the impact of the external environment on operations, and improve the accuracy and flexibility of control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120559862A_ABST
    Figure CN120559862A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent glasses control method, electronic equipment and a storage medium, and relates to the technical field of intelligent glasses, and the method comprises the steps: judging that intelligent glasses enter a preset application scene, and obtaining interaction state configuration data corresponding to the application scene; acquiring voice data and sensing data of a glasses inertial measurement unit from the intelligent glasses; acquiring ring inertial measurement unit sensing data and ring key data from an intelligent ring; splicing the identification data of the application scene, the interaction state configuration data, the voice data, the sensing data of the glasses inertial measurement unit, the sensing data of the ring inertial measurement unit and the ring key data into a one-dimensional data set, and inputting the one-dimensional data set into a pre-trained large model for reasoning to obtain an interaction command of a user; and obtaining and executing an interaction command. According to the method, the real intention of the user can be decided in different application scenes, and the influence of the external environment on the glasses control operation of the user is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of smart glasses, and in particular to a smart glasses control method, electronic device, and storage medium. Background Art

[0002] In some smart glasses application scenarios where noise is present and manual operation is inconvenient, users often cannot use smart glasses properly. For example, in cycling scenarios, the strong wind noise makes voice control inaccurate, and it is difficult to touch the smart glasses' control buttons. Or in noisy environments, it is difficult to manually control smart glasses with objects in hand. No better solution has been proposed yet. Summary of the Invention

[0003] The embodiments of the present application provide a smart glasses control method, an electronic device, and a storage medium to solve one or more of the above-mentioned technical problems.

[0004] In a first aspect, an embodiment of the present application provides a method for controlling smart glasses, including: determining that the smart glasses have entered a preset application scenario, and obtaining interaction state configuration data corresponding to the application scenario; collecting voice data and glasses inertial measurement unit perception data from the smart glasses; collecting ring inertial measurement unit perception data and ring button data from a smart ring; splicing the recognition data of the application scenario, the interaction state configuration data, the voice data, the glasses inertial measurement unit perception data, the ring inertial measurement unit perception data, and the ring button data into a one-dimensional data group, inputting the one-dimensional data group into a pre-trained large model for inference to obtain the user's interaction command; and obtaining and executing the interaction command.

[0005] In a second aspect, an embodiment of the present application provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory, wherein the processor implements any of the above-described methods when executing the computer program.

[0006] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements any of the methods described above.

[0007] Compared with the related art, this application has the following advantages:

[0008] The present application provides a method for controlling smart glasses, an electronic device, and a storage medium. The method comprises: determining that the smart glasses have entered a preset application scenario and obtaining interaction state configuration data corresponding to the application scenario; collecting voice data and glasses inertial measurement unit perception data from the smart glasses; collecting ring inertial measurement unit perception data and ring button data from a smart ring; concatenating the application scenario recognition data, the interaction state configuration data, the voice data, the glasses inertial measurement unit perception data, the ring inertial measurement unit perception data, and the ring button data into a one-dimensional data set, inputting the one-dimensional data set into a pre-trained large model for inference to obtain a user interaction command; and obtaining and executing the interaction command. According to an embodiment of the present application, based on the characteristics of various operations performed by the user and the different application scenarios entered by the smart glasses and the different interaction state configuration data, the user interaction command is automatically generated, and the smart glasses obtain and execute the interaction command, thereby ensuring that the interaction between the smart glasses and the user can be as normal as possible in different application scenarios, determining the user's true intention, and reducing the impact of the external environment on the user's control of the glasses.

[0009] The above description is only an overview of the technical solution of this application. In order to more clearly understand the technical means of this application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of this application more obvious and easy to understand, the specific implementation methods of this application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In the accompanying drawings, unless otherwise specified, the same reference numerals throughout the multiple drawings represent the same or similar components or elements. These drawings are not necessarily drawn to scale. It should be understood that these drawings only depict some embodiments according to the present application and should not be regarded as limiting the scope of the present application.

[0011] Figure 1 A flow chart of a smart glasses control method provided in an embodiment of the present application is shown;

[0012] Figure 2 An interactive intention decision diagram based on deep learning provided in an embodiment of the present application is shown;

[0013] Figure 3 A flowchart of integrating multiple detection methods provided in an embodiment of the present application is shown;

[0014] Figure 4 A flowchart of finger tapping detection for a smart ring provided in an embodiment of the present application is shown;

[0015] Figure 5 A waveform detection method provided in an embodiment of the present application is shown. Figure 1 ;

[0016] Figure 6 A waveform detection method provided in an embodiment of the present application is shown. Figure 2 ;

[0017] Figure 7 A waveform detection method provided in an embodiment of the present application is shown. Figure 3 ;

[0018] Figure 8 A structural block diagram of a smart glasses control device provided in an embodiment of the present application is shown;

[0019] Figure 9 A block diagram of an electronic device used to implement an embodiment of the present application is shown. DETAILED DESCRIPTION

[0020] Hereinafter, only certain exemplary embodiments are briefly described. As will be appreciated by those skilled in the art, the described embodiments may be modified in various ways without departing from the spirit or scope of the present application. Therefore, the drawings and description are to be regarded as illustrative in nature and not restrictive.

[0021] To facilitate understanding of the technical solutions of the embodiments of the present application, the following describes the related technologies of the embodiments of the present application. The following related technologies can be combined with the technical solutions of the embodiments of the present application as optional solutions, and all of them fall within the scope of protection of the embodiments of the present application.

[0022] Currently, voice interaction is the typical control method for smart glasses. Most smart glasses rely on voice control, but voice control can fail in noisy environments, such as train stations and shopping malls. It's also not suitable for quieter environments, such as movie theaters or meetings.

[0023] In addition, some smart glasses offer physical buttons that support basic actions like single-click, double-click, scrolling up and down, and long-press. Smart rings typically have one or two buttons that also support single-click, double-click, scrolling up and down, and long-press. Smart rings can connect to phones or smart glasses via Bluetooth. Smart rings often replace the functions of smart wristbands, detecting movement and providing health monitoring capabilities. They often include accelerometers and gyroscopes.

[0024] In addition to button control, the smart ring can also detect the movements of the finger wearing the ring to control smart glasses. Tapping with the finger on the smart ring is much more convenient than pressing buttons on smart glasses. Even if the user is holding other objects, the tapping action of individual fingers is not affected. The smart ring's inertial sensor detects single, double, and continuous tapping movements of the finger. These tapping movements are mapped to various smart glasses operations, effectively replacing the buttons on the smart glasses.

[0025] Therefore, smart glasses can be controlled based on multimodal data to perform operations required by the user. However, existing multimodal input data is sometimes redundant and sometimes missing. Existing solutions are to fuse multimodal data, perform feature engineering to extract features from the multimodal data, and use machine learning or deep learning to determine the interaction intention. Existing multimodal interaction data fusion decision-making solutions are firstly professional and difficult in terms of feature engineering. Secondly, when fusing multiple features, they use specific formulas set by expert systems, and the results are unpredictable. In particular, when constants are present, the selection of constant values ​​is accidental and highly correlated with individual users and usage habits. The generalization performance is weak and not universal.

[0026] Based on the above background, the present application provides a smart glasses control method, electronic device and storage medium, which integrate the multimodal methods of multiple portable smart devices to control smart glasses. According to the characteristics of various control methods, a system is designed to automatically select the appropriate control method according to different scenarios, integrate the multimodal input data of multiple devices, and use specific logic to arbitrate more appropriate interaction intentions. This application uses a neural network-based approach, without the need for feature engineering, and directly uses the various modal data of each device as the input of the neural network to determine the user's true intention.

[0027] The following describes in detail the technical solution of this application and how it solves the aforementioned technical problems using specific embodiments. The several specific embodiments listed can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments. The following describes the embodiments of this application in detail with reference to the accompanying drawings.

[0028] The embodiment of the present application provides a method for controlling smart glasses. Figure 1 FIG2 is a flow chart of a method for controlling smart glasses according to an embodiment of the present application. The method may include:

[0029] S101: Determine that the smart glasses have entered a preset application scenario, and obtain interaction state configuration data corresponding to the application scenario.

[0030] In the embodiments of the present application, smart glasses can be used to implement different functions in different application scenarios. The pre-set application scenarios of smart glasses may include, but are not limited to, the following scenarios: navigation, work assistance, healthcare, education and training, entertainment and gaming, social interaction, shopping and retail, security monitoring, information acquisition, productivity tools, and Internet of Things integration.

[0031] Navigation includes but is not limited to: outdoor navigation, where smart glasses can provide real-time navigation information outdoors, displaying routes, directions, and destination locations to help users find the right path; it can also include indoor navigation, providing navigation services in complex indoor environments (such as shopping malls, airports, and hospitals) to help users find specific stores, boarding gates, or clinics. Work assistance includes but is not limited to: industry and manufacturing, where smart glasses can display operating instructions, equipment status, and fault diagnosis information in factories or production lines to help workers improve work efficiency and safety; repair and maintenance, where technicians can use smart glasses to obtain real-time repair instructions, view virtual models of equipment, and troubleshooting steps to reduce repair time and error rates; logistics and warehousing, where smart glasses can help staff quickly find the location of items, optimize routes, and manage inventory.

[0032] Medical and health monitoring, including but not limited to: telemedicine, where doctors can use smart glasses to conduct remote diagnosis and consultation, view patients' real-time data and medical records, and make video calls and provide guidance; surgical assistance, where smart glasses can display patients' medical information, surgical steps, and real-time images during surgery to help surgeons perform precise operations; health monitoring, where smart glasses can monitor users' physiological data (such as heart rate, number of steps, and calorie consumption) and provide health advice and alerts. Education and training, including but not limited to: distance learning, where students can use smart glasses to participate in virtual classes, watch instructional videos, and interactive exercises to enhance their learning experience; and vocational training, where smart glasses can provide real-time guidance and simulated exercises to help trainees master skills and knowledge.

[0033] Entertainment and gaming, including but not limited to: AR (Augmented Reality) gaming, where smart glasses can overlay virtual elements onto real-world environments to provide an immersive gaming experience; video and music playback, where users can watch videos, live broadcasts, and concerts through smart glasses and enjoy high-quality entertainment content; social interaction, including but not limited to: real-time translation, where smart glasses can provide real-time language translation to help users communicate with friends and colleagues who speak different languages; and social media, where users can view and update social media statuses, send messages, and make video calls through smart glasses.

[0034] Shopping and retail, including but not limited to: enhancing the shopping experience: in stores, smart glasses can display product information, price comparisons, and recommendations to help users make purchasing decisions; virtual try-on: users can use smart glasses to try on clothes and accessories in a virtual environment and see how they look. Security and surveillance, including but not limited to: security patrols: security personnel can use smart glasses to conduct real-time monitoring, view camera feeds and alarm information, and improve safety; identity verification: smart glasses can perform facial recognition and identity verification to ensure access rights and security.

[0035] Information access and notifications, including but not limited to: real-time notifications, such as receiving and displaying notifications from mobile phones or other devices, such as messages, emails, and calendar reminders; information query, such as weather, news, and stock information, quickly querying through voice commands or gestures. Productivity tools, including but not limited to: document viewing and editing, such as viewing and editing documents, spreadsheets, and presentations to improve office efficiency; task management, such as managing schedules, tasks, and to-do lists to keep work organized. Internet of Things (IoT) integration, including but not limited to: smart home control, such as controlling smart devices in the home, such as lights, thermostats, and security systems, through smart glasses; and environmental monitoring, such as real-time monitoring and display of environmental data, such as temperature, humidity, and air quality.

[0036] Each preset application scenario corresponds to a set of interaction state configuration data. The application scenario describes the application currently running on the smart glasses. The interaction state configuration data describes the interaction currently waiting for the user, i.e., the interaction state.

[0037] In this step, after determining that the smart glasses have entered a preset application scenario, the current working status and interaction status of the smart glasses are obtained by acquiring the interaction status configuration data corresponding to the application scenario, so as to provide richer input information for the large model and obtain interaction commands that are more in line with the user's intentions.

[0038] S102: Collect voice data and glasses inertial measurement unit perception data from the smart glasses.

[0039] In the embodiment of the present application, the user can issue control commands to the smart glasses by speaking, or by nodding, shaking the head, etc. The smart glasses include a voice acquisition device and an inertial measurement unit. Therefore, after the user issues a control command, the smart glasses can collect voice data and perception data from the glasses' inertial measurement unit.

[0040] It should be noted that the voice acquisition device of the glasses and / or the inertial measurement unit of the glasses can be set inside or outside the glasses frame, and can be set according to actual needs. The embodiments of the present application do not specifically limit this.

[0041] In this step, by collecting voice data and perception data from the smart glasses' inertial measurement unit, multimodal control instructions can be obtained, which facilitates decision-making based on the user's voice data and head control data, and obtains interactive commands that are more adapted to the user's true intentions.

[0042] S103: Collect ring inertial measurement unit perception data and ring button data from a smart ring.

[0043] In an embodiment of the present application, a user can control a smart ring, for example, by tapping or shaking a finger, to issue control commands to the smart glasses. The smart ring includes an inertial measurement unit (IMU) and a case data acquisition unit. Thus, the smart ring can acquire its own IMU sensing data and ring key press data. The smart glasses include a data receiving device that communicates with the smart ring through this data receiving device, collecting IMU sensing data and ring key press data from the smart ring.

[0044] In this step, by collecting the ring inertial measurement unit perception data and ring button data from the smart ring, multimodal control instructions can be obtained, which facilitates decision-making based on the user's finger control data and obtains interactive commands that are more suitable for the user's true intentions.

[0045] S104, splicing the recognition data of the application scenario, the interaction state configuration data, the voice data, the glasses inertial measurement unit perception data, the ring inertial measurement unit perception data, and the ring button data into a one-dimensional data group, and inputting the one-dimensional data group into a pre-trained large model for inference to obtain the user's interaction command.

[0046] In an embodiment of the present application, the identification data of each application scenario can be used to uniquely identify the application scenario. The identification data may include but is not limited to the following types of data: text data, label data, numerical data, and image data. In one possible implementation, the identification data of the application scenario can be set as one-dimensional data. For example, it can include at least one-dimensional data corresponding to the following typical applications: no application: 0; music: 1; navigation: 2; voice assistant: 3; translation: 4; phone transcription: 5; speech: 6. The interaction state configuration data can also be one-dimensional data. In one possible implementation, the voice data, the glasses inertial measurement unit perception data, the ring inertial measurement unit perception data, and the ring button data can be multi-dimensional data. For example, the voice data is 35-dimensional data, the glasses inertial measurement unit perception data is 3-dimensional data, the ring inertial measurement unit perception data is 6-dimensional data, and the ring button data is 6-dimensional data.

[0047] In a possible implementation, the ring tapping result data is shown in Table 1:

[0048] No knock Click 2 consecutive hits 3 consecutive hits 4 consecutive hits 5 consecutive hits

[0049] Table 1

[0050] The ring's finger tap data consists of six elements (6 dimensions). For example, if the ring taps three times in a row, the resulting data is: [0,0,0,1,0,0]. For another example, if no taps are detected, the data is: [1,0,0,0,0,0]. All data can be formatted as integers, or as single-precision floating-point numbers to facilitate integration with other data.

[0051] In a possible implementation, the ring key press result data is shown in Table 2:

[0052] No action Click double click Slide up Downward Long press

[0053] Table 2

[0054] The ring button data consists of six (6-dimensional) values. For example, if you scroll down, the result is [0,0,0,0,1,0]. For another example, if no button is pressed, the result is [1,0,0,0,0,0]. All data can be formatted as integers. To integrate with other data, all data can be formatted as single-precision floating-point numbers.

[0055] On this basis, the application scenario recognition data, the interaction state configuration data, the voice data, the glasses inertial measurement unit perception data, the ring inertial measurement unit perception data, and the ring button data are spliced ​​into a one-dimensional data set. That is, all input data are spliced ​​together to form a 6+6+3+35+1+1=52-dimensional one-dimensional floating-point array as the input of the neural network. When splicing, the splicing order of the input data can be set according to actual needs and is not specifically limited in this embodiment of the application.

[0056] In the embodiment of the present application, multi-dimensional data can be used to represent voice data, multi-dimensional data can be used to represent glasses inertial measurement unit perception data, multi-dimensional data can be used to represent finger ring inertial measurement unit perception data, and multi-dimensional data can be used to represent finger ring button data. After obtaining the recognition data of the application scenario, the interaction state configuration data, the voice data, the glasses inertial measurement unit perception data, the finger ring inertial measurement unit perception data, and the finger ring button data, they are spliced ​​into a one-dimensional data group, and the one-dimensional data group is input into a pre-trained large model for inference to obtain the user's interaction command. Figure 2The deep learning-based interaction intention decision diagram shown above uses a neural network to infer the concatenated one-dimensional data set. The neural network's inference output is the eyewear application interaction intention. In one possible implementation, the neural network outputs the intent as shown in Table 3 below. The output has 10 labels, and the dimension of the final fully connected layer output is 10. The label corresponding to the maximum softmax output is the user's interaction intention.

[0057] It should be noted that in the embodiments of the present application, the neural network can be deployed on the glasses or on other terminals that are connected to the glasses, such as a mobile phone, to use the powerful computing power of the mobile phone for reasoning. In daily use, the embedded points in the glasses will collect the input data and results of each interaction and transmit them to the mobile phone or the cloud, with each user having a copy of the data. The model will be fine-tuned regularly, and the fine-tuned model will be sent to the mobile phone. The inference model will be updated regularly, so that the algorithm can better understand the user based on the user's usage habits.

[0058] Output Label Output label number explain No operation 1 confirm 2 Cancel 3 Play 4 stop 5 Options 6 For example, when selecting a scene, the user says "the third one" Slide up 7 Increase, flip up, brighten, make the sound louder, etc. Downward 8 Reduce, scroll down, darken, lower the volume, etc. Shutdown 9 Restart 10

[0059] Table 3

[0060] It should also be noted that, in the embodiment of the present application, the pre-trained large model can be a large language model (Large Language Models, LLM), including but not limited to any one of the following models: GPT (Generative Pre-trained Transformer) series models, BERT (Bidirectional Encoder Representations from Transformers) model, T5 (Text-to-Text Transfer Transformer) model and PaLM (Pathways Language Model) model.

[0061] In this step, the multimodal data can be simply spliced ​​together and used as the input of the pre-trained large model, and then the user's interaction commands can be inferred. There is no need to perform feature engineering on the multimodal data. The extracted features are fused and merged as the input of the large model, which reduces the difficulty of data processing, ensures the effect of reasoning, reduces randomness, and improves versatility.

[0062] S105: Acquire and execute the interactive command.

[0063] In an embodiment of the present application, if the pre-trained large model is deployed in the glasses, the interaction command is obtained from the smart glasses. If the pre-trained large model is deployed in another device that is communicatively connected to the glasses, the interaction command is obtained from the other device. After obtaining the interaction command, the smart glasses are controlled to execute the interaction command, thereby performing the corresponding operation according to the user's intention.

[0064] In one possible implementation, the smart glasses application interaction commands may include but are not limited to the following commands: confirm, cancel, play, stop, volume up, volume down, screen brighter, screen darker, power on, power off.

[0065] The voice data collected from the smart glasses can be set to support all the above interactive commands. The button mode of the smart bracelet can be set so that the smart ring supports all the above interactive commands. Smart rings usually have more than three buttons. One of the buttons can be used to support the instructions of confirm, cancel, play, stop, power on, and power off. For example, pressing once is confirm / play / stop, pressing twice in a row is cancel, pressing for 3 seconds is power off, pressing for 5 seconds is power on, etc. The other two buttons can be used to support volume adjustment and screen brightness adjustment. The IMU (Inertial Measurement Unit) sensor on the smart glasses can detect the movement of the head. Therefore, the data sensed by the glasses' inertial measurement unit can also be mapped to the above interactive commands, such as shaking the head (mapped to confirm / play / stop) and nodding (can be mapped to cancel). The tapping action of the finger wearing the ring can also be mapped to different commands. For example, you can refer to the following Table 4 for configuration:

[0066]

[0067] Table 4

[0068] The present application provides a method for controlling smart glasses, the method comprising: determining that the smart glasses have entered a preset application scenario, obtaining interaction state configuration data corresponding to the application scenario; collecting voice data and glasses inertial measurement unit perception data from the smart glasses; collecting ring inertial measurement unit perception data and ring button data from a smart ring; concatenating the application scenario recognition data, the interaction state configuration data, the voice data, the glasses inertial measurement unit perception data, the ring inertial measurement unit perception data, and the ring button data into a one-dimensional data set, inputting the one-dimensional data set into a pre-trained large model for inference to obtain a user interaction command; and obtaining and executing the interaction command. According to an embodiment of the present application, based on the characteristics of various operations performed by the user and the different application scenarios entered by the smart glasses and the different interaction state configuration data, the user interaction command is automatically generated, and the smart glasses obtain and execute the interaction command, thereby ensuring that the interaction between the smart glasses and the user can be as normal as possible in different application scenarios, determining the user's true intention, and reducing the impact of the external environment on the user's control of the glasses.

[0069] In a possible implementation, collecting the perception data of the ring inertial measurement unit can be performed according to the following steps: collecting sensor data of the ring inertial measurement unit of the smart ring and generating waveform data corresponding to the sensor data; detecting the waveform data, and in response to detecting a finger tapping action, setting the current cumulative number of taps of the ring; if no finger tapping action is detected again within a specified time period, the current cumulative number of taps is used as the number of ring taps, otherwise the current cumulative number of taps of the ring is reset; and setting the number of ring taps as the perception data of the ring inertial measurement unit.

[0070] In this possible implementation, the sensor data from the ring's inertial measurement unit is directly collected by the ring's inertial measurement unit. The collected sensor data is processed to generate waveform data. By detecting the waveform data, the user's finger movements are determined. When a finger tapping action is detected, the number of taps can be measured. Figure 5-Figure 7 Different waveform detection results are shown.

[0071] Set the current cumulative number of taps of the ring. If no finger tapping action is detected again within the specified time, the current cumulative number of taps will be used as the number of ring taps. Otherwise, the current cumulative number of taps of the ring will be reset. It should be noted that in the embodiment of the present application, the value of the specified time can be set according to actual needs and experience, and the embodiment of the present application does not make any specific restrictions on this.

[0072] The following table 5 and table 6 are combined to Figure 4The parameters in the description are as follows, where the parameter Default is not in Figure 4 It shows:

[0073] state describe Default Initial state, default state Stage 1 Detect a single click, wait for a timeout or click again Stage 2 Detect double click, wait for timeout or click again Stage 3 Detect 3 clicks, wait for timeout or click again Stage 4 Detect 4 clicks, wait for timeout or click again

[0074] Table 5

[0075] state T3 timeout event Detecting a click event Default Change the state to stage1 and start timer T3 Stage 1 End of detection: Click Change the state to stage2 and start timer T3 Stage 2 End of detection: Double-click Change the state to stage3 and start timer T3 Stage 3 Detection end: 3 hits Change the state to stage4 and start timer T3 Stage 4 Detection ended: 4 hits Change the state to stage5 and start timer T3

[0076] Table 6

[0077] For details on this step, see Figure 4 The smart ring finger tap detection flow chart shown in the figure is executed, including:

[0078] S1, start waveform detection.

[0079] S2, state setting.

[0080] S3 starts timer T3. Since a single click takes approximately Tr = 300ms and the interval between clicks is Tc = 300ms, the timer T3 expires at 300 + 300 = 600ms. Each time a single click is detected, timer T3 starts. If no more single clicks are detected after T3 times out, the operation ends. Five consecutive clicks are an exception, as they can be set to a maximum of five. Detection is disabled for more than five consecutive clicks, so the operation ends immediately after the fifth single click is detected.

[0081] S4, T3 timeout is detected. T3 timeout means that no finger tapping action is detected again within 600ms, that is, the user action is completed, and this time it is a single-click operation.

[0082] S5: If no finger tapping is detected within the specified time, the current cumulative tapping count (i.e., one in S2) is used as the ring tapping count. Entering this step indicates that the user has single-clicked their finger, which corresponds to a single-click operation on the glasses, notifying the glasses to perform the corresponding action. The state is reset to default, and the next single-click operation is detected. Table 4 shows the corresponding operations of the ring and the smart glasses:

[0083] Steps S6, S11, S16, and S21 waveform detection: same as step S1.

[0084] Steps S7, S12, and S17 state settings: same as step S2.

[0085] Steps S8, S13, and S18 start timer T3: same as step S3.

[0086] Steps S9, S14, S19 T3 timeout: same as step S4.

[0087] Step S5: Double-click: This step indicates that the user has double-clicked their finger. This corresponds to a double-click on the glasses, notifying them to perform the corresponding action. The glasses are reset to default and continue detecting the next single-click.

[0088] Step S5: Continuous 3-click: Entering this step indicates that the user has 3 consecutive clicks, which corresponds to the upward swipe operation of the glasses, notifying the glasses to perform the corresponding action. Reset the state to default and continue to detect the next single-click operation.

[0089] Step S5: Continuous 4-click: Entering this step indicates that the user has clicked their finger 4 times in a row, which corresponds to a swipe down operation on the glasses, notifying the glasses to perform the corresponding action. Reset the state to default and continue detecting the next single-click operation.

[0090] Step S5: 5 consecutive taps: This step indicates that the user has tapped their finger five times in a row, which corresponds to a long press on the glasses, notifying the glasses to perform the corresponding action. The state is reset to default, and the next single tap is detected.

[0091] In this possible implementation, a specified time period is set, finger tapping actions within the specified time period are detected, and the cumulative number of tappings is gradually determined to identify control instructions transmitted by the user through different numbers of tappings.

[0092] In one possible embodiment, collecting voice data and glasses inertial measurement unit perception data from the smart glasses can be performed according to the following steps: collecting voice commands from the smart glasses, using a pre-trained neural network to recognize the voice commands, obtaining command words and the confidence of the command words, determining the noise scene category and noise energy according to the voice commands, and splicing the command words, the confidence, the noise scene category and the noise energy to obtain voice data; collecting glasses motion data from the glasses inertial measurement unit, and determining the head motion vector as the glasses inertial measurement unit perception data according to the glasses motion data and a preset glasses motion type.

[0093] In this possible implementation, the voice commands collected from the smart glasses are input into a pre-trained neural network. The pre-trained neural network is used to recognize the voice commands and obtain the command words and the confidence level of the command words. The execution level is used to characterize the reliability of the command words. The pre-trained neural network can be used to process audio data, including but not limited to the following neural networks: recurrent neural networks (RNNs), Transformers (a neural network architecture for processing sequence data), etc. In one possible implementation, for example, the recognized command word is "confirm" and the confidence level is 0.9.

[0094] In a voice command, when the noise intensity exceeds the intensity threshold or the noise type is greater than the type threshold, the noise scene category of the voice command can be determined as a noise scene; otherwise, the noise scene category of the voice command can be determined as a quiet scene, that is, the noise scene category is determined according to the voice command. The noise energy can be used to characterize the intensity of the noise, and the signal-to-noise ratio can be used as the noise energy. The command word, the confidence level, the noise scene category, and the noise energy are concatenated to obtain voice data. For example, in a possible implementation, referring to Table 7, according to the voice command of the smart glasses, information such as the command word, the voice recognition confidence level, the noise scene, and the signal-to-noise ratio can be obtained. After concatenating these information, the dimension of the voice data is 35. Among them, the recognized command word is "confirm", the confidence level is 0.9, the noise scene category is 1 (0 is the quiet scene, 1 is the noise scene), and the signal-to-noise ratio is 80 dB. Then the voice data is: [101, 202, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0.9, 1, 80]. Assume that 101 is the character code of "que" and 202 is the character code of "ren".

[0095] command word Command word confidence Noise scene Noise Energy Dimensions 32 1 1 1

[0096] Table 7

[0097] In an embodiment of the present application, the glasses inertial measurement unit collects glasses motion data. According to the motion of the glasses, the actual operation of the user is determined among the preset glasses motion types, and then the head motion vector is determined. For example, the preset motion types may include, but are not limited to, one or several of: no motion, nodding, and shaking the head. The dimension of the head motion vector is determined according to the preset motion types. For example, the preset motion types include the data in Table 8:

[0098] No action nod Shake your head

[0099] Table 8

[0100] The head motion vector is 3 (3-dimensional). For example, when the glasses motion data is shaking the head, the head motion vector is: [0, 0, 1]. Another example is that when the user does not operate the glasses at all, the head motion vector is: [1, 0, 0]. It should be noted here that all data formats can be integer types for easy fusion with other data, and all are set to single-precision floating-point formats.

[0101] It should be noted that in an embodiment of the present application, the glasses IMU sensor includes a gyroscope and an accelerometer. The gyroscope measures the angular velocity and has high dynamic characteristics. It is a device that indirectly measures the angle. It measures the derivative of the angle,

[0102] That is, angular velocity. Angles are obtained by integrating angular velocity over time. Due to the effects of noise and other errors, these errors accumulate during integration, ultimately causing low-frequency interference and drift in the gyroscope. Gyroscopes measure instantaneous angular velocity, but drift errors can occur due to factors such as temperature fluctuations, friction, and unstable torques. Integrating gyroscope measurement data yields tilt information relative to the vertical direction, providing fast dynamic response. However, drift errors accumulate over time and grow indefinitely. While gyroscopes offer sufficient bandwidth and excellent dynamic performance, their static output is significantly affected by drift errors.

[0103] Accelerometers measure the direction of current acceleration (including acceleration due to gravity). Their measurement principle makes them sensitive to high-frequency signals, making them susceptible to high-frequency interference in vibrating environments. Accelerometers measure acceleration related to inertia, including rotation, gravity, and linear acceleration. Integrating the measured data yields linear velocity, and quadratic integration yields linear displacement. However, the drift error introduced by integration accumulates over time, growing indefinitely and causing inaccurate data. Accelerometers can also calculate inclination angles through trigonometric calculations, but the output signal is susceptible to noise.

[0104] The advantages and disadvantages of gyroscopes and accelerometers are compared in Table 9:

[0105]

[0106] Table 9

[0107] Since the gyroscope and accelerometer have their own advantages and disadvantages, the fusion of these two sensors can obtain more accurate attitude data. The essence of attitude solution is to fuse the 3-axis (X, Y, Z axis) gyroscope acceleration sensor to obtain the 3-axis attitude of the current object. In the embodiment of the present application, the attitude solution algorithm can adopt the following algorithms, including but not limited to: complementary filtering algorithm, gradient descent algorithm and Kalman filtering algorithm. For example, the typical complementary filtering fusion algorithm formula is: Angle = 0.98*(Angle+Gyro*dt)+0.02*Acc. Among them: Angle is the angle obtained after complementary filtering, Gyro is the angular velocity obtained by the gyroscope part, Acc is the angle of the acceleration sensor part after the unit is converted by the inverse tangent function atan(), dt is the operating period of the filter, 0.98 and 0.02 are weighting coefficients α and (1-α).

[0108] It should be noted that the smart glasses themselves have an IMU (Inertial Measurement Unit) that can detect basic head movements, such as nodding and shaking the head. Leveraging this IMU, when the noise energy in a noisy scene exceeds a preset threshold, the user can nod or shake their head for secondary confirmation when a command word is detected by voice.

[0109] In this step, multi-dimensional data can be used to more comprehensively describe the language data and the glasses inertial measurement unit perception data, thereby facilitating the pre-trained large model to determine the user's true intention.

[0110] Therefore, in one possible embodiment, the ring inertial measurement unit includes an accelerometer and a gyroscope; the waveform data is detected according to the following steps: if any dimension sensor data of the accelerometer has a fluctuation with an amplitude greater than a first given threshold within a given time length, and the amplitude of the fluctuation of all dimension sensor data of the gyroscope is less than a second given threshold, it is determined that a finger tapping action is detected; wherein the first given threshold is greater than the second given threshold.

[0111] In this possible implementation, the ring inertial measurement unit includes an accelerometer and a gyroscope. Therefore, the accelerometer can be used to collect data in three directions (or three dimensions), and the gyroscope can be used to collect data in three directions (or three dimensions). On this basis, the movement of the finger can be judged by the fluctuation of each dimension. Specifically, if the sensor data of any dimension of the accelerometer has a fluctuation with an amplitude greater than a first given threshold within a given time length, and the amplitude of the fluctuation of the sensor data of all dimensions of the gyroscope is less than a second given threshold, it is determined that a finger tapping action is detected; wherein, the first given threshold is greater than the second given threshold. In an embodiment of the present application, the fluctuation can be the appearance of a peak or a trough. The given time length can be determined based on the range duration of the peak or trough, for example, it can be 300 milliseconds.

[0112] The following is a specific example. The smart ring has an accelerometer and a gyroscope, each with three dimensions x / y / z. After testing, when the finger taps, the accelerometer is more sensitive and the amplitude is larger. Figure 5 A waveform detection shown Figure 1 , Figure 5 The data of the three axes of the accelerometer and gyroscope collected by four single clicks are shown in the figure, among which the amplitude of the acceleration x data is the largest, and each peak-to-valley is relatively greater than 2. During finger tapping, there is always one vibration range that is relatively the largest among the three collected data of the accelerometer. Since the finger movements in the relevant scenarios are known in advance, that is, it is known when the user uses the ring instead of the glasses button, there are only single-click and multi-click movements for a period of time before and after, and the fingers will not make other movements, so there is relatively little interference. Through experimental observation, the basis for detection is that any dimension of the accelerometer has an obvious and relatively large peak or trough within a certain period of time, the peak or trough range time Tr is limited to 300ms, and the peak-to-trough amplitude is greater than the threshold value V (V=2). The interval between multiple clicks will not exceed the threshold value Tc (Tc=300ms).

[0113] In this possible implementation, it is determined whether there is a tapping action by the finger by comprehensively analyzing the data of each dimension collected by the accelerometer and the gyroscope.

[0114] In one possible implementation, collecting sensor data from the inertial measurement unit of the smart ring and generating waveform data corresponding to the sensor data can be performed according to the following steps: collecting the accelerometer and gyroscope sensor data; using a low-pass filter to filter out fluctuations within a given fluctuation range in the accelerometer and gyroscope sensor data; and generating waveform data based on the filtered sensor data.

[0115] In this possible implementation, the x-axis data collected by the accelerometer may be as follows: 2.14, 2.14, 2.21, 2.14, 2.15, 2.17, 2.13, 2.14, 2.14, 2.21, 2.14, 2.15, 2.17, 2.14, 2.14, 2.21, 2.14, 2.15, 2.17, 2.13, 2.14. From the above accelerometer sampling values, there is very little prominent high-frequency interference or noise, except for slight errors on some values. For example, when the finger is not moving, the x-axis value of the accelerometer has a small fluctuation, and the fluctuation range is usually within 10%. Considering that these slight fluctuations have little effect on the results, the slight fluctuations can be directly ignored, for example, ignoring the value after the floating-point value in the sampling result, or using a low-pass filter to filter out similar high-frequency interference. In the embodiment of the present application, the given fluctuation range can be set according to experience or actual needs, and the embodiment of the present application does not specifically limit this.

[0116] In this step, a low-pass filter is used to filter out a certain degree of high-frequency interference, which can reduce the impact of noise on the data processing process and improve the accuracy of the processing results.

[0117] In order to facilitate the comparison of the peak and trough values ​​with the preset threshold values, in a possible implementation, the waveform data corresponding to the sensor data is generated, which can be performed according to the following steps: obtaining given sliding window parameters, and determining the sensor data within a sliding window in the sensor data according to the sliding window parameters; calculating the average value of the sensor data within the sliding window, and taking the difference between the sampling value of the current sensor data and the average value as the normalization result corresponding to the current sensor data; and generating waveform data according to the normalization result.

[0118] In this possible implementation, the sliding window parameter is used to determine the number of sensor data included in the sliding window. For example, if the sliding window parameter is 20, 20 sensor data are processed using the sliding window each time, that is, the given sliding window parameter is obtained, and according to the sliding window parameter, the sensor data within a sliding window is determined in the sensor data. Afterwards, each currently adopted value is calculated according to the following formula: current sampling value = current sampling value - average value of 20 values ​​in the sliding window, that is, the average value of the sensor data in the sliding window is calculated, and the difference between the sampling value of the current sensor data and the average value is used as the normalization result corresponding to the current sensor data. Waveform data is generated according to the normalization result, so that the center point of the collected data is at point 0, which makes it easier to compare data. It should be noted that low-pass filtering and normalization operations can be implemented in combination.

[0119] In this step, the processing efficiency of the data comparison process can be improved by normalization. For example, in one possible implementation, a new sequence of sampling values ​​is obtained after low-pass filtering and normalization. A given detection window, for example, a 300ms detection window is used to move from front to back on the sampling value, one sampling value at a time, to find out whether the absolute value of the maximum value in the detection window exceeds the threshold V. If the threshold V is exceeded, the sampling value before V changes from small to large, and the sampling value after V changes from large to small. If it is a trough, on the contrary, the value before V changes from large to small, and the value after V changes from small to large. If it conforms to this rule, it is considered that a finger tapping action is detected.

[0120] In a possible implementation, obtaining the interaction state configuration data corresponding to the application scenario may be performed according to the following steps: determining the current state of the application scenario; and determining the interaction state configuration data as no interaction, waiting for confirmation, or waiting for selection based on the current state.

[0121] In this possible implementation, the interaction state configuration data may also be configured as one-dimensional data. For example, it may include one-dimensional data for at least the following states: default interaction (no specific interaction): 0; waiting for confirmation: 1; waiting for selection: 2. The current state of the application scenario is determined based on the interaction data between the user and the smart glasses. For example, if the current state requires the user to confirm whether to perform the next operation, the interaction state configuration data is determined to be waiting for confirmation: 1. If the current state requires the user to confirm which operation to perform, the interaction state configuration data is determined to be waiting for selection: 2.

[0122] In one possible implementation, different application scenarios are configured with different control instruction sets; the control instruction sets include at least one or more of the following operations: no operation, confirm, cancel, play, stop, option, slide up, slide down, shut down and restart.

[0123] In this possible implementation, for example, if the current application scenario is music playback, the control command set supported by this application scenario may include the following commands: "Play, Stop, Volume Up, Volume Down." If the current application scenario is screen-brightening mode, the control command set supported by this application scenario may include "Screen Brighter, Screen Dim." The control command set supported by all application scenarios is "Power On, Power Off." If the application scenario requires user confirmation, such as a confirmation scenario at the start of navigation, the control command set supported by this application scenario is "Confirm, Cancel."

[0124] When smart glasses enter a preset application scenario, they respond to specific commands. When entering a specific scenario, various devices and sensors begin detecting commands. For example, in a music player interface, they start detecting commands such as "play, stop, volume up, volume down."

[0125] The following combination Figure 3 The flowchart shown integrates multiple detection methods to illustrate the possible arbitration process. It should be noted that in the embodiments of this application, different interactive command detection methods have their own characteristics. Among them, the ring button control method is the fastest and most accurate, and will not cause misjudgment; voice commands are easily affected by environmental noise; and the tapping action of the finger wearing the ring takes the longest time to detect. Several control command detection methods can be used in full or in part, and their priorities can be manually preset or automatically and dynamically adjusted.

[0126] Step S01 triggers detection:

[0127] Smart glasses enter specific application scenarios and respond to specific commands. For example, in the music playback scenario, they support "play, stop, volume up, volume down" commands. Bright screen mode supports "brighten screen, dim screen," and all scenarios support "power on, power off." Confirmation scenarios, such as those at the start of navigation, support "confirm, cancel."

[0128] When entering a specific scenario, various devices and sensors start to detect commands. For example, in the music playback interface, they start detecting commands such as "play, stop, volume up, volume down".

[0129] After detection begins, two timers, T1 and T2, are started: T1 = 2 seconds, T2 = 3 seconds. T1 is the key detection timeout timer. If no key is detected for more than 2 seconds, it is assumed that the user has not pressed a key. T2 is the timeout timer for voice command recognition, head movement detection, and finger tap detection. If it exceeds 3 seconds, it means that there is no detection result for these.

[0130] Step S02: Ring button detection:

[0131] There are usually three or more buttons, one of which supports "confirm, cancel, play, stop, power on, and power off". For example, on the music playback interface, pressing once is "confirm / play / stop", pressing twice is "cancel", pressing and holding for 3 seconds is "power off", and pressing and holding for 5 seconds is "power on". The other two support "volume adjustment and screen brightness adjustment".

[0132] The key detection is relatively fast, and the whole process usually takes less than 2 seconds. Therefore, a 2-second timer T1 is started after the detection begins. When the timer expires, it is considered that the user has not pressed a key.

[0133] Step S03: Ring finger tapping detection:

[0134] Step S04: Voice command detection:

[0135] The glasses start voice command recognition detection and start a 3-second timer T2.

[0136] Command word recognition is usually implemented based on a small-scale neural network and includes the following steps:

[0137] (1) Microphone recording, obtain audio data frame by frame, a frame of data can be 10ms, 14ms, 20ms.

[0138] (2) Audio feature extraction: Calculate MFCC (Mel Frequency Cepstral Coefficents) or FBANK (Filter Bank, a type of audio feature) features for each frame of data, convert the time domain data to the frequency domain, and obtain the frequency domain features of each frame of audio. MFCCs are cepstral parameters extracted in the Mel-scale frequency domain and are a feature widely used in automatic speech and speaker recognition.

[0139] (3) MFCC / FBANK is input to the trained neural network. Depending on the neural network, the input can be a single frame of FBANK features or multiple frames of FBANK features. If the input is a single frame, the neural network infers the probability of which word the current frame belongs to. If the input is multiple frames, such as 1 second of audio data, the neural network infers the probability of which command the data belongs to.

[0140] This step detects the results of the command words recognized in the current time period. If they match the current application, the command word recognition is considered successful.

[0141] Step S05 T1 times out or there is a result judgment:

[0142] Entering this step means that T1 has timed out or the key detection has a result. If it times out, it means that the user has not pressed the key and the result is ignored; if the key operation is detected and the result matches the current application scenario, the result is considered valid and the process proceeds to step S011.

[0143] Step S06: Notify other devices to detect and execute control commands:

[0144] Entering this step means that a key operation command has been detected. Because the key has the highest priority and the shortest detection time, the other detection methods have not yet produced any results. First, stop the other three detection methods to save power. Then notify the smart glasses to execute the corresponding control command.

[0145] Step S07 T2 times out or there is a result judgment:

[0146] If T2 times out, it means that no result is detected. If there is a command result, it is combined with the current application scenario to see if it matches. If it matches, the result is considered valid and the process goes to step S011.

[0147] For example, in a music playback scenario, if the user says "turn up the volume", the command is considered valid, but the "confirm" is unreasonable and invalid.

[0148] Step S08 T2 times out or there is a result judgment:

[0149] Same as S07, if there is a result, go to step S09.

[0150] Step S09: Noise scene judgment:

[0151] This step determines whether the current scene is noisy. Noise determination can be a separate module, or in the simplest form, it can directly determine the sound energy over a period of time when the wearer is not speaking, or the energy of certain frequency bands, such as the average energy within 500ms. If it exceeds a certain threshold, it is considered a noise scene.

[0152] If there is noise, go to step S010; if there is no noise, go directly to step S011.

[0153] Step S010: Glasses IMU detection secondary confirmation (head movement detection):

[0154] In this step, the glasses enable the IMU to detect head movements for secondary confirmation.

[0155] By detecting the vertical and horizontal movements of the gyroscope and accelerometer on the glasses within a specific time interval, it can be detected whether the head movement is a nod or a shake. If the head nods, the next step will be taken; if the head shakes, the result of the current voice command will be ignored.

[0156] The time interval is within 3 seconds from the start of detection by the glasses application.

[0157] Step S011: Multimodal result arbitration:

[0158] Entering this step indicates that all or some of the finger tap and voice command detection results are available. These results need to be arbitrated and the best, unique result selected for processing. Considering that there may be inconsistencies in the detection results from multiple methods, you can manually set priorities or preset conditions to allow the system to automatically arbitrate multiple results.

[0159] The following are several arbitration methods based on preset logic:

[0160] (1) If the user presets the same priority for voice interaction commands and finger tapping, the detection results of the two methods are considered and the same result is selected as the final result. For example, if the music playback scene interface detects that the voice command is "play", and the finger tapping is detected as a single click (corresponding to "play"), it is finally considered to be the "play" command. For another example, if the results of the two detection methods are different, no command is executed, or the two detection results are displayed on the glasses, and the user is prompted by voice to select one of them. The IMU of the glasses can also be used for secondary confirmation of the head movement, similar to step S010.

[0161] (2) If the user presets different priorities for voice interaction commands and finger tapping, the detected results will be used to select the final command according to the preset priority. For example, if the preset priority is: voice interaction command > finger tapping, and both methods have results, the result of the voice interaction command will be selected first, and if there is no voice command, the result of the finger tapping will be selected.

[0162] (3) Automatically adjust the priority based on the noise environment. If the current environment is noisy, the voice command priority is lowered to the lowest, and the finger tapping detection result is given priority. If the scene is quiet, the voice command priority is adjusted to the highest. This requires another audio module to detect the current environmental noise level, which usually runs on smart glasses because they have microphones.

[0163] In addition to the above arbitration method, a neural network based on deep learning can also be used to infer the user's interaction intention. For specific steps, please refer to the above steps S101-S105, which will not be repeated in this embodiment of the present application.

[0164] Step S012 executes the control command.

[0165] Corresponding to the application scenario and method of the method provided in the embodiment of the present application, the embodiment of the present application also provides a smart glasses control device. Figure 8 FIG2 is a block diagram of a smart glasses control device according to an embodiment of the present application. The device may include:

[0166] The determination module 801 is used to determine whether the smart glasses have entered a preset application scenario and obtain the interaction state configuration data corresponding to the application scenario; the first acquisition module 802 is used to collect voice data and glasses inertial measurement unit perception data from the smart glasses; the second acquisition module 803 is used to collect ring inertial measurement unit perception data and ring button data from a smart ring; the splicing module 804 is used to splice the recognition data of the application scenario, the interaction state configuration data, the voice data, the glasses inertial measurement unit perception data, the ring inertial measurement unit perception data and the ring button data into a one-dimensional data group, and input the one-dimensional data group into a pre-trained large model for inference to obtain the user's interaction command; the execution module 805 is used to obtain and execute the interaction command.

[0167] According to the embodiments of the present application, based on the characteristics of various operations performed by the user, based on the different application scenarios entered by the smart glasses, and based on different interaction state configuration data, the user's interaction commands are automatically generated, and the smart glasses obtain and execute the interaction commands, thereby ensuring that in different application scenarios, the smart glasses and the user can interact as normally as possible, determine the user's true intentions, and reduce the impact of the external environment on the user's control of the glasses operation.

[0168] In one possible implementation, collecting the perception data of the finger ring inertial measurement unit includes: collecting sensor data of the finger ring inertial measurement unit of the smart finger ring, and generating waveform data corresponding to the sensor data; detecting the waveform data, and in response to detecting a finger tapping action, setting the current cumulative number of taps of the finger ring, if the finger tapping action is not detected again within a specified time period, using the current cumulative number of taps as the number of ring taps, otherwise resetting the current cumulative number of taps of the finger ring; and setting the number of ring taps as the perception data of the finger ring inertial measurement unit.

[0169] In one possible implementation, voice data and glasses inertial measurement unit perception data are collected from the smart glasses, including: collecting voice commands from the smart glasses, using a pre-trained neural network to recognize the voice commands, obtaining command words and confidence levels of the command words, determining noise scene categories and noise energy based on the voice commands, and splicing the command words, the confidence levels, the noise scene categories, and the noise energy to obtain voice data; collecting glasses motion data from the glasses inertial measurement unit, and determining a head motion vector as the glasses inertial measurement unit perception data based on the glasses motion data and a preset glasses motion type.

[0170] In one possible implementation, the ring inertial measurement unit includes an accelerometer and a gyroscope; the waveform data is detected, including: if any dimension sensor data of the accelerometer has a fluctuation with an amplitude greater than a first given threshold within a given time length, and the amplitude of the fluctuation of all dimension sensor data of the gyroscope is less than a second given threshold, it is determined that a finger tapping action is detected; wherein the first given threshold is greater than the second given threshold.

[0171] In one possible implementation, sensor data from an inertial measurement unit of the smart ring is collected to generate waveform data corresponding to the sensor data, including: collecting accelerometer and gyroscope sensor data; using a low-pass filter to filter out fluctuations within a given fluctuation range in the accelerometer and gyroscope sensor data; and generating waveform data based on the filtered sensor data.

[0172] In one possible implementation, waveform data corresponding to the sensor data is generated, including: obtaining given sliding window parameters, and determining sensor data within a sliding window in the sensor data according to the sliding window parameters; calculating an average value of the sensor data within the sliding window, and using a difference between a sampling value of current sensor data and the average value as a normalization result corresponding to the current sensor data; and generating waveform data based on the normalization result.

[0173] In a possible implementation, obtaining the interaction state configuration data corresponding to the application scenario includes: determining a current state of the application scenario; and determining, according to the current state, whether the interaction state configuration data is no interaction, waiting for confirmation, or waiting for selection.

[0174] In one possible implementation, different application scenarios are configured with different control instruction sets; the control instruction sets include at least one or more of the following operations: no operation, confirm, cancel, play, stop, option, slide up, slide down, shut down and restart.

[0175] The functions of each module in each device in the embodiment of the present application can be referred to the corresponding description in the above method, and have corresponding beneficial effects, which will not be repeated here.

[0176] Figure 9 FIG. 1 is a block diagram of an electronic device for implementing an embodiment of the present application. Figure 9 As shown, the electronic device includes: a memory 901 and a processor 902. The memory 901 stores a computer program that can be run on the processor 902. When the processor 902 executes the computer program, the method in the above embodiment is implemented. The number of the memory 901 and the processor 902 can be one or more.

[0177] The electronic device also includes:

[0178] The communication interface 903 is used to communicate with external devices and perform data exchange transmission.

[0179] If the memory 901, processor 902, and communication interface 903 are implemented independently, the memory 901, processor 902, and communication interface 903 can be connected to each other via a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 9 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.

[0180] Optionally, in a specific implementation, if the memory 901, the processor 902 and the communication interface 903 are integrated on a chip, the memory 901, the processor 902 and the communication interface 903 can communicate with each other through an internal interface.

[0181] An embodiment of the present application provides a computer-readable storage medium storing a computer program, which implements the method provided in the embodiment of the present application when the program is executed by a processor.

[0182] An embodiment of the present application provides a computer program product, wherein the computer program product includes a computer program, and when the computer program is executed by a processor, the method provided in the embodiment of the present application is implemented.

[0183] An embodiment of the present application also provides a chip, which includes a processor for calling and executing instructions stored in the memory from the memory, so that a communication device equipped with the chip executes the method provided in the embodiment of the present application.

[0184] An embodiment of the present application also provides a chip, including: an input interface, an output interface, a processor and a memory. The input interface, the output interface, the processor and the memory are connected through an internal connection path. The processor is used to execute the code in the memory. When the code is executed, the processor is used to execute the method provided in the embodiment of the application.

[0185] It should be understood that the processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc. It is worth noting that the processor may be a processor that supports the Advanced RISC Machines (ARM) architecture.

[0186] Furthermore, optionally, the above-mentioned memory may include a read-only memory and a random access memory. The memory may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may include a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may include a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available. For example, static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM) and direct memory bus random access memory (DR RAM).

[0187] In the above embodiments, all or part of the embodiments may be implemented using software, hardware, firmware, or any combination thereof. When implemented using software, all or part of the embodiments may be implemented in the form of a computer program product. A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another.

[0188] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0189] In the description of this specification, the reference terms "one embodiment," "some embodiments," "example," "specific example," or "some examples" mean that the specific features, structures, materials, or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. Moreover, the specific features, structures, materials, or characteristics described may be combined in any appropriate manner in any one or more embodiments or examples. In addition, those skilled in the art may combine and combine different embodiments or examples described in this specification, as well as features of different embodiments or examples, unless they are mutually inconsistent.

[0190] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one such feature. Throughout the description of this application, "plurality" means two or more, unless otherwise specifically defined.

[0191] Any process or method described in the flowchart or otherwise described herein can be understood to represent a module, segment or portion of code comprising one or more executable instructions for implementing the steps of a specific logical function or process. The scope of the preferred embodiments of the present application includes other implementations in which the functions may be performed in a different order than shown or discussed, including performing the functions substantially simultaneously or in reverse order depending on the functions involved.

[0192] The logic and / or steps described in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, apparatus or device (such as a computer-based system, a system including a processor or other system that can fetch instructions from an instruction execution system, apparatus or device and execute instructions), or used in combination with such instruction execution systems, apparatuses or devices.

[0193] It should be understood that various parts of the present application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. All or part of the steps of the above embodiment method can be completed by instructing the relevant hardware through a program, which can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.

[0194] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing module, or each unit may exist physically separately, or two or more units may be integrated into a single module. The aforementioned integrated modules may be implemented in the form of hardware or in the form of software functional modules. If the aforementioned integrated modules are implemented in the form of software functional modules and sold or used as independent products, they may also be stored in a computer-readable storage medium. The storage medium may be a read-only memory, a magnetic disk, or an optical disk, etc.

[0195] The above is merely an exemplary embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily conceive of various modifications or substitutions within the technical scope described in this application, and such modifications or substitutions should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A method for controlling smart glasses, comprising: Determining that the smart glasses have entered a preset application scenario, and obtaining interaction state configuration data corresponding to the application scenario; Collecting voice data and glasses inertial measurement unit perception data from the smart glasses; Collect ring inertial measurement unit perception data and ring button data from a smart ring; splicing the recognition data of the application scenario, the interaction state configuration data, the voice data, the glasses inertial measurement unit perception data, the ring inertial measurement unit perception data, and the ring button data into a one-dimensional data group, and inputting the one-dimensional data group into a pre-trained large model for inference to obtain the user's interaction command; Obtain and execute the interactive command.

2. The method according to claim 1, wherein Collects sensor data from the ring's inertial measurement unit, including: Collecting sensor data of an inertial measurement unit of the smart ring and generating waveform data corresponding to the sensor data; detecting the waveform data, and in response to detecting a finger tapping action, setting a current cumulative tapping count of the ring; if no finger tapping action is detected again within a specified time period, using the current cumulative tapping count as the ring tapping count; otherwise, resetting the current cumulative tapping count of the ring; and The number of ring taps is set as the perception data of the ring inertial measurement unit.

3. The method according to claim 1, wherein Collecting voice data and glasses inertial measurement unit perception data from the smart glasses, including: collecting voice commands from the smart glasses, recognizing the voice commands using a pre-trained neural network to obtain command words and confidence levels of the command words, determining noise scene categories and noise energy based on the voice commands, and concatenating the command words, the confidence levels, the noise scene categories, and the noise energy to obtain voice data; Glasses motion data is collected from the glasses inertial measurement unit, and a head motion vector is determined as the glasses inertial measurement unit perception data according to the glasses motion data and a preset glasses motion type.

4. The method according to claim 2, wherein: The finger ring inertial measurement unit includes an accelerometer and a gyroscope; Detecting the waveform data includes: If any dimension sensor data of the accelerometer fluctuates with an amplitude greater than a first given threshold within a given time period, and the amplitude of the fluctuation of all dimension sensor data of the gyroscope is less than a second given threshold, it is determined that a finger tapping action is detected; wherein the first given threshold is greater than the second given threshold.

5. The method according to claim 4, wherein Collecting sensor data of the ring inertial measurement unit of the smart ring and generating waveform data corresponding to the sensor data includes: Collecting the accelerometer and gyroscope sensor data; Using a low-pass filter, filtering out fluctuations within a given fluctuation range in the accelerometer and gyroscope sensor data; Waveform data is generated based on the sensor data after filtering.

6. The method according to claim 2, wherein: Generating waveform data corresponding to the sensor data, including: Obtaining given sliding window parameters, and determining sensor data within a sliding window in the sensor data according to the sliding window parameters; Calculating an average value of the sensor data within the sliding window, and taking a difference between a sampling value of the current sensor data and the average value as a normalized result corresponding to the current sensor data; Waveform data is generated according to the normalization result.

7. The method according to claim 1, wherein Obtaining the interaction state configuration data corresponding to the application scenario includes: Determining a current state of the application scenario; The interaction state configuration data is determined according to the current state to be no interaction, waiting for confirmation, or waiting for selection.

8. The method according to claim 7, wherein: Different application scenarios are configured with different control instruction sets; The control instruction set includes at least one or more of the following operations: no operation, confirm, cancel, play, stop, option, slide up, slide down, shut down and restart.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory, wherein the processor implements the method according to any one of claims 1 to 8 when executing the computer program.

10. A computer-readable storage medium, wherein a computer program is stored in the computer-readable storage medium, and when the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.