Interactive control system and interactive control method

TWI934799BActive Publication Date: 2026-08-01COMPAL ELECTRONICS INC
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
TW · TW
Patent Type
Patents
Current Assignee / Owner
COMPAL ELECTRONICS INC
Filing Date
2025-10-03
Publication Date
2026-08-01

AI Technical Summary

Technical Problem

Current head-mounted devices face challenges in interactive control due to accidental triggering from unintentional movements, leading to power consumption and inconvenience.

Method used

An interactive control system combining a head-mounted device with auxiliary sensors and a processor that uses a dual-condition triggering mechanism, where a user action detected by a motion sensor serves as a pre-activation condition, and auxiliary sensor information confirms the intention, reducing accidental activations.

Benefits of technology

The system provides a more accurate and intuitive interaction experience by filtering out unconscious actions, minimizing power consumption and improving response accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure TWG2TB001904144_001
    Figure TWG2TB001904144_001
  • Figure TWG2TB001904144_002
    Figure TWG2TB001904144_002
  • Figure TWG2TB001904144_003
    Figure TWG2TB001904144_003
Patent Text Reader

Abstract

This invention provides an interactive control system and an interactive control method. The interactive control system includes a head-mounted device, auxiliary sensors, and a processor. The processor identifies motion information detected by the motion sensors of the head-mounted device as corresponding to a target behavior, and obtains auxiliary information through the auxiliary sensors based on the target behavior. Next, the processor determines that the auxiliary information corresponds to a target condition, and obtains a test image through the image capturing device of the head-mounted device based on the target condition. The processor identifies a target object in the test image and generates display content for an information image based on the target object. This enhances the interactive experience.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This invention relates to a human-computer interaction technology, and more particularly to an interactive control method and an interactive control system. [Previous Technology]

[0002] The development of head-mounted devices such as smart glasses has brought users the convenience of overlaying digital information onto the real world. However, current devices still face challenges in interactive control. Traditional activation mechanisms, such as manually touching the frame or specific voice commands, may not be intuitive or convenient to operate in certain situations. For example, a user may simply be fixing their hair or adjusting their glasses but accidentally touch the touchpad; or when they need to quickly capture the scene in front of them, they may not have enough time to give an accurate voice command.

[0003] On the other hand, if the function is triggered solely by the motion sensor of the head-mounted device (e.g., detecting when the user looks up or turns their head), it is easy to cause frequent false activations due to unintentional daily actions, which not only disturbs the user but also causes unnecessary power consumption. [Summary of the Invention]

[0004] The present invention provides an interactive control method and an interactive control system, which can solve the problem of head-mounted devices being easily triggered or not starting up in time, and provide a more intuitive, accurate and low-power interactive experience.

[0005] The interactive control system of this invention includes a head-mounted device, an auxiliary sensor, and a processor. The head-mounted device includes a motion sensor, an image capturing device, and a display. The processor is communicatively connected to the head-mounted device and the auxiliary sensor. The processor is configured to: identify that the motion information detected by the motion sensor corresponds to a target behavior, and obtain auxiliary information through the auxiliary sensor based on the target behavior; determine that the auxiliary information corresponds to a target condition, and obtain a test image through the image capturing device based on the target condition; and identify a target object in the test image, and generate display content of an information image displayed on the display based on the target object.

[0006] The interactive control method of this invention includes (but is not limited to) the following steps: providing a head-mounted device, an auxiliary sensor, and a processor. The head-mounted device includes a motion sensor, an image capturing device, and a display. The processor identifies that the motion information detected by the motion sensor corresponds to a target behavior, and obtains auxiliary information through the auxiliary sensor based on this target behavior. The processor determines that the auxiliary information corresponds to a target condition, and obtains a test image through the image capturing device based on this target condition. The processor identifies a target object in the test image, and generates display content of an information image displayed on the display based on the target object.

[0007] Based on the above, the interactive control system and interactive control method of this invention improve the accuracy of interaction through a dual-condition triggering mechanism. The first user action (i.e., the target behavior) detected by the motion sensor of the head-mounted device serves as the "pre-start" condition, while the second condition (i.e., the target condition) detected by the auxiliary sensor serves as the final confirmation trigger signal. This judgment method, which combines the user's own posture and external sensing information, can effectively filter out simple, unconscious actions, and only activate image capture and information display when the system highly confirms the user's intention, thereby solving the problems of easy accidental touches, power consumption, and untimely response.

[0008] In order to make the above features and advantages of the present invention more apparent and understandable, specific embodiments are described below, and detailed descriptions are given in conjunction with the accompanying drawings.

Implementation Method

[0010] FIG1A is a block diagram of an interactive control system 100 according to an embodiment of the present invention. Referring to FIG1A, the interactive control system 100 includes (but is not limited to) a head-mounted device 110, an auxiliary sensor 120, and a processor 170.

[0011] The head-mounted device 110 is, for example, a head-mounted display, smart glasses, or a virtual reality device. The head-mounted device 110 includes (but is not limited to) a motion sensor 111, an image capturing device 112, a distance sensor 113, and a display 114.

[0012] Figures 1B and 1C are schematic diagrams of the appearance of a head-mounted device 110 (taking smart glasses as an example) according to an embodiment of the present invention. Referring to Figures 1A, 1B, and 1C, the motion sensor 111 may be a gravitational sensor (G-sensor), an accelerometer, a magnetometer, a gyroscope, an inertial measurement unit (IMU), or a combination of the foregoing. In one embodiment, the motion sensor 111 is used to measure the physical state (attitude, or orientation) of the head-mounted device 110 in three-dimensional space and output motion information. The motion information may be, for example, the acceleration vector generated by the accelerometer, the angular velocity generated by the gyroscope, the magnetic field vector generated by the magnetometer, or the comprehensive measurement data generated by the IMU.

[0013] The image capturing device 112 is, for example, an RGB (red-green-blue) camera. In one embodiment, the image capturing device 112 is used to capture a specified environment and generate a test image. That is, the test image is an image generated by captured ambient light. As shown in FIG1B, the image captures the field of view in front of the glasses. The image capturing device 112 captures multiple complete test images at a fixed frame rate, for example, which can be used to identify visual / image features such as color, texture, and shape of target objects in the environment (e.g., a user, or body parts of the user such as head, hands, or torso).

[0014] The distance sensor 113 may be a proximity sensor, an infrared sensor, a LiDAR, or a radar. In one embodiment, the distance sensor 113 is used to detect distance information between the head-mounted device 110 and the wearable device 130. The distance information indicates the distance value between the head-mounted device 110 and the wearable device 130.

[0015] The display 114 may be an LCD, LED display, OLED display, or electronic paper display. In one embodiment, the display 114 is used to display images.

[0016] The head-mounted device 110 may provide an input device for receiving user commands. Referring to FIG1C, in one embodiment, the head-mounted device 110 may further include a power switch 115. The power switch 115 is used to receive a user's pressing operation and thereby turn the display 114 on or off.

[0017] In one embodiment, the head-mounted device 110 may further include a touch panel 116. The touch panel 116 is used to receive touch operations from the user. For example, pressing, swiping, or clicking operations. Touch operations are used, for example, to trigger taking a photo (corresponding to a pressing operation), face search (corresponding to a long press operation), left operation (corresponding to a left swipe operation), right operation (corresponding to a right swipe operation), selection operation (corresponding to a single click operation), returning to the home page or back operation (corresponding to a double click operation), or activating the microphone module (not shown) to perform recording operations to receive voice commands (corresponding to a long press operation).

[0018] In one embodiment, the head-mounted device 110 may further include buttons 117 and 118, which receive press operations from the user. Buttons 117 and 118 may be preset or customized to trigger specific functions or applications of the head-mounted device 110. For example, work schedules, messages, stop, shopping lists, or traffic information.

[0019] Referring to FIG1A, the assistive sensor 120 includes a distance sensor 113 and a second motion sensor 131 of the wearable device 130. In one embodiment, the assistive sensor 120 is used to detect assistive information. The assistive information is, for example, distance information of the distance sensor 113. Alternatively, the assistive information is, for example, second motion information of the second motion sensor 131 (but described below).

[0020] The processor 170 communicates with the head-mounted device 110 and the auxiliary sensor 120, for example, via Wi-Fi, Bluetooth, or other wireless communication technologies. One or more processors 170 may be processors of a smartphone, tablet, and / or head-mounted device 110. The processor 170 may be a central processing unit (CPU), or other programmable general-purpose or special-purpose microprocessor, digital signal processor (DSP), programmable controller, application-specific integrated circuit (ASIC), or other similar components or combinations thereof. In some embodiments, some functions of the processor 170 may be implemented on multiple devices.

[0021] In one embodiment, the interactive control system 100 may further include a wearable device 130. The wearable device 130 is communicatively connected to the head-mounted device 110 and the processor 170, for example, via Wi-Fi, Bluetooth, or other communication technologies. The wearable device 130 is, for example, a smartwatch, smart bracelet, smart ring, or smart necklace, and is intended to be worn on a user's wrist. The wearable device 130 includes (but is not limited to) a second motion sensor 131.

[0022] The implementation and function of the second motion sensor 131 can be referred to the description of the motion sensor 111 above, and will not be repeated here. In one embodiment, the second motion sensor 131 is used to measure the physical state (attitude, or orientation) of the wearable device 130 in three-dimensional space and output second motion information. The second motion information is, for example, the acceleration vector generated by the accelerometer, the angular velocity generated by the gyroscope, the magnetic field vector generated by the magnetometer, or the comprehensive measurement data generated by the IMU.

[0023] Figures 2A and 2B are schematic diagrams illustrating the usage scenarios of the interactive control systems 100-1 and 100-2 according to an embodiment of the present invention. Referring to Figure 2A, the interactive control system 100-1 includes a head-mounted device 110, a regular watch 135 (taking a non-smartwatch as an example), a smartphone 140, and a cloud server 150. The head-mounted device 110 and the smartphone 140 are connected via Bluetooth, and the smartphone 140 can connect to the cloud server 150 via mobile communication (e.g., 2G, 3G, 4G, or 5G mobile communication). The cloud server 150 can run artificial intelligence models. In addition, the head-mounted device 110 can perform image recognition on images captured by the regular watch 130.

[0024] Referring to Figure 2B, the interactive control system 100-2 includes a head-mounted device 110, a wearable device 130 (taking a smartwatch as an example), a smartphone 140, and a cloud server 150. The head-mounted device 110, the wearable device 130, and the smartphone 140 can be connected via Bluetooth, and the smartphone 140 can connect to the cloud server via mobile communication (e.g., 2G, 3G, 4G, or 5G mobile communication).

[0025] Hereinafter, the methods described in embodiments of the present invention will be described in conjunction with the apparatus, elements, or modules shown in Figures 1 to 2B. Each step of this method may be adjusted according to the implementation situation, and is not limited thereto.

[0026] One or more key concepts of this invention are: based on the natural behavior of a user raising their hand to look at a watch, when the user makes this action, the head-mounted device 110 simultaneously displays the corresponding information interface to provide the user with visual information, and through the auxiliary sensor 120 and the image capturing device 112, it avoids other actions from unintentionally activating the head-mounted device 110. Several embodiments will be described in detail below.

[0027] Figure 3 is a flowchart of an interactive control method according to an embodiment of the present invention. Referring to Figure 3, the processor 170 identifies motion information corresponding to a target behavior and obtains auxiliary information through the auxiliary sensor 120 based on the target behavior (step S310). Specifically, the motion information represents the posture or orientation of the head-mounted device 110 or the head wearing the head-mounted device 110 in three-dimensional space. The posture or orientation may correspond to a specific orientation, and the orientation is quantified into an angle value.

[0028] On the other hand, the target behavior corresponds to a specific action / behavior of the head-mounted device 110 or the head wearing the head-mounted device 110. In one embodiment, the target behavior is a head-down behavior. For example, a continuous movement of the head lowering from a forward-facing position. Generally, a head-down behavior occurs when a user raises their hand to look at a watch on their wrist. Therefore, the processor 170 detects this head-down behavior.

[0029] Figure 4 is a flowchart of target behavior detection according to an embodiment of the present invention. Referring to Figure 4, the processor 170 can determine whether the angle value of the motion information is within the angle range of the target behavior (step S410). Specifically, the angle range of the target behavior is, for example, 15~45, 20~40, or 10~35 degrees relative to the head's facing direction, but can still be changed according to design requirements. The processor 170 converts the motion information into an angle value (e.g., through formula conversion, lookup table, or artificial intelligence model inference) and compares this angle value with the angle range of the target behavior.

[0030] When the angle value of the motion information is within the angle range of the target behavior, the processor 170 can obtain auxiliary information through the auxiliary sensor 120 (step S420). For example, the angle value corresponding to the motion information measured by the motion sensor 111 is 35 degrees relative to the reference direction (i.e., the direction the head is facing), and 35 degrees is within the angle range of 15 to 45 degrees. Therefore, the processor 170 sends an auxiliary request command to the auxiliary sensor 120 (e.g., the distance sensor 113 of the head-mounted device 110 and / or the second motion sensor 131 of the wearable device 130), so that the auxiliary sensor 120 performs the measurement operation and sends auxiliary information to the processor 170.

[0031] On the other hand, when the angle value of the motion information is not within the angle range of the target behavior, the processor 170 may prohibit the acquisition of auxiliary information (step S430). For example, the angle value corresponding to the motion information measured by the motion sensor 111 is 10 degrees relative to the reference direction (i.e., the direction the head is facing), and 10 degrees is not within the angle range of 20 to 50 degrees. Therefore, the processor 170 prohibits / does not transmit the auxiliary request command to the auxiliary sensor 120 (e.g., the distance sensor 113 of the head-mounted device 110 and / or the second motion sensor 131 of the wearable device 130), and the auxiliary sensor 120 will not perform measurement operations or transmit auxiliary information, causing the processor 170 to be unable to acquire auxiliary information.

[0032] For example, Figures 5A to 5D are schematic diagrams of target behavior and information images according to an embodiment of the present invention. Referring to Figure 5A, a user wears a head-mounted device 110 (taking smart glasses as an example). Referring to Figure 5B, the head-mounted device 110 first detects a head-down behavior LHH. The motion sensor 111 measures this head-down behavior LHH, and the processor 170 determines that the angle value D1 corresponding to the measured motion information is within the angle range corresponding to the head-down behavior LHH.

[0033] Referring to Figure 3, the processor 170 determines that the auxiliary information corresponds to the target condition and obtains the image to be tested based on the target condition (step S320). Specifically, the target condition is related to the user's continuous movements of rotating the wrist and / or raising the hand. Generally, when a user intends to look at a watch, they will rotate their wrist and / or raise their hand so that their eyes can look directly at the watch face.

[0034] Figure 6 is a flowchart of depth detection according to an embodiment of the present invention. Referring to Figure 6, the auxiliary sensor 120 is the distance sensor 113 of the head-mounted device 110. This embodiment is applicable to the usage scenarios of Figures 2A and 2B. In this case, the auxiliary information is distance information (representing the distance / depth value of the target object, such as the wearable device 130 or other general-purpose watch, relative to the head-mounted device 110), and the target condition is the target depth range. This target depth range corresponds to the distance / depth range of the wearable device 130 or other general-purpose watch relative to the head (e.g., eyes, or forehead) during the continuous action of rotating the wrist and / or raising the hand. For example, 35 to 60 cm, 20 to 50 cm, or 15 to 60 cm, and can be changed according to design requirements.

[0035] The processor 170 may determine whether the distance information corresponds to the target depth range (step S610). Specifically, the processor 170 converts the distance information into a depth / distance value (e.g., through formula conversion, lookup table or artificial intelligence model inference) and compares this depth / distance value with the target depth range of the target conditions.

[0036] When the distance information corresponds to the target depth range, the processor 170 can acquire the image to be measured through the image capturing device 112 (step S620). For example, the depth / distance value corresponding to the distance information (e.g., sensing intensity value or round-trip time) measured by the distance sensor 113 is 40 cm, and 40 cm is within the target depth range of 35 to 45 cm. Therefore, the processor 170 sends an identification request instruction to the image capturing device 112 of the head-mounted device 110, causing the image capturing device 112 to transmit the image to be measured to the processor 170.

[0037] On the other hand, when the distance information does not correspond to the target depth range, the processor 170 may prevent the acquisition of the image to be measured acquired by the image capturing device 112 (step S630). For example, the depth / distance value corresponding to the distance information (e.g., sensing intensity value or round-trip time) measured by the distance sensor 113 is 60 cm, and 60 cm is not within the target depth range of 20 to 50 cm. Therefore, the processor 170 prevents / does not transmit the identification request instruction to the image capturing device 112 of the head-mounted device 110, so that the image capturing device 112 prevents / does not transmit the image to be measured to the processor 170, and causes the processor 170 to be unable to acquire the image to be measured or to ignore the image to be measured.

[0038] Taking Figure 5B as an example, the head-mounted device 110 detects the behavior RHH corresponding to wrist rotation and / or hand raising. The distance sensor 113 measures this behavior RHH, and the processor 170 determines that the depth / distance value DI corresponding to the measured distance information is within the target depth range corresponding to the behavior RHH. Referring to Figure 5C, the image capturing device 112 acquires the image TIM to be measured.

[0039] Figure 7 is a flowchart of motion detection according to an embodiment of the present invention. Referring to Figure 7, the auxiliary sensor 120 is the second motion sensor 131 of the wearable device 130. This embodiment is applicable to the usage scenario of Figure 2B. In this case, the auxiliary information is the second motion information (the posture or orientation of the wearable device 130 in three-dimensional space), and the target condition is the second target behavior. For example, the second target behavior corresponds to the continuous motion of rotating the wrist and / or raising the hand. This second target behavior corresponds to the rotation range of the wearable device 130 relative to a reference direction (e.g., the viewing direction of glasses) during the continuous motion of rotating the wrist and / or raising the hand. For example, 85 to 100 degrees, 90 to 105 degrees, and can be changed according to design requirements.

[0040] The processor 170 may determine whether the second action information corresponds to the second target behavior (step S710). Specifically, the processor 170 converts the second action information into an angle value (e.g., through formula conversion, lookup table or artificial intelligence model inference) and compares this angle value with the rotation range of the second target behavior.

[0041] In one embodiment, the processor 170 may determine whether the angle value of the second action information (e.g., gravitational acceleration, magnetic force, or angular velocity) is within the rotation range of the second target behavior.

[0042] When the second motion information corresponds to the second target behavior, the processor 170 can acquire the image to be measured through the image capturing device 112 (step S720). For example, the angle value corresponding to the second motion information (e.g., gravitational acceleration, magnetic force, or angular velocity) measured by the second motion sensor 131 is 92 degrees relative to the reference direction, and 92 degrees is within the rotation range of 90 to 100 degrees. Therefore, the processor 170 sends an identification request instruction to the image capturing device 112 of the head-mounted device 110, so that the image capturing device 112 transmits the image to be measured to the processor 170.

[0043] On the other hand, when the second motion information does not correspond to the second target behavior, the processor 170 may prevent the acquisition of the image to be tested acquired by the image capturing device 112 (step S730). For example, the angle value corresponding to the second motion information (e.g., gravitational acceleration, magnetic force, or angular velocity) measured by the second motion sensor 131 is 5 degrees relative to the reference direction, and 5 degrees is not within the rotation range of 95 to 110. Therefore, the processor 170 prevents / does not transmit the identification request instruction to the image capturing device 112 of the head-mounted device 110, so that the image capturing device 112 prevents / does not transmit the image to be tested to the processor 170, and causes the processor 170 to be unable to acquire the image to be tested or ignore the image to be tested.

[0044] Taking Figure 5B as an example, the wearable device 130 detects the behavior RHH corresponding to wrist rotation and / or hand raising. The second motion sensor 131 measures this behavior RHH, and the processor 170 determines that the angle value D2 corresponding to the measured second motion information is within the rotation range corresponding to the behavior RHH. Referring to Figure 5C, the image capturing device 112 acquires the image TIM to be measured.

[0045] Referring to Figure 3, the processor 170 identifies the target object in the image to be tested and generates the display content of the information image based on the target object (step S330). Specifically, the target object can be a wearable device 130, a regular watch, a user's hand, a smartphone, a computer, medicine, kitchen utensils, or any predefined identifiable object.

[0046] Target object identification can be based on object detection technology. For example, processor 170 can apply neural network-based algorithms (e.g., YOLO (You Only Look Once), Region Based Convolutional Neural Networks (R-CNN), or Fast R-CNN) or feature matching-based algorithms (e.g., Histogram of Oriented Gradient (HOG), Scale-Invariant Feature Transform (SIFT), Haar, or Speeded Up Robust Features (SURF) feature matching), or computer vision algorithms based on surface color information to achieve object detection. Processor 170 can determine whether the type of object in the image to be tested is a preset target object.

[0047] Figure 8 is a flowchart of image recognition according to an embodiment of the present invention. Referring to Figure 8, the processor 170 can determine whether a target object exists in the image to be tested (step S810). Specifically, the processor 170 can make the determination through image recognition technology (e.g., reference template matching, feature point comparison, or object detection model based on deep learning). When a target object exists in the image to be tested, the processor 170 can generate and display the display content corresponding to the target object through the display 114 (step S820). On the other hand, when no target object exists in the image to be tested, for example, if the image is blurry or an unexpected object is captured, the processor 170 can prohibit / not generate the display content of the information image to avoid presenting irrelevant information (step S830).

[0048] Taking Figure 5C as an example, the image under test TIM contains a target object TO1 (taking a watch or the wearable device 130 in Figure 5B as an example). The display 114 can then display the content of the information image IIM.

[0049] The content displayed in the information image may vary depending on design and / or application requirements. For example, Figures 9A to 9D are schematic diagrams of the display content in different application scenarios according to an embodiment of the present invention. Referring to Figure 9A, the processor 170 detects that the user is looking down and holding a mobile phone (i.e., the target object). The head-mounted device 110 can connect to the application program of the mobile phone system. The display 114 of the head-mounted device 110 displays information related to the current mobile phone activity. For example, mobile phone usage reminders, music lyrics.

[0050] Referring to Figure 9B, the processor 170 detects that the user is looking down and holding a pillbox (i.e., the target object). The display 114 of the head-mounted device 110 displays the medication schedule for the day. In addition, the processor 170 can check and confirm the records, and provide more detailed information display or voice broadcast, which is helpful for elderly people who are forgetful and have poor eyesight.

[0051] Referring to Figure 9C, the processor 170 detects the studio environment (e.g., a computer as the target object). The head-mounted device 110 switches to working mode. The display 114 displays user-customized gadgets, such as meetings, calendars, and / or stocks.

[0052] Referring to Figure 9D, the processor 170 detects movement on the countertop and in the hands (i.e., the target object). The head-mounted device 110 switches to cooking mode. The display 114 shows user-customized applications, such as a cooking timer, recipe notes, and / or nutritional measurements.

[0053] When the target object is a wearable device 130 or a regular watch, the processor 170 can identify that this is a "time query" or "personal information query" scenario. In one embodiment, the content displayed in the information image is one of calendar information, game information, to-do list information, and stock information. For example, FIG10 is a schematic diagram of the display content of an information image according to an embodiment of the present invention. Referring to FIG10, calendar information is, for example, calendar content, work schedule, or to-do list. Widgets are, for example, stock information, shopping lists, or news headlines. Game information is, for example, related to baseball, football, racing, or basketball. These display contents correspond to the applications of the connected wearable device 130 or other mobile devices (e.g., smartphones or tablets). The wearable device 130 or other mobile device can receive a user's selection operation for specific display content, causing the display 114 to display the content corresponding to the selection operation. For example, calendar, widgets, and game information corresponding to a specific time period.

[0054] Figures 11A and 11B are schematic diagrams of the information image activation process according to an embodiment of the present invention. Referring to Figure 11A, this embodiment illustrates how to activate the display of information images by combining head posture and hand movements. The display 114 of the head-mounted device 110 can initially be in a standby state (step S1101), for example, with the screen turned off or at low brightness to save power.

[0055] When the user makes the action of "raising hand and looking down" (step S1102), this action combines the aforementioned head-down behavior with the action of raising the wrist of the wearable device 130 (such as a smartwatch). The motion sensor 111 of the head-mounted device 110 detects the motion information of the head-down behavior, and the second motion sensor 131 of the wearable device 130 also detects the second motion information of raising the hand, and the processor 170 analyzes the image to be tested captured by the image capturing device 112 within the field of view.

[0056] When the processor 170 determines that the combined action meets the preset target behavior and target conditions and that there is a target object in the image to be tested, it will trigger the display of the information image. Finally, the display 114 will be activated and display the corresponding widget content (i.e., the display content of the information image) (step S1103), such as schedule, match, stock, shopping list, or traffic information. In the figure, the widget Wid1 shown on the display 114 is the schedule widget, and the widget Wid2 is the sports match widget.

[0057] Referring to FIG11B, this embodiment illustrates a method of activation via voice command. The display 114 may be in the Home state (step S1111), displaying basic information such as time and battery level. The user can activate the Listening mode by long-pressing the touch panel 116 on the headset 110 (step S1112). In Listening mode, the microphone module of the headset 110 receives the user's voice command (step S1113). For example, the user can say keywords such as "Instant widget," "Start instant widget," "Open instant widget," or "Open widget." After the processor 170 recognizes a valid voice command, it will trigger and display the corresponding widget content on the display 114 after the user releases the touch panel 116 (step S1114) (step S1115).

[0058] Figure 12 is a schematic diagram of the information image browsing process according to an embodiment of the present invention. Referring to Figure 12, the display 114 of the head-mounted device 110 can initially be in a standby state (step S1201). When the user makes the action of "raising hand and looking down" (step S1202), the processor 170 analyzes the image to be tested captured by the image capturing device 112 within the field of view. When the processor 170 determines that this combination of actions meets the preset target behavior and target conditions and that there is a target object in the image to be tested, the display of the information image will be triggered. For example, the display 114 will be activated and display the corresponding widget content (step S1203). After the information image (widget) is activated, the user can further interact with it. For example, in the state of displaying widgets, the user can browse different widgets by swiping left or right on the touch panel 116 of the head-mounted device 110 (step S1205). The information display may switch from a screen showing schedules and match information to a screen showing match and stock information. Users can double-tap the touch panel 116 (step S1206) while browsing widgets to close the information display and return to the home screen (step S1207).

[0059] Alternatively, while browsing the calendar (step S1208), the user can double-click the touch panel 116 (step S1209) to switch to the widget page (step S1210). To return to the previous calendar browsing screen, the user can double-click the touch panel 116 again (step S1211) to return to the calendar screen (step S1208).

[0060] Figure 13 is a flowchart of gesture control according to an embodiment of the present invention. Referring to Figure 13, the processor 170 identifies gesture information of a target object in the image to be tested (step S1310). At this time, the target object is a hand. The gesture information includes, for example, the skeletal shape of the hand, the position of joints / parts, and / or the length and orientation of the joint lines.

[0061] Upon recognizing the gesture information, the processor 170 can then execute the target program corresponding to the reference gesture (step S1320). Figure 14 is a schematic diagram of a reference gesture according to an embodiment of the present invention. As shown in Figure 14, the user can browse the next widget by using the "thumbs-up" gesture G1, or browse the previous widget by using the "thumbs-down" gesture G2 (steps S1401, S1402).

[0062] In one embodiment, the target program may be a schedule browsing, missed call browsing, and unread message browsing, allowing users to interact with information using more intuitive gestures. Reference gestures may also include clenching a fist, extending the thumb and little finger, extending the index finger, and may vary depending on actual needs.

[0063] For example, when the processor 170 recognizes a watch image and a fist gesture from the image under test, the display 114 can display the time schedule. When the processor 170 recognizes a watch image and a fist gesture with the thumb and little finger extended from the image under test, the display 114 can display the call interface. And when the processor 170 recognizes a watch image and a fist gesture with the thumb and index finger extended from the image under test, the display 114 can display the call interface.

[0064] It should be noted that the reference gestures and the content of the corresponding target program can be customized according to actual needs, and the embodiments of the present invention are not limited thereto.

[0065] The wearable device 130 can also be connected to the head-mounted device 110 or other mobile devices (e.g., smartphones or tablets). Users can configure settings on their smartphone screens to select which gadgets to display on the head-mounted device 110. Figures 15A and 15B are schematic diagrams of a user interface according to an embodiment of the present invention. As shown in Figures 15A and 15B, a gadget Wid3 in a game state and a menu providing gadget Wid4 are respectively displayed.

[0066] In one embodiment, the information image may include extended information. For example, Figures 16 and 17 are schematic diagrams of extended information according to an embodiment of the present invention. Referring to Figures 16 and 17, in this scenario, the processor 170 can identify the target person in the image to be tested (e.g., through facial recognition) and generate a transcript and extended information for conversation suggestions based on the target person and voice data (e.g., live dialogue recorded through a microphone module). As shown in Figure 17, when a user meets a customer, the display 114 can display the customer's personal profile (e.g., name, title, or professional experience), summaries of past conversations, and feasible chat topic suggestions. It can even combine voice input to generate response suggestions in real time through an artificial intelligence module.

[0067] Figures 18A to 18D are schematic diagrams of another interactive process according to an embodiment of the present invention. Referring to Figure 18A, this process illustrates a voice-activated method. First, when the display 114 displays the Home screen (step S1801), the user can long press the touch panel 116 on the head-mounted device 110 (step S1802) to put the system into Listening mode (step S1803). In this mode, the user can speak a voice command, such as "Take a look and tell me who she is." After receiving the command and the user releases the touch panel 116 (step S1804), the processor 170 will start and enter the face search application (step S1805), and then perform face detection (step S1806).

[0068] Referring to Figure 18B, this process demonstrates another way to initiate the operation. Similarly, when the display 114 is in the Home state (step S1811), the user can long press the button corresponding to taking a photo (e.g., button 117) (step S1812). This operation will trigger the processor 170 to start and enter the face search application or mode (step S1813), and then perform face detection (step S1814).

[0069] Referring to Figure 18C, this process demonstrates another way to activate the face search function through multi-layered touch operations. In the Home screen state (step S1821), the user can call up the function menu (Show menu) by tapping or other preset gestures (e.g., swiping left or right) (step S1822) (step S1823). In this menu, the user can select the "Friend Memo" option by swiping or tapping (steps S1824, S1825). After entering the Friend Memo interface (step S1826), the user can further select the "Face Search" function by tapping (steps S1827, S1828). Then, the face search mode will be entered (step S1829), and face detection will then be performed (step S1830).

[0070] Referring to Figure 18D, this flowchart details the specific operation steps after entering the face search mode. Regardless of how the face search mode is entered (step S1841), the processor 170 drives the image capturing device 112 to begin face detection (step S1842). The processor 170 continuously analyzes the image to be tested and marks one or more detected faces on the display 114 with bounding boxes or the like, and can assign numbers (step S1843) for the user to select.

[0071] After detecting a face, the user can select a target in two ways. The first is touch selection: the user can swipe left / right to choose (step S1844) on the touch panel 116 on the head-mounted device 110 to browse and select (step S1845) the corresponding bounding box. After receiving the user's operation command for a specific bounding box or face, the processor 170 requests the cloud server 150 to process the corresponding request (step S1846). The second is voice selection: the user can long press the touch panel 116 on the head-mounted device 110 to activate the listening mode (step S1847). In listening mode, the microphone module of the head-mounted device 110 will receive voice commands, such as "Select number 1" (step S1848). After the processor 170 recognizes a valid voice command, it will trigger and request the cloud server 150 to process the corresponding request after the user releases the touch panel 116 (step S1849) (step S1846).

[0072] After receiving the user's selection of a face or its bounding box, the processor 170 searches for the corresponding person's information in a backend database (e.g., cloud server 150 or local storage). Then, the processor 170 makes a judgment to confirm whether matching information has been found (step S1850).

[0073] If the data is successfully found, the processor 170 generates an information image display based on the person's personal information and displays it on the display 114 (step S1851). The user can browse the personal information by swiping left and right on the touch panel 116 on the head-mounted device 110 (step S1852). The user can double-tap the touch panel 116 twice while viewing the personal information (step S1853) to close the information image and return to the home screen (step S1854). On the other hand, if the corresponding data is not found in the database (step S1855), the display 114 displays a pop-up message "No data found" (step S1855). After 3 seconds, the user returns to the home screen (step S1854).

[0074] Figures 19A to 19C are schematic diagrams of another interactive process according to an embodiment of the present invention, specifically illustrating how a user can further interact with derived information after displaying the target person's personal information. Referring to Figure 19A, in the state of displaying personal information (Personal info p1) (step S1901), the user can browse to other pages' personal information (Personal info p2) (step S1903) by swiping left or right (step S1902).

[0075] On another page of personal information, the display 114 provides several suggested interactive options, such as "Show more topic" (step S1904). When the user selects this option, the display 114 pops up a window containing several suggested topics, such as "beach," "castle," and "wine tour." The user can select one of the topics via voice or touch (step S1904). Then, the processor 170 generates more detailed conversation suggestions or extended information based on this topic (step S1905). The user can double-tap the touch panel 116 twice while in personal information mode (step S1906) to close the information image of the conversation suggestions or extended information and return to the interactive options screen (step S1907).

[0076] Referring to Figure 19B, this process demonstrates how to generate a conversation summary. With interactive options displayed (step S1911), the user can tap twice on the touch panel 116 (step S1912). A pop-up window appears on the display 114. The user can swipe left / right to choose on the touch panel 116 on the head-mounted device 110 (step S1913) and tap (step S1914) to select the "Show summary" option. Upon receiving the operation command for this option, the processor 170 begins processing (step S1915), integrating past conversation records or related information with the target person into a summary text, which is then displayed in the pop-up window (step S1916). At this time, the microphone module stops recording. After reviewing the summary, the user can tap twice on the touch panel 116 (step S1916) to perform a face search again (step S1919).

[0077] Referring to Figure 19C, this process demonstrates how to trigger other applications. With interactive options displayed (step S1921), the user can long press the touch panel 116 on the headset 110 (step S1922) to enter Listening mode (step S1923). In this mode, the user can speak a voice command, such as "Play some music," to trigger a multimedia application. After receiving the command and the user releases the touch panel 116 (step S1924), the display 140 displays a pop-up message asking "Are you sure to exit?" (step S1925). The user can then confirm to exit the application or choose Cancel to remain on the current screen. Users can swipe left / right to choose on the touch panel 116 on the head-mounted device 110 (step S1926) and can select the "Cancel" option by tapping (step S1927). After confirmation, the user remains on the music app screen (step S1928) and starts playing music.

[0078] In one embodiment, the processor 170 can identify a target product in the image to be tested and generate derived information based on the target product. The type of the target product is, for example, a handbag or a lunchbox, but is not limited thereto. In this case, the derived information is used for product description.

[0079] For example, Figures 20A and 20B are schematic diagrams of a target object and corresponding display content according to an embodiment of the present invention. Referring to Figure 20A, when a user looks at a lunchbox (i.e., the target product), the image capturing device 112 captures the image to be tested of the lunchbox (step S2001). After the processor 170 identifies the contents of the meal (step S2002), it can obtain its nutritional information from the cloud server 150 or the local database, and display derived information such as calories (e.g., 248 kcal) on the display 114 (step S2003).

[0080] Referring to Figure 20B, when the user gazes at a handbag (i.e., the target product), the image retrieval device 112 retrieves the image to be tested for the handbag (step S2011). After recognizing the handbag (step S2012), the processor 170 may obtain its purchase information from the cloud server 150 or a local database and display derived information such as price, shopping platform (including purchase links), etc. on the display 114 (step S2013).

[0081] In one embodiment, the processor 170 can recognize the original text information in the image to be tested and generate the derived information based on the original text information. The language type of the original information is, for example, English, Chinese, and is not limited to this. At this point, the derived information is used for text translation.

[0082] FIG. Referring to Figure 21A, when the user gazes at a menu containing a foreign language, the image retrieval device 112 retrieves the image to be tested for this menu (step S2101). The processor 170 identifies the original information in the image to be tested via Optical Character Recognition (OCR) technology (step S2102).

[0083] After identifying the original text, processor 170 calls the translation engine to generate the translation. The original text may be displayed by default on the display 114 . The user can switch the display content to translated extended information (text can be arranged into a horizontal presentation format) by performing a right-to-left swipe gesture (From R to L) on the touch panel 116 (step S2103). Conversely, users can switch the display back to the original information by sliding gestures from left to right. This type of interaction allows users to conveniently compare the original text with the translation. Alternatively, the extended information may be displayed on a blank space in the translation area and the position of the translation area may be changed as the posture of the head-mounted device 110 is changed.

[0084] Referring to Figure 21B, there are other extended applications of the translation function. Users can store handwritten or notated note data into the application interface, and the application can classify the person being interviewed according to the time and place.

[0085] For translation and note-taking applications: Users can use handwriting to circle or underline slang terms on paper to inquire about their meaning. For example, (Context: Inquiring about the meaning of a slang term?) User: Tell me what's the meaning of the sentence I marked? Furthermore, users can request more example sentences from the AI ​​model, and these generated example sentences can be added to the user's note database. For example: Give me an example sentences to let me know how Americans use it.

[0086] It should be noted that the generation of the aforementioned extended information may be executed through the artificial intelligence model of the cloud server 170, or it may be executed through the local model of one or more devices of the interactive control system 100. The artificial intelligence model is, for example, a large language model (LLM) or other natural language model, but is not limited thereto.

[0087] In summary, the interactive control system and interactive control method of the present invention establish a dual-condition triggering mechanism by combining the motion sensor and auxiliary sensor of the head-mounted device. This mechanism can accurately determine the user's interactive intention and only activate the image capture and information display function when a specific combination of behaviors such as "looking down at the wrist" or "turning the head to look at other targets" is detected.

[0088] Furthermore, embodiments of the present invention provide at least the following features: Interaction between traditional mechanical watches and AI glasses in virtual reality. Meeting Records: Quick display of social information: Displaying relevant information about individuals through facial recognition; Real-time recording of conversations: Generating verbatim transcripts of conversations; Social record aggregation: Aggregating conversation records and incorporating them into a social interaction database. AI Object Recognition: Smart shopping recognition: Identifying products and displaying product descriptions and shopping links; Smart nutritionist: Identifying meals, providing nutritional analysis and records, and offering dietary recommendations based on individual health / sub-health conditions. AI Translation: AI text and image analysis (finger, hand-drawn underline), translation of unfamiliar words and phrases, AI sentence construction and extended usage, and learning notes.

[0089] This not only significantly reduces accidental triggering and power waste caused by unconscious actions, but also provides a seamless, intuitive and real-time information interaction experience, effectively solving the problems faced by previous technologies.

[0090] Although the present invention has been disclosed above by way of embodiments, it is not intended to limit the present invention. Anyone skilled in the art can make some modifications and refinements without departing from the spirit and scope of the present invention. Therefore, the scope of protection of the present invention shall be determined by the appended claims. [Simplified Explanation of the Diagram]

[0009] Figure 1A is a block diagram of an interactive control system according to an embodiment of the present invention. Figures 1B and 1C are schematic diagrams of the appearance of a head-mounted device according to an embodiment of the present invention. Figures 2A and 2B are schematic diagrams of usage scenarios of the interactive control system according to an embodiment of the present invention. Figure 3 is a flowchart of an interactive control method according to an embodiment of the present invention. Figure 4 is a flowchart of target behavior detection according to an embodiment of the present invention. Figures 5A to 5D are schematic diagrams of target behavior and information images according to an embodiment of the present invention. Figure 6 is a flowchart of depth detection according to an embodiment of the present invention. Figure 7 is a flowchart of motion detection according to an embodiment of the present invention. Figure 8 is a flowchart of image recognition according to an embodiment of the present invention. Figures 9A to 9D are schematic diagrams of display content in different application scenarios according to an embodiment of the present invention. Figure 10 is a schematic diagram of the display content of information images according to an embodiment of the present invention. Figures 11A and 11B are schematic diagrams of the information image startup process according to an embodiment of the present invention. Figure 12 is a schematic diagram of the information image browsing process according to an embodiment of the present invention. Figure 13 is a flowchart of gesture control according to an embodiment of the present invention. Figure 14 is a schematic diagram of reference gestures according to an embodiment of the present invention. Figures 15A and 15B are schematic diagrams of a user interface according to an embodiment of the present invention. Figures 16 and 17 are schematic diagrams of extended information according to an embodiment of the present invention. Figures 18A to 18D are schematic diagrams of another interactive flow according to an embodiment of the present invention. Figures 19A to 19C are schematic diagrams of another interactive flow according to an embodiment of the present invention. Figures 20A and 20B are schematic diagrams of a target object and corresponding displayed content according to an embodiment of the present invention. Figures 21A and 21B are schematic diagrams of another interactive flow according to an embodiment of the present invention.

Claims

1. An interactive control system, comprising: A head-mounted device includes: a motion sensor for detecting motion information; an image capturing device for acquiring an image to be tested; a display for displaying an information image; an auxiliary sensor for detecting auxiliary information; and a processor communicatively connected to the head-mounted device and the auxiliary sensor, and configured to: identify that the motion information corresponds to a target behavior, and acquire the auxiliary information based on the target behavior; determine that the auxiliary information corresponds to a target condition, and acquire the image to be tested based on the target condition; and identify a target object in the image to be tested, and generate display content of the information image based on the target object.

2. The interactive control system as claimed in claim 1, wherein the head-mounted device includes the auxiliary sensor, the auxiliary sensor being a distance sensor, the auxiliary information being distance information, the target condition being a target depth range, and the processor further being configured to: determine whether the distance information corresponds to the target depth range; when the distance information corresponds to the target depth range, acquire the image to be tested through the image capturing device; and when the distance information does not correspond to the target depth range, prohibit the acquisition of the image to be tested.

3. The interactive control system as claimed in claim 1, wherein the auxiliary sensor is a second motion sensor of a wearable device, the auxiliary information is second motion information, the target condition is a second target behavior, and the processor is further configured to: determine whether the second motion information corresponds to the second target behavior; when the second motion information corresponds to the second target behavior, acquire the image to be tested through the image capturing device; and when the second motion information does not correspond to the second target behavior, prohibit the acquisition of the image to be tested.

4. The interactive control system as claimed in claim 3, wherein the processor is further configured to: determine whether the angle value of the second motion information is within the rotation range of the second target behavior.

5. The interactive control system as claimed in claim 1, wherein the processor is further configured to: determine whether the angle value of the motion information is within the angle range of the target behavior; when the angle value of the motion information is within the angle range of the target behavior, acquire the auxiliary information through the auxiliary sensor; and when the angle value of the motion information is not within the angle range of the target behavior, prohibit the acquisition of the auxiliary information.

6. The interactive control system as claimed in claim 1, wherein the processor is further configured to: determine whether the target object exists in the image under test; when the target object exists in the image under test, display content corresponding to the target object; and when the target object does not exist in the image under test, prohibit the generation of display content of the information image.

7. The interactive control system as claimed in claim 1, wherein the target behavior is a head-down behavior, and the content displayed in the information image is one of calendar information, competition information, to-do information, and stock information.

8. The interactive control system of claim 1, wherein the processor is further configured to: identify a gesture information of a target object in the image to be tested, wherein the target object is a hand; and execute a target program of a reference gesture corresponding to the gesture information.

9. The interactive control system as described in claim 8, wherein the target program is one of trip browsing, missed call browsing, and unread message browsing.

10. The interactive control system of claim 1, wherein the information image includes derived information, and the processor is further configured to: identify a target person in the image under test and generate the derived information based on the target person and voice data, wherein the derived information is used for conversation suggestions; or identify a target product in the image under test and generate the derived information based on the target product, wherein the derived information is used for product description; or identify original text information in the image under test and generate the derived information based on the original text information, wherein the derived information is used for text translation.

11. An interactive control method, comprising: A head-mounted device, an auxiliary sensor, and a processor are provided. The head-mounted device includes a motion sensor, an image capturing device, and a display. The processor identifies motion information detected by the motion sensor as corresponding to a target behavior and obtains auxiliary information through the auxiliary sensor based on the target behavior. The processor determines that the auxiliary information corresponds to a target condition and obtains a test image through the image capturing device based on the target condition. The processor also identifies a target object in the test image and generates display content of an information image displayed on the display based on the target object.

12. The interactive control method as claimed in claim 11, wherein the head-mounted device includes the auxiliary sensor, the auxiliary sensor being a distance sensor, the auxiliary information being distance information, the target condition being a target depth range, and obtaining the image to be measured based on the target condition includes: Determine whether the distance information corresponds to the target depth range; When the distance information corresponds to the target depth range, the image to be tested is acquired through the image capturing device; and when the distance information does not correspond to the target depth range, the acquisition of the image to be tested is prohibited.

13. The interactive control method of claim 11, wherein the auxiliary sensor is a second motion sensor of a wearable device, the auxiliary information is second motion information, the target condition is a second target behavior, and the image to be tested is obtained according to the target condition: determining whether the second motion information corresponds to the second target behavior; when the second motion information corresponds to the second target behavior, obtaining the image to be tested through the image capturing device; and when the second motion information does not correspond to the second target behavior, prohibiting the acquisition of the image to be tested.

14. The interactive control method as described in claim 13, wherein determining whether the second action information corresponds to the second target behavior includes: Determine whether the angle value of the second action information is within the rotation range of the second target behavior.

15. The interactive control method as described in claim 11, wherein identifying the action information corresponding to the target behavior includes: Determine whether the angle value of the action information is within the angle range of the target behavior; When the angle value of the motion information is within the angle range of the target behavior, the auxiliary information is obtained through the auxiliary sensor; and when the angle value of the motion information is not within the angle range of the target behavior, the acquisition of the auxiliary information is prohibited.

16. The interactive control method as described in claim 11, wherein generating display content of information images displayed through the display based on the target object includes: Determine whether the target object exists in the image to be tested; When the target object exists in the image to be tested, the display content corresponding to the target object is displayed; And when the target object does not exist in the image under test, the display content of the information image is prohibited.

17. The interactive control method as described in claim 11, wherein the target behavior is a head-down behavior, and the content displayed in the information image is one of calendar information, competition information, to-do information, and stock information.

18. The interactive control method as described in claim 11 further includes: The processor identifies a gesture information of a target object in the image under test, wherein the target object is a hand; And a target program that executes the reference gesture corresponding to the gesture information through the processor.

19. The interactive control method as described in claim 18, wherein the target program is one of schedule browsing, missed call browsing, and unread message browsing.

20. The interactive control method as claimed in claim 11, wherein the information image includes derived information, and the interactive control method further includes: The processor identifies a target person in the image under test and generates derived information based on the target person and voice data, wherein the derived information is used for conversation suggestions; or the processor identifies a target product in the image under test and generates derived information based on the target product, wherein the derived information is used for product description; or the processor identifies original text information in the image under test and generates derived information based on the original text information, wherein the derived information is used for text translation.