Equipment control methods, apparatus, computer equipment, readable storage media and program products

Smart glasses automatically calculate the gaze direction and determine the display area of ​​the device by sensing eye movement and head posture data in real time, enabling seamless switching of control. This solves the problem of cumbersome operation between multiple devices and provides an efficient and smooth interactive experience.

CN121578896BActive Publication Date: 2026-04-17SHENZHEN JOOAN TECH CO LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHENZHEN JOOAN TECH CO LTD
Filing Date
2026-01-29
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

In existing technologies, users frequently perform explicit, manual, or voice focus switching operations between multiple devices, resulting in cumbersome and inefficient operations that fail to meet the needs for efficient, smooth, and seamless multi-device collaborative interaction.

Method used

By sensing the wearer's eye movement and head posture data in real time through smart glasses, the system automatically calculates the gaze direction and performs spatial intersection judgment with the display areas of multiple candidate devices in the environment to filter, evaluate and determine the target control device, thereby achieving seamless switching of control.

Benefits of technology

It significantly reduces the cognitive and operational burden of interaction, achieving a continuous, smooth, and highly collaborative seamless multi-device interaction experience, and solving the problems of low efficiency and fragmented experience caused by traditional manual switching methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121578896B_ABST
    Figure CN121578896B_ABST
Patent Text Reader

Abstract

This application relates to a device control method, apparatus, computer device, readable storage medium, and program product. The method includes: in response to a change in the motion state of a wearer corresponding to target smart glasses, acquiring the wearer's current motion state; determining the wearer's gaze direction based on the current motion state; filtering possible selectable devices from multiple candidate controllable devices based on the intersection of the gaze direction and the display area position; calculating the wearer's selection probability for each possible selectable device based on the relative positional relationship between the display area position of each possible selectable device and the gaze direction, as well as historical interaction information between the target smart glasses and the possible selectable devices; and determining the wearer's current target control device from the possible selectable devices. This method can improve the efficiency of device control and the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of equipment control technology, and in particular to an equipment control method, apparatus, computer equipment, readable storage medium, and program product. Background Technology

[0002] With the development of the Internet of Things and wearable computing technology, smart glasses technology has emerged that utilizes biosignals such as eye tracking for human-computer interaction. The key feature of this technology is its ability to infer a user's attention and intentions by collecting and analyzing physiological data such as eye movements and gaze direction, thereby providing wearers with a hands-free interactive experience.

[0003] In related technologies, when users need to operate multiple independent electronic devices (such as mobile phones, computers, and smart TVs), they generally need to manually switch control. This involves users using physical contact (such as touching a screen or clicking a mouse) or explicit switching commands (such as remote control input or the voice command "switch to TV") to shift the focus of currently acceptable input from one device to another. For example, in cross-device workflows, users need to repeatedly move their hands and shift their gaze between the input devices of different devices to switch control targets.

[0004] The problems with the related technologies include at least the following: users need to frequently perform explicit, manual, or voice-based focus switching operations between multiple devices. This not only disrupts the continuity of users' work or entertainment, causing cumbersome and inefficient operations, but also makes the interactive experience fragmented and disjointed. In particular, when the number of devices increases or their layout becomes more dispersed, the cognitive load and operational burden brought about by manual switching increase sharply, making it difficult to meet users' needs for efficient, smooth, and seamless multi-device collaborative interaction. Summary of the Invention

[0005] Therefore, it is necessary to provide a device control method, apparatus, computer equipment, readable storage medium, and program product that can improve device control efficiency and user experience in response to the above-mentioned technical problems.

[0006] In a first aspect, this application provides a device control method, the method comprising:

[0007] In response to changes in the motion state of the wearer corresponding to the target smart glasses, the current motion state of the wearer is obtained; the current motion state includes the wearer's eye movement data and head posture data.

[0008] The gaze direction of the wearer is determined based on the current motion state;

[0009] Identify multiple candidate controllable devices corresponding to the target smart glasses, and obtain the display area position of each candidate controllable device;

[0010] Based on the intersection of the gaze direction and the display area position, the possible selection device for the wearer is filtered from the plurality of candidate controllable devices;

[0011] For each of the possible selected devices, the probability of the wearer selecting each of the possible selected devices is calculated based on the relative positional relationship between the display area position of each of the possible selected devices and the gaze direction, as well as the historical interaction information between the target smart glasses and the possible selected devices.

[0012] Based on the selection probability of each of the possible selectable devices, determine the current target control device of the wearer from the possible selectable devices;

[0013] The target control device is configured as a master control device with the authority to receive and respond to input commands from the target smart glasses.

[0014] In some embodiments, the eye-tracking data is acquired through the target smart glasses; determining the gaze direction of the wearer based on the current motion state includes:

[0015] The eye-tracking data is converted into a gaze vector in the glasses coordinate system corresponding to the target smart glasses;

[0016] Based on the head pose data, the gaze vector is converted into a gaze vector in the world coordinate system to obtain the gaze direction;

[0017] The step of determining multiple candidate controllable devices corresponding to the target smart glasses and obtaining the display area position of each candidate controllable device includes:

[0018] For each of the candidate controllable devices, calculate the intersection point between the gaze direction and the plane where the display area of ​​the current candidate controllable device is located;

[0019] The candidate controllable devices whose intersection point and the display area position meet the preset distance condition are identified as the possible selected devices.

[0020] In some embodiments, the relative positional relationship includes the distance between the intersection point and the display area location, and the angle of incidence of the gaze direction relative to the plane where the display area location is located;

[0021] For each of the possible selected devices, the probability of the wearer selecting each of the possible selected devices is calculated based on the relative positional relationship between the display area position of each of the possible selected devices and the gaze direction, as well as the historical interaction information between the target smart glasses and the possible selected devices. This includes:

[0022] For each of the possible selected devices, the initial selection probability of the wearer for the current possible selected device is determined based on the distance and the incident angle; wherein, the initial selection probability is negatively correlated with the distance and the initial selection probability is negatively correlated with the incident angle;

[0023] The historical interaction information is analyzed to obtain the interaction frequency of the target smart glasses with the currently selectable device within a preset historical time period;

[0024] The initial selection probability is obtained by weighting the interaction frequency; wherein the selection probability is positively correlated with the interaction frequency and the initial selection probability.

[0025] In some embodiments, configuring the target control device as a master control device with the authority to receive and respond to input instructions from the target smart glasses includes:

[0026] Within multiple preset monitoring windows, the changes in the selection probability of each candidate controllable device are monitored;

[0027] If the change in the selection probability of the target control device within the plurality of listening windows meets a preset condition, a control grant instruction is sent to the target control device.

[0028] In some embodiments, before obtaining the current motion state of the wearer in response to a change in the motion state of the wearer corresponding to the target smart glasses, the method further includes:

[0029] In response to the device registration requests of the plurality of candidate controllable devices; the device registration request includes the device identifier, device location, display area location, and communication protocol of the candidate controllable devices;

[0030] The device list information corresponding to the multiple candidate controllable devices is constructed based on the device registration request, and the device list information is maintained based on the spatial position and orientation of each candidate controllable device in a preset world coordinate system.

[0031] In some embodiments, the method further includes:

[0032] Responding to an interaction event of the wearer; the interaction event includes at least one of a gaze pause, a blink, and a voice command detected by the smart glasses corresponding to the wearer;

[0033] According to the communication protocol supported by the main control device, the interactive event is converted into a target control command;

[0034] The target control command is routed to the main control device so that the main control device executes the target control command.

[0035] Secondly, this application also provides a device control apparatus, the apparatus comprising:

[0036] The response module is used to respond to changes in the motion state of the wearer corresponding to the target smart glasses and obtain the current motion state of the wearer; the current motion state includes the wearer's eye movement data and head posture data.

[0037] The first determining module is used to determine the gaze direction of the wearer based on the current motion state;

[0038] The acquisition module is used to determine multiple candidate controllable devices corresponding to the target smart glasses and acquire the display area position of each candidate controllable device;

[0039] The filtering module is used to filter out possible devices for the wearer from the plurality of candidate controllable devices based on the intersection of the gaze direction and the position of the display area;

[0040] The calculation module is used to calculate the probability of the wearer selecting each of the possible devices, based on the relative positional relationship between the display area position of each of the possible devices and the gaze direction, as well as the historical interaction information between the target smart glasses and the possible devices.

[0041] The second determining module is used to determine the current target control device of the wearer from the possible selectable devices based on the selection probability of each of the possible selectable devices;

[0042] The setting module is used to set the target control device as a master control device with the authority to receive and respond to input commands from the target smart glasses.

[0043] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps included in any of the foregoing device control method embodiments.

[0044] Fourthly, this application also provides a readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps included in any of the aforementioned device control method embodiments.

[0045] Fifthly, this application also provides a program product, including a computer program that, when executed by a processor, implements the steps included in any of the foregoing device control method embodiments.

[0046] The aforementioned device control method, apparatus, computer equipment, readable storage medium, and program product automatically calculate the wearer's gaze direction by sensing the wearer's eye movement and head posture data in real time. Based on this, they perform spatial intersection judgment with the display areas of multiple candidate devices in the environment, thereby filtering, evaluating, and determining the target control device, ultimately automatically and seamlessly switching control. This eliminates the need for users to explicitly switch between multiple devices; users simply need to naturally look at the device they wish to operate, and control is automatically and accurately assigned to that device. This significantly reduces the cognitive and operational burden of interaction, achieving a continuous, smooth, and highly collaborative seamless multi-device interactive experience, effectively solving the inefficiency and fragmented experience problems caused by traditional manual switching methods. Attached Figure Description

[0047] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0048] Figure 1 This is an application environment diagram of the device control method in one embodiment;

[0049] Figure 2 This is a flowchart illustrating a device control method in one embodiment;

[0050] Figure 3 This is a structural block diagram of the device control apparatus in one embodiment;

[0051] Figure 4 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0052] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0053] It should be noted that the terms "first," "second," etc., used in this application can be used to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish the first element from the second element. The terms used in this application include and have, and any variations thereof, are intended to cover non-exclusive inclusion. The term "multiple" in this application refers to two or more. The terms used in this application and / or refer to one of the embodiments, or any combination of multiple embodiments.

[0054] The device control method provided in this application embodiment can be applied to, for example, Figure 1 In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be integrated onto server 104 or located on the cloud or other network servers. Terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, drones, low-altitude aircraft, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, projection devices, etc. Portable wearable devices can include smartwatches, smart bracelets, head-mounted devices, etc. Head-mounted devices can be virtual reality (VR) devices, augmented reality (AR) devices, smart glasses, etc. Server 104 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.

[0055] In one exemplary embodiment, such as Figure 2 As shown, a device control method is provided, which is applied to... Figure 1 Taking smart glasses as an example, the explanation includes the following steps:

[0056] Step 202: In response to the change in motion state of the wearer corresponding to the target smart glasses, obtain the current motion state of the wearer; the current motion state includes the wearer's eye movement data and head posture data.

[0057] In cross-device interaction scenarios, users typically don't constantly switch focus; their gaze may linger on the same device for extended periods, performing in-depth operations. Therefore, continuously performing all calculations at a high frequency would result in unnecessary computational and energy consumption, failing to meet practical engineering application requirements and potentially shortening the smart glasses' battery life. Thus, in this embodiment of the invention, the device switching event is triggered by detecting changes in the motion state of the configured object.

[0058] Changes in motion state can be more than just minor tremors; they can represent significant behavioral patterns indicating a potential shift in gaze between devices. Considering user behavior patterns, when a user intends to switch operational targets, it is usually accompanied by head rotation and rapid eye movements (saccades). Therefore, by monitoring head angular velocity / acceleration exceeding thresholds or rapid, large-amplitude pupil position shifts, the moments most likely to involve focus switching can be efficiently identified, achieving an optimal balance between efficiency and real-time performance.

[0059] Specifically, a change in motion state refers to a change in the wearer's eye movement or head posture that exceeds a preset threshold, as detected by the target smart glasses. For example, when the inertial measurement unit (IMU) built into the smart glasses detects that the head angular velocity or acceleration exceeds the threshold, or when the eye-tracking sensor detects a significant displacement of the pupil position, it is determined as a change in state, triggering the acquisition and reporting of current motion state data. Eye-tracking data can be acquired by a high-precision eye-tracking sensor on the target smart glasses. This data can include the two-dimensional coordinates (x_left, y_left), (x_right, y_right) of the left and right pupils on the glasses' imaging plane, pupil size, blink status markers (such as blink_flag), and a timestamp t_gaze. Head posture data can also be acquired by the inertial measurement unit (IMU) on the target smart glasses. This data is typically represented in Euler angles (yaw, pitch, roll) or quaternions, and includes a timestamp t_pose. Optionally, eye-tracking data and head pose data can be timestamped and synchronized to ensure that eye-tracking and head pose data at the same moment are paired. Subsequently, preliminary filtering (such as Kalman filtering) can be performed to smooth noise, forming a unified dataset of current motion states arranged in time series.

[0060] Step 204: Determine the gaze direction of the wearer based on the current motion state.

[0061] In order to convert the wearer's original physiological signals (eye movements, head posture) into user interaction intentions, and considering the user's control intentions in the field of device control... Figure 1Generally, eye movements involve pointing the gaze towards the desired target object (device). However, simple pupil coordinates or head angles are isolated measurements centered on the user (with the glasses or head as a reference), and they cannot be directly correlated with multiple devices in the external environment. Therefore, in this embodiment of the invention, the two-dimensional coordinates of the pupil in the eye-tracking data are combined with the intrinsic parameters of the smart glasses (such as camera focal length and optical center) to convert them into a three-dimensional gaze vector V_glass in the glasses coordinate system with the optical center of the smart glasses as the origin. This vector indicates the direction of the pupil's gaze in the glasses coordinate system. Then, using the rotation matrix R_head corresponding to the head pose data (Euler angles or quaternions), the gaze vector V_glass in the glasses coordinate system is rotated to the world coordinate system. The calculation method is: V_world = R_head * V_glass. The resulting V_world is the gaze direction representing where the user is actually looking in the world coordinate system. This direction can be represented as a three-dimensional ray consisting of a starting point (such as the position of the smart glasses or the user's eye position) and a direction vector.

[0062] The gaze direction (V_world) in the world coordinate system is obtained through coordinate system transformation and fusion calculation. The gaze direction is used to represent the ray that originates from the user's eye and points to the three-dimensional spatial location where their attention is focused. It can be understood that this ray can connect the user's subjective intention with the objective physical world of the device.

[0063] Step 206: Determine multiple candidate controllable devices corresponding to the target smart glasses, and obtain the display area position of each candidate controllable device.

[0064] Candidate controllable devices can include electronic devices that are registered in the system, currently connected (e.g., online), and capable of accepting control permissions. For example, candidate controllable devices can include smartphones, tablets, laptops, smart TVs, and car infotainment screens corresponding to the wearer. A dynamic device list can be maintained, recording the unique identifier (device_id), device type, network address, and other information for each candidate controllable device. By defining the candidate controllable devices corresponding to the target smart glasses, it is clear which devices are likely to be selected by the user (i.e., the wearer of the target smart glasses), avoiding the processing of irrelevant or uncontrollable devices in subsequent steps, and ensuring the effectiveness and security of the interaction.

[0065] The display area location contains information for spatial intersection determination, such as the equation of the plane on which the screen lies and the bounding box coordinates of the effective display area in the world coordinate system. Considering that the device's screen is a bounded plane in three-dimensional space, based on the precise spatial description of this plane, accurate spatial geometric calculations (intersection and interior determination) can be performed with the aforementioned line-of-sight ray, thus determining in principle whether the user's gaze is actually falling on the device. This achieves precise focus recognition, which is fundamentally different from related technologies based on color, templates, and other two-dimensional image recognition methods, the latter being prone to failure when devices have similar appearances or when ambient light changes.

[0066] For each candidate controllable device, it is necessary to obtain the spatial position and geometric information of its display screen in a unified world coordinate system. Specifically, this can be achieved as follows: When the device is first connected, the spatial equation of its screen plane (e.g., a plane point normal equation Ax + By + Cz + D = 0) and the physical boundaries of the screen (usually defined by one or more rectangular areas) are determined through auxiliary means (such as the user scanning the four corners of the device screen with smart glasses, or using UWB / Bluetooth beacons). Optionally, for mobile devices (such as mobile phones), the position of its display area can be updated in real time by combining the device's own sensors (such as IMU) and relative positioning technologies (such as visual SLAM, UWB).

[0067] Step 208: Based on the intersection of the gaze direction and the display area position, select the possible device for the wearer from the plurality of candidate controllable devices.

[0068] In particular, if a ray originating from the viewpoint does not intersect with the representative plane (screen) of an object, or if the intersection point is outside the effective range (screen boundary) of the object, then physically, the user's gaze is less likely to actually see or focus on the object. Seeing and focusing are the physical basis for interaction. Therefore, these devices are excluded from the current focus candidate list, thereby narrowing the scope of subsequent fine-grained calculations and concentrating computing resources on a small number (usually 0-2) of devices that are truly likely to be viewed, thus avoiding the performance bottleneck caused by traversal intent matching in a massive device environment.

[0069] Specifically, for each candidate controllable device, the intersection point P_intersect of the gaze direction ray and the plane containing the device's screen is calculated, and it is determined whether this intersection point falls within the effective display area of ​​the device's screen. This effective display area can be the entire physical rectangular area of ​​the screen, or the core area minus the non-interactive edge areas. If P_intersect is within the effective display area: the user's gaze direction is considered to have hit the device's screen, the device is marked as a possible selection device, and added to a candidate list. If the ray is parallel to the plane (no intersection) or the intersection point P_intersect is outside the effective display area: the user is considered not currently looking at the device, and it is excluded from this focus determination. Finally, a set containing zero, one, or more possible selection devices is obtained. Considering that in most normal situations, the user's gaze can only fall on the screen of one device at a time, this set usually contains 0 or 1 devices. When the user's gaze falls on the physical boundary between two screens, it is possible to hit two devices simultaneously.

[0070] Step 210: For each of the possible selection devices, calculate the selection probability of the wearer for each of the possible selection devices based on the relative positional relationship between the display area position of each of the possible selection devices and the gaze direction, as well as the historical interaction information between the target smart glasses and the possible selection devices.

[0071] To further refine the differentiation of a user's primary intent among a limited number of candidates, based on the behavioral characteristics of human visual attention, when a user intends to interact with an object (such as operating a device), they subconsciously and steadily focus their gaze on the center or target area of ​​that object (a manifestation of Fitts's Law). A gaze deviating from the center often indicates scanning, browsing, or an undetermined target. Therefore, the distance from the intersection point to the screen center (d_center) is a direct spatial indicator quantifying the stability and purposefulness of the current gaze. Correspondingly, when a user is looking at a screen directly and attentively, their gaze direction will be as perpendicular to the screen plane as possible. A large angle of incidence (θ) usually indicates that the user is glancing from the side, a viewing state that often does not align with a deep interaction intent. Therefore, the angle θ is an angular indicator quantifying the user's level of engagement and direct focus on the device.

[0072] Correspondingly, in real-world scenarios, users exhibit habits and contextual associations in their use of different devices. For example, a user might frequently switch to their phone to reply to messages while looking at a computer. A high frequency of historical interactions means that, in similar current scenarios, the user is more likely to choose that device again. Using historical interaction information as a weighting factor allows the system to learn and adapt to users' individual behavioral patterns, making predictions that better align with user habits and expectations. This can resolve the issue of false triggers on users' habitual scanning paths that might arise from relying solely on instantaneous spatial relationships.

[0073] Specifically, in this embodiment of the invention, the selected candidates are given a refined score to quantify the intensity of the user's intention to select each device.

[0074] Relative positional relationship refers to a quantitative indicator of spatial geometric relationship, which may include at least one of the following: Distance deviation: Calculate the Euclidean distance d_center from the intersection point P_intersect to the center point of the screen. The closer the distance, the more focused the gaze is on the center of the screen, and the stronger the selection intention. Angle deviation: Calculate the angle θ between the gaze direction ray and the normal vector of the device's screen plane. The smaller the angle, the more perpendicular the gaze is to the screen, the more direct the gaze, and the stronger the selection intention. A comprehensive spatial relationship score S_spatial can be calculated based on d_center and θ, and its value is negatively correlated with these two indicators.

[0075] Furthermore, retrieving the interaction records between the target smart glasses (or the corresponding user account) and the currently selected device in the database in recent times (e.g., within the past hour or the current day) can include at least one of the following: Number of gazes: the number of times the device was identified as the focus; gaze duration: the cumulative or average duration of a single gaze; and control operation frequency: the number of operations initiated by the user through the smart glasses when the device is the master device. Based on the above statistical values, a historical preference score S_history is calculated, whose value is positively correlated with the frequency of interaction.

[0076] The spatial relationship score S_spatial and the historical preference score S_history are then weighted and fused, and normalized (e.g., using the Softmax function) to obtain the probability value P_select of each possible device being selected by the user's intention. For the case with only one possible device, P_select may be set to 1 or a high-confidence value.

[0077] Step 212: Determine the current target control device for the wearer from the possible selection devices based on the selection probability of each of the possible selection devices.

[0078] In this process, the device with the highest P_select value can be selected as the target control device. If all P_select values ​​are below a confidence threshold (e.g., 0.5), it can be determined that there is no clear focus, and the current master control device remains unchanged or enters a waiting state. Optionally, to address erroneous switching caused by rapid scanning of devices, this step can integrate a time window filtering mechanism. For example, it is required that the same device is calculated to have the highest P_select value in multiple consecutive processing cycles (e.g., the most recent 5 frames, 50ms per frame), and its average probability exceeds a threshold, in order to identify it as the target control device.

[0079] This invention, by combining the precision of spatial geometry with the regularity of long-term behavioral patterns, uses probabilistic fusion to accurately determine the device control intent of the smart glasses' configuration target, which is more resistant to noise and more consistent with cognitive logic than relying solely on single-frame data.

[0080] Step 214: Configure the target control device as a master control device with the authority to receive and respond to input commands from the target smart glasses.

[0081] The process involves comparing the newly identified target control device with the current main control device corresponding to the target smart glasses. If they differ, a control release command is sent to the current main control device, notifying it to enter a passive state. Simultaneously, a control grant command is sent to the target control device, notifying it to prepare to receive input commands from the smart glasses.

[0082] Optionally, the status record can be updated on the server side, changing the identifier (device_id) of the main control device to the identifier of the target control device. This ensures that all subsequent user interaction events received from the target smart glasses (such as gaze confirmation, blinking, and voice commands) are forwarded to the new main control device, which then performs the corresponding operations according to its own application interface protocol.

[0083] Through the above steps, the embodiments of the present invention can construct a system that can perceive and accurately understand the user's visual focus in real time, and automatically, smoothly, and intelligently transfer control between multiple physical devices accordingly. Ultimately, it creates a seamless interactive experience for the user that is controllable as long as the user's gaze is directed, solving the problem of fragmented operation in multi-device collaboration.

[0084] In some embodiments, the eye-tracking data is acquired through the target smart glasses; determining the gaze direction of the wearer based on the current motion state includes:

[0085] The eye-tracking data is converted into a gaze vector in the glasses coordinate system corresponding to the target smart glasses;

[0086] Based on the head pose data, the gaze vector is converted into a gaze vector in the world coordinate system to obtain the gaze direction;

[0087] The step of determining multiple candidate controllable devices corresponding to the target smart glasses and obtaining the display area position of each candidate controllable device includes:

[0088] For each of the candidate controllable devices, calculate the intersection point between the gaze direction and the plane where the display area of ​​the current candidate controllable device is located;

[0089] The candidate controllable devices whose intersection point and the display area position meet the preset distance condition are identified as the possible selected devices.

[0090] The eye-tracking data (such as the two-dimensional coordinates of the pupil (x_eye, y_eye)) can be measurements taken in the image coordinate system of the eye-tracking sensor itself, with the sensor chip plane as the reference. To establish a correlation with head posture, all measurements need to be unified into a fixed glasses coordinate system with the optical center of the smart glasses or a preset reference point as the origin. The eye-tracking sensor can be built into the temples or frame of the target smart glasses, with its imaging plane fixed relative to the wearer's eyes.

[0091] Specifically, by using the pre-calibrated intrinsic parameter matrix of the eye-tracking sensor (containing focal lengths f_x, f_y and principal points c_x, c_y) and the extrinsic parameter matrix (describing the rotation and translation relationship from the sensor coordinate system to the eyeglass coordinate system), the pixel coordinates (x_eye, y_eye) can be converted into a three-dimensional vector V_glass that originates from the approximate position of the eye (or the center of the sensor) and points in the direction of the pupil's gaze. This means that the movement of the pupil on the sensor image corresponds to the change in the direction of the gaze in the fixed coordinate system of the eyeglasses, and V_glass precisely quantifies this direction.

[0092] Considering that V_glass represents the gaze direction relative to the head, and that when a user turns their head, even if their eyes remain stationary, their actual world position changes, head pose data needs to be introduced for compensation. Specifically, head pose data (such as [yaw, pitch, roll] from the IMU) defines the rotation transformation of the glasses coordinate system relative to the world coordinate system (a preset, fixed global reference system). By calculating the rotation matrix R_head corresponding to this pose data, and left-multiplying the vector V_glass in the glasses coordinate system by this matrix: V_world = R_head * V_glass, the gaze vector V_world in the world coordinate system can be obtained.

[0093] Through a coordinate transformation process from the sensor to the glasses coordinate system and then to the world coordinate system, local and relative biometric signals are decoupled and fused into global and absolute spatial pointing information V_world. The starting point of V_world can be approximated as the position of the user's eyes in the world (obtained through initial glasses positioning or SLAM technology), and its direction accurately represents the ray of the user's line of sight in real three-dimensional space, laying a rigorous mathematical foundation for subsequent precise spatial intersection with the device screen.

[0094] Furthermore, after obtaining the gaze direction V_world in the world coordinate system (represented as a ray Ray: P_eye + t * V_dir, t >= 0) and the plane equation of the screen for each candidate controllable device (e.g., the general form Ax + By + Cz + D = 0, where (A, B, C) are the unit normal vectors of the screen plane, and D is a constant term), for the i-th candidate device, the ray equation is substituted into its plane equation to solve for the parameter t_intersect. If t_intersect exists and is positive (indicating that the ray originates from the eye and intersects the plane forward), then the intersection point coordinates P_intersect_i = P_eye + t_intersect_i * V_dir are calculated.

[0095] After calculating the intersection point P_intersect_i, it is necessary to determine whether it hits the effective area of ​​the screen. For example, the 2D bounding box of the device screen can be obtained ([x_min, x_max, y_min, y_max], where these coordinates are defined in the screen's own 2D planar coordinate system UV). Using a predefined projection transformation matrix M_proj_i from world coordinates to screen UV coordinates, the 3D intersection point P_intersect_i is projected onto the screen's UV coordinate system, resulting in (u_i, v_i). Then, it is determined whether (u_i, v_i) satisfies u_i ∈ [x_min, x_max] and v_i ∈ [y_min, y_max]. If these conditions are met, it indicates that the line-of-sight ray not only intersects the screen plane, but the intersection point also falls within the screen's display area, satisfying the physical prerequisite that the user can see the device.

[0096] Unlike the ambiguity and uncertainty commonly found in related technologies based on image recognition and template matching, the embodiments of this invention elevate the judgment criterion from appearance to geometric intersection, thereby ensuring high accuracy and robustness of focus recognition in principle, unaffected by factors such as ambient lighting, device appearance, and partial occlusion.

[0097] In some embodiments, the relative positional relationship includes the distance between the intersection point and the display area location, and the angle of incidence of the gaze direction relative to the plane where the display area location is located;

[0098] For each of the possible selected devices, the probability of the wearer selecting each of the possible selected devices is calculated based on the relative positional relationship between the display area position of each of the possible selected devices and the gaze direction, as well as the historical interaction information between the target smart glasses and the possible selected devices. This includes:

[0099] For each of the possible selected devices, the initial selection probability of the wearer for the current possible selected device is determined based on the distance and the incident angle; wherein, the initial selection probability is negatively correlated with the distance and the initial selection probability is negatively correlated with the incident angle;

[0100] The historical interaction information is analyzed to obtain the interaction frequency of the target smart glasses with the currently selectable device within a preset historical time period;

[0101] The initial selection probability is obtained by weighting the interaction frequency; wherein the selection probability is positively correlated with the interaction frequency and the initial selection probability.

[0102] In order to more accurately assess the control intention of the wearer, this embodiment of the invention introduces a multi-dimensional and quantifiable intention assessment index to perform a refined intention intensity score on the user's instantaneous gaze behavior and make personalized corrections in combination with long-term behavior patterns.

[0103] Specifically, distance (d_center) can refer to the two-dimensional Euclidean distance from the screen intersection point P_intersect_i to the center point of the device's screen display area. Considering that in real-world cross-device interaction scenarios, the user's gaze placement carries rich behavioral semantics—for example, when a user intends to deeply manipulate a device (such as clicking an icon or reading text), their gaze will generally instinctively and stably focus on the center area of ​​the screen; conversely, when the gaze wanders to the screen edges, it often indicates browsing, scanning, or an undetermined target, indicating a weaker interaction intent—d_center is an objective spatial indicator that quantifies the stability and purposefulness of the current gaze. By setting d_center negatively correlated with the initial selection probability (e.g., mapping it using the function f(d) = exp(-λ * d)), the more focused the gaze is on the center, the higher the score, thus effectively filtering out false gaze events caused by the gaze sweeping across the screen edges.

[0104] The incident angle is the angle between the gaze direction vector V_dir in the world coordinate system and the unit normal vector N_i of the device's screen plane. θ reflects the degree to which the user is directly facing the screen. Considering that when a user interacts with the device head-on and attentively, they will subconsciously adjust their posture to make their gaze as perpendicular to the screen as possible (θ ≈ 0°) to achieve optimal viewing and operational precision. A larger θ usually indicates side-looking, glancing, or non-primary focus, indicating a lower sustained interaction intent. Therefore, θ is an angular indicator that quantifies the user's level of attention and interaction comfort with the device. Setting θ negatively correlated with the initial selection probability (e.g., through g(θ) = cos(θ) mapping) effectively distinguishes between head-on operation and side-looking browsing, improving the dimensionality and accuracy of intent judgment.

[0105] Furthermore, considering that in complex multi-device workflows (e.g., a developer simultaneously coding on a computer, reviewing documents on a mobile phone, and taking notes on a tablet), user device switching is generally not completely random, but follows certain personal habits and contextual logic. Simply relying on instantaneous spatial relationships may lead to misjudgments based on the user's habitual gaze path (e.g., glancing from the computer screen to rest outside the window). To address this issue, this embodiment of the invention introduces time-series-based behavioral memory. Historical interaction information records the recent (e.g., the past 15 minutes, 1 hour, or a custom period) interactions between the target smart glasses (or associated user) and various potentially selectable devices. By analyzing historical interaction information, key statistical features can be extracted, such as gaze frequency (the number of times the device is identified as the focus and control is successfully switched), effective gaze duration (the cumulative duration of a single gaze exceeding a stable threshold), and control command frequency (the number of operation commands issued by the user through the smart glasses when the device is the master device). These features collectively constitute the interaction frequency F_i. A high F_i value indicates that the user frequently selects and operates the device in the current or similar context.

[0106] Furthermore, based on the instantaneous spatial features d_center and incident angle θ, the initial selection probability P_init_i is calculated using a predefined fusion function H(d, θ) (e.g., a weighted product H = w1 * f(d) * w2 * g(θ), where w1 + w2 = 1). P_init_i reflects the instantaneous intent confidence based on the geometric relationship of the current frame. To incorporate long-term behavioral patterns into the decision, a weight factor α_i positively correlated with the interaction frequency F_i can be used (e.g., α_i = sigmoid(β * F_i), where β is an adjustment factor). The initial probability is then weighted and corrected, for example, using the formula P_final_i = α_i * P_init_i + (1 - α_i) * P_baseline, where P_baseline is a base probability value. Alternatively, before Softmax normalization, P_init_i can be added to a reward term derived from F_i. After the above fusion and weighting, the final selection probability P_select_i for each possible device is obtained. This probability simultaneously embodies the instantaneous spatial precision (d_center, θ) and the long-term behavioral pattern (F_i).

[0107] Based on the complexity of intent in cross-device visual interaction, which depends on both the current physical orientation and historical operating habits, this embodiment of the invention upgrades focus recognition from geometric hit detection to context-aware intent inference, thereby enabling a natural, smooth, and personalized cross-device interactive experience.

[0108] In some embodiments, configuring the target control device as a master control device with the authority to receive and respond to input instructions from the target smart glasses includes:

[0109] Within multiple preset monitoring windows, the changes in the selection probability of each candidate controllable device are monitored;

[0110] If the change in the selection probability of the target control device within the plurality of listening windows meets a preset condition, a control grant instruction is sent to the target control device.

[0111] To ensure that the transfer of control is a true reflection of the user's stable and continuous intent, rather than a false trigger caused by transient actions or line-of-sight noise, thereby improving the reliability of device control and user experience, this embodiment of the invention further introduces a time-series-based stability judgment mechanism to perform anti-jitter and intent confirmation processing on focus switching events.

[0112] Specifically, considering that in real-world dynamic interactive environments, user gaze behavior is not an ideal, steady-state signal. For example, when a user's gaze quickly moves from one device (such as a computer) to another (such as a mobile phone), the gaze briefly passes over the screen of an intermediate device (such as a tablet). If left unprocessed, the system might mistakenly interpret this as the user intending to operate the tablet, causing control to flicker. During decision-making, a user's gaze may briefly linger between two devices or produce a brief, less than one-second, curious glance at a particular device. Immediately switching control in these situations would be abrupt and disruptive to the user. Furthermore, minute, unconscious head or eye movements can cause the gaze point to rapidly shift within a small area, triggering unnecessary probability recalculation. Therefore, using the highest probability calculation result in a single instance as the basis for switching is unreliable.

[0113] To address the aforementioned issues, this invention presents a time-series-based filtering and decision mechanism. Multiple monitoring windows can be designed as continuous or overlapping sliding time windows. For example, a short-term observation window (e.g., the past 3 frames, 50ms per frame) can be set to quickly capture the initial formation of intent. Simultaneously or in cascade, a mid-term confirmation window (e.g., the past 10 frames) can be set to verify the persistence of intent. Within each monitoring window, the selection probability P_select_i(t) of each candidate controllable device is continuously received and recorded.

[0114] Optionally, within each monitoring window, the trend and stability of the selection probability can also be analyzed. For example, it can be monitored whether the target control device maintains the highest probability continuously (e.g., for more than 80% of the frames) within the window. Whether the probability value shows a monotonically increasing trend or remains stable at a high level, and optionally, whether the probability value is significantly and stably higher than the second-highest (e.g., the difference is consistently greater than the threshold Δ_thresh).

[0115] The preset conditions are a comprehensive judgment standard set for the changes detected above. For example, it can be defined as follows: within the mid-term confirmation window, the selection probability of device D must remain the highest for more than 8 consecutive frames, and the average probability value must exceed 0.7, while the average probability difference between it and the second highest device must be greater than 0.2.

[0116] When the target control device exhibits sufficient stability, consistency, and superiority over time, the user's switching intention is determined to be genuine, clear, and ready. This avoids situations that may lead to misidentification, such as: glancing at passing devices, whose high-probability states are fleeting and cannot meet the condition of consistency; hesitant or curious stares, whose probability may briefly increase but cannot maintain sufficient duration or stable advantage; and minor fluctuations in probability caused by slight jitter.

[0117] If the changes of the target control device in multiple monitoring windows simultaneously meet the preset conditions, the system confirms that the focus switching event has been established. At this time, a control granting instruction is sent to the target control device, reducing the false trigger rate of control switching to a low level and avoiding device switching caused by unconscious user actions. The switching action is aligned with the user's stable gaze behavior in time, and the interaction process is natural, smooth, and predictable.

[0118] Optionally, preset conditions can be dynamically adjusted and personalized according to device type (e.g., longer confirmation time can be required when switching to a TV), application scenario (faster response is required in game mode, and higher stability is required in reading mode) or user preferences, so that the system can flexibly adapt to various complex scenarios.

[0119] In some embodiments, before obtaining the current motion state of the wearer in response to a change in the motion state of the wearer corresponding to the target smart glasses, the method further includes:

[0120] In response to the device registration requests of the plurality of candidate controllable devices; the device registration request includes the device identifier, device location, display area location, and communication protocol of the candidate controllable devices;

[0121] The device list information corresponding to the multiple candidate controllable devices is constructed based on the device registration request, and the device list information is maintained based on the spatial position and orientation of each candidate controllable device in a preset world coordinate system.

[0122] In open IoT or smart spaces (such as smart homes, multi-screen offices, and smart cockpits), the interactive devices are not pre-fixed or known. New devices may be added (e.g., a user brings a new tablet into the room), and existing devices may be removed or turned off. Therefore, a standardized candidate device access mechanism can be established, allowing devices to announce their presence and interactive capabilities.

[0123] A device registration request can be a self-describing information packet sent by a candidate controllable device to a central server (or a coordinator in a peer-to-peer network). Fields included in this packet may include: Device ID (device_id): such as a unique UUID, MAC address, or user-defined alias. Used to uniquely identify the device throughout the system. Device Position (device_position): The device's three-dimensional coordinates in the world coordinate system (e.g., (x, y, z)). For fixed devices (e.g., TVs, desktops), this position can be determined during initial installation via user guidance (e.g., scanning screen corners with smart glasses) or pre-configuration. For mobile devices (e.g., phones, tablets), this position can be dynamically estimated and updated in real-time by fusing the device's own sensors (IMU, UWB) with environmental awareness technologies (e.g., visual SLAM, Bluetooth AoA). And display area position (display_info): This is the core geometric description of the interactive plane. It not only includes the device position but, more importantly, defines the spatial equations of the screen plane (normal vectors and offsets) and the effective display boundaries of the screen (e.g., the coordinates of the four corners of a rectangular area in the world coordinate system). This provides the geometric basis for the precise ray-plane intersection in step S220.

[0124] A communication protocol is used to define the communication specifications for subsequent control commands and event routing, such as MQTT topics, REST API endpoints, WebSocket addresses, or custom binary protocols. This ensures that the system can accurately transmit natural interaction commands generated by the smart glasses (such as "blink confirmation") to the target device for execution.

[0125] The central server's device discovery and space mapping module continuously listens for registration requests within the network. Upon receiving a request, it first performs identity and security authentication (e.g., based on certificates), then parses the aforementioned information, creates or updates the device's file entry in memory and / or the database, ultimately forming and maintaining a dynamic device list.

[0126] Since the core of visual focus calculation is spatial relationships, all calculations (line of sight, intersection point, distance) must be performed within the same global reference system (i.e., the preset world coordinate system); otherwise, the results will be meaningless. The world coordinate system is typically established during environment initialization (e.g., with a corner of the room as the origin). All devices' `device_position` and `display_info` values ​​need to be transformed and expressed in this unified coordinate system. This requires the system to have coordinate transformation capabilities, able to map various local coordinates reported by devices (such as coordinates based on the device's own IMU) to world coordinates through known transformation relationships.

[0127] The dynamic maintenance process of the list information may include the following: For mobile devices, the system needs to periodically or triggerically receive their location updates and refresh the `device_position` and `display_info` in the list. For example, when a user picks up their phone and moves around, the phone's UWB module will continuously report its new location. This includes device online / offline status, battery level, and currently running application context (optional). Devices that have not been updated for a long time or have been actively logged out are removed from the active list to avoid unnecessary calculations for invalid devices.

[0128] For example, in a smart meeting room, projectors, multiple participants' laptops, and smart whiteboards are automatically registered at the start of the meeting. The system-built list allows the speaker to route PowerPoint slide navigation commands to the corresponding devices simply by looking at any screen with their glasses. In a smart home living room, after TVs, game consoles, speakers, and lighting panels are registered, users can switch control to the entertainment system by watching the set-top box, or control the home environment by watching the smart panel.

[0129] This invention reduces the complexity of device integration, allowing devices conforming to the registration protocol to seamlessly connect, thus improving the system's openness and ecosystem compatibility. Unlike simple device discovery, this invention constructs a spatially aware, dynamic, and standardized device interaction context. It abstracts and organizes discrete, heterogeneous physical devices into a unified digital set that can be understood and manipulated by upper-layer visual interaction logic.

[0130] In some embodiments, the method further includes:

[0131] Responding to an interaction event of the wearer; the interaction event includes at least one of a gaze pause, a blink, and a voice command detected by the smart glasses corresponding to the wearer;

[0132] According to the communication protocol supported by the main control device, the interactive event is converted into a target control command;

[0133] The target control command is routed to the main control device so that the main control device executes the target control command.

[0134] In order to provide a unified and convenient way to issue specific commands to the controlled device, and considering that requiring users to use a keyboard, mouse, or touchscreen in scenarios where smart glasses are worn would interrupt the experience, this invention defines the following natural, contactless, and low-cognitive-load interaction modality based on the sensing capabilities of smart glasses:

[0135] Eye-dwelling event: When a user gazes at a specific UI element (such as a button or icon) on the main control device screen for more than a preset time (such as 500ms), a "select" or "confirm" event is triggered. This simulates the experience of "mouse hover and click," which is suitable for precise selection.

[0136] Specific blinks: such as two rapid, consecutive blinks or prolonged eye closure, serve as general commands like "confirm," "cancel," or "return." Blinking is an active and clear way of expressing intent, highly resistant to interference, and suitable as a primary confirmation instruction.

[0137] Voice commands: The glasses receive the user's voice through the built-in microphone and recognize it as high-level semantic commands such as "play," "pause," "increase volume," and "open app X." Voice is suitable for conveying complex intentions and parameters.

[0138] The aforementioned interactive events are generated by the interaction detection module of the smart glasses client through real-time analysis of raw sensor data (eye tracking, camera images, microphone audio), and reported to the central server via a low-latency wireless link (such as Wi-Fi 6, UWB), thereby ensuring the instant capture of interactive intentions.

[0139] Considering that different candidate controllable devices (such as Windows computers, Android TVs, and custom IoT devices) have drastically different control interfaces and protocols, a "blink confirmation" event might simulate a left mouse click for a computer, send an infrared-coded "OK" command for a TV, and trigger an MQTT message for a smart light. Therefore, upon receiving a standardized "interaction event," the module queries the registration information of the current master control device, particularly its supported "communication protocols." For each protocol, a conversion mapping table or adapter can be pre-configured or dynamically loaded. For example, for a computer supporting the HID (Human Interface Device) protocol, the module combines the "eye movement at coordinates (x,y)" and "two blinks" events, converts and encapsulates them into a standard "mouse moves to (x,y) and left-clicks" HID report, and sends it to the computer via a virtual HID driver or network HID.

[0140] For smart TVs that support REST APIs, the module may convert "voice command: increase volume" into an HTTP POST request sent to a TV-specific REST endpoint.

[0141] For dedicated devices with custom protocols, the adapter will convert events into binary or JSON format instructions defined by the device manufacturer.

[0142] This decouples the user interaction layer from the device control layer. Users only need to learn a unified set of natural interaction methods (looking, blinking, speaking), which can automatically handle complex adaptations to various device protocols, greatly reducing the user's learning cost and interaction complexity.

[0143] Based on the currently maintained "Master Control Device" identifier, the network address (IP, port, topic, etc.) is retrieved from the device list, and the converted "Target Control Command" is sent to that device via the corresponding network protocol. Upon receiving the command, the Master Control Device executes the corresponding operation via its local client or operating system interface, such as simulating a click, adjusting volume, or switching applications. Optionally, the device can feed back the execution result or status change to the central server, which can then provide confirmation feedback to the user (such as a "beep" sound or visual cues) through the smart glasses' speaker or display screen (e.g., a miniature projector), forming an interactive closed loop.

[0144] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps. It is understood that the steps in different embodiments can be freely combined as needed, and all non-contradictory solutions formed by such combinations are within the scope of protection of this application.

[0145] Based on the same inventive concept, this application also provides a device control apparatus for implementing the device control method described above. The solution provided by this apparatus is similar to the implementation scheme described in the above method; therefore, the specific limitations in one or more device control apparatus embodiments provided below can be found in the limitations of the device control method described above, and will not be repeated here.

[0146] In one exemplary embodiment, such as Figure 3 As shown, a device control device 300 is provided, comprising:

[0147] Response module 302 is used to respond to changes in the motion state of the wearer corresponding to the target smart glasses and obtain the current motion state of the wearer; the current motion state includes the wearer's eye movement data and head posture data;

[0148] The first determining module 304 is used to determine the gaze direction of the wearer based on the current motion state;

[0149] The acquisition module 306 is used to determine multiple candidate controllable devices corresponding to the target smart glasses and acquire the display area position of each candidate controllable device;

[0150] The filtering module 308 is used to filter out possible devices for the wearer from the plurality of candidate controllable devices based on the intersection of the gaze direction and the position of the display area;

[0151] The calculation module 310 is used to calculate the probability of the wearer selecting each of the possible devices based on the relative positional relationship between the display area position of each of the possible devices and the gaze direction, as well as the historical interaction information between the target smart glasses and the possible devices.

[0152] The second determining module 312 is used to determine the current target control device of the wearer from the possible selectable devices based on the selection probability of each of the possible selectable devices;

[0153] Setting module 314 is used to set the target control device as a master control device with the authority to receive and respond to input commands from the target smart glasses.

[0154] Each module in the aforementioned equipment control device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0155] In one exemplary embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 4As shown, the computer device includes a processor, memory, input / output interfaces, a communication interface, a display unit, and an input device. The processor, memory, and input / output interfaces are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interfaces. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The input / output interfaces are used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, Near Field Communication (NFC), or other technologies. When the computer program is executed by the processor, it implements a device control method. The display unit is used to form a visually visible image and can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the computer device, or external keyboards, touchpads, or mice, etc.

[0156] Those skilled in the art will understand that Figure 4 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0157] In one exemplary embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps included in any of the foregoing device control method embodiments.

[0158] In one embodiment, a readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps included in any of the foregoing device control method embodiments.

[0159] In one embodiment, a program product is provided, including a computer program that, when executed by a processor, implements the steps included in any of the foregoing device control method embodiments.

[0160] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0161] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.

[0162] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0163] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A device control method, characterized in that, The method includes: In response to changes in the motion state of the wearer corresponding to the target smart glasses, the current motion state of the wearer is obtained; the changes in motion state include the wearer's head angular velocity or acceleration exceeding a first threshold or the pupil position displacement exceeding a second threshold and the rate of change of the pupil position displacement being greater than a rate of change threshold; the current motion state includes the wearer's eye movement data and head posture data; The gaze direction of the wearer is determined based on the current motion state; Identify multiple candidate controllable devices corresponding to the target smart glasses, and obtain the display area position of each candidate controllable device; Based on the intersection of the gaze direction and the display area position, the possible selection device for the wearer is filtered from the plurality of candidate controllable devices; For each of the possible selection devices, based on the relative positional relationship between the display area position and the gaze direction of each of the possible selection devices and the historical interaction information between the target smart glasses and the possible selection devices, the selection probability of the wearer for each of the possible selection devices is calculated; based on the selection probability of each of the possible selection devices, the current target control device of the wearer is determined from the possible selection devices; the target control device is set as the master control device with the authority to receive and respond to input commands from the target smart glasses; wherein, within a preset plurality of listening windows, the changes in the selection probability of each of the candidate controllable devices are monitored; If the change in the selection probability of the target control device within the plurality of listening windows meets a preset condition, a control grant instruction is sent to the target control device; the preset condition includes the selection probability remaining the highest among the plurality of possible selection devices within the plurality of listening windows, the selection probability showing a monotonically increasing trend within the plurality of listening windows, or the difference in selection probability between the selected device and the possible selection device with the second highest selection probability continuously being greater than a preset threshold within the plurality of listening windows.

2. The method according to claim 1, characterized in that, The eye-tracking data is acquired through the target smart glasses; determining the gaze direction of the wearer based on the current motion state includes: The eye-tracking data is converted into a gaze vector in the glasses coordinate system corresponding to the target smart glasses; Based on the head pose data, the gaze vector is converted into a gaze vector in the world coordinate system to obtain the gaze direction; The step of determining multiple candidate controllable devices corresponding to the target smart glasses and obtaining the display area position of each candidate controllable device includes: For each of the candidate controllable devices, calculate the intersection point between the gaze direction and the plane where the display area of ​​the current candidate controllable device is located; The candidate controllable devices whose intersection point and the display area position meet the preset distance condition are identified as the possible selected devices.

3. The method according to claim 2, characterized in that, The relative positional relationship includes the distance between the intersection point and the display area location, and the angle of incidence of the gaze direction relative to the plane where the display area location is located; For each of the possible selected devices, the probability of the wearer selecting each of the possible selected devices is calculated based on the relative positional relationship between the display area position of each of the possible selected devices and the gaze direction, as well as the historical interaction information between the target smart glasses and the possible selected devices. This includes: For each of the possible selected devices, the initial selection probability of the wearer for the current possible selected device is determined based on the distance and the incident angle; wherein, the initial selection probability is negatively correlated with the distance and the initial selection probability is negatively correlated with the incident angle; The historical interaction information is analyzed to obtain the interaction frequency of the target smart glasses with the currently selectable device within a preset historical time period; the initial selection probability is weighted according to the interaction frequency to obtain the selection probability; wherein, the selection probability is positively correlated with the interaction frequency and the initial selection probability.

4. The method according to claim 3, characterized in that, The distance includes the two-dimensional Euclidean distance from the intersection of the screens to the center point of the device screen display area; the incident angle includes the angle between the gaze direction vector in the world coordinate system and the unit normal vector of the device screen plane; the interaction frequency includes the number of times the video is viewed, the effective gaze duration, and the frequency of control commands.

5. The method according to claim 1, characterized in that, Before obtaining the current motion state of the wearer in response to a change in the motion state of the wearer corresponding to the target smart glasses, the method further includes: In response to the device registration requests of the plurality of candidate controllable devices; the device registration request includes the device identifier, device location, display area location, and communication protocol of the candidate controllable devices; The device list information corresponding to the multiple candidate controllable devices is constructed according to the device registration request, and the device list information is maintained according to the spatial position and orientation of each candidate controllable device in a preset world coordinate system.

6. The method according to claim 5, characterized in that, The method further includes: Responding to an interaction event of the wearer; the interaction event includes at least one of a gaze pause, a blink, and a voice command detected by the smart glasses corresponding to the wearer; According to the communication protocol supported by the main control device, the interactive event is converted into a target control command; The target control command is routed to the main control device so that the main control device executes the target control command.

7. A device control apparatus, characterized in that, The device includes: The response module is used to respond to changes in the motion state of the wearer corresponding to the target smart glasses and obtain the current motion state of the wearer; the changes in motion state include the wearer's head angular velocity or acceleration exceeding a first threshold or the pupil position displacement exceeding a second threshold and the rate of change of the pupil position displacement being greater than a rate of change threshold; the current motion state includes the wearer's eye movement data and head posture data; The first determining module is used to determine the gaze direction of the wearer based on the current motion state; The acquisition module is used to determine multiple candidate controllable devices corresponding to the target smart glasses and acquire the display area position of each candidate controllable device; The filtering module is used to filter out possible devices for the wearer from the plurality of candidate controllable devices based on the intersection of the gaze direction and the position of the display area; The calculation module is used to calculate the probability of the wearer selecting each of the possible devices based on the relative positional relationship between the display area position of each of the possible devices and the gaze direction, as well as the historical interaction information between the target smart glasses and the possible devices. The second determining module is used to determine the current target control device of the wearer from the possible selectable devices based on the selection probability of each of the possible selectable devices; The setting module is used to set the target control device as a main control device with the authority to receive and respond to input commands from the target smart glasses; wherein, within a plurality of preset listening windows, the changes in the selection probability of each of the candidate controllable devices are monitored; If the change in the selection probability of the target control device within the plurality of listening windows meets a preset condition, a control grant instruction is sent to the target control device; the preset condition includes the selection probability remaining the highest among the plurality of possible selection devices within the plurality of listening windows, the selection probability showing a monotonically increasing trend within the plurality of listening windows, or the difference in selection probability between the selected device and the possible selection device with the second highest selection probability continuously being greater than a preset threshold within the plurality of listening windows.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.

9. A readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.

10. A program product comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Intelligent equipment control method and system based on intelligent glasses

    CN115291734A

  • Control method and device

    CN120686970A

  • Information processing apparatus, information processing method, and program

    US20190243460A1