Method for controlling computing device using lip reading and eye tracking

WO2026205614A1PCT designated stage Publication Date: 2026-10-01TERAMIME INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/003915
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2026-10-01

Smart Images

  • Figure KR2025003915_01102026_PF_FP_ABST
    Figure KR2025003915_01102026_PF_FP_ABST
Patent Text Reader

Abstract

A method for controlling a computing device according to an embodiment of the present disclosure relates to a method for controlling a computing device including a camera configured to acquire an image of a user's face and an artificial intelligence module configured to read the user's lips to determine a control command of the user. The method for controlling a computing device according to the present disclosure may comprise: (A) a step in which a camera acquires standard word image data, which is an image of a user pronouncing a standard word, and transmits the standard word image data to an artificial intelligence module; (B) a step in which the artificial intelligence module is trained on the standard word image data to derive first feature points corresponding to changes in the user's lip shape in the standard word image data; (C) a step in which the camera acquires image data of the user's face and transmits the image data to the artificial intelligence module; (D) a step in which the artificial intelligence module determines, on the basis of the image data, whether the user is speaking; (E) a step in which the artificial intelligence module, when it is determined that the user is speaking, derives second feature points corresponding to changes in the user's lip shape in the image data; (F) a step in which the artificial intelligence module derives the user's control command by comparing the second feature points with the first feature points; and (G) a step in which the artificial intelligence module controls the computing device on the basis of the derived control command of the user.
Need to check novelty before this filing date? Find Prior Art

Description

Computing device control method using lip reading and eye tracking

[0001] The disclosed content relates to a method for controlling a computing device using lip reading and eye tracking.

[0002] Unless otherwise indicated in this specification, the contents described in this section are not prior art for the claims of this application, and are not to be recognized as prior art simply because they are included in this section.

[0003] Technologies for controlling computing devices typically utilize input devices such as mice or keyboards. However, in recent years, technologies for controlling computing devices through voice recognition have been gaining prominence. Voice recognition is a technology that analyzes a user's voice to execute specific commands, and it is rapidly advancing alongside the development of artificial intelligence (AI) and natural language processing (NLP). While early voice recognition technologies operated based on simple acoustic pattern matching, recent advancements in machine learning and deep neural networks (DNNs) enable the understanding of context and the execution of more sophisticated commands.

[0004] Representative voice recognition systems include Apple's Siri, Google's Google Assistant, Amazon's Alexa, and Microsoft's Cortana. These systems analyze the user's voice in real time and integrate with cloud-based data to provide appropriate responses. This technology is utilized in various fields, such as smartphones, smart speakers, car navigation systems, and IoT devices, and is particularly useful in environments where hands are not free.

[0005] However, there are several drawbacks to computing control using voice recognition. Misrecognition can occur in cases of severe background noise or inaccurate pronunciation, and security issues may arise if others are nearby, as control commands for computing devices may be unintentionally heard by them.

[0006] The disclosed content aims to provide a method for controlling computing using lip reading in a computing device including a camera in order to solve the aforementioned problem.

[0007] A method for controlling a computing device according to an embodiment of the present disclosure relates to a method for controlling a computing device comprising a camera capable of acquiring an image of a user's face and an artificial intelligence module capable of reading the user's lips to determine a user's control command. The method for controlling a computing device according to the present disclosure comprises: (A) a step in which the camera acquires standard word image data, which is an image of a user pronouncing a standard word, and transmits it to the artificial intelligence module; (B) a step in which the artificial intelligence module learns the standard word image data and derives a first feature point regarding a change in the shape of the user's lips within the standard word image data; (C) a step in which the camera acquires image data of the user's face and transmits it to the artificial intelligence module; (D) a step in which the artificial intelligence module determines whether the user is speaking based on the image data; (E) a step in which the artificial intelligence module derives a second feature point regarding a change in the shape of the user's lips within the image data when it is determined that the user is speaking; (F) a step in which the artificial intelligence module derives a user's control command by comparing the second feature point and the first feature point. and (G) the artificial intelligence module may include the step of controlling the computing device based on the derived user's control command.

[0008] In one embodiment, the standard word may be characterized by being configured to include a bilabial plosive, a bilabial nasal, an alveolar nasal, an open vowel, an unrounded vowel, and a rounded vowel.

[0009] In one embodiment, in the step of determining whether the user (D) is speaking, the artificial intelligence module may be characterized by dividing the video data into frames of a certain length and, within the frames, determining that the user is speaking if there is movement of the user's lips, and determining that the user is not speaking if there is no movement of the user's lips.

[0010] In one embodiment, the artificial intelligence module determines whether a user is speaking based on a plurality of frames, compares a first number of frames determined to be spoken by the user with a second number of frames determined not to be spoken by the user, determines that the user is speaking if the first number is greater than the second number, determines that the user is speaking if the first number is less than the second number, determines that the user is not speaking if the first number is less than the second number, and determines that the previous state is maintained if the first number and the second number are the same.

[0011] In one embodiment, in the step (B) of deriving a first feature point regarding a change in the user's lip shape and the step (E) of deriving a second feature point regarding a change in the user's lip shape, the artificial intelligence module may be characterized by extracting the movement of points on the inner contour of the user's lips to derive the first feature point and the second feature point.

[0012] In one embodiment, the computing device control method of the present disclosure further comprises the step (H) of the artificial intelligence module correcting the image data; and in the step (H), the artificial intelligence module derives a first triangle including the center point of the user's two eyes and the tip point of the nose for the user's face in the standard word image data and a perpendicular line of the centroid of the first triangle, and derives a second triangle including the center point of the user's two eyes and the tip point of the nose for the user's face in the image data and a perpendicular line of the centroid of the second triangle, and compares the shape and size of the first triangle and the shape and size of the second triangle, and compares the slope of the perpendicular line of the centroid of the first triangle with respect to a straight line perpendicular to the ground, thereby correcting the image data by considering the distance at which the user's face is located and the angle of the user's face.

[0013] As an embodiment, the computing device control method of the present disclosure may further include: (I) the artificial intelligence module tracking the user's gaze within the image data; and (J) the artificial intelligence module deriving a region of interest that the user's gaze is looking at among the screens of the computing device based on the tracked user's gaze.

[0014] In one embodiment, in the step of deriving the area of ​​interest that the user (J) gazes at, the artificial intelligence module may be characterized by applying an autoregressive Transformer-based gaze prediction model that receives data such as past user gaze coordinates, changes in the user's pupil size, and the user's head movements to predict the gaze at the next point in time.

[0015] In one embodiment, in the step of deriving the user's control command (F), the artificial intelligence module may derive the user's control command by considering the correlation with the derived region of interest.

[0016] The computing device control method according to the embodiment of the present disclosure can increase the recognition rate by controlling the computing device using lip reading rather than voice, and can also recognize the user's control commands even when the user moves only their lips without speaking, thereby preventing security problems that may arise from others hearing the user's voice.

[0017] The effects of the present invention are not limited to the effects described above, and should be understood to include all effects that can be inferred from the configuration of the invention described in the detailed description of the invention or the claims.

[0018] FIG. 1 is a configuration diagram of a computing device according to an embodiment of the present disclosure.

[0019] FIG. 2 is a flowchart of a computing device control method according to an embodiment of the present disclosure.

[0020] FIG. 3 is a description of an algorithm for determining whether a user has uttered a sound in a computing device control method according to an embodiment of the present disclosure.

[0021] FIG. 4 is an example of a point on the contour of the lip for analyzing changes in the shape of the user's lips in order to extract feature points regarding changes in the shape of the user's lips in a computing device control method according to an embodiment of the present disclosure.

[0022] FIG. 5 is a flowchart of a computing device control method according to one embodiment of the present disclosure.

[0023] FIG. 6 is an explanation of the correction of image data of a computing device control method according to an embodiment of the present disclosure.

[0024] Hereinafter, embodiments disclosed in this specification will be described in detail with reference to the attached drawings. Identical or similar components regardless of drawing symbols are assigned the same reference number, and redundant descriptions thereof will be omitted. The suffixes "module" and "part" used for components in the following description are assigned or used interchangeably solely for the ease of drafting the specification and do not inherently possess distinct meanings or roles. Furthermore, in describing the embodiments disclosed in this specification, if it is determined that a detailed description of related prior art could obscure the essence of the embodiments disclosed in this specification, such detailed description will be omitted. Additionally, the attached drawings are intended only to facilitate understanding of the embodiments disclosed in this specification; the technical concept disclosed in this specification is not limited by the attached drawings, and it should be understood that they include all modifications, equivalents, and substitutions that fall within the spirit and technical scope of the invention.

[0025] Terms including ordinal numbers, such as first, second, etc., may be used to describe various components, but said components are not limited by said terms. These terms are used solely for the purpose of distinguishing one component from another.

[0026] In this application, terms such as “comprising” or “having” are intended to specify the existence of the features, numbers, steps, actions, components, parts, or combinations thereof described in the specification, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.

[0027]

[0028] Hereinafter, a method for controlling a computing device (100) according to an embodiment of the present disclosure will be described in detail with reference to the drawings.

[0029] FIG. 1 is a configuration diagram of a computing device (100) according to an embodiment of the present disclosure, FIG. 2 is a flowchart of a control method for a computing device (100) according to an embodiment of the present disclosure, FIG. 3 is a description of an algorithm for determining whether a user speaks according to a control method for a computing device (100) according to an embodiment of the present disclosure, FIG. 4 is an example of a point on the contour of a lip for analyzing a change in the shape of a user's lip in order to extract a feature point regarding a change in the shape of a user's lip according to an embodiment of the present disclosure.

[0030] Referring to FIG. 1, the computing device (100) includes a camera (110), an artificial intelligence module (120), and a screen (130). The computing device (100) includes at least one CPU (Central Process Unit), at least one memory device, and at least one GPU (Graphic Process Unit), and may be any one of a desktop PC, a laptop PC, a tablet PC, a smartphone, a car navigation system, or a VR device, but is not limited thereto. Each step of the method for controlling the computing device (100) disclosed herein may be implemented in the computing device (100) by a hardware or software method.

[0031] The camera (110) may be installed on the upper part of the screen (130) of the computing device (100), but is not limited thereto. The camera (110) may be a type of web camera (110) capable of acquiring an image of a user's face, but is not limited thereto. The camera (110) may acquire an image of a user's face, convert it into image data, and then transmit it to an artificial intelligence module to be described later via a wired or wireless communication method.

[0032] The artificial intelligence module (120) may be an on-device artificial intelligence module (120) installed in the memory device of the computing device (100), or a cloud-based artificial intelligence module (120) that is cloud-based and can communicate with the computing device (100) through a network. The artificial intelligence module (120) may be a neural network based on deep learning.

[0033] Deep learning is a technology that learns data and recognizes patterns using deep neural networks (DNN). Unlike conventional machine learning, deep learning can analyze large amounts of data and extract features on its own without human intervention. The artificial intelligence module (120) of the present disclosure may be any one of a multilayer perceptron (MLP), a convolutional neural network (CNN), or a recurrent neural network (RNN), or a combination thereof.

[0034] One of the core technologies of deep learning is the Convolutional Neural Network (CNN), which is primarily used for image recognition and video processing. CNNs operate by extracting image features through multiple layers and classifying objects based on them. For example, Google's image search technology and object detection systems in autonomous vehicles utilize CNNs to perform accurate analysis.

[0035] In addition, Recurrent Neural Networks (RNNs) and their variant models, such as Long Short-Term Memory (LSTM) and Gated Recurrent Unit (GRU), play important roles in Natural Language Processing (NLP). These models are suitable for processing time-series data or continuous sentences and are widely used in applications such as speech recognition, machine translation, and chatbots. Recently, the emergence of Transformer models has enabled more sophisticated natural language processing, which has become the foundation for large-scale language models such as GPT.

[0036] The screen (130) may be a display unit of the computing device (100). In the method for controlling the computing device (100) of the present disclosure, the artificial intelligence module (120) of the computing device (100) controls the computer device according to the user's control command, thereby performing a task corresponding to the user's control command and displaying it on the screen (130).

[0037] Referring to FIG. 2, a method for controlling a computing device (100) according to an embodiment of the present disclosure is

[0038] (A) A step in which a camera (110) acquires standard word image data, which is a video of a user pronouncing a standard word, and transmits it to an artificial intelligence module (120);

[0039] (B) The artificial intelligence module (120) learns standard word image data and derives a first feature point for changes in the user's lip shape within the standard word image data;

[0040] (C) A step in which the camera (110) acquires image data of the user's face and transmits it to the artificial intelligence module (120);

[0041] (D) The artificial intelligence module (120) determines whether the user speaks based on video data;

[0042] (E) The artificial intelligence module (120) derives a second feature point for a change in the shape of the user's lips in the video data when it is determined that the user is speaking;

[0043] (F) The artificial intelligence module (120) derives a user's control command by comparing a second feature point and a first feature point; and

[0044] (G) The artificial intelligence module (120) may include the step of controlling the computing device (100) based on the derived user control command.

[0045] The aforementioned steps (A) and (B) are processes for training the artificial intelligence module (120) on changes in the user's lip shape for each word when the user speaks.

[0046] (A) In step (A), as an example, a computing device (100) may display standard words on a screen (130) and have a user pronounce the standard words only with their lips without sound, and while the user pronounces the standard words, standard word image data is acquired through a camera (110), and the aforementioned standard word image data may be transmitted to an artificial intelligence module (120). Here, the standard word image data may be a video of the user pronouncing the standard words. In one embodiment, the standard words may be characterized by being configured to include a bilabial plosive, a bilabial nasal, an alveolar nasal, an open vowel, an unrounded vowel, and a rounded vowel.

[0047] Regarding consonants, standard words may include bilabial plosives, bilabial nasals, and alveolar nasals. Bilabial plosives are sounds produced when the lips are completely closed and then released. Examples include p (ㅍ) and b (ㅂ). Bilabial nasals are sounds produced when the lips are completely closed and then released along with a nasal sound. Examples include m (ㅁ). Alveolar nasals are sounds produced when the lips open and the tongue touches the alveolar ridge to finish. Examples include n (ㄴ).

[0048] Regarding vowels, standard words may include open vowels, unrounded vowels, and rounded vowels. Open vowels are sounds produced by opening the mouth wide during pronunciation. Examples include a (ㅏ) and ya (ㅑ). Unrounded vowels are sounds produced by spreading the lips as far to the left and right as possible, resembling a smile. Examples include i (ㅣ). Rounded vowels are sounds produced by rounding and protruding the lips forward during pronunciation. Examples include u (ㅜ) and yu (ㅠ).

[0049] For example, in English, standard words may include pop, beep, and moon, while in Korean, standard words may include rice, mouth, and spring. However, standard words are not limited to the aforementioned examples.

[0050] In step (B), the artificial intelligence module (120) learns standard word image data and can derive a first feature point regarding the change in the user's lip shape for the corresponding standard word. Referring to FIG. 4, when deriving the first feature point, the artificial intelligence module (120) can derive the first feature point by, for example, extracting the movement of points on the inner contour of the user's lips. The reason for deriving the first feature point by extracting the movement of points on the inner contour of the user's lips is that the movement of points on the inner contour of the user's lips is most significantly involved when the user pronounces. In addition, if the artificial intelligence module (120) extracts the movement of points on the entire lip image, it takes a long time and consumes a lot of resources because it has to process a large amount of data. By extracting the movement of points on the inner contour of the user's lips to derive the first feature point, the artificial intelligence module (120) has the advantage of saving time and resources in deriving the first feature point. In particular, in the case of an on-device artificial intelligence module (120), it may be more advantageous because the resources of the computing device (100) are limited.

[0051] (C) In step (C), the camera (110) can acquire image data of the user's face and transmit it to the artificial intelligence module (120).

[0052] (D) In ​​step (D), the artificial intelligence module (120) can determine whether the user is speaking based on video data. In one embodiment, the artificial intelligence module (120) can divide the video data received from the camera (110) into frames of a certain length. Within a frame, the artificial intelligence module (120) can determine that the user is speaking if there is movement of the user's lips, and determine that the user is not speaking if there is no movement of the user's lips.

[0053] Referring to FIG. 3, the artificial intelligence module (120) is an algorithm that determines whether a user is speaking based on video data. The artificial intelligence module (120) determines whether a user is speaking based on a plurality of frames and compares a first number of frames determined to be in which the user is speaking with a second number of frames determined not to be in which the user is speaking. If the first number is greater than the second number, it is determined that the user is speaking; if the first number is less than the second number, it is determined that the user is not speaking; and if the first number and the second number are the same, it is determined that the previous state is maintained. In the control method of the computing device (100) disclosed in the present disclosure, determining whether a user is speaking is very important as an initial process for deriving a control command of the user. As described above, by determining whether a user is speaking based on a plurality of frames, there is an advantage of being able to determine the time of the user's speech more accurately.

[0054] (E) In step (E), the artificial intelligence module (120) can derive a second feature point regarding the change in the shape of the user's lips in the image data received from the camera (110) when it determines that the user is speaking.

[0055] As described above in the explanation of step (B), when the artificial intelligence module (120) derives the second feature point, for example, it may derive the second feature point by extracting the movement of points on the inner contour of the user's lips. The reason for deriving the second feature point by extracting the movement of points on the inner contour of the user's lips is that, as described above, the movement of points on the inner contour of the user's lips is most significantly involved when the user pronounces. Furthermore, since the artificial intelligence module (120) must process a large amount of data when extracting the movement of points on the entire lip image, it takes a long time and consumes a lot of resources. Therefore, by extracting the movement of points on the inner contour of the user's lips to derive the second feature point, the artificial intelligence module (120) has the advantage of saving time and resources in deriving the second feature point. In particular, if it takes a long time to derive the second feature point, it will take a long time to derive the control command for the computing device (100) in step (F) to be described later, which may cause inconvenience to the user. Therefore, there is an advantage in that the movement of points on the inner contour of the user's lips is extracted to derive the second feature point, thereby preventing such inconvenience to the user.

[0056] (F) In step, the artificial intelligence module (120) can derive a user's control command by comparing the second feature point and the first feature point. The artificial intelligence module (120) can derive a user's control command by deriving a first feature point that matches the second feature point and deriving a phoneme among the standard words corresponding to the derived first feature point.

[0057] (G) In step, the artificial intelligence module (120) can control the computing device (100) based on the derived user control command. The user control command may be a click, double-click, scroll down, scroll up, close a window, etc., but is not limited thereto. Accordingly, the computing device (100) can implement a screen (130) corresponding to the user control command.

[0058] FIG. 5 is a flowchart of a method for controlling a computing device (100) according to one embodiment of the present disclosure, and FIG. 6 is an explanation of the correction of image data of a method for controlling a computing device (100) according to an embodiment of the present disclosure.

[0059] Referring to FIG. 5, the method for controlling a computing device (100) of the present disclosure is,

[0060] (H) The artificial intelligence module (120) may further include a step of correcting image data.

[0061] Referring to FIG. 6, in step (H), the artificial intelligence module (120) derives a first triangle including the center point of the user's two eyes' pupils and the tip point of the nose in the standard word image data and a perpendicular line of the first triangle's centroid, and derives a second triangle including the center point of the user's two eyes' pupils and the tip point of the nose in the image data and a perpendicular line of the second triangle's centroid, and compares the shape and size of the first triangle and the shape and size of the second triangle, and compares the slope of the perpendicular line of the centroid of the first triangle with respect to a straight line perpendicular to the ground, thereby correcting the image data by considering the distance at which the user's face is located and the angle of the user's face.

[0062] For example, the artificial intelligence module (120) can derive the difference in lateral tilt of the user's face by using the angle formed by the perpendicular line of the centroid of the first triangle and the perpendicular line of the centroid of the second triangle. For example, the artificial intelligence module (120) can derive the difference in distance from the camera (110) to the user's face by using the difference in length of one side connecting the pupil points of the user's two eyes in the first triangle and the length of one side connecting the pupil points of the user's two eyes in the second triangle, and the difference in area between the first triangle and the second triangle. For example, the artificial intelligence module (120) can derive the difference in front and back tilt of the user's face by using the difference in angle formed by two sides connecting the pupil points of the user's two eyes in the first triangle and the tip of the user's nose in the second triangle and the angle formed by two sides connecting the pupil points of the user's two eyes and the tip of the user's nose. The image data can be corrected by using the standard word image data derived in this way, the difference in lateral tilt of the user's face within the image data, the difference in distance from the camera (110) to the user's face, and the difference in front and back tilt of the user's face, to convert the inner contour of the user's lips within the image data into coordinates of the user's face within the standard word image data.

[0063] Referring to FIG. 5, the method for controlling a computing device (100) of the present disclosure is,

[0064] (I) The artificial intelligence module (120) tracks the user's gaze within the video data; and

[0065] (J) The artificial intelligence module (120) may further include the step of deriving the area of ​​interest that the user’s gaze is looking at among the screen (130) of the computing device (100) based on the tracked user’s gaze.

[0066] Eye tracking is a technology that analyzes a user's eye movements to predict where their gaze will rest in the future, and it is utilized in various fields such as human-computer interaction (HCI), virtual reality (VR), augmented reality (AR), marketing analytics, and assistive systems for the visually impaired.

[0067] In one embodiment, in step (J), the artificial intelligence module (120) may apply an autoregressive Transformer-based gaze prediction model that receives data such as past user gaze coordinates, changes in the user's pupil size, and the user's head movements to predict the gaze at the next point in time. Here, the area of ​​interest may be an icon, a window minimize, maximize, close, and scroll area within the screen (130) of the computing device (100), and text within the window. The autoregressive Transformer model has a structure that predicts step by step based on past data and can progressively predict the gaze at the next point in time by utilizing gaze data from the previous point in time.

[0068] The basic structure of this model can be broadly divided into four stages. First, during the input data preprocessing stage, the user's eye tracking information (e.g., pupil position, head movement, context within the screen (130), etc.) is refined and converted into input in a time-series format. Subsequently, since the Transformer itself does not recognize the order, position encoding is applied to the input data to preserve position information. In the encoder-decoder structure, which is the core of the model, the encoder processes the input data to extract features, and the decoder predicts the next time point step by step based on previous eye data using an autoregressive method. During this process, an attention mechanism is applied to assign weights to important information at specific times points. Finally, the eye coordinates predicted by the model are post-processed to be visually represented or converted for use in specific applications. As a result, uncertainty in eye tracking caused by the slight trembling of the pupils even when the user is looking at one spot can be prevented, and user fatigue caused by the burden of having to stare intently at one spot can also be prevented.

[0069] As an example, the method for controlling a computing device (100) of the present disclosure may be characterized in that, in step (F) described above, the artificial intelligence module (120) derives a user's control command by considering the correlation with the derived region of interest. In this way, by deriving the user's control command by the artificial intelligence module (120) considering the correlation with the region of interest, the user's control command can be recognized and derived more accurately by considering the user's context, compared to deriving a control command simply by lip reading. As an example, when the region of interest that the user is looking at is "window" and the user pronounces "close," the artificial intelligence module (120) can derive a "control command to close the window" with a high probability.

[0070] In one embodiment, the method for controlling a computing device (100) disclosed in this invention may further include a step of recognizing a user's face. The user's face recognition process involves initially receiving a first photograph of a plurality of user faces from a camera (110) of the computing device (100), after which an artificial intelligence module (120) derives feature points of the user's face. Subsequently, when a second photograph of the user's face is received from the camera (110), the artificial intelligence module (120) derives feature points of the user's face within the second photograph and authenticates the user based on whether these feature points match the feature points of the user's face within the first photograph. As a result, there is an advantage that only an authenticated user can control the computing device (100) using the method for controlling the computing device (100) disclosed in this invention, and that a person other than the user cannot control the computing device (100).

[0071]

[0072] The disclosed content is merely illustrative and can be modified and implemented in various ways by a person skilled in the art without departing from the gist of the claim in the patent claims; therefore, the scope of protection of the disclosed content is not limited to the specific embodiments described above.

[0073]

[0074] [Explanation of the symbol]

[0075] 100: Computing device

[0076] 110: Camera

[0077] 120: Artificial Intelligence Module

[0078] 130: Screen

[0079] The present disclosure is available for use in the computer industry.

Claims

1. A method for controlling a computing device using lip reading, The above computing device includes a camera capable of acquiring an image of a user's face and an artificial intelligence module capable of reading the user's lips to determine the user's control command. The above camera acquires standard word video data, which is a video of a user pronouncing a standard word, and transmits it to the above artificial intelligence module; The above artificial intelligence module learns the above standard word image data and derives a first feature point regarding a change in the user's lip shape within the above standard word image data; The above camera acquires image data of the user's face and transmits it to the artificial intelligence module; The above artificial intelligence module includes a step of determining whether the user speaks based on the above video data; The above artificial intelligence module includes the step of deriving a second feature point regarding a change in the shape of the user's lips within the image data when it is determined that the user is speaking; The artificial intelligence module comprises the step of deriving a user's control command by comparing the second feature point and the first feature point; and A method for controlling a computing device comprising the step of the artificial intelligence module controlling the computing device based on the derived user's control command.

2. In Claim 1, A method for controlling a computing device characterized by the above standard words being configured to include bilabial plosives, bilabial nasals, alveolar nasals, open vowels, unrounded vowels, and rounded vowels.

3. In Claim 1, In the step of determining whether the above user has spoken, The artificial intelligence module divides the video data into frames of a certain length, and within the frames, If there is movement of the user's lips, it is determined that the user is speaking, and A method for controlling a computing device characterized by determining that the user is not speaking when there is no movement of the user's lips.

4. In Claim 3, The artificial intelligence module determines whether the user is speaking based on a plurality of the frames, and compares a first number of frames determined to be in which the user is speaking with a second number of frames determined not to be in which the user is speaking, If the first number is greater than the second number, it is determined that the user is speaking, and If the first number is smaller than the second number, it is determined that the user is not speaking, and A computing device control method characterized by determining that the previous state is maintained when the first number and the second number are the same.

5. In Claim 1, In the step of deriving a first feature point regarding the change in the user's lip shape and the step of deriving a second feature point regarding the change in the user's lip shape, A computing device control method characterized by the artificial intelligence module extracting the movement of points on the inner contour of the user's lips to derive the first feature point and the second feature point.

6. In Claim 1, The above artificial intelligence module further includes a step of correcting the image data; and In the step of correcting the above image data, A computing device control method characterized by the above artificial intelligence module deriving a first triangle including the center point of the user's two pupils and the tip point of the nose for the user's face within the standard word image data and a perpendicular line to the centroid of the first triangle, deriving a second triangle including the center point of the user's two pupils and the tip point of the nose for the user's face within the image data and a perpendicular line to the centroid of the second triangle, comparing the shape and size of the first triangle and the shape and size of the second triangle, and comparing the slope of the perpendicular line to the centroid of the first triangle with respect to a straight line perpendicular to the ground, thereby considering the distance at which the user's face is located and the angle of the user's face.

7. In Claim 1, The above artificial intelligence module includes the step of tracking the user's gaze within the above video data; and A computing device control method characterized by further including the step of the artificial intelligence module deriving an area of ​​interest that the user’s gaze is looking at among the screens of the computing device based on the tracked user’s gaze.

8. In Claim 7, In the step of deriving the area of ​​interest that the aforementioned user's gaze is looking at, A computing device control method characterized by the above artificial intelligence module receiving data such as past user gaze coordinates, changes in user pupil size, and user head movements, and applying an autoregressive Transformer-based gaze prediction model to predict the gaze at the next point in time.

9. In Claim 7, In the step of deriving the control command of the above user, A computing device control method characterized by the above artificial intelligence module deriving the user's control command by considering the correlation with the derived region of interest.