Balance management system and method for generating balance state information and executing balance rehabilitation program by tracking change of eyeballs and head circumference in video, recording medium storing program for implementing same, and computer program stored in recording medium
The balance management system uses artificial neural networks to analyze head and eye movements from video, addressing the limitations of existing vestibular rehabilitation and nystagmus testing, enabling accurate diagnosis and rehabilitation of dizziness.
Patent Information
- Application Number
- JP2024179367
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-05
- Filing Date
- 2024-10-11
- Publication Date
- 2026-02-18
- Estimated Expiration
- 2044-10-11
AI Technical Summary
Existing vestibular rehabilitation methods lack accuracy and require additional equipment, while nystagmus testing is difficult to administer and analyze, leading to inadequate diagnosis and delayed recovery from dizziness.
A balance management system using artificial neural networks to analyze head and eye movements from video footage, generating balance function status information without additional medical equipment, and providing a rehabilitation program.
Enables accurate balance rehabilitation and rapid diagnosis of dizziness using a smartphone, distinguishing between peripheral and central dizziness without separate devices, and offering a remote rehabilitation program.
Smart Images

Figure 2026027160000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a balance management system and method for generating balance function status information and implementing a balance function rehabilitation program. More specifically, the present invention relates to a balance management system and method that generates balance function status information for a subject by tracking changes in the subject's eyeballs and head circumference in video footage of the subject taken using at least one camera, and implements a program for the subject's balance function rehabilitation. The present invention also relates to a recording medium on which a program for implementing the system and method is stored, and a computer program stored on the recording medium. [Background technology]
[0002] The material described in this section merely provides background information for the embodiments described herein and does not necessarily constitute prior art.
[0003] The human body maintains balance by fixing the gaze regardless of head movement. The inner ear detects head movement and transmits the detected signal to the brain. When the signal detected by the inner ear is transmitted to the brain, the brain stimulates the extraocular muscles of the eyeball, inducing the vestibulo-ocular reflex, which moves the eyes in the opposite direction to the head movement, thereby contributing to maintaining balance. Therefore, if the inner ear is damaged or there is a problem with the signal integration process in which head movement detected by the inner ear is transmitted to the brain, the vestibulo-ocular reflex is not triggered, and the human body cannot maintain balance, resulting in severe dizziness.
[0004] In clinical practice, balance abnormalities are determined through tests that evaluate head movement versus eye movement. To do this, inertial measurement units (IMUs) are used to measure the speed of head movement versus eye movement, and eye-tracking video cameras are used to record eye movement, to determine whether a patient has a balance abnormality.
[0005] Meanwhile, when balance dysfunction occurs, acute dizziness occurs initially, but gradually improves over time due to the compensatory action of the vestibular system. Improvement of dizziness occurs due to the compensatory action of the vestibular system, and through this compensatory action, the human body activates the proprioception of the eyes and muscles instead of the damaged balance function, i.e., the inner ear, contributing to symptom improvement. Such vestibular compensation accelerates the more the eyes and head are moved, and clinical vestibular rehabilitation treatment is based on the principle of activating vestibular compensation by moving the eyes and head.
[0006] Existing vestibular rehabilitation methods include therapists using business cards to guide the patient's gaze and head movements, or more recently, virtual reality devices to guide the patient's head and eye movements. However, existing rehabilitation exercise methods using business cards cannot guarantee accurate rehabilitation exercises. Furthermore, rehabilitation exercise methods using virtual reality devices can track the patient's head and eye movements and provide feedback to the patient, allowing for more accurate rehabilitation exercises, but they have limitations in clinical practice due to the need for additional equipment such as virtual reality devices.
[0007] Nystagmus testing, which evaluates the vestibulo-ocular reflex when dizziness occurs, is not only cost-effective but has also been proven to have superior diagnostic sensitivity than MRI. However, it is difficult to administer and analyze, and there is no equipment to perform the test. As a result, nystagmus testing is not widely used in emergency rooms and primary care hospitals, where the most dizziness patients visit, and dizziness is often diagnosed through unnecessary imaging and blood tests. This leads to wasteful medical expenses in terms of cost-effectiveness and often leads to poor patient outcomes due to inappropriate diagnoses. In particular, vestibular rehabilitation treatment, which is intended to relieve symptoms of acute dizziness, is difficult for patients to perform on their own, delaying recovery from dizziness and reducing quality of life. [Prior art documents] [Patent documents]
[0008] [Patent Document 1] Korean Patent Publication No. 10-2004-0107677 Summary of the Invention [Problem to be solved by the invention]
[0009] The present specification aims to provide a balance management system for generating balance status information and implementing a balance rehabilitation program.
[0010] The present specification is not limited to the above-mentioned problems, and other problems not mentioned will be clearly understood by those skilled in the art from the following description. [Means for solving the problem]
[0011] The balance function management system according to the present specification, which solves the above-mentioned problems, includes at least one processor and a memory storing instructions executable by the processor and storing at least one artificial neural network model executed on a computing device, wherein the at least one processor inputs n video frame images of a subject captured through n (a natural number) cameras into at least one artificial neural network model to obtain at least one of information related to the subject's head coordinates, pupil center coordinates, and eye phase changes in the order of the m-th video frame images (a natural number from 1 to n), and uses the information to generate balance function status information or information related to head movement and eye movement for carrying out a balance function rehabilitation program.
[0012] According to one embodiment of the present specification, the at least one processor may include a head coordinate acquisition unit that executes a first artificial neural network model stored in the memory, inputting a frame image of an mth video or a multi-frame image obtained by concatenating frame images of n videos into the first artificial neural network model, and acquiring information related to head coordinates according to the mth video; an eye coordinate acquisition unit that executes a second artificial neural network model stored in the memory, inputting information related to the head coordinates into the second artificial neural network model, and acquiring information related to pupil center coordinates according to the mth video; and a phase change acquisition unit that executes a third artificial neural network model stored in the memory, inputting information related to pupil coordinates according to the time order of the frame image of the mth video or the multi-frame image into the third artificial neural network model, and acquiring information related to eye phase changes according to the mth video.
[0013] According to one embodiment of the present specification, the first artificial neural network model may be an artificial neural network model trained by the at least one processor using facial feature points and coordinates of the feature points extracted from at least one frame image of a video in which a human face is captured or a multi-frame image in which multiple frame images of a video are connected as training data.
[0014] In this case, the feature points may be feature points located within a predetermined area in the frame image or the multiple frame image.
[0015] According to one embodiment of the present specification, the second artificial neural network model may be an artificial neural network model trained by the at least one processor to generate information related to the coordinates of the pupil center by using, as training data, an image of an eyeball region extracted to include an eyeball from at least one frame image of a video in which a human face is captured or a multi-frame image in which a plurality of frame images of a video are connected.
[0016] In this case, the at least one processor generates a pupil region image in which the region occupied by the pupil and the remaining region in the eyeball region image have different pixel values, and can train the second artificial neural network model using the pupil region image or an array of pixel values of the pupil region image as training data.
[0017] In this case, the at least one processor can generate feature points of the eyeball and coordinate information of the feature points in the eyeball region image, and can train the second artificial neural network model to generate horizontal coordinate values and vertical coordinate values of the pupil center using coordinates of a plurality of preset feature points.
[0018] According to another embodiment of the present specification, the memory may store data of at least one virtual object, and the second artificial neural network model may be an artificial neural network model trained by the at least one processor to generate information related to coordinates of a pupil center using training data including parameter values obtained by changing at least one of parameters related to head rotation, eye rotation, and camera settings of the virtual object, and images of the virtual object acquired according to the parameter values.
[0019] According to one embodiment of the present specification, the third artificial neural network model may be an artificial neural network model trained by the at least one processor to generate an eyeball rotation value by using, as training data, information related to eyeball phase changes generated according to the time sequence of eyeball region images extracted to include eyeballs from frame images of a video in which a human face is captured or a multi-frame image obtained by connecting multiple frame images of a video.
[0020] According to one embodiment of the present specification, the third artificial neural network model may be an artificial neural network model trained by the at least one processor using information obtained by comparing pixel values of the area occupied by the iris between eye area images corresponding to each frame image of a video or a multi-frame image of a video in which the person's face is captured.
[0021] In this case, the at least one processor may calculate the size of the area occupied by the pupil in the eye region image, and adjust the size of the target eye region image using the size of the area occupied by the pupil in the previous eye region image.
[0022] According to one embodiment of the present specification, the at least one processor may include a head movement generation unit that generates information related to head movement in the mth video (a natural number from 1 to n) using information related to head coordinates generated by at least one artificial neural network model; an eye movement generation unit that generates information related to eye movement in the mth video using information related to pupil center coordinates or eye phase changes generated by at least one artificial neural network model; a velocity information generation unit that generates information related to the head and eye movement velocity in the mth video using the information related to the head movement and eye movement; and a balance function status information generation unit that generates status information of the subject's balance function using information related to the head and eye movement velocity.
[0023] According to one embodiment of the present specification, the speed information generating unit may generate information related to the speed of head movement and eye movement within a predetermined time based on a point in time when the head movement exceeds a predetermined threshold value.
[0024] According to one embodiment of the present specification, the balance function status information generating unit can calculate a gain value using a time value at which the head movement speed is at its maximum and a time value at which the eye movement speed is at its maximum.
[0025] According to one embodiment of the present specification, when n is 2 or more, the head movement generation unit may further generate reference head movement information by calculating statistics of information related to head coordinates according to the mth video, and the eye movement generation unit may further generate reference eye movement information by calculating statistics of information related to pupil center coordinates and eye phase changes according to the mth video.
[0026] According to another embodiment of the present specification, the at least one processor may include a head movement generation unit that generates head movement information in the mth video (a natural number from 1 to n) using information related to head coordinates obtained from at least one artificial neural network model; an eye movement generation unit that generates eye movement information in the mth video using information related to pupil center coordinates or eye phase changes obtained from at least one artificial neural network model; a target output unit that outputs a virtual target to a display device; a head direction providing unit that provides directional information regarding head movement to a subject; and a feedback providing unit that provides feedback based on the subject's head movement and eye movement.
[0027] A balance function status information generating method according to the present specification for solving the above-mentioned problems is a balance function management system including at least one processor and a memory storing instructions executable by the processor and storing at least one artificial neural network model executed on a computing device, and includes: an information acquisition step in which the at least one processor inputs n video frame images of a subject through n (a natural number) cameras into at least one artificial neural network model to acquire information related to the subject's head coordinates, pupil center coordinates, and eye phase changes in the order of the m-th video frame image (a natural number from 1 to n); and a balance function status information generation step in which the at least one processor generates head movement information and eye movement information using the information generated in the information acquisition step, calculates head movement speed and eye movement speed, and generates information related to the state of balance function using the information related to the head movement speed and eye movement speed.
[0028] To solve the above-mentioned problems, the balance rehabilitation method according to the present specification provides a balance management system including at least one processor and a memory storing instructions executable by the processor and storing at least one artificial neural network model executed on a computing device. The balance rehabilitation method may include: an information acquisition step in which the at least one processor inputs frame images of n videos of a subject captured through n (a natural number) cameras into at least one artificial neural network model to acquire information related to the subject's head coordinates, pupil center coordinates, and eye phase changes according to the frame order of the mth video (a natural number from 1 to n); and a balance rehabilitation step in which the at least one processor generates head movement information and eye movement information using the information generated in the information acquisition step, and tracks the subject's head and eye movements to perform a balance rehabilitation program.
[0029] The balance function status information generating method according to the present specification can be realized in the form of a computer program that is written to execute each step in a computer and is recorded on a computer-readable recording medium.
[0030] The balance function rehabilitation method according to the present specification can be realized in the form of a computer program that is created to perform each step on a computer and is recorded on a computer-readable recording medium.
[0031] Other details of the invention are included in the detailed description and drawings. [Effects of the Invention]
[0032] According to one aspect of the present specification, a balance function management system may input video of a subject undergoing a balance function status test using a smartphone, webcam, etc. into an artificial neural network model, calculate eye movement and head movement without the need for additional medical equipment, and generate balance function status information.
[0033] According to another aspect of the present disclosure, the balance management system can be utilized as a clinical decision support system for dizziness patients that generates the balance status information.
[0034] According to another aspect of the present specification, the balance function management system can provide important information for quickly distinguishing between peripheral and central dizziness using only a single smartphone in an emergency room environment, without the need for a separate medical device or for the patient to wear separate equipment.
[0035] According to another aspect of the present specification, the balance function management system may provide a remote balance function rehabilitation program by outputting information on changes in head circumference and eye movement of a subject to a display device to enable more accurate balance rehabilitation.
[0036] The effects of the present invention are not limited to those mentioned above, and other effects not mentioned will be clearly understood by those skilled in the art from the following description. [Brief explanation of the drawings]
[0037] [Figure 1] FIG. 1 is a reference diagram illustrating a scene in which a subject is photographed by at least one camera. [Figure 2] 1 is a reference diagram illustrating a scene in which a subject is photographed using a camera included in a terminal; [Figure 3] 1 is a block diagram of a balance function management system according to a first embodiment of the present specification. [Figure 4] 10 is an example of a scene in which a subject is being tested for balance status and / or undergoing a balance rehabilitation program according to one embodiment of the present disclosure. [Figure 5] 10 is an example of a scene in which a subject is being tested for balance status and / or undergoing a balance rehabilitation program according to another embodiment of the present disclosure. [Figure 6] 1 is a diagram illustrating an example of pre-processing an eye area image according to an embodiment of the present disclosure; [Figure 7] 10 is a diagram illustrating an example of generating training data for a second artificial neural network model according to another embodiment of the present specification. [Figure 8] FIG. 1 is a block diagram of a balance management system for generating balance status information according to an embodiment of the present disclosure. [Figure 9] 10 is an example of balance function status information output on a display according to one embodiment of the present specification. [Figure 10] FIG. 1 is a block diagram of a balance management system for implementing a balance rehabilitation program according to an embodiment of the present specification. [Figure 11] 1 is a diagram illustrating an example of a scene in which a balance rehabilitation program is performed according to an embodiment of the present specification; [Figure 12] 1 is a block diagram of a balance function management system that generates balance function status information and performs a balance function rehabilitation program according to an embodiment of the present specification. [Figure 13] FIG. 10 is a block diagram of a balance function management system according to a second embodiment of the present specification. [Figure 14] FIG. 10 is a reference diagram illustrating an example in which a head coordinate learning unit according to an embodiment of the present specification connects frame images. [Figure 15] 1 is an illustration of a multi-frame image of a subject undergoing a balance status test and / or balance rehabilitation program according to an embodiment of the present disclosure; [Figure 16] 1 is an illustration of a multi-frame image of a subject undergoing a balance status test and / or balance rehabilitation program according to an embodiment of the present disclosure; [Figure 17] 1 is a diagram illustrating an example of pre-processing an eye area image according to an embodiment of the present disclosure; [Figure 18] FIG. 1 is a block diagram of a balance management system for generating balance status information according to an embodiment of the present disclosure. [Figure 19]FIG. 1 is a block diagram of a balance management system for implementing a balance rehabilitation program according to an embodiment of the present specification. [Figure 20] 1 is a block diagram of a balance function management system that generates balance function status information and performs a balance function rehabilitation program according to an embodiment of the present specification. [Figure 21] 1 is a flowchart of a balance function status information generating method according to an embodiment of the present specification. [Figure 22] 1 is a flowchart of a balance rehabilitation method according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0038] The advantages and features of the invention disclosed herein, as well as methods for achieving them, will become apparent from the following detailed description of the embodiments in conjunction with the accompanying drawings. However, the present specification is not limited to the embodiments disclosed below, and may be realized in various different forms. The present embodiments are provided merely to complete the disclosure of the specification and to fully inform those skilled in the art (hereinafter referred to as "those skilled in the art") of the scope of the specification, and the scope of the specification is defined only by the scope of the claims.
[0039] The terms used in this specification are for the purpose of describing the embodiments and are not intended to limit the scope of the specification. In this specification, the singular form includes the plural form unless otherwise stated in the phrase. When used in this specification, "comprises" and / or "comprising" do not exclude the presence or addition of one or more other elements other than the elements mentioned.
[0040] Throughout the specification, the same reference numerals refer to the same components, and the term "and / or" includes each and every combination of one or more of the referenced components. Even if "first," "second," etc. are used to describe various components, these components are not limited by these terms. These terms are merely used to distinguish one component from another. Therefore, a first component referred to below may also be a second component within the technical spirit of the present invention.
[0041] Unless otherwise defined, all terms (including technical and scientific terms) used herein are used as they are commonly understood by those of ordinary skill in the art to which this specification pertains. Furthermore, terms defined in commonly used dictionaries are not interpreted ideally or excessively unless explicitly defined otherwise.
[0042] An artificial neural network (ANN) is an artificial intelligence that is realized by interconnecting artificial neurons that are mathematically modeled after the neurons that make up the human brain.
[0043] As used herein, an "artificial neural network model" may be composed of a collection of interconnected computational units that may generally be referred to as nodes. Such nodes may also be referred to as neurons. A neural network is composed of at least one or more nodes. The nodes (or neurons) that make up a neural network may be interconnected by one or more links.
[0044] In a neural network, one or more nodes connected via links can form a relative input node-output node relationship. The concepts of input node and output node are relative, and any node that has an output node relationship with one node can also have an input node relationship with another node, and vice versa. As described above, the input node-output node relationship can be generated around links. One or more output nodes can be connected to one input node via links, and vice versa.
[0045] The initial input node may refer to one or more nodes in a neural network to which data is directly input without passing through a link in relation to other nodes. Alternatively, the initial input node may refer to a node in a neural network that does not have other input nodes connected by a link in relation to other nodes based on links. Similarly, the final output node may refer to one or more nodes in a neural network that do not have output nodes in relation to other nodes. Furthermore, the hidden node may refer to a node that constitutes a neural network, rather than the initial input node or the final output node.
[0046] As used herein, "inputting" data into an artificial neural network model means that a value is input to the first input node. As used herein, "obtaining a value," "outputting data," "obtaining information," etc. from an artificial neural network means that data is output from the last output node.
[0047] A deep neural network (DNN) may refer to a neural network that includes multiple hidden layers in addition to an input layer and an output layer. Deep neural networks may include convolutional neural networks (CNNs), recurrent neural networks (RNNs), autoencoders, generative adversarial networks (GANs), restricted Boltzmann machines (RBMs), deep belief networks (DBNs), Q-networks, U-networks, Siamese networks, generative adversarial networks (GANs), etc. The above descriptions of deep neural networks are merely examples, and the present disclosure is not limited thereto.
[0048] A neural network can be trained using at least one of supervised learning, unsupervised learning, semisupervised learning, or reinforcement learning. Training a neural network can be a process of applying knowledge to the neural network to enable the neural network to perform a specific operation.
[0049] Neural networks can be trained to minimize output errors. Training a neural network involves repeatedly inputting training data into the neural network, calculating the error between the neural network's output and the target for the training data, and backpropagating the neural network's error from the output layer to the input layer of the neural network in a direction that reduces the error, thereby updating the weights of each node in the neural network. In supervised learning, training data in which a correct answer is labeled is used (i.e., labeled training data). In unsupervised learning, the correct answer may not be labeled. For example, in supervised learning for data classification, the training data may be data in which a category is labeled. Labeled training data is input to the neural network, and the error can be calculated by comparing the neural network's output (category) with the label of the training data. As another example, in unsupervised learning for data classification, the error can be calculated by comparing the input training data with the neural network output. The calculated error is backpropagated in the reverse direction (i.e., from the output layer to the input layer) in the neural network, and the connection weight value of each node in each layer of the neural network can be updated through backpropagation. The amount of change in the connection weight value of each node to be updated can be determined by the learning rate. The neural network calculation for input data and backpropagation of the error can constitute a learning cycle (epoch). The learning rate can be applied differently depending on the number of iterations of the neural network learning cycle. For example, in the early stages of learning of the neural network, a high learning rate can be used to quickly ensure a certain level of performance and improve efficiency, and in the later stages of learning, a low learning rate can be used to improve accuracy.
[0050] In this specification, "learning" of an artificial neural network model means that the neural network updates the connection weights of each node so that the output error is minimized, and "learning" in this specification is not limited to a specific learning method.
[0051] As used herein, a "processor" may be composed of one or more cores and may include a processor for data analysis and deep learning, such as a central processing unit (CPU), a general-purpose graphics processing unit (GPGPU), or a tensor processing unit (TPU) of a computing device. The processor may read a computer program stored in a memory and perform data processing for machine learning according to an embodiment of the present specification. According to an embodiment of the present specification, the processor may perform calculations for neural network training. The processor may perform calculations for neural network training, such as processing input data for deep learning (DL) training, feature extraction from the input data, error calculation, and updating neural network weights using backpropagation. At least one of the CPU, GPGPU, and TPU of the processor may process network function training. For example, the CPU and GPGPU may jointly process network function training and data classification using the network function. Furthermore, in an embodiment of the present specification, processors of multiple computing devices may be used together to process network function training and data classification using the network function. Furthermore, the computer program running on a computing device according to an embodiment of the present disclosure may be a CPU, GPGPU, or TPU executable program.
[0052] Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings.
[0053] FIG. 1 is a reference diagram illustrating a scene in which a subject is photographed by at least one camera.
[0054] Referring to FIG. 1, a balance function management system according to an embodiment of the present specification can generate head movement and eye movement information of a subject using n videos of the subject taken through n (a natural number) cameras.
[0055] The n cameras are installed at predetermined positions and can capture the face of the subject, and the n cameras can capture the face of the subject from different angles.
[0056] For example, one camera may be placed in front of the subject to capture the subject's face, and the front of the subject may refer to the direction in which the front of the body faces.
[0057] As another example, two cameras can be installed at predetermined positions to capture the subject's face, with one camera capturing the subject's face on the left side in front of the subject, and the other camera capturing the subject's face on the right side in front of the subject.
[0058] As another example, three cameras may be installed at predetermined positions to capture the subject's face, with one camera installed in front of the subject, one on the left side of the subject, and one on the right side of the subject.
[0059] The number and positions of the cameras are merely examples and are not limited thereto. Depending on the number, positions, and distance of the cameras from the subject, n cameras can capture the subject in various directions and distances.
[0060] The balance function management system can receive video data of the subject captured by the n cameras. The balance function management system can display the received video data on a display device. The balance function management system can display m-th videos of the subject captured by the m-th (a natural number from 1 to n) camera on the display device. Hereinafter, the m-th video can refer to all videos from the first video to the n-th video.
[0061] The balance function management system may generate information related to head and eye movements of the subject using n input videos. The balance function management system may display the information related to head and eye movements on a display device. The balance function management system may also generate information related to head and eye movement speeds using the information related to head and eye movements and display the information on a display device.
[0062] The screen shown on the display device in FIG. 1 corresponds to an example, and the present invention is not limited to this screen.
[0063] FIG. 2 is a reference diagram illustrating a scene in which a subject is photographed by a camera included in a terminal.
[0064] Referring to FIG. 2, according to another embodiment of the present specification, the balance management system may generate information related to head and eye movements using video captured through a front camera (left side of FIG. 2) and / or a rear camera (right side of FIG. 2) of a terminal such as a smartphone or a tablet computer. When capturing an image of a subject using the rear camera of the terminal, the balance management system may display, on a display device, each of the images captured by at least one camera included in the rear of the terminal. The balance management system may generate information related to the subject's head and eye movements from each of the videos captured by at least one camera included in the rear of the terminal.
[0065] The images captured by the front camera and / or rear camera of the terminal may be displayed on the display of the terminal and / or a display device connected to the terminal.
[0066] The number, positions, and orientation of the front and / or rear cameras of the terminal illustrated in Fig. 2 are merely examples and are not intended to be limiting. In addition, at least one additional camera for capturing an image of a subject may be installed.
[0067] The balance function management system according to the present specification may be implemented in the form of a computing device such as a computer, laptop, smartphone, and / or tablet computer, which is an example and is not limited to such devices.
[0068] The subject may be photographed through a camera included in the computing device and / or a camera connected to the computing device via a wired and / or wireless connection, and various types of cameras, such as a webcam and an action camera, may be used, but this is by way of example only and is not intended to be limiting.
[0069] The balance function management system may generate information related to the balance function status of a subject using n videos of the subject undergoing a video head impulse test, a spontaneous nystagmus test, a saccade test, etc., which are merely examples and are not limited to the above tests.
[0070] The following describes a balance function management system according to a first embodiment of the present specification. According to the first embodiment of the present specification, the balance function management system generates head and eye movement information of a subject using at least one video captured by at least one camera, and can generate balance function status information and / or implement a balance function rehabilitation program.
[0071] The balance function management system according to the first embodiment of the present specification may include at least one processor and a memory storing instructions executable by the processor and storing at least one artificial neural network model executed on a computing device.
[0072] The balance function management system may include a processor that inputs n video frame images of a subject captured through n cameras into at least one artificial neural network model to acquire at least one of information related to the subject's head coordinates, pupil center coordinates, and eye phase change in the order of the frame images, and the processor may generate head movement information and eye movement information using the information.
[0073] FIG. 3 is a block diagram of a balance function management system according to the first embodiment of the present specification.
[0074] Referring to FIG. 3, the balance function management system 10 according to the first embodiment of the present specification may include a memory 100, a head coordinate learning unit 110, an eyeball coordinate learning unit 120, an eyeball rotation learning unit 130, a head coordinate acquisition unit 140, an eyeball coordinate acquisition unit 150, and a phase change acquisition unit 160.
[0075] The memory 100 may store at least one of a first artificial neural network model that generates information related to head coordinates, a second artificial neural network model that acquires information related to pupil center coordinates, and a third artificial neural network model that acquires information related to eye phase changes.
[0076] According to one embodiment of the present specification, the first artificial neural network model may be trained by the head coordinate training unit 110 using facial feature points and their coordinates extracted from at least one frame image of a video in which a human face is captured as training data.
[0077] The head coordinate learning unit 110 can learn the first artificial neural network model using at least one video of a person's face. The at least one video of a person's face may be a video of the person turning their head.
[0078] As an example, the video of the person's face may be captured by a single camera. The video of the person's face may include an image of the person's entire face. The video of the person's face may be a video of a scene in which the person's face is facing forward and moving their head from side to side. Furthermore, the video of the person's face may be a video of a scene in which the person's face is facing forward and moving their head up and down. The state in which the person's face is facing forward may be a state in which the person's line of sight and the front direction of their body are aligned.
[0079] The video of the person's face may be a video of the person turning their head to the right and then moving their head up and down. The video of the person's face may be a video of the person turning their head to the left and then moving their head up and down.
[0080] Furthermore, the video in which the person's face is captured may be a video in which the person is filmed turning his or her head while gazing at a specific gaze point.
[0081] A video of a person's face may be captured by a camera placed in front of the person, or may be captured by cameras placed at various angles in front of the person, such as on the left, right, upper left, or upper right side.
[0082] In addition, the video of the person's face may be captured by a plurality of cameras, which may be installed at various positions and angles to capture the person. The video captured by a plurality of cameras may refer to a video of the same situation captured by a plurality of cameras.
[0083] The videos of the human faces may all be captured by cameras with the same specifications. Alternatively, the videos of the human faces may be captured by cameras with different specifications. Alternatively, the videos of the human faces may be captured by cameras with different settings.
[0084] The video of a person described above is merely an example, and any video of a person's head and eyeballs or any video of a person's head may be used. Furthermore, the video is not limited by specific camera specifications or settings.
[0085] Preferably, the video of the person may be video taken with a camera having the same settings.
[0086] The head coordinate learning unit 110 may extract feature points on the person's face from each frame image of a video in which the person's face is captured. For example, the head coordinate learning unit 110 may extract feature points located on the person's face from each frame image. For example, the feature points may include feature points on the person's nose tip, left outer canthus (where the outermost eyelid meets the left eye), right outer canthus (where the outermost partial eyelid meets the right eye), and forehead. The feature points may not change position even when the person blinks. The feature points are merely examples and are not limited to these feature points and may further include feature points used for conventional face recognition. Techniques for generating feature points and feature point coordinates on a person's head are widely known to those skilled in the art, so a detailed description will be omitted.
[0087] The head coordinate learning unit 110 may generate coordinate information of feature points extracted from each frame image. The head coordinate learning unit 110 may train a first artificial neural network model to generate 3D head coordinates for each frame image using the feature points and their coordinates extracted from each frame image as learning data. The 3D head coordinates may refer to the 3D coordinates of the head according to a 3D standard head model. The 3D head coordinates may include coordinates for all feature points of the head. The 3D head coordinates may include information related to index numbers for each feature point. In this case, the head coordinate learning unit 110 may train the first artificial neural network model to generate index numbers based on the feature point of the nose. The head coordinate learning unit 110 may also train the first artificial neural network model to further generate head movement information using the 3D coordinates extracted from each frame image.
[0088] In addition, the head coordinate learning unit 110 may further use data on a 3D standard head model such as a frame image from which feature points have been extracted, a 3D Morphable Model (3DMM), or a FLAME (Faces Learned with an Articulated Model and Expressions) model as learning data to train the first artificial neural network model. This is just an example, and various learning data may be further used without being limited by the learning data.
[0089] According to one embodiment of the present specification, when extracting feature points from a plurality of videos captured using a plurality of cameras, the head coordinate learning unit 110 may extract feature points after adjusting the sync of the plurality of videos. The head coordinate learning unit 110 may adjust the sync of the plurality of videos using an algorithm such as a specific audio signal of the plurality of videos or feature point matching-based synchronization, but this is by way of example only and is not limiting, and various techniques widely known to those skilled in the art may be used.
[0090] According to an embodiment of the present specification, the head coordinate learning unit 110 can use feature points having coordinate values according to a preset criterion as learning data. The head coordinate learning unit 110 can select feature points having coordinate values according to a preset criterion as learning data.
[0091] FIG. 4 is an example of a scene in which a subject is being tested for balance status and / or undergoing a balance rehabilitation program according to one embodiment of the present disclosure.
[0092] 4, the video of a person may be a video of a balance function status test and / or a balance function rehabilitation program being performed using the balance function management system 10. The video may be a video of only the subject 200. Alternatively, the video may be a video of the subject 200 and the examiner 201. Hereinafter, the video of a person will be described as a video of a balance function status test and / or a balance function rehabilitation program being performed. However, this is just an example and is not limited to the video.
[0093] When a balance function status test is performed, the subject 200 may be seated in a chair, and the examiner 201 may be standing behind the subject. The examiner 201 may be a medical professional. The balance function management system 10 analyzes the eye movements caused by the head movements of the subject 200 to determine whether the subject has an abnormality in balance function, and can analyze only the head movements and eye movements of the subject 200 to proceed with a rehabilitation program.
[0094] The head coordinate learning unit 110 can extract feature points and the coordinates of the feature points on the face of the target person 200 from the video and train the first artificial neural network model.
[0095] As described above, since the subject 200 is sitting in a chair and the examiner 201 is standing, the feature points extracted from the face of the subject 200 may be located relatively lower than the feature points extracted from the face of the examiner 201. This may mean that the y coordinate values of the feature points extracted from the face of the subject 200 are relatively smaller than the y coordinate values of the feature points extracted from the face of the examiner 201.
[0096] The head coordinate learning unit 110 can extract, as learning data, feature points with relatively small y-coordinate values from feature points at the same parts extracted from the faces of the subject 200 and the examiner 201. The head coordinate learning unit 110 can extract, as learning data, feature points located in an area below a preset y-coordinate value.
[0097] In addition, the head coordinate learning unit 110 may cluster feature points extracted from the faces of the subject 200 and the examiner 201. Thereafter, the head coordinate learning unit 110 may extract a set of feature points located relatively lower as learning data.
[0098] FIG. 5 is an example of a scene in which a subject is being tested for balance status and / or undergoes a balance rehabilitation program according to another embodiment of the present disclosure.
[0099] 5, a balance function status test and / or balance function rehabilitation program may be performed in the moving image while the subject 200' and the examiner 201' are both standing. In this case, feature points extracted from the face of the subject 200' may be closer to the center of the display screen than feature points extracted from the face of the examiner 201'. The head coordinate learning unit 110 may extract feature points having coordinate values relatively closer to the coordinate values of the center of the display screen as learning data.
[0100] As another example, the subject may be closer to the camera than the examiner in actuality. As a result, the subject's face may occupy a larger area in the video than the examiner's face. The head coordinate learning unit 110 may extract, as learning data, a set of feature points that are distributed over a larger area from the set of feature points extracted from the subject's face and the set of feature points extracted from the examiner's face.
[0101] The above-described process in which the head coordinate learning unit 110 extracts only feature points extracted from the subject's face as learning data is merely an example, and is not limited thereto. Various criteria can be set depending on the subject's position, the examiner's position, the camera position, angle, etc. Also, only feature points extracted from the subject's face can be extracted as learning data using markers for distinguishing the subject from the examiner, as well as the positions of the subject's and examiner's faces. Therefore, various embodiments can be realized according to various situations in which a balance function status test and / or rehabilitation test program is performed.
[0102] If the subject rapidly turns his / her head in the video, the subject's face may not be clearly captured in the frame image. In this case, the head coordinate learning unit 110 may not be able to extract feature points from the subject's face. The head coordinate learning unit 110 trains the first artificial neural network model using feature points extracted from the examiner's face, and the first artificial neural network model may calculate inaccurate results.
[0103] According to one embodiment of the present specification, the head coordinate learning unit 110 can track the coordinates of feature points extracted from a frame image preceding a given frame image. As described above, if feature points cannot be extracted from the subject's face in a given frame image, the head coordinate learning unit 110 can extract learning data using an image of the subject in a preceding frame image. The preceding frame image may refer to the closest frame image from which feature points can be extracted from the subject's face among the frame images preceding the given frame image.
[0104] According to one embodiment of the present specification, the second artificial neural network model may be trained by the eyeball coordinate training unit 120 to generate information related to the coordinates of the pupil center by using an eyeball area image, which is an image of an area including an eyeball in a video frame image in which a person's face is captured, as training data.
[0105] The eyeball coordinate learning unit 120 may extract an image of the subject's eyeball region, which is an image of an area including the subject's eyeball, from the video. The eyeball coordinate learning unit 120 may use the image of the subject's eyeball region extracted from each frame image as learning data, and may train the second artificial neural network model to generate information related to the coordinates of the pupil center.
[0106] FIG. 6 is a diagram illustrating an example of pre-processing an eye area image according to an embodiment of the present specification.
[0107] 6, the eyeball coordinate learning unit 120 can extract an eyeball region image 203 according to each frame image 202. The eyeball coordinate learning unit 120 can extract an image inside a bounding box of the eyeball region in each frame image 202 as the eyeball region image 203. The eyeball coordinate learning unit 120 can segment an iris and a pupil region in the eyeball region image 203. Techniques for extracting an eyeball region from a human face and segmenting the iris and the pupil region are well known to those skilled in the art, and therefore, detailed description thereof will be omitted.
[0108] The eyeball coordinate learning unit 120 may estimate a region where the iris and / or pupil is hidden by the eyelid. As shown in FIG. 6, a portion of the iris may be hidden by the upper and lower eyelids. The eyeball coordinate learning unit 120 may estimate the hidden portion using an ellipse fitting algorithm, a circle Hough transform algorithm, or the like. Alternatively, the eyeball coordinate learning unit 120 may segment the iris and pupil regions using an artificial neural network model that has been trained in advance to segment the iris and pupil regions. This is merely an example, and the present invention is not limited to this method.
[0109] The eyeball coordinate learning unit 120 can learn the second artificial neural network model using data in which the iris and / or pupil region is divided in the eyeball region image 203. As an example, the eyeball coordinate learning unit 120 can generate a mask image (204) by dividing the iris and / or pupil region in the eyeball region image 203. The eyeball coordinate learning unit 120 can generate a mask image in which a region occupied by the iris and / or pupil and the remaining region in the eyeball region image 203 have different pixel values. In the mask image, a region occupied by the iris and / or pupil may be displayed in white or black, and the remaining region may be displayed in black or white, which is an example and not limited thereto.
[0110] The eyeball coordinate training unit 120 may train the second artificial neural network model to generate information related to the coordinates of the pupil center using the mask image 204 as training data, or may train the second artificial neural network model using two-dimensional pixel values of the mask image 204 as training data, or may train the second artificial neural network model using the mask image 204 and two-dimensional pixel values as training data.
[0111] As another example, the eyeball coordinate learning unit 120 can train the second artificial neural network model using a heatmap model that divides and shows the iris and / or pupil area in the eyeball area image 203.
[0112] The eyeball coordinate learning unit 120 can learn the second artificial neural network model using at least one of the mask image and the heat map model.
[0113] According to an embodiment of the present specification, the eyeball coordinate learning unit 120 can learn the second artificial neural network model to generate eyeball feature points and feature point coordinate information in the eyeball region image 203 and generate horizontal and vertical coordinate values of the pupil center using the coordinates of a plurality of preset feature points. The eyeball coordinate learning unit 120 can extract normalized pupil center coordinates using the coordinates of a plurality of feature points extracted from the eyeball region image 203.
[0114] The plurality of feature point coordinates may include a feature point with a relatively smallest x-axis coordinate value and a feature point with a relatively largest x-axis coordinate value among feature points whose y-axis coordinates are within a preset range in the eyeball region image 203.
[0115] The plurality of feature point coordinates may include a feature point with a relatively smallest y-axis coordinate value and a feature point with a relatively largest y-axis coordinate value among feature points whose x-axis coordinates are within a preset range in the eyeball region image 203 captured from the front of the person.
[0116] For example, the eyeball coordinate learning unit 120 may extract the horizontal coordinate of the normalized pupil center using the feature point 203-1 for the inner canthus of the eye and the feature point 203-2 for the outer canthus of the eye among the feature points extracted from the eyeball region image 203. The eyeball coordinate learning unit 120 may use a line segment connecting the feature point 203-1 for the inner canthus of the eye and the feature point 203-2 for the outer canthus of the eye as a horizontal axis for the pupil center coordinate. The x-coordinate of the feature point 203-1 for the inner canthus of the eye and the x-coordinate of the feature point 203-2 for the outer canthus of the eye may correspond to the extreme values of the horizontal axis. The difference between the x-coordinate of the feature point 203-1 for the inner canthus of the eye and the x-coordinate of the feature point 203-2 for the outer canthus of the eye may represent the length of the entire horizontal axis. The eyeball coordinate learning unit 120 may calculate the horizontal coordinate of the normalized pupil center using the horizontal coordinate value of the pupil center relative to the length of the entire horizontal axis.
[0117] In addition, the eyeball coordinate learning unit 120 may extract the vertical coordinate of the normalized pupil center using feature points associated with the upper eyelid and the lower eyelid among the feature points extracted from the eyeball region image 203. In this case, the feature point associated with the upper eyelid may be a feature point 203-3 having the largest y-axis coordinate among the feature points extracted from the eyeball. Hereinafter, the feature point 203-3 will be referred to as an upper eyelid feature point.
[0118] The feature point associated with the lower eyelid may be the feature point 203-4 with the smallest y-axis coordinate among the feature points extracted from the eyeball, which will be referred to as the lower eyelid feature point hereinafter.
[0119] The eyeball coordinate learning unit 120 may use a line segment connecting the upper eyelid feature point 203-3 and the lower eyelid feature point 203-4 as a vertical axis for the pupil center coordinate. The y coordinate of the upper eyelid feature point 203-3 and the y coordinate of the lower eyelid feature point 203-4 may correspond to the extreme values of the vertical axis. The difference between the y coordinate of the upper eyelid feature point 203-3 and the y coordinate of the lower eyelid feature point 203-4 may represent the entire length of the vertical axis. The eyeball coordinate learning unit 120 may calculate a normalized horizontal coordinate of the pupil center using the vertical coordinate value of the pupil center relative to the entire length of the vertical axis. This is merely an example and is not limited by the feature points.
[0120] 6, the feature point 203-1 for the inner canthus of the eye and the feature point 203-2 for the outer canthus of the eye are shown to be located on the same horizontal line, and the feature point 203-3 for the upper eyelid and the feature point 203-4 for the lower eyelid are shown to be located on the same vertical line, but this may change depending on the shooting angle of the camera, the angle of the person's head, etc.
[0121] For example, in an eyeball region image rotated 30° clockwise based on the eyeball region image 203, the feature point 203-1 for the inner canthus of the eye and the feature point 203-2 for the outer canthus of the eye may not be located on the same horizontal line, and the upper eyelid feature point 203-3 and the lower eyelid feature point 203-4 may not be located on the same vertical line. In this case, the eyeball coordinate learning unit 120 may transform the image so that the feature point 203-1 for the inner canthus of the eye and the feature point 203-2 for the outer canthus of the eye are located on the same horizontal line, and the upper eyelid feature point 203-3 and the lower eyelid feature point 203-4 are located on the same vertical line in the rotated eyeball region image. In this case, the eyeball coordinate learning unit 120 may transform the rotated eyeball region image using an affine transform, etc., which is an example and is not limited thereto.
[0122] The eyeball coordinate learning unit 120 can further use the generated normalized horizontal and vertical coordinate values of the pupil center as learning data to learn the second artificial neural network model.
[0123] In addition, the eyeball coordinate learning unit 120 may generate horizontal and vertical coordinate values of the pupil center using frame images of a plurality of videos captured by a plurality of cameras. The plurality of videos may be synchronized videos. The eyeball coordinate learning unit 120 may calculate average values of horizontal and vertical coordinate values of the pupil center calculated from a plurality of frame images of a plurality of videos in which synchronization is achieved. The eyeball coordinate learning unit 120 may further use the average values of the horizontal and vertical coordinate values of the pupil center as learning data to train the second artificial neural network model. As a result, the second artificial neural network model may generate more accurate information related to pupil center coordinates. The information related to the pupil center coordinates may include vertical and horizontal coordinate values of the pupil center according to each frame image, information on vertical and horizontal movement of the pupil center, etc. The information related to the pupil center coordinates may include information on two-dimensional coordinates.
[0124] The eyeball coordinate learning unit 120 can extract eyeball region images for the left eye and the right eye from the frame image and train the second artificial neural network model to generate information related to the pupil center coordinates of the left eye and information related to the pupil center coordinates of the right eye.
[0125] According to another embodiment of the present specification, the memory 100 may store data of at least one virtual object. The second artificial neural network model may be trained by the eyeball coordinate training unit 120 to generate information related to the coordinates of the pupil center using training data including parameter values obtained by changing at least one of parameters related to head rotation, eyeball rotation, and camera setting of virtual object parameters, and images of the virtual object acquired according to the parameter values.
[0126] FIG. 7 is a diagram illustrating an example of generating training data for a second artificial neural network model according to another embodiment of the present specification.
[0127] 7, the eyeball coordinate learning unit 120 may change at least one of parameters related to head rotation, eyeball rotation, and camera setting of a virtual object. For example, the eyeball coordinate learning unit 120 may set a parameter value so that a virtual camera photographs a virtual object from the front (upper drawing of FIG. 7). The eyeball coordinate learning unit 120 may change a parameter value so that a virtual camera photographs a virtual object from the right (lower drawing of FIG. 7). In addition, the eyeball coordinate learning unit 120 may change a parameter value related to the distance between the virtual camera and the virtual object.
[0128] The eyeball coordinate learning unit 120 may set a parameter value so that the head of the virtual object rotates in at least one direction of roll, pitch, and yaw.
[0129] The eyeball coordinate learning unit 120 may set a parameter value so that the eyeball of the virtual object rotates in at least one direction of roll, pitch, and yaw.
[0130] The eyeball coordinate learning unit 120 can change at least one of the parameters to acquire an image of the virtual object. The eyeball coordinate learning unit 120 can learn the second artificial neural network model using the parameter values and the virtual object according to the parameter values. In this case, the second artificial neural network model can be trained to generate information on two-dimensional coordinates and / or three-dimensional coordinates of the pupil center.
[0131] The virtual object may be a Gaussian avatar generated using 3D Gaussian splatter. The eyeball coordinate learning unit 120 may change the Euler coordinates of the head and pupil of the virtual object by controlling the latent vector of the virtual object. The Gaussian avatar is an example, and is not limited thereto. Any virtual object generated by a technique well known to those skilled in the art may be used.
[0132] The memory 100 may also store labeling data including at least one of head coordinates, pupil center coordinates, and camera setting information according to an image of a human face. The eyeball coordinate learning unit 120 may learn the second artificial neural network model using the labeling data.
[0133] In addition, the eyeball coordinate learning unit 120 can learn the second artificial neural network model using at least one of learning data using the eyeball area image, learning data acquired according to parameter changes of the virtual object, and labeling data.
[0134] According to an embodiment of the present specification, the third artificial neural network model may extract an eyeball region image, which is an image of a region including an eyeball extracted from each frame image of a video in which a human face is captured, by the eyeball rotation learning unit 130. The eyeball rotation learning unit 130 may train the third artificial neural network model to generate an eyeball rotation value using learning data including information related to eyeball phase changes generated according to the time sequence of the eyeball region image.
[0135] The eyeball rotation learning unit 130 may extract the eyeball region image from each frame image of the video. The eyeball rotation learning unit 130 may generate information related to eyeball phase change by comparing the eyeball region image extracted from an arbitrary frame image with the eyeball region image extracted from a frame image within a predetermined time range based on the corresponding frame image. As an example, the eyeball rotation learning unit 130 may generate information related to eyeball phase change by comparing the eyeball region image extracted from an arbitrary frame image with the eyeball region image extracted from a frame image acquired around 0.1 seconds based on the corresponding frame image. This is an example and is not limited by the time.
[0136] According to an embodiment of the present specification, the eyeball rotation learning unit 130 may extract an iris region image, which is an image of an area occupied by the iris, from the eyeball region image. The iris region image may refer to an image within a bounding box including the iris in the eyeball region image. The eyeball rotation learning unit 130 may train the third artificial neural network model using information related to a phase change of the iris according to time in the iris region image.
[0137] More specifically, the eyeball rotation learning unit 130 may compare an iris region image extracted from an arbitrary frame image with an iris region image extracted from a frame image within a predetermined time range based on the arbitrary frame image. For example, the iris region image may be a mask image in which an area occupied by the iris is divided into pixel values different from other areas, but this is just an example and is not limited thereto.
[0138] In addition, the eyeball rotation learning unit 130 can generate information related to a phase change using the iris mask image generated by the eyeball coordinate learning unit 120 .
[0139] The eyeball rotation learning unit 130 may generate information related to a phase change of the iris by comparing pixel values of an iris region image extracted from an arbitrary frame image with pixel values of another iris region image. The eyeball rotation learning unit 130 may generate a phase cross correlation value for pixel values of an iris region image extracted from an arbitrary frame image with pixel values of another iris region image using phase cross correlation analysis. The information related to the phase change calculated using the phase cross correlation analysis may include information related to an angle change of the iris.
[0140] The eyeball rotation learning unit 130 may generate the phase cross-correlation value by a method of obtaining a cross-correlation value upsampled by a fast Fourier transform (FFT). The eyeball rotation learning unit 130 may calculate an initial cross-correlation peak estimate using an FFT, and then precisely estimate a phase shift of the upsampled signal by a discrete Fourier transform (DFT) in a predetermined region based on the estimated value, thereby generating the phase cross-correlation value. This is just an example, and the present invention is not limited to this method.
[0141] The eyeball rotation learning unit 130 can use the image of the iris region and the phase cross correlation value as learning data to train the third artificial neural network model.
[0142] According to an embodiment of the present specification, the eyeball rotation learning unit 130 may calculate the size of an area occupied by the pupil in an eyeball region image extracted from a frame image of the mth video. The eyeball rotation learning unit 130 may adjust the size of a target eyeball region image according to a preset criterion. The target eyeball region image refers to an eyeball region image extracted from an arbitrary frame image, and is not a term referring to a specific eyeball region image.
[0143] The eyeball rotation learning unit 130 may compare the size of the pupil area calculated from a target eyeball region image extracted from an arbitrary frame image with the size of the pupil area calculated from a previous eyeball region image extracted from the immediately preceding frame image. The eyeball rotation learning unit 130 may adjust the size of the target eyeball region image so that the size of the pupil area extracted from the target eyeball region image has a value within a predetermined difference value from the size of the pupil area extracted from the previous target eyeball region image. The eyeball rotation learning unit 130 may calculate the phase cross-correlation value after adjusting the size of each eyeball region image.
[0144] In addition, the eyeball rotation learning unit 130 may generate a bounding box of an area including an eyeball in the frame image. The eyeball rotation learning unit 130 may extract a pupil center from within the bounding box. Alternatively, the eyeball rotation learning unit 130 may receive information related to the coordinates of the pupil center generated by the second artificial neural network model.
[0145] The eyeball rotation learning unit 130 may adjust the bounding box so that the pupil center is located at the center of the bounding box. The eyeball rotation learning unit 130 may extract an image inside the adjusted bounding box as an eyeball region image. After adjusting the size of the eyeball region image according to the above-described method, the eyeball rotation learning unit 130 may extract the iris region image and calculate the phase cross-correlation value.
[0146] The eyeball rotation learning unit 130 can train the third artificial neural network model to generate an eyeball rotation value using the iris region image and the phase cross-correlation value. The eyeball rotation value may represent an angle rotated clockwise or counterclockwise around the central axis of the eyeball.
[0147] The eyeball rotation learning unit 130 can extract eyeball region images for the left eye and the right eye from the frame image and train the third artificial neural network model to generate information related to phase changes of the left eye and the right eye.
[0148] According to an embodiment of the present specification, the eyeball coordinate learning unit 120 can train the second artificial neural network model using information generated by the first artificial neural network model, and the eyeball rotation learning unit 130 can train the third artificial neural network model using information generated by the second artificial neural network model.
[0149] For example, the eyeball coordinate learning unit 120 may generate learning data using a plurality of frame images input to the first artificial neural network model and information related to head coordinates extracted from each frame image. The eyeball coordinate learning unit 120 may generate an eyeball region image for each frame image using each frame image and information related to head coordinates extracted from each frame image. The eyeball coordinate learning unit 120 may generate an eyeball region image for each frame image using information related to eye coordinates in the information related to head coordinates. The eyeball coordinate learning unit 120 may train the second artificial neural network model according to the process described above.
[0150] The eyeball rotation learning unit 130 can further use information related to the coordinates of the pupil center according to each frame image generated by the second artificial neural network model to train the third artificial neural network model. Also, the eyeball rotation learning unit 130 can use the iris and / or pupil images segmented by the eyeball coordinate learning unit 120 to generate training data according to the above-described process to train the third artificial neural network model.
[0151] The first to third artificial neural network models can be trained independently of each other, or can be trained using information generated by each artificial neural network model.
[0152] The following describes the process in which the balance function management system 10 according to the first embodiment generates balance function status information and executes a balance function rehabilitation program using a trained artificial neural network model.
[0153] The balance function management system 10 can acquire n video frame images in real time from n cameras, which capture a subject undergoing a balance function status test and / or a balance function rehabilitation program.
[0154] When the subject is photographed with one camera, the camera may be set to a value equal to or greater than a preset FPS (Frames Per Second). For example, one camera may photograph the subject at 240 FPS, which is an example and is not limited by the above value.
[0155] When the subject is photographed using multiple cameras, the balance function management system 10 can adjust the synchronization of the multiple cameras through at least one processor. For example, the at least one processor can adjust the synchronization of the multiple cameras in real time using technology such as genlock, which is one example, and can adjust the synchronization of the multiple cameras using technology widely known to those skilled in the art.
[0156] Furthermore, the balance function management system 10 may sample frame images from multiple cameras through at least one processor. For example, one camera may capture the subject at 100 FPS, and another camera may capture the subject at 50 FPS. In this case, the balance function management system 10 may use at least one processor to downsample the video captured at 100 FPS by 1 / 2, or upsample the video captured at 50 FPS by 2. This is merely an example and is not limiting.
[0157] Preferably, the subject can be photographed by multiple cameras with the same frame rate.
[0158] The balance function management system 10 can acquire at least one of information related to the subject's head coordinates, pupil center coordinates, and eyeball phase changes by using at least one of the first to third artificial neural network models.
[0159] The balance function management system 10 can acquire information related to the head coordinates, pupil center coordinates and / or eyeball phase changes using the first to third artificial neural network models.
[0160] In addition, the balance function management system 10 can generate information related to the head coordinates, pupil center coordinates, and eyeball phase changes using an algorithm in which at least one processor generates learning data for the first to third artificial neural network models described above.
[0161] In the following description, the first to third artificial neural network models are used to generate information related to the head coordinates, pupil center coordinates, and eyeball phase changes, but it is not necessary to use an artificial neural network model to generate the information.
[0162] The head coordinate acquisition unit 140 executes the first artificial neural network model stored in the memory 100, and inputs frame images of the mth video into the first artificial neural network model to acquire information related to head coordinates according to the mth video. The information related to head coordinates according to the mth video may refer to information related to head coordinates generated according to the order of frame images of the mth video.
[0163] As an example, when the subject is photographed using two cameras, the head coordinate acquisition unit 140 may acquire information related to head coordinates according to the order of frame images of the first video and information related to head coordinates according to the order of frame images of the second video, which is an example and is not limited by the number of cameras and videos.
[0164] When the subject is photographed by multiple cameras, the head coordinate acquisition unit 140 can input the frame images of the mth video independently to the first artificial neural network model, or the head coordinate acquisition unit 140 can input the frame images of the mth video sequentially to the first artificial neural network model.
[0165] For example, when the subject is photographed with two cameras, the head coordinate acquisition unit 140 may input a frame image of a first video to the first artificial neural network model, and then the head coordinate acquisition unit 140 may input a frame image of a second video to the first artificial neural network model.
[0166] Alternatively, the head coordinate acquiring unit 140 may input frame images of the second video into the first artificial neural network model, and then input frame images of the first video into the first artificial neural network model.
[0167] The head coordinate acquisition unit 140 may input frame images of the mth video to the first artificial neural network model sequentially in a predetermined order. For example, the head coordinate acquisition unit 140 may input the first frame image of the first video to the first artificial neural network model and the first frame image of the second video to the first artificial neural network model. Thereafter, the head coordinate acquisition unit 140 may input the second frame image of the first video to the first artificial neural network model and the second frame image of the second video to the first artificial neural network model. This is an example and is not limited by the above order. In this case, the frame images input to the first artificial neural network model may include information of the extracted video.
[0168] The eyeball coordinate acquiring unit 150 executes the second artificial neural network model stored in the memory 100, and inputs information related to the head coordinates into the second artificial neural network model to acquire information related to pupil center coordinates according to the mth video. The information related to pupil center coordinates according to the mth video may mean information related to pupil center coordinates generated according to the order of frame images of the mth video.
[0169] As an example, when the subject is photographed using two cameras, the eyeball coordinate acquisition unit 150 may acquire information related to pupil center coordinates according to the order of frame images of a first video and information related to pupil center coordinates according to the order of frame images of a second video, which corresponds to one example and is not limited by the number of cameras and videos.
[0170] The phase change acquisition unit 160 executes the third artificial neural network model stored in the memory 100, and inputs information related to pupil center coordinates according to the time sequence of the frame images of the mth video into the third artificial neural network model, thereby acquiring information related to the phase change of the eyeball according to the mth video. The information related to the phase change of the eyeball according to the mth video may mean information related to the phase change of the eyeball generated according to the sequence of the frame images of the mth video.
[0171] As an example, when the subject is photographed using two cameras, the phase change acquisition unit 160 can acquire information related to the phase change of the eyeball according to the order of frame images of the first video and information related to the phase change of the eyeball according to the order of frame images of the second video, which is an example and is not limited by the number of cameras and videos.
[0172] FIG. 8 is a block diagram of a balance management system for generating balance status information according to one embodiment of the present disclosure.
[0173] Referring to FIG. 8, a balance function management system 10-1 for generating balance function status information according to one embodiment of the present specification may include a memory 100, a head coordinate learning unit 110, an eyeball coordinate learning unit 120, an eyeball rotation learning unit 130, a head coordinate acquisition unit 140, an eyeball coordinate acquisition unit 150, a phase change acquisition unit 160, a head movement generation unit 1100, an eyeball movement generation unit 1110, a velocity information generation unit 1120, and a balance function status information generation unit 1130.
[0174] The memory 100, head coordinate learning unit 110, eyeball coordinate learning unit 120, eyeball rotation learning unit 130, head coordinate acquiring unit 140, eyeball coordinate acquiring unit 150, and phase change acquiring unit 160 have been described above, so repeated description will be omitted.
[0175] The head movement generation unit 1100 may generate information related to head movement in the mth video using information related to head coordinates acquired from the first artificial neural network model. The information related to head movement may include horizontal and vertical head movement over time, and the degree of head rotation. The degree of head rotation may refer to the rotation angle in the roll, pitch, and yaw directions of the head. The information related to head movement may be expressed as a graph of horizontal and vertical head coordinate values over time.
[0176] According to an embodiment of the present specification, the head movement generation unit 1100 may calculate a normal vector of the subject's head using head feature points and their coordinates generated in each frame image. The direction of the normal vector may indicate a direction in which the subject's head is facing forward. The direction in which the subject's head is facing forward may indicate a direction in which the tip of the nose is facing. The head movement generation unit 1100 may calculate the head normal vector using the head feature points based on the tip of the nose feature point among the head feature points.
[0177] Alternatively, the direction of the subject's head may refer to a direction toward any feature point (such as the edge of the forehead or the center of the lips) that is on a vertical line based on the feature point of the nose. The head movement generation unit 1100 may calculate a normal vector of the head based on any one of the feature points.
[0178] Calculating a normal vector for the front of the head using feature points is a technique well known to those skilled in the art, and therefore a detailed description thereof will be omitted.
[0179] The head movement generation unit 1100 may generate information related to head movement over time using 3D head coordinate information and normal vectors according to frame images of the mth moving image, and may output the information related to head movement in the form of a graph.
[0180] The eye movement generation unit 1110 may generate information related to eye movement in the mth video using information related to pupil center coordinates and eye phase changes generated by the second and third artificial neural network models. The information related to eye movement may include information related to vertical and horizontal movements of the pupil center and eye rotation values over time. The information related to eye movement may be expressed as a graph of vertical coordinate values, horizontal coordinate values, and rotation angles of the pupil centers of the left and right eyes over time.
[0181] The eyeball movement generator 1110 may calculate a gaze vector of the eyeball using vertical coordinate values, horizontal coordinate values, and rotation values of the pupil center. Calculating a gaze vector using the coordinate values and rotation values of the pupil center is a technique well known to those skilled in the art, and therefore, detailed description thereof will be omitted.
[0182] The eye movement generating unit 1110 can output information related to the eye movement in the form of a graph.
[0183] According to one embodiment of the present specification, the head movement generation unit 1100 and the eye movement generation unit 1110 can correct errors between head movement and eye movement information based on learning data and the head movement and eye movement information of the subject.
[0184] The head movement and eye movement information based on the training data may represent actual data values for training the artificial neural network model.
[0185] As an example, the actual data values may refer to parameter values of the Euler angles of the head and eyes acquired from the eye coordinate learning unit 120. The eye coordinate learning unit 120 may set parameter values of a virtual camera to resemble actual settings in a balance function status test and / or a balance function rehabilitation program. The eye coordinate learning unit 120 may change parameter values of the Euler angles to resemble head and eye movements of a subject in a balance function status test and / or a balance function rehabilitation program. In this case, head and eye movement information of a virtual object resulting from changes in head and eye parameter values may refer to actual data values. This is an example, and the actual data values may refer to head and eye movement information based on actual data values that can be used to train the first to third artificial neural network models.
[0186] The head movement generation unit 1100 and the eye movement generation unit 1110 can correct errors in head movement and eye movement between frame images acquired within a preset time. For example, when the preset time is 1.5 seconds and a camera captures an image of a subject at 100 FPS, the head movement generation unit 1100 and the eye movement generation unit 1110 can correct errors in head movement and eye movement between 150 frame images. This is just an example and is not limited by the time and frame rate.
[0187] The head movement generator 1100 and the eye movement generator 1110 can correct the head movement and eye movement errors so that they have values within a preset range.
[0188] For example, the head movement generation unit 1100 and the eye movement generation unit 1110 may generate information related to head and eye movements using information related to head coordinates, pupil center coordinates, and / or eye phase changes acquired from 150 frame images (frame images acquired 1.5 seconds after an arbitrary frame image) in a video captured by a camera at 100 FPS of the subject. The information related to head and eye movements may be calculated as the amount of head and eye movements (vertical, horizontal, and rotational) over time according to the order of the frame images. In this case, the amount of head and eye movements at the time corresponding to the 50th frame image (frame image acquired 0.5 seconds after an arbitrary frame image) may deviate from a predetermined error range. In this case, the head movement generation unit 1100 and the eye movement generation unit 1110 may calculate statistics such as the average or median of the amount of head and eye movements generated using information from previous frame images, and replace the amount of movement at the time corresponding to the 50th frame image. Although it has been stated that errors are corrected between frame images acquired for 1.5 seconds based on an arbitrary frame image, this is merely an example, and various embodiments are possible, such as frame images acquired before 1.5 seconds based on the arbitrary frame image, frame images acquired around 1.5 seconds, etc. This is merely an example, and is not limited by the time, frame rate, statistical values, etc.
[0189] As another example, the head movement generation unit 1100 and the eye movement generation unit 1110 may apply a filter to correct errors in the head and eye movements. For example, the head movement generation unit 1100 and the eye movement generation unit 1110 may correct the errors using a filter such as a Chaining Kalman filter, a Moving Average filter, a Savitzky-Golay filter, a High Pass Filter, a Low Pass Filter, or a Band Pass Filter. This is just an example, and various types of filters may be used without being limited thereto.
[0190] The velocity information generator 1120 may generate information related to the velocity of head and eye movement in the mth moving image using information related to the head movement and eye movement. The velocity information generator 1120 may calculate a vertical movement velocity, a horizontal movement velocity, and / or a head rotation velocity over time using the information related to the head movement. The velocity information generator 1120 may calculate a vertical movement velocity, a horizontal movement velocity, and / or an eye rotation velocity over time using the information related to the eye movement.
[0191] According to one embodiment of the present disclosure, the velocity information generator 1120 may filter noise values from the information related to the velocity of head and eye movement. For example, when performing a balance function status test, noise values that prevent the velocity of head and eye movement from being accurately calculated may be generated if the subject's eyes are covered, if the subject rotates their head quickly, if the subject rotates their head slowly, or if the head position changes. The velocity information generator 1120 may remove noise from the information related to the velocity of head and eye movement using a filter such as a Chaining Kalman filter, a Moving Average filter, a Savitzky-Golay filter, a High Pass filter, a Low Pass filter, or a Band Pass filter. This is by way of example only, and various noise processing methods may be used.
[0192] According to one embodiment of the present disclosure, the speed information generator 1120 may generate information related to the speed of head movement and eye movement within a predetermined time based on the time point when the head movement exceeds a predetermined threshold. In a balance function status test, the subject may move their head horizontally (lateral left, lateral right). In this case, the speed information generator 1120 may generate information related to the speed of head and eye movement when the horizontal head movement exceeds a predetermined threshold.
[0193] In addition, the subject may move their head in a downward and upward right direction while turning the right side of their face so that the right anterior semicircular canal and the left posterior semicircular canal (Right Anterior, Left Posterior, RALP) are stimulated. In addition, the subject may move their head in a downward and upward left direction while turning the left side of their face so that the right anterior semicircular canal and the left posterior semicircular canal (Left Anterior, Right Posterior, LARP) are stimulated. In this case, the speed information generating unit 1120 may generate information related to the speed of head and eye movement when the vertical movement of the head is equal to or greater than a preset threshold value.
[0194] In the following, the direction in which the subject rotates their head to stimulate the RALP will be referred to as the RALP direction, and the direction in which the subject rotates their head to stimulate the LARP will be referred to as the LARP direction.
[0195] For example, the speed information generator 1120 may determine whether the head movement is equal to or greater than a threshold value using head movement information according to frame images existing within a preset time based on the last input frame image. The speed information generator 1120 may determine whether the head movement is equal to or greater than a threshold value by calculating the difference between the maximum and minimum values of head feature point coordinates in the head movement information according to frame images existing within a preset time. The preset threshold value may vary depending on the frame rate of a video, the size of a frame image, etc.
[0196] As another example, the memory 100 may further store a fourth artificial neural network model that generates information about head movement. The fourth artificial neural network model may be trained by at least one processor using frame images of a video for performing a balance function status test and data on head pitch and yaw values in the corresponding frame images. The fourth artificial neural network model may be a time series model or a Transformer model, which are just examples and are not limited to the model. In this case, the frame images may be labeled with information about the lateral, RALP, and LARP directions. The velocity information generator 1120 may input frame images present within a predetermined time based on the last acquired frame image to the fourth artificial neural network model to determine whether head movement is above a threshold value.
[0197] The speed information generator 1120 may calculate the speed of head and eye movement using information related to head and eye movement generated within a predetermined time after the head movement exceeds a critical value. As an example, the speed information generator 1120 may calculate the speed of head and eye movement using information related to head and eye movement generated within 1.5 seconds after the head movement exceeds a critical value, which is an example and is not limited by the time. The speed information generator 1120 may calculate the speed of head and eye movement in the mth moving image.
[0198] The velocity information generator 1120 may display information related to the velocity of the head and eye movement on a display device. The velocity information generator 1120 may display information related to the velocity of the head and eye movement from which noise has been removed and / or information related to the velocity of the head and eye movement from which noise has not been removed on a display.
[0199] The balance function status information generating unit 1130 may generate balance function status information of the subject using information related to the movement speed of the head and eyeballs.
[0200] According to an embodiment of the present specification, the balance function status information generator 1130 may calculate a gain coefficient using a time value (Head Peak Index) when the speed of the subject's head movement is relatively the fastest within a predetermined analysis window and a time value (Eye Peak Index) when the speed of the subject's eye movement is relatively the fastest when the subject's eyes move in the direction of the head movement and then return to their original positions. The analysis window may refer to a predetermined time range based on a point in time when the head movement exceeds a threshold value. The size of the analysis window may correspond to a time range in which the speed information generator 1120 generates information related to the speed of head and eye movement.
[0201] For example, the size of the analysis window may be 1.5 seconds after the head movement reaches a critical value, but this is just an example and is not limited thereto.
[0202] The balance function status information generator 1130 may calculate a gain coefficient using the Head Peak Index and Eye Peak Index, which are noise-removed information related to the movement speed of the head and eyeballs, using Equation 1.
number
[0203] Thereafter, the balance function status information generator 1130 can calculate a gain value using the gain coefficient.
number
[0204] The balance function status information generator 1130 can calculate gain values for the left and right eyes, respectively.
[0205] In addition, the balance function state information generating unit 1130 may calculate gain values for each of the mth videos. For example, if two cameras are used to capture images of the subject, gain values for the two videos may be calculated. In this case, gain values for the left eye and right eye in the first video and gain values for the left eye and right eye in the second video may be calculated. The balance function state information generating unit 1130 may calculate a statistical value for the gain values for the left eye calculated for the first and second videos, and may calculate at least one statistical value for the gain values for the right eye calculated for the first and second videos. The statistical value may correspond to an average value, a median value, a minimum value, a maximum value, a standard deviation, etc., but this is by way of example and is not limited thereto.
[0206] According to an embodiment of the present specification, when the subject is photographed by a plurality of cameras (n is 2 or more), the head movement generation unit 1100 may further generate reference head movement information by calculating statistics of information related to head coordinates according to the mth video. The eye movement generation unit 1110 may further generate reference eye movement information by calculating statistics of information related to pupil center coordinates and eye phase changes according to the mth video.
[0207] When the subject is photographed using multiple cameras, the 3D head coordinate values generated in the frame image of the mth video may differ depending on the camera position, angle, etc.
[0208] For example, a first camera may photograph the subject on the right side of the subject, and a second camera may photograph the subject on the left side of the subject. In this case, if the subject turns his / her head to the right, the coordinates of the feature points located on the right side of the subject's face in the frame image of the first video captured by the first camera may be calculated relatively more accurately than the coordinates of the feature points located on the left side. Also, if the subject turns his / her head to the left, the coordinates of the feature points located on the left side of the subject's face in the frame image of the second video captured by the second camera may be calculated relatively more accurately than the coordinates of the feature points located on the right side.
[0209] As another example, multiple cameras may be installed around the subject at 1 meter distances from the subject, at 15° intervals. In this case, the multiple cameras may capture the subject at eye level. In this case, the coordinates of feature points measured relatively more accurately for each camera may be generated depending on the direction of the subject's head movement. Thus, differences may occur between the 2D head coordinate information calculated depending on the positions, angles, etc. of the multiple cameras.
[0210] The head movement generation unit 1100 may generate reference head coordinate information by calculating an average value for each feature point in the 3D head coordinates generated in the frame images where the syncs are aligned in the mth moving image. The head movement generation unit 1100 may further generate information related to the reference head movement using the reference head coordinate information over time.
[0211] The eyeball movement generation unit 1110 may generate reference pupil center coordinates and information related to a phase change of the reference eyeball by calculating an average value of information related to pupil center coordinates and a phase change of the reference eyeball, which are generated in frame images whose syncs are aligned in the mth moving image. The eyeball movement generation unit 1110 may generate a reference gaze vector of the eyeball using the reference pupil center coordinates and information related to a phase change of the reference eyeball. The eyeball movement generation unit 1110 may further generate information related to the movement of the reference eyeball using the reference pupil center coordinates, information related to a phase change of the reference eyeball, and the reference gaze vector.
[0212] In this case, the velocity information generating unit 1120 may generate information related to the head movement velocity and eye movement velocity for the mth moving image within a predetermined time based on the point in time when the reference head movement becomes equal to or greater than a predetermined critical value.
[0213] The balance function status information generating unit 1130 may calculate gain values using a Head Peak Index and an Eye Peak Index for the mth moving image within a predetermined time based on the point at which the reference head movement exceeds a predetermined critical value.
[0214] FIG. 9 is an example of balance function status information output on a display according to one embodiment of the present specification.
[0215] 9, the head movement generator 1100, eye movement generator 1110, velocity information generator 1120, and balance function status information generator 1130 can output calculated information on a display screen. Videos captured by a plurality of cameras can be output on the display screen.
[0216] The velocity information generating unit 1120 may output head and eye movement velocity information for the mth moving image as a graph 205. The head and eye movement velocity graph 205 may display velocity information according to the number of balance function status tests superimposed. The time values at which peaks and / or valleys appear in the head and eye movement velocity graph 205 may correspond to the Head Peak Index and / or Eye Peak Index. In the graph, the vertical axis may represent velocity values, and the horizontal axis may represent time values. While FIG. 9 illustrates a graph for the horizontal velocity of the head and eye, this is merely an example, and graphs for velocity in the vertical and rotational directions may also be output depending on the type of test.
[0217] The head movement generation unit 1100 and the eye movement generation unit 1110 can output head and eye movement information in the form of graphs. The head movement generation unit 1100 and the eye movement generation unit 1110 can output a horizontal head and eye movement graph 206 and a vertical head and eye movement graph 207. In addition, they can further output movement graphs for the rotation direction of the head and / or eyes. In addition, the head and eye movement graphs can include movement information for the mth moving image, and can include the reference head movement and reference eye movement information.
[0218] The balance function status information generator 1130 may output information 208 related to the balance function status test for the subject to a display device. The information related to the balance function status test may include at least one of head rotation direction (lateral, RALP, LARP), gain value and standard deviation according to head rotation direction, number of balance function status tests according to head rotation direction, number of gain value calculation successes, and number of gain value calculation failures.
[0219] A failure in the calculation of the gain value may occur when the subject moves their head faster than the standard of the test. In this case, the difference between the Head Peak Index and the Eye Peak Index may exceed the middle of the analysis window size. In this case, the balance function status information generator 1130 may fail to calculate the gain value. The balance function status information generator 1130 may generate information on the number of successes and failures in the calculation of the gain value, thereby enabling the subject and / or the examiner to make a more accurate judgment.
[0220] In addition, the balance function status information generating unit 1130 may generate information regarding the presence or absence of an abnormality in the semicircular canals according to the gain value. For example, the balance function status information generating unit 1130 may generate information regarding an abnormality in the semicircular canals when the gain values of the left and right eyes are equal to or less than a preset value in the balance function status test.
[0221] In addition, the subject may rotate his / her head laterally during the balance function status test. In this case, if the difference in gain value between the left eye and the right eye when the subject rotates his / her head to the left and right is equal to or greater than a preset value, the balance function status information generator 1130 may generate information about abnormality of the semicircular canals.
[0222] In addition, in the balance function status test, the subject may rotate his / her head in the RALP or LARP direction. In this case, if the difference in gain value between the left eye and the right eye when the subject rotates his / her head upward and downward is equal to or greater than a preset value, the balance function status information generating unit 1130 may generate abnormal information about the RALP or LARP.
[0223] In addition, the balance function status information generating unit 1130 can further generate information regarding the presence or absence of abnormalities in central vestibular nerve function and peripheral vestibular nerve function using information related to the eye movement and gain value information.
[0224] FIG. 10 is a block diagram of a balance function management system for implementing a balance function rehabilitation program according to one embodiment of the present specification.
[0225] 10, a balance function management system 10-2 for performing a balance function rehabilitation program according to an embodiment of the present specification may include a memory 100, a head coordinate learning unit 110, an eye coordinate learning unit 120, an eye rotation learning unit 130, a head coordinate acquiring unit 140, an eye coordinate acquiring unit 150, a phase change acquiring unit 160, a head movement generating unit 1100, an eye movement generating unit 1110, a target output unit 1140, a head direction providing unit 1150, and a feedback providing unit 1160. The memory 100, the head coordinate learning unit 110, the eye coordinate learning unit 120, the eye rotation learning unit 130, the head coordinate acquiring unit 140, the eye coordinate acquiring unit 150, the phase change acquiring unit 160, the head movement generating unit 1100, and the eye movement generating unit 1110 have been described above, and therefore, a repeated description will not be provided.
[0226] FIG. 11 is a diagram illustrating an example of a scene in which a balance rehabilitation program according to an embodiment of the present specification is performed.
[0227] 11, the target output unit 1140 can output a virtual target 209 to a display device. The virtual target can be displayed at any position on the display device. In FIG. 11, the target is illustrated in the shape of a playing card, but this is merely an example and the target is not limited to this shape.
[0228] The head direction providing unit 1150 may provide the subject with information regarding the head rotation direction according to a balance rehabilitation protocol. For example, the head direction providing unit 1150 may provide information to the subject to rotate their head in a lateral direction.
[0229] Alternatively, the head direction providing unit 1150 may provide information to the subject to rotate their head in a RALP or LARP direction.
[0230] In this case, the head direction providing unit 1150 may provide information to the subject so that the subject rotates his / her head only up or down while the right side of his / her head faces forward. Alternatively, the head direction providing unit 1150 may provide information to the subject so that the subject rotates his / her head only up or down while the left side of his / her head faces forward.
[0231] In addition, the head direction providing unit 1150 may provide information to restore the state before the subject's head is rotated within a predetermined time after the subject has rotated their head. For example, the head direction providing unit 1150 may provide information to restore the state before the subject's head is rotated within one second.
[0232] The head direction providing unit 1150 can visually display the information on a display device, and can also output the information to an audio device.
[0233] The feedback providing unit 1160 may provide feedback according to the subject's head and eye movements.
[0234] A standard angle by which the subject should rotate his / her head may be predetermined according to the balance rehabilitation protocol. The feedback providing unit 1160 may compare the subject's head rotation angle generated by the head movement generating unit 1100 with the standard angle. The feedback providing unit 1160 may provide auditory feedback and / or visual feedback if the subject's head rotation angle satisfies the standard angle. Furthermore, the feedback providing unit 1160 may provide auditory feedback and / or visual feedback if the subject's head rotation angle does not satisfy the standard angle. The feedback providing unit 1160 may provide different feedback depending on whether the subject's head rotation angle satisfies the standard angle.
[0235] In addition, the eye movement generating unit 1110 may generate coordinate information of a gaze point to which the subject's gaze is directed on a display device using an eye gaze vector. The feedback providing unit 1160 may compare the coordinate information of the gaze point generated by the eye movement generating unit 1110 with coordinate information of the virtual target 209.
[0236] When the coordinates of the gaze point are located within the area of the virtual target 209, the feedback providing unit 1160 may change the hue value of the virtual target 209. In addition, the feedback providing unit 1160 may further indicate the gaze point within the virtual target 209.
[0237] When the coordinates of the gaze point are located outside the area of the virtual target 209, the feedback providing unit 1160 can display the position of the gaze point on a display device.
[0238] When the subject is photographed by a plurality of cameras, the head movement generation unit 1100 and the eye movement generation unit 1110 may generate information related to the reference head movement and the reference eye movement as described above, and the feedback providing unit 1160 may provide feedback according to the information related to the reference head movement and the reference eye movement.
[0239] FIG. 12 is a block diagram of a balance function management system that generates balance function status information and executes a balance function rehabilitation program according to one embodiment of the present specification.
[0240] 12, the balance function management system 10-3 may include a memory 100, a head coordinate learning unit 110, an eye coordinate learning unit 120, an eye rotation learning unit 130, a head coordinate acquiring unit 140, an eye coordinate acquiring unit 150, a phase change acquiring unit 160, a head movement generating unit 1100, an eye movement generating unit 1110, a velocity information generating unit 1120, a balance function status information generating unit 1130, a target output unit 1140, a head direction providing unit 1150, and a feedback providing unit 1160. The balance function management system 10-3 may generate a gain value and provide a balance function rehabilitation program depending on whether or not there is a balance function abnormality.
[0241] A balance function management system according to a second embodiment of the present specification will be described below. According to the second embodiment of the present specification, the balance function management system generates head and eye movement information of a subject using multiple videos captured by multiple cameras, and can generate balance function status information and / or implement a balance function rehabilitation program. In the following, n can represent a natural number of 2 or greater.
[0242] FIG. 13 is a block diagram of a balance function management system according to the second embodiment of the present specification.
[0243] Referring to FIG. 13, the balance function management system 10′ according to the second embodiment of the present specification may include a memory 100′, a head coordinate learning unit 110′, an eyeball coordinate learning unit 120′, an eyeball rotation learning unit 130′, a head coordinate acquisition unit 140′, an eyeball coordinate acquisition unit 150′, and a phase change acquisition unit 160′.
[0244] The memory 100' may store at least one of a first artificial neural network model that generates information related to head coordinates, a second artificial neural network model that acquires information related to pupil center coordinates, and a third artificial neural network model that acquires information related to eye phase changes.
[0245] According to an embodiment of the present specification, the first artificial neural network model may be trained by the head coordinate training unit 110′ using training data including facial feature points and the coordinates of the feature points extracted from a plurality of frame images of a video in which a human face is captured.
[0246] The head coordinate learning unit 110' can learn the first artificial neural network model using at least a plurality of videos in which a person's face is captured, and the plurality of videos in which a person's face is captured may be videos in which a person is filmed turning their head.
[0247] For example, the video of the person's face may be captured by a plurality of cameras. The plurality of cameras may be installed at various positions and angles to capture the person. The video of the person's face may be captured by cameras installed at various positions and angles in front of the person, such as the left side, right side, upper left side, or upper right side. The video captured by a plurality of cameras may refer to a video captured by a plurality of cameras of the same situation.
[0248] The video of the person's face may be a video of a scene in which the person's face is facing forward and moving their head from side to side. Also, the video of the person's face may be a video of a scene in which the person's face is facing forward and moving their head up and down. The state in which the person's face is facing forward may be a state in which the person's line of sight and the direction of the front of their body are aligned.
[0249] The video of the person's face may be a video of the person turning their head to the right and then moving their head up and down. The video of the person's face may be a video of the person turning their head to the left and then moving their head up and down.
[0250] Furthermore, the video in which the person's face is captured may be a video in which the person is filmed turning his or her head while gazing at a specific gaze point.
[0251] The videos of the human faces may all be captured by cameras with the same specifications. Alternatively, the videos of the human faces may be captured by cameras with different specifications. Alternatively, the videos of the human faces may be captured by cameras with different settings.
[0252] The video of a person described above is merely an example, and any video of a person's head and eyeballs or any video of a person's head may be used. Furthermore, the video is not limited by specific camera specifications or settings.
[0253] Preferably, the video of the person may be video taken with a camera having the same settings.
[0254] According to one embodiment of the present disclosure, when extracting feature points from a plurality of videos captured using a plurality of cameras, the head coordinate learning unit 110' may adjust the sync of the plurality of videos and then extract the feature points. The head coordinate learning unit 110' may adjust the sync of the plurality of videos using an algorithm such as a specific audio signal of the plurality of videos or feature point matching-based synchronization. This is merely an example, and various techniques widely known to those skilled in the art may be used without being limited thereto.
[0255] FIG. 14 is a reference diagram illustrating an example in which the head coordinate learning unit according to an embodiment of the present specification connects frame images.
[0256] Referring to FIG. 14, the video in which the person's face is captured may be a video captured through two cameras. The head coordinate learning unit 110' may adjust the synchronization of a first video 300 captured by a first camera and a second video 301 captured by a second camera. The head coordinate learning unit 110' may generate a multiple frame image 310 by concatenating frame images having the same synchronization between the first video 300 and the second video 301. The head coordinate learning unit 110' may generate a multiple frame image 310 by concatenating frame images having the same synchronization. For example, the first frame image 300-1 of the first video 300 and the first frame image 301-1 of the second video 301 may be concatenated to generate a first multiple frame image 310-1. The multiple frame image may refer to a single image generated by concatenating a plurality of frame images having the same synchronization. The multiple frame image may refer to a single image in which the plurality of frame images are arranged.
[0257] In the following description, it is assumed that a multiple frame image is formed by concatenating frame images of an m-th moving image having the same sync among frame images of n moving images.
[0258] The head coordinate learning unit 110' may generate a multiple frame image by connecting frame images of an mth video having the same sink among n videos captured by n cameras. The head coordinate learning unit 110' may connect frame images of the mth video according to a preset criterion. In addition, the head coordinate learning unit 110' may label information about the video from which the frame images are extracted.
[0259] The head coordinate learning unit 110' may extract feature points on a person's face from the multiple frame images. For example, the head coordinate learning unit 110' may extract feature points located on a person's face from each of the multiple frame images. For example, the feature points may include feature points on the tip of the person's nose, the left outer corner of the eye (where the outermost eyelid meets the left eye), the right outer corner of the eye (where the outermost partial eyelid meets the right eye), and the forehead. The feature points may not change position even when the person blinks. The feature points are merely examples and are not limited to these feature points and may further include feature points used for conventional face recognition. Techniques for generating feature points and feature point coordinates on a person's head are widely known to those skilled in the art, and therefore, detailed description thereof will be omitted.
[0260] The head coordinate learning unit 110' may generate coordinate information of feature points extracted from each frame image. The head coordinate learning unit 110' may train a first artificial neural network model to generate 3D head coordinates for each multi-frame image using learning data including feature points extracted from each multi-frame image and their coordinates. The 3D head coordinates may refer to the 3D coordinates of the head according to a 3D standard head model. The 3D head coordinates may include coordinates for all feature points of the head. The 3D head coordinates may include information related to index numbers for each feature point. In this case, the head coordinate learning unit 110' may train the first artificial neural network model to generate an index number based on the feature point of the nose.
[0261] The head coordinate learning unit 110' can train the first artificial neural network model to generate 3D coordinates of the head for a frame image of the mth video in the multi-frame image. For example, when a multi-frame image is formed by connecting frame images of two videos, the head coordinate learning unit 110' can train the first artificial neural network model to generate 3D coordinates of the head for a frame image of the first video and 3D coordinates of the head for a frame image of the second video. This is just an example and is not limited by the number of videos.
[0262] In addition, the head coordinate learning unit 110' can train the first artificial neural network model to generate reference head coordinates for the 3D coordinates of the head. The reference head coordinates may mean average values of coordinates of each feature point in the 3D coordinates of the head generated in the frame image of the mth video. The head coordinate learning unit 110' calculates the average values and trains the first artificial neural network model to generate information related to the reference head coordinates in which differences in head coordinates are corrected according to the position and angle of the camera.
[0263] Also, the head coordinate learning unit 110' can learn the first artificial neural network model to further generate head movement information using three-dimensional coordinates extracted from each multi-frame image.
[0264] In addition, the head coordinate learning unit 110' may further use data on a 3D standard head model such as a frame image from which feature points have been extracted, a 3D Morphable Model (3DMM), or a FLAME (Faces Learned with an Articulated Model and Expressions) model as learning data to train the first artificial neural network model. This is just an example, and various additional learning data may be used without being limited by the learning data.
[0265] According to an embodiment of the present specification, the head coordinate learning unit 110' can use feature points having coordinate values according to a preset criterion as learning data. The head coordinate learning unit 110' can select feature points having coordinate values according to a preset criterion as learning data.
[0266] FIG. 15 is an illustration of a multi-frame image of a subject undergoing a balance status test and / or balance rehabilitation program according to one embodiment of the present disclosure.
[0267] 15, the plurality of videos of a person may be a plurality of videos of a balance function status test and / or a balance function rehabilitation program being performed using the balance function management system 10'. The videos may be videos of only the subject 400. Alternatively, the videos may be videos of the subject 400 and the examiner 401. Hereinafter, the videos of a person will be referred to as videos of a balance function status test and / or a balance function rehabilitation program being performed. However, this is merely an example and is not limited to the videos.
[0268] When a balance function status test is performed, the subject 400 may be sitting in a chair, and the examiner 401 may be standing behind the subject. The examiner 401 may be a medical professional. The balance function management system 10' may analyze the eye movements caused by the head movements of the subject 400 to determine whether the subject has a balance function abnormality, and may analyze only the head movements and eye movements of the subject 400 to proceed with a rehabilitation program.
[0269] The head coordinate learning unit 110' can extract facial feature points and coordinates of the feature points of the subject 400 from the multiple frame images to train the first artificial neural network model. The head coordinate learning unit 110' can extract facial feature points and coordinates of the feature points of the subject 400 from frame images of the m-th moving image in the multiple frame images to train the first artificial neural network model.
[0270] As described above, since the subject 400 is sitting in a chair and the examiner 401 is standing, the feature points extracted from the face of the subject 400 may be located relatively lower than the feature points extracted from the face of the examiner 401. This may mean that the y coordinate values of the feature points extracted from the face of the subject 400 are relatively smaller than the y coordinate values of the feature points extracted from the face of the examiner 401.
[0271] The head coordinate learning unit 110' can extract, as learning data, feature points with relatively small y-coordinate values from feature points at the same parts extracted from the faces of the subject 400 and the examiner 401. The head coordinate learning unit 110' can extract, as learning data, feature points located in an area below a preset y-coordinate value.
[0272] In addition, the head coordinate learning unit 110′ may cluster feature points extracted from the faces of the subject 400 and the examiner 401. Thereafter, the head coordinate learning unit 110′ may extract a set of feature points located relatively lower as learning data.
[0273] FIG. 16 is an illustration of a multi-frame image of a subject undergoing a balance status test and / or balance rehabilitation program according to one embodiment of the present disclosure.
[0274] 16, a balance function status test and / or balance function rehabilitation program may be performed in the moving image while the subject 400' and the examiner 401' are both standing. In this case, feature points extracted from the face of the subject 400' may be closer to the center of the display screen than feature points extracted from the face of the examiner 401'. The head coordinate learning unit 110' may extract feature points having coordinate values relatively closer to the coordinate values of the center of the display screen as learning data.
[0275] As another example, the subject may be closer to the camera than the examiner. As a result, the subject's face may occupy a larger area in the video than the examiner's face. The head coordinate learning unit 110′ may extract, as learning data, a set of feature points that are distributed over a larger area from the set of feature points extracted from the subject's face and the set of feature points extracted from the examiner's face.
[0276] The head coordinate learning unit 110' may generate feature points and coordinate information of the feature points on the face of the subject included in the frame image of the mth moving image in the multiple frame image.
[0277] The above-described process in which the head coordinate learning unit 110' extracts only feature points extracted from the subject's face as learning data is merely an example, and is not limited thereto. Various criteria can be set depending on the subject's position, the examiner's position, the camera position, angle, etc. Also, only feature points extracted from the subject's face can be extracted as learning data using markers for distinguishing the subject from the examiner, as well as the positions of the subject's and examiner's faces. Therefore, various embodiments can be realized according to various situations in which a balance function status test and / or rehabilitation test program is performed.
[0278] If the subject rapidly turns his / her head in the video, the subject's face may not be clearly captured in at least one frame image of the mth video. In this case, the head coordinate learning unit 110' may not be able to extract feature points from the subject's face. In this case, the head coordinate learning unit 110' trains the first artificial neural network model using feature points extracted from the examiner's face, and the first artificial neural network model may calculate inaccurate results.
[0279] According to one embodiment of the present specification, the head coordinate learning unit 110' may track coordinates of feature points extracted from a frame image preceding an arbitrary frame image of the mth video. As described above, if feature points cannot be extracted from the subject's face in the frame image of the mth video, the head coordinate learning unit 110' may extract learning data using the subject's image in the preceding frame image. The preceding frame image may refer to the closest frame image preceding the arbitrary frame image from which feature points can be extracted from the subject's face.
[0280] According to one embodiment of the present specification, the second artificial neural network model may be trained by the eyeball coordinate training unit 120′ to generate information related to the coordinates of the pupil center using training data including eyeball area images, which are images of areas including eyeballs in multiple frame images of a video in which a person's face is captured.
[0281] The eye coordinate learning unit 120' can generate a multi-frame image by adjusting the sync of a plurality of moving images, or can receive the multi-frame image generated by the head coordinate learning unit 110'.
[0282] The eyeball coordinate learning unit 120' can extract an image of an eyeball region of the subject, which is an image of a region including the subject's eyeball, from multiple frame images. The eyeball coordinate learning unit 120' can use the image of the eyeball region of the subject extracted from each frame image as learning data, and can train the second artificial neural network model to generate information related to the coordinates of the pupil center.
[0283] FIG. 17 is a diagram illustrating an example of pre-processing an eye area image according to an embodiment of the present specification.
[0284] 17, the eyeball coordinate learning unit 120' can extract eyeball region images 403-1 and 403-2 according to frame images of the mth moving image from each multi-frame image 402. The eyeball coordinate learning unit 120' can extract images within a bounding box of the eyeball region from each multi-frame image 402 as eyeball region images 403-1 and 403-2. The eyeball coordinate learning unit 120' can segment iris and pupil regions from the eyeball region images 403-1 and 403-2. Techniques for extracting eyeball regions from a human face and segmenting iris and pupil regions are well known to those skilled in the art, and therefore will not be described in detail.
[0285] The eyeball coordinate learning unit 120' may estimate a region where the iris and / or pupil is hidden by the eyelid. As shown in FIG. 17, a portion of the iris may be hidden by the upper and lower eyelids. The eyeball coordinate learning unit 120' may estimate the hidden portion using an ellipse fitting algorithm, a circle Hough transform algorithm, or the like. Alternatively, the eyeball coordinate learning unit 120' may segment the iris and pupil regions using an artificial neural network model that has been trained in advance to segment the iris and pupil regions. This is merely an example, and the present invention is not limited to this method.
[0286] The eyeball coordinate learning unit 120' can learn the second artificial neural network model using data obtained by dividing the iris and / or pupil regions in the eyeball region images 403-1 and 403-2. As an example, the eyeball coordinate learning unit 120' can divide the iris and / or pupil regions in the eyeball region images 403-1 and 403-2 to generate mask images 404-1 and 404-2. The eyeball coordinate learning unit 120' can generate a mask image in which the regions occupied by the iris and / or pupil and the remaining regions in the eyeball region images 403-1 and 403-2 have different pixel values. In the mask image, the regions occupied by the iris and / or pupil may be displayed in white or black, and the remaining regions may be displayed in black or white. This is an example and is not limited thereto.
[0287] The eyeball coordinate training unit 120' can train the second artificial neural network model to generate information related to the coordinates of the pupil centers for each frame image of a moving image using the mask images 404-1 and 404-2 as training data. Alternatively, the eyeball coordinate training unit 120' can train the second artificial neural network model using two-dimensional pixel values of the mask images 404-1 and 404-2 as training data. Alternatively, the eyeball coordinate training unit 120' can train the second artificial neural network model using the mask images 404-1 and 404-2 and two-dimensional pixel values as training data.
[0288] As another example, the eyeball coordinate learning unit 120' may train the second artificial neural network model using a heatmap model that divides and shows the iris and / or pupil regions in the eyeball region images 403-1 and 403-2.
[0289] The eyeball coordinate learning unit 120' can learn the second artificial neural network model using at least one of the mask image and the heat map model.
[0290] Although FIG. 17 illustrates a pre-processing process using a multi-frame image in which frame images of two moving images are concatenated, this is merely an example and is not limited by the number of moving images.
[0291] According to an embodiment of the present specification, the eyeball coordinate learning unit 120' can generate eyeball feature points and feature point coordinate information in the eyeball region images 403-1 and 403-2, and can train the second artificial neural network model to generate horizontal and vertical coordinate values of the pupil center for each video frame image using the coordinates of a plurality of preset feature points. The eyeball coordinate learning unit 120' can extract normalized pupil center coordinates using the coordinates of a plurality of feature points extracted from the eyeball region images 403-1 and 403-2.
[0292] As an example, the plurality of feature point coordinates may include a feature point with a relatively smallest x-axis coordinate value and a feature point with a relatively largest x-axis coordinate value among feature points whose y-axis coordinates are within a preset range in the eyeball region image 403-1.
[0293] The plurality of feature point coordinates may include a feature point with a relatively smallest y-axis coordinate value and a feature point with a relatively largest y-axis coordinate value among feature points whose x-axis coordinates are within a preset range in the eyeball region image 403-1 captured from the front of the person.
[0294] For example, the eyeball coordinate learning unit 120′ may extract the horizontal coordinate of the normalized pupil center using the feature point 403-10 for the inner canthus of the eye and the feature point 403-11 for the outer canthus of the eye among the feature points extracted from the eyeball region image 403-1. The eyeball coordinate learning unit 120′ may use a line segment connecting the feature point 403-10 for the inner canthus of the eye and the feature point 403-11 for the outer canthus of the eye as a horizontal axis for the pupil center coordinate. The x-coordinate of the feature point 403-10 for the inner canthus of the eye and the x-coordinate of the feature point 403-11 for the outer canthus of the eye may correspond to the extreme values of the horizontal axis. The difference between the x-coordinate of the feature point 403-10 for the inner canthus of the eye and the x-coordinate of the feature point 403-11 for the outer canthus of the eye may represent the length of the entire horizontal axis. The eyeball coordinate learning unit 120' can calculate the normalized horizontal coordinate of the pupil center using the horizontal coordinate value of the pupil center relative to the entire length of the horizontal axis.
[0295] In addition, the eyeball coordinate learning unit 120' may extract the vertical coordinate of the normalized pupil center using feature points associated with the upper eyelid and lower eyelid among the feature points extracted from the eyeball region image 403-1. In this case, the feature point associated with the upper eyelid may be the feature point 403-12 having the largest y-axis coordinate among the feature points extracted from the eyeball. Hereinafter, the feature point 403-12 will be referred to as the upper eyelid feature point.
[0296] The feature point associated with the lower eyelid may be the feature point 403-13 having the smallest y-axis coordinate among the feature points extracted from the eyeball, which will be referred to as the lower eyelid feature point hereinafter.
[0297] The eyeball coordinate learning unit 120' may use a line segment connecting the upper eyelid feature point 403-12 and the lower eyelid feature point 403-13 as a vertical axis for the pupil center coordinate. The y coordinate of the upper eyelid feature point 403-12 and the y coordinate of the lower eyelid feature point 403-13 may correspond to the extreme values of the vertical axis. The difference between the y coordinate of the upper eyelid feature point 403-12 and the y coordinate of the lower eyelid feature point 403-13 may represent the entire length of the vertical axis. The eyeball coordinate learning unit 120' may calculate a normalized horizontal coordinate of the pupil center using the vertical coordinate value of the pupil center relative to the entire length of the vertical axis. This is merely an example and is not limited by the feature points.
[0298] 17, the feature point 403-10 for the inner canthus of the eye and the feature point 403-11 for the outer canthus of the eye are shown to be located on the same horizontal line, and the feature point 403-12 for the upper eyelid and the feature point 403-13 for the lower eyelid are shown to be located on the same vertical line, but this may change depending on the shooting angle of the camera, the angle of the person's head, etc.
[0299] For example, in an eyeball region image rotated 30° clockwise from the eyeball region image 403-10, the feature point 403-10 for the inner canthus of the eye and the feature point 403-11 for the outer canthus of the eye may not be located on the same horizontal line, and the upper eyelid feature point 403-12 and the lower eyelid feature point 403-13 may not be located on the same vertical line. In this case, the eyeball coordinate learning unit 120' may transform the image so that the feature point 403-10 for the inner canthus of the eye and the feature point 403-11 for the outer canthus of the eye are located on the same horizontal line, and the upper eyelid feature point 403-12 and the lower eyelid feature point 403-13 are located on the same vertical line in the rotated eyeball region image. In this case, the eyeball coordinate learning unit 120' may transform the rotated eyeball region image using an affine transform, etc., which is an example and is not limited thereto.
[0300] The eyeball coordinate learning unit 120' can extract eyeball feature points and their coordinates from each of the eyeball region images 403-1 and 403-2, and generate coordinate information of a normalized pupil center for each video frame image. The eyeball coordinate learning unit 120' can further use the generated horizontal and vertical coordinate values of the normalized pupil center for each video frame image as learning data to learn the second artificial neural network model.
[0301] In addition, the eyeball coordinate learning unit 120' may calculate average values of horizontal and vertical coordinate values of the pupil center calculated from multiple frame images. The eyeball coordinate learning unit 120' may further use the average values of the horizontal and vertical coordinate values of the pupil center as learning data to learn the second artificial neural network model. The eyeball coordinate learning unit 120' may use the average values to learn the second artificial neural network model to generate information related to reference pupil center coordinates in which differences in pupil center coordinates are corrected depending on the position and angle of the camera.
[0302] The information related to the pupil center coordinates may include a vertical coordinate value, a horizontal coordinate value, vertical movement information of the pupil center according to the frame image of the mth moving image, horizontal movement information of the pupil center, etc. The information related to the pupil center coordinates may include content regarding two-dimensional coordinates.
[0303] The eyeball coordinate learning unit 120' can extract eyeball area images for the left eye and the right eye from the multiple frame images and train the second artificial neural network model to generate information related to the pupil center coordinates of the left eye and information related to the pupil center coordinates of the right eye in the frame image of the mth video.
[0304] According to another embodiment of the present disclosure, the memory 100' may store data of at least one virtual object. The second artificial neural network model may be trained by the eyeball coordinate training unit 120' to generate information related to the coordinates of the pupil center using training data including parameter values obtained by changing at least one of parameters related to head rotation, eyeball rotation, and camera settings of virtual object parameters, and images of the virtual object acquired according to the parameter values.
[0305] 7, the eyeball coordinate learning unit 120' can change at least one of parameters related to head rotation, eyeball rotation, and camera settings of a virtual object. For example, the eyeball coordinate learning unit 120' can set parameter values so that a virtual camera photographs a virtual object from the front (upper diagram of FIG. 7). The eyeball coordinate learning unit 120' can change parameter values so that a virtual camera photographs a virtual object from the right (lower diagram of FIG. 7). In addition, the eyeball coordinate learning unit 120' can change a parameter value related to the distance between the virtual camera and the virtual object.
[0306] The eyeball coordinate learning unit 120' may set parameter values so that the head of the virtual object rotates in at least one direction of roll, pitch, and yaw.
[0307] The eyeball coordinate learning unit 120' may set a parameter value so that the eyeball of the virtual object rotates in at least one direction of roll, pitch, and yaw.
[0308] The eyeball coordinate learning unit 120' can change at least one of the parameters to acquire an image of the virtual object. The eyeball coordinate learning unit 120' can learn the second artificial neural network model using the parameter values and the virtual object according to the parameter values. In this case, the second artificial neural network model can be trained to generate information on two-dimensional coordinates and / or three-dimensional coordinates of the pupil center.
[0309] The virtual object may be a Gaussian avatar generated using 3D Gaussian splatter. The eyeball coordinate learning unit 120' may change the Euler coordinates of the head and pupil of the virtual object by controlling the latent vector of the virtual object. The Gaussian avatar is an example, and is not limited thereto. Any virtual object generated by a technique well known to those skilled in the art may be used.
[0310] The memory 100' may also store labeling data including at least one of head coordinates, pupil center coordinates, and camera setting information according to an image of a human face. The eyeball coordinate learning unit 120' may use the labeling data to learn the second artificial neural network model.
[0311] In addition, the eyeball coordinate learning unit 120' can learn the second artificial neural network model using at least one of learning data using the eyeball area image, learning data acquired according to parameter changes of the virtual object, and labeling data.
[0312] According to an embodiment of the present specification, the third artificial neural network model may extract an eyeball region image, which is an image of a region including an eyeball, from each multi-frame image of a plurality of videos in which a human face is captured, by the eyeball rotation learning unit 130'. The eyeball rotation learning unit 130' may train the third artificial neural network model to generate an eyeball rotation value using learning data including information related to eyeball phase changes generated according to the time sequence of eyeball region images from each video.
[0313] The eyeball rotation learning unit 130' may generate the multi-frame image by adjusting the sync of a plurality of moving images, or may receive the multi-frame image generated by the head coordinate learning unit 110'.
[0314] The eyeball rotation learning unit 130' may extract an eyeball region image for a frame image of the mth video from the multi-frame images. The eyeball rotation learning unit 130' may compare the eyeball region image extracted from the multi-frame images with an eyeball region image extracted from a multi-frame image within a predetermined time range based on the corresponding multi-frame image to generate information related to an eyeball phase change. As an example, the eyeball rotation learning unit 130' may compare an eyeball region image extracted from an arbitrary multi-frame image with an eyeball region image extracted from a multi-frame image acquired around 0.1 seconds based on the corresponding multi-frame image to generate information related to an eyeball phase change. In this case, the eyeball rotation learning unit 130' may generate information related to an eyeball phase change due to the mth video using the eyeball region image for the mth video. This is an example and is not limited by the time.
[0315] According to an embodiment of the present specification, the eyeball rotation learning unit 130' may extract an iris region image, which is an image of an area occupied by the iris, from the eyeball region image. The iris region image may refer to an iris region image in a frame image of the mth video. The iris region image may refer to an image within a bounding box including the iris in the eyeball region image. The eyeball rotation learning unit 130' may train the third artificial neural network model using information related to a phase change of the iris according to time sequence in the iris region image.
[0316] More specifically, the eyeball rotation learning unit 130' may compare an iris region image extracted from an arbitrary multi-frame image with an iris region image extracted from a multi-frame image within a predetermined time range based on the arbitrary multi-frame image. The iris region image may be a mask image in which an area occupied by the iris is divided into pixel values different from other areas, but this is just an example and is not limited thereto.
[0317] In addition, the eyeball rotation learning unit 130' can generate information related to phase change using the iris mask image generated by the eyeball coordinate learning unit 120'.
[0318] The eyeball rotation learning unit 130' may generate information related to a phase change of the iris by comparing pixel values of an iris region image extracted from any multi-frame image with other iris region images. The eyeball rotation learning unit 130' may generate a phase cross correlation value for pixel values of an iris region image extracted from any multi-frame image with other iris region images using phase cross correlation analysis. The information related to the phase change calculated using the phase cross correlation analysis may include information related to an angle change of the iris.
[0319] The eyeball rotation learning unit 130' may generate the phase cross-correlation value by a method of obtaining a cross-correlation value upsampled by a fast Fourier transform (FFT). The eyeball rotation learning unit 130' may calculate an initial cross-correlation peak estimate using an FFT, and then precisely estimate a phase shift of the upsampled signal by a discrete Fourier transform (DFT) in a predetermined region based on the estimate, thereby generating the phase cross-correlation value. This is merely an example, and the present invention is not limited to this method.
[0320] The eyeball rotation learning unit 130' can use the image of the iris region and the phase cross-correlation value as learning data to train the third artificial neural network model.
[0321] According to an embodiment of the present specification, the eyeball rotation learning unit 130' may calculate the size of an area occupied by the pupil in the eyeball region image. The eyeball rotation learning unit 130' may adjust the size of the target eyeball region image according to a preset criterion. The target eyeball region image refers to an eyeball region image extracted from any multi-frame image, and is not a term referring to a specific eyeball region image.
[0322] The eyeball rotation learning unit 130' may compare the size of the pupil area calculated from a target eyeball region image extracted from any multi-frame image with the size of the pupil area calculated from a previous eyeball region image extracted from the immediately preceding multi-frame image. The eyeball rotation learning unit 130' may adjust the size of the target eyeball region image so that the size of the pupil area extracted from the target eyeball region image has a value within a predetermined difference value from the size of the pupil area extracted from the previous target eyeball region image. In this case, the eyeball rotation learning unit 130' may adjust the size of the target eyeball region image extracted from the mth moving image by comparing the size of the pupil area extracted from an eyeball region image corresponding to the mth moving image. The eyeball rotation learning unit 130' may calculate the phase cross-correlation value after adjusting the size of each eyeball region image.
[0323] In addition, the eyeball rotation learning unit 130' may generate a bounding box of an area including an eyeball in the multi-frame image. The eyeball rotation learning unit 130' may extract a pupil center from within the bounding box. The eyeball rotation learning unit 130' may adjust the bounding box so that the pupil center is located at the center of the bounding box. Alternatively, the eyeball rotation learning unit 130' may receive information related to the coordinates of the pupil center generated by the second artificial neural network model.
[0324] The eyeball rotation learning unit 130' may extract an image inside the adjusted bounding box as an eyeball region image. After adjusting the size of the eyeball region image according to the above-described method, the eyeball rotation learning unit 130' may extract the iris region image and calculate the phase cross-correlation value.
[0325] The eyeball rotation learning unit 130' may train the third artificial neural network model to generate an eyeball rotation value for the mth video using the iris region image and the phase cross-correlation value. The eyeball rotation value may mean an angle rotated clockwise or counterclockwise around the central axis of the eyeball. In addition, the eyeball rotation learning unit 130' may calculate an average value of eyeball rotation values calculated from frame images over time in the mth video. The eyeball rotation learning unit 130' may train the third artificial neural network model to generate a reference eyeball rotation value using the average value. The reference eyeball rotation value may mean a rotation value obtained by correcting a difference in rotation values calculated for each video depending on the position, angle, etc. of the camera.
[0326] The eyeball rotation learning unit 130' can extract eyeball region images for the left eye and the right eye from the frame image and train the third artificial neural network model to generate information related to phase changes of the left eye and the right eye.
[0327] According to an embodiment of the present specification, the eyeball coordinate learning unit 120' can train the second artificial neural network model using information generated by the first artificial neural network model, and the eyeball rotation learning unit 130' can train the third artificial neural network model using information generated by the second artificial neural network model.
[0328] For example, the eyeball coordinate learning unit 120' may generate learning data using information related to head coordinates in a plurality of multi-frame images input to the first artificial neural network model and a frame image of an mth moving image extracted from each multi-frame image. The eyeball coordinate learning unit 120' may generate an eyeball region image in each multi-frame image using information related to each multi-frame image and head coordinates extracted from each multi-frame image. The eyeball coordinate learning unit 120' may generate an eyeball region image in each multi-frame image using information related to eye coordinates in the information related to the head coordinates. The eyeball coordinate learning unit 120' may train the second artificial neural network model according to the process described above.
[0329] The eyeball rotation learning unit 130' can further use information related to the coordinates of the pupil center from each multi-frame image generated by the second artificial neural network model to train the third artificial neural network model. Also, the eyeball rotation learning unit 130' can use the iris and / or pupil images segmented by the eyeball coordinate learning unit 120' to generate training data according to the above-described process to train the third artificial neural network model.
[0330] The first to third artificial neural network models can be trained independently of each other, or can be trained using information generated by each artificial neural network model.
[0331] The following describes how the balance management system 10' generates balance status information and executes a balance rehabilitation program using the trained artificial neural network model.
[0332] The balance function management system 10' can acquire n video frame images in real time from n cameras, which capture a subject undergoing a balance function status test and / or a balance function rehabilitation program.
[0333] The balance function management system 10' can adjust the synchronization of the multiple cameras through at least one processor. For example, the at least one processor can adjust the synchronization of the multiple cameras in real time using technology such as genlock, which is one example, and can adjust the synchronization of the multiple cameras using technology widely known to those skilled in the art.
[0334] In addition, the balance function management system 10' may sample frame images from multiple cameras through at least one processor. For example, one camera may capture the subject at 100 FPS, and another camera may capture the subject at 50 FPS. In this case, the balance function management system 10' may use at least one processor to downsample a video captured at 100 FPS by 1 / 2, or upsample a video captured at 50 FPS by 2. This is merely an example and is not limiting.
[0335] Preferably, the balance management system 10' can capture real-time video of a subject undergoing a balance status test and / or a balance rehabilitation program through multiple cameras with the same FPS setting.
[0336] The balance function management system 10' can acquire at least one of information related to the subject's head coordinates, pupil center coordinates, and eyeball phase changes by using at least one of the first to third artificial neural network models.
[0337] The balance function management system 10' can acquire information related to the head coordinates, pupil center coordinates and / or eyeball phase changes using the first to third artificial neural network models.
[0338] In addition, the balance function management system 10' can also generate information related to the head coordinates, pupil center coordinates, and eyeball phase changes by using an algorithm in which at least one processor generates learning data for the first to third artificial neural network models described above.
[0339] In the following description, the first to third artificial neural network models are used to generate information related to the head coordinates, pupil center coordinates, and eyeball phase changes, but it is not necessary to use an artificial neural network model to generate the information.
[0340] The head coordinate acquisition unit 140' executes the first artificial neural network model stored in the memory 100' and inputs a multi-frame image obtained by connecting frame images of n moving images to the first artificial neural network model to acquire information related to head coordinates according to the mth moving image. The information related to head coordinates according to the mth moving image may mean information related to head coordinates for the mth moving image generated according to the order of the multi-frame images.
[0341] The eyeball coordinate acquisition unit 150' executes the second artificial neural network model stored in the memory 100' and inputs information related to the head coordinates into the second artificial neural network model to acquire information related to pupil center coordinates according to the mth moving image. The information related to pupil center coordinates according to the mth moving image may mean information related to pupil center coordinates of the mth moving image generated in the order of multiple frame images.
[0342] The phase change acquisition unit 160' executes the third artificial neural network model stored in the memory 100' and inputs information related to the pupil center coordinates of the mth moving image in the time order of the multiple frame images to the third artificial neural network model to acquire information related to the phase change of the eyeball due to the mth moving image. The information related to the phase change of the eyeball due to the mth moving image may mean information related to the phase change of the eyeball generated according to the order of the frame images of the mth moving image.
[0343] FIG. 18 is a block diagram of a balance management system for generating balance status information according to one embodiment of the present disclosure.
[0344] Referring to FIG. 18, a balance function management system 10'-1 for generating balance function status information according to one embodiment of the present specification may include a memory 100', a head coordinate learning unit 110', an eyeball coordinate learning unit 120', an eyeball rotation learning unit 130', a head coordinate acquisition unit 140', an eyeball coordinate acquisition unit 150', a phase change acquisition unit 160', a head movement generation unit 1100', an eyeball movement generation unit 1110', a velocity information generation unit 1120', and a balance function status information generation unit 1130'.
[0345] The memory 100', head coordinate learning unit 110', eyeball coordinate learning unit 120', eyeball rotation learning unit 130', head coordinate acquisition unit 140', eyeball coordinate acquisition unit 150', and phase change acquisition unit 160' have been described above, so repeated description will be omitted.
[0346] The head movement generation unit 1100' may generate information related to head movement in the mth video using information related to head coordinates acquired from the first artificial neural network model. The information related to head movement may include horizontal and vertical head movement over time, and the degree of head rotation. The degree of head rotation may refer to the rotation angle in the roll, pitch, and yaw directions of the head. The information related to head movement may be expressed as a graph of horizontal and vertical head coordinate values over time.
[0347] According to an embodiment of the present specification, the head movement generation unit 1100' may calculate a normal vector of the subject's head using head feature points and their coordinates generated in each frame image. The direction of the normal vector may indicate a direction in which the subject's head is facing forward. The direction in which the subject's head is facing forward may indicate a direction in which the tip of the nose is facing. The head movement generation unit 1100' may calculate the head normal vector using the head feature points based on the tip of the nose feature point among the head feature points.
[0348] Alternatively, the direction of the subject's head may refer to the direction of any feature point (such as the edge of the forehead or the center of the lips) on a straight line perpendicular to the nose tip feature point. The head movement generation unit 1100' may calculate a normal vector of the head based on any one of the feature points.
[0349] Calculating a normal vector for the front of the head using feature points is a technique well known to those skilled in the art, and therefore a detailed description thereof will be omitted.
[0350] The head movement generation unit 1100' may generate information related to head movement over time using 3D head coordinate information and normal vectors according to frame images of the mth moving image, and may output the information related to head movement in the form of a graph.
[0351] The eye movement generating unit 1110' may generate information related to eye movement in the mth moving image using information related to pupil center coordinates and eye phase changes generated by the second and third artificial neural network models. The information related to eye movement may include information related to vertical and horizontal movements of the pupil center over time and eye rotation values. The information related to eye movement may be expressed as a graph of vertical coordinate values, horizontal coordinate values, and rotation angles of the pupil centers of the left and right eyes over time.
[0352] The eyeball movement generator 1110' may calculate a gaze vector of the eyeball using vertical coordinate values, horizontal coordinate values, and rotation values of the pupil center. Calculating a gaze vector using the coordinate values and rotation values of the pupil center is a technique well known to those skilled in the art, and therefore, detailed description thereof will be omitted.
[0353] The eye movement generating unit 1110' can output information related to the eye movement in the form of a graph.
[0354] According to one embodiment of the present specification, the head movement generation unit 1100' and eye movement generation unit 1110' can correct errors between head movement and eye movement information based on learning data and the subject's head movement and eye movement information.
[0355] The head movement and eye movement information based on the training data may represent actual data values for training the artificial neural network model.
[0356] As an example, the actual data values may refer to parameter values of the Euler angles of the head and eyes acquired from the eye coordinate learning unit 120′. The eye coordinate learning unit 120′ may set parameter values of a virtual camera to resemble actual settings in a balance function status test and / or a balance function rehabilitation program. The eye coordinate learning unit 120′ may change parameter values of the Euler angles to resemble head and eye movements of a subject in a balance function status test and / or a balance function rehabilitation program. In this case, head and eye movement information of a virtual object resulting from changes in head and eye parameter values may refer to actual data values. This corresponds to one example, and the actual data values may refer to head and eye movement information based on actual data values that can be used to train the first to third artificial neural network models.
[0357] The head movement generation unit 1100' and the eye movement generation unit 1110' can correct errors in head movement and eye movement between frame images acquired within a preset time. For example, when the preset time is 1.5 seconds and a camera captures an image of a subject at 100 FPS, the head movement generation unit 1100' and the eye movement generation unit 1110' can correct errors in head movement and eye movement between 150 frame images. This is just an example and is not limited by the time and frame rate.
[0358] The head movement generator 1100' and eye movement generator 1110' can correct the head movement and eye movement errors so that they have values within a preset range.
[0359] For example, the head movement generation unit 1100' and the eye movement generation unit 1110' may generate information related to head and eye movements using information related to head coordinates, pupil center coordinates, and / or eye phase changes acquired from 150 frame images (frame images acquired for 1.5 seconds based on an arbitrary frame image) in a video in which multiple cameras capture the subject at 100 FPS. The information related to head and eye movements may be calculated as the amount of head and eye movements (vertical, horizontal, and rotational) over time according to the order of the frame images. In this case, the amount of head and eye movements at the time corresponding to the 50th frame image (frame image acquired 0.5 seconds after an arbitrary frame image) may deviate from a predetermined error range. In this case, the head movement generation unit 1100' and the eye movement generation unit 1110' may calculate statistics such as the average or median of the amount of head and eye movements generated using information from previous frame images, and replace the amount of movement at the time corresponding to the 50th frame image.
[0360] Although it has been stated that errors are corrected between frame images acquired for 1.5 seconds based on an arbitrary frame image, this is merely an example, and various embodiments are possible, such as frame images acquired before 1.5 seconds based on the arbitrary frame image, frame images acquired around 1.5 seconds, etc. This is merely an example, and is not limited by the time, frame rate, statistical values, etc.
[0361] As another example, the head movement generator 1100' and the eye movement generator 1110' may apply a filter to correct errors in the head and eye movements. For example, the head movement generator 1100' and the eye movement generator 1110' may correct the errors using a filter such as a Chaining Kalman filter, a Moving Average filter, a Savitzky-Golay filter, a High Pass filter, a Low Pass filter, or a Band Pass filter. This is just an example and is not limited thereto, and various types of filters may be used.
[0362] According to another embodiment of the present specification, the head coordinate learning unit 110', the eyeball coordinate learning unit 120', and the eyeball rotation learning unit 130' can train the first to third artificial neural network models to correct the errors. The first to third artificial neural networks can generate information related to the error-corrected head coordinates, pupil center coordinates, and eyeball phase changes.
[0363] The velocity information generator 1120' may generate information related to the velocity of head and eye movement in the mth moving image using information related to the head movement and eye movement. The velocity information generator 1120' may calculate a vertical movement velocity of the head, a horizontal movement velocity, and / or a head rotation velocity over time using the information related to the head movement. The velocity information generator 1120' may calculate a vertical movement velocity of the eye, a horizontal movement velocity, and / or an eye rotation velocity over time using the information related to the eye movement.
[0364] According to one embodiment of the present disclosure, the velocity information generator 1120′ may filter noise values from the information related to the velocity of head and eye movement. For example, when performing a balance function status test, noise values that prevent the velocity of head and eye movement from being accurately calculated may be generated if the subject's eyes are covered, if the subject rotates their head quickly, if the subject rotates their head slowly, or if the head position changes. The velocity information generator 1120′ may remove noise from the information related to the velocity of head and eye movement using a filter such as a Chaining Kalman filter, a Moving Average filter, a Savitzky-Golay filter, a High Pass filter, a Low Pass filter, or a Band Pass filter. This is by way of example only, and various noise processing methods may be used.
[0365] According to one embodiment of the present disclosure, the speed information generator 1120' may generate information related to the speed of head movement and eye movement within a predetermined time based on the time point when the head movement exceeds a predetermined threshold. In a balance function status test, the subject may move their head horizontally (lateral left, lateral right). In this case, the speed information generator 1120' may generate information related to the speed of head and eye movement when the horizontal head movement exceeds a predetermined threshold.
[0366] In addition, the subject may move their head in a downward and upward right direction while turning the right side of their face so that the right anterior semicircular canal and the left posterior semicircular canal (Right Anterior, Left Posterior, RALP) are stimulated. In addition, the subject may move their head in a downward and upward left direction while turning the left side of their face so that the right anterior semicircular canal and the left anterior semicircular canal (Left Anterior, Right Posterior, RALP) are stimulated. In this case, the velocity information generating unit 1120' may generate information related to the velocity of head and eye movement when the vertical movement of the head is equal to or greater than a preset threshold value.
[0367] In the following, the direction in which the subject rotates their head to stimulate RALP will be referred to as the RALP direction, and the direction in which the subject rotates their head to stimulate LARP will be referred to as the LARP direction.
[0368] For example, the speed information generator 1120' may determine whether the head movement is equal to or greater than a threshold value using head movement information according to frame images existing within a preset time based on the last input frame image. The speed information generator 1120' may determine whether the head movement is equal to or greater than a threshold value by calculating the difference between the maximum and minimum values of head feature point coordinates in the head movement information according to frame images existing within a preset time. The preset threshold value may vary depending on the frame rate of a video, the size of a frame image, etc.
[0369] As another example, the memory 100' may further store a fourth artificial neural network model that generates information about head movement. The fourth artificial neural network model may be trained by at least one processor using frame images of a video for performing a balance function status test and data on head pitch and yaw values in the corresponding frame images. The fourth artificial neural network model may be a time series model or a Transformer model, which are just examples and are not limited to the model. In this case, the frame images may be labeled with information about the LAteral, RALP, and LARP directions. The velocity information generator 1120' may input frame images present within a predetermined time based on the last acquired frame image to the fourth artificial neural network model to determine whether head movement is above a threshold value.
[0370] The speed information generator 1120' may calculate the speed of head and eye movement using information related to head and eye movement generated within a predetermined time after the head movement exceeds a critical value. As an example, the speed information generator 1120' may calculate the speed of head and eye movement using information related to head and eye movement generated within 1.5 seconds after the head movement exceeds a critical value, but this is just an example and is not limited by the time. The speed information generator 1120' may calculate the speed of head and eye movement in the mth moving image.
[0371] The velocity information generator 1120' may display information related to the velocity of the head and eye movement on a display device. The velocity information generator 1120' may display information related to the velocity of the head and eye movement with noise removed and / or information related to the velocity of the head and eye movement without noise removed on a display.
[0372] The balance function status information generating unit 1130' may generate balance function status information of the subject using information related to the movement speed of the head and eyeballs.
[0373] According to an embodiment of the present specification, the balance function status information generator 1130' may calculate a gain coefficient using a time value (Head Peak Index) when the speed of the subject's head movement is relatively the fastest within a preset analysis window and a time value (Eye Peak Index) when the speed of the subject's eye movement is relatively the fastest when the subject's eyes move in the direction of the head movement and then return to their original positions. The analysis window may refer to a preset time range based on a point in time when the head movement exceeds a threshold value. The size of the analysis window may correspond to a time range in which the speed information generator 1120' generates information related to the speed of head and eye movement.
[0374] For example, the size of the analysis window may be 1.5 seconds after the head movement reaches a critical value, but this is just an example and is not limited thereto.
[0375] The balance function status information generator 1130' can calculate a gain coefficient using the Head Peak Index and Eye Peak Index, which are noise-removed information related to the movement speed of the head and eyeballs, using Equation 1.
[0376] Thereafter, the balance function status information generator 1130' can calculate a gain value using Equation 2.
[0377] The balance function status information generator 1130' can calculate gain values for the left and right eyes, respectively.
[0378] The balance function status information generator 1130′ may calculate gain values for each of the m-th videos. For example, if two cameras are used to capture images of the subject, gain values for the two videos may be calculated. In this case, gain values for the left eye and right eye in the first video and gain values for the left eye and right eye in the second video may be calculated. The balance function status information generator 1130′ may calculate statistics for the gain values for the left eye calculated for the first and second videos, and may calculate at least one statistics for the gain values for the right eye calculated for the first and second videos. The statistics may correspond to an average value, a median value, a minimum value, a maximum value, a standard deviation, etc., but are not limited thereto.
[0379] According to an embodiment of the present specification, the head movement generation unit 1100' may further generate reference head movement information by calculating statistics of information related to head coordinates according to the mth moving image. The head movement generation unit 1100' may further generate reference head movement information by calculating statistics of information related to head coordinates according to the mth moving image generated from a multi-frame image.
[0380] The eyeball movement generating unit 1110' may further generate reference eyeball movement information by calculating statistics of information related to pupil center coordinates and eyeball phase changes according to the mth moving image. The eyeball movement generating unit 1110' may further generate reference eyeball movement information by calculating statistics of information related to pupil center coordinates and eyeball phase changes according to each moving image generated from multi-frame images.
[0381] When the subject is photographed using multiple cameras, the 3D head coordinate values generated in the frame image of the mth video may differ depending on the camera position, angle, etc.
[0382] For example, a first camera may photograph the subject on the right side of the subject, and a second camera may photograph the subject on the left side of the subject. In this case, if the subject turns his / her head to the right, the coordinates of the feature points located on the right side of the subject's face in the frame image of the first video captured by the first camera may be calculated relatively more accurately than the coordinates of the feature points located on the left side. Also, if the subject turns his / her head to the left, the coordinates of the feature points located on the left side of the subject's face in the frame image of the second video captured by the second camera may be calculated relatively more accurately than the coordinates of the feature points located on the right side.
[0383] As another example, multiple cameras may be installed around the subject at 15° intervals, with the subject at a distance of 1 m from the center. The multiple cameras may capture the subject at eye level. In this case, the coordinates of feature points may be generated for each camera, which may be measured relatively more accurately depending on the direction of the subject's head movement.
[0384] In this way, the two-dimensional head coordinate information calculated based on the positions and angles of multiple cameras may differ from each other.
[0385] The head movement generation unit 1100' may generate reference head coordinate information by calculating an average value of 3D head coordinates generated in frame images where syncs match in the mth moving image. The head movement generation unit 1100' may further generate information related to reference head movement using the reference head coordinate information over time.
[0386] Alternatively, the head coordinate acquisition unit 110' may acquire the reference head coordinates from the first artificial neural network model. The head movement generation unit 1100' may further generate information related to the reference head movement using the reference head coordinates.
[0387] The eyeball movement generation unit 1110' may generate reference pupil center coordinates and information related to a phase change of the reference eyeball by calculating an average value of information related to pupil center coordinates and a phase change of the reference eyeball generated in frame images whose syncs are aligned in the mth moving image. The eyeball movement generation unit 1110' may generate a reference gaze vector of the eyeball using the reference pupil center coordinates and information related to a phase change of the reference eyeball. The eyeball movement generation unit 1110' may further generate information related to the movement of the reference eyeball using the reference pupil center coordinates, information related to a phase change of the reference eyeball, and the reference gaze vector.
[0388] Alternatively, the eyeball coordinate acquiring unit 150' and the phase change acquiring unit 160' may acquire information related to the reference pupil center coordinates and information on the rotation value of the reference eyeball from the second and third artificial neural network models. The eyeball movement generating unit 1110' may generate information related to the movement of the reference eyeball using the information related to the reference pupil center coordinates and information on the rotation value of the reference eyeball.
[0389] In this case, the speed information generating unit 1120' may generate information related to the head movement speed and eye movement speed for the mth moving image within a predetermined time based on the point in time when the reference head movement becomes equal to or greater than a predetermined critical value.
[0390] The balance function status information generating unit 1130' may calculate gain values using a Head Peak Index and an Eye Peak Index for the mth moving image within a predetermined time based on the point at which the reference head movement exceeds a predetermined critical value.
[0391] 9, the head movement generator 1100', eye movement generator 1110', velocity information generator 1120', and balance function status information generator 1130' can output calculated information on a display screen. Videos captured by a plurality of cameras can be output on the display screen.
[0392] The velocity information generator 1120′ may output head and eye movement velocity information for the mth moving image as a graph 205. The head and eye movement velocity graph 205 may display velocity information according to the number of balance function status tests superimposed. The time values at which peaks and / or valleys appear in the head and eye movement velocity graph 205 may correspond to the Head Peak Index and / or Eye Peak Index. In the graph, the vertical axis may represent velocity values, and the horizontal axis may represent time values. While FIG. 9 illustrates a graph for the horizontal velocity of the head and eye, this is merely an example, and graphs for velocity in the vertical and rotational directions may also be output depending on the type of test.
[0393] The head movement generator 1100' and the eye movement generator 1110' can output head and eye movement information in the form of graphs. The head movement generator 1100' and the eye movement generator 1110' can output a horizontal head and eye movement graph 206 and a vertical head and eye movement graph 207. In addition, they can further output movement graphs for the rotation direction of the head and / or eyes. In addition, the head and eye movement graphs can include movement information for the mth moving image, and can include the reference head movement and reference eye movement information.
[0394] The balance function status information generator 1130′ may output information 208 related to the balance function status test for the subject to a display device. The information related to the balance function status test may include at least one of head rotation direction (lateral, RALP, LARP), gain value and standard deviation according to head rotation direction, number of balance function status tests according to head rotation direction, number of gain value calculation successes, and number of gain value calculation failures.
[0395] A failure in the calculation of the gain value may occur when the subject moves their head faster than the standard of the test. In this case, the difference between the Head Peak Index and the Eye Peak Index may exceed the middle of the analysis window size. In this case, the balance function status information generator 1130' may fail to calculate the gain value. The balance function status information generator 1130' may generate information on the number of successes and failures in the calculation of the gain value, thereby enabling the subject and / or the examiner to make a more accurate judgment.
[0396] In addition, the balance function status information generating unit 1130' may generate information regarding the presence or absence of a semicircular canal abnormality according to the gain value. For example, the balance function status information generating unit 1130' may generate information regarding the semicircular canal abnormality when the gain values of the left and right eyes are equal to or less than a preset value in the balance function status test.
[0397] In addition, the subject may rotate his / her head laterally during the balance function status test. In this case, if the difference in gain value between the left eye and the right eye when the subject rotates his / her head to the left and right is equal to or greater than a preset value, the balance function status information generator 1130′ may generate information about abnormality of the semicircular canals.
[0398] In addition, in the balance function status test, the subject may rotate his / her head in the RALP or LARP direction. In this case, if the difference in gain value between the left eye and the right eye when the subject rotates his / her head upward and downward is equal to or greater than a preset value, the balance function status information generator 1130' may generate information about abnormality of the semicircular canals.
[0399] In addition, the balance function status information generating unit 1130' can further generate information regarding the presence or absence of abnormalities in central balance nerve function and peripheral balance nerve function using information related to the eye movement and gain value information.
[0400] FIG. 19 is a block diagram of a balance function management system for implementing a balance function rehabilitation program according to one embodiment of the present specification.
[0401] 19, a balance function management system 10'-2 for performing a balance function rehabilitation program according to an embodiment of the present specification may include a memory 100', a head coordinate learning unit 110', an eye coordinate learning unit 120', an eye rotation learning unit 130', a head coordinate acquiring unit 140', an eye coordinate acquiring unit 150', a phase change acquiring unit 160', a head movement generating unit 1100', an eye movement generating unit 1110', a target output unit 1140', a head direction providing unit 1150', and a feedback providing unit 1160'. The memory 100', head coordinate learning unit 110', eye coordinate learning unit 120', eye rotation learning unit 130', head coordinate acquiring unit 140', eye coordinate acquiring unit 150', phase change acquiring unit 160', head movement generating unit 1100', and eye movement generating unit 1110' have been described above, and therefore will not be described again.
[0402] As shown in Fig. 11, the target output unit 1140' can output a virtual target 209 to a display device. The virtual target can be displayed at any position on the display device. In Fig. 11, the target is illustrated in the shape of a playing card, but this is merely an example and the shape is not limited thereto.
[0403] The head direction providing unit 1150' may provide the subject with information regarding the head rotation direction according to a balance rehabilitation protocol. For example, the head direction providing unit 1150' may provide information to the subject to rotate their head in a lateral direction.
[0404] Alternatively, the head direction providing unit 1150' may provide information to the subject to rotate their head in a RALP or LARP direction.
[0405] In this case, the head direction providing unit 1150' may provide information to the subject so that the subject rotates his / her head only up or down while the right side of his / her head faces forward. Alternatively, the head direction providing unit 1150' may provide information to the subject so that the subject rotates his / her head only up or down while the left side of his / her head faces forward.
[0406] In addition, the head direction providing unit 1150' may provide information to restore the state before the subject's head was rotated within a predetermined time after the subject has rotated their head. For example, the head direction providing unit 1150' may provide information to restore the state before the subject's head was rotated within one second.
[0407] The head direction providing unit 1150' can visually display the information on a display device, and can also output the information to an audio device.
[0408] The feedback providing unit 1160' may provide feedback according to the subject's head and eye movements.
[0409] A standard angle by which the subject should rotate his / her head may be predetermined according to the balance rehabilitation protocol. The feedback providing unit 1160′ may compare the subject's head rotation angle generated by the head movement generating unit 1100′ with the standard angle. The feedback providing unit 1160′ may provide auditory feedback and / or visual feedback if the subject's head rotation angle satisfies the standard angle. Furthermore, the feedback providing unit 1160′ may provide auditory feedback and / or visual feedback if the subject's head rotation angle does not satisfy the standard angle. The feedback providing unit 1160′ may provide different feedback depending on whether the subject's head rotation angle satisfies the standard angle.
[0410] In addition, the eye movement generating unit 1110' may generate coordinate information of a gaze point to which the subject's gaze is directed on a display device using an eye gaze vector. The feedback providing unit 1160' may compare the coordinate information of the gaze point generated by the eye movement generating unit 1110' with coordinate information of the virtual target 209.
[0411] When the coordinates of the gaze point are located within the area of the virtual target 209, the feedback providing unit 1160′ may change the hue value of the virtual target 209. In addition, the feedback providing unit 1160′ may further indicate the gaze point within the virtual target 209.
[0412] When the coordinates of the gaze point are located outside the area of the virtual target 209, the feedback providing unit 1160' can display the position of the gaze point on a display device.
[0413] When the subject is photographed by a plurality of cameras, the head movement generation unit 1100' and the eye movement generation unit 1110' may generate information related to the reference head movement and the reference eye movement as described above, and the feedback providing unit 1160' may provide feedback according to the information related to the reference head movement and the reference eye movement.
[0414] FIG. 20 is a block diagram of a balance function management system that generates balance function status information and executes a balance function rehabilitation program according to one embodiment of the present specification.
[0415] 20, the balance function management system 10'-3 may include a memory 100', a head coordinate learning unit 110', an eye coordinate learning unit 120', an eye rotation learning unit 130', a head coordinate acquiring unit 140', an eye coordinate acquiring unit 150', a phase change acquiring unit 160', a head movement generating unit 1100', an eye movement generating unit 1110', a velocity information generating unit 1120', a balance function status information generating unit 1130', a target output unit 1140', a head direction providing unit 1150', and a feedback providing unit 1160'. The balance function management system 10'-3 may generate gain values and provide a balance function rehabilitation program depending on whether or not there is a balance function abnormality.
[0416] The head coordinate learning units 110, 110', eye coordinate learning units 120, 120', eye rotation learning units 130, 130', head coordinate acquiring units 140, 140', eye coordinate acquiring units 150, 150', phase change acquiring units 160, 160', head movement generating units 1100, 1100', eye movement generating units 1110, 1110', velocity information generating units 1120, 1120', balance function status information generating units 1130, 1130', target output units 1140, 1140', head direction providing units 1150, 1150', and feedback providing units 1160, 1160' may include a processor, an application-specific integrated circuit (ASIC), other chipsets, logic circuits, registers, communication modems, data processing devices, etc., known in the art to which the present invention pertains, for performing calculations and various control logic. Furthermore, when the above-described control logic is implemented in software, the head coordinate learning unit 110, 110', eye coordinate learning unit 120, 120', eye rotation learning unit 130, 130', head coordinate acquiring unit 140, 140', eye coordinate acquiring unit 150, 150', phase change acquiring unit 160, 160', head movement generating unit 1100, 1100', eye movement generating unit 1110, 1110', velocity information generating unit 1120, 1120', balance function status information generating unit 1130, 1130', target output unit 1140, 1140', head direction providing unit 1150, 1150', and feedback providing unit 1160, 1160' may be implemented as a set of program modules. In this case, the program modules may be stored in the memory device and executed by the processor.
[0417] Hereinafter, a method for generating balance status information and a method for balance rehabilitation using a balance management system according to the present invention will be described. However, in describing the method for generating balance status information and the method for balance rehabilitation according to the present invention, repeated descriptions of each component will be omitted.
[0418] FIG. 21 is a flowchart of a balance function status information generating method according to one embodiment of the present specification.
[0419] Referring to FIG. 21, in step S10, at least one processor may input frame images of n videos to a first artificial neural network model to acquire information related to head coordinates. The n videos may refer to videos captured by n cameras of a subject undergoing a balance function status test. The frame images of the n videos may be input sequentially as described above, and multiple frame images for the n videos may be input. Thereafter, the at least one processor may input information related to head coordinates for the n videos to a second artificial neural network model to acquire information related to pupil center coordinates. Thereafter, the at least one processor may input information related to pupil center coordinates for the n videos to the third artificial neural network model to acquire information related to eye phase changes. The learning processes of the first to third artificial neural network models have been described above, so repeated description will be omitted.
[0420] In step S11, the at least one processor may generate head and eye movement information using information related to head coordinates, pupil center coordinates, and eye phase changes. The at least one processor may generate head and eye movement information for the mth moving image, respectively. Alternatively, the at least one processor may further generate reference head movement information and reference eye movement information for the mth moving image.
[0421] In step S12, when the head movement in the mth video of the subject is equal to or greater than a predetermined threshold, the at least one processor may generate velocity information of head and eye movement for the mth video. Also, when the reference head movement of the subject is equal to or greater than a predetermined threshold, the at least one processor may generate velocity information of head and eye movement for the mth video.
[0422] In step S13, the at least one processor may calculate a gain value using Equation 1 and Equation 2. The at least one processor may generate balance function status information using the gain value.
[0423] FIG. 22 is a flowchart of a balance function rehabilitation method according to one embodiment of the present specification.
[0424] Referring to FIG. 22, in step S20, at least one processor may input frame images of n videos to a first artificial neural network model to acquire information related to head coordinates. The n videos may refer to videos captured by n cameras of a subject undergoing a balance rehabilitation program. The frame images of the n videos may be input sequentially as described above, and multiple frame images for the n videos may be input. Thereafter, the at least one processor may input information related to head coordinates for the n videos to a second artificial neural network model to acquire information related to pupil center coordinates. Thereafter, the at least one processor may input information related to pupil center coordinates for the n videos to the third artificial neural network model to acquire information related to eye phase changes. The learning processes of the first to third artificial neural network models have been described above, so repeated description will be omitted.
[0425] In step S21, the at least one processor may generate head and eye movement information using information related to head coordinates, pupil center coordinates, and eye phase changes. The at least one processor may generate head and eye movement information for the mth video, respectively. The at least one processor may further generate reference head movement information and reference eye movement information for the mth video. The at least one processor may generate coordinate information of a gaze point according to the subject's gaze using the reference eye movement information.
[0426] In step S22, the at least one processor may output the virtual target on a display.
[0427] In step S23, the at least one processor may provide the subject with information regarding the direction of head rotation, which has been described above and will not be described again.
[0428] In step S24, the at least one processor may utilize head and eye movement information to provide auditory and / or visual feedback to the subject.
[0429] The balance function status information generating method and balance function rehabilitation method can be implemented in the form of a computer program written to execute each step on a computer and recorded on a computer-readable recording medium. The computer program may include code written in a computer language, such as C / C++, C#, JAVA, Python, or machine language, that can be read by the computer's processor (CPU) through the computer's device interface so that the computer can load the program and execute the method implemented in the program. This code may include functional code associated with functions defining the functions required to execute the method, and execution procedure-related control code required for the computer's processor to execute the function according to a predetermined procedure. This code may also include memory reference-related code indicating at which location (address) in the computer's internal or external memory additional information or media required for the computer's processor to execute the function should be referenced. In addition, if the computer processor needs to communicate with any other remote computers or servers to perform the function, the code may further include communication-related code for determining how to communicate with any other remote computers or servers using the computer's communication module, and what information or media to send and receive during communication.
[0430] The storage medium refers to a medium that stores data semi-permanently and can be read by a device, rather than a medium that stores data for a short period of time, such as a register, cache, or memory. Specific examples of the storage medium include, but are not limited to, ROM, RAM, CD-ROM, magnetic tape, floppy disk, and optical data storage device. That is, the program can be stored in various recording media on various servers to which the computer can connect, or various recording media on the user's computer. Furthermore, the medium can be distributed among computer systems connected via a network, and computer-readable code can be stored in a distributed manner.
[0431] Although the embodiments of the present specification have been described above with reference to the accompanying drawings, those skilled in the art will understand that the present invention can be embodied in other specific forms without changing the technical spirit or essential features thereof. Therefore, it should be understood that the embodiments described above are illustrative in all respects and are not limiting. [Explanation of symbols]
[0432] 10,10' balance function management system, 100,100' memory, 110,110' head coordinate learning unit, 120,120' eye coordinate learning unit, 130,130' eye rotation learning unit, 140,140' head coordinate acquisition unit, 150,150' eye coordinate acquisition unit, 160,160' phase change acquisition unit, 1100,1100' head movement generation unit, 1110,1110' eye movement generation unit, 1120,1120' velocity information generation unit, 1130,1130' balance function status information generation unit, 1140,1140' target output unit, 1150,1150' head direction providing unit, 1160,1160' feedback providing unit.
Claims
1. a memory storing at least one processor and instructions executable by said processor and storing at least one artificial neural network model to be executed by the computing device; The at least one processor inputs n video frame images obtained by photographing the subject through n (a natural number) cameras into at least one artificial neural network model to acquire at least one of information related to the subject's head coordinates, pupil center coordinates, and eyeball phase change in the order of the m-th (a natural number from 1 to n) video frame images; Using the information, balance function status information is generated or information related to head and eye movements is generated for carrying out a balance function rehabilitation program. A balance function management system characterized by:
2. The at least one processor a head coordinate acquisition unit that executes a first artificial neural network model stored in the memory, inputs a frame image of an m-th video or a multi-frame image obtained by connecting frame images of n videos to the first artificial neural network model, and acquires information related to a head coordinate according to the m-th video; an eyeball coordinate acquisition unit that executes a second artificial neural network model stored in the memory, inputs information related to the head coordinates into the second artificial neural network model, and acquires information related to pupil center coordinates according to the mth moving image; a phase change acquiring unit that executes the third artificial neural network model stored in the memory, and inputs information related to pupil coordinates in time order of the frame image or multi-frame image of the mth moving image into the third artificial neural network model to acquire information related to a phase change of the eyeball due to the mth moving image, The balance function management system according to claim 1 .
3. The first artificial neural network model is an artificial neural network model trained by the at least one processor using, as training data, facial feature points and coordinates of the feature points extracted from at least one frame image of a video in which a human face is captured or a multi-frame image in which a plurality of frame images of a video are connected; The balance function management system according to claim 2 .
4. The feature points are located within a predetermined area in the frame image or the multi-frame image. The balance function management system according to claim 3 .
5. The second artificial neural network model is an artificial neural network model trained by the at least one processor to generate information related to coordinates of a pupil center by using, as training data, an eyeball region image extracted to include an eyeball from at least one frame image of a video in which a human face is captured or a multi-frame image obtained by connecting a plurality of frame images of a video. The balance function management system according to claim 2 .
6. the at least one processor generates a pupil region image in which a region occupied by a pupil and a remaining region in the eyeball region image have different pixel values, and uses the pupil region image or an array of pixel values of the pupil region image as training data to train the second artificial neural network model; The balance function management system according to claim 5 .
7. the at least one processor generates feature points and coordinate information of the feature points in the eyeball region image, and trains the second artificial neural network model to generate horizontal and vertical coordinate values of the pupil center using coordinates of a plurality of preset feature points. The balance function management system according to claim 5 .
8. the memory stores data of at least one virtual object; The second artificial neural network model is an artificial neural network model trained by the at least one processor to generate information related to the coordinates of the pupil center using training data including parameter values obtained by changing at least one of parameters related to head rotation, eye rotation, and camera settings of the virtual object and images of the virtual object acquired according to the parameter values; The balance function management system according to claim 2 .
9. The third artificial neural network model is an artificial neural network model trained by the at least one processor to generate an eyeball rotation value by using, as training data, information related to eyeball phase changes generated according to a time sequence of an eyeball region image extracted to include an eyeball from a frame image of a video in which a human face is captured or a multi-frame image obtained by connecting a plurality of video frame images; The balance function management system according to claim 2 .
10. The third artificial neural network model is an artificial neural network model trained by the at least one processor using information obtained by comparing pixel values of an iris area between eyeball area images corresponding to each frame image of a video or a multi-frame image of the person's face; The balance function management system according to claim 9 .
11. The at least one processor calculating a size of an area occupied by a pupil in the eye-region image, and adjusting a size of the target eye-region image using the size of the area occupied by the pupil in the previous eye-region image; The balance function management system according to claim 9 .
12. The at least one processor a head movement generator for generating information related to head movements in an m-th video (a natural number from 1 to n) using information related to head coordinates generated by at least one artificial neural network model; an eye movement generation unit that generates information related to eye movement in the m-th video by using information related to pupil center coordinates or eye phase changes generated by at least one artificial neural network model; a speed information generating unit that generates information related to the speed of head and eye movements in the m-th moving image using the information related to the head and eye movements; a balance function status information generating unit that generates balance function status information of the subject using information related to head and eye movement speed; The balance function management system according to any one of claims 1 to 11, comprising:
13. The velocity information generating unit generates information related to the velocity of head movement and the velocity of eye movement within a predetermined time based on a point in time when the head movement exceeds a predetermined threshold value. The balance function management system according to claim 12.
14. the balance function status information generating unit calculates a gain value using a time value at which the head movement speed is at its maximum and a time value at which the eye movement speed is at its maximum; 13. The balance function management system according to claim 12.
15. When n is 2 or more, The head movement generation unit Calculating statistics of information related to head coordinates according to the mth video to further generate reference head movement information; The eyeball movement generation unit Calculating statistics of information related to pupil center coordinates and eyeball phase changes according to the mth video to further generate reference eyeball movement information; The balance function management system according to claim 12.
16. The at least one processor a head movement generator for generating head movement information for an m-th video (a natural number from 1 to n) using information related to head coordinates acquired from at least one artificial neural network model; an eye movement generating unit that generates eye movement information for the m-th moving image by using information related to pupil center coordinates or eye phase changes acquired from at least one artificial neural network model; a target output unit that outputs a virtual target to a display device; a head direction providing unit that provides directional information regarding head movement to a subject; a feedback providing unit that provides feedback based on the subject's head movement and eye movement, The balance function management system according to any one of claims 1 to 11.
17. 1. A balance function management system including at least one processor and a memory storing instructions executable by said processor and storing at least one artificial neural network model executed by a computing device, comprising: an information acquiring step of inputting n video frame images, obtained by the at least one processor through n (a natural number) cameras, into at least one artificial neural network model to acquire information related to the head coordinates, pupil center coordinates, and eyeball phase changes of the subject in the order of the m-th video frame images (a natural number from 1 to n); a balance function status information generating step in which the at least one processor generates head movement information and eye movement information using the information generated in the information acquiring step, calculates head movement speeds and eye movement speeds, and generates information related to a balance function status using information related to the head movement speeds and eye movement speeds. A method for generating balance function status information.
18. 1. A balance function management system including at least one processor and a memory storing instructions executable by said processor and storing at least one artificial neural network model executed by a computing device, comprising: an information acquiring step of inputting n video frame images of the subject captured through n (a natural number) cameras by the at least one processor into at least one artificial neural network model to acquire information related to the subject's head coordinates, pupil center coordinates, and eyeball phase changes according to the frame order of the m-th video (a natural number from 1 to n); a balance rehabilitation step in which the at least one processor generates head movement information and eye movement information using the information generated in the information acquisition step, and tracks the head and eye movements of the subject to perform a balance rehabilitation program. A method for rehabilitation of balance function.
19. A computer program recorded on a computer-readable recording medium, the computer program being created to cause a computer to perform the steps of the balance function status information generating method according to claim 17.
20. A computer program recorded on a computer-readable recording medium, the computer program being created to cause a computer to perform each step of the balance function rehabilitation method according to claim 18.
Citation Information
Patent Citations
Dizziness rehabilitation evaluation device
JP2005319027A
Estimation device, estimation system, estimation method, and program for estimation
JP2021041067A
Data processing device, eyeball movement data processing system, data processing method and program
JP2023058161A
Vertigo diagnostic device and method using ultra-fast lightweight deep running model for tracking eyeball and head position change in video image, recording medium storing program for achieving the same and computer program stored in recording medium
JP2025036125A
Eye movement measurement system combined with video- and elctro-oculograph
KR1020040107677A
Cited By
Three-dimensional flash memory having improved integration density
US12628348B2