Method and apparatus for determining degree of dementia of user
A device and method using multimodal data analysis through voice, gaze, and touch reactions processed by neural networks addresses the limitations of existing dementia diagnosis methods, enabling early and accurate assessment of dementia stages.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- AIBLE THERAPEUTICS CO LTD
- Filing Date
- 2025-03-13
- Publication Date
- 2026-05-21
AI Technical Summary
Existing methods for diagnosing dementia, such as neurocognitive tests and imaging techniques, require specialized medical professionals and are costly or inconvenient for patients, often leading to delayed diagnosis for mild cognitive impairment or Alzheimer's disease.
A method and device using an electronic device to determine dementia level through multimodal data analysis, including voice, gaze, and touch reactions to pre-made content, processed by neural networks to classify dementia stages.
Enables early and accurate dementia diagnosis without specialized medical intervention, providing a cost-effective and convenient means to assess dementia levels in various stages.
Smart Images

Figure KR2025099741_21052026_PF_FP_ABST
Abstract
Description
Method and device for determining the degree of dementia in a user
[0001] The technology field relates to a technology for determining a user's dementia level, and in particular, to a device and method for providing content to a user and determining the degree of dementia based on the user's reaction to the provided content.
[0002] Dementia has become the most serious disease affecting the elderly alongside the aging of society, showing a rapid increase over the past decade and consequently a surge in socio-economic costs. Furthermore, it is a disease that prevents patients from living independently and causes immense suffering not only to the patients themselves but also to their caregivers, leading to risks such as disappearance and suicide. While early diagnosis and appropriate treatment can prevent or delay further decline in cognitive function, there are limitations to existing early diagnostic methods. Since patients are required to visit specialized medical institutions such as hospitals, many patients who visit due to worsening forgetfulness have already progressed to mild cognitive impairment (MCI) or Alzheimer's disease (AD). Additionally, highly reliable neurocognitive function tests (such as SNSB-II and CERAD-K) require medical professionals with sufficient experience and expertise, while diagnostic methods like MRI, SPECT, PET, and cerebrospinal fluid analysis are not only expensive but also cause significant inconvenience for the patients undergoing the diagnosis.
[0003] One embodiment may provide a device and method for determining the degree of dementia in a user.
[0004] One embodiment may provide a device and method for determining the degree of dementia of a user based on multimodal data regarding the user's reaction.
[0005] However, technical challenges are not limited to the technical challenges described above, and other technical challenges may exist.
[0006] A method for determining the degree of dementia of a user, performed by an electronic device according to one embodiment, comprises: an operation of outputting a first content pre-made to determine the degree of dementia of a user through a user terminal; an operation of obtaining first reaction information of the user regarding the first content through the user terminal, wherein the first reaction information includes first voice information of the user corresponding to the first content; an operation of outputting a second content pre-made to determine the degree of dementia of the user through the user terminal; an operation of obtaining second reaction information of the user regarding the second content through the user terminal, wherein the second reaction information includes first gaze information regarding coordinates on a display where the user's eyes gaze, corresponding to the second content; an operation of outputting a third content pre-made to determine the degree of dementia of the user through the user terminal; an operation of obtaining third reaction information of the user regarding the third content through the user terminal, wherein the third reaction information includes second voice information of the user corresponding to the third content and second gaze information regarding coordinates on a display where the user's eyes gaze. The method may include the operation of outputting a pre-made fourth content to determine the degree of dementia of the user through the user terminal, the operation of obtaining the user's fourth reaction information regarding the fourth content through the user terminal—wherein the fourth reaction information includes the user's first touch action information corresponding to the fourth content—and the operation of determining the degree of dementia of the user by inputting the first reaction information, the second reaction information, the third reaction information, and the fourth reaction information into a dementia degree classification model based on an artificial neural network.
[0007] The first content above may be content including instructions to induce reading a presented text aloud, content including instructions to induce describing a presented image, content including instructions to induce associating and saying aloud words corresponding to the presented text, content including instructions to induce performing multiple arithmetic operations aloud, or content including instructions to induce describing a past anecdote.
[0008] The second content above may be one of the following: content including instructions to induce gaze at or not gaze at a target area on the display in response to a presented rule, or content including instructions to induce gaze movement in response to a presented image.
[0009] The above third content may be one of the following: content including instructions to induce reading the presented text aloud, content including instructions to induce reading the presented text without speaking, or content including instructions to induce describing the presented image.
[0010] The above-mentioned fourth content may be one of the following: content including instructions to guide connecting multiple points on the display according to a presented rule, content including instructions to guide drawing a pattern corresponding to a presented pattern, or content including instructions to guide drawing a shape corresponding to a presented rule.
[0011] The operation of determining the degree of dementia of the user by inputting the first reaction information, the second reaction information, the third reaction information, and the fourth reaction information into the dementia degree classification model comprises: an operation of obtaining a first biomarker based on the first reaction information; an operation of obtaining a preset number of first features for the first reaction information by inputting the first biomarker into a pre-updated first model; an operation of obtaining a second biomarker based on the second reaction information; an operation of obtaining a preset number of second features for the second reaction information by inputting the second biomarker into a pre-updated second model; an operation of obtaining a third biomarker based on the third reaction information; an operation of obtaining a preset number of third features for the third reaction information by inputting the third biomarker into a pre-updated third model; an operation of obtaining a fourth biomarker based on the fourth reaction information; and a preset number of features for the fourth reaction information by inputting the fourth biomarker into a pre-updated fourth model. It may include an operation of acquiring a number of fourth features and an operation of determining the degree of dementia of the user by inputting the first features, the second features, the third features, and the fourth features into a pre-updated fifth model.
[0012] The first biomarker is a first spectrogram image that visualizes at least one characteristic of the first voice information, and the first model may include a deep neural network (DNN) model that is pre-updated to generate the first features for the first voice information by processing the first spectrogram image.
[0013] The second biomarker is first time series data obtained for the user's gaze of the first gaze information, and the second model may include a convolutional neural network (CNN) model that is pre-updated to generate the second features for the first gaze information by processing the first time series data.
[0014] The third biomarker is a second spectrogram image visualizing at least one characteristic of the second voice information and second time-series data obtained for the user's gaze of the second gaze information, and the third model may include a pre-updated association acquisition model to generate third features regarding the association between the second voice information and the second gaze information by processing the second spectrogram image and the second time-series data.
[0015] The above-mentioned fourth biomarker includes at least one of first picture information for a first picture represented by the first touch action information, third time series data obtained for the first touch action information, and first structured data associated with the first picture, and the fourth model includes at least one of a CNN model pre-updated to generate first sub-features for the first touch action information by processing the first picture information, a recurrent neural network (RNN) model pre-updated to generate second sub-features for the first touch action information by processing the third time series data, or a fully connected layer pre-updated to generate third sub-features for the first touch action information by processing the first structured data, and the fourth features can be obtained based on at least one of the first sub-features, the second sub-features, or the third sub-features.
[0016] The operation of determining the degree of dementia of the user by inputting the first features, the second features, the third features, and the fourth features into the fifth model may include the operation of generating multi-modal data in which the first features, the second features, the third features, and the fourth features are concatenated—where the first features, the second features, the third features, and the fourth features are flattened to the same dimension—and the operation of determining the degree of dementia of the user by inputting the multi-modal data into a pre-updated DNN model.
[0017] The first content, the second content, the third content, and the fourth content may include instructions appearing as voice or text that instruct the user to perform an action.
[0018] The above degree of dementia may be any one of normal, mild cognitive impairment (MCI), and Alzheimer's disease (AD).
[0019] The degree of dementia determined above can be output through the user terminal.
[0020] According to one embodiment, an electronic device for determining the degree of dementia of a user comprises at least one processor including a processing circuit and a memory including one or more storage media for storing instructions, and when the instructions are executed individually or collectively by the at least one processor, the electronic device may: output a first content pre-made to determine the degree of dementia of a user through a user terminal, and obtain first reaction information of the user regarding the first content through the user terminal, wherein the first reaction information includes first voice information of the user; output a second content pre-made to determine the degree of dementia of the user through the user terminal, and obtain second reaction information of the user regarding the second content through the user terminal, wherein the second reaction information includes first gaze information regarding coordinates on a display where the user's eyes gaze; output a third content pre-made to determine the degree of dementia of the user through the user terminal, and obtain third reaction information of the user regarding the third content through the user terminal, wherein the third reaction information includes the user's second Including voice information and second gaze information regarding coordinates on a display where the user's eyes gaze—, outputting a pre-made fourth content to determine the degree of dementia of the user through the user terminal, and obtaining the user's fourth reaction information regarding the fourth content through the user terminal—the fourth reaction information includes the user's first touch action information corresponding to the fourth content—, and determining the degree of dementia of the user by inputting the first reaction information, the second reaction information, the third reaction information, and the fourth reaction information into a dementia degree classification model based on an artificial neural network.
[0021] A device and method for determining the degree of dementia in a user may be provided.
[0022] A device and method for determining the degree of dementia of a user based on at least one of the user's voice, gaze, or drawing may be provided.
[0023] Figure 1 is a configuration diagram of a system for determining the degree of dementia in a user according to one example.
[0024] FIG. 2 illustrates images output to a user terminal to determine the degree of dementia of a user according to one example.
[0025] FIG. 3 is a configuration diagram of an electronic device for determining the degree of dementia in a user according to one embodiment.
[0026] FIG. 4 is a flowchart of a method for determining the degree of dementia in a user according to one embodiment.
[0027] FIGS. 5 to 9 illustrate pre-made content for obtaining a user's voice reaction according to various examples.
[0028] FIGS. 10 and 11 illustrate pre-made content for obtaining a user's gaze reaction according to various examples.
[0029] FIGS. 12 to 14 illustrate pre-made content for obtaining a user's voice and gaze reactions according to various examples.
[0030] FIGS. 15 to 17 illustrate pre-made content for obtaining a user's picture reaction according to various examples.
[0031] FIG. 18 is a flowchart of a method for determining the degree of dementia of a user based on a plurality of reaction information according to one embodiment.
[0032] Figure 19 illustrates a spectrogram image generated for speech according to one example.
[0033] FIGS. 20a to 20c illustrate examples of spectrogram images generated for voices acquired by users with different degrees of dementia.
[0034] FIG. 21 illustrates a model for processing a user's voice biomarker according to one example.
[0035] FIG. 22 illustrates a model for processing a user's gaze biomarker according to one example.
[0036] FIG. 23 illustrates a model for processing a user's voice and gaze biomarkers according to one example.
[0037] FIG. 24 illustrates a model for processing a user's picture biomarker according to one example.
[0038] FIG. 25 is a flowchart of a method for determining the degree of dementia in a user based on multimodal data according to one embodiment.
[0039] FIG. 26 illustrates a model for processing biomarkers of various modalities according to one example.
[0040] Embodiments are described in detail below with reference to the attached drawings. However, the scope of the patent application is not limited or restricted by these embodiments. Identical reference numerals in each drawing indicate identical components.
[0041] Various modifications may be made to the embodiments described below. The embodiments described below are not intended to limit the forms of practice and should be understood to include all modifications, equivalents, and substitutions thereof.
[0042] The terms used in the embodiments are used merely to describe specific embodiments and are not intended to limit the embodiments. Singular expressions include plural expressions unless the context clearly indicates otherwise. In this specification, terms such as "comprising" or "having" are intended to indicate the presence of the features, numbers, steps, actions, components, parts, or combinations thereof described in the specification, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.
[0043] Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as generally understood by those skilled in the art to which the embodiments pertain. Terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant technology, and should not be interpreted in an ideal or overly formal sense unless explicitly defined in this application.
[0044] In addition, when describing with reference to the attached drawings, identical components are assigned the same reference numeral regardless of drawing symbols, and redundant descriptions thereof are omitted. When describing embodiments, if it is determined that a detailed description of related prior art could unnecessarily obscure the essence of the embodiment, such detailed description is omitted.
[0045]
[0046] Figure 1 is a configuration diagram of a system for determining the degree of dementia in a user according to one example.
[0047] According to one embodiment, a system for determining the degree of dementia of a user may include an electronic device (110) for determining the degree of dementia of a user, a user terminal (120) for outputting content, and a monitoring terminal (130) of a medical institution. For example, the electronic device (110) may be a server.
[0048] The electronic device (110) may provide pre-made content to the user terminal (120) to determine the degree of dementia of the user. For example, the content may be content for obtaining at least one of voice and gaze from the user. For example, the content may be content for obtaining a picture (or drawing) from the user. Content output through the user terminal (120) to obtain a reaction from the user is described in detail below with reference to FIGS. 5 through 17.
[0049] According to one embodiment, the content may instruct the user to perform a task via a voice response. The user may record a voice response using a user terminal (120) as a task. For example, the user terminal (120) may be a mobile terminal such as a tablet PC (personal computer) and a smartphone that includes a microphone. As another example, the user terminal (120) may be an electronic device connected to a microphone as an input tool.
[0050] According to one embodiment, the content may instruct the user to perform a task that induces eye movement. As a task, the user may record the eye movement using a user terminal (120). For example, the user terminal (120) may be a mobile terminal such as a tablet PC and a smartphone that includes a camera capable of photographing the user's eyeball. As another example, the user terminal (120) may be an electronic device connected to a camera as an input tool.
[0051] According to one embodiment, the content may instruct the user to perform a task through a touch action. As a task, the user may draw a picture using the user terminal (120). For example, the user terminal (120) may be a mobile terminal such as a tablet PC and a smartphone that includes a touch display. As another example, the user terminal (120) may be an electronic device connected to a tablet as an input tool. The user may draw a picture through the user terminal (120) by using a tool such as a digital pen or by touching it directly with their hand.
[0052] The user terminal (120) can communicate with the electronic device (110) by being connected offline or online. The electronic device (110) provides content to the user terminal (120), and the user terminal (120) outputs the content to the user through a display.
[0053] For example, the user terminal (120) can acquire the user's voice as a voice reaction to the content through a microphone and transmit the acquired voice to the electronic device (110).
[0054] For example, the user terminal (120) can capture user images as gaze reactions to content through a camera, continuously determine coordinates on a display that the user's eyes are looking at based on the captured user images, and transmit the determined coordinates to an electronic device (110).
[0055] For example, the user terminal (120) can receive touch action information as a drawing reaction to content through a touch display or tablet. The user terminal (120) can output a drawing drawn by the user through the display (or touch display) of the user terminal (120). If the user draws a drawing using a digital pen, the user terminal (120) can receive pen information from the digital pen. The user terminal (120) can transmit the drawing information and pen information regarding the acquired drawing to the electronic device (110) as touch action information.
[0056] The electronic device (110) can determine the degree of dementia of a user based on at least one of the user's voice reaction, gaze reaction, and drawing reaction, and transmit the determined degree of dementia to the user terminal (120).
[0057] If the user terminal (120) is a mobile terminal, the user is not restricted by time and place and can measure the degree of dementia at a low cost.
[0058] Below, a method for determining the degree of dementia in a user is described in detail with reference to FIGS. 2 to 25.
[0059]
[0060] FIG. 2 illustrates images output to a user terminal to determine the degree of dementia of a user according to one example.
[0061] The images below (210 to 240) may be images of an application for determining the degree of dementia. For example, a user of an electronic device (e.g., the electronic device (110) of FIG. 1) may create and distribute an application, and the user may run the application through a user terminal (e.g., the user terminal (120)).
[0062] The first video (210) is the start screen of the application.
[0063] The second image (220) shows the functions supported by the application.
[0064] The third video (230) is an example of content provided to the user. One or more types of content may be provided to the user.
[0065] The fourth image (240) indicates the determined degree of dementia of the user. For example, the degree of attention regarding dementia may be output as a result of the user's test. For example, normal, mild cognitive impairment (MCI), or Alzheimer's disease (AD) determined as the degree of dementia of the user may be output. For example, the degree of attention regarding the user's dementia may be output as a score or probability. In addition to the degree of attention regarding individual diseases, a comprehensive judgment may also be output.
[0066]
[0067] FIG. 3 is a configuration diagram of an electronic device for determining the degree of dementia in a user according to one embodiment.
[0068] The electronic device (300) includes a communication unit (310), a processor (320), and a memory (330). For example, the electronic device (300) may be the electronic device (110) described above with reference to FIG. 1.
[0069] The communication unit (310) is connected to the processor (320) and memory (330) to transmit and receive data. The communication unit (310) may be connected to other external devices to transmit and receive data. In the following, the expression "transmit and receive A" may indicate transmitting and receiving "information or data representing A".
[0070] The communication unit (310) may be implemented as a circuitry within the electronic device (300). For example, the communication unit (310) may include an internal bus and an external bus. As another example, the communication unit (310) may be an element connecting the electronic device (300) and an external device. The communication unit (310) may be an interface. The communication unit (310) may receive data from an external device and transmit the data to the processor (320) and memory (330).
[0071] The processor (320) processes data received by the communication unit (310) and data stored in memory (330). The "processor" may be a data processing device implemented in hardware having a circuit having a physical structure for executing desired operations. For example, the desired operations may include code or instructions included in a program. For example, the data processing device implemented in hardware may include a microprocessor, a central processing unit, a processor core, a multi-core processor, a multiprocessor, an Application-Specific Integrated Circuit (ASIC), or a Field Programmable Gate Array (FPGA).
[0072] The processor (320) executes computer-readable code (e.g., software) stored in memory (e.g., memory (330)) and instructions triggered by the processor (320).
[0073] The memory (330) stores data received by the communication unit (310) and data processed by the processor (320). For example, the memory (330) may store a program (or application, software). The program to be stored may be a set of syntaxes that are coded to determine the degree of dementia of the user and can be executed by the processor (320).
[0074] According to one aspect, the memory (330) may include one or more volatile memory, non-volatile memory and RAM (Random Access Memory), flash memory, hard disk drive and optical disk drive.
[0075] The memory (330) stores a set of instructions (e.g., software) for operating the electronic device (300). The set of instructions for operating the electronic device (300) is executed by the processor (320).
[0076]
[0077] FIG. 4 is a flowchart of a method for determining the degree of dementia in a user according to one embodiment.
[0078] The following operations 410 to 490 can be performed by the electronic device (300) described above with reference to FIG. 3 (hereinafter, electronic device).
[0079] In operation 410, the electronic device may output a first content through a user terminal. For example, the first content may include instructions appearing as voice or text that instruct the user to perform an action in order to obtain a voice reaction from the user. The first content is output to the user terminal, and the user may generate first voice information by recording first voice information as a reaction to the first content using the user terminal. The generated first voice information may be in the form of a data file.
[0080] According to one embodiment, one or more contents for obtaining a voice reaction from a user are provided, and voice information for each of the one or more contents may be generated. For example, the first content may be one of a plurality of voice response contents for generating a voice reaction from the user. For example, voice response contents for generating a voice reaction from the user are described using [Table 1] below.
[0081]
[0082] In operation 420, the electronic device may obtain first reaction information of a user regarding the first content through a user terminal. The first reaction information may include first voice information of the user. For example, the electronic device may receive a voice reaction of a user regarding the first content from a user terminal as first reaction information.
[0083] When multiple voice response contents are created to obtain a user's voice reaction, operations 410 and 420 may be performed repeatedly. By performing operations 410 and 420 repeatedly, the user's voice information regarding the multiple voice response contents is received.
[0084] In operation 430, the electronic device may output a second content through a user terminal. For example, the second content may include instructions appearing as voice or text that instruct the user to perform an action in order to obtain a user's gaze reaction. When the second content is output to the user terminal, first gaze information, which is the movement of the user's gaze appearing as a reaction to the second content, may be generated using the user terminal.
[0085] According to one embodiment, the camera of the user terminal may generate user images including the user's eyes at a preset interval while the second content is being displayed. The generated user images may be in the form of data files. Based on the user images, the user terminal continuously determines the coordinates on the display that the user's eyes are gazing at. For example, the coordinates on the display may be determined in pixel units or by area.
[0086] For example, to continuously determine the coordinates on a display that the user's eye gazes at based on user images, the user terminal may detect the user's first eye within a first user image among the user images received for the content, determine the user's first gaze direction based on the first eye, determine the first coordinates on the display based on the first gaze direction, and associate the first coordinates with a first time stamp of the first user image. For example, the position of the user's eye in three-dimensional space may be determined based on the coordinates of the first eye within the first user image, and the first gaze direction may be determined based on the position of the eye and the direction of the pupil, etc. Software that determines the user's first gaze direction based on user images may be used, and is not limited to the described embodiments. For example, the first coordinates may be pixel coordinates on the display. As another example, the first coordinates may correspond to any one of the regions into which the entire area of the display is divided into a plurality of regions. The first time stamp indicates the time at which the first user image was generated. Actions to continuously determine the coordinates on the display that the user's eye gazes at based on the user images can be repeatedly performed for each user image. As the actions are repeatedly performed, coordinates for the content can be generated, and a timestamp can be associated with each.
[0087] According to one embodiment, the electronic device can control the user terminal to execute an additional application for determining coordinates on a display that the user's gaze is looking at. For example, the additional application can generate user images of a user observing content using the camera of the user terminal, and determine first gaze information regarding coordinates on a display that the user's eyes within the user images are looking at.
[0088] According to one embodiment, one or more contents for obtaining a gaze reaction from a user are provided, and user gaze information for each of the one or more contents may be generated. For example, the second content may be one of a plurality of gaze response contents for generating a user's gaze reaction. For example, gaze response contents for generating a user's gaze reaction are described using [Table 2] below.
[0089]
[0090] According to one embodiment, the gaze response content for generating user images may be pre-made content for measuring saccade, anti-saccade, fixation duration, and eye vergence as user eye movements.
[0091] In operation 440, the electronic device may obtain a user's second reaction information regarding the second content through a user terminal. The second reaction information may include first gaze information regarding coordinates on a display that the user's eyes are gazing at. For example, the electronic device may receive the user's gaze reaction regarding the second content from the user terminal as the second reaction information.
[0092] When multiple gaze response contents are created to obtain a user's gaze reaction, actions 430 and 440 may be performed repeatedly. By performing actions 430 and 440 repeatedly, user gaze information regarding the multiple gaze response contents is received.
[0093] In operation 450, the electronic device may output a third content through a user terminal. For example, the third content may include instructions appearing as voice or text that instruct the user to perform an action in order to obtain a voice-gaze composite reaction of the user. When the third content is output to the user terminal, second voice information, which is the user's voice appearing as a voice-gaze composite reaction to the third content, and second gaze information, which is the movement of the user's gaze, may be generated using the user terminal.
[0094] The description of the process of generating the first voice information corresponding to the first content and the description of the process of generating the first gaze information corresponding to the second content, as described by referring to operations 410 to 440 regarding the process of generating the second voice information and the second gaze information corresponding to the third content, may be applied with similar modifications.
[0095] According to one embodiment, one or more contents for obtaining a voice-gaze combined reaction from a user are provided, and the user's voice information and gaze information for each of the one or more contents may be generated. For example, the third content may be one of a plurality of voice-gaze response contents for generating the user's voice-gaze combined reaction. For example, voice-gaze response contents for generating the user's voice-gaze combined reaction are described using [Table 3] below.
[0096]
[0097] In operation 460, the electronic device may obtain a user's third reaction information regarding the third content through a user terminal. The third reaction information may include the user's second voice information and second gaze information regarding coordinates on the display that the user's eyes are gazing at. For example, the electronic device may receive the user's voice-gaze combined reaction to the third content from the user terminal as the third reaction information. For example, the third reaction information may include a timestamp indicating the time when the voice of the second voice information or the gaze coordinates of the second gaze information were obtained.
[0098] When multiple voice-eye response contents are created to obtain a user's voice-eye combined reaction, actions 450 and 460 may be performed repeatedly. By performing actions 450 and 460 repeatedly, the user's voice information and gaze information regarding the multiple voice-eye response contents are received.
[0099] In operation 470, the electronic device may output a fourth content through a user terminal. The fourth content may include instructions appearing as voice or text that instruct the user to perform an action in order to obtain a drawing reaction from the user. The fourth content is output to the user terminal, and first touch action information for a drawing reaction to the fourth content may be generated at the user terminal.
[0100] For example, the first touch action information can be obtained by the user drawing a picture by directly touching the touch display or tablet of the user terminal with their hand. For example, the first touch action information can be obtained by the user drawing a picture on the touch display or tablet of the user terminal using a digital pen.
[0101] According to one embodiment, the first touch action information may include a first drawing and first pen information. A user terminal may receive a first drawing generated as a drawing reaction to third content via a touch display or tablet as the first touch action information. Additionally, the user terminal may receive first pen information as the first touch action information from a digital pen connected to the user terminal via a wired or wireless connection.
[0102] According to one embodiment, the first pen information may include the characteristic elements of [Table 4] below regarding the time-series movement of the pen.
[0103]
[0104] In [Table 4], time on surface may be the time the digital pen touches the touch display of the user terminal. Time in air may be the time the digital pen does not touch the touch display.
[0105] According to one embodiment, the first pen information may include the characteristic elements of [Table 5] below regarding the operation parameters of the pen.
[0106]
[0107] In [Table 5], total time may be the sum of time on surface and time in air. Velocity on surface may be the first movement speed of the digital pen during the time on surface when the digital pen is in contact with the touch display of the user terminal. Velocity in air may be the second movement speed of the digital pen during the time in air. The second movement speed may be calculated based on time in air and pen-up stroke length. Pressure may be the pen pressure during time on surface. Strokes per minute may be the number of lines drawn during a preset time (e.g., 1 minute). Time not painting may be the time when no drawing is done.
[0108] According to one embodiment, one or more contents for obtaining a picture reaction from a user are provided, and user gaze information for each of the one or more contents may be generated. For example, the fourth content may be one of a plurality of picture response contents for generating a picture reaction from the user. For example, picture response contents for generating a picture reaction from the user are described using [Table 6] below.
[0109]
[0110] In operation 480, the electronic device can obtain user's fourth reaction information regarding the fourth content through a user terminal. The fourth reaction information may include user's first touch action information corresponding to the fourth content.
[0111] When multiple picture response contents are created to obtain a user's picture reaction, actions 470 and 480 may be performed repeatedly. By performing actions 470 and 480 repeatedly, user touch action information for the multiple picture response contents is received.
[0112] According to one embodiment, a content set including a first content, a second content, a third content, and a fourth content may be configured so that the user’s language ability, short-term memory, long-term memory, concentration, calculation ability, conceptual thinking ability, and visual comprehension related to the degree of dementia can be measured through the user’s reaction to each of the contents included in the content set. As the contents of the content set are determined so that various symptoms according to the user’s degree of dementia can be effectively evaluated within a short period of time, a method for determining the degree of dementia can be performed on elderly people or dementia patients with reduced physical strength and concentration. For example, the content set may be composed of 11 contents for which the total time for obtaining user reactions corresponding to the contents is 15 minutes or less.
[0113] In operation 490, the electronic device can determine the degree of dementia of the user by inputting first reaction information, second reaction information, third reaction information, and fourth reaction information into a dementia degree classification model based on an artificial neural network.
[0114] A method for determining a user's degree of dementia by inputting first reaction information, second reaction information, third reaction information, and fourth reaction information into a dementia degree classification model based on an artificial neural network according to one embodiment is described in detail with reference to FIGS. 18 to 25.
[0115] According to one embodiment, the degree of dementia may be any one of normal, mild cognitive impairment (MCI), and Alzheimer's disease (AD). For example, the degree of dementia may include a score or probability for at least one of normal, SCI, MCI, and AD.
[0116] After operation 490 is performed, the electronic device can output the degree of dementia determined through the user terminal. The electronic device provides multiple contents capable of obtaining reactions of various modalities, and by reflecting the reactions of various modalities in the determination of the user's degree of dementia, it can comprehensively evaluate various symptoms that may appear in the user depending on the degree of dementia and accurately determine the user's degree of dementia. A modality may represent a type of input data.
[0117]
[0118] FIGS. 5 to 9 illustrate pre-made content for obtaining a user's voice reaction according to various examples.
[0119] Content (500, 600, 700, 800, 900) can be displayed on a display (e.g., a touch display) of a user terminal (e.g., the user terminal (120) of FIG. 1). Content (500, 600, 700, 800, 900) can convey instructions regarding the content to the user. Content (500, 600, 700, 800, 900) can instruct the user to perform a task via voice recording. For example, instructions can be displayed via text (511, 521, 610, 710, 810, 910). As another example, instructions can be displayed via sound.
[0120] Referring to FIG. 5, a first voice response content (500) provided to a user according to one embodiment is illustrated. The first voice response content (500) may be content that includes instructions (511, 521) that induce the user to read aloud text presented to the user.
[0121] According to one embodiment, the first voice response content (500) may present text for the user to read aloud on the first screen (510). For example, the text for the user to read aloud may be output as voice. For example, the text for the user to read aloud may be output on a separate screen.
[0122] According to one embodiment, the first voice response content (500) can generate a voice reaction from the user on the second screen (520). The user can generate a voice reaction by recording a voice of reading the presented text aloud using a microphone included in the user terminal.
[0123] According to one embodiment, the user may terminate the provision of the first voice response content (500) by touching a submit button after recording a voice corresponding to the first voice response content (500). The user terminal may generate a voice reaction to the voice recorded by the user.
[0124] According to one embodiment, if the user fails to complete the task within a set time, the provision of the first voice response content (500) may be forcibly terminated when the set time has elapsed. The user terminal may generate a voice reaction to the user's voice until the provision of the first voice response content (500) is terminated. The user terminal may transmit the voice reaction to an electronic device.
[0125] The electronic device can measure the user's language ability, short-term memory, and concentration by inducing the user to read aloud the text presented to the user through the first voice response content (500). The voices obtained for each user may have various characteristics, but the higher the degree of dementia, the longer the time taken by the user to perform the task or the more inaccurate the voice reaction may appear.
[0126] Referring to FIG. 6, a second voice response content (600) provided to a user according to one embodiment is illustrated. The second voice response content (600) may be content that includes instructions (610) that induce the user to explain an image (620) presented to the user.
[0127] According to one embodiment, the second voice response content (600) may present an image (620) that the user explains. The user may generate a voice reaction by recording a voice explaining the presented image (620) using a microphone included in the user terminal.
[0128] According to one embodiment, the user may terminate the provision of the second voice response content (600) by touching the submit button (630) after recording a voice corresponding to the second voice response content (600). The user terminal may generate a voice reaction to the voice recorded by the user.
[0129] According to one embodiment, if the user fails to complete the task within a set time, the provision of the second voice response content (600) may be forcibly terminated when the set time has elapsed. The user terminal may generate a voice reaction to the user's voice until the provision of the second voice response content (600) is terminated. The user terminal may transmit the voice reaction to an electronic device.
[0130] The electronic device can evaluate the user's language ability, short-term memory, concentration, and visual comprehension by inducing the user to explain an image (620) presented to the user through the second voice response content (600). The voices obtained for each user may have various characteristics, but as the degree of dementia increases, the interval between speech may be longer, or the explanation may focus on the moment when a problem occurs within the image (620).
[0131] Referring to FIG. 7, a third voice response content (700) provided to a user according to one embodiment is illustrated. The third voice response content (700) may be content that includes instructions (710) that induce the user to associate and speak out words corresponding to the presented text.
[0132] According to one embodiment, a user can generate a voice reaction by using a microphone included in a user terminal to record a voice describing words associated with a presented instruction (710).
[0133] According to one embodiment, the user may terminate the provision of the third voice response content (700) by touching the submit button (720) after recording a voice corresponding to the third voice response content (700). The user terminal may generate a voice reaction to the voice recorded by the user.
[0134] According to one embodiment, the provision of the third voice response content (700) may be forcibly terminated when a set time has elapsed. The user terminal may generate a voice reaction to the user's voice until the provision of the third voice response content (700) is terminated. The user terminal may transmit the voice reaction to an electronic device.
[0135] The electronic device can evaluate the user's language ability, long-term memory, and concentration by inducing the user to associate and speak aloud words corresponding to the text presented to the user through the third voice response content (700). The voices acquired for each user may have various characteristics, but as the degree of dementia increases, characteristics such as longer intervals between speech or a decrease in the number of associated words may appear.
[0136] Referring to FIG. 8, a fourth voice response content (800) provided to a user according to one embodiment is illustrated. The fourth voice response content (800) may be content that includes instructions (810) that induce the user to perform a plurality of arithmetic operations by speaking them aloud.
[0137] According to one embodiment, a user can generate a voice reaction by using a microphone included in the user terminal to record a voice listing numbers calculated in response to a presented instruction (810).
[0138] According to one embodiment, the user may terminate the provision of the fourth voice response content (800) by touching the submit button (820) after recording a voice corresponding to the fourth voice response content (800). The user terminal may generate a voice reaction to the voice recorded by the user.
[0139] According to one embodiment, the provision of the fourth voice response content (800) may be forcibly terminated when a set time has elapsed. The user terminal may generate a voice reaction to the user's voice until the provision of the fourth voice response content (800) is terminated. The user terminal may transmit the voice reaction to an electronic device.
[0140] The electronic device can evaluate the user's short-term memory, concentration, calculation ability, and conceptual thinking ability by inducing the user to perform multiple arithmetic operations aloud through the fourth voice response content (800). The voices obtained for each user may have various characteristics, but as the degree of dementia increases, the speech interval may be longer or the number of arithmetic operations performed may decrease.
[0141] Referring to FIG. 9, a fifth voice response content (900) provided to a user according to one embodiment is illustrated. The fifth voice response content (900) may be content including instructions (910) that induce the user to explain a past anecdote.
[0142] According to one embodiment, a user can generate a voice reaction by using a microphone included in the user terminal to record a voice listing numbers calculated in response to a presented instruction (910).
[0143] According to one embodiment, the user may terminate the provision of the fifth voice response content (900) by touching the submit button (920) after recording a voice corresponding to the fifth voice response content (900). The user terminal may generate a voice reaction to the voice recorded by the user.
[0144] According to one embodiment, the provision of the fifth voice response content (900) may be forcibly terminated when a set time has elapsed. The user terminal may generate a voice reaction to the user's voice until the provision of the fifth voice response content (900) is terminated. The user terminal may transmit the voice reaction to an electronic device.
[0145] The electronic device can evaluate the user's language ability, short-term memory, long-term memory, and concentration by inducing the user to explain past anecdotes through the fifth voice response content (900). The voices obtained for each user may have various characteristics, but as the degree of dementia increases, the speech interval may become longer or the average speech time may decrease.
[0146]
[0147] FIGS. 10 and 11 illustrate pre-made content for obtaining a user's gaze reaction according to various examples.
[0148] Content (1000, 1100) can be displayed on a display (e.g., a touch display) of a user terminal (e.g., the user terminal (120) of FIG. 1). Content (1000, 1100) can convey instructions regarding the content to the user. Content (1000, 1100) can instruct the user to perform a task through eye movements. For example, instructions can be displayed via text (1011, 1110). As another example, instructions can be displayed via sound.
[0149] Referring to FIG. 10, a first gaze response content (1000) provided to a user according to one embodiment is illustrated. The first gaze response content (1000) may be content that includes instructions (1011) that induce the user to gaze at or not gaze at a target area on a display in response to a presented rule.
[0150] According to one embodiment, the first gaze response content (1000) may present a rule for looking at or not looking at a part of the display on the first screen (1010). For example, a rule may be presented for looking at the area containing the icon when a white icon appears in either of the left or right divided areas, and looking at the empty space on the opposite side when a black icon appears.
[0151] According to one embodiment, the first gaze response content (1000) can output a screen in which a target area appears on the second screen (1020) and generate a gaze reaction of the user. A gaze reaction can be generated as the user's eyes are photographed by a camera included in the user terminal, and the coordinates on the display where the eyes gaze are determined based on the image of the captured eyes.
[0152] According to one embodiment, when a set time has elapsed, the provision of the second screen (1020) is forcibly terminated, and a third screen showing a new target area may be displayed. The user terminal may generate a gaze reaction for the user's gaze while the third screen is being displayed. The user terminal may transmit the gaze reaction generated in correspondence with at least one screen to an electronic device.
[0153] The electronic device can evaluate the user's short-term memory, concentration, and conceptual thinking ability by inducing the user to gaze at or not gaze at a target area on the display in response to rules presented to the user through the first gaze response content (1000). The gazes acquired by each user may have various characteristics, but as the degree of dementia increases, the time taken for the user to move their gaze may be longer or the gaze reaction may appear inaccurate.
[0154] Referring to FIG. 11, a second gaze response content (1100) provided to a user according to one embodiment is illustrated. The second gaze response content (1100) may be content that includes instructions (1110) that induce the user to move their gaze in response to a presented image (1120).
[0155] According to one embodiment, the second gaze response content (1100) can output a screen in which a presented image (1120) appears and generate a gaze reaction from the user. For example, the image (1120) may include a maze. A gaze reaction can be generated as the user's eyes are photographed by a camera included in the user terminal, and the coordinates on the display where the eyes are gazing are determined based on the image of the captured eyes.
[0156] According to one embodiment, the provision of the second gaze response content (1100) may be forcibly terminated when a set time has elapsed. The user terminal may generate a gaze reaction to the user's gaze until the provision of the second gaze response content (1100) is terminated. The user terminal may transmit the gaze reaction to an electronic device.
[0157] The second eye-tracking response content (1100) can evaluate the user's short-term memory and concentration by inducing the user to move their gaze in response to the presented image (1120). The gazes obtained from each user may have various characteristics, but as the degree of dementia increases, the time taken for the user to move their gaze may be longer or the gaze reaction may appear inaccurate.
[0158]
[0159] FIGS. 12 to 14 illustrate pre-made content for obtaining a user's voice and gaze reactions according to various examples.
[0160] Content (1200, 1300, 1400) can be displayed on a display (e.g., a touch display) of a user terminal (e.g., the user terminal (120) of FIG. 1). Content (1200, 1300, 1400) can convey instructions regarding the content to the user. Content (1200, 1300, 1400) can instruct the user to perform a task through eye movements. For example, instructions can be displayed via text (1210, 1310, 1410). As another example, instructions can be displayed via sound.
[0161] Referring to FIG. 12, a first voice-eye response content (1200) provided to a user according to one embodiment is illustrated. The first voice-eye response content (1200) may be content that includes instructions (1210) that induce the user to read aloud the presented text.
[0162] According to one embodiment, the first voice-eye response content (1200) outputs a fingerprint (1220) composed of multiple sentences and can generate a voice reaction and an eye reaction for the user. A voice reaction can be generated by recording the user's voice reading the fingerprint aloud through a microphone included in the user terminal. An eye reaction can be generated by photographing the user's eye through a camera included in the user terminal and determining the coordinates on the display where the eye gazes based on the image of the photographed eye.
[0163] According to one embodiment, the user may terminate the provision of the second voice-eye response content (1200) by touching the submit button (1230) after recording a voice corresponding to the first voice-eye response content (1200). The user terminal may generate a voice reaction to the voice recorded by the user.
[0164] According to one embodiment, the provision of the first voice-eye response content (1200) may be forcibly terminated when a set time has elapsed. Until the provision of the first voice-eye response content (1200) is terminated, the user terminal may generate voice reactions and gaze reactions for the user's voice and gaze. The user terminal may transmit the voice reactions and gaze reactions to an electronic device.
[0165] The electronic device can evaluate the user's language ability, concentration, conceptual thinking ability, and visual comprehension by inducing the user to read aloud the text presented to the user through the first voice-eye response content (1200). The voices and gazes acquired for each user may have various characteristics, but as the degree of dementia increases, the reading speed of the text may become slower, and characteristics such as irregular eye movements or a low correlation between voice and gaze may appear.
[0166] Referring to FIG. 13, a second voice-eye response content (1300) provided to a user according to one embodiment is illustrated. The second voice-eye response content (1300) may be content that includes instructions (1310) that induce the user to read the presented text without speaking aloud.
[0167] According to one embodiment, the second voice-gaze response content (1300) outputs a fingerprint (1320) consisting of at least one sentence and can generate a voice reaction and a gaze reaction for the user. A voice reaction can be generated by recording the voice of the user unconsciously reading the fingerprint aloud through a microphone included in the user terminal. A gaze reaction can be generated by photographing the user's eyes through a camera included in the user terminal and determining the coordinates on the display where the eyes are gazing based on the image of the captured eyes.
[0168] According to one embodiment, the user may terminate the provision of the second voice-eye response content (1300) by touching the submit button (1330) after recording a voice corresponding to the second voice-eye response content (1300). The user terminal may generate a voice reaction to the voice recorded by the user.
[0169] According to one embodiment, the provision of the second voice-eye response content (1300) may be forcibly terminated after a set time has elapsed. The user terminal may generate voice reactions and gaze reactions for the user's voice and gaze until the provision of the second voice-eye response content (1300) is terminated. The user terminal may transmit the voice reactions and gaze reactions to an electronic device.
[0170] The electronic device can evaluate the user's language ability, concentration, conceptual thinking ability, and visual comprehension by inducing the user to read text presented to the user through the second voice-eye response content (1300) without speaking. The voices and gazes acquired for each user may have various characteristics, but as the degree of dementia increases, the reading speed of the text may become slower, and characteristics such as irregular eye movements or a low correlation between voice and gaze may appear.
[0171] Referring to FIG. 14, a third voice-eye response content (1400) provided to a user according to one embodiment is illustrated. The third voice-eye response content (1400) may be content including instructions (1410) that induce the user to describe an image (1420) presented to the user.
[0172] According to one embodiment, the third voice-eye response content (1400) may present an image (1420) to be described by the user. The user may generate a voice reaction by recording a voice describing the presented image (1420) using a microphone included in the user terminal. An eye reaction may be generated as the user's eye is photographed by a camera included in the user terminal, and the coordinates on the display where the eye gazes are determined based on the image of the photographed eye.
[0173] According to one embodiment, the user may terminate the provision of the third voice-eye response content (1400) by touching the submit button (1430) after recording a voice corresponding to the third voice-eye response content (1400). The user terminal may generate a voice reaction to the voice recorded by the user.
[0174] According to one embodiment, if the user fails to complete the task within a set time, the provision of the third voice-eye response content (1400) may be forcibly terminated when the set time has elapsed. The user terminal may generate voice reactions and gaze reactions for the user's voice and gaze until the provision of the third voice-eye response content (1400) is terminated. The user terminal may transmit the voice reactions and gaze reactions to an electronic device.
[0175] The electronic device can evaluate the user's language ability, short-term memory, concentration, and visual comprehension by inducing the user to describe an image (1420) presented to the user through the third voice-eye response content (1400). The voices and gazes acquired for each user may have various characteristics, but as the degree of dementia increases, characteristics such as longer intervals between speech, prolonged gazing at the area where a problem situation occurs within the image, or a low correlation between voice and gaze may appear.
[0176]
[0177] FIGS. 15 to 17 illustrate pre-made content for obtaining a user's picture reaction according to various examples.
[0178] Content (1500, 1600, 1700) can be displayed on a display (e.g., a touch display) of a user terminal (e.g., the user terminal (120) of FIG. 1). Content (1500, 1600, 1700) can convey instructions regarding the content to the user. Content (1500, 1600, 1700) can instruct the user to perform a task through a touch action. For example, instructions can be displayed via text (1510, 1610, 1710). As another example, instructions can be displayed via sound.
[0179] Referring to FIG. 15, a first picture response content (1500) provided to a user according to one embodiment is illustrated. The first picture response content (1500) may be content that induces the user to connect multiple points on a display according to rules presented to the user.
[0180] According to one embodiment, the first picture response content (1500) may be content designed to induce a user to perform a task of drawing lines connecting numbers in order. The user may use a digital pen to draw lines connecting numbers in the order indicated by text (1510) through a display or a tablet connected to a user terminal. For example, the user may draw a picture (1530) connecting numbers in order. Depending on the task of the first picture response content (1500), the user may draw a picture connecting numbers in ascending or descending order. As another example, the user may draw a picture connecting numbers of different colors alternately in order among duplicate numbers of different colors.
[0181] According to one embodiment, if the user connects the numbers in the wrong order, the first picture response content (1500) may provide feedback, such as by making an alarm sound.
[0182] According to one embodiment, the user can terminate the provision of the first picture response content (1500) by touching the submit button (1520) after completing a picture in which lines connecting all the numbers are drawn. The user terminal can generate a picture reaction for the picture (1530) drawn by the user.
[0183] According to one embodiment, if the user fails to complete the task within a set time, the provision of the first picture response content (1500) may be forcibly terminated when the set time has elapsed. The user terminal may generate a picture reaction for an unfinished picture drawn by the user. The user terminal may transmit the picture reaction to an electronic device.
[0184] According to one embodiment, a user terminal can generate execution information for the first picture response content (1500). The execution information for the first picture response content (1500) may include data in the form of a scalar or vector for the lines of the picture. The execution information for the first picture response content (1500) may include at least one of information regarding whether the user has completed all tasks, the total time taken to perform the tasks, and to what level (or point) the user has performed the tasks. The user terminal can transmit the execution information for the first picture response content (1500) to an electronic device.
[0185] The electronic device can evaluate the user's concentration and conceptual thinking ability by inducing the user to connect multiple points on the display according to rules presented to the user through the first picture response content (1500). The pictures drawn by each user may have various characteristics, but the higher the degree of dementia, the longer the time taken by the user to perform the task, the longer the curvature of the lines and large patterns may appear, the hesitating patterns may appear, or the level of performance may appear low.
[0186] Referring to FIG. 16, a second picture response content (1600) provided to a user according to one embodiment is illustrated. The second picture response content (1600) may be content that includes instructions to induce the user to draw a pattern corresponding to a pattern presented to the user.
[0187] According to one embodiment, the second picture response content (1600) may be content intended to induce the user to perform a task of drawing a symbol (or symbol) corresponding to a specific number by looking at a given symbol table. The second picture response content (1600) may pre-output examples of a pre-set reference symbol table and pictures. For example, the second picture response content (1600) may pre-output a table divided into multiple areas for the user to draw each symbol.
[0188] According to one embodiment, a user can use a digital pen to draw pictures of symbols corresponding to each of the given numbers on a display or a tablet connected to a user terminal. For example, the user can draw a symbol corresponding to the number 7 in the area (1630). For example, after completing the drawings of symbols corresponding to all numbers, the user can end the provision of the second picture response content (1600) by touching the submit button (1620). The user terminal can generate a picture reaction for the drawings drawn by the user. The user terminal can transmit the picture reaction to an electronic device (e.g., electronic device (300)).
[0189] The electronic device can evaluate the user's short-term memory, concentration, conceptual thinking ability, and visual comprehension by inducing the user to draw a pattern corresponding to a pattern presented to the user through the second picture response content (1600). Although the drawings created by each user may have various characteristics, the performance score for correctly drawn symbols may appear lower as the degree of dementia increases.
[0190] Referring to FIG. 17, a third picture response content (1700) provided to a user according to one embodiment is illustrated. The third picture response content (1700) may be one of the contents that includes instructions to induce the user to draw a shape corresponding to a rule presented to the user.
[0191] According to one embodiment, the third drawing response content (1700) may be content that induces the task of drawing a clock with hands pointing to a specific time (e.g., 11:10). For example, the third drawing response content (1700) may pre-output the outer shape of the clock (e.g., a circle).
[0192] According to one embodiment, a user can use a digital pen to draw a picture (1730) of an analog clock pointing to a specific time on a display or a tablet connected to a user terminal. For example, after completing the drawing, the user can end the provision of the third picture response content (1700) by touching a submit button (1720). The user terminal can generate a picture reaction for the drawing (1730) drawn by the user. The user terminal can transmit the picture reaction to an electronic device.
[0193] According to one embodiment, the electronic device may determine execution information for a picture (hereinafter referred to as the third picture) created by a user as a task of the third picture response content (1700). The electronic device may normalize the third picture based on the third picture information for the third picture. The electronic device may set the area where the actual picture is drawn within the entire area of the third picture as the region of interest (ROI). The electronic device may normalize the picture by adjusting the ROI to a normalized size.
[0194] According to one embodiment, normalized information, in which the figure is normalized information, may be further generated. The size and / or ratio of the adjusted ROI may be generated as normalized information.
[0195] According to one embodiment, the electronic device may input a normalized figure to a pre-updated neural network. For example, the pre-updated neural network may be a neural network based on a CNN and / or a deep neural network (DNN). According to one embodiment, normalization information may be further input to the pre-updated neural network.
[0196] According to one embodiment, the electronic device may calculate a balance score for a normalized picture as execution information for the third picture response content (1700). For example, the calculated balance score may be a value within a preset range (e.g., 0 to 5). The electronic device may calculate the balance score by separating the elements of the picture and determining the degree of balance between the separated elements. Segmentation of the picture may be performed first to determine multiple elements. For example, a first element representing the arrangement of numbers indicating the hours of a clock and a second element representing the arrangement of hands may be determined, and a balance score may be calculated based on the balance between the first element and the second element. For example, the balance score may be higher as the symmetry of the arranged hours increases. As another example, the balance score may be higher as the size of the first element and the size of the second element are appropriate.
[0197] The electronic device can evaluate the user's concentration, conceptual thinking ability, and visual comprehension by inducing the user to draw a shape corresponding to a rule presented to the user through the third picture response content (1700). The drawings created by each user may have various characteristics, but the higher the degree of dementia, the more unclear the shape of the clock and the time indicated by the clock may appear.
[0198]
[0199] FIG. 18 is a flowchart of a method for determining the degree of dementia of a user based on a plurality of reaction information according to one embodiment.
[0200] The following operations 1810 to 1890 may be performed by the electronic device (300) (hereinafter, electronic device) described above with reference to FIG. 3. For example, the operation 490 described above with reference to FIG. 4 may include operations 1810 to 1890.
[0201] In operation 1810, the electronic device may acquire a first biomarker based on first reaction information. For example, the first biomarker may be a first spectrogram image visualizing at least one characteristic of the first voice information. The electronic device may generate a spectrogram image for the voice through a librosa tool. The first spectrogram image may be a mel-spectrogram image.
[0202] According to one embodiment, the process of generating a first biomarker may include a process of preprocessing first reaction information. The process of preprocessing first reaction information may include a process of removing noise and gaps from first voice information, a process of converting a data file for the first voice information into a spectrum image, and a process of adjusting the size of the spectrum image.
[0203] For example, when multiple voice response contents are created to obtain a user's voice reaction, spectrogram images corresponding to the user's voice information for each of the multiple voice response contents may be generated. Spectrogram images are described in detail below with reference to FIGS. 19 to 20c.
[0204] In operation 1820, the electronic device can obtain first features for first reaction information by inputting a first biomarker into a first model. For example, the first model may include a deep neural network (DNN) model that has been pre-updated to generate first features for first voice information by processing a first spectrogram image. A DNN model for processing the first biomarker is described in detail below with reference to FIG. 21.
[0205] In operation 1830, the electronic device may acquire a second biomarker based on second reaction information. For example, the second biomarker may be first time series data acquired for the user's gaze of the first gaze information. For example, the first time series data may include coordinates where the user's gaze is directed for each generated time stamp.
[0206] According to one embodiment, the process of generating a second biomarker may include a process of preprocessing second reaction information. The process of preprocessing second reaction information may include a process of removing or replacing missing values and outliers from first gaze information and a process of converting the first gaze information into time-series data by visualizing it.
[0207] For example, if multiple gaze response contents are created to acquire a user's gaze reaction, time-series data corresponding to the user's gaze information for each of the multiple gaze response contents can be generated.
[0208] In operation 1840, the electronic device can obtain second features for second reaction information by inputting a second biomarker into a second model. For example, the second model may include a convolutional neural network (CNN) model that is pre-updated to generate second features for first gaze information by processing first time series data. For example, the second model may include a CNN-LSTM combined model that combines a CNN model and a long short-term memory (LSTM) model. A CNN-LSTM combined model for processing the second biomarker is described in detail below with reference to FIG. 22.
[0209] In operation 1850, the electronic device may acquire a third biomarker based on third reaction information. For example, the third biomarker may be a second spectrogram image visualizing at least one characteristic of the second voice information and second time-series data acquired for the user's gaze of the second gaze information. For example, the third reaction information may include a timestamp indicating the time when the voice of the second voice information or the gaze coordinates of the second gaze information were acquired. For example, the second spectrogram image may be a mel-spectrogram image. For example, the second time-series data may include the coordinates where the user's gaze is directed for each generated timestamp. For example, the second spectrogram image and the second time-series data may be generated for the same timestamp interval.
[0210] According to one embodiment, the process of generating a third biomarker may include a process of preprocessing third reaction information. The process of preprocessing third reaction information may include a process of removing noise and gaps from second voice information, a process of converting a data file for second voice information into a spectrum image, a process of adjusting the size of the spectrum image, a process of removing or replacing missing values and outliers from second gaze information, and a process of converting second gaze information into time-series data by visualizing it.
[0211] For example, when multiple voice-gaze response contents are produced to obtain a user's voice-gaze combined reaction, spectrogram images corresponding to the user's voice information and time-series data corresponding to the gaze information can be generated for each of the multiple voice-gaze response contents.
[0212] In operation 1860, the electronic device can acquire third features for third reaction information by inputting a third biomarker into a third model. For example, the third model may include a pre-updated association acquisition model to generate third features for association between second speech information and second gaze information by processing a second spectrogram image and second time series data. For example, the association acquisition model may be a model that utilizes an attention algorithm. For example, the third features may represent speech-gaze association or speech-gaze association features between the second speech information and second gaze information at different time intervals. An association acquisition model for processing the third biomarker is described in detail below with reference to FIG. 23.
[0213] Third characteristics regarding the correlation between second voice information and second gaze information may appear differently depending on the degree of dementia of the user. For example, users in the normal group may show a high correlation between the second voice information and the second gaze information as the speech rate in the second voice information and the gaze movement rate in the second gaze information change similarly. For example, users in the MCI group may show a low correlation between the second voice information and the second gaze information as their gaze moves frequently, resulting in a high gaze movement rate, while the speech rate in the voice information appears normal or slow. The electronic device can more accurately determine the degree of dementia of the user by not only examining the user's voice reaction and gaze reaction separately using the user's first or second content, but also by acquiring the correlation between the user's voice reaction and gaze reaction using third content.
[0214] According to one embodiment, the third features may be results output in the complete form of the attention algorithm. For example, the association acquisition model may convert the input second spectrogram image and second time series data into sets of embedding vectors by time stamp, respectively, and input the set of embedding vectors for the second spectrogram image and the set of embedding vectors for the second time series data into the attention algorithm using a query or a key. The attention algorithm may determine the association between the query and the key and generate the results of applying softmax and weights to the determined association as the third features. For example, the third features generated by the attention algorithm may be a matrix representing the association between the second spectrogram image and the second time series data, or a feature vector of new time series data based on the second spectrogram image and the second time series data.
[0215] In operation 1870, the electronic device may acquire a fourth biomarker based on fourth reaction information. The fourth biomarker may include at least one of first picture information for a first picture generated by a user as a task, third time-series data acquired by first touch action information, and first structured data associated with the first picture.
[0216] According to one embodiment, the first picture information for the first picture may include information such as information about the overall size of the first picture and pixel information.
[0217] According to one embodiment, the third time series data obtained by the first touch action on the first picture may include the characteristic elements of [Table 4] for the time series movement of the pen.
[0218] According to one embodiment, the first structured data associated with the first figure may include at least one of performance information for a task and first parameter information. The first structured data may include data in the form of a scalar or a vector. The performance information for the fourth task may include at least one of a performed level, a performance score, or a balance score. The second parameter information may include the characteristic elements of [Table 5] for the operation parameters of the pen.
[0219] According to one embodiment, the process of generating a fourth biomarker may include a process of preprocessing fourth reaction information. The process of preprocessing fourth reaction information may include a process of obtaining first figure information by removing unnecessary regions from the first figure and normalizing it, a process of performing differencing on data representing the time-series movement of the pen, a process of converting the data representing the time-series movement of the pen into third time-series data, a process of normalizing variables appearing in the task performance information and first parameter information, and a process of removing variables among the variables that exhibit multicollinearity.
[0220] For example, when multiple picture response contents are created to obtain a user's picture reaction, picture information corresponding to the user's picture for each of the multiple picture response contents, time-series data corresponding to the time-series movement of the pen, or execution information for the task and structured data corresponding to variables appearing in the first parameter information may be generated.
[0221] In operation 1880, the electronic device can obtain fourth features for fourth reaction information by inputting a fourth biomarker into a fourth model. For example, the fourth model may include at least one of a pre-updated CNN model to generate first sub-features for first touch action information by processing first picture information, an RNN model (e.g., an LSTM model) to generate second sub-features for first touch action information by processing third time-series data, or a pre-updated fully connected layer to generate third sub-features for first touch action information by processing first structured data. The fourth model may be determined in correspondence with the type of information or data included in the fourth biomarker. The fourth features may be obtained based on at least one of the first sub-features, second sub-features, or third sub-features. The CNN model, RNN model, and fully connected layer for processing the fourth biomarker are described in detail below with reference to FIG. 24.
[0222] In operation 1890, the electronic device can determine the degree of dementia of the user by inputting the first features, the second features, the third features, and the fourth features into the fifth model. A method for determining the degree of dementia of the user by processing the first features, the second features, the third features, and the fourth features generated based on the first reaction information, the second reaction information, the third reaction information, and the fourth reaction information is described in detail below with reference to FIGS. 25 and 26.
[0223] The electronic device can acquire the characteristics of various modalities by evaluating symptoms that may appear in the user as the degree of dementia progresses from various perspectives, and can accurately determine the degree of dementia of the user based on models generated to effectively process each of the characteristics of various modalities.
[0224]
[0225] Figure 19 illustrates a spectrogram image generated for speech according to one example.
[0226] According to one embodiment, an electronic device (e.g., the electronic device (300) of FIG. 3) can generate a spectrogram image (1900) for speech through a Librosa tool. The horizontal axis of the spectrogram image (1900) may be the time axis, and the vertical axis may be the frequency axis. The spectrogram image (1900) represents the difference in amplitude according to changes in the time axis and the frequency axis as a difference in print density / display color. The display color at a corresponding location may be determined based on the magnitude of the changing amplitude difference. For example, a legend (1910) of the display color for the magnitude of the amplitude difference may be output along with the spectrogram image (1900). To display the determined color, the values of the R, G, and B channels of the pixel at the corresponding coordinate may be determined.
[0227] Multiple spectrogram images may be generated for each of the multiple voices. For example, a first spectrogram image may be generated for the first voice information, and a second spectrogram image may be generated for the second voice information. The scales of the time axis and the frequency axis of the spectrogram images may vary depending on the total time of the individual voice, but the sizes of the generated spectrogram images may be the same. For example, the size of the first spectrogram image and the size of the second spectrogram image may be 100x100, which is the same.
[0228]
[0229] FIGS. 20a to 20c illustrate examples of spectrogram images generated for voices acquired by users with different degrees of dementia.
[0230] According to one embodiment, the electronic device can generate spectrogram images (2010, 2020, 2030) for the voices of users as described above with reference to FIG. 19. The horizontal axis of the spectrogram images (2010, 2020, 2030) is the time axis and the vertical axis is the frequency axis, and the spectrogram images (2010, 2020, 2030) can represent differences in amplitude according to changes in the time axis and the frequency axis as differences in print density / display color. The spectrogram images may appear differently depending on the degree of dementia of the user.
[0231] Referring to FIG. 20a, a spectrogram image (2010) generated for a voice acquired by a user of the normal group is shown. In the spectrogram image (2010) generated for a voice acquired by a user of the normal group, features such as a high frequency of speech, regular characteristics, and large amplitude in the low-frequency band may appear.
[0232] Referring to FIG. 20b, a spectrogram image (2020) generated for a voice acquired by a user of the MCI group is shown. In the spectrogram image (2020) generated for a voice acquired by a user of the MCI group, features such as a high frequency of speech but irregular characteristics and small amplitude in the low-frequency band may appear.
[0233] Referring to FIG. 20c, a spectrogram image (2030) generated for a voice acquired by a user of the AD group is shown. In the spectrogram image (2030) generated for a voice acquired by a user of the AD group, features such as a low frequency of speech, irregular characteristics, and high amplitude in the low-frequency band may appear.
[0234] According to one embodiment, an electronic device can acquire a user's voice (e.g., first reaction information of FIG. 4) using content (e.g., first content of FIG. 4), generate a spectrogram image (e.g., first biomarker of FIG. 18) for the user's voice, acquire features for the spectrogram image (e.g., first features of FIG. 18), and determine the degree of dementia of the user based on the features for the spectrogram image. The electronic device can determine the degree of dementia based on the user's voice by acquiring features of spectrogram images (2010, 2020, 2030) that appear differently depending on the degree of dementia of the user.
[0235] According to one embodiment, an electronic device may acquire a user's voice and gaze (e.g., third reaction information in FIG. 4) using content (e.g., third content in FIG. 4), generate spectrogram images and time-series data (e.g., third biomarker in FIG. 18) for the user's voice and gaze, acquire an association between the voice and gaze (e.g., third features in FIG. 18) based on the spectrogram images and time-series data, and determine the degree of dementia of the user based on the association between the voice and gaze. For example, the association between the voice and gaze may appear high as the user's gaze also moves at a constant speed in the high frequency of speech in the spectrogram images (2010) generated for the voice acquired by a normal group user. For example, the association between the voice and gaze may appear low as the user's gaze moves fast or slow in the high frequency of speech in the spectrogram images (2020, 2030) generated for the voice acquired by an MCI or AD group user. The electronic device can determine the degree of dementia based on the user's voice and gaze by processing spectrogram images (2010, 2020, 2030) acquired for the user's voice together with time-series data of gaze acquired simultaneously with the voice to obtain the correlation between the voice and the gaze.
[0236]
[0237] FIG. 21 illustrates a model for processing a user's voice biomarker according to one example.
[0238] According to one embodiment, a DNN model (2100) for processing a user's voice biomarker (e.g., the first biomarker of FIG. 18) may include an input layer (2110), first hidden layers (2120), a second hidden layer (2130), a third hidden layer (2140), and an output layer (2150). The number of hidden layers included in the DNN model (2100) is not limited to the illustrated example.
[0239] A spectrogram image representing a voice reaction can be input to a DNN model (2100) through an input layer (2110). The DNN model (2100) can output the user's voice features (e.g., the first features of FIG. 18) as outputs for the input of voice biomarkers. As features related to the degree of dementia are learned in the DNN model (2100), voice features related to the degree of dementia can be extracted from the user's voice biomarkers.
[0240] According to one embodiment, the output layer (2150) may include a flatten layer. As the size of voice features obtained from voice biomarkers is adjusted through the flatten layer, the voice features can be concatenated with features of other modalities.
[0241]
[0242] FIG. 22 illustrates a model for processing a user's gaze biomarker according to one example.
[0243] According to one embodiment, a CNN-LSTM combined model (2200) (e.g., the second model of FIG. 18) for processing a user's gaze biomarker (e.g., the second biomarker of FIG. 18) may include an input layer (2210), a first convolution layer block (2220), a second convolution layer block (2230), a first LSTM layer (2240), a second LSTM layer (2250), a third LSTM layer (2260), and a flattening layer (2270). The structure of the model for processing the user's gaze biomarker is not limited to the illustrated example, and the model for processing the gaze biomarker may include only a CNN model or a model in which the CNN model is combined with a model other than an LSTM model.
[0244] According to one embodiment, a convolution layer block (2220, 2230) may include one or more convolution layers, a batch normalization layer, and a pooling layer.
[0245] Time series data representing gaze reactions can be input to a CNN-LSTM combined model (2200) through an input layer (2210). When the time series data is processed using the CNN-LSTM combined model (2200), short-term patterns appearing in the time series data are learned in one or more convolutional layer blocks, and long-term patterns appearing in the time series data are learned in one or more LSTM layers, so that complex patterns appearing in the gaze reactions can be extracted as gaze features (e.g., the second features of FIG. 18).
[0246] According to one embodiment, the CNN-LSTM combined model (2200) may include a flattening layer. As the size of gaze features obtained from gaze biomarkers is adjusted through the flattening layer (2270), the gaze features can be concatenated with features of other modalities.
[0247]
[0248] FIG. 23 illustrates a model for processing a user's voice and gaze biomarkers according to one example.
[0249] According to one embodiment, a model (2300) for processing a user’s voice-gaze combined reaction (e.g., the third model of FIG. 18) may include an association acquisition model (2300) for processing voice biomarkers (e.g., spectrogram images) and gaze biomarkers (e.g., time series data) obtained by the voice-gaze combined reaction.
[0250] Spectrogram images representing the voice reaction and time series data representing the gaze reaction among voice-gaze composite reactions can be input to the attention algorithm (2320) through the input layer (2310). Voice-gaze association features related to voice-gaze association or the degree of dementia can be extracted by the attention algorithm (2320).
[0251] According to one embodiment, the association acquisition model (2300) may include a flattening layer (2330). As the size of the voice-gaze association or voice-gaze association features is adjusted through the flattening layer (2330), the voice-gaze association or voice-gaze association features may be concatenated with features of other modalities.
[0252]
[0253] FIG. 24 illustrates a model for processing a user's picture biomarker according to one example.
[0254] According to one embodiment, a model (2400) for processing a user's picture reaction (e.g., the fourth model of FIG. 18) may include a neural network model comprising a CNN model for processing picture information about a picture obtained by the picture reaction, an RNN model for processing time-series data obtained by touch action information, and a fully connected layer (2432) for processing structured data associated with the picture. For example, the RNN model for processing time-series data may be an LSTM model.
[0255] The CNN model may include a first input layer (2411), a first convolution layer block (2412), and a first flattening layer (2413). Picture information may be input to the CNN model through the first input layer (2411). As spatial patterns of pictures related to the degree of dementia are learned by the CNN model, features related to the degree of dementia (e.g., the first sub-features of FIG. 18) may be extracted from the user's picture information.
[0256] The LSTM model may include a second input layer (2421), a first LSTM layer (2422), a second LSTM layer (2423), a third LSTM layer (2424), and a second flattening layer (2425). Time-series data regarding a user's touch action may be input to the LSTM model through the second input layer (2421). As the temporal patterns of the time-series data related to the degree of dementia are learned by the LSTM model, features related to the degree of dementia (e.g., the second sub-features of FIG. 18) may be extracted from the time-series data regarding the user's touch action.
[0257] A neural network model including a fully connected layer (2432) may include a third input layer (2431), a fully connected layer (2432), and a third flattening layer (2433). Structured data associated with a drawing may be input to the neural network model through the third input layer (2431). As patterns of structured data related to the degree of dementia are learned by the neural network model, features related to the degree of dementia (e.g., the third sub-features of FIG. 18) may be extracted from the structured data associated with the user's drawing.
[0258] According to one embodiment, each of the neural network models including a CNN model, an LSTM model, and a fully connected layer may include a flattening layer (2413, 2425, 2433). As the size of various features related to the picture reaction is adjusted through the flattening layer (2413, 2425, 2433), the various features may be concatenated with features of other modalities.
[0259]
[0260] FIG. 25 is a flowchart of a method for determining the degree of dementia of a user based on multimodal data according to one embodiment, and FIG. 26 illustrates a model for processing biomarkers of various modalities according to one example.
[0261] The following operations 2510 and 2520 may be performed by the electronic device (300) (hereinafter, electronic device) described above with reference to FIG. 3. For example, the operation 1890 described above with reference to FIG. 18 may include operations 2510 and 2520.
[0262] According to one embodiment, as illustrated in FIG. 26, a fifth model (2600) processing first features, second features, third features, and fourth features of various modalities may include a connection layer (2610), a first fully connected layer (2620), a dropout layer (2630), a second fully connected layer (2640), and a third fully connected layer (2650). The fifth model (2600) may be a dementia degree determination model that determines the degree of dementia of a user by processing the features of various modalities.
[0263] In operation 2510, the electronic device can generate multimodal data in which the first features, the second features, the third features, and the fourth features are connected. Each of the first features, the second features, the third features, and the fourth features may be flattened to the same dimension through a flattening layer. As the features of various modalities flattened in the connection layer (2610) are connected, multimodal data can be generated. When multimodal data in which features of various modalities are concatenated is generated, the relationships between the features of various modalities can be reflected in the generation of output data.
[0264] In operation 2520, the electronic device can determine the degree of dementia of a user by inputting multimodal data into a pre-updated DNN model. The DNN model may include a first fully connected layer (2620), a dropout layer (2630), a second fully connected layer (2640), and a third fully connected layer (2650). As patterns related to the degree of dementia are learned from the features of various modalities or the relationships between features by the DNN model, the degree of dementia of a user can be determined from the features of various modalities.
[0265] According to one embodiment, the fifth model may be a model capable of determining the degree of dementia of a user even when only some of the first features, second features, third features, and fourth features are acquired. The electronic device may also determine the degree of dementia of a user even when a user's reaction to some of the prepared contents is acquired.
[0266] The DNN model can output any one of a plurality of preset degrees of dementia. For example, the plurality of preset degrees of dementia may include determined normal, mild cognitive impairment (MCI), and Alzheimer's disease (AD).
[0267]
[0268] The embodiments described above may be implemented as hardware components, software components, and / or combinations of hardware and software components. For example, the devices, methods, and components described in the embodiments may be implemented using a general-purpose computer or a special-purpose computer, such as, for example, a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable gate array (FPGA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing and responding to instructions. The processing unit may execute an operating system (OS) and software applications executed on said operating system. Additionally, the processing unit may access, store, manipulate, process, and generate data in response to the execution of the software. For ease of understanding, the processing unit may be described as being used as a single unit, but those skilled in the art will understand that the processing unit may include multiple processing elements and / or multiple types of processing elements. For example, the processing unit may include multiple processors or one processor and one controller. In addition, other processing configurations, such as parallel processors, are also possible.
[0269] Software may include computer programs, code, instructions, or a combination of one or more of these, and may configure a processing unit to operate as desired or instruct the processing unit independently or collectively. Software and / or data may be stored on any type of machine, component, physical device, virtual equipment, computer storage medium, or device so as to be interpreted by the processing unit or to provide instructions or data to the processing unit. Software may be distributed over networked computer systems and stored or executed in a distributed manner. Software and data may be stored on computer-readable recording media.
[0270] The method according to the embodiment may be implemented in the form of program instructions that can be executed through various computer means and recorded on a computer-readable medium. The computer-readable medium may store program instructions, data files, data structures, etc., either individually or in combination, and the program instructions recorded on the medium may be those specifically designed and configured for the embodiment or those known and available to those skilled in the art of computer software. Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes; optical recording media such as CD-ROMs and DVDs; magneto-optical media such as floptical disks; and hardware devices specifically configured to store and execute program instructions, such as ROM, RAM, and flash memory. Examples of program instructions include machine code, such as that generated by a compiler, as well as high-level language code that can be executed by a computer using an interpreter, etc.
[0271] The hardware device described above may be configured to operate as one or more software modules to perform the operation of the embodiment, and vice versa.
[0272] Although the embodiments described above have been explained with reference to limited drawings, those skilled in the art can apply various technical modifications and variations based thereon. For example, appropriate results can be achieved even if the described techniques are performed in a different order than described, and / or if the components of the described system, structure, device, circuit, etc. are combined or assembled in a form different from described, or replaced or substituted by other components or equivalents.
[0273] Therefore, other implementations, other embodiments, and equivalents to the claims also fall within the scope of the claims set forth below.
Claims
1. A method for determining the degree of dementia in a user, performed by an electronic device, The operation of outputting a first content that has been pre-produced to determine the degree of dementia of a user through a user terminal; An operation of obtaining first reaction information of the user regarding the first content through the user terminal - the first reaction information includes first voice information of the user corresponding to the first content -; The operation of outputting a second content pre-produced to determine the degree of dementia of the user through the user terminal; An operation of obtaining second reaction information of the user regarding the second content through the user terminal - the second reaction information includes first gaze information regarding coordinates on a display where the user's eyes gaze in correspondence with the second content -; The operation of outputting a third content pre-produced to determine the degree of dementia of the user through the user terminal; An operation of obtaining third reaction information of the user regarding the third content through the user terminal - the third reaction information includes second voice information of the user corresponding to the third content and second gaze information regarding coordinates on the display where the user's eyes are gazing -; The operation of outputting a pre-made fourth content to determine the degree of dementia of the user through the user terminal; An operation of obtaining the user's fourth reaction information regarding the fourth content through the user terminal - the fourth reaction information includes the user's first touch action information corresponding to the fourth content -; and An operation to determine the degree of dementia of the user by inputting the first reaction information, the second reaction information, the third reaction information, and the fourth reaction information into a dementia degree classification model based on an artificial neural network. including, Method for determining the degree of dementia.
2. In Paragraph 1, The above-mentioned first content is content including instructions to induce reading a presented text aloud, content including instructions to induce explaining a presented image, content including instructions to induce associating and saying aloud words corresponding to the presented text, content including instructions to induce performing multiple arithmetic operations aloud, or content including instructions to induce explaining a past anecdote. Method for determining the degree of dementia.
3. In Paragraph 1, The second content above is one of the following: content including instructions to induce gaze at or not gaze at a target area on the display in response to a presented rule, or content including instructions to induce gaze movement in response to a presented image. Method for determining the degree of dementia.
4. In Paragraph 1, The above third content is one of the following: content including instructions to induce reading the presented text aloud, content including instructions to induce reading the presented text without aloud, or content including instructions to induce describing the presented image. Method for determining the degree of dementia.
5. In Paragraph 1, The above-mentioned fourth content is one of the following: content including instructions to guide connecting multiple points on the display according to a presented rule, content including instructions to guide drawing a pattern corresponding to a presented pattern, or content including instructions to guide drawing a shape corresponding to a presented rule. Method for determining the degree of dementia.
6. In Paragraph 1, The operation of determining the degree of dementia of the user by inputting the first reaction information, the second reaction information, the third reaction information, and the fourth reaction information into the dementia degree classification model is, The operation of obtaining a first biomarker based on the above first reaction information; The operation of obtaining a preset number of first features for the first reaction information by inputting the first biomarker into a previously updated first model; The operation of obtaining a second biomarker based on the above second reaction information; The operation of obtaining a preset number of second features for the second reaction information by inputting the second biomarker into a previously updated second model; The operation of obtaining a third biomarker based on the above third reaction information; The operation of obtaining a preset number of third features for the third reaction information by inputting the third biomarker into a previously updated third model; The operation of obtaining a fourth biomarker based on the above fourth reaction information; The operation of obtaining a preset number of fourth features for the fourth reaction information by inputting the above-mentioned fourth biomarker into a pre-updated fourth model; and An operation to determine the degree of dementia of the user by inputting the above first features, the above second features, the above third features, and the above fourth features into a pre-updated fifth model. including, Method for determining the degree of dementia.
7. In Paragraph 6, The first biomarker is a first spectrogram image that visualizes at least one characteristic of the first voice information, and The first model comprises a deep neural network (DNN) model that is pre-updated to generate the first features for the first voice information by processing the first spectrogram image, Method for determining the degree of dementia.
8. In Paragraph 6, The second biomarker above is first time-series data obtained for the gaze of the user of the first gaze information, and The second model comprises a pre-updated convolutional neural network (CNN) model for generating the second features for the first gaze information by processing the first time series data, Method for determining the degree of dementia.
9. In Paragraph 6, The third biomarker is a second spectrogram image visualizing at least one characteristic of the second voice information and second time-series data obtained for the user's gaze of the second gaze information, and The third model comprises a pre-updated association acquisition model for generating third features regarding the association between the second voice information and the second gaze information by processing the second spectrogram image and the second time series data. Method for determining the degree of dementia.
10. In Paragraph 6, The above-mentioned fourth biomarker includes at least one of first picture information for a first picture represented by the first touch action information, third time-series data obtained for the first touch action information, and first structured data associated with the first picture, and The above-mentioned fourth model includes at least one of a pre-updated CNN model for generating first sub-features for the first touch action information by processing the first picture information, a pre-updated recurrent neural network (RNN) model for generating second sub-features for the first touch action information by processing the third time series data, or a pre-updated fully connected layer for generating third sub-features for the first touch action information by processing the first structured data. The above fourth features are obtained based on at least one of the above first sub-features, the above second sub-features, or the above third sub-features, Method for determining the degree of dementia.
11. In Paragraph 6, The operation of determining the degree of dementia of the user by inputting the first features, the second features, the third features, and the fourth features into the fifth model is, The operation of generating multi-modal data in which the first features, the second features, the third features, and the fourth features are concatenated - the first features, the second features, the third features, and the fourth features are flattened to the same dimension -; and The operation of determining the degree of dementia of the user by inputting the above multimodal data into a pre-updated DNN model including, Method for determining the degree of dementia.
12. In Paragraph 1, The above first content, the above second content, the above third content and the above fourth content are, Including instructions appearing as voice or text that instruct the user to perform an action, Method for determining the degree of dementia.
13. In Paragraph 1, The above degree of dementia is any one of normal, mild cognitive impairment (MCI), and Alzheimer's disease (AD). Method for determining the degree of dementia.
14. In Paragraph 1, The degree of dementia determined above is output through the user terminal, Method for determining the degree of dementia.
15. A computer-readable recording medium storing a program for executing the method according to paragraph 1.
16. In an electronic device for determining the degree of dementia in a user, At least one processor including a processing circuit; and Memory comprising one or more storage media that store instructions Includes, When the above instructions are executed individually or collectively by the at least one processor, the electronic device causes at least: Outputting a first content that was pre-produced to determine the degree of dementia of the user through the user terminal, and Acquiring the user's first reaction information regarding the first content through the user terminal, wherein the first reaction information includes the user's first voice information corresponding to the first content. Outputting a second content that has been pre-produced to determine the degree of dementia of the user through the above user terminal, and Acquiring the user's second reaction information regarding the second content through the user terminal, wherein the second reaction information includes first gaze information regarding coordinates on a display where the user's eyes gaze in correspondence with the second content. Outputting a third content pre-produced to determine the degree of dementia of the user through the above user terminal, and Acquiring third reaction information of the user regarding the third content through the user terminal, wherein the third reaction information includes second voice information of the user corresponding to the third content and second gaze information regarding coordinates on the display where the user's eyes are gazing -, Outputting a pre-made fourth content through the above user terminal to determine the degree of dementia of the user, and Acquiring the user's fourth reaction information regarding the fourth content through the user terminal, wherein the fourth reaction information includes the user's first touch action information corresponding to the fourth content. The degree of dementia of the user is determined by inputting the first reaction information, the second reaction information, the third reaction information, and the fourth reaction information into a dementia degree classification model based on an artificial neural network. making, Electronic device.