Information processing system, information processing method, program, and information processing device
The information processing system enhances cognitive function assessment by integrating attribute and image information, improving detection accuracy for conditions like dementia.
Patent Information
- Application Number
- JP2024048810
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-25
- Publication Date
- 2025-10-07
AI Technical Summary
Existing technologies for outputting indicators of cognitive function based on images of a subject's appearance lack accuracy.
An information processing system that acquires attribute information and image information from a subject's face, using a combination of hardware components and software algorithms to output an index indicating the state of cognitive function, considering factors such as apparent age, lifestyle, medical history, and genetic information.
Enables the output of cognitive function indices with higher accuracy, facilitating early detection of conditions like dementia and mild cognitive impairment.
Smart Images

Figure 2025148182000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an information processing system, an information processing method, a program, and an information processing device. [Background technology]
[0002] Patent Document 1 discloses a diagnosis support information providing device that provides information useful for diagnosing brain function disorders such as dementia through simple measurements.
[0003] In the diagnostic assistance information providing device, a face direction measuring unit measures the face direction of the subject in each of the time-series image data. A change amount calculating unit calculates a time-series face direction change amount based on the face direction measured by the face direction measuring unit. A diagnostic assistance information generating unit calculates a percentage of face direction change amounts that are equal to or less than a predetermined value during a predetermined period based on the face direction change amount calculated by the change amount calculating unit, and generates diagnostic assistance information indicating the calculation result. An output unit outputs the diagnostic assistance information generated by the diagnostic assistance information generating unit. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Application Publication No. 2019-055064 Summary of the Invention [Problem to be solved by the invention]
[0005] However, there is still room for improvement in the technology for outputting indicators of cognitive function based on images of a subject's appearance. [Means for solving the problem]
[0006] According to one aspect of the present invention, there is provided an information processing system. The information processing system includes at least one processor. The processor is configured to execute a program to perform the following steps: In the acquisition step, attribute information relating to attributes of a subject and image information relating to an image including at least a portion of the subject's face are acquired; and in the status output step, an index indicating the state of the subject's cognitive function is output based on the image information and the attribute information.
[0007] With this configuration, it is possible to output an index related to cognitive function with higher accuracy. [Brief explanation of the drawings]
[0008] [Figure 1] 1 is a configuration diagram illustrating an example of an information processing system 1. FIG. [Figure 2] 2 is a block diagram showing an example of a hardware configuration of an information processing device 2. FIG. [Figure 3] 3 is a block diagram showing an example of the hardware configuration of a medical professional terminal 3. FIG. [Figure 4] FIG. 2 is a plan view showing an example of the setup of the imaging system 5. [Figure 5] FIG. 10 is a diagram showing another example of the setup of the imaging system 5. [Figure 6] 1 is an activity diagram showing an example of the flow of information processing executed in the information processing system 1. FIG. [Figure 7] 1 is an example of a scatter plot in which the apparent age and the degree of cognitive decline of a normal group and a group showing cognitive decline are plotted according to whether or not the group smokes. [Figure 8] 1 is a diagram showing an example of a front image 6, which is an image captured by a front camera 51. FIG. [Figure 9] 1 is a diagram showing an example of a side image 7, which is an image captured by a side camera 52. FIG. [Figure 10] 1 shows an example of an evaluation result display screen 8 displayed on the display unit 34 of the medical professional terminal 3. [Figure 11]FIG. 1 is an activity diagram showing an example of the flow of information processing when a subject 1000 is made to perform a task. [Figure 12] FIG. 10 is a diagram showing a recording screen 9, which is an example of a screen for presenting a task. DETAILED DESCRIPTION OF THE INVENTION
[0009] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS The present invention will be described below with reference to the accompanying drawings. Various features shown in the following embodiments can be combined with each other.
[0010] Incidentally, the program for realizing the software appearing in one embodiment may be provided as a non-transitory computer-readable medium, or may be provided so that it can be downloaded from an external server, or may be provided so that the program is started on an external computer and its functions are realized on a client terminal (so-called cloud computing).
[0011] Furthermore, various information processing according to an embodiment may realize input and output corresponding to the input. Here, the form of information referenced in such information processing (hereinafter referred to as reference information) is not limited as long as an output is obtained as a result of the input. The reference information may be, for example, rule-based information such as a database, a lookup table, or a predetermined function (including a decision formula such as a regression formula constructed using a statistical method), a trained model that has previously trained the correlation between input and output, or a large-scale language model that can output a desired result by inputting a prompt.
[0012] In one embodiment, a "unit" may include, for example, a combination of hardware resources implemented by a circuit in the broad sense and software information processing that can be specifically realized by these hardware resources. In one embodiment, various information is handled, and this information is represented, for example, by physical values of signal values representing voltage and current, high and low signal values as a binary bit set consisting of 0 or 1, or quantum superposition (so-called quantum bits), and communication and calculations can be performed on a circuit in the broad sense.
[0013] Furthermore, a circuit in the broad sense is a circuit realized by at least an appropriate combination of a circuit, circuitry, processor, memory, etc. The processor may be a general-purpose processor or a dedicated circuit. That is, it includes an application specific integrated circuit (ASIC), a programmable logic device (e.g., a simple programmable logic device (SPLD), a complex programmable logic device (CPLD), and a field programmable gate array (FPGA)), etc.
[0014] 1. Hardware Configuration This section explains the hardware configuration.
[0015] <Information Processing System 1> FIG. 1 is a configuration diagram showing an example of an information processing system 1. The information processing system 1 includes an information processing device 2 and a medical professional terminal 3, which is an example of a user terminal. The information processing system 1 may further include a sound collection device 4 and an imaging system 5. The information processing device 2, the medical professional terminal 3, the sound collection device 4, and the imaging system 5 are configured to be able to communicate with each other via a telecommunications line. In one embodiment, the information processing system 1 is made up of one or more devices or components. For example, if the information processing system 1 consists only of the information processing device 2, the information processing system 1 can be the information processing device 2. Similarly, if the information processing system 1 consists only of the medical professional terminal 3, the information processing system 1 can be the medical professional terminal 3.
[0016] <Information processing device 2> 2 is a block diagram showing an example of the hardware configuration of the information processing device 2. The information processing device 2 includes a communication bus 20, a communication unit 21, a storage unit 22, and at least one processor 23, and these components are electrically connected via the communication bus 20 inside the information processing device 2. Each component will be further described.
[0017] The communication unit 21 is preferably a wired communication means such as USB, IEEE1394, Thunderbolt (registered trademark), wired LAN network communication, etc., but may also include wireless LAN network communication, mobile communication such as 3G / LTE / 5G, BLUETOOTH (registered trademark) communication, etc. as needed. In other words, it is more preferable to implement it as a collection of multiple communication means. In other words, the information processing device 2 may communicate various information from the outside via the communication unit 21 and the network.
[0018] The storage unit 22 stores various information defined above. This may be implemented as a storage device such as a solid state drive (SSD), a solid state hybrid drive (SSHD), a hard disk drive (HDD), a universal serial bus (USB) flash drive (USB memory), an SD memory card, a CD, a DVD, or a Blu-ray (registered trademark) disc (BD) that stores various programs and the like related to the information processing device 2 executed by the processor 23, or as a memory such as a random access memory (RAM) that stores temporarily required information (arguments, arrays, etc.) related to program calculations. The storage unit 22 stores various programs, variables, etc. related to the information processing device 2 executed by the processor 23.
[0019] The processor 23 processes and controls the overall operations related to the information processing device 2. The processor 23 is, for example, a central processing unit (CPU) not shown. The processor 23 realizes various functions related to the information processing device 2 by reading out predetermined programs stored in the storage unit 22. In other words, information processing by software stored in the storage unit 22 is specifically realized by the processor 23, which is an example of hardware, and can be executed as each functional unit included in the processor 23. Note that the processor 23 is not limited to being single, and may be implemented with multiple processors 23 for each function. A combination of these may also be used.
[0020] The processor 23 is configured as an acquisition unit to acquire information from the medical professional terminal 3 or other devices. The processor 23 is configured to be able to acquire various pieces of information by reading out various pieces of information stored in a storage area that is at least a part of the memory unit 22 and writing the read out information in a working area that is at least a part of the memory unit 22. The storage area is, for example, an area of the memory unit 22 that is implemented as a storage device such as an SSD. The working area is, for example, an area that is implemented as a memory such as a RAM. Note that acquisition by the processor 23 includes acquiring the output results of each functional unit included in the processor 23.
[0021] The processor 23 is configured as a display processing unit to display various information. This information can be presented to the user via the display unit 34 of the medical practitioner terminal 3 (described later) or another device. In such a case, for example, the processor 23 controls the display unit 34 of the medical practitioner terminal 3 to display visual information such as images, icons, messages, and the like, including screens, still images, or moving images (continuous images). The processor 23 may generate only rendering information for displaying the visual information on the medical practitioner terminal 3. The processor 23 may also present the output information to the user without going through the user of the medical practitioner terminal 3 or another device.
[0022] <Medical Professional Terminal 3> 3 is a block diagram showing an example of the hardware configuration of the medical professional terminal 3. The medical professional terminal 3 includes a communication bus 30, a communication unit 31, a storage unit 32, a processor 33, a display unit 34, and an input unit 35, and these components are electrically connected via the communication bus 30 inside the medical professional terminal 3. The explanation of the communication unit 31, the storage unit 32, and the processor 33 is omitted because they are the same as the explanation of each unit in the information processing device 2.
[0023] The display unit 34 may be included in the housing of the medical professional terminal 3 or may be externally attached. The display unit 34 displays a graphical user interface (GUI) screen that can be operated by the user. This is preferably implemented by using display devices such as a CRT display, a liquid crystal display, an organic EL display, and a plasma display, which are selected depending on the type of medical professional terminal 3.
[0024] The input unit 35 is configured to accept input from the user. The input unit 35 may be included in the housing of the medical professional terminal 3 or may be externally attached. For example, the input unit 35 may be integrated with the display unit 34 and implemented as a touch panel. A touch panel allows the user to input tapping, swiping, and the like. Of course, instead of a touch panel, a switch button, a mouse, a QWERTY keyboard, a voice recognition device, a gesture detection device, a gaze detection device, a biosignal detection device, an imaging device, and the like may be used. That is, the input unit 35 accepts an input operation made by the user. In response, the input unit 35 transfers a signal corresponding to the input operation to the processor 33 via the communication bus 30. The processor 33 can execute predetermined control and calculations as necessary.
[0025] <Sound collection device 4> The sound collector 4 is a so-called microphone that is configured to be able to convert external sounds into signals. The sound collector 4 may be provided by directly connecting the sound collector 4 to the information processing device 2, but is provided in or connected to the medical professional terminal 3.
[0026] The sound collection device 4 is configured to generate voice data by collecting the user's speech. The voice data is temporarily stored in a memory in the user terminal and does not have to be stored non-volatilely in the storage unit 32. The voice data generated by the sound collection device 4 is configured to be transferable to the information processing device 2 via a network.
[0027] The sound collection device 4 collects, but is not limited to, at least sounds in the human audible range, sounds with frequencies between 20 Hz and 20,000 Hz, and converts them into electrical signals. The sound may be recorded in monaural or stereo. The sampling rate for digitally processing the sound data may be, for example, 48,000 Hz, 44,100 Hz, 32,000 Hz, 22,050 Hz, 16,000 Hz, 11,025 Hz, 11,000 Hz, 8,000 Hz, etc. The sampling rate may be within any of the ranges of values exemplified here. Increasing the sampling rate allows for more precise discretization of the temporal timing of the sound, improving the accuracy of voice recognition.
[0028] Furthermore, the data collected by the sound collection device 4 may be appropriately compressed by the processor 33 of the medical professional terminal 3, and the compression format at this time may be any of MP3, AAC, WMA, Vorbis, AC3, MP2, FLAC, TAK, etc. Compression can reduce communication traffic due to data transfer from the user terminal to the information processing device 2.
[0029] <Imaging System 5> The imaging system 5 is configured to capture an image of the subject 1000's appearance and generate image information related to an image including at least a portion of the subject's 1000 face. The imaging system 5 is configured to communicate with the medical professional terminal 3 via the communication bus 30. The imaging system 5 is configured to capture at least a front image of the subject 1000. The imaging system 5 may be further configured to capture a side image of the subject 1000. In this embodiment, the imaging system 5 includes a front camera 51 and a side camera 52 as imaging devices. Each of the front camera 51 and the side camera 52 is an optical camera configured to detect at least visible light. The images captured by the front camera 51 and the side camera 52 are transmitted to the medical professional terminal 3 via the communication bus 30. The image format may be any format, such as JPG, JPEG, PNG, BMP, or PDF. The image format may also be compatible with video formats such as MPG, MPEG, AVI, MOV, SMV, or FLV. At least one of the front camera 51 and the side camera 52 may be integrated with the medical professional terminal 3. The imaging device is not limited to an optical camera as long as it can capture an image of the subject's 1000 appearance, and may be an ultrasonic camera, an infrared camera, or any other suitable device. When capturing an image of the subject's appearance using the imaging system 5, it is more preferable to capture the image again after a predetermined time has elapsed, and estimate the cognitive function level based on changes in facial expression. That is, the imaging system 5 may be configured to capture continuous images (e.g., moving images).
[0030] FIG. 4 is a plan view showing an example of the setup of the imaging system 5. As shown in FIG. 4, the front camera 51 included in the imaging system 5 is disposed in the front region R1 and configured to capture an image of the frontal appearance of the face of the subject 1000. The side camera 52 included in the imaging system 5 is disposed in the side region R2 and configured to capture an image of the side appearance of the face of the subject 1000. Note that when the imaging device is a camera, the imaging system 5 is not limited to being implemented with multiple cameras 51 and 52 as shown in the figure. For example, the imaging system 5 may be implemented using a single wide-angle camera or a single camera that is movable to perform panoramic imaging. The imaging system 5 may also be implemented by combining an optical camera with a different type of imaging device (such as an ultrasonic camera or an infrared camera).
[0031] The front region R1 is defined to include a front direction D1. The front direction D1 is a direction defining the front of the subject 1000, for example, a direction from the center line of the subject 1000 toward the front of the subject 1000. The front region R1 is, for example, a region within a range from the front direction D1 to a direction intersecting the front direction D1 at an inclination angle θ1. The inclination angle θ1 may be any angle less than 45 degrees, specifically, for example, 0, 5, 10, 15, 20, 25, 30, 35, or 40 degrees, or may be within a range between any two of the values exemplified here. For example, the inclination angle θ1 is 0 to 45 degrees, preferably 0 to 40 degrees, and more preferably 0 to 30 degrees. The front region R1 is preferably configured to be able to capture images of both eyes of the subject 1000 using the front camera 51. In this embodiment, the front camera 51 is disposed in the front direction D1 of the subject 1000.
[0032] The side region R2 is defined to include a side direction D2. The side direction D2 is a direction defining the side of the subject 1000, for example, a direction from the center line of the subject 1000 toward either the left or right side of the subject 1000. The side direction D2 is a direction intersecting with the front direction D1, specifically a direction perpendicular to the front direction D1. The side region R2 is a region within a range from the side direction D2 to a direction intersecting at an inclination angle θ2. The inclination angle θ2 may be any angle less than 45 degrees, specifically, 0, 5, 10, 15, 20, 25, 30, 35, or 40 degrees, for example, or may be within a range between any two of the values exemplified here. For example, the inclination angle θ1 is 0 to 45 degrees, preferably 0 to 40 degrees, and more preferably 0 to 30 degrees. The side region R2 may be defined so as to be able to capture the three-dimensional appearance of the face from the side direction D2, for example, the outline of the center line of the face in the front direction D1, which is difficult to grasp from the front region R1. In particular, the side camera 52 may be defined so as to be able to capture the shape of the tip of the chin of the subject 1000 or the three-dimensional shape of the throat. The front region R1 and the side region R2 are defined so as not to overlap with each other. In this embodiment, the side camera 52 is disposed in the side direction D2 of the subject 1000.
[0033] The imaging system 5 is not limited to one including the front camera 51 and the side camera 52. For example, the imaging system 5 may include only the front camera 51 (i.e., one imaging device). FIG. 5 is a diagram showing another example of the setup of the imaging system 5. As shown in FIG. 5, the front camera 51 of the imaging system 5 is arranged in the front region R1. Therefore, the image information may include only the front image 6. The front camera 51 may also be configured to function as the side camera 52 by moving from the front region R1 to the side region R2. Furthermore, the imaging system 5 may be configured to measure a three-dimensional shape of the external appearance of the subject 1000, such as the face. The three-dimensional shape can be measured, for example, by using a time-of-flight (TOF) camera or the like as the imaging system 5.
[0034] 3. Information Processing In this section, information processing executed in the above-mentioned information processing system 1 will be described. Through this information processing, the information processing system 1 acquires attribute information related to the attributes of the subject 1000 and image information related to an image including the face of the subject 1000, and outputs an index indicating the state of cognitive function of the subject 1000 based on the image information and the attribute information. With this configuration, even if the criteria for whether cognitive function is normal or not differ depending on the attributes of the subject 1000, it is possible to output an index indicating the state of cognitive function with high accuracy. This makes it easier to detect, for example, dementia that may have developed in the subject 1000 or mild cognitive impairment, which is a precursor to mild dementia, at an earlier stage.
[0035] 3.1. Information processing flow FIG. 6 is an activity diagram showing an example of the flow of information processing executed in the information processing system 1. Note that the information processing may include any exception processing not shown. Exception processing includes interruption of the information processing or omission of each process. Selection or input performed in the information processing may be based on a user operation or may be performed automatically without relying on a user operation.
[0036] [Activity A1] First, in activity A1, the processor 23 acquires image information. The image information is generated, for example, as an image of the subject 1000 captured by the imaging system 5. Note that the image information is not limited to image information captured by the imaging system 5, and may be image information captured in advance using a camera or the like owned by the subject 1000.
[0037] [Activity A2] Next, in activity A2, the processor 23 calculates the apparent age of the subject 1000 by performing a predetermined image analysis on the acquired image information. The apparent age of the subject 1000 is an example of a feature extracted from an image including the face of the subject 1000. The apparent age as a feature has a positive correlation with the degree of decline in cognitive function. The feature can be extracted, for example, using a trained model that has been trained in advance. The state of cognitive function includes dementia (mild, moderate, severe) or mild cognitive impairment (pre-stage of mild dementia). It is particularly preferable that the feature be capable of predicting whether or not there is mild cognitive impairment.
[0038] [Activity A3] Next, in activity A3, the processor 23 performs a predetermined image analysis on the acquired image information to extract attribute information related to the attributes of the subject 1000. The feature amount can be extracted, for example, using a trained model that has been trained in advance. Note that attribute information that cannot be acquired from image analysis may be identified based on the subject 1000's self-reporting.
[0039] The attribute information may include information about the living environment or lifestyle of the subject 1000. With this configuration, even if the living environment or lifestyle of the subject 1000 changes the appearance of the subject 1000 and there is a large discrepancy between the subject's 1000 apparent age and his / her actual age, an index indicating the state of cognitive function can be accurately output from the image information. Hereinafter, for convenience of explanation, information about the living environment or lifestyle of the subject 1000 is referred to as "lifestyle information." The lifestyle information is not limited to the current living environment or lifestyle, but may also include information about the living environment or lifestyle of the subject 1000 in the past (e.g., within a predetermined period from the present). The lifestyle information may include, for example, the address of residence, temperature, humidity, etc., as information about the living environment. Furthermore, the lifestyle information may include, as information about lifestyle habits, drinking habits, exercise habits, and, in particular, information about the subject's 1000 smoking. With this configuration, since smoking tends to increase the age estimated from the subject's 1000 appearance, an index indicating the state of cognitive function can be output taking into account the impact of smoking on the subject's 1000 appearance. Information regarding the subject's 1000 smoking may include, for example, the amount of smoking, the frequency of smoking, the duration of smoking, and the like.
[0040] The attribute information may include information about the birth of the subject 1000. With this configuration, it is possible to output an index indicating the state of cognitive function after taking into consideration the influence of information about the birth of the subject 1000, such as race, ethnicity, genetics, etc., on the appearance of the subject 1000. Hereinafter, for convenience of explanation, information about the birth of the subject 1000 will be referred to as birth information. The birth information may include, for example, information about the genetics of the subject 1000. The genetic information may include, for example, information such as skin color (white, yellow, black, etc.) and eye shape (e.g., single or double eyelids) derived from genes. In other words, the birth information is information derived from the birth of the subject 1000 that may affect the appearance of the subject 1000. The birth information may indicate, for example, whether the subject 1000 is genetically susceptible to ultraviolet rays and prone to wrinkles when exposed to ultraviolet rays, or whether the subject is genetically prone to developing muscles such as the jaw. Furthermore, the birth information may indicate that, for example, genetic skin color may affect the wrinkle detection rate when image information is acquired using different tones. Furthermore, the birth information may indicate that, for example, for a subject with congenital ptosis, information obtained from a facial image may be predicted to indicate a state of further cognitive decline.
[0041] The attribute information may include information about the medical history of the subject 1000. With this configuration, even if the subject's 1000 appearance may change depending on the medical history, an index indicating the state of cognitive function can be output with high accuracy. Hereinafter, for convenience of explanation, information about the subject's 1000 medical history will be referred to as medical history information. The medical history information may include, for example, the subject's 1000 medical history, surgical history, pregnancy history, dialysis history, allergy history, etc. In particular, the medical history information may include information about cosmetic treatments performed on the subject 1000. With this configuration, for example, even if the subject's apparent age is estimated to be lower than their actual age due to cosmetic treatments, the apparent age can be estimated while taking into account the change in apparent age due to the cosmetic treatments. The cosmetic treatments are not limited to treatments performed to treat an illness or injury suffered by the subject 1000, but may also include treatments performed for cosmetic purposes. Cosmetic procedures are not limited to surgical procedures, but may include any procedure, such as oral medication, laser treatment on the skin, or injection of medication into the affected area. In particular, cosmetic procedures may include procedures performed on the face (especially the eyes or mouth). Note that cosmetic procedures (or plastic surgery) refer to procedures that can change the appearance of the face, and may include, for example, cataract surgery and ptosis surgery.
[0042] The attribute information may include information regarding the condition of the oral cavity of the subject 1000. With this configuration, the condition of the oral cavity may reflect the chewing function and thus correlate with the swallowing function. Here, dysphagia caused by a decline in swallowing function is one of the symptoms commonly seen in patients with dementia or decline in cognitive function, and therefore correlates with decline in cognitive function. Therefore, by incorporating information regarding the condition of the oral cavity of the subject 1000 as attribute information, an index related to cognitive function can be output with greater accuracy. Hereinafter, for convenience of explanation, information regarding the condition of the oral cavity of the subject 1000 will be referred to as oral information. The oral information may include information regarding the oral cavity, particularly the teeth or gums, such as the state of tooth discoloration, the state of gums, the state of caries treatment, the state of orthodontic treatment, and the state of dry mouth. In particular, the oral information includes information regarding the remaining teeth of the subject 1000. With this configuration, even if the apparent age of the subject 1000 is likely to be estimated higher than the actual age due to the condition of the remaining teeth, such as the number of remaining teeth, the health condition of the remaining teeth, etc., it is possible to output an index related to cognitive function with higher accuracy. The information related to the remaining teeth may include any information, such as the number of remaining teeth, the positions of the remaining teeth, the health condition of the remaining teeth, etc.
[0043] The attribute information may include at least one of the above information, i.e., information on the living environment or lifestyle of the subject 1000, information on the birth of the subject 1000, information on the medical history of the subject 1000, and information on the condition of the oral cavity of the subject 1000.
[0044] [Activity A4] Next, in activity A4, the processor 23 extracts at least a portion of the attribute information based on the image information. This configuration reduces the amount of work required by the user to acquire the attribute information, compared to when the subject 1000, the subject's caregiver, a doctor, or other user manually inputs all of the attribute information. For example, a system can be provided that easily outputs indicators related to cognitive function, even for a subject 1000 whose cognitive function has declined and who has difficulty inputting attribute information by himself or herself. For example, the processor 23 extracts the subject's 1000's race (e.g., Caucasian, Asian, Black, etc.) as the subject's 1000's birth information based on the image information. Furthermore, if the image information includes an area related to the subject's 1000's mouth and the area includes the appearance of the subject's 1000's teeth, information related to the subject's 1000's smoking (e.g., whether or not the subject smokes) can be extracted based on the appearance of the teeth. Extracting attribute information also means acquiring attribute information. An example of a method for extracting information related to smoking is to input information about the tooth color of the subject 1000 into a pre-trained machine learning model to extract whether or not the subject smokes, which is an example of information related to smoking. The machine learning model can be obtained by training the model with data on tooth color and training data indicating whether or not the subject smokes, using the data on tooth color as an input variable and whether or not the subject smokes as an output. Note that, because a history of smoking causes yellowish staining of teeth, it is also possible to determine whether or not the subject smokes based on the tooth staining.
[0045] [Activity A5] Next, in activity A5, processor 23 accepts input of further attribute information. Processor 23, for example, uses display unit 34 to present an input screen that prompts the person making the input to input attribute information. The input screen is configured to accept text input from the person making the input, for example, in the form of a questionnaire such as a medical questionnaire. The input screen may also be configured to accept voice input from the person making the input. Furthermore, the input screen may be configured to accept corrections to the attribute information extracted in activity A3. The display of the input screen may be executed by processor 33. If processor 23 determines that a sufficient amount of attribute information has been obtained, this process may be omitted.
[0046] [Activity A6] Next, in activity A6, the processor 23 identifies a comparison target based on the attribute information. The comparison target is a parameter (in other words, a threshold) indicating whether the cognitive function of the subject 1000 is within the range of the normal group. The comparison target is a parameter corresponding to the degree of decline in cognitive function. A method for identifying the comparison target will be described later.
[0047] [Activity A7] Next, in activity A7, the processor 23 calculates an index indicating the degree of dementia or cognitive function state based on the apparent age and the comparison subject. The index indicating the degree of dementia is an example of an index indicating the state of cognitive function of the subject 1000. As described above, the apparent age has a positive correlation with the degree of cognitive decline. Therefore, the processor 23 may convert the apparent age into the degree of cognitive decline using predetermined reference information indicating the relationship between the apparent age and the degree of cognitive decline. The reference information may be set, for example, based on test data indicating the apparent age and the degree of cognitive decline of randomly selected subjects in a cognitive function test. More specifically, the reference information may be defined by a relational equation (e.g., a regression equation), a lookup table, a trained model, or the like calculated based on the test data. The trained model may be obtained, for example, by supervised learning in which the apparent age included in the test data is used as input and the degree of cognitive decline is used as output. Note that the reference information may be set for each different attribute information. In this case, the test data may further include labels corresponding to each of the attribute information. At this time, the processor 23 can obtain reference information for each attribute information by changing the test data used to calculate the regression equation and learn the learned model for each label.
[0048] The processor 23 compares the degree of cognitive decline after conversion with that of the comparison subject. If the degree of cognitive decline after conversion is less than that of the comparison subject, the processor 23 calculates that the cognitive function of the subject 1000 is within the range of the normal group. On the other hand, if the degree of cognitive decline after conversion is equal to or greater than that of the comparison subject, the processor 23 calculates that the subject 1000 may have dementia. The processor 23 may also calculate the degree of dementia (mild, moderate, severe, etc.) or the degree of cognitive decline (whether or not the subject has mild cognitive impairment, etc.) that the subject 1000 may have, depending on the magnitude of the difference between the degree of cognitive decline after conversion and that of the comparison subject. The type of dementia may be any type, such as Alzheimer's disease, vascular dementia, dementia with Lewy bodies, or frontotemporal dementia. As described above, being able to predict the degree of dementia or cognitive decline can change the required treatment method depending on the state of cognitive impairment or dementia, so it is preferable to be able to predict the state of cognitive function. Furthermore, it is more preferable that the prediction of the level of cognitive decline can accurately predict the state of mild cognitive impairment. Accurate detection of mild cognitive impairment is preferable because mild cognitive impairment is particularly reversible to a healthy state.
[0049] [Activity A8] Next, in activity A8, processor 23 outputs the calculation result from activity A6. In other words, processor 23 outputs an index indicating the state of cognitive function of subject 1000 based on the image information and attribute information. The output result is presented to subject 1000, for example, via display unit 34.
[0050] 3.2. How to identify comparison targets based on attribute information Next, we will explain how to identify a comparison subject based on attribute information in activity A5. Here, we will explain a method of identifying a comparison subject based on whether or not the subject smokes, as an example of attribute information. Figure 7 is an example of a scatter plot in which the apparent age and the degree of cognitive decline for a normal group and a group showing cognitive decline are plotted by whether or not the subject smokes. The group showing cognitive decline corresponds to a group of subjects who are judged to be at risk of developing dementia.
[0051] FIG. 7 shows a first non-smoking group G11, a second non-smoking group G12, a first smoking group G21, and a second smoking group G22. The first non-smoking group G11 is a normal group composed of non-smoking subjects. The second non-smoking group G12 is a group of non-smoking subjects who exhibit cognitive decline. The first smoking group G21 is a normal group composed of smoking subjects. The second smoking group G22 is a group of smoking subjects who exhibit cognitive decline. These groups exhibit a so-called positive correlation, in which the degree of cognitive decline increases as the apparent age increases. As shown in FIG. 7, the boundary between the first non-smoking group G11 and the second non-smoking group G12 can be expressed by thresholds Th11 and Th21.
[0052] The threshold Th11 is a value indicating the boundary between the first non-smoking group G11 and the second non-smoking group G12 based on apparent age, and is an example of a comparison target. The threshold Th11 may be determined, for example, based on the apparent age values of the individuals included in the second non-smoking group G12. The threshold Th11 may be set, for example, to the lowest apparent age value among the individuals in the second non-smoking group G12. This reduces the possibility that a subject 1000 belonging to the second non-smoking group G12 will be determined to belong to the first non-smoking group G11 based on their apparent age. The threshold Th11 may be set lower or higher than the minimum apparent age value of the individuals included in the second non-smoking group G12.
[0053] The threshold Th21 is a value indicating the boundary between the first non-smoking group G11 and the second non-smoking group G12 based on the degree of cognitive decline, and is an example of a comparison target. In the present embodiment, the threshold Th21 is a comparison target to be compared with the degree of cognitive decline of the subject 1000 in activity A6 when the subject 1000 has attribute information indicating that he or she is not a smoker. The threshold Th21 may be set, for example, to a value indicating the lowest degree of cognitive decline among the subjects in the second non-smoking group G12. This reduces the possibility that a subject 1000 belonging to the second non-smoking group G12 will be determined to belong to the first non-smoking group G11 based on the degree of cognitive decline. The threshold Th11 may be set lower or higher than the minimum apparent age included in the second non-smoking group G12. The boundary between the first smoking group G21 and the second smoking group G22 may be expressed by the threshold Th12 and / or the threshold Th22.
[0054] The threshold value Th12 is a value indicating the boundary between the first smoking group G21 and the second smoking group G22 based on apparent age, and is an example of a comparison target. The threshold value Th12 may be determined, for example, based on the apparent age values of the individuals in the second smoking group G22. The threshold value Th12 may be set, for example, to the value of the lowest apparent age among the individuals in the second smoking group G22. This reduces the possibility that a subject 1000 belonging to the second smoking group G22 will be determined to belong to the first smoking group G21 based on their apparent age. The threshold value Th12 may be set lower or higher than the minimum apparent age value of the individuals in the second smoking group G22.
[0055] The threshold Th22 is a value indicating the boundary between G21 and the second smoking group G22 based on the degree of cognitive decline, and is an example of a comparison target. In the present embodiment, the threshold Th22 is a comparison target to be compared with the degree of cognitive decline of the subject 1000 in activity A6 when the subject 1000 has attribute information indicating that he or she is not a smoker. The threshold Th22 may be set, for example, to a value indicating the lowest degree of cognitive decline among the second smoking group G22. This reduces the possibility that the subject 1000 belonging to the second smoking group G22 will be determined to belong to the first smoking group G21 based on the degree of cognitive decline. The threshold Th21 may be set lower or higher than the minimum apparent age included in the second non-smoking group G12.
[0056] As shown in Figure 7, the first smoking group G21 and the second smoking group G22, which correspond to the attribute information of smoking, tend to have a higher apparent age overall than the first non-smoking group G11 and the first non-smoking group G11, which correspond to the attribute information of not smoking. However, the difference in the degree of cognitive decline is not as great as the influence of apparent age. Therefore, the thresholds Th11 and Th21 for the non-smoking subject 1000 tend to be lower than the thresholds Th12 and Th22 for the smoking subject 1000. The values suitable for comparison may vary depending on the attribute information.
[0057] In this way, if the degree of cognitive decline is evaluated based on appearance features obtained from image information such as apparent age without considering attribute information, the accuracy of the index showing the degree of dementia may decrease. Furthermore, compared to when a normal group and a group showing cognitive decline are classified without considering attribute information (for example, when the first non-smoking group G11 and the first smoking group G21 are combined, and the first smoking group G21 and the second smoking group G22 are combined), by specifying a threshold (comparison target) showing whether or not a subject is in the normal group according to attribute information, it is possible to specify a threshold (comparison target) with high accuracy according to the attributes of the 1,000 individual subjects.
[0058] 3.3.Examples of image information In this section, an example of an image defined by the image information described above will be described.
[0059] FIG. 8 is a diagram illustrating an example of a front image 6 captured by the front camera 51. As illustrated in FIG. 8, the front image 6 is an image obtained by capturing an image of the subject 1000 from a front region R1 using the front camera 51, and is an image in which the front appearance of the face of the subject 1000 can be seen. The front image 6 includes a face region 61 and a neck vicinity region 62. The face region 61 corresponds to the face of the subject 1000. The face region 61 may include a region corresponding to the eyes or mouth of the subject 1000. For example, the face region 61 may include at least one of an eye region 611 corresponding to the eyes of the subject 1000 and a mouth region 612 corresponding to the mouth of the subject 1000. With this configuration, changes in appearance are most noticeable in the eyes or mouth, making it possible to output an index indicating the state of cognitive function with greater accuracy. In this embodiment, the eye region 611 is defined to include both eyes of the subject 1000. The eye region 611 may include the eyebrow region.
[0060] The front image 6 may further include a neck vicinity region 62 as an example of a region related to the jaw or throat of the subject 1000. According to this configuration, the jaw or throat of the subject 1000 is an important part of the subject's 1000 swallowing function. Because a decline in swallowing function may be correlated with a decline in cognitive function, incorporating the jaw or throat of the subject 1000 as an image can more accurately output an index related to cognitive function. This can assist, for example, in the prevention or early detection of dementia, which is particularly prone to a decline in swallowing function. The neck vicinity region 62 is defined, for example, to include the vicinity of the subject's 1000's neck and is located below the face region 61. The face region 61 and the neck vicinity region 62 may partially overlap each other. For example, the face region 61 and the neck vicinity region 62 may both be defined to include the jaw of the subject 1000. Furthermore, since the jaw or throat are areas that tend to be less frequently subjected to plastic surgery than the eyes or mouth, even if, for example, subject 1000 appears younger due to cosmetic surgery, an index showing the state of cognitive function can be output with high accuracy.
[0061] FIG. 9 is a diagram showing an example of side image 7, which is an image captured by side camera 52. As shown in FIG. 9, side image 7 is an image in which the side of the face of subject 1000 is visible, obtained by using side camera 52 to capture an image of subject 1000 from side region R2, which is different from front region R1. Side image 7 is configured to make it possible to visually recognize the three-dimensional effect of the front of the face in front image 6, for example, the unevenness in the front direction D1 caused by the nose and chin. Side image 7 may include a face region 71 and a neck vicinity region 72. Face region 71 is a region obtained by capturing an image of a part of the face included in face region 61 from side direction D2. Face region 71 may include eye region 711 and mouth region 712. Eye region 711 may also include an eyebrow region.
[0062] The eye region 711 is a region obtained by capturing an image of the part of the face included in the eye region 611 from the side direction D2. The mouth region 712 is a region obtained by capturing an image of the part of the face included in the mouth region 612 from the side direction D2.
[0063] The neck vicinity region 72 is a region obtained by capturing an image of the face portion included in the neck vicinity region 62 from the lateral direction D2. The face region 71 and the neck vicinity region 72 may partially overlap each other, similar to the face region 61 and the neck vicinity region 62. The region near the neck is unlikely to be the subject of cosmetic surgery or the like, and it is therefore preferable for image information to include the region near the neck, as this makes the region less susceptible to cosmetic surgery or the like.
[0064] In summary, the image may include a front image 6 in which the front appearance of the subject's face is visible, and a side image 7 in which the side of the subject's face is visible, obtained by capturing an image from a side region R2 different from the front region R1. This configuration allows for a more accurate output of an index related to cognitive function by evaluating the facial appearance of the subject from multiple angles. The areas included in the front image 6 and the side image 7 may be preset or automatically identified by image recognition or the like. Eye movements may also be measured simultaneously when acquiring the images. If the eye movements of a subject with impaired cognitive function include characteristic movements that differ from those of subjects belonging to a normal group, the processor 23 may further calculate an index related to the state of cognitive function based on the eye movements of the subject (subject 1000). The processor 23 may also calculate the index based on any information related to the subject's movements, such as data representing the subject's walking state (e.g., fluctuations in the pressure applied from the feet to the floor, fluctuations in the movement of the hips during walking movements corresponding to each step, etc.). Furthermore, the method of calculating the index is not limited to the method described below, and may be, for example, a method that combines prediction of cognitive function based on spontaneous speech.
[0065] 3.4 Screen examples Next, an example of a screen displayed on the display unit 34 by the above information processing will be described.
[0066] (Evaluation result display screen 8) 10 shows an example of the evaluation result display screen 8 displayed on the display unit 34 of the medical professional terminal 3. For ease of explanation, the subject 1000 may be referred to as the test subject. The evaluation result display screen 8 includes an area 80, an area 81, an area 82, an area 83, and a remeasurement button 84.
[0067] Area 80 is an area where information on cognitive decline, which is an example of an index showing the state of cognitive function, is displayed. Hereinafter, for the sake of convenience, an index showing the state of cognitive function may be referred to as an evaluation result. Calculating the index can also be said to be evaluating the state of cognitive function. Area 80 displays the "cognitive function evaluation result" as the evaluation target.
[0068] Area 81 is an area where acquired image information, attribute information, etc. are displayed. For example, area 81 may display the subject's 1000 origins, smoking habits, etc. as input attribute information. Area 81 may also display a front image 6, a side image 7, etc. captured by the imaging system 5.
[0069] The area 82 is an area where the evaluation result 821 is displayed. For example, the area 82 displays information such as "cognitive function is at a normal level" as the evaluation result 821.
[0070] More preferably, area 83 displays a comparison result 831 showing a comparison between the current evaluation result and the previous evaluation result when the subject previously underwent the same evaluation. For example, comparison result 831 is evaluated based on the amount of change in current evaluation value 834 relative to the previous evaluation value 835, which will be described later. In other words, processor 23 further acquires evaluation information including past evaluation results of the specific user who is the subject. Processor 23 further evaluates the amount of change in the degree of cognitive decline of the specific user based on the evaluation information. For example, area 83 displays information indicating "no change" as comparison result 831 showing a comparison with the previous evaluation. When comparing multiple types of image data over time as described above, it is preferable that the image data be stored with features extracted. As described above, storing the data after feature extraction can improve security. If multiple types of image data are stored as is, there is a risk that if the stored information is leaked, it could be used for biometric authentication of bank accounts using AI or the like. Therefore, it is more preferable that the stored data be stored in a form in which only features are extracted.
[0071] Preferably, the area 83 further displays an item 832 "this time", an item 833 "last time", a current evaluation value 834, and a previous evaluation value 835 so that they can be seen by the subject.
[0072] This allows the subject to understand the level of their own cognitive function, and also allows the progress of the user's cognitive decline to be quantitatively evaluated.
[0073] 4. Information processing when subject 1000 performs a task In the information processing in Section 3, the apparent age of the subject 1000 is identified based on image information, the degree of decline in cognitive function is estimated from the apparent age, and the degree of decline in cognitive function is compared with that of a comparison subject to calculate an index related to the degree of dementia, but this is not limited to this. For example, the processor 23 may measure the degree of decline in cognitive function separately from the apparent age based on the result of execution of a task by the subject 1000 to measure the degree of cognitive function, and calculate an index related to the degree of dementia based on the apparent age and the degree of decline in cognitive function.
[0074] 4.1. Information processing flow Here, an example of information processing when the subject 1000 is made to perform a task will be described. For convenience of explanation, the information processing described in Section 3 will be referred to as the first information processing, and the information processing described in Section 4 will be referred to as the second information processing. In this section, unless otherwise specified, "information processing" means the second information processing. FIG. 11 is an activity diagram showing an example of the flow of information processing when the subject 1000 is made to perform a task.
[0075] [Activity A101] First, in activity A101, the processor 23 presents a task to the subject 1000 via the display unit 34 or the like. The task is set so that the score of the task correlates with the level of the subject 1000's cognitive function. For example, the task may be defined to require motor movements that decline with a decline in cognitive function. The task may be defined to require eye or mouth movements. For example, a task requiring mouth movements may include a speech task that has the subject 1000 speak a specific pattern. This configuration may enable more accurate output of an index indicating cognitive function based on, for example, facial movements resulting from the speech task, such as mouth movements and changes in facial expressions. Preferably, the speech task may be defined to have the subject 1000 repeatedly speak a pattern. This configuration may reduce the burden on the subject 1000 when performing the speech task. For example, the speech task may be defined to have the subject 1000 repeatedly speak a three-syllable pattern (e.g., "katama") multiple times (e.g., five times). The tasks may also include visual tasks that require the user to select specific graphic patterns displayed on the display 34 .
[0076] [Activity A102] Next, in activity A102, the processor 23 receives input from the subject 1000 according to the task. The input may be a touch operation on the input unit 35 or a voice input to the sound collection device 4. For example, in the case of a speaking task, the subject 1000 inputs voice according to the speaking task to the sound collection device 4, and the processor 23 receives the voice input to the sound collection device 4. As a result, the subject 1000 performs the presented task. The sound collection device 4 generates the input voice as voice data. The processor 23 acquires the generated voice data from the sound collection device 4.
[0077] [Activity A103] Meanwhile, while the subject 1000 is performing a task (in other words, while input corresponding to the task is being received), the imaging system 5 captures the appearance (particularly the face) of the subject 1000. As a result, the imaging system 5 acquires, as image information, successive images that allow the subject 1000 to visually recognize the facial movements of the subject 1000 while the subject 1000 is performing a predetermined task. In other words, the image information may include successive images that allow the subject 1000 to visually recognize the facial movements of the subject 1000 while the subject 1000 is performing a predetermined task. Hereinafter, for convenience of explanation, these images will be referred to as task images. With this configuration, signs of cognitive function decline can be more easily detected from the facial movements of the subject 1000 while performing a task, thereby enabling more accurate output of an index related to cognitive function. The frame rate of the successive images is arbitrary, but specifically may be, for example, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 fps, or may be within a range between any two of the values exemplified here.
[0078] [Activity A104] When the task is completed, the processor 23 ends the processing of activities A102 and A103. After that, in activity A104, the processor 23 calculates the degree of decline in cognitive function based on the result of the task performed by the subject 1000.
[0079] [Activity A105] Next, in activity A105, the processor 23 acquires image information. The specific aspects of this process are similar to those of activity A1, but here the task image captured in activity A103 is acquired as image information.
[0080] [Activity A106] Next, in activity A106, the processor 23 calculates the apparent age based on the image information. The specific aspects of this process are the same as those in activity A2.
[0081] [Activity A107] Next, in activity A107, the processor 23 performs a predetermined image analysis on the acquired image information to extract attribute information related to the attributes of the subject 1000. The specific aspects of this processing are the same as those in activity A3.
[0082] [Activity A108] Next, in activity A108, the processor 23 extracts at least a part of the attribute information based on the image information. The specific aspects of this process are the same as those in activity A4.
[0083] [Activity A109] Next, in activity A109, the processor 23 accepts input of further attribute information. The specific aspects of this process are the same as those in activity A5.
[0084] [Activity A110] Next, in activity A110, the processor 23 identifies a comparison target based on the attribute information. The specific aspects of this process are the same as those in activity A6.
[0085] [Activity A111] Next, in activity A111, the processor 23 calculates an index indicating the degree of dementia or cognitive impairment (whether it is mild cognitive impairment or not) based on the apparent age and the comparison subject. The specific aspect of this process is the same as that of activity A7.
[0086] 4.2. Screen examples Next, an example of a screen displayed in the second information processing will be described.
[0087] (Recording screen 9) 12 is a diagram showing a recording screen 9, which is an example of a screen for presenting a task. The recording screen 9 is an example of a screen displayed in activity A101. As shown in FIG. 12, the recording screen 9 includes an area 91, an area 92, a record button 93, a stop button 94, an area 95, an area 96, and an area 97.
[0088] Area 91 is an area where instructions regarding speech are displayed to the subject. Area 91 displays the instruction "After tapping the record button, please read the following aloud." Area 92 is an area where the specific speech content to be spoken by the subject is displayed. Area 92 displays the speech content "katama (5 times)." Note that the method of instructions regarding speech may be, for example, audio instructions. As described above, audio instructions make it easier for subjects with impaired visual function due to cataracts, glaucoma, etc. to understand the content of instructions. Audio instructions and visual instructions may also be combined. Furthermore, if a subject has glaucoma, for example, there is a risk that the instructions may not be accurately read due to visual field defects. Therefore, the position of the display portion of the instructions may be movable.
[0089] The record button 93 is a button for starting recording. The processor 23 acquires audio data of the subject's speech by accepting input to the record button 93 from the subject or medical professional. The stop button 94 is a button for stopping recording. The processor 23 stops acquiring audio data by accepting input to the stop button 94 from the subject or medical professional.
[0090] Area 95 is an area where it is visibly displayed that the subject's voice is being recorded. When processor 23 receives input from the subject or medical professional via record button 93, it displays "REC" in area 95. When processor 23 receives input from the subject or medical professional via stop button 94, it stops displaying "REC" in area 95.
[0091] Area 96 is an area in which the elapsed time from the start of recording is visibly displayed. The elapsed time may be displayed in area 96 using object IF3. Furthermore, area 96 may display the elapsed time using text information IF4 "00:02.00". Area 96 may also visibly display the remaining time.
[0092] Area 97 is an area where information indicating the timing of the subject's speech is displayed. In other words, processor 23 further presents information including instructions indicating the timing of speech to the user who is the subject in a comprehensible manner. In yet other words, processor 23 further presents guide information including information guiding the timing of speech to the specific user who is the subject in a comprehensible manner. Area 97 displays information "katamakatama" indicating the content of the subject's speech. Furthermore, the text information displayed in area 97 includes black text information IF1 and white text information IF2. As an example, when the timing to speak arrives, the style of the text indicating the content of the speech is changed from white text information IF2 to black text information IF1. For example, if the next word to be spoken by the subject is "ma," the white text information IF2 "ma" changes to black text information when the timing to speak arrives.
[0093] Here, the manner in which the timing of speech can be grasped is not limited to this. For example, the manner in which the characters indicating the content of speech are emphasized may change as the timing of speech arrives. Furthermore, the manner in which the characters are emphasized may change to return to their original manner as the timing of speech ends.
[0094] Furthermore, the instruction indicating the timing of speech is not limited to a change in the appearance of characters. For example, an object with the appearance of a metronome may be displayed in area 97. The object displayed in area 97 may be presented to the user in the form of a pendulum needle moving left and right in time with the timing of the subject's speech. Also, for example, an object with the appearance of clapping hands may be displayed in area 97. The object displayed in area 97 may be presented to the user in the form of clapping hands in time with the timing of the subject's speech. According to such an embodiment, the subject can speak while understanding the timing and rhythm of the speech. Therefore, it is possible to obtain voice data that reflects the subject's original speaking state.
[0095] Note that the information shown in Figure 12 is for illustrative purposes only and is not limited thereto. The information indicating the presence or absence or degree of cognitive decline is not limited in format, but preferably includes information indicating a stage in numbers as an evaluation value. More preferably, the evaluation value may be classified into "normal," "cognitive decline," "risk of cognitive decline," or "risk of dementia" depending on the stage.
[0096] Here, an example of a method for evaluating the degree of cognitive decline of a user from voice data will be described. This method includes a method for measuring and evaluating fluctuations in the speed of each "katama" repetition over and over again (a so-called fixed speech method). Furthermore, the method for evaluating the degree of cognitive decline of a user from voice data is not limited to the fixed speech method, and may be combined with various speech formats. Furthermore, it may be combined with various evaluation methods, such as fluctuations in the movement of the feet and hips when walking.
[0097] [others] The above embodiment may be modified as follows.
[0098] Dementia associated with Lewy body dementia and severe Parkinson's disease may be accompanied by physical tremors, such as tremors in the hands. Therefore, as described above, the state of physical tremors may be extracted as a feature, and the type of dementia or mild dementia may be further predicted. As described above, if it is possible to predict the type of dementia, it is even more preferable to be able to estimate the type that causes cognitive decline, since treatment methods and medications vary depending on the type of dementia. That is, the information processing device 2 preferably includes: (A) a means for predicting or inputting a person's attributes; (B) an image acquisition means capable of acquiring information about the person's appearance (e.g., face, full-body image, upper body image); (C) a feature extraction means using the acquired information; (D) a means for predicting a cognitive function state by combining the extracted content and the person's attributes; and (E) a display means for displaying the predicted content. In other words, the information processing device 2 includes at least one processor. The processor 23 includes a program for performing the following steps. Specifically, in the acquisition step, the processor 23 is configured to acquire appearance information regarding the subject 1000's appearance. In the estimation step, attribute information related to the subject's attributes is estimated by inputting appearance information into a pre-trained machine learning model. In the extraction step, features correlated with the subject's cognitive function are extracted from the acquired appearance information. In the calculation step, an index related to the subject's cognitive function is calculated based on the extracted features and the predicted attribute information. In the display step, the index related to the cognitive function is displayed. In other words, in the input acceptance step, the processor 23 accepts input of attribute information related to the subject's attributes. In the acquisition step, appearance information related to the subject's appearance is acquired. In the extraction step, features correlated with the subject's cognitive function are extracted based on the acquired appearance information. In the calculation step, an index related to the subject's cognitive function is calculated based on the extracted features and the predicted attribute information. In the display step, the index related to the cognitive function is displayed.
[0099] The configurations (A) to (E) described above allow cognitive function to be accurately predicted and visually recognized using images. Furthermore, it is more preferable that the appearance information acquisition means acquires images using video or acquires images multiple times over time. As described above, time-series data can improve the accuracy of the image acquisition. In other words, the appearance information is configured to include visual information representing the subject's appearance, and the visual information is preferably a series of images including the subject's appearance.
[0100] The image information may be any information that defines an image including the face of the subject 1000. For example, the image information may be only one of the front image 6 and the side image 7. Furthermore, the image information may be an image captured using the imaging system 5 from a position that is not included in either the front region R1 or the side region R2, as long as it defines an image that includes the face of the subject 1000.
[0101] The user terminal is not limited to the medical professional terminal 3 operated by a diagnosing professional such as a doctor, but may be a terminal operated by the subject 1000 himself or herself, or a terminal operated by a caregiver of the subject 1000, etc.
[0102] The information processing device 2 may be an on-premise type or a cloud type. As the information processing device 2 in the cloud type, the above-mentioned functions and processes may be provided in the form of, for example, SaaS (Software as a Service) or cloud computing.
[0103] In the above embodiment, the information processing device 2 performs various storage and control operations, but multiple external devices may be used instead of the information processing device 2. That is, various information and programs may be distributed and stored in multiple external devices using block chain technology or the like.
[0104] The above embodiment is not limited to the information processing system 1, and may be an information processing method or an information processing program. The information processing method includes each step of the information processing system 1. The program causes at least one computer to execute each step of the information processing system 1.
[0105] The information processing system 1 and the like may be provided in the following aspects.
[0106] (1) An information processing system comprising at least one processor, the processor being configured to execute a program to perform the following steps: in an acquisition step, attribute information relating to attributes of a subject and image information relating to an image including at least a portion of the subject's face are acquired; and in a status output step, an indicator indicating the state of the subject's cognitive function is output based on the image information and the attribute information.
[0107] With this configuration, even if the criteria for determining whether cognitive function is normal or not differ depending on the attributes of the subject, it is possible to output an index that accurately indicates the state of cognitive function, which can facilitate, for example, the early detection of dementia that may be present in the subject.
[0108] (2) In the information processing system described in (1) above, the image includes an area corresponding to the subject's eyes or mouth.
[0109] With this configuration, changes in appearance are most noticeable in the eyes or mouth, making it possible to output an index that indicates the state of cognitive function with greater accuracy.
[0110] (3) In the information processing system described in (1) or (2) above, the attribute information includes information about the living environment or lifestyle of the subject.
[0111] With this configuration, even if the subject's living environment or lifestyle changes the subject's appearance and there is a large discrepancy between the subject's apparent age and their actual age, it is possible to accurately output an indicator showing the state of cognitive function from the image information.
[0112] (4) In the information processing system described in (3) above, the information regarding the living environment or lifestyle habits of the subject includes information regarding the subject's smoking.
[0113] With this configuration, since smoking tends to increase the age that can be inferred from the subject's appearance, it is possible to output an index showing the state of cognitive function while taking into account the impact that smoking has on the subject's appearance.
[0114] (5) In the information processing system according to any one of (1) to (4) above, the attribute information includes information about the birth of the subject.
[0115] With this configuration, it is possible to output an index showing the state of cognitive function while taking into account the impact that information about the subject's birth, such as race, ethnicity, genetics, etc., has on the subject's appearance.
[0116] (6) In the information processing system according to any one of (1) to (5) above, the attribute information includes information relating to the medical history of the subject.
[0117] With this configuration, even if there is a possibility that the subject's appearance will change depending on their medical history, an index showing the state of cognitive function can be output with high accuracy.
[0118] (7) In the information processing system described in (6) above, the information regarding the subject's medical history includes information regarding cosmetic procedures performed on the subject.
[0119] With this configuration, even if the apparent age is estimated to be lower than the actual age due to cosmetic surgery, the apparent age can be estimated taking into account the change in apparent age due to cosmetic surgery.
[0120] (8) In the information processing system according to any one of (1) to (7) above, the attribute information includes information relating to the condition of the oral cavity of the subject.
[0121] According to this configuration, the state of the oral cavity may reflect the masticatory function and thus correlate with the swallowing function. Here, dysphagia caused by a decline in swallowing function is one of the symptoms of dementia, and therefore correlates with a decline in cognitive function. Therefore, by incorporating information about the state of the oral cavity of the subject as attribute information, it is possible to output an index related to cognitive function with greater accuracy.
[0122] (9) In the information processing system described in (8) above, the information regarding the subject's oral condition includes information regarding the subject's remaining teeth.
[0123] With this configuration, even if the subject's apparent age is likely to be estimated higher than their actual age due to the condition of their remaining teeth, such as the number of remaining teeth or the health of their remaining teeth, it is possible to output more accurate indicators of cognitive function.
[0124] (10) In the information processing system according to any one of (1) to (9) above, in the attribute extraction step, at least a part of the attribute information is extracted based on the image information.
[0125] With this configuration, it is possible to reduce the amount of work required by a user to acquire attribute information compared to when all attribute information is manually input by a user such as a subject, the subject's caregiver, a doctor, etc. For example, it is possible to provide a system that easily outputs indices related to cognitive function, even for a subject whose cognitive function has declined and who has difficulty inputting attribute information by himself / herself.
[0126] (11) In the information processing system according to any one of (1) to (10) above, the image further includes an area relating to the subject's jaw or throat.
[0127] According to this configuration, the subject's jaw and throat are important parts when the subject exerts swallowing function. Since a decline in swallowing function may be correlated with a decline in cognitive function, by incorporating the subject's jaw or throat as an image, it is possible to output an index related to cognitive function with greater accuracy. This can support, for example, the prevention or early detection of dementia, which is particularly prone to a decline in swallowing function.
[0128] (12) In the information processing system described in any one of (1) to (11) above, the image includes a front image obtained by capturing an image from a frontal area, in which the frontal appearance of the subject's face is visible, and a side image obtained by capturing an image from a side area different from the frontal area, in which the side of the subject's face is visible.
[0129] According to this configuration, by evaluating the facial appearance of the subject from multiple angles, it is possible to output more accurate indicators related to cognitive function.
[0130] (13) In the information processing system described in any one of (1) to (12) above, the images include a series of images in which the facial movements of the subject are visible while the subject is performing a predetermined task.
[0131] With this configuration, it becomes easier to detect signs of decline in cognitive function from the way the face moves while performing a task, making it possible to output indicators related to cognitive function with greater accuracy.
[0132] (14) In the information processing system described in (13) above, the task includes a speech task that requires the subject to speak in a specific pattern.
[0133] With this configuration, for example, it is possible to output an index showing cognitive function with higher accuracy based on facial movements during a speech task, such as mouth movements and changes in facial expression.
[0134] (15) In the information processing system described in (14) above, the speech task is defined to have the subject repeatedly speak the pattern.
[0135] This configuration can reduce the burden on the subject when performing a speech task.
[0136] (16) An information processing method, comprising the steps of the information processing system according to any one of (1) to (15) above.
[0137] (17) A program that causes at least one computer to execute each step of the information processing system according to any one of (1) to (15) above.
[0138] (18) An information processing device comprising at least one processor, the processor being configured to execute a program to perform the following steps: in the acquisition step, appearance information relating to the appearance of a subject is acquired; in the estimation step, attribute information relating to the attributes of the subject is estimated by inputting the appearance information into a pre-trained machine learning model; in the extraction step, features correlated with the cognitive function of the subject are extracted from the acquired appearance information; in the calculation step, an index indicating the state of the cognitive function of the subject is calculated based on the extracted features and predicted attribute information; and in the display step, the index indicating the state of the cognitive function is displayed.
[0139] (19) An information processing device comprising at least one processor, the processor being configured to execute a program to perform the following steps: an input receiving step receiving input of attribute information relating to the attributes of a subject; an acquisition step acquiring appearance information relating to the appearance of the subject; an extraction step extracting features correlated with the cognitive function of the subject based on the acquired appearance information; a calculation step calculating an index indicating the state of the cognitive function of the subject based on the extracted features and predicted attribute information; and a display step displaying the index indicating the state of the cognitive function.
[0140] (20) In the information processing device described in (18) or (19) above, the appearance information is configured to include visual information representing the appearance of the subject, and the visual information is a series of images including the appearance of the subject. Of course, this is not the case.
[0141] Finally, while various embodiments of the present disclosure have been described, they are presented as examples and are not intended to limit the scope of the invention. The novel embodiments may be embodied in various other forms, and various omissions, substitutions, and modifications may be made without departing from the spirit of the invention. Such embodiments and modifications are intended to be included within the scope and spirit of the invention, as well as within the scope of the inventions and their equivalents as defined in the claims. [Explanation of symbols]
[0142] 1: Information processing system 2: Information processing equipment 20: Communication bus 21: Communications Department 22: Storage section 23: Processor 3: Medical device 30: Communication bus 31: Communications Department 32: Storage section 33: Processor 34: Display section 35: Input section 4: Sound collection device 5: Imaging system 51: Front camera 52: Side camera 6: Front image 61: Face area 611 :Eye area 612: Mouth area 62: Area near the neck 7: Side view 71: Face area 72: Area near the neck 8: Evaluation result display screen 80: area 81 :Area 82 :Area 821: Evaluation results 83 :Area 831: Comparison result 832 :Item 833 :Item 834: Rating value 835: Rating value 84: Remeasurement button 9: Recording screen 91: area 92: area 93: Record button 94: Stop button 95 :Area 96 :Area 97 :Area 1000: Target D1: Front direction D2: Lateral direction G11: First non-smoking group G12: Second non-smoking group G21: First smoking group G22: Second smoking group IF1: Text information IF2: Text information IF3 :Object IF4: Text information L1: 1st layer L2: 2nd layer R1: Front area R2: Side area Th11: Threshold Th12: Threshold Th21: Threshold Th22: Threshold w: weight θ1: Inclination angle θ2: Inclination angle
Claims
1. An information processing system, at least one processor; The processor is configured to execute a program to perform the following steps: In the acquiring step, attribute information relating to attributes of the subject and image information relating to an image including at least a part of a face of the subject are acquired; In the status output step, the system outputs an index indicating the state of the subject's cognitive function based on the image information and the attribute information.
2. 2. The information processing system according to claim 1, The image includes regions corresponding to the subject's eyes or mouth.
3. 2. The information processing system according to claim 1, The attribute information includes information about the subject's living environment or lifestyle habits.
4. 4. The information processing system according to claim 3, A system wherein the information regarding the subject's living environment or lifestyle habits includes information regarding the subject's smoking.
5. 2. The information processing system according to claim 1, The system, wherein the attribute information includes information about the subject's birth.
6. 2. The information processing system according to claim 1, The system, wherein the attribute information includes information regarding the subject's medical history.
7. 7. The information processing system according to claim 6, The system, wherein the information regarding the subject's medical history includes information regarding cosmetic procedures performed on the subject.
8. 2. The information processing system according to claim 1, The system, wherein the attribute information includes information regarding the subject's oral condition.
9. 9. The information processing system according to claim 8, A system wherein the information regarding the subject's oral condition includes information regarding the subject's remaining teeth.
10. 2. The information processing system according to claim 1, In the attribute extraction step, at least a portion of the attribute information is extracted based on the image information.
11. 2. The information processing system according to claim 1, The system, wherein the image further includes an area relating to the subject's chin or throat.
12. 2. The information processing system according to claim 1, The system includes a front image, obtained by imaging from a frontal region, in which the frontal appearance of the subject's face is visible, and a side image, obtained by imaging from a side region different from the frontal region, in which the side of the subject's face is visible.
13. 2. The information processing system according to claim 1, The system, wherein the images include a sequence of images in which the subject's facial movements are visible while the subject is performing a predetermined task.
14. 14. The information processing system according to claim 13, The system, wherein the task includes a speech task that requires the subject to produce a specific pattern of speech.
15. 15. The information processing system according to claim 14, The speech task is defined to have the subject repeatedly speak the pattern.
16. An information processing method, comprising: A method comprising the steps of the information processing system according to any one of claims 1 to 15.
17. A program, A program that causes at least one computer to execute each step of the information processing system according to any one of claims 1 to 15.
18. An information processing device, at least one processor; The processor is configured to execute a program to perform the following steps: The acquiring step is configured to acquire appearance information regarding the appearance of the subject; In the estimation step, attribute information regarding attributes of the subject is estimated by inputting the appearance information into a pre-trained machine learning model; In the extraction step, a feature correlated with a cognitive function of the subject is extracted from the acquired appearance information; In the calculation step, an index indicating a state of cognitive function of the subject is calculated based on the extracted feature amount and the predicted attribute information; In the display step, the information processing device displays an index indicating the state of the cognitive function.
19. An information processing device, at least one processor; The processor is configured to execute a program to perform the following steps: In the input reception step, input of attribute information regarding the attributes of the target person is received, In the acquisition step, appearance information regarding the appearance of the subject is acquired; In the extraction step, a feature amount correlated with a cognitive function of the subject is extracted based on the acquired appearance information; In the calculation step, an index indicating a state of cognitive function of the subject is calculated based on the extracted feature amount and the predicted attribute information; In the display step, the information processing device displays an index indicating the state of the cognitive function.
20. 19. The information processing device according to claim 18, the appearance information is configured to include visual information representing the subject's appearance; An information processing device, wherein the visual information is a series of images including the subject's appearance.
Citation Information
Patent Citations
Diagnosis support information providing device and diagnosis support information providing method
JP2019055064A