Program, information processing system, and information processing method

By superimposing a guide image on a user's face to align the user's position, the program improves sound collection accuracy for predicting cognitive function, addressing the challenge of varying sound collection positions in existing devices.

JP2025148227AActive Publication Date: 2025-10-07SEKISUI CHEMICAL CO LTD
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
JP2024204174
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-11-22
Publication Date
2025-10-07
Estimated Expiration
2044-03-25

AI Technical Summary

Technical Problem

Existing devices face challenges in accurately acquiring audio information due to variations in sound collection positions, which can affect the prediction of cognitive function.

Method used

A program that superimposes a guide image on a captured user's face to align the user's position, allowing continuous display and improved sound collection of voice information for predicting cognitive function.

Benefits of technology

Enhances the accuracy of sound collection by optimizing the user's position relative to the device, facilitating better prediction of cognitive function through continuous alignment and improved audio information acquisition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025148227000001_ABST
    Figure 2025148227000001_ABST
Patent Text Reader

Abstract

To provide a program or the like capable of increasing a voice collection property of voice information for predicting a cognitive function.SOLUTION: According to an aspect of the present invention, a program is provided that causes at least one computer to execute the following steps. A first acquisition step includes acquiring a captured image in which a face of a user appears. A display control step includes continuously displaying, on the captured image during capturing, a superimposition image obtained by superimposing a guide part for positioning the face of the user. A second acquisition step includes acquiring voice information of voice of the user. A cognitive function can be predicted from the acquired information.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a program, an information processing system, and an information processing method. [Background technology]

[0002] Patent Document 1 discloses a device for evaluating the degree of cognitive decline from a user's speech. The device disclosed in Patent Document 1 acquires speech information from a subject by displaying an instruction image that instructs the subject to speak a predetermined phrase, and analyzes this acquired speech information to test cognitive function. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Publication No. 2019-083903 Summary of the Invention [Problem to be solved by the invention]

[0004] Since the content of the audio information can change depending on the sound collection position, it may be difficult for the device of the known technique disclosed in Patent Document 1, for example, to properly acquire the audio information.

[0005] In view of the above circumstances, the present invention provides a program or the like that can improve the sound collection of voice information for predicting cognitive function. [Means for solving the problem]

[0006] According to one aspect of the present invention, a program is provided. The program causes at least one computer to execute the following steps: In a first acquisition step, a captured image of a user's face is acquired; In a display control step, a superimposed image in which a guide portion for aligning the user's face is superimposed on the captured image being captured is continuously displayed; In a second acquisition step, voice information uttered by the user is acquired; and Cognitive function can be predicted based on the acquired information.

[0007] According to the present disclosure, it is possible to continuously display a superimposed image with a guide section superimposed on the captured image for alignment, and to provide a program or the like that can improve the sound collection ability of audio information for predicting cognitive function. [Brief explanation of the drawings]

[0008] [Figure 1] 1 is a configuration diagram illustrating an information processing system 1. FIG. [Figure 2] FIG. 2 is a block diagram showing a hardware configuration of an information processing device 2. [Figure 3] FIG. 2 is a block diagram showing the hardware configuration of a user terminal 3. [Figure 4] 3 is a functional block diagram showing functions of a processor 33 of a user terminal 3. FIG. [Figure 5] FIG. 2 is an explanatory diagram showing an example of the flow of various information for each functional block. [Figure 6] FIG. 2 is an activity diagram showing an example of the flow of information processing when a program according to the embodiment is executed. [Figure 7] 10 shows an example of the configuration of the guide portion G of the guide image d5 displayed on the display unit 34 of the user terminal 3. [Figure 8] This is an example of a superimposed image d4, showing an example of a state in which the state of the user (subject) has been optimized. [Figure 9] This shows an example of a state in which most of the face of the user (subject) is outside the guide portion G and the user's state is not optimized. [Figure 10]This shows an example of a state in which the face (mouth) of the user (subject) is far away from the user terminal 3, and the user's state is not optimized. [Figure 11] This shows an example of a state in which the user (subject)'s face is not facing forward (front), and the user's state is not optimized. [Figure 12] 1 shows an example of a screen 7 on which the results (level of cognitive function) are displayed on the display unit 34 of the user terminal 3. FIG. [Figure 13] FIG. 10 is a block diagram showing a hardware configuration of a user terminal 3 according to a modified example. DETAILED DESCRIPTION OF THE INVENTION

[0009] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS The present invention will be described below with reference to the accompanying drawings. Various features shown in the following embodiments can be combined with each other.

[0010] Incidentally, the program for realizing the software appearing in one embodiment may be provided as a non-transitory computer-readable medium, or may be provided so that it can be downloaded from an external server, or may be provided so that the program is started on an external computer and its functions are realized on a client terminal (so-called cloud computing).

[0011] Furthermore, various information processing according to an embodiment may realize input and output corresponding to the input. Here, the form of information referenced in such information processing (hereinafter referred to as reference information) is not limited as long as an output is obtained as a result of the input. The reference information may be, for example, rule-based information such as a database, a lookup table, or a predetermined function (including a decision formula such as a regression formula constructed using a statistical method), a trained model that has previously trained the correlation between input and output, or a large-scale language model that can output a desired result by inputting a prompt.

[0012] In one embodiment, a "unit" may include, for example, a combination of hardware resources implemented by a circuit in the broad sense and software information processing that can be specifically realized by these hardware resources. In one embodiment, various information is handled, and this information is represented, for example, by physical values ​​of signal values ​​representing voltage and current, high and low signal values ​​as a binary bit set consisting of 0 or 1, or quantum superposition (so-called quantum bits), and communication and calculations can be performed on a circuit in the broad sense.

[0013] Furthermore, a circuit in the broad sense is a circuit realized by at least an appropriate combination of a circuit, circuitry, processor, memory, etc. The processor may be a general-purpose processor or a dedicated circuit. That is, it includes an application specific integrated circuit (ASIC), a programmable logic device (e.g., a simple programmable logic device (SPLD), a complex programmable logic device (CPLD), and a field programmable gate array (FPGA)), etc.

[0014] [Embodiment] 1. Hardware Configuration This section explains the hardware configuration.

[0015] 1.1 Information Processing System 1 FIG. 1 is a configuration diagram illustrating an information processing system 1. The information processing system 1 includes an information processing device 2 and a user terminal 3. The information processing device 2 and the user terminal 3 are configured to be able to communicate with each other via a telecommunications line. In one embodiment, the information processing system 1 is composed of one or more devices or components. For example, if the information processing system 1 includes only the information processing device 2, the information processing system 1 can be the information processing device 2. More specifically, the information processing system 1 may include an element selected from the group consisting of the information processing device 2 and the user terminal 3. Furthermore, multiple information processing devices 2 or user terminals 3 may be used. The unselected elements may not be included in the information processing system 1, but may be electrically connected to the selected elements as external elements. These components will be described below.

[0016] 1.2 Information processing device 2 2 is a block diagram showing the hardware configuration of the information processing device 2. The information processing device 2 includes a communication bus 20, a communication unit 21, a storage unit 22, and a processor 23. The communication unit 21, the storage unit 22, and the processor 23 are electrically connected via the communication bus 20 inside the information processing device 2.

[0017] The communication unit 21 is preferably a wired communication means such as USB, IEEE1394, Thunderbolt (registered trademark), wired LAN network communication, etc., but may also include wireless LAN network communication, mobile communication such as 3G / LTE / 5G, BLUETOOTH (registered trademark) communication, etc. as needed. In other words, it is more preferable to implement it as a collection of multiple communication means. In other words, the information processing device 2 may communicate various information from the outside via the communication unit 21 and the network.

[0018] The storage unit 22 stores various pieces of information defined above. This may be implemented as a storage device such as a solid state drive (SSD), a solid state hybrid drive (SSHD), a hard disk drive (HDD), a universal serial bus (USB) flash drive (USB memory), an SD memory card, a CD, a DVD, or a Blu-ray (registered trademark) disc (BD) that stores various programs and the like related to the information processing device 2 executed by the processor 23, or as a memory such as a random access memory (RAM) that stores temporarily required information (arguments, arrays, etc.) related to program calculations. The storage unit 22 stores various programs, variables, etc. related to the information processing device 2 executed by the processor 23.

[0019] The processor 23 processes and controls the overall operations related to the information processing device 2. The processor 23 is, for example, a central processing unit (CPU) not shown. The processor 23 realizes various functions related to the information processing device 2 by reading out predetermined programs stored in the storage unit 22. In other words, information processing by software stored in the storage unit 22 is specifically realized by the processor 23, which is an example of hardware, and can be executed as each functional unit included in the processor 23. Note that the processor 23 is not limited to being single, and may be implemented with multiple processors 23 for each function. A combination of these may also be used.

[0020] 1.3 User terminal 3 The user terminal 3 is a terminal carried by a user. The user terminal 3 may be a terminal operated by a doctor, nurse, caregiver at an elderly care facility or the like, or a medical professional (hereinafter also referred to as a medical professional), or may be a terminal operated by a subject whose cognitive function is to be evaluated. The user terminal 3 may be a portable terminal such as a smartphone or tablet terminal, a computer, or any other device that can access the information processing device 2 via a telecommunications line. In the embodiment, a case where the user terminal is a portable terminal will be described as an example.

[0021] 3 is a block diagram showing the hardware configuration of the user terminal 3. The following description will be given taking the user terminal 3 as an example. The user terminal 3 includes a communication bus 30, a communication unit 31, a storage unit 32, a processor 33, a display unit 34, an input unit 35, a sound collection unit 36, and an imaging unit 37. The communication unit 31, the storage unit 32, the processor 33, the display unit 34, the input unit 35, the sound collection unit 36, and the imaging unit 37 are electrically connected via the communication bus 30 in the user terminal 3. The description of the communication unit 31, the storage unit 32, and the processor 33 is omitted because it is the same as the description of each unit in the information processing device 2.

[0022] The display unit 34 displays a screen of a graphical user interface (GUI) that can be operated by the user. The display unit 34 may be included in the housing of the user terminal 3 or may be externally attached. Specifically, the display unit 34 may be implemented as a display device such as a CRT display, a liquid crystal display, an organic EL display, or a plasma display. It is preferable that these display devices are implemented by selectively using them depending on the type of the user terminal 3.

[0023] The input unit 35 accepts operation inputs made by a medical professional or a subject. The operation inputs are transferred as command signals to the processor 33 via the communication bus 30. The processor 33 can execute predetermined control or calculations based on the transferred command signals as necessary. The input unit 35 may be included in the housing of the user terminal 3 or may be externally attached. In this embodiment, since the user terminal 3 is a portable terminal, the input unit 35 is implemented as a touch panel integrated with the display unit 34. When the input unit 35 is implemented as a touch panel in this way, the medical professional or the subject can input tap operations, swipe operations, etc. to the input unit 35. Note that, instead of a touch panel, a switch button, a mouse, a QWERTY keyboard, etc. can be used as the input unit 35.

[0024] The sound collection unit 36 ​​is a so-called microphone configured to be able to convert external sounds into signals. The microphone function of the sound collection unit 36 ​​may be an external one. The sound collection unit 36 ​​is configured to generate voice data by collecting the speech of the subject. The voice data may be temporarily stored in a memory in the user terminal and may not be stored non-volatilely in the storage unit 32. The voice data generated by the sound collection unit 36 ​​is configured to be transferable to the information processing device 2 via a network.

[0025] The sound collection unit 36 ​​collects, but is not limited to, at least sounds in the human audible range, sounds with frequencies between 20 Hz and 20,000 Hz, and converts them into electrical signals. The sound may be recorded in monaural or stereo. The sampling rate for digitally processing the sound data may be, for example, 48,000 Hz, 44,100 Hz, 32,000 Hz, 22,050 Hz, 16,000 Hz, 11,025 Hz, 11,000 Hz, 8,000 Hz, or any of the ranges of values ​​exemplified here. Increasing the sampling rate allows for more precise discretization of the temporal timing of the sound, improving the accuracy of voice recognition.

[0026] The data collected by the sound collection unit 36 ​​may be appropriately compressed by the processor 33 of the user terminal 3, and the compression format at this time may be any of MP3, AAC, WMA, Vorbis, AC3, MP2, FLAC, TAK, etc. Compression can reduce communication traffic caused by data transfer from the user terminal to the information processing device 2.

[0027] The imaging unit 37 is a so-called vision sensor (camera) configured to be able to acquire information of the external world as an image. The resolution of the imaging unit 37 is not particularly limited, and an imaging unit 37 with high spatial resolution (resolution) may be adopted. It is preferable that the imaging unit 37 be an imaging unit 37 with high temporal resolution (frame rate).

[0028] The resolution of the imaging unit 37 may be full HD or less, full HD, WQHD, 2K, 4K, 8K, or higher. Specifically, the frame rate of the imaging unit 37 is, for example, 1, 2, 45, 10, 15, 20, 25, 30, 60, 100, 125, 150, 175, 200, 225, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475, 500, 525, 550, 575, 600, 625, 650, 675, 700, 725, 750, 775, 800, 825, 850, 875, 900, 925, 950, 975, 1000, 1025, 1050, 1075, 110 The frame rate may be 0, 1125, 1150, 1175, 1200, 1225, 1250, 1275, 1300, 1325, 1350, 1375, 1400, 1425, 1450, 1475, 1500, 1525, 1550, 1575, 1600, 1625, 1650, 1675, 1700, 1725, 1750, 1775, 1800, 1825, 1850, 1875, 1900, 1925, 1950, 1975, or 2000 fps, or may be within a range between any two of the values ​​shown here. Alternatively, the frame rate may be one second or more.

[0029] 2. Functional configuration In this section, the functional configuration will be described. Fig. 4 is a functional block diagram showing the functions of the processor 33 of the user terminal 3. Information processing by software stored in the storage unit 32 can be specifically realized by hardware (specifically, the processor 33), and executed as each functional unit included in the processor 33. In other words, the information processing system 1 has each functional unit in a program.

[0030] Specifically, the processor 33 has, as its functional units, a reception unit 331, an image acquisition unit 332, a status acquisition unit 333, a status determination unit 334, a display control unit 335, an audio acquisition unit 336, an identification unit 337, a prediction information acquisition unit 338, and a presentation unit 339.

[0031] The reception unit 331 is configured to receive various types of information. The various types of information received by the reception unit 331 include, for example, basic information d1. The reception unit 331 is configured to receive information via, for example, the communication unit 21, the storage unit 22, etc., and to be able to read this information into a working memory.

[0032] The image acquiring unit 332 is configured to acquire a video (image, video) output from the imaging unit 37. In other words, the image acquiring unit 332 accepts, via the imaging unit 37, an input of a captured image d2 which is an image of the subject.

[0033] The state acquisition unit 333 is configured to acquire features of the user (subject). The user features may include various features, for example, details related to the user's face. In the embodiment, the features acquired by the state acquisition unit 333 correspond to feature information d21. The feature information d21 includes information related to the user's features in the captured image d2 (for example, information specifying where the face area is, etc.).

[0034] The state determination unit 334 has the function of determining whether the user's state (such as the positional relationship between the user and the user terminal 3) is appropriate when acquiring the user's voice information d3, based on the characteristics of the user (subject) acquired by the state acquisition unit 333.

[0035] The display control unit 335 is configured to superimpose another image on the captured image d2 captured by the imaging unit 37. In this embodiment, an image having a guide portion G (see FIG. 7) is used as the other image to be superimposed. The guide portion G is displayed on the display unit 34 as a guide for the user (subject) who is the subject to be imaged to ensure that the positional relationship between the user and the user terminal 3 falls within an appropriate range. The guide portion G will be described later.

[0036] The voice acquiring unit 336 is configured to acquire the voice output from the sound collecting unit 36. In other words, the voice acquiring unit 336 receives input of voice information d3 related to the voice uttered by the subject via the sound collecting unit 36.

[0037] The specifying unit 337 is configured to specify various pieces of information. Specifically, the specifying unit 337 specifies a presentation mode for various instructions to the subject, such as prompting the subject to correct his / her position or speak, based on the information held by the subject.

[0038] The prediction information acquisition unit 338 is configured to acquire prediction information d31 including a result of prediction regarding cognitive function. The prediction information acquisition unit 338 is configured to acquire the prediction information d31 based on the audio information d3. In particular, in predicting cognitive function, it is preferable to distinguish between mild cognitive impairment (a precursor to mild dementia) and dementia (mild, moderate, severe) among the declines in cognitive function and make the prediction.

[0039] The presentation unit 339 presents various pieces of information stored in the storage unit 32 in a form that can be understood by the presentation target. In the embodiment, the presentation by the presentation unit 339 is performed only by the user terminal 3, but is not limited to this. For example, the information may be presented to the presentation target, including the user's family, etc., via email or a communication application that allows mutual message exchange. The presentation target may be, for example, the subject, the subject's family, etc., or a medical professional. Furthermore, the subject's family, etc. may include not only the subject's family, but also relatives and those in a position to protect and support the subject (for example, an adult guardian, etc.).

[0040] For example, the presentation unit 339 may display a screen or the like containing various information in a visible manner on various terminals such as the user terminal 3. This allows various information to be presented to a person operating the various terminals such as the user terminal 3. Preferably, the presentation unit 339 controls the display of the various terminals described above to display visual information such as a screen, an image, an icon, or a message. The presentation unit 339 may generate only rendering information for displaying the visual information on the terminal. The presentation unit 339 may also present various information to the presentation subject as auditory information. Furthermore, the presentation unit 339 may present various information to the subject as uneven information (Braille information) that presents an uneven tactile sensation to the finger. In this case, it is preferable that the display is compatible with Braille.

[0041] 3. Operation of Information Processing System 1 Fig. 5 is an explanatory diagram showing an example of the flow of various information for each functional block. Fig. 6 is an activity diagram showing an example of the flow of information processing when a program according to an embodiment is executed. In this section, the flow of an information processing method and information used in the information processing method will be described mainly with reference to Figs. 5 and 6.

[0042] 3.1 Overview of Information in Information Processing Various types of information exchanged in information processing will be described with reference to FIG.

[0043] The basic information d1 includes the subject's personal ID and personal information about the subject (for example, the subject's name, age, and gender). The basic information d1 may also include biometric authentication information such as fingerprints and retina. The basic information d1 may also include information indicating whether or not the subject has a visual impairment or hearing impairment. The visual impairment and hearing impairment in the basic information d1 may be classified. For example, visual impairments may include total blindness, glaucoma, and cataracts.

[0044] The identification information d11 is information identified by the identification unit 337 based on the basic information d1. The identification information d11 can include, for example, information regarding age, as well as information indicating whether the user has a visual impairment or a hearing impairment. The identification information d11 can be used to present instructions to the user. For example, if the user has a visual impairment, it is preferable that the instructions be presented to the user in the form of voice or Braille. Furthermore, if the user has a hearing impairment, it is preferable that the instructions be presented to the user in the form of a display. In this way, the identification information d11 can be used to provide appropriate instructions taking into account the user's situation.

[0045] The captured image d2 is an image acquired through the imaging unit 37. The captured image d2 is used to correct the position of the subject when speaking in order to predict cognitive function. In the embodiment, the captured image d2 is a moving image in order to grasp the position of the subject in real time.

[0046] The feature information d21 is information related to the features of the user in the captured image d2, and is used to understand the state of the user (the position and orientation of the face). That is, the feature information d21 is information related to the state of the user that is acquired by the state acquisition unit 333 analyzing the captured image d2. The feature information d21 can be set to various contents. For example, the feature information d21 can be information for identifying the area of ​​the user's face, information for identifying the outline of the user's face, information for identifying features of the user's face (e.g., eyes, nose, mouth, eyebrows, etc.), information for identifying any other characteristic part of the face, etc.

[0047] The state determination information d22 is information indicating whether the position of the subject when speaking is appropriate in predicting cognitive function. If the position is appropriate (OK), an instruction to speak is presented to the user. If the position is not appropriate (NG), an instruction to correct the state is presented to the user.

[0048] The voice information d3 is information for acquiring the prediction information d31, and corresponds to the speech voice acquired from the sound collection unit 36 ​​in this embodiment.

[0049] The prediction information d31 is information including the result of prediction regarding the cognitive function of the user (subject) and can be acquired based on the voice information d3. The prediction information d31 includes the result of prediction of cognitive function by utterance of a fixed phrase (in one embodiment, "katama"). The prediction result of the prediction information d31 is level information indicating the degree of the cognitive function of the subject, and the form is not particularly limited, and may be expressed, for example, as a number or may be indicated as high, medium, or low. In addition, classification such as mild cognitive impairment risk, mild dementia risk, moderate dementia risk, and severe dementia risk may be performed.

[0050] The superimposed image d4 is generated based on the captured image d2 and a guide image d5 in which the guide portion G (see FIG. 7) is displayed. Specifically, the superimposed image d4 is an image in which the guide image d5 is superimposed on the captured image d2.

[0051] Guide image d5 has a guide portion G (see FIG. 7) as a displayed image. Guide image d5 is displayed superimposed on the front side (top side) of captured image d2, so that the portion (pixels) where guide portion G of guide image d5 is displayed hides the portion of captured image d2 and becomes invisible. Note that the transparency of guide portion G of guide image d5 may be adjustable, and in this case, the content of captured image d2 can be seen through the portion where guide portion G is displayed.

[0052] 3.2 Overview of Information Processing An overview of information processing will be described with reference to Figures 5 and 6. In the embodiment, the information processing system 1 is configured to be able to execute each activity (step) related to the information processing method described below. That is, the information processing system 1 includes at least one processor (processor 33 in one embodiment) that can execute a program so that each step of the program is performed. In other words, the program according to the embodiment causes at least one computer (processor 33 in one embodiment) to execute each step corresponding to the activity described below, and the information processing method includes each step of the program. Note that the order of the processes can be changed as appropriate, multiple processes may be executed simultaneously, or some processes may be omitted.

[0053] (Activity A001: Receiving basic information d1) The receiving unit 331 receives basic information d1 (such as the subject's personal ID). The receiving unit 331 receives the basic information d1, thereby enabling the subject to be authenticated. When the receiving unit 331 receives the basic information d1, it is stored in the storage unit 32.

[0054] (Activity A002: Start acquiring image d2) When the processor 33 activates the imaging unit 37, the image acquisition unit 332 acquires a captured image d2 capturing an image of the face of the user (subject). This activity A002 is an example of a first acquisition step. This activity A002 is performed continuously. Specifically, it is preferable that this activity A002 be performed continuously from the start of acquisition of the captured image d2 until acceptance of the voice input of activity A008, which will be described later, is completed. The timing of activating the imaging unit 37 (the timing of starting acquisition of the captured image d2) is arbitrary, and may be before or during input of the basic information d1, for example. However, since it is necessary to generate the superimposed image d4, this is before the start of displaying the superimposed image d4.

[0055] The captured image d2 is used to understand the user's state in the subsequent activity A004, and in this regard, the image acquisition unit 332 may perform preprocessing to acquire (generate) the captured image d2. For example, since the captured image d2 output from the imaging unit 37 may contain noise or unnecessary information, the image acquisition unit 332 may acquire the captured image d2 after performing processing such as filtering to remove unnecessary information in advance.

[0056] (Activity A003: Start displaying superimposed image d4) The display control unit 335 continuously displays a superimposed image d4 in which a guide portion G for aligning the face of the user (subject) is superimposed on a captured image d2 being captured. In other words, the display control unit 335 generates the superimposed image d4 based on the captured image d2 continuously acquired in the activity A002 and the guide image d5 stored in the storage unit 32, and continuously displays the superimposed image d4 on the display unit 34. This activity A003 is an example of a display control step.

[0057] According to this aspect, the following can be achieved. Elderly users are often unfamiliar with how to use devices. According to this aspect, a superimposed image d4, in which a guide portion G is superimposed on a captured image d2, is displayed, prompting the user to move relative to the position of the guide portion G in the superimposed image d4 (movement of the user or movement of the terminal). This relative movement is easy even for elderly users. This relative movement optimizes the sound collection position, thereby improving the sound collection of audio information for predicting cognitive function. Furthermore, in this aspect, the guide portion G (superimposed image d4) is continuously displayed. This device may be used by people who have difficulty focusing their gaze (face direction) in one direction (e.g., people with declining cognitive function tend to do this). However, the continuous display of the guide portion G (superimposed image d4) helps the user focus their gaze, making it easier to keep their mouth facing the sound collection portion 36, thereby improving the sound collection of audio information for predicting cognitive function.

[0058] This activity A003 is performed continuously, preferably from the start of displaying the superimposed image d4 until acceptance of voice input in activity A008 (described later) is completed.

[0059] (Activity A004: Obtaining user characteristics) The state acquisition unit 333 can acquire feature information d21 obtained by analyzing the features of the user (subject) based on the captured image d2. The feature information d21 includes information for identifying the user's face, etc.

[0060] (Activity A005: Determine user status) The state determination unit 334 has a function of determining whether the user's state (such as the positional relationship between the user and the user terminal 3) is appropriate when acquiring the user's voice information d3, based on the feature information d21 acquired by the state acquisition unit 333. In other words, the state determination unit 334 generates state determination information d22 indicating the result of determining whether the user's state is appropriate. The determination may be made from one or more perspectives. In the embodiment, multiple determination perspectives are provided, one of which is whether the user's face is out of the guide portion G. This determination perspective will be described later.

[0061] (Activity A006: Correction Instructions) If it is determined that the user's state is not appropriate (if the determination in activity A005 is NG), the process proceeds to activity A006. The presentation unit 339 presents instructions to the user to adapt the user's state so that the audio information can be acquired appropriately. In other words, activity A006 presents instructions to the user to prompt the user to correct the user's state in the captured image d2. According to this embodiment, the user can understand how to make adjustments from the presentation results (for example, visual, auditory, braille, etc. information).

[0062] The content of the instruction depends on the state of the user. For example, if the user's face is outside the guide portion G, an instruction to fit the face into the frame is displayed (see FIG. 9). Activity A006 is an example of a presentation step.

[0063] The presentation mode is predetermined and specified by the specification unit 337. Specifically, when basic information d1 is accepted in activity A001, the specification unit 337 generates specification information d11 using the basic information d1 and specifies the presentation mode when presenting some information to the subject.

[0064] For example, if the basic information d1 indicates that the subject has a visual impairment, the identification unit 337 can identify that the presentation format is to be auditory information. In this case, auditory information can be output from an output unit (e.g., a speaker) not shown in the figure of the user terminal 3. Furthermore, if the basic information d1 indicates that the subject has a hearing impairment, the identification unit 337 can identify that the presentation format is to be visual information. In this case, visual information can be output using the display unit 34 of the user terminal 3. Furthermore, if the subject's age is equal to or greater than a predetermined threshold, there is a possibility that the subject's hearing or vision may be impaired. Therefore, the identification unit 337 can identify that both auditory information and visual information should be used as the presentation format. The volume of the auditory information and the font size of the visual information may be adjusted depending on the subject's age. Of course, the presentation format is not limited to these. For example, even if the subject has a visual impairment, a medical professional may be accompanying the subject. In this case, it is sufficient for the medical professional to be able to understand the instructions. In the embodiment, the presentation unit 339 presents information to the user (subject) based on the identified presentation format.

[0065] In this way, the instructions in activity A006 (an example of a presentation step) are presented using at least one of visual information and auditory information. According to this aspect, if a user has a visual impairment, for example, the auditory information can help the user understand how to make adjustments, and if a user has a hearing impairment, for example, the visual information can help the user understand.

[0066] When activity A006 is completed, the process moves to activity A004, where the user's status is acquired again.

[0067] (Activity A007: Present voice input instructions, etc.) If it is determined that the user's condition has been optimized (if the determination in activity A005 is OK), the process proceeds to activity A007. The presentation unit 339 presents that the user's condition has been optimized. For example, in the case of visual information, this can be realized by displaying the word "OK" (see FIG. 8). In addition, the presentation unit 339 presents an instruction to encourage the user to speak. For example, one example of the instruction can be "Please speak the phrase" (see FIG. 8). The presentation format is the same as that described in the above-mentioned activity A006. In other words, the instruction in activity A007 is presented based on the presentation format previously specified by the specification unit 337.

[0068] (Activity A008: Accepting voice input) The user (subject) will speak after confirming the instructions in activity A007. Then, the voice acquisition unit 336 acquires voice information d3 uttered by the user (subject). This makes it possible to predict cognitive function based on the acquired information. For example, it is possible to predict cognitive function by inputting the acquired voice information d3 into a trained model. Note that the prediction of cognitive function is performed in activity A009, which will be described later. Furthermore, in the embodiment, the voice acquisition unit 336 acquires voice information d3 from the sound collection unit 36. Activity A008 is an example of a second acquisition step.

[0069] Furthermore, it is preferable that the display control (display control step) of the superimposed image started in activity A003 continues even while acquiring voice information d3 in activity A008. In other words, the superimposed image d4 is continuously displayed while acquiring voice information d3 uttered by the user (subject) in activity A008 (an example of a second acquisition step). According to this aspect, it is easy to continuously suppress misalignment of the face even while speaking, that is, the sound collection position is continuously optimized.

[0070] (Activity A009: Obtaining forecast information d31) The prediction information acquisition unit 338 acquires (generates) prediction information d31 based on the voice information d3 and a preset trained model. That is, in the embodiment, the prediction information acquisition unit 338 has a function of predicting the level of cognitive function of a certain subject by analyzing the subject's speech.

[0071] (Activity A010: Presenting forecast information d31) In the activity A010, the presentation unit 339 presents (notifies) the subject of the prediction information d31. The presentation format is the same as that described for the activity A006. That is, the presentation of the activity A010 is performed based on the presentation format previously specified by the specification unit 337.

[0072] 3.3 Details of information processing The above-mentioned information processing will be described in detail with reference to FIGS.

[0073] (Regarding guide part G, etc.) Fig. 7 shows a configuration example of the guide section G of the guide image d5 displayed on the display section 34 of the user terminal 3. Note that in Fig. 7, the user terminal 3 is a portable terminal, and the rectangular display section 34, the sound collection section 36 arranged below the display section 34, and the imaging section 37 arranged above the display section 34 are shown schematically, but the configuration is not limited to this, and is merely an example.

[0074] Area Rg1 on the display unit 34 is an area where visual information related to, for example, activity A006 (modifying the user's state) or activity A007 (instructing voice input) is displayed. Area Rg2 on the display unit 34 is an area where visual information (for example, OK or NG) is displayed to inform the user whether or not the user's state is appropriate in activity A007. The visual information displayed in areas Rg1 and Rg2 may be changed depending on the user's state (whether the user's state is appropriate or not). For example, if the user's state is appropriate, "OK" may be displayed in red in area Rg2, and if it is not appropriate, "NG" may be displayed in blue in area Rg2.

[0075] The guide portion G has a contour shape portion G1 that corresponds to the contour of the user's (subject's) face. In addition to the contour shape portion G1, the guide portion G also has a contour shape portion G2 and a center portion G3. The contour shape portion G1 has a shape that follows the contour of the subject's face, who is the subject to be imaged. The contour shape portion G2 has a shape that follows the contour of the subject's neck, who is the subject to be imaged. The center portion G3 is a guide that makes it easier for the subject to position their face within the contour shape portion G1; for example, by positioning the center portion G3 over the nose or forehead (the position between the eyebrows), the subject's face can be smoothly placed within the contour shape portion G1. It is optional whether the guide portion G has the contour shape portion G2 and the center portion G3.

[0076] In this embodiment, the contour portion G1 is used to determine the user's position (determine the state of activity A005), and the contour portion G2 and center portion G3 are displayed solely for user friendliness. Of course, the contour portion G2 and center portion G3 can also be used for this determination. In this case, the instruction to correct the user's state can be changed accordingly.

[0077] The guide portion G may have only the contour shape portion G2 or only the center portion G3. The user may be able to select which of the contour shape portion G1, the contour shape portion G2, and the center portion G3 to use.

[0078] Furthermore, the components of the guide portion G (in the embodiment, the outline shape portion G1, the outline shape portion G2, and the center portion G3) may be partially separated or may be represented by dashed lines. Furthermore, the components of the guide portion G do not necessarily need to be displayed in black, and any color can be used. It is preferable that they are easy for the user (subject) to understand, and for example, they may flash or change color.

[0079] By displaying the guide section G, the subject can make an utterance for predicting cognitive function while checking his or her own position displayed on the display section 34. As a result, the subject's position when speaking is optimized, and the sound collection of voice information for predicting cognitive function is improved.

[0080] (Regarding the user's status and its determination) In one example of the embodiment, the state of the user (subject) includes at least one of the positional relationship (first perspective) between the guide unit G and the user's face (described later) and the orientation of the user's face (second perspective). According to this aspect, it is possible to acquire the audio information d3 taking into consideration the state of the user (subject) that affects the sound collection performance. In the embodiment, state determination (determination by the state determination unit 334) is performed from these first and second perspectives. Note that it is not necessarily necessary to perform state determination from both perspectives, and this is merely an example of the embodiment.

[0081] Fig. 8 is an example of a superimposed image d4, showing an example of a state in which the state of the user (subject) is optimized. Fig. 9 shows an example of a state in which a large part of the user's (subject's) face is outside the guide portion G, and the user's state is not optimized. Fig. 10 shows an example of a state in which the user's (subject's) face (mouth) is far from the user terminal 3, and the user's state is not optimized.

[0082] The first aspect (positional relationship) corresponds to whether the position of the user's face is displaced vertically or horizontally relative to the user terminal 3, and whether the user is too close (or too far) from the user terminal 3. The state determination related to the first aspect (positional relationship) can include, for example, a state determination a that the area of ​​the user's face within the contour shape portion G1 is equal to or greater than a predetermined ratio with respect to the area of ​​the contour shape portion G1. The state determination related to the first aspect (positional relationship) can also include, for example, a state determination b that the area of ​​the user's face within the contour shape portion G1 is unevenly distributed within the contour shape portion G1. The state determination related to the first aspect (positional relationship) can also include, for example, a state determination c that the contour shape portion G1 includes a pair of eye areas (which may be substituted with eyebrows), a nose area, and a mouth area. As described above, controlling the distance between the user's sound collection unit 4 and the mouth and the orientation of the mouth to be within a certain range can improve the accuracy of cognitive function diagnosis during sound collection. Since the distance between the user terminal 3 and the sound collection unit 4 may vary depending on the type of terminal, a screen for inputting the type of terminal may be displayed before the image is captured.

[0083] For example, if the state determinations b and c are satisfied but the state determination a is not satisfied, it is conceivable that the user is too far away from the image capture unit 37 (see FIG. 10), and by using these state determinations a to c, this situation can be resolved. That is, it is possible to prompt the user to correct their state by presenting an instruction to correct their state ("please move closer" in the example of FIG. 10) to the user. Also, for example, if the state determinations a and b are satisfied but the state determination c is not satisfied, it is conceivable that the user is too close to the image capture unit 37 (not shown), and by using these state determinations a to c, this situation can be resolved. That is, it is possible to prompt the user to give an instruction to correct their state (for example, presenting an instruction such as "please move away").

[0084] Furthermore, by using state determinations a to c, it is possible to determine that, as shown in FIG. 9, the distance to the user terminal 3 is fine, but the face is not contained within the guide portion G and is misaligned (left and right position in the example of FIG. 9). For example, FIG. 9 shows a case where state determinations a and c are satisfied, but state determination b is not. In other words, in FIG. 9, the left region of the face is unevenly positioned within the contour shape portion G1. Therefore, by using state determinations a to c, this situation can be resolved. In other words, it is possible to prompt the user to correct their state by presenting them with an instruction to correct their state (in the example of FIG. 9, "Please place your face within the frame").

[0085] As described above, in the embodiment, the positional relationship (first aspect) between the guide unit G and the user's face is based on how the user's (subject's) face fits into the contour shape unit G1. If the fit is inappropriate, instructions are presented to the user based on the superimposed image d4 to prompt the user to correct the positional relationship, as shown in the instructions in FIGS. 9 and 10. According to this aspect, for example, if the face area is small, the distance is far, so the face can be moved closer, thereby improving sound collection accuracy. Also, for example, if the face is shifted left / right or up / down from the guide unit, the face is not positioned in the center, so the face can be positioned in the center (within the guide unit), thereby improving sound collection accuracy.

[0086] 11 shows an example of a state in which the user's (subject's) face is not facing forward (front), and the user's state is not optimized. Even if the user satisfies the state determination from the first perspective, if the direction of the mouth is not facing the user terminal 3, the sound collection performance may be reduced. For this reason, in the embodiment, the user's state also includes the perspective of the direction of the face.

[0087] The second perspective (face direction) corresponds to whether the mouth is facing the user terminal 3. The state determination related to the second perspective (face direction) can be performed, for example, using the positions and shapes of various facial features (e.g., eyes, nose, mouth, eyebrows, etc.) or the coordinate positions of any characteristic part of the face. Note that the state determination related to the second perspective is preferably performed after the determination of the first perspective is completed.

[0088] As described above, the embodiment also takes into consideration the second viewpoint (face direction). Note that information on the guide unit G is not necessarily required for determining the face direction. Therefore, the state determination unit 334 may determine the state based on the captured image d2 or the superimposed image d4. In other words, the above-described activity A005 and activity A006 (an example of a presentation step) can be implemented by presenting instructions to the user (subject) to prompt them to correct the face direction based on the captured image or the superimposed image. According to this aspect, even if the face is within the guide unit G, the sound collection accuracy can be further improved by orienting the face forward.

[0089] As described above, the state of the user includes both the positional relationship (first perspective) and the facial direction (second perspective), but of course this is not limited to this and only one of these perspectives may be used. Also, another perspective may be further provided. As one example, another perspective may be to detect that the user is nervous and may not be able to speak appropriately and issue instructions to relax, thereby relieving the tension, or to detect behavior such as playfully deforming the face in front of the user terminal 3, which may prevent the user from speaking appropriately, and to urge the user to refrain from such behavior.

[0090] In addition, in the example of the embodiment, the state determinations a to c have been described as examples, but the present invention is not limited to these. As long as it is possible to determine the positional relationship (first viewpoint) between the guide unit G and the user's face and the orientation of the user's face (second viewpoint), various methods can be used (for example, for an aspect in which a trained model is used, see Modification 2 described later).

[0091] (Regarding obtaining forecast information d31) In activity A009, an example of a method in which the prediction information acquisition unit 338 acquires prediction information d31 including the results of predictions regarding cognitive function from the speech information d3 will be described. One method is to measure and evaluate fluctuations in the speed of each "katama" repetition. Furthermore, this method may be combined with various speech formats, not limited to the fixed speech. Furthermore, it may be further combined with various evaluation methods, such as fluctuations in the movement of the feet and hips while walking, or the movement of the gaze.

[0092] (Details of the presentation of forecast information d31) FIG. 12 shows an example of a screen 7 displaying the results (level of cognitive function) displayed on the display unit 34 of the user terminal 3. Presentation in activity A010 can be implemented by displaying a screen 7 such as that shown in FIG. 12. The screen 7 includes an area 71. The area 71 is an area where prediction information d31 including the results of prediction regarding the user's cognitive function is displayed. In the example shown in FIG. 12, the prediction information d31 displays "Your level is XXX," where XXX is, for example, a number indicating the level of cognitive function.

[0093] 4. Variations The information processing system 1 may have the following configuration.

[0094] (Variation 1: Transmitting Device 4) FIG. 13 is a block diagram showing the hardware configuration of a user terminal 3 according to a modified example. As shown in FIG. 13, the user terminal 3 is connected to a transmitting device 4 as an external device. The transmitting device 4 can be configured, for example, by a millimeter-wave / ultrasonic laser transmitting device. The transmitting device 4 is a device for acquiring distance information (a distance corresponding to the distance between the sound collecting unit 36 ​​and the mouth of the user (subject)). The distance information may be the actual distance between the sound collecting unit 36 ​​and the mouth, or a distance that is correlated with the distance between the sound collecting unit 36 ​​and the mouth. As the correlated distance, for example, the distance from the transmitting device 4 to any part of the face can be used. Of course, this is not limited to this. The distance information is preferably displayed on the display unit 34.

[0095] Although not shown in the figures, it is preferable that an activity for acquiring distance be added before the timing of determining the state of activity A005 described in Fig. 6. This activity (an example of a third acquisition step) acquires distance information corresponding to the distance between sound collection unit 36 ​​and the user's mouth. Activity A005 then uses this distance information to determine the state.

[0096] If the distance between the sound collection unit 36 ​​and the user's mouth is outside a predetermined range, the presentation unit 339 presents the instruction "please move closer" or "please move away." That is, in activity A006 (an example of a presentation step), an instruction to prompt the user to correct the distance is presented to the user based on the distance information.

[0097] In the embodiment, as described in Fig. 10, the determination is made using state determinations a to c, so it can be considered that distance information is acquired indirectly. On the other hand, if distance information is acquired directly using transmitting device 4 as in this modification, it can be expected that the user's speaking position will be more appropriate. In addition, the determination in activity A005 may employ both the determination regarding the positional relationship (distance) using state determinations a to c described in the embodiment and the determination regarding the distance using transmitting device 4 according to this modification, or it may employ either one of them.

[0098] (Variation 2: State Determination Method) In the embodiment, the case where the user's state is determined using an algorithm using state determinations a to c, etc., has been described as an example, but the present invention is not limited thereto. For example, the state may be determined using a trained model (neural network) that outputs the user's state when the superimposed image d4 is input. That is, the state determination may be, for example, normal, NG because the user is outside the frame, or NG because the user is too far away. In the case of Modification 2, the activities A004 and A005 in FIG. 6 described in the embodiment can be replaced with the state determination method of Modification 2. Furthermore, such a trained model may be included in any component of the information processing system 1, or may be provided as an external service not included in the information processing system 1. In either case, the prediction information acquisition unit 338 acquires prediction information d31 including the results of a cognitive function test from the audio information d3.

[0099] (Variation 3: Acquisition of prediction information d31) In the embodiment, a case where cognitive function is predicted using the voice information d3 has been described as an example, but the present invention is not limited to this. In activity A008, while receiving the user's (subject's) speech, a captured image (which may be a still image or a video) of the user while he or she is speaking may be acquired. For example, when a fixed phrase such as "katama" is spoken, if a specific movement (e.g., trembling of a body part such as the lips) is observed, it is possible to suggest the possibility of cognitive decline due to dementia with Lewy bodies or Parkinson's disease. Therefore, in activity A009, the prediction information acquisition unit 338 analyzes not only the voice information d3 but also the captured image during the speech, thereby making it possible to predict whether the decline in cognitive function is due to dementia with Lewy bodies or Parkinson's disease.

[0100] [others] In the first modification, the distance information is acquired using the transmitting device 4, but the present invention is not limited to this. For example, the distance information may be acquired by predicting it from information that can be known in advance, such as the angle of view of the camera, the size of the guide part G, and the size of the face in the captured image (or parts such as the eyes, nose, and mouth).

[0101] In the embodiment, an example has been described in which the text "OK" or "NG" is displayed and presented in the region Rg2. However, the present invention is not limited to this. For example, instead of or in addition to displaying "OK," most of the screen of the display unit 34 may be displayed in blue. Similarly, instead of or in addition to displaying "NG," most of the screen may be displayed in red to actively encourage correction of the status. Furthermore, the determination of the status (such as OK or NG) may not be limited to a display on the screen, but may also be notified by voice. The voice notification method may use, for example, a bone conduction microphone. As described above, by providing both a display and a voice notification, accurate voice data can be easily acquired even for subjects with impaired visual or auditory functions.

[0102] In the embodiment, a guide image having a guide portion G is used to generate the superimposed image d4, but the present invention is not limited to this. For example, a film corresponding to the guide portion G may be attached to the display unit 34 of the user terminal 3. In this case, the superimposed image d4 can be obtained by displaying the captured image d2 on the display unit 34. Note that when the guide portion G is a film, there are factors such as errors in the size of the user terminal 3 and the position when the film is attached, so it is preferable to use the guide image d5 as described in the embodiment.

[0103] In the embodiment, when the user terminal 3 presents instructions or the like as auditory information, the speaker of the user terminal 3 is used, but this is not limited to this and earphones or headphones may also be used. Furthermore, if the subject is elderly, it is expected that they may have hearing loss, so it is more preferable to use bone conduction earphones or bone conduction headphones.

[0104] The user terminal 3 as an information processing device may be in an on-premise form or in a cloud form. In the cloud form, for example, the above functions and processes may be provided in the form of SaaS (Software as a Service) or cloud computing.

[0105] In the above embodiment, the user terminal 3 performs various storage and control operations, but multiple external devices may be used instead of the user terminal 3. That is, various information and programs may be distributed and stored in multiple external devices using blockchain technology or the like.

[0106] At least one of the devices included in the information processing system 1 may be installed outside Japan. For example, the information processing device 2 or server may be installed outside Japan, and the user terminal 3 may be installed in Japan. Similarly, a medical professional may access the information processing device 2 installed in Japan from outside Japan using his or her own user terminal 3. According to such an embodiment, a more convenient experience can be provided to the user through various management forms.

[0107] Furthermore, for example, some of the functional units (FIG. 4) of the user terminal 3 may be included in the processor 23 of the information processing device 2. For example, at least one of the state acquisition unit 333, the state determination unit 334, the display control unit 335, and the predicted information acquisition unit 338 may be included in the functional units of the processor 23 of the information processing device 2, and the functional units described in FIG. 4 may be distributed between the information processing device 2 and the user terminal 3.

[0108] Furthermore, it may be provided in the following aspects.

[0109] (1) A program that causes at least one computer to execute the following steps: a first acquisition step acquires an image capturing a user's face; a display control step continuously displays a superimposed image in which a guide section for aligning the user's face is superimposed on the captured image during capture; and a second acquisition step acquires voice information uttered by the user and enables prediction of cognitive function based on the acquired information.

[0110] According to this aspect, the following can be achieved. Elderly users are often unfamiliar with how to use devices. According to the above aspect, a superimposed image in which a guide section is superimposed on a captured image is displayed, prompting the user to make a relative movement (movement of the user or movement of the terminal) that corresponds to the position of the guide section of the superimposed image. Even elderly people can easily perform such relative movement. This relative movement optimizes the sound collection position, thereby improving the sound collection of audio information for cognitive function testing. Furthermore, in this aspect, the guide section (superimposed image) is continuously displayed. This device may be used by people who have difficulty focusing their gaze (face direction) in one direction (e.g., people with declining cognitive function tend to do this). However, the continuous display of the guide section (superimposed image) allows the user to focus their gaze, making it easier to keep their mouth facing the sound collection section, thereby improving the sound collection of audio information for cognitive function testing.

[0111] (2) In the program according to (1) above, in the presenting step, an instruction for prompting the user to correct the state of the user in the captured image is presented to the user.

[0112] According to this aspect, the user can understand how to make adjustments from the presented results (visual or auditory information).

[0113] (3) In the program described in (2) above, the state of the user includes at least one of the positional relationship between the guide unit and the user's face and the direction of the user's face.

[0114] According to this aspect, it is possible to acquire voice information while taking into consideration the state of the user that affects the sound collection ability.

[0115] (4) In the program described in (3) above, the guide unit has a contour shape portion corresponding to the contour of the user's face, the positional relationship is based on how the user's face fits into the contour shape portion, and in the presentation step, the instructions to prompt the user to correct the positional relationship are presented to the user based on the superimposed image.

[0116] According to this aspect, for example, if the face area is small, the distance is far, so by moving it closer, the sound collection accuracy can be improved. Also, for example, if the face is shifted left / right or up / down from the guide unit, the face is not positioned in the center, so by positioning it in the center (bringing it inside the guide unit), the sound collection accuracy can be improved.

[0117] (5) In the program described in (3) or (4) above, in the presentation step, the instruction to prompt the user to correct the direction of the user's face is presented to the user based on the captured image or the superimposed image.

[0118] According to this aspect, even if the face is contained within the guide portion, the sound collection accuracy can be further improved by directing the face to face forward.

[0119] (6) A program according to any one of (3) to (5) above, wherein in the third acquisition step, distance information corresponding to the distance between a sound collection unit and the user's mouth is acquired, and in the presentation step, the instruction to prompt the user to correct the distance is presented to the user based on the distance information.

[0120] According to this aspect, it is expected that the user's speaking position can be made more appropriate using distance information.

[0121] (7) The program according to any one of (2) to (6) above, wherein the instruction in the presenting step is presented by at least one of visual information and auditory information.

[0122] According to this aspect, if the user has a visual impairment, for example, the auditory information can help the user understand how to make adjustments, and if the user has a hearing impairment, for example, the visual information can help the user understand.

[0123] (8) In the program described in any one of (2) to (7) above, in the display control step, the superimposed image is continuously displayed while the voice information uttered by the user is being acquired in the second acquisition step.

[0124] According to this aspect, it is easy to continuously suppress the displacement of the face even while speaking, that is, the sound collection position is continuously optimized.

[0125] (9) An information processing system, comprising at least one processor capable of executing the program described in any one of (1) to (8) above so as to perform each step of the program.

[0126] According to this embodiment, it is possible to provide a technology that is used in cognitive function testing that is easier to use.

[0127] (10) An information processing method, comprising the steps of the program described in any one of (1) to (8) above.

[0128] According to this embodiment, it is possible to provide a technology that is used in cognitive function testing that is easier to use. Of course, this is not the case.

[0129] Finally, while various embodiments of the present disclosure have been described, they are presented as examples and are not intended to limit the scope of the invention. The novel embodiments may be embodied in various other forms, and various omissions, substitutions, and modifications may be made without departing from the spirit of the invention. Such embodiments and modifications are intended to be included within the scope and spirit of the invention, as well as within the scope of the inventions and their equivalents as defined in the claims. [Explanation of symbols]

[0130] 1: Information processing system 2: Information processing equipment 20: Communication bus 21: Communications Department 22: Storage section 23: Processor 3: User terminal 30: Communication bus 31: Communications Department 32: Storage section 33: Processor 331: Reception 332: Image acquisition unit 333: Status acquisition unit 334: Status determination unit 335: Display control unit 336: Voice acquisition unit 337: Specific part 338: Inspection information acquisition unit 339: Presentation section 34:Display section 35: Input section 36: Sound collection section 37: Imaging unit 4: Transmitting device 7: Screen 71 :Area G: Guide part G1: Contour shape part G2: Contour shape part G3: Center section Rg1: area Rg2: area d1: Basic information d11 :Specific information d2: Captured image d21:Feature information d22: Status determination information d3: Audio information d31: Forecast information d4: Superimposed image d5: Guide image

Claims

1. A program, causing at least one computer to perform the following steps: In the first acquisition step, a captured image of a user's face is acquired; In the display control step, a superimposed image in which a guide portion for aligning the face of the user is superimposed on the captured image during image capture is continuously displayed; In the second acquiring step, voice information uttered by the user is acquired; A program that makes it possible to predict cognitive function based on the information acquired.

2. 2. The program according to claim 1, In the presenting step, an instruction for prompting the user to correct the state of the user in the captured image is presented to the user.

3. 3. The program according to claim 2, The state of the user includes at least one of a positional relationship between the guide unit and the user's face and a direction of the user's face.

4. 4. The program according to claim 3, the guide portion has a contour shape portion corresponding to the contour of the user's face, the positional relationship is based on how the user's face fits into the contour shape portion; The presenting step presents the instruction to the user to prompt the user to correct the positional relationship based on the superimposed image.

5. 4. The program according to claim 3, In the presenting step, the instruction for prompting the user to correct the orientation of the user's face is presented to the user based on the captured image or the superimposed image.

6. 4. The program according to claim 3, In the third obtaining step, distance information corresponding to the distance between the sound collecting unit and the user's mouth is obtained; The presenting step presents the instruction to the user to prompt the user to correct the distance based on the distance information.

7. 3. The program according to claim 2, The instruction in the presenting step is presented by at least one of visual information and auditory information.

8. 3. The program according to claim 2, In the display control step, the superimposed image is continuously displayed while the voice information uttered by the user is being acquired in the second acquisition step.

9. An information processing system, A system comprising at least one processor capable of executing the program according to any one of claims 1 to 8 so as to perform each step of the program.

10. An information processing method, comprising: A method comprising the steps of the program according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Speech recognition device

    JP1995064595A

  • Subject identification method, subject identification system, blood pressure measurement state determination method, blood pressure measurement state determination device, and blood pressure measurement state determination program

    JP2018023768A

  • Cognitive function evaluation apparatus, cognitive function evaluation system, cognitive function evaluation method, and program

    JP2019083903A

  • Information processing device, information processing system, and program

    JP2020081088A

  • Biometric authentication implementation

    JP2020525868A