Assistance device, assistance method, assistance system, and program
Patent Information
- Application Number
- JP2022137905
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-08-31
- Filing Date
- 2022-08-31
- Publication Date
- 2025-07-02
AI Technical Summary
Caregivers face challenges in grasping the overall level of their mastery of multimodal care techniques in Humanitude, which combines skills like 'look', 'touch', and 'talk', making it difficult to improve their proficiency.
A support device and method that includes sensors to detect and analyze caregiver actions, providing real-time feedback on the mastery level of multimodal skills through a display, allowing caregivers to understand their performance and areas for improvement.
Enables caregivers to assess and improve their mastery of multimodal care techniques in Humanitude, facilitating better care provision by providing actionable feedback on their skills.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a support device, a support method, a support system, and a program. More specifically, it relates to a support device, a support method, a support system, and a program suitable for improving the acquisition level of multimodal techniques in Humani-chude (registered trademark).
Background Art
[0002] There is Humani-chude, which is one of the care communication techniques. "Humani-chude" is a technique for assisting in life or providing nursing care to a person whose daily life has become difficult due to cognitive decline or the like. Humani-chude includes a multimodal care technique that combines multiple of the four skills of "looking", "touching", "talking", and "standing up" simultaneously.
[0003] There is known a support device that gives instruction information for instructing a caregiver (hereinafter referred to as a caregiver) to make eye contact with a care recipient (hereinafter referred to as a care recipient), to give a conversational explanation (live relay) of care operations, or to prompt a conversational explanation of care operations (see Patent Document 1, specifically pages 12-14, Figures 10-12).
[0004] There is also known a support device that feeds back to the caregiver the relationship for Humani-chude between the caregiver and the care recipient (see Patent Document 2, specifically paragraph
[0038] , Figure 5). For example, when a speech is detected, there is eye contact, and the head-to-head distance is within 20 cm, feedback such as "wonderful!" is given.
[0005] Furthermore, there is known a system that divides the caregiver's "looking" action into elements of distance, line of sight, and angle, provides an evaluation value to the caregiver, and provides it as the expression of the avatar's face for the "touching" action (see Non-Patent Document 1).
Prior Art Documents
Patent Documents
[0006] [Patent Document 1] Japanese Patent Publication No. 2020-201793 [Patent Document 2] Japanese Patent Publication No. 2019-185634 [Non-patent literature]
[0007] [Non-Patent Document 1] Tomoki Hiramatsu, Masaya Kamei, Daiji Inoue, Akihiro Kawamura, Qi An, and Ryo Kurazume, "Development of dementia care training system based on augmented reality and whole body wearable tactile sensor", 2020 IEEE / RSJ International Conference on Intelligent Robots and Systems (IROS 2020), IEEE, February 10, 2021. [Overview of the project] [Problems that the invention aims to solve]
[0008] Traditional Humanitude support had challenges in terms of mastering the Humanitude technique. Specifically, because the Humanitude technique combines multiple skills, similar to multimodal care techniques, it was difficult for caregivers themselves to assess the overall level of their caregiving actions. Therefore, there was room for improvement in the caregivers' own ability to enhance their multimodal care techniques.
[0009] These challenges are not limited to caregiving; they arise similarly in the acquisition of techniques that combine multiple skills, such as Humanitude Multimodal.
[0010] In view of this, the present invention aims to provide a support device, support method, support system, and program that output the level of mastery of the Humanitude multimodal technique, enabling Humanitude providers to improve their Humanitude multimodal technique.
[0011] Furthermore, another objective of the present invention is to provide a multimodal movement training device and method that demonstrates the level of proficiency in multimodal care techniques. [Means for solving the problem]
[0012] To achieve this objective, a support device according to one aspect of the present invention comprises: a plurality of detection means for detecting the status of multiple skills within Humanitude Multimodal by a Humanitude provider; an acquisition means for acquiring degree information indicating the degree of mastery of Humanitude Multimodal based on the status of the multiple skills detected by the plurality of detection means; and an output means for outputting the degree information acquired by the acquisition means to the Humanitude provider.
[0013] Another aspect of the present invention provides a support method comprising the steps of: using a central processing unit to detect the status of multiple skills within Humanitude Multimodal performed by a Humanitude provider; acquiring degree information indicating the degree of mastery of Humanitude Multimodal based on the detected status of multiple skills; and outputting the acquired degree information to a display used by the Humanitude provider.
[0014] Another aspect of the present invention is a support system comprising a server and equipment, wherein the server detects the status of multiple skills within Humanitude Multimodal provided by a Humanitude provider, acquires degree information indicating the degree of Humanitude Multimodal proficiency based on the detected skill statuses, transmits the degree information to the equipment, and the equipment is configured to output the transmitted degree information to the Humanitude provider.
[0015] A program in another aspect of the present invention, when executed by a central processing unit, causes the central processing unit to perform operations including detecting the status of multiple skills within Humanitude Multimodal as performed by the Humanitude provider, acquiring degree information indicating the degree of mastery of Humanitude Multimodal based on the detected status of multiple skills, and outputting the acquired degree information to a display used by the Humanitude provider.
[0016] Another aspect of the present invention comprises a plurality of detection means for detecting the status of multiple skills within Humanitude Multimodal, provided by a Humanitude provider; an acquisition means for acquiring degree information indicating the degree of mastery of Humanitude Multimodal based on the status of the multiple skills detected by the plurality of detection means; an accumulation means for accumulating the degree information acquired by the acquisition means; and an output means for outputting the degree information accumulated by the accumulation means to the Humanitude provider.
[0017] Another aspect of the present invention comprises the steps of: detecting the status of multiple skills within Humanitude Multimodal, performed by a Humanitude provider, using a central processing unit; acquiring degree information indicating the degree of mastery of Humanitude Multimodal based on the detected status of multiple skills; accumulating the acquired degree information; and outputting the accumulated degree information in chronological order to a display used by the Humanitude provider.
[0018] Another aspect of the present invention is a support system comprising a server and equipment, wherein the server detects the status of multiple skills within Humanitude Multimodal provided by a Humanitude provider, acquires degree information indicating the degree of Humanitude Multimodal proficiency based on the detected skill statuses, stores the acquired degree information, transmits the stored degree information to the equipment, and the equipment outputs the transmitted degree information to the Humanitude provider.
[0019] Another aspect of the present invention, when performed by a central processing unit, involves causing the central processing unit to perform operations including: detecting the status of multiple skills within Humanitude Multimodal by the Humanitude provider; acquiring degree information indicating the degree of mastery of Humanitude Multimodal based on the detected status of multiple skills; storing the acquired degree information; and outputting the stored degree information in chronological order to a display used by the Humanitude provider.
[0020] Another aspect of the present invention is a multimodal motion training device for indicating the level of mastery of multimodal care techniques in Humanitude, comprising: a plurality of sensors; means for acquiring a plurality of data from the plurality of sensors; means for determining whether a plurality of predefined conditions are met based on the acquired plurality of data; means for extracting the time in which it is determined that two or more of the plurality of predefined conditions are met simultaneously; means for calculating the ratio of the extracted time to the time from the start time of acquisition of the plurality of data to the current time; and means for outputting the calculated ratio.
[0021] Another aspect of the present invention is a multimodal motion training method for indicating the level of mastery of multimodal care techniques in Humanitude, comprising the steps of: acquiring multiple data from multiple sensors using a central processing unit; determining whether multiple predefined conditions are met based on the acquired data; extracting the time in which two or more of the multiple predefined conditions are determined to be met simultaneously; calculating the ratio of the extracted time to the time from the start time of data acquisition to the current time; and outputting the calculated ratio to a display.
[0022] Another aspect of the present invention is a system for indicating the acquisition level of multimodal operation, comprising a server and devices, wherein the server acquires a plurality of data from a plurality of sensors, determines whether the acquired plurality of data satisfy predefined conditions for each data, and when the first data among the acquired plurality of data satisfies the first condition and the second data satisfies the second condition, calculates the degree to which the first data satisfies the first condition and the degree to which the second data satisfies the second condition, calculates a numerical value indicating the acquisition level based on the first degree and the second degree, transmits the calculated numerical value to the devices, and the devices are configured to receive the calculated numerical value transmitted from the server and output the received numerical value to a display.
[0023] Another aspect of the present invention is a method for indicating the acquisition level of multimodal operation, comprising steps of: acquiring a plurality of data from a plurality of sensors by a central processing unit; determining whether the acquired plurality of data satisfy predefined conditions for each data; calculating the degree to which the first data satisfies the first condition and the degree to which the second data satisfies the second condition when the first data among the acquired plurality of data satisfies the first condition and the second data satisfies the second condition; calculating a numerical value indicating the acquisition level based on the first degree and the second degree; and outputting the calculated numerical value to a display.
Advantages of the Invention
[0024] According to the present invention, it is possible to provide an assistance device, an assistance method, an assistance system, and a program that output the acquisition level so that a humanitude provider can improve humanitude multimodal techniques.
[0025] Also, according to another invention of the present application, a caregiver can know the acquisition level of multimodal care techniques.
Brief Description of the Drawings
[0026] [Figure 1] This is a diagram showing a multimodal motion training device according to one aspect of the present invention. [Figure 2] This is a functional block diagram of a multimodal motion training device according to one aspect of the present invention. [Figure 3] This flowchart shows a method relating to individual scores for a multimodal motion training device according to one aspect of the present invention. [Figure 4] This flowchart shows a method relating to the multimodal score of a multimodal motion training device according to one aspect of the present invention. [Figure 5] This figure illustrates the action time for each skill in Humanitude according to one aspect of the present invention. [Figure 6] This is a schematic diagram illustrating an entire system according to one aspect of the present invention. [Figure 7] This figure illustrates a screen image displayed on a display according to one aspect of the present invention. [Figure 8] This figure illustrates a screen image displayed on a display according to one aspect of the present invention. [Figure 9] This is a schematic diagram illustrating an entire system according to one aspect of the present invention. [Figure 10] This is a functional block diagram of a multimodal motion training device according to one aspect of the present invention. [Modes for carrying out the invention]
[0027] The embodiments of the present invention will be described in detail below with reference to the drawings.
[0028] Figure 1 is a configuration diagram showing a multimodal motion training device according to one aspect of the present invention. The multimodal motion training device 100 is interconnected via a bus 109 and includes a CPU 101, which is a central processing unit that performs calculations related to the multimodal motion training method; a GPU 102 that performs image processing and screen output processing for the multimodal motion training method; a sensor input interface 103 that receives signals from a contact sensor, a vision sensor, an inertial sensor, and a microphone; a display 104 that outputs image data generated by the GPU 102; a speaker 105 that outputs audio data generated by the CPU 101; a main memory 106 that the CPU 101 reads and writes programs or data related to the multimodal motion training method; a secondary memory 107 that permanently stores programs or data related to the multimodal motion training method; and a network interface 108 that connects to a computer network such as a LAN.
[0029] The contact sensor is used to measure the duration of the "touching" action in Humanitude. In one embodiment, the contact sensor is a tactile sensor glove worn on the caregiver's hand when the caregiver uses the multimodal movement training device 100. Alternatively, the contact sensor may be a distributed wearable whole-body tactile sensor attached to a life-size mannequin. Here, the life-size mannequin is used when the caregiver uses the multimodal movement training device 100 and plays the role of the person being cared for. Alternatively, the caregiver may attach the distributed wearable whole-body tactile sensor directly to the person being cared for without using the life-size mannequin. The tactile sensor glove and distributed wearable tactile sensor are examples, and any contact sensor capable of inputting the necessary tactile data into the multimodal movement training device 100 may be used.
[0030] The visual sensor is used to measure the duration of the "gazing" action in Humanitude. The visual sensor can be any visual sensor that can input the necessary visual data to the multimodal motion training device 100. Examples of visual sensors include visible light sensors and infrared sensors. In one embodiment, the multimodal motion training device 100 may include a distance depth sensor for detecting the distance between the caregiver and the person being cared for.
[0031] The inertial sensor is used in the process by which the CPU 101 generates drawing data, as will be described later with reference to Figures 3 and 4. The inertial sensor can be any inertial sensor that can input the inertial data required for the multimodal motion training device 100. In one embodiment, the multimodal motion training device 100 may include an accelerometer, magnetometer, gyroscope, etc., as the inertial sensor.
[0032] The microphone is used to measure the duration of the "speaking" action in Humanitude. In one embodiment, the multimodal motion training device 100 may include a sound pressure sensor instead of a microphone.
[0033] The contact sensor, vision sensor, inertial sensor, and microphone can each be directly mounted on the multimodal motion training device 100. Alternatively, the contact sensor, vision sensor, inertial sensor, and microphone may each be located outside the multimodal motion training device 100, in which case they can be connected to the multimodal motion training device 100 via a communication connection such as Bluetooth®. In one embodiment, the vision sensor, inertial sensor, and microphone are directly mounted on the multimodal motion training device 100, while the contact sensor is located outside the multimodal motion training device 100.
[0034] The CPU 101 reads the program and data for the multimodal motion training method stored in the secondary memory 107 into the main memory 106. Using the read program and data, the CPU 101 receives tactile data, visual data, and auditory data from the sensor input interface 103, processes each of the received data as described later with reference to Figures 3 and 4, and sends the drawing data to the GPU 102 and the audio data to the speaker 105.
[0035] The GPU 102 receives drawing data from the CPU 101, performs image processing and image output processing using the received drawing data to generate image data, and transmits the generated image data to the display 104.
[0036] Here, the calculations for the multimodal motion training method are performed by a single central processing unit, but alternatively, they may be performed by multiple CPUs. In another embodiment, for example, using a different computing system from the multimodal motion training device 100, such as a server, some or all of the processing performed by CPU 101 may be performed by the central processing unit of the other computing system via the network interface 108. In another embodiment, CPU 101 processes visual and auditory data, and a different computing system from the multimodal motion training device 100 processes tactile data and transmits the processed data to the multimodal motion training device 100. Alternatively, CPU 101 may perform the processing that GPU 102 would normally perform, in which case GPU 102 is not required.
[0037] The sensor input interface 103 can receive tactile data from the contact sensor, visual data from the visual sensor, inertial data from the inertial sensor, and auditory data from the microphone, and transmit each of the received data to the CPU 101. In another embodiment, if a different computing system than the multimodal motion training device 100 processes the tactile data, the tactile data is transmitted to the different computing system without going through the sensor input interface 103.
[0038] Display 104 receives image data from GPU 102 and outputs the received image data to the screen of display 104. In one embodiment, display 104 is a wearable transparent display, such as a head-mounted display. In one embodiment, display 104 is directly mounted on the multimodal motion training device 100. In another embodiment, display 104 may be externally attached to the multimodal motion training device 100. Alternatively, display 104 can be any display that outputs image data from GPU 102 and simultaneously displays the scene on the other side of the screen.
[0039] Speaker 105 receives audio data from CPU 101 and outputs the received audio data. In one embodiment, speaker 105 is directly mounted on the multimodal motion training device 100. In another embodiment, speaker 105 may be located outside the multimodal motion training device 100. Alternatively, speaker 105 can be any device that can emit the audio data generated by CPU 101 as sound.
[0040] The main memory 106 reads programs and data necessary for processing the multimodal operation training method from the secondary memory 107 according to instructions from the CPU 101, and reads and writes the read programs and data according to instructions from the CPU 101. The main memory 106 is a storage device such as RAM (Random Access Memory), DRAM (Dynamic RAM), or SDRAM (Synchronous DRAM).
[0041] The secondary storage device 107 permanently stores programs and data necessary for processing the multimodal operation training method, and the stored programs and data are read into the main memory device 106. The secondary storage device 107 is, for example, a storage device such as a magnetic disk, flash memory, or optical disk.
[0042] The data stored in the secondary storage device 107 includes an upper threshold value (hereinafter referred to as the contact pressure threshold) that the person being cared for finds comfortable with respect to the amount of pressure the caregiver applies when touching the person being cared for.
[0043] The network interface 108 is an interface for connecting the multimodal motion training device 100 to a wired network, a wireless network, or a network combining both. In one embodiment, when a separate computing system different from the multimodal motion training device 100 processes tactile data, the network interface 108 is used to receive the data processed by the separate computing system. In one embodiment, the network interface 108 is used for program updates or data updates of the multimodal motion training method.
[0044] Contact sensors, vision sensors, inertial sensors, and microphones enable multiple detection means to detect the state of multiple skills within the Humanitude multimodal approach, provided by the Humanitude provider.
[0045] Figure 2 is a functional block diagram of a multimodal motion training device according to one aspect of the present invention.
[0046] When a program for a multimodal motion training method according to one aspect of the present invention is executed, the functional blocks on the CPU 101 include: an input data receiving unit 201 that receives data from a sensor input interface 103 or a network interface 108; a condition determination unit 202 that performs condition determination on each data received by the input data receiving unit 201, as described later with reference to Figures 3 and 4; an individual score calculation unit 203 that calculates an individual score, as described later with reference to Figures 3 and 4, based on the determination of the condition determination unit 202; a multimodal score calculation unit 204 that calculates a multimodal score, as described later with reference to Figures 3 and 4, based on the determination of the condition determination unit 202; and an output data transmission unit 205 that generates output data, as described later with reference to Figures 3 and 4, based on the calculation of the individual score calculation unit 203 or the multimodal score calculation unit 204, and transmits the generated output data to the GPU 102 or speaker 105.
[0047] The functional blocks shown in Figure 2 are an example of a multimodal motion training device. Functional blocks in a device implementing the method of the present invention are not limited to those shown above.
[0048] This configuration allows us to demonstrate the level of proficiency in multimodal care techniques within the Humanitude approach.
[0049] Figure 3 is a flowchart showing a method for individual scoring of a multimodal motion training device according to one aspect of the present invention.
[0050] When a program for a multimodal motion training method according to one aspect of the present invention is executed, the contact sensor, visual sensor, inertial sensor, and microphone begin measuring tactile data, visual data, inertial data, and auditory data, respectively.
[0051] When a life-sized doll is used to represent the person being cared for, the CPU 101 creates drawing data to display the doll's head as the head of the person being cared for, and sends the created drawing data to the GPU 102. The GPU 102 creates image data from the received drawing data and sends the created image data to the display 104. On the screen of the display 104, the caregiver can see the doll's head with the image of the person being cared for superimposed onto the scene in front of them, including the doll. Processing by the CPU 101 or GPU 102 starts from step S301.
[0052] Here, the drawing data to be displayed on the head of the person being cared for can be created by the CPU 101 using, for example, visual data acquired by an infrared sensor included in the visual sensor and inertial data acquired by an inertial sensor. The infrared sensor can track the caregiver's gaze, and the inertial sensor can accurately measure the movement of the display 104 in three dimensions, making it possible to superimpose an image of the head of the person being cared for.
[0053] In other words, infrared sensors and inertial sensors make it possible to track movement relative to the direction of the display 104, which corresponds to the direction of the caregiver's gaze.
[0054] If a real person is playing the role of the person being cared for, the CPU 101 creates rendering data to show the real person's head on the head of the person being cared for, and sends the created rendering data to the GPU 102. The GPU 102 creates image data from the received rendering data and sends the created image data to the display 104. Then, on the screen of the display 104, the caregiver can see the image of the real person's head superimposed on the scene in front of them, which includes the real person. The reason why the same processing is performed even when a real person is playing the role of the person being cared for as when a life-sized doll is used is because the real person may not be accustomed to playing the role of the person being cared for. Processing by the CPU 101 or GPU 102 starts from step S301.
[0055] In this specification, thereafter, when referring to a person being cared for displayed on the screen of display 104, there will be no distinction between a life-sized doll with an image of the person being cared for superimposed on it and an actual person.
[0056] In step S301, the CPU 101 receives tactile data measured by the contact sensor, visual data measured by the visual sensor, and auditory data measured by the microphone in real time via the sensor input interface 103, and processing proceeds to step S302.
[0057] The received tactile data includes the measurement time and data on the contact pressure at the time of measurement. The data on contact pressure includes the magnitude of the pressure exerted when the caregiver's hand touches the person being cared for; if the caregiver's hand is not touching the person being cared for, the magnitude of the pressure is zero.
[0058] More specifically, if the contact sensor is a distributed wearable whole-body tactile sensor, the received tactile data may include data regarding the measurement time, the contact pressure at the time of measurement, and the contact location at the time of measurement.
[0059] The received visual data includes the measurement time and data related to the gaze at the time of measurement. The gaze data includes data for determining whether the caregiver's gaze is aligned with the care recipient's gaze. More specifically, the gaze data includes data necessary to determine the direction and height of the caregiver's gaze using machine learning-based eye-tracking techniques well known to those skilled in the art.
[0060] The received auditory data includes the measurement time and data on sound pressure at the time of measurement. The sound pressure data includes the magnitude value of the sound pressure of the caregiver's voice; if the caregiver is not speaking, the magnitude value of the sound pressure is zero. In another embodiment, the auditory data may include data on the content of the conversation in addition to the sound pressure data.
[0061] In step S302, the CPU 101 determines whether the pressure value is greater than zero and less than or equal to the contact pressure threshold. If the result of the determination is that the pressure value is greater than zero and less than or equal to the contact pressure threshold, the process proceeds to step S303; otherwise, the process proceeds to step S304.
[0062] In step S302, the CPU 101 determines whether the direction and height of the caregiver's gaze, determined using the eye-tracking method described above, matches the direction and height of the eyes included in the image of the care recipient's head created by the GPU 102. If both the direction and height of the caregiver's gaze and the direction and height of the care recipient's eyes match, the CPU 101 further determines whether the distance between the caregiver's eyes and the care recipient's eyes is 20 cm or less. If the distance is 20 cm or less, the process proceeds to step S303; if the distance is greater than 20 cm, the process proceeds to step S304. If even one of the directions and heights of the caregiver's gaze and the care recipient's eyes do not match, the process proceeds to step S304. To determine this distance, the distance between the caregiver and the care recipient is detected by the depth sensor included in the visual sensor of this embodiment.
[0063] Here, when determining whether the distance between the caregiver's eyes and the care recipient's eyes is 20 cm or less, a threshold other than 20 cm, such as 70 cm, can be used. This threshold should be determined by considering the height and posture of the caregiver or care recipient.
[0064] In step S302, the CPU 101 determines whether the sound pressure value is greater than zero based on the sound pressure data. If the result of the determination is that the sound pressure value is greater than zero, the process proceeds to step S303; otherwise, the process proceeds to step S304.
[0065] In step S303, the CPU 101 extracts the time for each of the data related to contact pressure, line of sight, and sound pressure. Next, for each of the data mentioned above, the CPU 101 calculates the percentage of time from the time the program of the multimodal motion training method according to one aspect of the present invention was executed until the present time, and the process proceeds to step S304.
[0066] In this specification, each calculated percentage will be referred to as an individual score. The individual score for data related to contact pressure is the individual score for the "touching" skill, the individual score for data related to gaze is the individual score for the "gazing" skill, and the individual score for data related to sound pressure is the individual score for the "talking" skill. For example, the higher the individual score for data related to contact pressure, the longer the caregiver is touching the person being cared for, and the better the "touching" skill is performed. The higher the individual score for data related to gaze, the longer the caregiver's gaze and the person being cared for are aligned, and the better the "gazing" skill is performed. The higher the individual score for data related to sound pressure, the longer the caregiver is talking to the person being cared for, and the better the "talking" skill is performed.
[0067] In step S304, if there are individual scores for contact pressure data, gaze data, or sound pressure data, CPU 101 creates drawing data that includes each individual score as an individual score for each skill, and further creates audio data, and the process proceeds to step S305. For gaze data, the drawing data includes data that allows GPU 102 to generate, for example, image data of the person being cared for smiling. For gaze data, the audio data includes, for example, audio data of the person being cared for sounding happy.
[0068] To explain in more detail, for gaze data, the rendering data may include data that GPU102 can generate, for example, image data representing a smile or an expression of relief. For gaze data, the audio data may include audio data representing a smile or an expression of "hee hee." For sound pressure data, the rendering data may include data that GPU102 can generate, for example, image data representing a normal facial expression of the person being cared for. For audio data, the audio data may include audio data indicating silence, for example, that the person being cared for is silent.
[0069] If there are no individual scores for data related to contact pressure, gaze, or sound pressure, the CPU 101 creates drawing data that includes, for example, "0%" as the individual score for each skill, and the process proceeds to step S305. If the CPU 101 determines in step S302 that the pressure value is greater than the contact pressure threshold for the data related to contact pressure, the drawing data includes data that allows the GPU 102 to generate, for example, image data showing that the person being cared for has an unhappy expression. If the CPU 101 determines in step S302 that the pressure value is greater than the contact pressure threshold for the data related to contact pressure, the CPU 101 creates audio data that shows the person being cared for is angry.
[0070] More specifically, if the CPU 101 determines in step S302 that the pressure value is greater than the contact pressure threshold, the drawing data may include data that allows the GPU 102 to generate image data representing, for example, an angry facial expression of the person being cared for. If the CPU 101 determines in step S302 that the pressure value is greater than the contact pressure threshold, the CPU 101 can create audio data representing an angry "ugh" sound from the person being cared for.
[0071] To explain in more detail regarding the data on contact pressure, if, after the CPU 101 determines in step S302 that the magnitude of the pressure value is zero, then the CPU 101 determines that the magnitude of the pressure value is not zero, and if in step S303 that no time has been extracted for 10 seconds or more, then the drawing data may include data that the GPU 102 can generate, for example, image data representing a surprised expression on the face of the person being cared for. Regarding the data on contact pressure, if, after the CPU 101 determines in step S302 that the magnitude of the pressure value is zero, then the CPU 101 determines that the magnitude of the pressure value is not zero, and in step S303 that no time has been extracted for 10 seconds or more, then the CPU 101 can create audio data representing a surprised "Ah" from the person being cared for.
[0072] In step S305, the CPU 101 sends drawing data to the GPU 102 and audio data to the speaker 105. Based on the received drawing data, the GPU 102 creates image data and sends the created image data to the display 104. On the display 104, the caregiver can simultaneously view an image including the head of the person being cared for, as well as an image containing individual scores for the "gazing" skill, the "touching" skill, and the "talking" skill. Furthermore, the caregiver can hear audio from the speaker 105 that expresses the emotions of the person being cared for.
[0073] It will be readily apparent to those skilled in the art that, while the program for the multimodal motion training method is being executed, and the contact sensor, visual sensor, inertial sensor, and microphone are measuring tactile data, visual data, inertial data, and auditory data, respectively, the processing of steps S301 to S305 is constantly repeated in real time, and the caregiver can always see the latest individual scores on the screen of the display 104.
[0074] According to this embodiment, caregivers can determine to what extent they are able to perform each of the "gazing," "touching," and "talking" skills in Humanitude.
[0075] Furthermore, according to this embodiment, the caregiver can also find out how many of the three skills the person has become able to perform simultaneously.
[0076] Figure 4 is a flowchart showing a method relating to the multimodal score of a multimodal motion training device according to one aspect of the present invention.
[0077] When a program for a multimodal motion training method according to one aspect of the present invention is executed, the contact sensor, visual sensor, inertial sensor, and microphone begin measuring tactile data, visual data, inertial data, and auditory data, respectively.
[0078] The CPU 101 creates drawing data to display a life-sized doll or the head of a real person on the head of the person being cared for, and sends the created drawing data to the GPU 102. The GPU 102 creates image data from the received drawing data and sends the created image data to the display 104. Then, on the screen of the display 104, the caregiver can see a life-sized doll or the head of a real person superimposed on the image of the person being cared for within the scene in front of them, which includes the life-sized doll or the head of a real person. Processing by the CPU 101 or GPU 102 starts from step S401.
[0079] In step S401, the CPU 101 receives tactile data measured by the contact sensor, visual data measured by the vision sensor, and auditory data measured by the microphone in real time via the sensor input interface 103, and processing proceeds to step S402. Each of the received data is identical to the data described with reference to Figure 3.
[0080] In step S402, the CPU 101 determines whether the pressure value of the tactile data is greater than zero and less than or equal to the contact pressure threshold. If the result of the determination is that the pressure value is greater than zero and less than or equal to the contact pressure threshold, the process proceeds to step S403; otherwise, the process proceeds to step S405.
[0081] In step S402, the CPU 101 determines whether the direction and height of the caregiver's gaze, determined using the eye-tracking method described above, matches the direction and height of the eyes included in the image of the care recipient's head created by the GPU 102, based on the gaze data of the visual data. If the determination finds that both the direction and height of the caregiver's gaze and the direction and height of the care recipient's eyes match, the CPU 101 further determines whether the distance between the caregiver's eyes and the care recipient's eyes is 20 cm or less. If the distance is 20 cm or less, the process proceeds to step S403; if the distance is greater than 20 cm, the process proceeds to step S405. If even one of the direction and height of the caregiver's gaze and the direction and height of the care recipient's eyes do not match, the process proceeds to step S405.
[0082] In determining whether the distance between the caregiver's eyes and the care recipient's eyes is 20 cm or less, as mentioned above, a threshold other than 20 cm, such as 70 cm, can be used.
[0083] In step S402, the CPU 101 determines whether the sound pressure value of the auditory data is greater than zero. If the result of the determination is that the sound pressure value is greater than zero, the process proceeds to step S403; otherwise, the process proceeds to step S405.
[0084] In step S403, the CPU 101 extracts the time for each of the data related to contact pressure, gaze, and sound pressure. If the same time is extracted from two or more data points, the process proceeds to step S404. If the same time is not extracted from two or more data points, the process proceeds to step S405. Here, the extraction of the same time from two or more data points means that when there is a first data point that satisfies the first condition, there is at least a second data point that satisfies the second condition. The data related to contact pressure, gaze, and sound pressure occur at times when the tactile data measured by the contact sensor, the visual data measured by the visual sensor, and the auditory data measured by the microphone each satisfy the aforementioned conditions.
[0085] In step S404, the CPU 101 calculates the percentage of time from the time the program for the multimodal motion training method according to one aspect of the present invention was executed until the present time, based on the extracted time, and the process proceeds to step S405.
[0086] In other words, in step S404, a multimodal score can be calculated based on a value that combines three ratios: 1) the ratio of the length of time during which both contact pressure data and gaze data occur simultaneously to the length of time since the multimodal movement training program was executed until the present; 2) the ratio of the length of time during which both gaze data and sound pressure data occur simultaneously to the length of time since the multimodal movement training program was executed until the present; and 3) the ratio of the length of time during which both sound pressure data and contact pressure data occur simultaneously to the length of time since the multimodal movement training program was executed until the present. Various calculations such as logical AND and logical OR can be used for this combination. It is sufficient to calculate a score that indicates, or is associated with, the long duration for which the caregiver is simultaneously using two or more of the following skills: "gazing," "touching," or "talking."
[0087] In this specification, the percentage calculated in step S404 will be referred to as the multimodal score. A higher multimodal score indicates that the caregiver is simultaneously using two or more of the following skills: "gazing," "touching," or "talking." In other words, a higher multimodal score indicates that the caregiving action involves the use of multimodal care techniques for a longer period of time.
[0088] Based on the calculated multimodal score, the CPU 101 creates drawing data including the multimodal score in step S405, and the process proceeds to step S406.
[0089] If there is no multimodal score, CPU 101 creates drawing data that includes, for example, "0%" as the multimodal score, and the process proceeds to step S406.
[0090] To explain in more detail, if all individual scores are unavailable, the CPU 101 may create drawing data including, for example, "0%", and audio data for a warning sound indicating that there is no interaction between the caregiver and the person being cared for, and the process may proceed to step S406. The warning sound may include not just one type, but multiple types, such as "No communication", "No communication", and a beep. "No communication" here means a state in which there is no "looking", "touching", and "speaking".
[0091] In step S406, the CPU 101 sends drawing data to the GPU 102. Based on the received drawing data, the GPU 102 creates image data and sends the created image data to the display 104. As a result, the caregiver can simultaneously view on the screen of the display 104 an image including the head of the person being cared for and an image including the multimodal score.
[0092] If not all individual scores are available, to elaborate further, caregivers can also hear warning sounds.
[0093] While the program for multimodal motion training is executed, and the contact sensor, visual sensor, inertial sensor, and microphone measure tactile data, visual data, inertial data, and auditory data, respectively, the processing of steps S401 to S406 is continuously repeated in real time, and the caregiver can always check the latest multimodal score on the screen of display 104.
[0094] According to this embodiment, caregivers can determine to what extent they are able to perform multimodal care techniques in Humanitude.
[0095] Furthermore, according to this embodiment, caregivers can know how much more training is needed.
[0096] By combining the processes described with reference to Figure 3 and Figure 4, the caregiver can simultaneously view on the screen of display 104 an image including the head of the person being cared for that represents the person's emotions, a multimodal score, and individual scores for the "gazing" skill, the "touching" skill, and the "talking" skill, and can also hear audio representing the person being cared for.
[0097] According to this embodiment, caregivers can independently determine their own level of proficiency in multimodal care techniques.
[0098] In step S401, the CPU 101 implements means (input data receiving unit 201) for acquiring multiple data from multiple sensors. In step S402, the CPU 101 implements means (condition determination unit 202) for determining whether multiple predefined conditions are met based on the acquired multiple data. In steps S402 to S404, the CPU 101 implements acquisition means (condition determination unit 202 and multimodal score calculation unit 204) for acquiring degree information indicating the degree of mastery of Humanitude Multimodal based on the status of multiple skills detected by multiple detection means. In steps S405 and S406, the CPU 101 implements output means (output data transmission unit 205) for outputting the degree information acquired by the acquisition means to the Humanitude provider, and means (output data transmission unit 205) for outputting the calculated percentage.
[0099] The CPU 101, executing steps S402 and S403, implements means (condition determination unit 202 and multimodal score calculation unit 204) for obtaining degree information indicating the degree of acquisition of multiple individual skills. The CPU 101, executing steps S402 to S404, implements means (condition determination unit 202 and multimodal score calculation unit 204) for obtaining degree information indicating the degree of acquisition of multiple skills combined. The CPU 101, executing steps S405 and S406, implements means (output data transmission unit 205) for outputting a display or sound to the Humanitude provider that allows the user to grasp the degree of acquisition of multiple individual skills combined. The CPU 101, executing steps S405 and S406, implements means (output data transmission unit 205) for outputting a display or sound to the Humanitude provider that allows the user to grasp the degree of acquisition of multiple skills combined.
[0100] The CPU 101 executing step S403 implements a means (multimodal score calculation unit 204) for extracting time periods in which two or more of a plurality of predefined conditions are simultaneously met. The CPU 101 executing step S404 implements a means (multimodal score calculation unit 204) for calculating the ratio of the extracted time to the time from the start time of acquisition of multiple data to the current time.
[0101] Figure 5 is a diagram illustrating the action time for each skill in Humanitude according to one aspect of the present invention.
[0102] The first row of Table 501 represents the time in seconds that has elapsed since the program for the multimodal motion training method according to one aspect of the present invention was executed and each sensor of the multimodal motion training device began measuring. The second, third, and fourth rows of Table 502 represent the time in seconds that corresponds to actions using the "gazing" skill, the "touching" skill, and the "talking" skill in Humanitude, respectively.
[0103] For example, row 2 of Table 501 indicates that the caregiver's actions involved using the "gazing" skill for 10-25 seconds, 30-35 seconds, and 40-50 seconds. Furthermore, including rows 3 and 4 of Table 501, it indicates that the caregiver's actions involved using only the "gazing" skill for 15-20 seconds, a multimodal action using the "gazing" and "touching" skills for 30-35 seconds, and a multimodal action using the "gazing," "touching," and "talking" skills for 40-50 seconds.
[0104] Table 501 is an example, showing examples of the action times for each skill in Humanitude, and is an example to help understand the multimodal actions in Humanitude in a concrete way.
[0105] Figure 6 is a schematic diagram illustrating an entire system according to one aspect of the present invention. The exemplary system according to one aspect of the present invention includes a life-size doll 601 that plays the role of a person being cared for, a tactile sensor glove 602 which is a contact sensor, a computing system 603 that processes tactile data, and a head-mounted display 604 equipped with a visual sensor, an inertial sensor, and a microphone, and having a central processing unit that processes visual data, inertial data, and auditory data, and outputs image data to a screen, on which an image 621 is displayed. The data of the exemplary system according to one aspect of the present invention includes tactile data 611 transmitted from the tactile sensor glove 602 to the computing system 603, and individual score data 612 of "touch" skills calculated by the computing system 603.
[0106] In Image 621, an image of a person being cared for is shown leaning against a care bed with its back raised, and the animation of the person's facial expression and individual score data 612 are superimposed on this image. For example, the animation of the person being cared for can represent emotions such as relief, anger, surprise, sadness, and anxiety.
[0107] The CPU 101 of the multimodal motion training device 100, as described with reference to Figures 3 and 4, receives tactile data, visual data, and auditory data via the sensor input interface 103 and processes the multimodal motion training method. The central processing unit of the head-mounted display 604 in Figure 6 receives visual and auditory data via the sensor input interface 103, calculates individual scores for the "staring" skill and the "talking" skill, and receives individual score data 612 for the "touching" skill calculated by another computing system 603 via the network interface 108 and can process the multimodal motion training method.
[0108] With regard to gaze data and sound pressure data, in step S301, the central processing unit of the head-mounted display 604 receives the visual data measured by the visual sensor and the auditory data measured by the microphone in real time via the sensor input interface 103, and processing can proceed to step S302.
[0109] Regarding data related to contact pressure, in step S301, another computing system 603 receives tactile data 611 transmitted from the tactile sensor glove 602 in real time, and processing can proceed to step S302. The individual score data 612 of the "touch" skill calculated in step S303 is transmitted by the other computing system 603 to the central processing unit of the head-mounted display 604, and processing proceeds to step S304. Processing from step S304 onward can be performed by the central processing unit of the head-mounted display 604.
[0110] With regard to gaze data and sound pressure data, in step S401, the central processing unit of the head-mounted display 604 receives the visual data measured by the visual sensor and the auditory data measured by the microphone in real time via the sensor input interface 103, and processing can proceed to step S402.
[0111] Regarding data related to contact pressure, in step S401, another computing system 603 receives the tactile data 611 transmitted from the tactile sensor glove 602 in real time, and transmits the received tactile data 611 to the central processing unit of the head-mounted display 604, allowing the process to proceed to step S402. Processing from step S402 onward can be performed by the central processing unit of the head-mounted display 604.
[0112] This configuration makes it possible to load balance the processing in the CPU 101 of the multimodal motion training device 100 with the CPU of another computing system.
[0113] To explain the data regarding gaze in more detail, in step S304, if the direction and height of the caregiver's gaze and the direction and height of the care recipient's eyes match, it can be determined whether the distance between the caregiver's eyes and the care recipient's eyes is 140 cm or less. If the distance is 140 cm or less, it is further determined whether the distance between the caregiver's eyes and the care recipient's eyes is 70 cm or less. When the distance between the caregiver's eyes and the care recipient's eyes is 70 cm or less, the distance evaluation value can be set to 3; when it is greater than 70 cm and 140 cm or less, the distance evaluation value can be set to 2; and when it is greater than 140 cm, the distance evaluation value can be set to 1. If the direction and height of the caregiver's gaze and the direction and height of the care recipient's eyes do not match, the distance evaluation value can be set to 0.
[0114] Here, the two lengths, 70 cm and 140 cm, are merely examples; the distance between the caregiver's eyes and the care recipient's eyes can be any length, and the number of lengths used to determine the distance between the caregiver's eyes and the care recipient's eyes can be any number. The four values, 0, 1, 2, and 3, are merely examples; the distance evaluation values can be any value, and the number of distance evaluation values can be the same as the number of lengths used to determine the distance between the caregiver's eyes and the care recipient's eyes.
[0115] This configuration allows for the inclusion of distance evaluation in the individual scores for the "gazing" skill, enabling caregivers to independently assess the appropriateness of their distance from the person they are caring for.
[0116] In another embodiment, in step S302, the contact pressure can be evaluated in multiple stages using multiple thresholds for the data relating to the contact pressure, and in step S304, the drawing data created by the CPU 101 may include data that allows the GPU 102 to create image data representing the contact pressure using, for example, three stages: "None" indicating zero contact pressure, "Suitable" indicating appropriate contact pressure, and "Strong" indicating excessive contact pressure.
[0117] To explain in more detail in step S304, the drawing data created by the CPU 101 may also include data that allows the GPU 102 to create image data related to contact pressure, using multiple stages such as "No Contact" indicating zero contact pressure, "In Contact" indicating appropriate contact pressure, "Too Much Force" indicating excessive contact pressure, "Touch Shoulder or Back" indicating zero contact pressure only for the care recipient's shoulder and back if the contact sensor is a distributed wearable whole-body tactile sensor, and "Not Connected" indicating disconnection between the multimodal motion training device 100 and the contact sensor.
[0118] This configuration allows caregivers to independently determine whether their own touching actions are appropriate.
[0119] In another embodiment, in step S302, the caregiver's touch can be evaluated based on machine learning with respect to data related to contact pressure, and in step S304, the drawing data created by the CPU 101 may include data that allows the GPU 102 to create image data indicating whether the caregiver's touch is considered good in Humanitude. Here, a good touch in Humanitude is, for example, a touch in which the contact pressure transitions from the fingertips to the palm, in other words, a touch in which the caregiver's hand lands on the body of the person being cared for. A touch in which the contact pressure transitions from the fingertips to the palm can be determined by the time change in the position of the part where pressure is detected by the contact sensor (the pressure detection part moves from the fingertips to the palm as time passes), the time change in the distribution of the part (the distribution of pressure detection parts changes from a small area to a large area as time passes), etc. The determination described above may also be performed, for example, by the CPU 101 determining whether a predetermined criterion for "time change of pressure" is met based on the tactile data received from the tactile sensor, in the same manner as in step S302.
[0120] This configuration allows caregivers to independently determine whether their own touching actions conform to the principles of good touching in Humanitude.
[0121] In another embodiment, if the auditory data in step S302 includes data relating to the content of the conversation, the caregiver's way of speaking can be evaluated based on machine learning, and in step S304, the drawing data created by the CPU 101 may include data that allows the GPU 102 to create image data representing, for example, whether the caregiver's way of speaking is considered good in Humanitude. Here, good speaking in Humanitude is, for example, speaking calmly, slowly, and with positive language.
[0122] This configuration allows caregivers to independently determine whether their own speaking actions conform to the recommended way of speaking in Humanitude.
[0123] In step S404, to explain in more detail, if a time is extracted from the gaze data, the eye contact evaluation value can be set to 1. If the same time is extracted from both the gaze data and the sound pressure data, and then no time is extracted from the gaze data but the time for which a time is extracted from the sound pressure data is within 3 seconds, the eye contact evaluation value can be set to 1. Otherwise, the eye contact evaluation value can be set to 0.
[0124] Here, the evaluation values of 0 and 1 for eye contact are merely examples, and various values can be used. Similarly, the time of 3 seconds is merely an example, and various times can be used.
[0125] This configuration allows for a more lenient assessment of the "gazing" skill.
[0126] In step S404, to explain in more detail, if time is extracted for gaze data, the eye contact evaluation value can be set to 2. If the same time is extracted for both gaze data and sound pressure data, and the time interval for which time is extracted for gaze data but time is extracted for sound pressure data is 3 seconds or less, the eye contact evaluation value can be set to 1. Otherwise, the eye contact evaluation value can be set to 0.
[0127] Here, the eye contact evaluation values of 0, 1, and 2 are merely examples, and various values can be used. Similarly, the time of 3 seconds is also merely an example, and various times can be used.
[0128] This configuration allows caregivers to assess their level of "observational" skills on their own.
[0129] In step S303, to explain in more detail, if the time extracted from the sound pressure data is 1 second or longer, it can be determined that a conversation is taking place; otherwise, it can be determined that a conversation is not taking place. 1 second is merely an example; various time periods can be used.
[0130] This configuration allows caregivers to determine whether they are able to communicate with the person they are caring for.
[0131] In step S403, if the time is not extracted from the gaze data and the sound pressure data, more specifically in step S405, the CPU 101 can create warning sound data and generate a warning sound from the speaker 105.
[0132] This configuration allows caregivers to immediately recognize when they are unable to "look at" or "talk to" the person they are caring for.
[0133] The multimodal motion training device 100, as described with reference to Figures 3 and 4, may use either a tactile sensor glove or a distributed wearable whole-body contact sensor as the contact sensor. In another embodiment, using a tactile sensor glove as the contact sensor allows caregivers to assess the level of proficiency in multimodal care techniques in Humanitude at home.
[0134] In this specification, the multimodal movement training apparatus and method have been described in relation to indicating the level of acquisition of multimodal care techniques in Humanitude. However, the multimodal movement training apparatus and method of the present invention are not limited to indicating the level of acquisition of multimodal care techniques in Humanitude, but can also be used to indicate the level of acquisition of multimodal movements in education that requires multimodal movements, such as medical education and nursing education. In this case, the term "caregiver" can be replaced with, for example, student, trainee, etc., and the term "person receiving care" can be replaced with, for example, patient, injured person, sick person, etc.
[0135] The multimodal motion training apparatus and method of the present invention can be used, for example, to determine how well a parent is interacting with their child, such as how much they are looking at the child's face while talking to them. In this case, the term "caregiver" can be replaced with "parent," and the term "care recipient" can be replaced with "child."
[0136] Those skilled in the art will understand that steps S301-S305 and S401-S406 can be freely combined as needed.
[0137] Figure 7 is a diagram illustrating a screen image displayed on a display according to one aspect of the present invention.
[0138] The image 621 displayed on the screen of the head-mounted display 604 includes a computer graphics (CG) image 701 of the care recipient's head, generated by the CPU 101 and GPU 102 and superimposed on the screen of the head-mounted display 604, a mode display 702 indicating the operating mode of the multimodal motion training device 100, an evaluation 703 regarding the "gazing" skill, an evaluation 704 regarding the "talking" skill, an evaluation 705 regarding the "touching" skill, and an overall evaluation 706 including the multimodal score.
[0139] CG image 701, like image 621, is an illustrative image showing a care recipient actually leaning against a care bed with its back raised.
[0140] In Figure 7, evaluation 706 illustrates that the multimodal score is 23%, the individual score for the "staring" skill is 38%, the individual score for the "talking" skill is 69%, and the individual score for the "touching" skill is 1%.
[0141] For evaluations 703-705, the CPU 101 creates the drawing data in S304 based on the judgment in step S302.
[0142] Furthermore, in Figure 7, evaluation 703 indicates that the caregiver is able to "stare," evaluation 704 indicates that the caregiver is able to "talk to" the person being cared for, and evaluation 705 indicates that the caregiver's contact pressure on the person being cared for is too strong. The animation of the person being cared for in CG image 701 indicates discomfort.
[0143] Since the CG image 701 representing the facial expression of the person receiving care is configured to change in conjunction with the caregiving actions, the caregiver does not need to avert their gaze from the person receiving care to the evaluation 706 during training.
[0144] This configuration allows caregivers to train caregiving actions while monitoring the level of proficiency in multimodal care techniques in real time through image 621.
[0145] Figure 8 is a diagram illustrating a screen image displayed on a display according to one aspect of the present invention.
[0146] The image 621 displayed on the screen of the head-mounted display 604 includes a CG image 701 of the care recipient's head, generated by the CPU 101 and GPU 102 and superimposed on the screen of the head-mounted display 604; a mode display 702 indicating the operating mode of the multimodal motion training device 100; an evaluation 703 regarding the "staring" skill; an evaluation 704 regarding the "talking" skill; an evaluation 705 regarding the "touching" skill; an overall evaluation 801 including a numerical value for the multimodal score; and a warning 802 indicating that the person is unable to "stare" or "talk."
[0147] CG image 701, like image 621, is an illustrative image showing a care recipient actually leaning against a care bed with its back raised.
[0148] Evaluation 801 exemplifies a multimodal score of 22%, an individual score of 37% for the "staring" skill, an individual score of 62% for the "talking" skill, and an individual score of 1% for the "touching" skill. In Figure 8, evaluation 703 indicates an inability to "stare," evaluation 704 indicates an inability to "talk," and evaluation 705 indicates an inability to "touch."
[0149] In Figure 8, a score of 37% for the "gazing" skill allows caregivers to understand that the level of proficiency in the "gazing" skill is approximately 37%. Similarly, a score of 62% for the "talking" skill allows caregivers to understand that the level of proficiency in "talking" is 62%, and a score of 1% for the "touching" skill allows caregivers to understand that the level of proficiency in the "touching" skill is 1%.
[0150] For evaluations 703-705, the CPU 101 creates the drawing data in S304 based on the judgment in step S302.
[0151] Furthermore, in Figure 8, warning 802 represents a state where the care recipient is unable to "look at," "talk to," or "touch," and the animation of the care recipient's facial expression in CG image 701 represents anxiety. The screen image explained with reference to Figure 8 is a specific example of a state that differs from the screen image explained with reference to Figure 6.
[0152] Figure 9 is a schematic diagram illustrating an entire system according to one aspect of the present invention. The exemplary system according to one aspect of the present invention includes a life-size doll 601 that plays the role of a person being cared for, a tactile sensor glove 602 which is a contact sensor, and a smartphone 901 which is an information terminal having a central processing unit that processes tactile data, visual data, inertial data, and auditory data, and outputs image data to a screen, with an image 921 displayed on the screen of the smartphone 901. The data of the exemplary system according to one aspect of the present invention includes tactile data 611 transmitted from the tactile sensor glove 602 to the smartphone 901.
[0153] When the caregiver points the smartphone 901 at the doll 601, the screen of the smartphone 901 displays an image 921 in which the animation of the care recipient's facial expression and a multimodal score are superimposed on an image of the care recipient actually leaning against a care bed with their back raised.
[0154] The smartphone 901 includes an optical sensor and a depth sensor that can be used as a visual sensor, an acceleration sensor, a magnetic sensor, a gyroscope sensor that can be used as an inertial sensor, and a microphone. The optical sensor and depth sensor use the image sensor of the camera on the back of the smartphone 901. The image sensor of the camera mounted on the front of the smartphone 901 and machine learning-based image processing are used to create an eye-tracking sensor that tracks the caregiver's gaze.
[0155] The smartphone 901 serves as both the computing system 603 and the head-mounted display 604, as described with reference to Figure 6. Therefore, the system described with reference to Figure 9 is an example of a system with a modified combination of components from the system described with reference to Figure 6.
[0156] This configuration allows caregivers to independently assess their own level of proficiency in multimodal care techniques (i.e., their level of mastery) without the need for a computing system or head-mounted display.
[0157] According to the disclosures herein, it can be understood that a system can be configured in which the device receives data from a sensor depending on its configuration, transmits the received data to a server, the server receives the data and performs the processing in steps S301-305 and S401-S406, and the device further receives and displays the output data from the server. Those skilled in the art will understand that the device can be developed not only to include the head-mounted displays and smartphones described herein, but also to be developed specifically for a particular system that demonstrates a level of mastery of multimodal operation.
[0158] Figure 10 is a functional block diagram of a multimodal motion training device. The functional block diagram shown in Figure 10 is an alternative example of the functional block diagram shown in Figure 2.
[0159] When the program for the multimodal motion training method is executed, the functional blocks on the CPU 101 include: an input data receiving unit 201 that receives data from the sensor input interface 103 or the network interface 108; a condition determination unit 202 that makes the condition determinations described above with reference to Figures 3 and 4 for each data received by the input data receiving unit 201; an individual score calculation unit 203 that calculates the individual scores described above with reference to Figures 3 and 4 based on the determinations of the condition determination unit 202; and a function that calculates the individual scores described above with reference to Figures 3 and 4 based on the determinations of the condition determination unit 202. The system includes a multimodal score calculation unit 204 that calculates the multimodal score described above, a score storage unit 1001 that stores and accumulates the results calculated by the individual score calculation unit 203 and the multimodal score calculation unit 204, a display switching unit 1002 that switches the screen display according to instructions from the caregiver, and an output data transmission unit 205 that generates the output data described above with reference to Figures 3 and 4 based on the calculations of the individual score calculation unit 203 or the multimodal score calculation unit 204, and transmits the generated output data to the GPU 102 or speaker 105.
[0160] The score storage unit 1001 can store and accumulate the time extracted in step S403. The display switching unit 1002 can instruct the output data transmission unit 205 to output image data for displaying the performance time of each skill in Humanitude on the display 104, for example, according to the caregiver's instructions. Alternatively, the display switching unit 1002 can instruct the output data transmission unit 205 to output image data for displaying the multimodal score on the display 104, for example, according to the caregiver's instructions.
[0161] This configuration allows caregivers to know the duration of each skill or the multimodal score as needed.
[0162] A storage means (score storage unit 1001) for storing degree information acquired by the acquisition means is realized by the CPU 101 that executes steps S402 and 403, and by at least one of the main memory 106 and the secondary memory 107. An output means (realized by an output data transmission unit 205) for outputting the degree information stored by the storage means to the Humanitude provider is realized by the CPU 101 that executes steps S405 and S406. A person skilled in the art will easily understand, by looking at the functional block diagram shown in Figure 10, that the time-series operation time of each skill, as exemplified in Figure 5, can be displayed on the display 104 according to the caregiver's instructions. [Industrial applicability]
[0163] The multimodal motion training apparatus and method of the present invention can be used in various fields where multimodal motion is required. [Explanation of Symbols]
[0164] 100 Multimodal motion training devices 101 CPU 102 GPU 103 Sensor Input Interface 104 displays 105 speakers 106 Main memory 107 Secondary storage 108 Network Interfaces 109 Bus 201 Input Data Receiving Unit 202 Condition judgment section 203 Individual Score Calculation Unit 204 Multimodal score calculation unit 205 Output Data Transmission Unit
Claims
1. A plurality of detection means for respectively detecting the states of a plurality of skills among the humanitude multimodals according to a humanitude provider; An acquisition means for acquiring degree information indicating the acquisition degree of the humanitude multimodal based on the states of the plurality of skills detected by the plurality of detection means; An output means for outputting the degree information acquired by the acquisition means An assistance device characterized by comprising the above.
2. The plurality of detection means for respectively detecting the states of a plurality of skills among the humanitude multimodals according to the humanitude provider include at least two of a contact sensor for detecting contact according to the humanitude provider, a visual sensor for detecting a staring motion, and an audio sensor for detecting a speaking motion. The assistance device according to Claim 1.
3. The acquisition means for acquiring degree information indicating the acquisition degree of the humanitude multimodal based on the states of the plurality of skills detected by the plurality of detection means includes means for obtaining degree information indicating the acquisition degree for each of the plurality of skills. The assistance device according to Claim 1 or 2.
4. The acquisition means for acquiring degree information indicating the acquisition degree of the humanitude multimodal based on the states of the plurality of skills detected by the plurality of detection means includes means for obtaining degree information indicating the acquisition degree of a combination of the plurality of skills. The assistance device according to Claim 1 or 2.
5. The output means for outputting the degree information acquired by the acquisition means to the humanitude provider includes means for outputting to the humanitude provider a display or sound for grasping the degree indicating the acquisition degree for each of the plurality of skills. The assistance device according to Claim 3.
6. The output means for outputting the degree information acquired by the acquisition means to the humanitude provider includes means for outputting to the humanitude provider a display or sound that can grasp the degree indicating the acquisition degree of a combination of the plurality of skills. The assistance device according to Claim 4.
7. By a central processing unit, A step of detecting the states of a plurality of skills among the humanitude multimodals according to a humanitude provider; A step of acquiring degree information indicating the acquisition degree of the humanitude multimodal based on the detected states of the plurality of skills; A step of causing the acquired degree information to be output A support method characterized by comprising
8. A support system comprising a server and a device, wherein the server detects the states of a plurality of skills among the humanitude multimodals according to a humanitude provider, obtains degree information indicating the acquisition degree of the humanitude multimodal based on the detected states of the plurality of skills, transmits the degree information to the device, wherein the device outputs the transmitted degree information A support system characterized by being configured as such.
9. When executed by a central processing unit, the central processing unit is caused to detect the states of a plurality of skills among the humanitude multimodals according to a humanitude provider, obtain degree information indicating the acquisition degree of the humanitude multimodal based on the detected states of the plurality of skills, output the obtained degree information, A program characterized by causing an operation including
10. A plurality of detection means for respectively detecting the states of a plurality of skills among the humanitude multimodals according to a humanitude provider, an acquisition means for obtaining degree information indicating the acquisition degree of the humanitude multimodal based on the states of the plurality of skills detected by the plurality of detection means, a storage means for storing the degree information obtained by the acquisition means, an output means for outputting the degree information stored by the storage means A support device characterized by comprising
11. By a central processing unit, a step of detecting the states of a plurality of skills among the humanitude multimodals according to a humanitude provider, a step of obtaining degree information indicating the acquisition degree of the humanitude multimodal based on the detected states of the plurality of skills, a step of storing the obtained degree information, a step of outputting the stored degree information in time series A support method characterized by comprising
12. A support system comprising a server and a device, wherein the server detects the states of a plurality of skills among the humanitude multimodals according to a humanitude provider, obtains degree information indicating the acquisition degree of the humanitude multimodal based on the detected states of the plurality of skills, stores the obtained degree information, transmits the stored degree information to the device, wherein the device configured to output the transmitted degree information A support system characterized by being configured as such.
13. When executed by a central processing unit, the central processing unit is caused to detect the states of a plurality of skills among the multimodal skills provided by a humanitude provider; acquire degree information indicating the degree of acquisition of the multimodal skills based on the detected states of the plurality of skills; accumulate the acquired degree information; cause the accumulated degree information to be output in time series A program characterized by causing an operation including the above to be performed.
14. A multimodal operation training device for indicating the acquisition level of multimodal care techniques in humanitude, comprising a plurality of sensors; means for acquiring a plurality of data from the plurality of sensors; means for determining whether a plurality of predefined conditions are satisfied based on the acquired plurality of data; means for extracting the time when two or more of the plurality of predefined conditions are determined to be satisfied simultaneously; means for calculating the ratio of the extracted time to the time from the start time of acquisition of the plurality of data to the current time; means for outputting the calculated ratio A multimodal operation training device characterized by comprising the above.
15. The multimodal operation training device according to claim 14, wherein the plurality of sensors include a contact sensor, a visual sensor, and a microphone.
16. The multimodal operation training device according to claim 14, wherein at least one of the data output from the plurality of sensors is based on machine learning.
17. The multimodal operation training device according to claim 14, wherein the means for outputting the calculated ratio includes means for displaying the calculated ratio on a transmissive display.
18. The multimodal operation training device according to claim 14, wherein the means for calculating the ratio of the extracted time includes means for calculating the ratio of the time when each of the acquired plurality of data satisfies each of the plurality of predefined conditions to the time from the start time of acquisition of the plurality of data to the current time.
19. The multimodal operation training device according to claim 14, wherein the means for outputting the calculated ratio includes means for displaying an image representing human emotion.
20. The means for outputting the calculated ratio includes means for outputting voice expressing human emotions, the multimodal motion training device according to claim 14.
21. The multimodal motion training device according to claim 15, wherein the contact sensor is a tactile sensor glove.
22. A multimodal motion training method for indicating the acquisition level of multimodal care techniques in humanitude, comprising: by a central processing unit, acquiring a plurality of data from a plurality of sensors; determining whether a plurality of predefined conditions are satisfied based on the acquired plurality of data; extracting the time when it is determined that two or more of the plurality of predefined conditions are simultaneously satisfied; calculating a ratio of the extracted time to the time from the acquisition start time of the plurality of data to the current time; outputting the calculated ratio; A multimodal motion training method characterized by comprising the steps of.
23. A multimodal motion training device for indicating the acquisition level of multimodal motion in education, comprising: a plurality of sensors; means for acquiring a plurality of data from the plurality of sensors; means for determining whether a plurality of predefined conditions are satisfied based on the acquired plurality of data; means for extracting the time when it is determined that two or more of the plurality of predefined conditions are simultaneously satisfied; means for calculating a ratio of the extracted time to the time from the acquisition start time of the plurality of data to the current time; means for outputting the calculated ratio; A multimodal motion training device characterized by comprising the steps of.
24. A multimodal motion training method for indicating the acquisition level of multimodal motion in education, comprising: by a central processing unit, acquiring a plurality of data from a plurality of sensors; determining whether a plurality of predefined conditions are satisfied based on the acquired plurality of data; extracting the time when it is determined that two or more of the plurality of predefined conditions are simultaneously satisfied; calculating a ratio of the extracted time to the time from the acquisition start time of the plurality of data to the current time; outputting the calculated ratio; A multimodal motion training method characterized by comprising the steps of.
25. A system for indicating the acquisition level of multimodal operation, comprising a server and devices, wherein the server acquires a plurality of data from a plurality of sensors, determines whether the acquired plurality of data satisfy predefined conditions for each data, when a first data among the acquired plurality of data satisfies a first condition and a second data satisfies a second condition, calculates a first degree to which the first data satisfies the first condition and a second degree to which the second data satisfies the second condition, calculates a numerical value indicating the acquisition level based on the first degree and the second degree, transmits the calculated numerical value to the devices, wherein the devices receive the calculated numerical value transmitted from the server, and output the received numerical value A system characterized by being configured as described above.
26. The server creates image data based on the calculated numerical value, transmits the created image data to the devices, wherein the devices receive the image data transmitted from the server, and output an image based on the received image data The system according to claim 25, further characterized by being configured as described above.
27. A method for indicating the acquisition level of multimodal operation, comprising: acquiring a plurality of data from a plurality of sensors by a central processing unit; determining whether the acquired plurality of data satisfy predefined conditions for each data; when a first data among the acquired plurality of data satisfies a first condition and a second data satisfies a second condition, calculating a first degree to which the first data satisfies the first condition and a second degree to which the second data satisfies the second condition; calculating a numerical value indicating the acquisition level based on the first degree and the second degree; and outputting the calculated numerical value A method characterized by comprising the above steps.
28. The central processing unit further comprises creating image data based on the calculated numerical value, and outputting an image based on the created image data The method according to claim 27, further characterized by comprising the above steps.
29. The method according to claim 27, characterized in that the central processing unit is included in an information terminal.