Nursing care support device, nursing care support method, and program

JPWO2024190274A5Pending Publication Date: 2025-11-07
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025506617
Authority / Receiving Office
JP · JP
Patent Type
Applications
Priority Date
2024-02-16
Filing Date
2024-02-16
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing nursing care support systems fail to evaluate the emotional state of care recipients during voice interactions, limiting caregivers' understanding of the care recipients' happiness, energy, distress, or sadness, which is crucial for effective nursing care and reducing caregiver burden.

Method used

A nursing care support device and method that acquires video data with audio, uses voice recognition to estimate emotions, and combines this with video-based emotion estimation to generate a summary sentence that includes emotional expressions, enabling caregivers to objectively assess the care recipient's emotional state.

Benefits of technology

Enables caregivers to easily understand the emotional situation of care recipients, improving nursing care by incorporating emotional evaluations beyond keyword-based assessments.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

A care assistance device 10 is provided with a data acquisition unit 11 that acquires moving image data with sound of a care recipient, a first emotion estimation unit 12 that estimates an emotion of the care recipient on the basis of a result of voice recognition processing performed on a voice in the moving image data, a second emotion estimation unit 13 that estimates an emotion of the care recipient on the basis of the moving image data, an expression determination unit 14 that determines an expression representing the state of the care recipient, using the emotion estimated by the first emotion estimation unit 12 and the emotion estimated by the second emotion estimation unit 13, and a summary generation unit 15 that generates a summary from the result of the voice recognition processing and adds the determined expression to the generated summary.
Need to check novelty before this filing date? Find Prior Art

Description

Care support device, care support method, and computer-readable recording medium

[0001] The present disclosure relates to a care support device and a care support method for supporting caregivers in care settings, and further to a computer-readable recording medium on which a program for realizing these is recorded.

[0002] In recent years, systems for supporting care have been proposed to reduce the burden on caregivers. For example, Patent Literature 1 discloses a system that conducts a dialogue with a care recipient and presents information obtained through the dialogue to the caregiver.

[0003] Specifically, the system disclosed in Patent Document 1 first executes a question-based voice dialogue between the robot and the care recipient. The system then performs voice recognition on the dialogue to generate text data and determines whether a trigger keyword is present in the text data. If a trigger keyword is present, the system then identifies a corresponding question and uses the keyword as the answer to the identified question.

[0004] The system then compares the answer to the question with a table in which points are set for each answer, calculates the points for the question, and uses the calculated points as evaluation information on the care situation.The system then transmits the question and its evaluation information to the caregiver's terminal.

[0005] Japanese Patent Application Laid-Open No. 2021-162929

[0006] As described above, the system disclosed in Patent Document 1 allows a caregiver to objectively evaluate the condition of a care recipient. However, the evaluation information is obtained only from keywords extracted from speech-recognized text, and the care recipient's emotions during the voice dialogue are not evaluated at all. In other words, the system disclosed in Patent Document 1 has difficulty in grasping the condition of a care recipient, such as whether the care recipient appears happy, energetic, in pain, or sad.

[0007] On the other hand, in caregiving, it is important to understand the situation based on the emotions of the person being cared for, such as whether they seem happy, energetic, in pain, or sad. It is only when the system can understand these emotions that the burden on caregivers can be reduced.

[0008] An example of an objective of the present disclosure is to enable a caregiver to easily understand the situation of a care recipient in light of their emotions.

[0009] In order to achieve the above object, a care support device according to one aspect of the present disclosure comprises: a data acquisition unit that acquires video data with audio of a care recipient; a first emotion estimation unit that estimates an emotion of the care recipient based on a result of speech recognition processing of the audio in the video data; a second emotion estimation unit that estimates an emotion of the care recipient based on the video data; an expression determination unit that determines an expression that expresses a state of the care recipient using the emotion estimated by the first emotion estimation unit and the emotion estimated by the second emotion estimation unit; and a summary generation unit that generates a summary from the result of the speech recognition processing and adds the determined expression to the generated summary.

[0010] Furthermore, in order to achieve the above object, a care support method according to one aspect of the present disclosure is characterized by comprising: a data acquisition step of acquiring video data with audio of a care recipient; a first emotion estimation step of estimating the emotion of the care recipient based on a result of speech recognition processing for the audio in the video data; a second emotion estimation step of estimating the emotion of the care recipient based on the video data; an expression determination step of determining an expression that expresses the condition of the care recipient using the emotion estimated in the first emotion estimation step and the emotion estimated in the second emotion estimation step; and a summary generation step of generating a summary from the result of the speech recognition processing and adding the determined expression to the generated summary.

[0011] Furthermore, in order to achieve the above object, a computer-readable recording medium according to one aspect of the present disclosure is characterized in that it records a program including instructions for causing a computer to execute: a data acquisition step of acquiring video data with audio of a care recipient; a first emotion estimation step of estimating the emotion of the care recipient based on the result of a voice recognition process for the audio in the video data; a second emotion estimation step of estimating the emotion of the care recipient based on the video data; an expression determination step of determining an expression that expresses the state of the care recipient using the emotion estimated in the first emotion estimation step and the emotion estimated in the second emotion estimation step; and a summary generation step of generating a summary from the result of the voice recognition process and adding the determined expression to the generated summary.

[0012] As described above, according to the present disclosure, it is possible for a caregiver to easily grasp the situation of a care recipient in consideration of their emotions.

[0013] FIG. 1 is a configuration diagram showing a schematic configuration of a first example of a care support device. FIG. 2 is a configuration diagram showing the configuration of the first example of the care support device in more detail. FIG. 3 is a diagram showing an example of text data obtained by speech recognition and an estimated emotional expression. FIG. 4 is a diagram showing an example of an estimated emotion comparison table constituting an emotional expression reverse lookup dictionary. FIG. 5 is a diagram showing an example of an emotional expression table constituting an emotional expression reverse lookup dictionary. FIG. 6 is a diagram showing an example of a generated summary sentence. FIG. 7 is a flow diagram showing a first example of the operation of the care support device. FIG. 8 is a configuration diagram showing the configuration of a second example of the care support device. FIG. 9 is a diagram showing an example of a question type list. FIG. 10 is a diagram showing an example of a document expression table. FIG. 11 is a flow diagram showing a second example of the operation of the care support device. FIG. 12 is a block diagram showing an example of a computer that realizes the care support device.

[0014] First Embodiment Hereinafter, a care support device, a care support method, and a program according to a first embodiment will be described with reference to FIGS.

[0015] [Device Configuration] First, the schematic configuration of a care support device in embodiment 1 will be described with reference to Fig. 1. Fig. 1 is a configuration diagram showing the schematic configuration of a first example of a care support device. The care support device 10 in embodiment 1 shown in Fig. 1 is a device for creating a document indicating the condition of a care recipient and supporting a caregiver.

[0016] As shown in FIG. 1 , the care support device 10 includes a data acquisition unit 11, a first emotion deduction unit 12, a second emotion deduction unit 13, an expression determination unit 14, and a summary sentence generation unit 15.

[0017] Of these, the data acquisition unit 11 acquires video data with audio of the care recipient. The first emotion deduction unit 12 estimates the emotion of the care recipient (hereinafter referred to as "text estimated emotion") based on the result of speech recognition processing on the audio in the video data. The second emotion deduction unit 13 estimates the emotion of the care recipient (hereinafter referred to as "video estimated emotion") based on the video data.

[0018] The expression determination unit 14 determines an expression expressing the state of the care recipient, using the text estimated emotion estimated by the first emotion deduction unit 12 and the video estimated emotion estimated by the second emotion deduction unit 13. The summary generation unit 15 first generates a summary from the result of the speech recognition processing. Furthermore, the summary generation unit 15 adds the expression determined by the expression determination unit 14 to the generated summary.

[0019] In this way, the care support device 10 generates a summary from video data with audio of the care recipient, and can further add to the summary a situation derived from the care recipient's emotions (e.g., happy, energetic, distressed, sad, etc.). This allows the caregiver to easily understand the situation based on the care recipient's emotions from the summary.

[0020] Next, the configuration and functions of the care support device 10 will be specifically described with reference to Fig. 2. Fig. 2 is a configuration diagram showing the configuration of the first example of the care support device in more detail.

[0021] As shown in FIG. 2 , the care support device 10 includes a speech recognition unit 16, a learning model 17, and an emotional expression reverse lookup dictionary 18 in addition to the data acquisition unit 11, the first emotion deduction unit 12, the second emotion deduction unit 13, the expression determination unit 14, and the summary sentence generation unit 15 described above.

[0022] In the first embodiment, the data acquiring unit 11 acquires video data with audio of the care recipient from the video database 30 or the imaging device 31. In the first embodiment, the video data is acquired by capturing a conversation situation consisting of questions to the care recipient and answers from the care recipient. That is, the video captures the care recipient responding to several questions. The data acquiring unit 11 inputs the acquired video data to the voice recognition unit 16 and the second emotion estimation unit 13.

[0023] The video database 30 stores video data with audio of the care recipient in advance. The imaging device 31 is positioned to capture an image of the care recipient and outputs the video data of the care recipient. The data acquisition unit 11 may acquire video data from both the video database 30 and the imaging device 31.

[0024] The speech recognition unit 16 performs speech recognition processing on speech included in the input video data, generates text data indicating the results of the speech recognition, and inputs the generated text data to the first feeling deduction unit 12 and the summary sentence generation unit 15. Note that the speech recognition processing in the first embodiment can be performed using existing technology.

[0025] In the first embodiment, the first feeling deduction unit 12 first extracts keywords from text data obtained by speech recognition processing. Next, the first feeling deduction unit 12 determines which of preset states the care recipient falls into based on the extracted keywords, and deduces the determination result as the feeling of the care recipient (text estimated feeling).

[0026] Specifically, the first emotion deduction unit 12 can deduce emotions using the emotion analysis library "ML-Ask" disclosed in the following Reference Document 1. In this case, the first emotion deduction unit 12 uses a dictionary in which emotions (joy, anger, sadness, fear, shame, like, disgust, excitement, relief, surprise) are set in advance for each keyword to deduce the emotion corresponding to each extracted keyword. Then, the first emotion deduction unit 12 determines the emotion (negative, positive, neutral) expressed by the entire text data using the emotion estimated for each keyword. The determined emotion becomes the text estimated emotion. [Reference Document 1] http: / / arakilab.media.eng.hokudai.ac.jp / ~ptaszynski / repository / mlask.htm

[0027] In the first embodiment, the second feeling estimation unit 13 inputs video data to the learning model 17 and estimates the emotion of the care recipient (video estimated emotion) using the output result of the learning model 17. The output result of the learning model 17 is input to the expression determination unit 14.

[0028] The learning model 17 is constructed by machine learning the relationship between human face images, human voices, and human emotions. A specific example of the learning model 17 is the learning model shown in the following Reference 2. In this case, the learning model 17 outputs the output result, i.e., the video estimated emotion, which is the type of emotion such as joy, anger, sadness, fear, disgust, surprise, and contempt, and a numerical value indicating the degree of the emotion. [Reference 2] https: / / www.cac.co.jp / trends / trend27.html

[0029] Furthermore, the learning model shown in the following Reference Document 3 may be used as the learning model 17. In this case, the learning model 17 outputs, as an output result, one of "joy," "anger," "sadness," and "normal," and a numerical value indicating the degree of the corresponding emotion on a scale of 1 to 10. [Reference Document 3] https: / / www.agi-web.co.jp /

[0030] The learning model 17 is, for example, a neural network. The learning model 17 is actually implemented by a machine learning program executed on a computer. The learning model 17 may also be implemented in a device (computer) separate from the care support device 10.

[0031] In the first embodiment, the expression determination unit 14 determines an expression that expresses the situation of the care recipient by comparing the text estimated emotion estimated by the first emotion deduction unit 12 and the video estimated emotion estimated by the second emotion deduction unit 13 with a table that defines the relationship between a person's emotion and an expression that expresses a person's situation (hereinafter referred to as a “situation expression”). In the first embodiment, the table is prepared as an emotional expression reverse lookup dictionary 18.

[0032] The processing of the expression determination unit 14 will now be described in more detail with reference to Figures 3 to 5. Figure 3 is a diagram showing an example of text data obtained by speech recognition and an estimated emotional expression. Figure 4 is a diagram showing an example of an estimated emotion comparison table constituting the emotional expression reverse lookup dictionary. Figure 5 is a diagram showing an example of an emotional expression table constituting the emotional expression reverse lookup dictionary.

[0033] The example of Fig. 3 shows text data obtained by speech recognition, a text estimated emotion estimated by the first emotion deduction unit 12, and a video estimated emotion estimated by the second emotion deduction unit 13. As described above, the video shows the care recipient responding to several questions. Therefore, in Fig. 3, the text data is shown divided into questions to the care recipient and the corresponding responses from the care recipient. Furthermore, in the example of Fig. 3, the text estimated emotion is expressed as one of NEGATIVE, POSITIVE, and NEUTRAL, and the video estimated emotion is expressed by the type of emotion (joy, anger, sadness, fear, disgust, surprise, contempt, calm) and a numerical value.

[0034] The emotional expression reverse dictionary is composed of an estimated emotion comparison table shown in Fig. 4 and an emotional expression table shown in Fig. 5. The estimated emotion comparison table shown in Fig. 4 defines estimated emotions (hereinafter referred to as "combined estimated emotions") obtained from combinations of the text estimated emotions and the emotion types in the video estimated emotions. The emotional expression table shown in Fig. 5 defines situational expressions corresponding to the estimated emotions.

[0035] Therefore, the expression determination unit 14 first identifies a combined estimated emotion by comparing the emotion types of the text estimated emotion and the video estimated emotion with the estimated emotion comparison table shown in Fig. 4. Then, the expression determination unit 14 identifies a corresponding situation expression by comparing the identified combined estimated emotion with the emotion expression table shown in Fig. 5, and inputs the identified situation expression to the summary sentence generation unit 15.

[0036] The summary generation unit 15 first generates a summary using the text data input from the speech recognition unit 16. Specifically, the summary generation unit 15 generates a summary of the input text data by using, for example, a learning model that has learned the relationship between a document and its summary through machine learning. Note that the method for generating a summary is not particularly limited, and existing technology may be used as the method.

[0037] Next, the summary generation unit 15 adds the situation expression received from the expression determination unit 14 to the generated summary. Fig. 6 is a diagram showing an example of a generated summary. In the example of Fig. 6, the summary generation unit 15 adds the situation expression and the sentence "The care recipient answered" to the summary to form a final summary. In Fig. 6, the situation expression is the part surrounded by a box.

[0038] The summary generator 15 then transmits the final summary to the terminal device 40 used by the caregiver. As a result, the summary shown in FIG. 6 is presented to the caregiver. According to the summary shown in FIG. 6, the situation derived from the care recipient's emotions, such as "smiling," "wry smile," "surprised," and "irritated," is added. Therefore, the caregiver can easily understand the situation based on the care recipient's emotions from the summary.

[0039] [Device Operation] Next, the operation of the care support device 10 will be described with reference to FIG. 7. FIG. 7 is a flow diagram showing a first example of the operation of the care support device. In the following description, reference will be made to FIGS. 1 to 6 as appropriate. In addition, in the first embodiment, a care support method is implemented by operating the care support device 10. Therefore, the description of the care support method in the first embodiment will be replaced by the following description of the operation of the care support device 10.

[0040] 7 , first, the data acquiring unit 11 acquires video data with the voice of the care recipient from the video database 30 or the imaging device 31 (step A1). The data acquiring unit 11 also inputs the acquired video data to the voice recognition unit 16 and the second emotion estimation unit 13.

[0041] Next, the speech recognition unit 16 performs speech recognition processing on the speech included in the video data acquired in step A1, and generates text data indicating the speech recognition result (step A2). The speech recognition unit 16 also inputs the generated text data to the first emotion deduction unit 12 and the summary sentence generation unit 15.

[0042] Next, the first emotion estimation unit 12 extracts keywords from the text data obtained in step A2, determines which of the pre-set states the care recipient falls into based on the extracted keywords, and estimates the determination result as the emotion of the care recipient (text estimated emotion) (step A3).

[0043] Next, the second emotion estimation unit 13 inputs the video data acquired in step A1 into the learning model 17, and estimates the emotion of the care recipient (video estimated emotion) using the output result of the learning model (step A4). The output result of the learning model 17 is input to the expression determination unit 14.

[0044] Next, the expression determination unit 14 determines an expression (situation expression) that expresses the situation of the care recipient by comparing the text estimated emotion estimated in step A3 and the video estimated emotion estimated in step A4 with an emotional expression reverse dictionary (step A5). The expression determination unit 14 also inputs the determined situation expression to the summary sentence generation unit 15.

[0045] Next, the summary generator 15 generates a summary using the text data generated in step A2, and then adds the situation expression determined in step A5 to the generated summary to produce a final summary (step A6).The summary generator 15 then transmits the generated final summary to the terminal device 40 used by the caregiver (step A7).The final summary is then displayed on the screen of the terminal device 40.

[0046] The summary displayed on the screen includes information about the situation based on the care recipient's emotions, such as "smiling," "wry smile," "surprised," and "irritated." Therefore, according to the first embodiment, the caregiver can easily understand the situation based on the care recipient's emotions from the summary displayed on the screen.

[0047] [Program] The program in the first embodiment may be a program that causes a computer to execute steps A1 to A7 shown in Fig. 7. By installing and executing this program on a computer, the care support device 10 and the care support method in the first embodiment can be realized. In this case, the processor of the computer functions as a data acquisition unit 11, a first emotion deduction unit 12, a second emotion deduction unit 13, an expression determination unit 14, a summary sentence generation unit 15, and a speech recognition unit 16 to perform processing. Examples of the computer include a general-purpose PC, a smartphone, and a tablet terminal device.

[0048] The program in Embodiment 1 may be executed by a computer system constructed by a plurality of computers. In this case, for example, each computer may function as any one of the data acquiring unit 11, the first emotion deduction unit 12, the second emotion deduction unit 13, the expression determining unit 14, the summary generating unit 15, and the speech recognizing unit 16.

[0049] Second Embodiment Next, a care support device, a care support method, and a program according to a second embodiment will be described with reference to FIGS.

[0050] [Device Configuration] First, the configuration of a care support device in embodiment 2 will be described with reference to Fig. 8. Fig. 8 is a configuration diagram showing the configuration of a second example of a care support device. Like the care support device 10 in embodiment 1, the care support device 20 in embodiment 2 shown in Fig. 8 is a device for creating a document indicating the condition of a care recipient and supporting a caregiver.

[0051] 8 , the care support device 20 according to the second embodiment has the same configuration as the care support device 10 according to the first embodiment. That is, the care support device 20 includes a data acquiring unit 11, a first emotion deduction unit 12, a second emotion deduction unit 13, an expression determining unit 14, a summary generating unit 15, a speech recognizing unit 16, a learning model 17, and an emotional expression reverse lookup dictionary 18.

[0052] Also in the second embodiment, similarly to the first embodiment, the video data is obtained by capturing a conversation situation consisting of questions to the care recipient and answers from the care recipient. Therefore, also in the second embodiment, the first feeling deduction unit 12 and the second feeling deduction unit 13 deduce the feeling of the care recipient for each question.

[0053] However, in the second embodiment, in addition to the above configuration, the care support device 20 further includes a question type estimation unit 21, an impression generation unit 22, a question type list 23, and a document expression table 24. The following will specifically explain the differences from the first embodiment.

[0054] The question type estimation unit 21 estimates the type of question from the video data. In the second embodiment, the question type estimation unit 21 first executes the same speech recognition process as the speech recognition unit 16 to generate text data from the video data. Then, the question type estimation unit 21 identifies text data corresponding to each question from the generated text data. Then, the question type estimation unit 21 estimates the type of each question using the question type list 23. Note that the question type estimation unit 21 can also use the text data generated by the speech recognition unit 16.

[0055] Specifically, the question type list 23 is a list in which question samples and their types are registered in association with each other, as shown in Fig. 9. Fig. 9 is a diagram showing an example of the question type list. For each question, the question type estimation unit 21 calculates the similarity between the corresponding text data and each sample, and determines the type of the sample with the highest similarity as the question type.

[0056] An example of the similarity to be calculated is cosine similarity. Further, methods for calculating cosine similarity include existing methods, such as a method of vectorizing documents and a method of using a learning model that learns relationships between words through machine learning.

[0057] For example, for the question "Activities you have given up" among the questions shown in Figure 3, the question type estimation unit 21 calculates the highest similarity with "Are there any activities or roles you have currently given up on?" in the question type list 23, and therefore estimates the type to be "Role / purpose in life."

[0058] The impression generating unit 22 generates a document showing an impression about the conversation by using the type of question estimated by the question type estimating unit 21. The impression summarizes the overall response of the care recipient and is useful for understanding the condition of the care recipient.

[0059] Specifically, the impression generation unit 22 extracts the question type that has been estimated the most frequently from among the estimated question types. Furthermore, for each question, the impression generation unit 22 extracts the emotion that has been estimated the most frequently from among the emotions estimated by either or both of the first emotion deduction unit 12 and the second emotion deduction unit 13. In the second embodiment, the impression generation unit 22 identifies, for each question, both the text estimated emotion estimated by the first emotion deduction unit 12 and the video estimated emotion estimated by the second emotion deduction unit 13, and extracts the emotion that has been estimated the most frequently. In other words, the impression generation unit 22 extracts the emotion that has been estimated the most frequently, regardless of the type of the text estimated emotion and the video estimated emotion.

[0060] Next, the impression generation unit 22 generates a document showing an impression about the conversation by applying the extracted questions and the extracted emotions to a preset document expression table 24. As shown in FIG. 10 , the document expression table 24 is a table that defines questions, emotions, and corresponding sentence expressions. FIG. 10 is a diagram showing an example of the document expression table. As shown in FIG. 10 , document templates are defined according to the number of extracted emotions.

[0061] For example, suppose that the emotion that was estimated the most frequently has not been extracted, but "swallowing" has been extracted as the type of question that was estimated the most frequently. In this case, the impression generation unit 22 selects "This time, the conversation was centered on (question type)" as the sentence template, and inputs the type "swallowing" into it. As a result, the impression generation unit 22 generates the document "This time, the conversation was centered on swallowing."

[0062] Also, assume that only one emotion that has been estimated the most frequently is extracted. In this case, the impression generation unit 22 selects, as a document template, "This time, the conversation was centered on (question type), with (emotion) as the focus," or "This time, the conversation was centered on (question type), with (emotion) as the focus." Note that which of these templates is selected may be randomly selected.

[0063] If only one emotion is extracted, the impression generation unit 22 determines an emotional expression appropriate for the extracted emotion. For example, the impression generation unit 22 determines an emotional expression appropriate for the extracted emotion using a table that defines the relationship between emotions and emotional expressions, or a learning model that has learned the relationship between emotions and emotional expressions through machine learning. In this case, the impression generation unit 22 may determine an emotional expression using the emotional expression table shown in FIG. 5.

[0064] For example, if "troubled" is determined as the emotion expression suitable for the emotion most frequently estimated, and "sleep" is extracted as the type of question most frequently estimated, the impression generating unit 22 generates a document such as "This time, the conversation was focused on sleep, and the conversation was troubled."

[0065] Furthermore, suppose that "awkward" is determined as the emotion expression suitable for the emotion most frequently estimated, and "meal" is extracted as the type of question most frequently estimated. In this case, the impression generating unit 22 generates a document saying, "This time, the conversation about the meal seemed awkward."

[0066] Furthermore, it is assumed that two or more emotions that have been estimated the most frequently are extracted. In this case, the impression generation unit 22 selects, as a document template, "This time, the conversation was centered on (question type) and was emotionally rich." or "This time, the conversation was centered on (question type) and was emotionally rich." Note that which of these emotions is selected may be randomly selected.

[0067] In this case, it is not necessary to determine the emotional expression. Therefore, for example, if "medication" is extracted as the type of question that has been estimated the most, the impression generation unit 22 generates a document such as "This time, the conversation was about medication and was emotionally rich." Furthermore, if "relationship" is extracted as the type of question that has been estimated the most, the impression generation unit 22 generates a document such as "This time, the conversation was about relationships and was emotionally rich."

[0068] The impression generator 22 then transmits the document generated as the impression to the terminal device 40 used by the caregiver. The impression is then presented to the caregiver. Therefore, the caregiver can grasp the overall situation of the care recipient from the impression.

[0069] [Device Operation] Next, the operation of the care support device 20 will be described with reference to FIG. 11. FIG. 11 is a flow diagram showing a second example of the operation of the care support device. In the following description, reference will be made to FIGS. 8 to 10 as appropriate. In addition, in the first embodiment, the care support method is implemented by operating the care support device 20. Therefore, the description of the care support method in the second embodiment will be replaced by the following description of the operation of the care support device 20.

[0070] First, the care support device 20 executes steps similar to steps A1 to A7 shown in Fig. 7 in the first embodiment to generate a summary and transmits the summary to the terminal device 40 used by the caregiver. In the second embodiment, the description of the process of generating the summary will be omitted.

[0071] 11, the question type estimation unit 21 first identifies text data corresponding to each question in the video data (step B1). Specifically, in step B1, the question type estimation unit 21 executes a speech recognition process to generate text data from the video data.

[0072] Next, the question type estimation unit 21 estimates the type of each question by comparing the text data of each question with the question type list 23 (step B2). Specifically, the question type estimation unit 21 calculates the similarity between the corresponding text data and each sample in the question type list 23 for each question, and determines the type of the sample with the highest similarity as the question type.

[0073] Next, the impression generating unit 22 extracts the question type that has been estimated the most frequently from among the question types estimated in step B1 (step B3).

[0074] Next, for each question, the impression generation unit 22 extracts the emotion that has been estimated most frequently from among the emotions estimated by either or both of the first emotion deduction unit 12 and the second emotion deduction unit 13 (step B4).

[0075] Next, the impression generation unit 22 applies the question type extracted in step B3 and the emotion extracted in step B4 to a pre-set document expression table 24 to generate a document indicating an impression about the conversation (step B5).

[0076] Thereafter, the impression generator 22 transmits the document generated as the impression to the terminal device 40 used by the caregiver (step B6).

[0077] As described above, according to the second embodiment, the caregiver is presented with the non-caregiver's impressions from the interview in addition to the summary described in the second embodiment. Therefore, the caregiver can understand the situation based on the care recipient's emotions from the summary displayed on the screen, and can also understand the overall situation of the care recipient from the impressions.

[0078] [Program] The program in the second embodiment may be a program that causes a computer to execute steps A1 to A7 shown in Fig. 7 and steps B1 to B6 shown in Fig. 11. By installing and executing this program on a computer, the care support device 10 and the care support method in the second embodiment can be realized. In this case, the processor of the computer functions as a data acquisition unit 11, a first emotion deduction unit 12, a second emotion deduction unit 13, an expression determination unit 14, a summary sentence generation unit 15, a speech recognition unit 16, a question type deduction unit 21, and an impression generation unit 22, and performs processing. Examples of the computer include a general-purpose PC, a smartphone, and a tablet terminal device.

[0079] The program in the second embodiment may be executed by a computer system constructed by a plurality of computers. In this case, for example, each computer may function as any one of the data acquiring unit 11, the first feeling deduction unit 12, the second feeling deduction unit 13, the expression determining unit 14, the summary generating unit 15, the speech recognizing unit 16, the question type deduction unit 21, and the impression generating unit 22.

[0080] [Physical Configuration] A computer that realizes the care support device by executing the programs in the first and second embodiments will now be described with reference to Fig. 12. Fig. 12 is a block diagram showing an example of a computer that realizes the care support device.

[0081] 12, the computer 110 includes a CPU (Central Processing Unit) 111, a main memory 112, a storage device 113, an input interface 114, a display controller 115, a data reader / writer 116, and a communication interface 117. These components are connected to each other via a bus 121 so as to be able to communicate data with each other.

[0082] Furthermore, the computer 110 may include a GPU (Graphics Processing Unit) or an FPGA (Field-Programmable Gate Array) in addition to or instead of the CPU 111. In this aspect, the GPU or FPGA can execute the programs in the embodiments.

[0083] The CPU 111 loads a program in the embodiment, which is composed of a group of codes and stored in the storage device 113, into the main memory 112 and executes each code in a predetermined order to perform various calculations. The main memory 112 is typically a volatile storage device such as a DRAM (Dynamic Random Access Memory).

[0084] The program in the embodiment is provided in a state stored in a computer-readable recording medium 120. The program in the embodiment may be distributed over the Internet connected via the communication interface 117.

[0085] Specific examples of the storage device 113 include a hard disk drive and a semiconductor storage device such as a flash memory. The input interface 114 mediates data transmission between the CPU 111 and input devices 118 such as a keyboard and a mouse. The display controller 115 is connected to a display device 119 and controls the display on the display device 119.

[0086] The data reader / writer 116 mediates data transmission between the CPU 111 and the recording medium 120, reads programs from the recording medium 120, and writes processing results from the computer 110 to the recording medium 120. The communication interface 117 mediates data transmission between the CPU 111 and other computers.

[0087] Specific examples of the recording medium 120 include general-purpose semiconductor storage devices such as CF (Compact Flash (registered trademark)) and SD (Secure Digital), magnetic recording media such as flexible disks, or optical recording media such as CD-ROMs (Compact Disk Read Only Memory).

[0088] The care support device in the embodiment can be realized not by a computer with a program installed, but by hardware corresponding to each unit, such as an electronic circuit. Furthermore, the care support device may be realized in part by a program and in the remaining part by hardware. In the embodiment, the computer is not limited to the computer shown in FIG. 12 .

[0089] Some or all of the above-described embodiments can be expressed by (Supplementary Note 1) to (Supplementary Note 18) described below, but are not limited to the following descriptions.

[0090] (Supplementary Note 1) A care support device comprising: a data acquisition unit that acquires video data with audio of a care recipient; a first emotion estimation unit that estimates an emotion of the care recipient based on a result of a voice recognition process on the audio in the video data; a second emotion estimation unit that estimates an emotion of the care recipient based on the video data; an expression determination unit that determines an expression that expresses a state of the care recipient using the emotion estimated by the first emotion estimation unit and the emotion estimated by the second emotion estimation unit; and a summary generation unit that generates a summary from the result of the voice recognition process and adds the determined expression to the generated summary.

[0091] (Supplementary Note 2) The care support device described in Supplementary Note 1, wherein the first emotion estimation unit extracts keywords from the text data obtained by the voice recognition processing, determines which of pre-set states the care recipient falls into based on the extracted keywords, and sets the determination result as the emotion of the care recipient.

[0092] (Supplementary Note 3) The care support device according to Supplementary Note 1, wherein the second emotion estimation unit inputs the video data into a learning model that has machine-learned the relationship between human facial images, human voices, and human emotions, and estimates the emotion of the care recipient using an output result of the learning model.

[0093] (Supplementary Note 4) The care support device according to Supplementary Note 1, wherein the expression determination unit determines an expression that expresses the state of the care recipient by comparing the emotion estimated by the first emotion estimation unit and the emotion estimated by the second emotion estimation unit with a table that defines a relationship between a person's emotion and an expression that expresses a person's state.

[0094] (Supplementary Note 5) The nursing care support device described in Supplementary Note 1, wherein the video data is video data obtained by filming a conversation situation consisting of questions to the care recipient and answers by the care recipient, and the nursing care support device further comprises: a question type estimation unit that estimates the type of the question from the video data; and an impression generation unit that generates a document indicating impressions about the conversation using the type of question estimated by the question type estimation unit.

[0095] (Supplementary Note 6) The care support device according to Supplementary Note 5, wherein the first emotion deduction unit and the second emotion deduction unit deduce an emotion of the care recipient for each of the questions, the question type deduction unit deduce a question type for each of the questions, the impression generation unit extracts the question type that has been estimated the most frequently, and further extracts the emotion that has been estimated the most frequently from the emotions deduced by either or both of the first emotion deduction unit and the second emotion deduction unit, and generates a document indicating an impression about the conversation by applying the extracted question type and the extracted emotion to a table that defines sentence expressions corresponding to questions and emotions.

[0096] (Supplementary Note 7) A care support method comprising: a data acquisition step of acquiring video data with audio of a care recipient; a first emotion estimation step of estimating an emotion of the care recipient based on a result of speech recognition processing of the audio in the video data; a second emotion estimation step of estimating an emotion of the care recipient based on the video data; an expression determination step of determining an expression that expresses a state of the care recipient using the emotion estimated in the first emotion estimation step and the emotion estimated in the second emotion estimation step; and a summary generation step of generating a summary from the result of the speech recognition processing and adding the determined expression to the generated summary.

[0097] (Appendix 8) A care support method as described in Appendix 7, wherein in the first emotion estimation step, keywords are extracted from the text data obtained by the voice recognition processing, and based on the extracted keywords, it is determined which of the pre-set states the care recipient falls into, and the determination result is used as the emotion of the care recipient.

[0098] (Supplementary Note 9) The care support method described in Supplementary Note 7, wherein in the second emotion estimation step, the video data is input into a learning model that has machine-learned the relationship between human facial images, human voices, and human emotions, and the emotion of the care recipient is estimated using the output result of the learning model.

[0099] (Appendix 10) The care support method described in Appendix 7, wherein in the expression determination step, an expression representing the state of the care recipient is determined by comparing the emotion estimated by the first emotion estimation step and the emotion estimated by the second emotion estimation step with a table that defines the relationship between a person's emotion and an expression representing a person's state.

[0100] (Supplementary Note 11) The nursing care support method described in Supplementary Note 7, wherein the video data is video data obtained by filming a conversation situation consisting of questions to the care recipient and answers from the care recipient, and the nursing care support method further includes: a question type estimation step of estimating the type of the question from the video data; and an impression generation step of generating a document indicating impressions about the conversation using the question type estimated by the question type estimation step.

[0101] (Supplementary Note 12) The care support method according to Supplementary Note 11, wherein in the first emotion estimation step and the second emotion estimation step, an emotion of the care recipient is estimated for each of the questions; in the question type estimation step, a question type is estimated for each of the questions; in the impression generation step, the question type that has been estimated the most times is extracted; and further, from the emotions estimated in either or both of the first emotion estimation step and the second emotion estimation step, the emotion that has been estimated the most times is extracted; and a document indicating impressions about the conversation is generated by applying the extracted question type and the extracted emotion to a table that specifies sentence expressions corresponding to questions and emotions.

[0102] (Supplementary Note 13) A computer-readable recording medium having recorded thereon a program including instructions for causing a computer to execute: a data acquisition step of acquiring video data with audio of a care recipient; a first emotion estimation step of estimating an emotion of the care recipient based on a result of speech recognition processing for the audio in the video data; a second emotion estimation step of estimating an emotion of the care recipient based on the video data; an expression determination step of determining an expression that expresses a state of the care recipient using the emotion estimated in the first emotion estimation step and the emotion estimated in the second emotion estimation step; and a summary generation step of generating a summary from the result of the speech recognition processing and adding the determined expression to the generated summary.

[0103] (Appendix 14) A computer-readable recording medium as described in Appendix 13, wherein in the first emotion estimation step, keywords are extracted from the text data obtained by the voice recognition processing, and based on the extracted keywords, it is determined which of the pre-set states the care recipient falls into, and the determination result is used as the emotion of the care recipient.

[0104] (Appendix 15) The computer-readable recording medium described in Appendix 13, wherein in the second emotion estimation step, the video data is input into a learning model that has machine-learned the relationship between human facial images, human voices, and human emotions, and the emotion of the care recipient is estimated using the output result of the learning model.

[0105] (Appendix 16) A computer-readable recording medium as described in Appendix 13, wherein in the expression determination step, an expression representing the state of the care recipient is determined by comparing the emotion estimated by the first emotion estimation step and the emotion estimated by the second emotion estimation step with a table that defines the relationship between a person's emotion and an expression representing a person's state.

[0106] (Supplementary Note 17) The computer-readable recording medium described in Supplementary Note 13, wherein the video data is video data obtained by filming a conversation situation consisting of questions to the care recipient and answers by the care recipient, and the program further includes instructions to cause the computer to execute: a question type estimation step of estimating the type of the question from the video data; and an impression generation step of generating a document indicating impressions about the conversation using the question type estimated by the question type estimation step.

[0107] (Supplementary Note 18) The computer-readable recording medium according to Supplementary Note 17, wherein in the first emotion estimation step and the second emotion estimation step, an emotion of the care recipient is estimated for each question; in the question type estimation step, a question type is estimated for each question; in the impression generation step, the question type that has been estimated the most times is extracted; and further, from the emotions estimated in either or both of the first emotion estimation step and the second emotion estimation step, the emotion that has been estimated the most times is extracted; and a document indicating impressions about the conversation is generated by applying the extracted question type and the extracted emotion to a table that specifies sentence expressions corresponding to questions and emotions.

[0108] Although the present invention has been described above with reference to the embodiments, the present invention is not limited to the above-described embodiments. Various modifications that can be understood by those skilled in the art can be made to the configuration and details of the present invention within the scope of the present invention.

[0109] This application claims priority based on Japanese Patent Application No. 2023-040942, filed on March 15, 2023, the disclosure of which is incorporated herein in its entirety.

[0110] According to the present disclosure, it is possible for a caregiver to easily grasp the situation of a care recipient based on the care recipient's emotions.The present disclosure is useful in caregiving settings.

[0111] DESCRIPTION OF SYMBOLS 10 Care support device (first embodiment) 11 Data acquisition unit 12 First emotion deduction unit 13 Second emotion deduction unit 14 Expression determination unit 15 Summary sentence generation unit 16 Speech recognition unit 17 Learning model 18 Emotion expression reverse dictionary 20 Care support device (second embodiment) 21 Question type deduction unit 22 Impression generation unit 23 Question type list 24 Document expression table 30 Video database 31 Imaging device 40 Terminal device 110 Computer 111 CPU 112 Main memory 113 Storage device 114 Input interface 115 Display controller 116 Data reader / writer 117 Communication interface 118 Input device 119 Display device 120 Recording medium 121 Bus

Claims

1. a data acquisition unit that acquires video data with audio of the care recipient; a first emotion estimation unit that estimates an emotion of the care recipient based on a result of a speech recognition process performed on a speech in the video data; a second emotion estimation unit that estimates an emotion of the care recipient based on the video data; an expression determination unit that determines an expression that expresses a state of the care receiver by using the emotion estimated by the first emotion estimation unit and the emotion estimated by the second emotion estimation unit; a summary generation unit that generates a summary from the result of the speech recognition processing and adds the determined expression to the generated summary; A care support device comprising:

2. the first emotion estimation unit extracts keywords from the text data obtained by the speech recognition processing, determines which of predetermined states the care recipient falls into based on the extracted keywords, and sets the determination result as the emotion of the care recipient. The care support device according to claim 1 .

3. the second emotion estimation unit inputs the video data into a learning model that has undergone machine learning to learn the relationship between human face images, human voices, and human emotions, and estimates the emotion of the care recipient using an output result of the learning model. The care support device according to claim 1 .

4. the expression determination unit determines an expression that expresses the state of the care receiver by comparing the emotion estimated by the first emotion deduction unit and the emotion estimated by the second emotion deduction unit with a table that defines a relationship between a person's emotion and an expression that expresses a person's state; The care support device according to claim 1 .

5. the video data is video data obtained by filming a conversation situation consisting of questions to the care recipient and answers from the care recipient, The care support device, a question type estimation unit that estimates the type of the question from the video data; an impression generation unit that generates a document indicating impressions about the conversation by using the type of question estimated by the question type estimation unit; Further comprising: The care support device according to claim 1 .

6. the first feeling deduction unit and the second feeling deduction unit deduce a feeling of the care receiver for each of the questions; the question type estimation unit estimates the type of the question for each of the questions; the impression generation unit extracts the type of the question that has been estimated the most frequently, and further extracts the emotion that has been estimated the most frequently from among the emotions estimated by either or both of the first emotion deduction unit and the second emotion deduction unit, and applies the extracted type of question and the extracted emotion to a table that defines sentence expressions corresponding to questions and emotions, thereby generating a document indicating impressions about the conversation. The care support device according to claim 5.

7. Acquire video data with audio of the care recipient, estimating an emotion of the care recipient based on a result of a voice recognition process performed on the voice in the video data; Estimating an emotion of the care recipient based on the video data; determining an expression that expresses a state of the care recipient using the emotion estimated by the first emotion estimation and the emotion estimated by the second emotion estimation; generating a summary from the result of the speech recognition processing, and adding the determined expression to the generated summary; A care support method characterized by:

8. In the first emotion estimation, a keyword is extracted from the text data obtained by the speech recognition processing, and a determination is made as to which of predetermined states the care recipient falls based on the extracted keyword, and the determination result is set as the emotion of the care recipient. The care support method according to claim 7.

9. In the second emotion estimation, the video data is input to a learning model that has undergone machine learning to learn the relationship between human face images, human voices, and human emotions, and the emotion of the care recipient is estimated using an output result of the learning model. The care support method according to claim 7.

10. In determining the expression, the expression representing the state of the care recipient is determined by comparing the emotion estimated by the first emotion estimation and the emotion estimated by the second emotion estimation with a table that defines a relationship between a person's emotion and an expression representing a person's state. The care support method according to claim 7.

11. the video data is video data obtained by filming a conversation situation consisting of questions to the care recipient and answers from the care recipient, The care support method further comprises: Estimating the type of the question from the video data; generating a document indicating impressions about the conversation using the question type estimated in the question type estimation; The care support method according to claim 7.

12. In the first emotion estimation and the second emotion estimation, an emotion of the care receiver is estimated for each of the questions; In the estimating of the question type, the question type is estimated for each of the questions; In generating the impressions, the type of question that has been estimated the most frequently is extracted, and further, from among the emotions estimated in either or both of the first emotion estimation and the second emotion estimation, the emotion that has been estimated the most frequently is extracted, and a document indicating impressions about the conversation is generated by applying the extracted type of question and the extracted emotion to a table that defines sentence expressions corresponding to questions and emotions. The care support method according to claim 11.

13. On the computer, Acquire video data with audio of the person receiving care, estimating the emotion of the care recipient based on a result of a voice recognition process performed on the voice in the video data; Estimating the emotion of the care recipient based on the video data; determining an expression that expresses a state of the care recipient using the emotion estimated by the first emotion estimation and the emotion estimated by the second emotion estimation; generating a summary from the result of the speech recognition processing, and adding the determined expression to the generated summary; program.

14. In the first emotion estimation, a keyword is extracted from the text data obtained by the speech recognition processing, and a determination is made as to which of predetermined states the care recipient falls based on the extracted keyword, and the determination result is set as the emotion of the care recipient. The program according to claim 13.

15. In the second emotion estimation, the video data is input to a learning model that has undergone machine learning to learn the relationship between human face images, human voices, and human emotions, and the emotion of the care recipient is estimated using an output result of the learning model. The program according to claim 13.

16. In determining the expression, the expression representing the state of the care recipient is determined by comparing the emotion estimated by the first emotion estimation and the emotion estimated by the second emotion estimation with a table that defines a relationship between a person's emotion and an expression representing a person's state. The program according to claim 13.

17. the video data is video data obtained by filming a conversation situation consisting of questions to the care recipient and answers from the care recipient, The computer further comprises: A type of the question is estimated from the video data; generating a document indicating impressions about the conversation using the question type estimated in the question type estimation; The program according to claim 13.

18. In the first emotion estimation and the second emotion estimation, an emotion of the care receiver is estimated for each of the questions; In the estimating of the question type, the question type is estimated for each of the questions; In generating the impressions, the type of question that has been estimated the most frequently is extracted, and further, from among the emotions estimated in either or both of the first emotion estimation and the second emotion estimation, the emotion that has been estimated the most frequently is extracted, and a document indicating impressions about the conversation is generated by applying the extracted type of question and the extracted emotion to a table that defines sentence expressions corresponding to questions and emotions. The program according to claim 17.