Letter creation system, letter creation method, and program

The AI-powered letter creation system addresses the challenge of frequent updates in care facilities by automatically generating letters with video excerpts and text, enhancing communication and reducing staff workload.

JP7723047B2Active Publication Date: 2025-08-13TEPCO TOWN PLANNING CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023129976
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-08-09
Publication Date
2025-08-13
Estimated Expiration
2043-08-09

AI Technical Summary

Technical Problem

Families of residents in residential care facilities, such as nursing homes, are unable to frequently receive updates on their loved ones' daily activities due to the time-consuming nature of traditional reporting methods, exacerbated by the COVID-19 pandemic, leading to a lack of easy access to their residents' status information.

Method used

A letter creation system and method utilizing AI to detect smiling video frames, create a video excerpt, and generate accompanying text from document data, incorporating audio content, to automatically produce high-quality letters that include videos and images, reducing staff workload.

Benefits of technology

The system enables the efficient creation of high-value-added letters that include videos, significantly reducing staff workload and ensuring families can easily access residents' recent activities, enhancing communication in care facilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007723047000001
    Figure 0007723047000001
  • Figure 0007723047000002
    Figure 0007723047000002
  • Figure 0007723047000003
    Figure 0007723047000003
Patent Text Reader

Abstract

To provide a letter creation system, a letter creation method, and a program for easily creating high-quality letters including video.SOLUTION: A letter creation system 1 comprises a control unit 2 and a storage unit 3. The control unit 2 detects video frames in which a target person is smiling using AI from video data 5 stored in the storage unit 3, and extracts the video data 5 before and after the video frames in which the target person is smiling, thereby creating a provided video 7, and uses AI to create provided text 8 related to document data 6 and the provided video 7 stored in the storage unit 3 from the document data 6 and the provided video 7.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a letter creation system, a letter creation method, and a program that use AI to automatically create documents such as letters from video data and document data. [Background technology]

[0002] Conventionally, there is known a technique for using AI to recognize facial expressions of a target person and detect smiles in images of video data, etc. (for example, Non-Patent Documents 1 and 2). [Prior art documents] [Non-patent literature]

[0003] [Non-Patent Document 1] "How does image recognition technology work? Types, latest use cases, and changes in the use of AI (deep learning)," [online], System Integrator Co., Ltd. [Retrieved July 18, 2023], Internet<URL:https: / / products.sint.co.jp / aisia-ad / blog / image-recognition-ai> [Non-patent document 2] "AI Image Recognition Solutions," [online], NEC Corporation [Retrieved July 18, 2023], Internet<URL:https: / / jpn.nec.com / bv / hoso / ai_recognition.html> Summary of the Invention [Problem to be solved by the invention]

[0004] Meanwhile, in residential care facilities such as nursing homes, for example, families have been unable to visit residents due to the impact of the COVID-19 pandemic. Meanwhile, recreational activities are held daily in care facilities, and reports are sometimes sent to families, but due to the time-consuming nature of this process, it was not possible to send out such reports frequently. Under these circumstances, families were unable to easily find out about the daily status of their residents.

[0005] The present invention has been made in view of the above circumstances, and provides a letter creation system, a letter creation method, and a program that enable users to easily create high-quality letters including videos. [Means for solving the problem]

[0006] In order to solve such problems, the present invention provides a letter creation system equipped with a control unit and a memory unit, and a letter creation method for creating letters using the same, in which the control unit uses AI to detect video frames in which a target person is smiling from video data stored in the memory unit, and creates a provided video by extracting video data before and after the video frame in which the target person is smiling, and the control unit uses AI to create a provided text with content related to the document data and the provided video from the document data stored in the memory unit and the provided video.

[0007] In the letter creation system and the letter creation method, when creating the provided text, the provided text can be created using AI from audio data included in the provided video.

[0008] In the letter creation system and letter creation method, when creating the provided text, AI derives multiple top first keywords that are highly relevant to the keywords included in the document data in each category of action verb words, state verb words, emotional adjective words, and noun words, and from among the derived first keywords, AI derives top second keywords that are highly relevant to the provided video, and AI can create the provided text based on the derived second keywords.

[0009] In the letter creation system and the letter creation method, when detecting video frames in which the target person is smiling, one video frame per any number of seconds can be used and the remaining video frames can be skipped.

[0010] In the letter creation system and the letter creation method, when detecting video frames in which the target person is smiling, the image size of the video data is divided, and from each of the divided video data, a video frame in which the target person is smiling can be detected using AI.

[0011] The letter creation system and the letter creation method can create a document that includes the provided text, a still image extracted from the video data, and an identification code that is an encrypted URL where the provided video is saved.

[0012] The present invention also provides a program for causing a computer to execute the letter creation method. [Effects of the Invention]

[0013] The letter creation system, letter creation method, and program of the present invention make it possible to easily create high-value-added letters that include video viewing. These letters can be created automatically using AI, significantly reducing the workload of nursing care facility staff. [Brief explanation of the drawings]

[0014] [Figure 1] 1 is a block diagram illustrating a letter writing system according to a preferred embodiment of the present invention. [Figure 2] FIG. 2 is a schematic diagram showing an example of the configuration of a video frame. [Figure 3] 10A and 10B are schematic diagrams illustrating an example of detecting video frames in which a target person is smiling. [Figure 4] FIG. 10 is a schematic diagram illustrating an example of a method for dividing the image size of video data. [Figure 5] 10 is a flowchart showing a flow of detecting a smile and creating a video to be provided. [Figure 6] 10 is a flowchart showing the flow of a masking process. [Figure 7] FIG. 10 is a schematic diagram for explaining a means for creating a provided text. [Figure 8] FIG. 10 is a schematic diagram showing an example of creating a provided sentence. [Figure 9] FIG. 10 is a schematic diagram for explaining a means for creating a provided text using a first keyword and a second keyword. [Figure 10] FIG. 10 is a schematic diagram showing an example of a provided sentence created using a first keyword and a second keyword. [Figure 11] FIG. 10 is a schematic diagram showing an example of a letter document. DETAILED DESCRIPTION OF THE INVENTION

[0015] Hereinafter, preferred embodiments of the present invention will be described with reference to the drawings. However, the present invention is not limited to the following description, and various modifications and changes may be made by those skilled in the art based on the gist of the invention as set forth in the claims or disclosed in the detailed description. Such modifications and changes are also included within the scope of the present invention.

[0016] <Overall Overview>

[0017] The letter creation system 1, letter creation method, and program 4 of this embodiment shown in Figure 1 enable staff at residential care facilities such as nursing homes to create letters to inform residents' families of their recent status. Because the letters are created automatically by AI, the workload of care facility staff can be significantly reduced.

[0018] Nursing care facilities hold recreational activities every day. In the letter creation system 1, letter creation method, and program 4 of this embodiment, AI is used to detect video frames in which residents are smiling from video data 5 of such recreational activities, and a provided video 7 with many smiling faces is created. Furthermore, AI is used to create a provided text 8 describing the resident's enjoyment from the provided video 7 and document data 6 outlining the recreational activities. Residents' families can easily find out about their recent activities by viewing the provided video 7 and provided text 8.

[0019] As shown in Fig. 1, the letter creation system 1 includes at least a control unit 2 and a storage unit 3. The control unit 2 is, for example, at least one processing unit such as a CPU or GPU, and executes a program 4 stored in the storage unit 3. The storage unit 3 is, for example, at least one memory device such as a RAM or ROM, and stores the program 4 and data used by the letter creation system 1.

[0020] The letter creation system 1 is, for example, a computer, which may be operated directly by a user, or may be a server that communicates via a network such as the Internet. A program 4 causes the computer serving as the letter creation system 1 to execute the letter creation method of this embodiment, and a control unit 2 uses AI to create a provided video 7 and a provided text 8 based on video data 5 and document data 6. The control unit 2 also creates a document 33 containing the provided text 8, a still image 31, and an identification code 32. Resident information 9 is used to identify the target person when detecting a smile.

[0021] <Creating videos to be provided>

[0022] To create the provided video 7, first, the control unit 2 uses AI to detect video frames 11 (see FIG. 3) in which the target person is smiling from the video data 5 stored in the storage unit 3. The video data 5 is, for example, footage of a recreational activity.

[0023] To detect video frames 11 in which a target person is smiling using AI, the control unit 2 performs face detection and identification of the target person using AI face recognition, and recognizes the target person's facial expression using AI smile detection. For AI face recognition, open source FaceNet or Insightface, for example, can be used. For AI smile detection, open source Residual Masking Network, for example, can be used.

[0024] Conventionally, AI-based facial recognition and smile detection processes all video frames in video data 5. In this embodiment, when detecting video frames 11 in which a target person smiles, one video frame per any number of seconds is used, and the remaining video frames are skipped. For example, a setting may be made so that a user (a care facility staff member) can select one video frame per second, one video frame per two seconds, or one video frame per N seconds. Note that the number of seconds may be a decimal rather than an integer, and may be any real number. For example, as shown in FIG. 2, if the video data 5 is 30 fps and one video frame per second is used, 29 frames are skipped, and video frames 11 in which a target person smiles are detected every 30 frames. This enables AI processing of facial recognition and smile detection to be performed quickly and with a low load.

[0025] AI smile detection estimates seven facial expressions, for example, Neutral, Happy, Sad, Surprise, Angry, Fear, and Disgust.

[0026] Any method can be used to detect video frames 11 in which the target person is smiling using AI, and various methods are possible. For example, as shown in FIG. 3, if multiple consecutive video frames (e.g., five frames) of "Happy" in which the estimated smile level is, for example, 90% or higher are detected, the first video frame of the five frames is determined to be video frame 11 in which the target person is smiling, and video data, for example, 15 seconds before and 15 seconds after the video frame 11, is extracted to provide the video 7. In this way, the control unit 2 can create the provided video 7 by using AI to detect video frames 11 in which the target person is smiling from the video data 5 stored in the storage unit 3 and extracting video data 5 before and after the video frame 11 in which the target person is smiling.

[0027] An example of a method different from that shown in FIG. 3 for detecting video frames 11 in which a target person smiles using AI is described below. First, the estimated degree of facial expression is created for video frames in which AI-based smile detection has been performed throughout the video data 5. Next, from the list created, one or more video frames with the highest degree of smile are selected as video frames 11 in which the target person smiles. Then, video data, for example, 15 seconds before and 15 seconds after each of these video frames 11 is extracted and used as the provided video 7. In this way, the control unit 2 can create the provided video 7 by using AI to detect video frames 11 in which the target person smiles from the video data 5 stored in the memory unit 3 and extracting video data 5 before and after the video frame 11 in which the target person smiles.

[0028] In AI processing, for example, an image size of 1920 x 1080 (2K) is typically compressed to an image size of 300 x 300 for facial recognition and smile detection. In this embodiment, as shown in FIG. 4 , when detecting video frames 11 in which a target person smiles, the image size of video data 5 is divided into, for example, m x n, and video frames 11 in which the target person smiles are detected from each of the divided video data 51 using AI. The image size of video data 5 can be set to, for example, a maximum of 7 x 4, taking into consideration the border width and overlap width. This allows for flexible response to environmental changes such as weather and illuminance by changing the m x n setting. Furthermore, the accuracy of facial recognition and smile detection is improved, preventing missed detections and ensuring privacy considerations, for example, when masking processing, as described below, is required.

[0029] FIG. 5 is a flowchart showing the process of detecting a smile and creating the provided video 7.

[0030] First, in step S1, the control unit 2 acquires video data 5 from the memory unit 3. In step S2, the control unit 2 performs face detection using AI-based face recognition. Face detection involves sampling a certain number of faces from all video frames and processing them. Next, processing is performed for each detected face. In step S3, the control unit 2 uses AI-based face recognition to compare the detected face with the resident information 9 stored in the memory unit 3 to identify the target person. In step S4, the control unit 2 uses AI-based smile detection to recognize the facial expression of the identified target person. In step S5, if the facial expression determination is OK, that is, if the control unit 2 uses AI to detect a video frame 11 in which the target person is smiling, the process proceeds to step S6. In step S6, the control unit 2 extracts a still image 31 from the video data 5. In step S7, the control unit 2 extracts video data 5 before and after the video frame 11 in which the target person is smiling, and creates the provided video 7.

[0031] In step S8, a masking process such as mosaic processing is performed on residents who do not wish to appear in the video. In this process, as shown in FIG. 6, the control unit 2 performs face detection using AI facial recognition in step S81 and identifies the target person in step S82. In step S83, the control unit 2 performs masking such as mosaic processing on the identified target person. The masking process in step S8 is also performed on video frames skipped when detecting video frame 11 in which the target person is smiling. Therefore, an end determination is made in step S84, and steps S81 to S83 are repeated until processing is completed for all video frames. Note that masking such as mosaic processing may be performed not only on residents who do not wish to appear in the video, but also on all people other than the target person for whom a smile is detected. In this case, the processes from steps S81 to S83 are repeated for all people to be masked until processing is completed for all video frames. If masking is not necessary, step S8 can be omitted.

[0032] In step S9, the control unit 2 stores the still image 31 and the provided video 7 for each target person in the storage unit 3. Then, in step S10, an end determination is made, and the processing from steps S3 to S9 is repeated until processing is completed for all people with detected faces. The stored provided video 7 is used to create the provided text 8. Note that instead of storing the data of the still image 31 and the provided video 7 in the storage unit 3, information on the provided video 7 and information output by the AI may be stored in the storage unit 3, and the still image 31 and the provided video 7 may be created from this information when creating the provided text 8.

[0033] <Creating the provided text>

[0034] As shown in FIG. 7 , provided text 8 is created by AI from document data 6 and provided video 7 stored in storage unit 3, with content related to document data 6 and provided video 7. Control unit 2 uses AI to create provided text 8 describing the target person enjoying the recreation described in document data 6 from keywords described in document data 6 that describes an overview of the recreation and footage of provided video 7 created from video data 5 that filmed the recreation. Because provided video 7 is a video of the target person smiling and enjoying the recreation described in document data 6, provided text 8 is a text describing the target person enjoying the recreation described in document data 6.

[0035] Furthermore, when creating the provided text 8, the control unit 2 may also use AI to create the provided text 8 from the audio data 12 included in the provided video 7. In this case, the control unit 2 uses AI to create the provided text 8, which describes the target person enjoying the recreational activity described in the document data 6, from the keywords described in the document data 6, the video of the provided video 7, and the conversation in the audio data 12.

[0036] It is preferable to train the AI to create only positive sentences, omitting negative emotions such as sadness, pain, hatred, resentment, jealousy, anger, and anxiety, which will make it more likely to describe a person enjoying recreational activities.

[0037] An example of the provided text 8 created from the document data 6 and the provided video 7 is shown in Figure 8. As shown in Figure 8, the document data 6 includes information such as the content and date of the recreational activity, and the provided text 8 is created based on the keywords entered. In addition to template text, the provided text 8 includes text describing how the target person is enjoying the recreational activity described in the document data 6.

[0038] For example, OpenAI's GPT can be used to create sentences from document data 6 using AI. For example, OpenAI's Clip can be used to create sentences from provided videos 7 using AI. For example, OpenAI's Whisper can be used to create sentences from audio data 12 using AI.

[0039] FIG. 9 shows the configuration of a means for creating provided text 8 using first keywords 25 and second keywords 26. In this creation means, when creating provided text 8, the control unit 2 uses AI to derive multiple first keywords 25 that are highly relevant to keywords included in document data 6 in each category of action verb words 21, state verb words 22, emotional adjective words 23, and noun words 24. Then, from the derived first keywords 25, the control unit 2 derives second keywords 26 that are highly relevant to the provided video 7. Based on the derived second keywords 26, the control unit 2 creates provided text 8. This creation means uses AI to create provided text 8 with content related to the document data 6 and the provided video 7 from the document data 6 and the provided video 7 stored in the storage unit 3. While the diagram shows ten first keywords 25 and three second keywords 26 derived for each category, the number of these keywords is not limited.

[0040] FIG. 10 shows an example of a provided text 8 created by a creating means using a first keyword 25 and a second keyword 26.

[0041] First, the control unit 2 uses AI to derive ten primary keywords 25 that are highly relevant to the keyword "chorus" contained in the document data 6 in each category of action verb words 21, state verb words 22, emotional adjective words 23, and noun words 24. The derived primary keywords 25 are as shown in FIG. 10.

[0042] Next, the control unit 2 inputs the first keywords 25 and the provided video 7, and uses AI to derive the top three second keywords 26 in each category that are highly relevant to the provided video 7 from the derived ten first keywords 25. The provided video 7 used here is a video of multiple people singing in a circle. The derived second keywords 26 are as shown in Figure 10.

[0043] Finally, the control unit 2 uses AI to create the provided text 8 based on a total of 12 second keywords 26. The created provided text 8 is as shown in Figure 10. In addition to the template text, the created provided text 8 includes text that describes the target person enjoying the recreational activities described in the document data 6, and includes keywords included in the document data 6 and text that is highly relevant to the video of the provided video 7.

[0044] In addition, verification was also conducted using other recreation-related keywords such as "gymnastics," "origami," and "spending time in the living room," as well as related provided videos 7, and it was confirmed that appropriate provided sentences 8 could be created. Furthermore, verification was also conducted when two or five secondary keywords 26 were derived for each category, and it was confirmed that appropriate provided sentences 8 could be created.

[0045] <Creating a document>

[0046] As shown in FIG. 11 , the control unit 2 may create a document 33 that includes the provided text 8, a still image 31 extracted from the video data 5, and an identification code 32 that is an encrypted URL where the provided video 7 is saved. The still image 31 may be created, for example, from a video frame 11 in which the target person is smiling. The identification code 32 may be, for example, a QR code (registered trademark). When creating the document 33, the staff of the nursing facility may be able to edit the provided text 8. Furthermore, the staff of the nursing facility may be able to select the still image 31 at their own discretion.

[0047] The staff of the nursing care facility provide the created document 33 to the resident's family in paper form or electronic data. The resident's family can view the provided video 7 by reading the identification code 32 with a smartphone or the like and accessing the URL where the provided video 7 is saved.

[0048] <Summary of the embodiment>

[0049] As described above in detail, the letter creation system 1 and letter creation method of this embodiment are a letter creation system 1 equipped with a control unit 2 and a memory unit 3, and a letter creation method for creating a letter using the same, in which the control unit 2 uses AI to detect video frames 11 in which a target person is smiling from video data 5 stored in the memory unit 3, and extracts video data 5 before and after the video frame 11 in which the target person is smiling, thereby creating a provided video 7, and from document data 6 stored in the memory unit 3 and the provided video 7, uses AI to create provided text 8 with content related to the document data 6 and the provided video 7.

[0050] This configuration allows AI to automatically create letters, significantly reducing the workload of nursing home staff.

[0051] Furthermore, when creating the provided text 8, the control unit 2 also creates the provided text 8 from the audio data 12 included in the provided video 7 using AI.

[0052] This structure makes it possible to create a provided document 8 that more faithfully describes the residents' recreational activities.

[0053] In addition, when creating the provided text 8, the control unit 2 uses AI to derive multiple first keywords 25 that are highly relevant to the keywords included in the document data 6 in each category of action verb words 21, state verb words 22, emotional adjective words 23, and noun words 24, and from the derived first keywords 25, derives second keywords 26 that are highly relevant to the provided video 7 using AI, and creates the provided text 8 using AI based on the derived second keywords 26.

[0054] With this configuration, it is possible to create provided text 8 that has content highly related to the keywords contained in document data 6 and the video of provided video 7.

[0055] Furthermore, when detecting video frames 11 in which the target person is smiling, the control unit 2 uses one video frame per given number of seconds and skips the remaining video frames.

[0056] This configuration enables AI processing for facial recognition and smile detection to be performed quickly and with low load.

[0057] In addition, when detecting video frames 11 in which the target person is smiling, the control unit 2 divides the image size of the video data 5 and uses AI to detect video frames 11 in which the target person is smiling from each of the divided video data 5.

[0058] This configuration allows for flexible response to changes in the environment, while also improving the accuracy of face recognition and smile detection, preventing missed detections.

[0059] The control unit 2 also creates a document 33 that includes the provided text 8, a still image 31 extracted from the video data 5, and an identification code 32 that is an encrypted URL where the provided video 7 is saved.

[0060] This configuration allows us to provide newsletters with higher added value, by allowing users to easily watch videos of residents having fun.

[0061] Furthermore, the program 4 of this embodiment causes a computer serving as the letter creation system 1 to execute the letter creation method described above. [Explanation of symbols]

[0062] 1 Letter creation system 2. Control section 3 Storage section 4. Program 5. Video data 6 Document Data 7 Provided Videos 8 Provided text 9. Tenant Information 11 Video frames in which the subject is smiling 12 Audio data 21 Action Verb Words (Categories) 22 State verb words (categories) 23 Emotional adjective words (categories) 24 Noun words (categories) 25 First Keyword 26 Second Keyword 31 still images 32 Identification Code 33 documents

Claims

1. A letter creation system including a control unit and a storage unit, The control unit uses the video data stored in the storage unit as input to AI, detects video frames in which a target person is smiling using AI, and extracts video data before and after the video frames in which the target person is smiling, thereby creating a video to be provided; The control unit uses the document data stored in the storage unit as an input to AI, and derives first keywords related to the document data by the AI; The control unit uses the first keyword and the provided video as input to an AI, and derives a second keyword related to the first keyword and the provided video by the AI; The control unit uses the second keyword as an input to AI, and creates a provided sentence with content related to the second keyword by AI, The AI facial recognition for detecting the smile of the target person uses the open source FaceNet or Insightface. The open source Residual Masking Network is used for AI-based smile detection to detect the smile of the target person. The AI derives the first keyword and creates the provided sentences using OpenAI's GPT. The second keyword is derived using OpenAI's Clip. The control unit creates a document that describes the provided text. Letter creation system.

2. The letter creation system according to claim 1 , wherein when detecting a video frame in which the target person is smiling, one video frame per any number of seconds is used and the remaining video frames are skipped.

3. The letter creation system of claim 1, wherein when detecting video frames in which the target person is smiling, the image size of the video data is divided, and video frames in which the target person is smiling are detected from each of the divided video data using AI.

4. The control unit: The provided text; A still image extracted from the video data; and An identification code obtained by encrypting the URL where the provided video is stored; 2. The letter creation system according to claim 1, which creates a document containing the following:

5. A letter creation method for creating a letter using a letter creation system having a control unit and a storage unit, The control unit uses the video data stored in the storage unit as input to AI, detects video frames in which a target person is smiling using AI, and extracts video data before and after the video frames in which the target person is smiling, thereby creating a video to be provided; The control unit uses the document data stored in the storage unit as an input to AI, and derives first keywords related to the document data by the AI; The control unit uses the first keyword and the provided video as input to an AI, and derives a second keyword related to the first keyword and the provided video by the AI; The control unit uses the second keyword as an input to AI, and creates a provided sentence with content related to the second keyword by AI, The AI facial recognition for detecting the smile of the target person uses the open source FaceNet or Insightface. The open source Residual Masking Network is used for AI-based smile detection to detect the smile of the target person. The AI derives the first keyword and creates the provided sentences using OpenAI's GPT. The second keyword is derived using OpenAI's Clip. The control unit creates a document that describes the provided text. How to create a letter.

6. The letter creation method according to claim 5, wherein when detecting a video frame in which the target person is smiling, one video frame per given number of seconds is used and the remaining video frames are skipped.

7. 6. The letter creation method of claim 5, wherein when detecting video frames in which the target person is smiling, the image size of the video data is divided, and video frames in which the target person is smiling are detected from each of the divided video data using AI.

8. The control unit: The provided text; A still image extracted from the video data; and An identification code obtained by encrypting the URL where the provided video is stored; 6. The method for creating a letter according to claim 5, further comprising creating a document containing the above.

9. A program for causing a computer to execute the letter creation method according to any one of claims 5 to 8.

Citation Information

Patent Citations

  • Program, method and device for supporting creation of document

    JP2019091476A