Care information generation system

JP2025112374A5Active Publication Date: 2025-09-10MAI SYSTEM PLANNING CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024006560
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-19
Publication Date
2025-09-10
Estimated Expiration
2044-01-19

AI Technical Summary

Technical Problem

Caregivers face a heavy burden in creating care records manually, which leads to overtime work, and there is a need for a system that can automatically generate care information using generative AI to support caregivers and caregiver leaders.

Method used

A care information generation system utilizing a terminal device with a microphone to convert voices into voice signals, employing energy-based voice activity detection to recognize voice activity data, and a generation AI engine to generate care information, which is displayed on a care monitor, trained with care and medicine-related corpora, and adapted for different environments.

Benefits of technology

The system efficiently recognizes necessary voices and generates care information in real-time, reducing caregiver workload and enabling automatic care record generation, data analysis, and flexible information retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To provide a care information generation system which recognizes required voice from voice of a care site and can generate care information.SOLUTION: A care information generation system 1000 includes a generative AI engine G and at least one of terminal devices 10A and 10B for generating care information from voice of a care site. The terminal devices A and B include microphones 11A and 11B for converting voice of the care site into a voice signal, recognize generation of voice activity data based on the voice signal and transmit the voice activity data to the generative AI engine. The voice activity data received from the terminal devices is inputted to the generative AI engine G and the generative AI engine G generates care information based on the voice activity data.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a care information generation system using generative AI.

Background Art

[0002] In the care system described in Patent Document 1, it is possible to display, on a touch panel, information obtained by voice recognition of the content of a call made through a receiver and a handset through which a care staff and a care recipient can communicate. However, there is a large amount of information obtained by care staff through care services, and the types of that information are diverse, including those requiring urgency, those becoming care records, and those such as deferrals and business contacts. It is necessary to utilize them in an appropriate output method depending on the nature of the input information. Also, it is a form in which recording is started by pushing a button, and generally, operation of a terminal is also required for smartphones and tablets.

[0003] In addition, natural language processing technology has highly developed in source separation, speech recognition, and Transformer for language processing. However, due to the emergence of generative AI in recent years, the development environment has changed greatly.

[0004] The voice recognition method used in the care information generation system described in Patent Document 1 utilizes a conversation app, and it is known that the recognition rate decreases for utterances with a high degree of specialization such as medical and care terms.

Prior Art Documents

Patent Documents

[0005]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0006] Caregivers create care records after providing care services to care recipients. The methods of creating these care records are handwritten on paper media or manually input into a personal computer, tablet, etc. This work places a heavy burden on caregivers and is the cause of overtime services. In addition, caregivers also follow up on foreign workers and part-timers, increasing the burden on caregivers.

[0007] During the provision of care services to care recipients, caregivers record information freely by voice without a time lag, and promptly convey urgent matters, deferrals, and business contacts to relevant parties. In addition, care records are analyzed and summarized in tables, etc., and then conveyed to relevant parties. Furthermore, it is desired that care record proposals are automatically generated by generative AI at the end of work.

[0008] The method of transcribing spoken voice data into text data is generally performed by a conversation application. By performing this transcription with generative AI, generative AI is trained to achieve "cultivating AI into top-notch caregivers". AI that has learned medical and care knowledge and the actual situation at the care site proposes care record drafts to caregivers.

[0009] Caregiver leaders monitor the care site in accordance with the care schedule, make appropriate greetings and give guidance to control the care site. However, the workload of caregiver leaders is large and it is a harsh role. A system that utilizes generative AI to support the work of caregiver leaders is needed.

[0010] Generally, care information such as data analysis data, tabulation data, drawing data, and text data of care services is registered in a DB (database), and is realized by starting a script. This requires a long development period and is also time-consuming to modify the script. The present invention provides a care information generation system capable of recognizing necessary voices from the voices at the care site and generating care information.

Means for Solving the Problems

[0011] Each aspect of the present invention is as follows. (Aspect Example 1) A care information generation system including at least one terminal device and a generation AI engine for generating care information from voices at the care site, wherein the terminal device includes a microphone that converts voices at the care site into voice signals, recognizes the occurrence of voice activity data based on the voice signals, and transmits the voice activity data to the generation AI engine, and the generation AI engine is input with the voice activity data received from the terminal device, and the generation AI engine generates the care information based on the voice activity data.

[0012] (Aspect Example 2) In the care information generation system according to Aspect Example 1, the care information is displayed on a care monitor. (Aspect Example 3) In the care information generation system according to Aspect Example 1, the generation AI engine is trained using a corpus related to care and / or medicine. (Aspect Example 4) In the care information generation system according to claim 1, the control unit of the terminal device monitors voice energy using energy-based voice activity detection (VAD) and determines the occurrence of the voice activity data when the threshold value of the voice energy is within a predetermined range.

[0013] (Aspect Example 5) In the care information generation system according to Aspect Example 4, the microphone of the terminal device is a wired microphone or a wireless microphone. The control unit of the terminal device includes a wired threshold for voice energy for the wired microphone and a wireless threshold for voice energy for the wireless microphone. A care information generation system. (Aspect Example 6) In the care information generation system according to Aspect Example 5, The wired threshold is 32000 ± 5000 (db), and the wireless threshold is 25000 ± 5000 (db). A care information generation system. (Aspect Example 7) In the care information generation system according to Aspect Example 1, The control unit of the terminal device recognizes a flicking sound generated by flicking the microphone with a finger and starts recording. A care information generation system.

[0014] (Aspect Example 8) In the care information generation system according to Aspect Example 7, The flicking sound has sound energy that is greater than the voice signal and lasts for a shorter time than the voice signal. (Aspect Example 9) In the care information generation system according to Aspect Example 7, After recognizing the sound, the control unit starts recording by activating a function of a Python script. A care information generation system. (Aspect Example 10) In the care information generation system according to Claim 1, During the day, the terminal device uses a buzzer to notify the caregiver of the care information, and at night, it uses a vibrator to notify the caregiver of the care information. A care information generation system.

[0015] (Aspect Example 11) In the care information generation system according to Claim 1, The system includes a PC / server that provides learning data to the generation AI engine. A care information generation system. (Aspect Example 12) In the care information generation system according to Claim 1, The care information generation system includes a data processing server capable of communicating with the terminal device and the AI engine. The care information generation system, wherein the data processing server receives the voice activity data from the terminal device and provides the voice activity data to the generation AI engine.

Effect of the Invention

[0016] The care information generation system of the present invention can recognize necessary voices from voices at the care site and generate care information.

Brief Description of the Drawings

[0017]

Figure 1

Figure 2

Figure 2A

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

Figure 25

Figure 26

Figure 27

Figure 28

Figure 29

Figure 30

Figure 31

Figure 32

Figure 33

Figure 34

Figure 35

Figure 36

Figure 37

Mode for Carrying Out the Invention

[0018] An embodiment of the care information generation system of the present invention will be described with reference to the drawings. Parts common to each figure are denoted by the same reference numerals and the description thereof will be omitted as appropriate. Note that the present invention is not limited to the care information generation system, and can also be a medical information generation system that can be used in a medical field, or a sports information generation system that can be used in a sports field such as fitness.

[0019] In the care information generation system 1000 of the present embodiment, text data, voice data, and image data including videos of care staff are input to the generation AI engine. Further, by inputting appropriate instructions by the care staff, the care information generation system 1000 automatically outputs care information such as making tables and drawings. These output materials are used as care summaries for daily operation management of care services, regular reviews, and care plans such as care plans.

[0020] Fig. 1 shows an overview of the developed care information input system. Caregiver A who provides care services at the site uses a microphone 11A on a portable communication terminal device 10A to record the situation of the care recipient. The recorded voice data is stored in the storage of a cloud server E (data processing server) through Internet communication H including satellite communication. The voice data is used to generate text data by utilizing a generation AI engine G set in the cloud server E with a speech-to-text script, and is stored in the storage. The generation AI engine G is provided in a generation AI server. The processed text data is registered / uploaded to a website F and can be browsed by the staff from a non-public URL. In addition, text data that requires urgency is converted back to voice by voice synthesis technology, downloaded to the communication terminal device 10B of caregiver B, and transmitted from a speaker 11B. A developer PC / server 10D such as a PC / server of developer D performs maintenance of the cloud server E. In addition, a learning program for the generation AI is executed, and highly specialized information such as medical care, nursing care, and legal regulations is provided to and cultivated in the generation AI.

[0021] As shown in Fig. 2, the inventors developed a freehand communication terminal device 10 that can be carried by caregivers. It enables voice recording without performing any operations such as pressing buttons or scrolling the screen during the provision of care services. When repeated tests were conducted at the care site, for example, in the method where recording starts with the utterance of "start", a time lag of more than 3 seconds occurred because the utterance was judged after being converted into text data.

[0022] Therefore, in this embodiment, the energy-based voice activity detection (VAD) using Simple Audio Processing is adopted to monitor the voice energy, and when it exceeds a certain threshold, it is specified to determine / recognize the generation of voice activity data. As a result, a performance of less than 0.1 second was obtained at the start of recording of speech. The energy-based voice activity detection (VAD) detects the presence of voice using a threshold of energy (dB). Specifically, the energy of the recorded voice signal is calculated, and it is determined that voice is present when it exceeds a preset energy threshold. This is an energy-based approach for determining whether voice is present and is used to monitor the ambient voice level before the start of recording.

[0023] Another VAD implemented the webrtcvad, a library that provides the VAD function of WebRTC prepared by Python, and was used to detect voice activity during recording. The sensitivity was initially set with vad = webrtcvad.Vad(3) and mode 3. The settings can be adjusted according to the on-site environment. Since the RATE is 16000 (Hz) and the 10-ms frame is 160 frames (16000 * 0.01), num_frames = int(RATE / 1000 * 10) was set.

[0024] The voice detection interval is based on the set RECORD_SECONDS value (20 seconds). This means that after voice is detected, recording will be performed for a maximum of 20 seconds, and the setting can be changed according to the site. The frequency range of the detected voice is determined by the sampling rate (RATE = 16000 Hz). Generally, the audible frequency range of humans is from 20 Hz to 20 kHz, but in this system, it is possible to detect up to the frequency component of 8 kHz at most. According to the Nyquist theorem, frequencies up to half of the sampling rate can be detected.

[0025] Assuming that the maximum frequency audible to the human ear is about 20 kHz, at least a sampling rate of 40 kHz is required to digitize this signal. In practice, it is common to use a sampling rate slightly higher than this theoretical value. For example, CD audio is sampled at 44.1 kHz, which is an appropriate rate to cover the human audible frequency range (20 Hz to 20 kHz). However, it was determined that a sampling rate of 16.0 kHz is sufficient for care records at the care site.

[0026] As a result of testing in the field environment, it was found that the energy threshold of the voice signal at the care site is preferably set to the following values. This makes it possible to set the energy threshold according to the field environment. Wired microphone: 32000 ± 5000 (db) Wireless microphone: 25000 ± 5000 (db)

[0027] Furthermore, it was found that having voice recognition such as "start" at the care site may make the care recipient feel suspicious and is not suitable for a quiet environment. In the VAD based on the energy of the voice signal developed this time, it was solved by starting the recording when there is no unnaturalness at the start of recording during normal conversation and by flicking the microphone with a finger in an environment where silence is required. The finger flicking sound generated when flicking the microphone was recorded by activating a function of a Python script. When the microphone is flicked with a finger, an energy of sound larger than the voice signal (for example, 32000 to 37000 (db)) is observed in a shorter time than the voice signal (for example, within 1 second, or between 0.1 and 1 second), so it is easy to recognize that the microphone is flicked with a finger.

[0028] That is, by taking advantage of the drawback of Simple Audio Processing, which is "susceptible to noise", and reacting to the largest noise, such as flicking the microphone itself with a finger, the specification is such that a function for recording with a Python script is activated.

[0029] In addition, in accordance with the caregiving site environment such as at night, the specification is such that the notification process from the portable terminal device 10 to the caregiver is switched from the buzzer type to the vibrator type. The control of the portable terminal device 10 is to notify the caregiver of caregiving information using a buzzer during the day and to notify the caregiver of caregiving information using a vibrator at night. Note that, for example, the night time can be set from 18:00 to 8:00. Or, at night, it can be set to a state where the caregiving staff is working in a night shift with fewer staff than during the day.

[0030] As shown in FIG. 2, in the caregiving information generation system according to the present embodiment, a method of activating a function for recording with a Python script installed in the freehand communication terminal device 10 that can be carried by the caregiving staff was considered. The contents considered are as follows.

[0031] 1. Vosk Vosk is a library or speech recognition API that supports offline speech recognition. It is based on the open-source speech recognition framework Kaldi. The advantages of Vosk are as follows. · Offline operation: An internet connection is not required, and the speech is converted into text data locally for judgment, so a reduction in time lag was expected. · In addition, many language models are available and can be customized (it is also possible to train and use your own Kaldi model). The results of Vosk are as follows. · Due to the performance limitation of the communication terminal device 10 using a control unit such as a lightweight and compact microcontroller for portability, the real-time performance decreased and it was not suitable, but it can be implemented if a control unit such as a microcontroller with higher processing performance is used.

[0032] 2. Speech Recognition Speech Recognition is a Python speech recognition library provided by Google and supports multiple speech recognition engines.

[0033] The advantages of Speech Recognition are as follows. · Simple setup and use · Support for multiple engines · Some engines (e.g., CMU Sphinx) support offline operation

[0034] The disadvantages of Speech Recognition are as follows. · Due to the performance limitations of the communication terminal device 10 using a control unit such as a lightweight and compact microcontroller for portability, the real-time performance decreased and it was not suitable. However, it can be implemented by using a control unit such as a microcontroller with higher processing performance.

[0035] 3.Deep Learning Based Approaches Deep learning was adopted for the speech recognition and language processing (OCR) of this embodiment. A keyword speech recognition model was trained using deep learning. The generation AI engine G of this embodiment is a Transformer-based model.

[0036] The advantages of deep learning are as follows. · High accuracy · Flexible to handle noise and different speakers · Can adapt to various situations and speakers if trained using a large amount of data

[0037] The disadvantages of deep learning are as follows. · Collection and preprocessing of training data are required, and a large amount of training data and training time are needed. · In the communication terminal device 10 using a control unit such as a lightweight and compact microcontroller for portability, it may be difficult to execute in real time. · Advanced knowledge and experience are required for model training and adjustment.

[0038] The inventors of the present invention were able to achieve performance that can be used at the care site by solving all the above-mentioned drawbacks through the optimization of the cloud server E and the generative AI engine G.

[0039] A flowchart of the care information generation system according to the present embodiment is shown in FIG. 2A. In step S001, with the microphone of the care staff activated, the control unit of the communication terminal device 10 determines whether it is a wired microphone. The determination as to whether it is a wired microphone is executed based on a type display value indicating wired or wireless, which is preset for each microphone. If the control unit determines in step S001 that it is a wired microphone, the process proceeds to step S002. If the control unit determines in step S001 that it is not a wired microphone, the process proceeds to step S005.

[0040] In step S0002, the control unit of the communication terminal device 10 determines whether the energy threshold value (wired threshold value) for the wired microphone is within a predetermined range (32000 ± 5000 (db)). If the wired threshold value is within the predetermined range, the control unit determines that the input data (voice activity data) of the wired microphone is voice and transmits it to the cloud server E via the Internet H, and the process proceeds to step S003. If the energy threshold value is not within the predetermined range, the process returns to step S001.

[0041] In step S003, the control unit starts recording the voice activity data, and after a lapse of a predetermined time (for example, 20 minutes) or a lapse of a predetermined silent time (for example, 2 minutes), the recorded voice data is transmitted to the cloud server E. In step S004, the generative AI engine G performs voice recognition on the voice data received via the cloud server E. In step S005, the generative AI engine G generates output data based on the recognized voice data. The generated output data is transmitted via the cloud server E to a personal computer to which the care monitor C is connected, and the personal computer displays the output data on the care monitor C. The output data (care information) displayed on the care monitor C can be, for example, those shown in FIGS. 6 to 18.

[0042] In step S006, the control unit of the communication terminal device 10 determines whether it is a wireless microphone. The determination of whether it is a wireless microphone is executed based on the type display value indicating wired or wireless, which is preset for each microphone. In step S007, the control unit determines whether the threshold value of the energy for the wireless microphone (wireless threshold value) is within a predetermined range (25000 ± 5000 (db)). If the wireless threshold value is within the predetermined range, the process proceeds to step S003. If the wireless threshold value is not within the predetermined range, the process returns to step S001.

[0043] As shown in FIG. 33, speech separation countermeasures were taken as preprocessing for speech recognition. Even when trying to input voice to a recording device at an actual site, various "noises" such as the following 1) to 3) will be mixed in. Therefore, this noise is removed by preprocessing.

[0044] 1) Interference sound: The voice of a person other than the person trying to input 2) Background noise: Noises such as machine sounds, air conditioner and ventilation fan sounds 3) Reverberation sound: The reverberation sound when the sound bounces back from walls, floors, and ceilings

[0045] The influence of insufficient preprocessing accumulates / propagates to subsequent processes. Therefore, it has been common sense until now that removing noise thoroughly by sound source separation improves the accuracy of the next speech recognition. However, it has been found that when directly loading voice into the generation AI engine G, for example, when identifying and transcribing the utterances of caregiver A and caregiver B from the conversations of multiple people, voice separation according to the care site environment is necessary.

[0046] In order to perform the above-mentioned voice separation in 2) and 3), the inventors implemented webrtcvad, a library that provides the VAD function of WebRTC provided by Python. The sensitivity was initially set with vad = webrtcvad.Vad(3) and mode 3, but it can be adjusted according to the on-site environment. Since RATE is 16000 (Hz) and a 10-ms frame is 160 frames (16000 * 0.01), num_frames = int(RATE / 1000 * 10) was set.

[0047] In order to solve 1) initially, the inventors established the specifications for the care site with two or more microphones using the attenuation / delay theory of the sound source signal based on the steering vectors in FIGS. 34, 35, 36, and 37. However, currently in the field of natural language processing (NLP), Transformer, an architecture of a neural network model that is widely used, has become the mainstream. Therefore, this time, sound source separation was performed using the generative AI engine G. As will be described later, rather than aiming to separate speakers, the inventors invented a processing / method suitable for the care site by raising the feelings that humans feel with their five senses in the care site environment to a level that the generative AI engine G can understand.

[0048] The inventors performed voice recognition using the generative AI engine G. Currently, the mainstream of the generative AI engine G is a Transformer-based model. That is, by training the Transformer-based model with a large amount of text data and voice data specialized for care, an effective and highly effective model leading to this invention was completed.

[0049] As a result, the memory of the control unit such as the GPU was reduced, and the accuracy of the recognition task for multiple people's conversations was improved. It is important to raise the feelings that humans feel with their five senses in the caregiving environment to a level that the generative AI engine G can understand. Based on the "Generative AI Learning Program" described later, a unique learning method with Figures 22 to 32 as the correct answers trained the generative AI engine G into a professional caregiver. Also, the voice data input daily continues to strengthen the generative AI engine G.

[0050] The generative AI learning program will be described below. The communication terminal device 10 in FIG. 2 mainly converts speech into voice data and uploads it to the cloud server E. In particular, the authentication systems 15A / 15B for authentication keys for obtaining connection authentication with the cloud server become important control points. With regard to the properties of the folders and files related to the authentication system, in addition to access control and hidden file settings, when opening a dummy Python script and a dummy cloud authentication key, a system is added in which voice data of malicious unauthorized accessers and sound source data of the surrounding environment are uploaded.

[0051] The microphone and speaker of the headset 11 are compatible with sensing technologies such as Bluetooth (registered trademark) - type headset sets, wireless pin microphones, wired pin microphones, small speakers, buzzers, and vibration sensors according to the caregiving environment.

[0052] The recorded voice data is immediately uploaded to the storage of the cloud server E. The voice data stored in the storage 17 is transcribed by "Script 1" 18A that utilizes the generative AI engine G without delay to generate text data.

[0053] In the case of multiple conversations, conventionally, the identification of the speaker has generally been carried out by using a stereo microphone to calculate the voice transmission distance between the speaker and microphones 1 and 2 for identification. The care information generation system 1000 of this embodiment causes the generation AI engine G to read voice data, and at the same time, as shown in FIG. 4, the inventors or individuals and organizations D commissioned by the inventors, the developer PC / server 10D, implements a learning program on the generation AI engine G, enabling the generation AI to determine the speaking style (habits, vocal range, speed, etc.) of the speaker and recognize who is speaking. For example, from the voice data obtained by recording the conversation of three care staff members, each speaker can be identified and transcribed into text.

[0054] The monitoring process or method by the generation AI learning program is detailed below. The generated text file is identified and selected by the generation AI engine G from its speech content into "sudden change", "emergency", "important", "postponement", "case report", "business contact", "care record", and other files.

[0055] The generated text data such as "sudden change", "emergency", "important", "postponement", "business contact", etc. is uploaded to a dedicated website F and displayed on a care monitor such as the on-site cafeteria as a quick report as shown in FIG. 6. The URL of the dedicated website can only be viewed by known persons.

[0056] Also, the generated text data such as "sudden change", "emergency", "important", "postponement", "business contact", etc. is converted back into voice data by voice synthesis technology, downloaded to the communication terminal device 10 of the care staff operating at the same care site, and transmitted as voice data.

[0057] Data such as "case report" and "care record" is analyzed, tabulated, and graphed by the generation AI engine G and can be displayed on the care monitor 10C. As a result, information on the situation of the care recipient, such as the amount of food and water intake today at the current time, can be confirmed at the care site.

[0058] The care information system / method developed this time was realized by a process / method of specifying the scope, type, etc. to be analyzed for data, and instructing the AI that generates the form to be output. By learning the standard format for data analysis, it can be output in a reproducible format.

[0059] In order to display the meal intake of all care recipients at the current time on the care monitor 10C, an example of the standard instruction manual "Data Analysis_Meal Intake" for the generation AI is shown in FIG. 7. And FIG. 8 is an example of the meal intake at 11:00 am actually displayed on the care monitor 10C. Since the care recipients with less meal intake are displayed on the care monitor 10C in order, the care staff who noticed it can respond quickly. FIG. 9 is also an example of the result of the meal intake at 8:00 pm.

[0060] For the water volume, the instruction manual and its result of the example of "Data Analysis_Water Volume" in FIG. 10 are shown in FIG. 11. It is set so that the care recipients with less water volume come to the upper row.

[0061] The meal intake, water volume, blood pressure, SPO2, and special notes for each care recipient are instructed in FIG. 12 "Data Analysis_User Status List", and as a result, FIG. 13 is displayed on the care monitor 10C.

[0062] The causes of abnormalities include recording omissions, recording errors, etc., and the purpose is to detect any abnormalities.

[0063] These care monitor displays are automatically updated, and since the current situation of the care recipients on site can be grasped, it can be said to be a static system.

[0064] After a care staff member or a care facility administrator speaks the words "Data Analysis" by voice into the communication terminal device 10 and then instructs the information they want to know, a flexible data analysis system that does not adhere only to the standard has been developed.

[0065] For example, when a caregiver wants to know the current situation of the care recipient "Mr. Higuchi", it can be achieved by the following two processes / methods.

[0066] First, if you want to receive a report of the situation along the "template" in Fig. 14, the result in Fig. 15 will be obtained. Second, if you give instructions freely without a template, the result as shown in Fig. 16 will be obtained.

[0067] Since it is left to the free will of the caregiver or the caregiver facility manager what kind of information they want to obtain, it can be said that it is a dynamic system for grasping the situation of the care services within the facility.

[0068] Data such as "case reports" and "care records" are transcribed by the generative AI engine G, processed by a morphological analysis engine such as MeCab, and registered in the database.

[0069] Figs. 17 and 18 are examples of browsing screens and printing screens. Fig. 19 is an example of creating a graph. The size of the circles indicates the difference in excretion volume.

[0070] Voice data that is not intentionally spoken other than the above contains the raw voices at the care site, so it contains important information leading to high-quality care services. Also, it serves as an information source for grasping the internal and external situations of the organization and becomes monitoring data for the organization's overall management system. A system has been constructed to select the information desired by the caregiver facility manager and perform data analysis.

[0071] The inventor or the developer PC / server 10D of the individual and organization D entrusted by the inventor manages the cloud server E as shown in Fig. 3. The developer PC / server 10D performs security management of the authentication system 15B, management of the storage 17, and management of various scripts and the website F.

[0072] The developer PC / server 10D of the present inventors or individuals and organizations D commissioned by the present inventors executes the learning program of the generation AI engine G as shown in FIG. 4. As a result, highly specialized information such as medical care, nursing care, and regulations can be provided to the generation AI engine G and it can be nurtured. Further, the developer PC / server 10D manages the communication terminal device 10 carried by the nursing staff and the script as shown in FIG. 5.

[0073] Generally, data analysis, table creation, and drawing of nursing care services are realized by registering text data in a DB (database) and performing data processing. This system inputs image data including text data and video into the generation AI engine G and appropriately instructs it to automatically execute table creation and drawing. These output materials are input as a nursing care summary into the daily operation management of nursing care services, regular reviews, and nursing care plans such as care plans.

[0074] The present inventors have developed a system that enables a conversation with the generation AI engine G by instructing the information to be known after a nursing staff or a nursing care facility manager speaks "question" into the communication terminal device 10 in voice.

[0075] FIG. 20 shows the answer when asking the generation AI engine G about the state of an initial pressure ulcer (bedsores).

[0076] Since it is left to the free will of the nursing staff or the nursing care facility manager what kind of information they want to obtain, it can be said to be a dynamic system. Further, since it is a voice-based conversation system with the generation AI engine G, various questions such as medical and nursing care specialized information and nursing care insurance are possible.

[0077] In the caregiving field, due to a shortage of manpower, foreign workers are employed. It has been made possible to create records corresponding to the native languages of foreign workers as one of the instructions. Figure 20 shows an example in which Indonesian care staff are made to explain to an AI that generates pressure ulcers (bedsores) and at the same time translated into the Indonesian language. This can prevent mistakes in caregiving services and can also be used for education and training. By connecting the monitor of the communication terminal device 10A (for example, 3.2 inches to 5 inches) and translating (for example, English or Indonesian) and displaying the results provided by the generation AI engine, communication becomes easier on the spot.

[0078] Figure 21 shows the result of transcribing the conversation between care staff A and office staff B at the caregiving site using the communication terminal device 10A by the generation AI engine G. Despite the fact that the microphone 11A has a microphone number = 1, the generation AI of the generation AI engine G can distinguish the voices of two people and accurately transcribes them. However, the generation AI misidentifies care staff A as the mother and office staff B as the daughter and makes the mistake of judging that it is a conversation between a mother and a daughter. Since office staff B is a foreign worker with a stuttering Japanese and a child-like voice quality, it is interesting that the generation AI engine G made such a judgment.

[0079] After that, by the generation AI learning program, the generation AI engine G was made to learn the voice of office staff B, and it became able to recognize office staff B. The generation AI learning program of the present invention is the most effective and reproducible invention developed by the inventors.

[0080] The inventors developed this system at the actual site of a care facility (social welfare corporation) with 160 care recipients and 150 care staff. The inventors investigated the care records (referred to as face sheets at the facility) of this care facility for 13 years and 2 months from October 30, 2008 to January 25, 2022 and summarized them in the following items.

[0081] 1. Corpus text (meal) 195 patterns (see Figure 22) 2. Corpus text (excretion) 232 patterns (see Figure 23) 3. Corpus text (bathing) 207 patterns (see Figure 24) 4. Corpus text (movement) 177 patterns (see Figure 25) 5. Corpus text (oral cavity) 78 patterns (see Figure 26) 6. Corpus text (care plan) 144 patterns (see Figure 27) 7. Corpus text (medical terms) 231 terms (see Figure 28) 8. Corpus text (caregiving terms) 58 terms (see Figure 29)

[0082] The inventors made the above patterns 1 to 8 as "correct answers" and let the generated AI engine G learn in two ways via the developer PC / server D in Figure 4.

[0083] 1) Learn as a text file 2) Learn as voice data For the method of creating the corpus, "Creating JUST Corpus, Shinnosuke Takamichi, Assistant Professor, Graduate School of Information Science and Technology, The University of Tokyo, 2023" was referred to.

[0084] The inventors let the generated AI engine G learn the following via the developer PC / server D in Figure 4. 1. The name and reading of the care recipient 2. The name and reading of the caregiver These need to be updated regularly.

[0085] The inventors let the generated AI engine G learn the following from the developer PC / server D in Figure 4. 1. For the speech in which the generated AI engine made a mistake in speech recognition, let it learn the correction table in the "replacements.json" file (see Figure 30). 2. Let it learn the glossary that must be kept in the "keep_patterns.json" file (see Figure 31). 3. When spoken, a collection of important terms (magic words) that are selected and identified from other audio files are shown in the "magic_words.json" file (see Figure 32). These need to be updated regularly.

Explanation of Symbols

[0086] 1000 Care Information Generation System A Care Staff B Care Staff C Care Monitor D Developer PC / Server E Cloud Server F Website G Generation AI Engine H Internet 10, 10A, 10B Communication Terminal Device 11, 11A, 11B Headset (Microphone, Speaker)

Claims

1. A care information generation system including at least one terminal device and a generation AI engine for generating care information from voices at a care site, The terminal device includes a microphone that converts voices at the care site into voice signals, determines the generation of voice activity data based on the voice signals, and transmits the voice activity data to the generation AI engine; A care information generation system in which the voice activity data received from the terminal device is input to the generation AI engine, and the generation AI engine generates the care information based on the voice activity data.

2. The care information generating system according to claim 1, A care information generating system, wherein the care information is displayed on a care monitor.

3. The care information generating system according to claim 1, A care information generation system in which the generative AI engine is trained using a corpus related to care and / or medical care.

4. The care information generating system according to claim 1, A care information generating system, wherein the control unit of the terminal device monitors voice energy using energy-based voice activity detection (VAD) and determines that the voice activity data has occurred when the voice energy threshold is within a predetermined range.

5. The care information generating system according to claim 4, the microphone of the terminal device is a wired microphone or a wireless microphone, A care information generating system, wherein the control unit of the terminal device is provided with a wired threshold of sound energy for the wired microphone and a wireless threshold of sound energy for the wireless microphone.

6. 6. The care information generating system according to claim 5, A care information generating system, wherein the wired threshold is 32000±5000 (db) and the wireless threshold is 25000±5000 (db).

7. The care information generating system according to claim 1, A care information generating system in which the control unit of the terminal device recognizes a finger flicking sound generated by flicking the microphone with a finger and starts recording.

8. The care information generating system according to claim 7, the finger snapping sound has a shorter duration than the audio signal and has greater sound energy than the audio signal;

9. The care information generating system according to claim 7, After recognizing the finger snapping sound, the control unit activates a Python script function to start recording.

10. The care information generating system according to claim 1, The terminal device notifies the caregiver of the care information by using a buzzer during the day and a vibrator at night.

11. The care information generating system according to claim 1, A care information generation system comprising a PC / server that provides learning data to the generation AI engine.

12. The care information generating system according to claim 1, The care information generation system includes a data processing server capable of communicating with the terminal device and the generation AI engine, A care information generation system in which the data processing server receives the voice activity data from the terminal device and provides the voice activity data to the generation AI engine.