Nursing care information generation system
The care information generation system addresses the burden on care staff by using a terminal device with voice activity detection and a generation AI engine to automate record creation and enhance communication efficiency, particularly in noisy environments, thereby improving care information generation accuracy and reducing staff workload.
Patent Information
- Application Number
- JP2024006560
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-01-19
- Publication Date
- 2026-01-29
- Estimated Expiration
- 2044-01-19
AI Technical Summary
Care staff face significant burdens in creating care records manually, leading to overtime work and increased workload, especially when dealing with urgent matters, handovers, and business communications, and existing voice recognition systems struggle with specialized medical and care terminology, necessitating a more efficient and accurate method for generating care information.
A care information generation system utilizing a terminal device with a microphone that converts voices into voice signals, employing energy-based voice activity detection and a generation AI engine to recognize and generate care information, including training with care and medical-related corpora, and implementing speech separation techniques to improve accuracy.
The system efficiently recognizes necessary voices and generates care information with high accuracy, reducing staff workload by automating record creation and improving communication efficiency, especially in noisy environments, and supports care staff leaders with real-time data analysis and monitoring.
Smart Images

Figure 0007808349000001 
Figure 0007808349000002 
Figure 0007808349000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to a care information generation system that uses generation AI. [Background technology]
[0002] The care system described in Patent Document 1 is capable of displaying information obtained by voice recognition of calls made between care staff and care recipients via receivers and handsets on a touch panel. However, care staff obtain a large amount of information through care services, and the types of information vary, including urgent matters, care records, handovers, and business communications, so it is necessary to utilize the information in an appropriate output method depending on the nature of the input information. In addition, recording is initiated by pressing a button, and generally requires device operation, even on smartphones and tablets.
[0003] In addition, natural language processing technology has advanced to the point where it has included sound source separation, speech recognition, and the Transformer language processing technology, but the development environment has changed dramatically in recent years with the emergence of generative AI.
[0004] The voice recognition method used in the care information generation system described in Patent Document 1 utilizes a conversation app, and it is known that the recognition rate drops for speech that is highly specialized, such as medical and care terminology. [Prior art documents] [Patent documents]
[0005] [Patent Document 1] Japanese Patent Publication No. 2022-56253 Summary of the Invention [Problem to be solved by the invention]
[0006] After providing care services to care recipients, care staff create care records. These records are created by hand on paper or manually input into a computer or tablet. This work places a heavy burden on care staff and causes overtime work. Care staff also have to follow up with foreign workers and part-timers, which increases the burden on care staff.
[0007] While providing care services to the care recipient, care staff will record information by voice freehand without any time lag, and urgent matters, handovers, and business communications will be promptly communicated to the relevant parties. In addition, the care records will be analyzed and compiled into tables before being communicated to the relevant parties. Furthermore, at the end of the shift, a proposed care record will be automatically generated by the generation AI.
[0008] The typical method of transcribing spoken voice data and converting it into text data is through a conversation app. By using a generative AI to transcribe this text, the generative AI can be trained to "develop AI into a first-class caregiver." The AI, having learned medical and nursing knowledge and the actual conditions of the nursing care field, will propose nursing record ideas to the care staff.
[0009] Care staff leaders monitor the care site in accordance with the care schedule, provide appropriate encouragement and guidance, and control the care site. However, the workload of care staff leaders is heavy and it is a demanding role. A system that utilizes generative AI to support the work of care staff leaders is needed.
[0010] Generally, this is achieved by registering nursing care information such as data analysis data, tabulation data, drawing data, and text data in a database and then running a script. This requires a long development period and modifying the script is also time-consuming. The present invention provides a care information generation system that can recognize necessary voices from voices in a care site and generate care information. [Means for solving the problem]
[0011] The aspects of the present invention are as follows. (Example 1) A care information generation system including at least one terminal device and a generation AI engine for generating care information from voices in a care site, The terminal device includes a microphone that converts voices at the care site into voice signals, recognizes the generation of voice activity data based on the voice signals, and transmits the voice activity data to the generation AI engine; A care information generation system in which the voice activity data received from the terminal device is input to the generation AI engine, and the generation AI engine generates the care information based on the voice activity data.
[0012] (Example 2) In the care information generation system according to the first aspect, A care information generating system, wherein the care information is displayed on a care monitor. (Example 3) In the care information generation system according to the first aspect, A care information generation system, wherein the generation AI engine is trained using a corpus related to care and / or medical care. (Example 4) In the care information generation system according to claim 1, A care information generation system in which the control unit of the terminal device monitors voice energy using energy-based voice activity detection (VAD) and determines that the voice activity data has occurred when the voice energy threshold is within a predetermined range.
[0013] (Example 5) In the care information generation system according to Example 4, the microphone of the terminal device is a wired microphone or a wireless microphone, A care information generating system, wherein the control unit of the terminal device is provided with a wired threshold of sound energy for the wired microphone and a wireless threshold of sound energy for the wireless microphone. (Example 6) In the care information generation system according to Example 5, A care information generating system, wherein the wired threshold is 32000±5000 (db) and the wireless threshold is 25000±5000 (db). (Example 7) In the care information generation system according to the first aspect, A care information generating system in which the control unit of the terminal device recognizes a finger flicking sound generated by flicking the microphone with a finger and starts recording.
[0014] (Example 8) In the care information generation system according to Example 7, the finger snapping sound has a shorter duration than the audio signal and has greater sound energy than the audio signal; (Example 9) In the care information generation system according to Example 7, After recognizing the sound, the control unit activates a function of a Python script to start recording. (Example 10) In the care information generation system according to claim 1, The terminal device notifies the caregiver of the care information by using a buzzer during the day and a vibrator at night.
[0015] (Example 11) In the care information generation system according to claim 1, A care information generation system comprising a PC / server that provides learning data to the generation AI engine. (Example 12) In the care information generation system according to claim 1, the care information generation system includes a data processing server capable of communicating with the terminal device and the AI engine, The data processing server receives the voice activity data from the terminal device and provides the voice activity data to the generation AI engine. [Effects of the Invention]
[0016] The care information generating system of the present invention can recognize necessary voices from voices in the care site and generate care information. [Brief explanation of the drawings]
[0017] [Figure 1] 1 is a schematic configuration diagram of a care information generation system according to an embodiment of the present invention. [Figure 2] FIG. 2 is a block diagram of the care information generation system of FIG. 1. [Figure 2A] 2 is a flowchart of the care information generation system of FIG. 1. [Figure 3] FIG. 2 is a configuration diagram showing the processing between the developer PC / server and the terminal device in FIG. [Figure 4] FIG. 2 is a configuration diagram showing the processing between the generation AI engine of FIG. 1 and the terminal device. [Figure 5] FIG. 2 is a configuration diagram showing the processing between the developer PC / server and the terminal device in FIG. [Figure 6] 2 is a schematic diagram showing an example of a news monitor display of the care information generation system of FIG. 1. FIG. [Figure 7] FIG. 10 is a schematic diagram showing an example of the display of the fixed instruction form "Data Analysis_Dietary Amount." [Figure 8] FIG. 10 is a schematic diagram showing a display example of an example of the amount of food eaten at 11:00 AM. [Figure 9] FIG. 10 is a schematic diagram showing a display example of an example of the amount of food eaten at 8 p.m. [Figure 10] FIG. 10 is a schematic diagram showing an example of the display of the fixed instruction form "Data Analysis_Moisture Content." [Figure 11] FIG. 10 is a schematic diagram showing a display example of a moisture content list monitor. [Figure 12] FIG. 10 is a schematic diagram showing a data analysis user status list. [Figure 13] FIG. 10 is a schematic diagram showing a display example of a user status list monitor. [Figure 14] FIG. 10 is a schematic diagram showing an example of a status report template. [Figure 15] FIG. 10 is a schematic diagram showing an example of input to a status report template. [Figure 16] FIG. 10 is a schematic diagram showing an example of a status report without a template. [Figure 17] FIG. 10 is a schematic diagram showing a viewing screen that displays the situation at each time point. [Figure 18] FIG. 10 is a schematic diagram showing a summary screen that displays the status of one day. [Figure 19] 1 is a graph showing data analysis. [Figure 20] This is a schematic diagram showing an example of the use of voice-generated AI. [Figure 21] FIG. 1 is a schematic diagram showing an example of speech recognition by a generation AI. [Figure 22] FIG. 1 is a schematic diagram showing a corpus text (food). [Figure 23] FIG. 1 is a schematic diagram showing corpus text (excretion). [Figure 24] FIG. 1 is a schematic diagram showing a corpus text (bathing). [Figure 25] FIG. 1 is a schematic diagram showing corpus text (movement). [Figure 26] FIG. 1 is a schematic diagram showing corpus text (oral cavity). [Figure 27] FIG. 1 is a schematic diagram showing a corpus text (care plan). [Figure 28] FIG. 1 is a schematic diagram showing corpus text (medical terms). [Figure 29] FIG. 1 is a schematic diagram showing corpus text (nursing care terms). [Figure 30] This is a schematic diagram showing an errata table that the generative AI engine learns. [Figure 31] FIG. 1 is a schematic diagram showing a glossary that a generative AI engine learns. [Figure 32]This is a schematic diagram showing a glossary of important terms that the generative AI engine learns. [Figure 33] FIG. 1 is a schematic diagram showing a voice separation process as a preprocessing of voice recognition. [Figure 34] FIG. 1 is an image diagram showing attenuation and delay of a sound source signal. [Figure 35] FIG. 1 is a schematic diagram showing the position of a sound source and sound propagation. [Figure 36] FIG. 10 is a schematic diagram showing a case where it is difficult to detect the direction of a sound source. [Figure 37] FIG. 1 is an image diagram showing the dereverberation effect using two microphones. DETAILED DESCRIPTION OF THE INVENTION
[0018] An embodiment of the care information generation system of the present invention will be described with reference to the drawings. Common parts in the drawings are designated by the same reference numerals, and descriptions thereof will be omitted as appropriate. The present invention is not limited to care information generation systems, but can also be applied to medical information generation systems that can be used in medical settings, and exercise information generation systems that can be used in fitness and other exercise settings.
[0019] In the care information generation system 1000 of this embodiment, image data including text data, voice data, and video data from care staff are input to a generation AI engine. Furthermore, when the care staff inputs appropriate instructions, the care information generation system 1000 automatically outputs care information such as tables and drawings. These output materials are used as care summaries for daily operational management of care services, periodic reviews, and care planning such as care plans.
[0020] Figure 1 shows an overview of the developed nursing care information input system. Caregiver A, who provides on-site nursing care services, uses a microphone 11A on a portable communication terminal device 10A to record the care recipient's condition and other information. The recorded audio data is stored in the storage of a cloud server E (a data processing server) via Internet communications H, including satellite communications. The audio data is then converted into text data using a generation AI engine G using a transcription script installed on the cloud server E, and stored in the storage. The generation AI engine G is installed on the generation AI server. The processed text data is registered / uploaded to website F, where staff can view it via a private URL. Urgent text data is converted back into speech using speech synthesis technology, downloaded to caregiver B's communication terminal device 10B, and transmitted via speaker 11B. Developer D's PC / server 10D, such as a developer PC / server, performs maintenance on the cloud server E. It also runs the generation AI's learning program and provides the generation AI with highly specialized information, such as medical care, nursing care, and legal regulations, to train it.
[0021] As shown in Figure 2, the inventors have developed a hands-free communication terminal device 10 that can be carried by care staff. It allows voice recording while providing care services without requiring any operations such as pressing buttons or scrolling the screen. When repeatedly tested in care settings, for example, the method of starting recording by uttering "start" resulted in a time lag of more than three seconds because the utterance had to be converted into text data before recording could be judged.
[0022] Therefore, in this embodiment, we use Simple Audio Processing to monitor voice energy using energy-based voice activity detection (VAD), and when a certain threshold is exceeded, the occurrence of voice activity data is determined / recognized. This resulted in a performance of less than 0.1 seconds from the start of speech recording. Energy-based voice activity detection (VAD) detects the presence of voice using an energy (db) threshold. Specifically, it calculates the energy of the recorded voice signal and determines that voice is present when it exceeds a pre-set energy threshold. This energy-based approach to determining whether voice is present was used to monitor the environmental voice level before recording began.
[0023] The other VAD was implemented using webrtcvad, a Python library that provides WebRTC VAD functionality, and was used to detect voice activity during recording. The sensitivity was initially set to mode 3 with vad=webrtcvad.Vad(3). The settings were adjusted to suit the on-site environment. Since the RATE is 16000 (Hz) and 10 ms frames are 160 frames (16000 * 0.01), num_frames = int(RATE / 1000 * 10) was used.
[0024] The audio detection interval is based on the set RECORD_SECONDS value (20 seconds). This means that after audio is detected, it will be recorded for a maximum of 20 seconds, and the setting can be changed to suit the site. The frequency range of the detected audio is determined by the sampling rate (RATE = 16000 Hz). Generally, the audible frequency range for humans is from 20 Hz to 20 kHz, but this system can detect frequency components up to 8 kHz. According to the Nyquist theorem, frequencies up to half the sampling rate can be detected.
[0025] Given that the maximum frequency audible to the human ear is approximately 20 kHz, a sampling rate of at least 40 kHz is required to digitize this signal. In practice, sampling rates slightly higher than this theoretical value are generally used. For example, CD audio is sampled at 44.1 kHz, which is an appropriate rate for covering the human audible frequency range (20 Hz to 20 kHz). However, we determined that a sampling rate of 16.0 kHz is sufficient for recording care in nursing care settings.
[0026] Testing in the nursing care environment revealed that the following values are preferable for the energy threshold of audio signals. This makes it possible to set the energy threshold according to the environment. Wired microphone: 32000±5000(dB) Wireless microphone: 25000±5000(db)
[0027] Furthermore, it was found that voice recognition of commands such as "start" in nursing care settings can make care recipients suspicious, making it unsuitable for quiet environments. The developed VAD based on the energy of the voice signal solves this problem by starting recording with normal conversation, which does not sound unnatural, and by starting recording in quiet environments by flicking the microphone with a finger. The finger flick sound generated when flicking the microphone was recorded by activating a Python script function. When a microphone is flicked with a finger, a sound energy (e.g., 32,000 to 37,000 (dB)) greater than the sound signal is observed for a shorter period of time than the sound signal (e.g., within 1 second, or between 0.1 and 1 second), making it easy to recognize the flicking of the microphone.
[0028] In other words, we took advantage of Simple Audio Processing's drawback of being susceptible to noise and made it react to the loudest noise, such as flicking the microphone with a finger, to activate a function for recording in a Python script.
[0029] Additionally, the notification process from the mobile terminal device 10 to the caregiver can be switched from a buzzer type to a vibrator type depending on the environment of the care site, such as at night. The mobile terminal device 10 is controlled so that it notifies the caregiver of care information using a buzzer during the day and a vibrator at night. Note that nighttime can be, for example, from 6:00 PM to 8:00 AM. Alternatively, at nighttime, the care staff can work a shorter night shift than during the daytime.
[0030] As shown in Figure 2, in the care information generation system according to this embodiment, we have investigated a method of activating a function for recording using a Python script implemented in a freehand communication terminal device 10 that can be carried by care staff. The details of the investigation are as follows.
[0031] 1. Vosk Vosk is a library or speech recognition API that supports offline speech recognition and is based on the open source speech recognition framework Kaldi. The advantages of Vosk are: Offline operation: No internet connection is required, and speech is converted into text data locally for judgment, which is expected to reduce time lag. Many other language models are available and customizable (you can even train and use Kaldi's models yourself) Vosk's results are as follows: Due to performance limitations of the communication terminal device 10, which uses a control unit such as a lightweight and compact microcontroller for portability, real-time performance is reduced and it is not suitable. However, it is possible to implement this by using a control unit such as a microcontroller with higher processing performance.
[0032] 2. Speech Recognition Speech Recognition is a Python speech recognition library provided by Google that supports multiple speech recognition engines.
[0033] The benefits of Speech Recognition include: Easy to set up and use Supports multiple engines Some engines (e.g. CMU Sphinx) support offline operation
[0034] The disadvantages of Speech Recognition are: Due to performance limitations of the communication terminal device 10, which uses a control unit such as a lightweight and compact microcontroller for portability, real-time performance is reduced and it is not suitable. However, it is possible to implement this by using a control unit such as a microcontroller with higher processing performance.
[0035] 3.Deep Learning Based Approaches Deep learning was employed for speech recognition and language processing (transcription) in this embodiment. Deep learning was used to train a keyword speech recognition model. The generative AI engine G in this embodiment is a Transformer-based model.
[0036] The advantages of deep learning are: High accuracy -Flexible to handle noise and different speakers -Trained with large amounts of data, it can adapt to a variety of situations and speakers
[0037] The drawbacks of deep learning are: -Training data collection and preprocessing are required, requiring a large amount of training data and training time. In the case of a communication terminal device 10 that uses a control unit such as a lightweight and compact microcontroller for portability, it may be difficult to execute the program in real time. -Training and adjusting models requires advanced knowledge and experience
[0038] The inventors have resolved all of the above drawbacks by optimizing the cloud server E and the generation AI engine G, thereby achieving performance that is usable in nursing care settings.
[0039] A flowchart of the care information generating system according to this embodiment is shown in FIG. 2A. In step S001, when the care staff member's microphone is activated, the control unit of the communication terminal device 10 determines whether it is a wired microphone. The determination of whether it is a wired microphone is made based on a type display value indicating wired or wireless, which is preset for each microphone. If the control unit determines in step S001 that it is a wired microphone, the process proceeds to step S002. If the control unit determines in step S001 that it is not a wired microphone, the process proceeds to step S005.
[0040] Steps S002 Then, the control unit of the communication terminal device 10 determines whether the energy threshold for the wired microphone (wired threshold) is within a predetermined range (32000±5000 (db)). If the wired threshold is within the predetermined range, the control unit determines that the input data (voice activity data) from the wired microphone is voice, transmits it to the cloud server E via the Internet H, and proceeds to step S003. If the energy threshold is not within the predetermined range, the process returns to step S001.
[0041] In step S003, the control unit starts recording voice activity data, and after a predetermined time (e.g., 20 minutes) or a predetermined silence period (e.g., 2 minutes) has elapsed, the control unit transmits the recorded voice data to the cloud server E. In step S004, the generation AI engine G performs voice recognition on the voice data received via the cloud server E. In step S005, the generation AI engine G generates output data based on the recognized voice data. The generated output data is transmitted via the cloud server E to a personal computer to which the care monitor C is connected, and the personal computer displays the output data on the care monitor C. The output data (care information) displayed on the care monitor C can be, for example, the data shown in FIGS. 6 to 18.
[0042] In step S006, the control unit of the communication terminal device 10 determines whether the microphone is a wireless microphone. The determination of whether the microphone is a wireless microphone is made based on a type display value indicating whether the microphone is wired or wireless, which is preset for each microphone. In step S007, the control unit determines whether the energy threshold for the wireless microphone (wireless threshold) is within a predetermined range (25000±5000 (db)). If the wireless threshold is within the predetermined range, the process proceeds to step S003. If the wireless threshold is not within the predetermined range, the process returns to step S001.
[0043] As shown in Figure 33, speech separation measures were taken as preprocessing for speech recognition. When trying to input audio into a recording device in the field, various "noises" such as 1) to 3) below will be mixed in. Therefore, this noise must be removed by pre-processing.
[0044] 1) Interference: Voices spoken by people other than the person trying to type 2) Background noise: Noise from machinery, air conditioners, ventilation fans, etc. 3) Reverberation: Sound bounces off the walls, floor, and ceiling
[0045] The effects of insufficient preprocessing accumulate and propagate to subsequent processing. Therefore, it has been common knowledge until now that thoroughly removing noise during sound source separation improves the accuracy of subsequent speech recognition. However, when directly loading audio into the generative AI engine G, for example, to distinguish and transcribe the speech of care staff A and B from a conversation between multiple people, we found that speech separation appropriate to the nursing care environment is necessary.
[0046] To perform the voice separation described in 2) and 3) above, the inventors implemented webrtcvad, a Python library that provides WebRTC's VAD functionality. The sensitivity was initially set to mode 3 (vad = webrtcvad.Vad(3)), but it can be adjusted according to the on-site environment. Since the RATE is 16000 (Hz) and there are 160 frames per 10 ms (16000 * 0.01), the num_frames value was set to int(RATE / 1000 * 10).
[0047] To solve problem 1), the inventors initially established specifications for nursing care settings with two or more microphones using the theory of sound source signal attenuation and delay using the steering vectors shown in Figures 34, 35, 36, and 37. However, in the field of natural language processing (NLP), Transformer, a widely used neural network model architecture, has become mainstream. Therefore, this time we performed sound source separation using the generative AI engine G. As will be described later, rather than aiming to separate speakers, we invented a processing / method suited to nursing care settings by raising the sensations felt by humans through their five senses in a nursing care environment to a level that can be understood by the generative AI engine G.
[0048] The inventors performed speech recognition using a generative AI engine G. Currently, the generative AI engine G is mainly based on a Transformer-based model. In other words, by training a Transformer-based model with a large amount of text data and speech data specialized for nursing care, an effective and highly efficient model that led to the present invention was completed.
[0049] As a result, memory requirements for control units such as the GPU were reduced, improving the accuracy of tasks that recognize conversations between multiple people. It is important to raise the level of understanding of the sensations experienced by humans through their five senses in a caregiving environment to a level that the Generative AI Engine G can understand. Based on the "Generative AI Learning Program" described below, a unique learning method that uses Figures 22 to 32 as correct answers has trained the Generative AI Engine G to become a professional caregiver. Furthermore, the voice data that is input daily continues to further strengthen the Generative AI Engine G.
[0050] The generative AI learning program is described below. The communication terminal device 10 in Fig. 2 mainly converts speech into audio data and uploads it to the cloud server E. In particular, the authentication systems 15A / 15B related to the authentication key for obtaining connection authentication with the cloud server are important control points. In addition to access control and hidden file settings in the properties of folders and files related to the authentication system, a system has been added in which, when a dummy Python script and a dummy cloud authentication key are opened, the voice of a malicious unauthorized accesser and sound source data of the surrounding environment are uploaded.
[0051] The microphone and speaker of the headset 11 are adapted to the nursing care environment and are compatible with sensing technologies such as a Bluetooth (registered trademark) headphone set, wireless pin microphone, wired pin microphone, small speaker, buzzer and vibration sensor.
[0052] The recorded voice data is immediately uploaded to the storage of the cloud server E. The voice data stored in the storage 17 is transcribed without delay by "Script 1" 18A, which utilizes the generation AI engine G, to generate text data.
[0053] In the case of a conversation between multiple people, conventionally, speakers were identified by using a stereo microphone and calculating the distance of sound transmission between the speaker and microphones 1 and 2. In the care information generation system 1000 of this embodiment, voice data is read into the generation AI engine G, and simultaneously, as shown in FIG. 4, the inventors or an individual or organization D commissioned by the inventors executes a learning program on the generation AI engine G via a developer PC / server 10D, allowing the generation AI to determine the speaker's speaking style (habits, range, speed, etc.) and recognize who is speaking. For example, from the voice data recorded during a conversation between three care staff members, each speaker can be identified and transcribed.
[0054] The monitoring process or method by the generative AI learning program is detailed below. The generated text files are identified and selected by the generation AI engine G based on the content of the speech as "sudden change," "urgent," "important," "handover," "case report," "business communication," "care record," and other files.
[0055] The generated text data, such as "sudden change," "urgent," "important," "transfer," and "business notice," is uploaded to a dedicated website F and displayed as breaking news on nursing monitors in the on-site cafeteria, etc., as shown in Figure 6. Only known individuals can view the dedicated website URL.
[0056] In addition, the generated text data such as "sudden change," "urgent," "important," "handover," and "business notice" is converted back into voice data using voice synthesis technology, downloaded to the communication terminal device 10 of the care staff working at the same care site, and transmitted as voice data.
[0057] Data such as "case reports" and "care records" can be analyzed, tabulated, and illustrated by the AI generation engine G, and displayed on the care monitor 10C. This allows the care recipient's current condition, such as the amount of food and water consumed that day, to be checked at the care site.
[0058] The newly developed nursing care information system / method is realized by specifying the range and type of data to be analyzed and instructing the generation AI on the format in which it should be output. By having the AI learn a standard format for data analysis, it is possible to output data in a reproducible format.
[0059] Figure 7 shows an example of a standard instruction to the generation AI, "Data Analysis_Amount of Food," to display the current amount of food eaten by all care recipients on the care monitor 10C. Figure 8 shows an example of the amount of food actually eaten as of 11:00 a.m., as displayed on the care monitor 10C. The amount of food eaten by care recipients with the least amount of food is displayed on the care monitor 10C in order, so that care staff who notice this can respond quickly. Figure 9 shows an example of the amount of food eaten as of 8:00 p.m.
[0060] The instructions for the example of "Data Analysis - Moisture Amount" in Figure 10 and the results are shown in Figure 11. The care recipients with low moisture amounts are set to be placed on the top shelf.
[0061] The amount of food, water, blood pressure, SPO2, and special notes for each care recipient are specified in "Data Analysis - User Status List" in FIG. 12, and the result, FIG. 13, is displayed on the care monitor 10C.
[0062] The cause of abnormalities includes missed or incorrect records, and the purpose is to detect any abnormalities.
[0063] These nursing care monitor displays are automatically updated, allowing the current status of the care recipient to be grasped, making them a static system.
[0064] We have developed a flexible data analysis system that does not stick to a fixed format, in which a care staff member or a care facility manager speaks "data analysis" into the communication terminal device 10 and then indicates the information they want to know.
[0065] For example, if a care staff member wants to know the current status of the care recipient, "Higuchi-san," this can be achieved using the following two processes / methods.
[0066] First, if you want the user to report the situation according to the template in Figure 14, you will get the result shown in Figure 15. Second, if you give instructions freely without a template, you will get the result shown in Figure 16.
[0067] Since care staff or care facility managers are free to choose what information they want to receive, this can be described as a dynamic system for understanding the status of care services within the facility.
[0068] Data such as "case reports" and "nursing records" are transcribed by the generation AI engine G, processed by a morphological analysis engine such as MeCab, and registered in a database.
[0069] Figures 17 and 18 show examples of the viewing screen and the printing screen. Figure 19 shows an example of a graph created. The size of the circles indicates the difference in excretion amount.
[0070] In addition to the above, unintentional voice data contains the raw voices of those in the nursing care field, and contains important information that will lead to high-quality nursing care services. It also serves as a source of information for understanding the internal and external situations of the organization, and serves as monitoring data for the organization's overall management system. We have built a system that allows nursing care facility managers to select the information they want and analyze the data.
[0071] A developer PC / server 10D of the inventor or an individual or organization D commissioned by the inventor manages a cloud server E as shown in Figure 3. The developer PC / server 10D manages the security of the authentication system 15B, manages the storage 17, and manages various scripts and the website F.
[0072] The developer PC / server 10D of the present inventors or an individual or organization D commissioned by the present inventors executes the learning program of the generation AI engine G as shown in Figure 4. This allows highly specialized information on medical care, nursing care, legal regulations, etc. to be provided to the generation AI engine G, enabling it to be developed. The developer PC / server 10D also manages the communication terminal devices 10 and scripts carried by the nursing staff as shown in Figure 5.
[0073] Generally, data analysis and tabulation / drawing for nursing care services are achieved by registering text data in a database (DB) and then processing the data. This system inputs image data, including text data and video, into a generation AI engine G, and by giving appropriate instructions, automatically creates tables and drawings. These output documents are used as nursing care summaries for the daily operational management of nursing care services, periodic reviews, and input into nursing care plans such as care plans.
[0074] The inventors have developed a system that allows care staff or care facility managers to converse with the generation AI engine G by uttering the word "question" into the communication terminal device 10 and then specifying the information they want to know.
[0075] Figure 20 shows the response when the generation AI engine G is asked about the state of an early stage bedsore.
[0076] It can be said to be a dynamic system, as it is up to the care staff or care facility administrator to freely decide what information they want to obtain.In addition, because it is a voice-based conversational system with the generation AI engine G, it is possible to ask a variety of questions, such as specialized information on medical care and nursing care, and nursing care insurance.
[0077] Due to a staff shortage in nursing care facilities, foreign workers are employed. Records can now be created in the foreign worker's native language with just one instruction. Figure 20 shows an example in which the generation AI explains bedsores to an Indonesian nursing staff member while simultaneously translating the explanation into Indonesian. This can prevent errors in nursing care services and can also be used for education and training. Connecting a monitor (e.g., 3.2 to 5 inches) to the communication terminal device 10A and displaying the translated results (e.g., English or Indonesian) provided by the generation AI engine facilitates on-the-spot communication.
[0078] Figure 21 shows the results of a conversation between caregiver A and office staff B at a care site, recorded by communication terminal device 10A and transcribed by generation AI engine G. Even though microphone 11A has a microphone count of 1, the generation AI of generation AI engine G was able to distinguish between the two voices and accurately transcribe them. However, the generation AI recognized caregiver A as the mother and office staff B as the daughter, and mistakenly determined that the conversation was between mother and daughter. It is interesting that generation AI engine G made such a judgment, as office staff B is a foreign worker who speaks Japanese haltingly and has a child-like voice quality.
[0079] Subsequently, the generative AI learning program was used to train generative AI engine G on the voice of office worker B, and it came to recognize office worker B. The generative AI learning program of the present invention is the most effective and reproducible invention developed by the inventors.
[0080] The inventors developed this system at a nursing care facility (social welfare corporation) with a scale of 160 care recipients and 150 care staff. The inventors investigated the nursing care records (called face sheets at the facility) of the nursing care facility for 13 years and 2 months from October 30, 2008 to January 25, 2022, and summarized the results in the following items.
[0081] 1. Corpus text (food) 195 patterns (see Figure 22) 2. Corpus text (excretion) 232 patterns (see Figure 23) 3. Corpus text (bathing) 207 patterns (see Figure 24) 4. Corpus text (movement) 177 patterns (see Figure 25) 5. Corpus text (oral) 78 patterns (see Figure 26) 6. Corpus text (care plan) 144 patterns (see Figure 27) 7. Corpus text (medical terms) 231 terms (see Figure 28) 8. Corpus text (caregiving terms) 58 terms (see Figure 29)
[0082] The inventors used the above patterns 1 to 8 as the "correct answers" and had the generation AI engine G learn them in two ways via the developer PC / server D in FIG.
[0083] 1) Training as a text file 2) Learning with audio data For information on how to create the corpus, we referred to "Creating the JUST Corpus, Takamichi Shinnosuke, Assistant Professor, Graduate School of Information Science and Technology, The University of Tokyo, 2023."
[0084] The inventors trained the generation AI engine G via the developer PC / server D in FIG. 4 to learn the following: 1. Name and pronunciation of the care recipient 2. Names and pronunciations of care staff These need to be updated regularly.
[0085] The inventors have trained the generation AI engine G via the developer PC / server D in FIG. 4 to learn the following: 1. For utterances that the generative AI engine mistranscribed, the errata was trained in the “replacements.json” file (see Figure 30). 2. The glossary that needs to be kept was trained in the "keep_patterns.json" file (see Figure 31). 3. The key terms (magic words) that are selected when spoken and distinguished from other audio files are shown in the “magic_words.json” file (see Figure 32). These need to be updated regularly. [Explanation of symbols]
[0086] 1000 Nursing care information generation system A. Nursing staff B. Nursing staff C. Nursing care monitor D Developer PC / Server E Cloud Server F website G Generative AI Engine H Internet 10, 10A, 10B Communication terminal equipment 11, 11A, 11B Headset (microphone, speaker)
Claims
1. A care information generation system including at least one terminal device and a generation AI engine for generating care information from voices at a care site, The terminal device includes a microphone that converts voices at the care site into voice signals, determines the generation of voice activity data based on the voice signals, and transmits the voice activity data to the generation AI engine; The voice activity data received from the terminal device is input to the generation AI engine, and the generation AI engine generates the care information based on the voice activity data; A care information generating system in which the control unit of the terminal device recognizes a finger flicking sound generated by flicking the microphone with a finger and starts recording.
2. The care information generating system according to claim 1, A care information generating system, wherein the care information is displayed on a care monitor.
3. The care information generating system according to claim 1, A care information generation system in which the generative AI engine is trained using a corpus related to care and / or medical care.
4. The care information generating system according to claim 1, A care information generating system, wherein the control unit of the terminal device monitors voice energy using energy-based voice activity detection (VAD) and determines that the voice activity data has occurred when the voice energy threshold is within a predetermined range.
5. The care information generating system according to claim 4, the microphone of the terminal device is a wired microphone or a wireless microphone, A care information generating system, wherein the control unit of the terminal device is provided with a wired threshold of sound energy for the wired microphone and a wireless threshold of sound energy for the wireless microphone.
6. 6. The care information generating system according to claim 5, A care information generating system, wherein the wired threshold is 32000±5000 (db) and the wireless threshold is 25000±5000 (db).
7. The care information generating system according to claim 1, the finger snapping sound has a shorter duration than the audio signal and has greater sound energy than the audio signal;
8. The care information generating system according to claim 1, After recognizing the finger snapping sound, the control unit activates a Python script function to start recording.
9. The care information generating system according to claim 1, The terminal device notifies the caregiver of the care information by using a buzzer during the day and a vibrator at night.
10. The care information generating system according to claim 1, A care information generation system comprising a PC / server that provides learning data to the generation AI engine.
11. The care information generating system according to claim 1, The care information generation system includes a data processing server capable of communicating with the terminal device and the generation AI engine, A care information generation system in which the data processing server receives the voice activity data from the terminal device and provides the voice activity data to the generation AI engine.
12. A care information generation system including at least one terminal device and a generation AI engine for generating care information from voices at a care site, The terminal device includes a microphone that converts voices at the care site into voice signals, determines the generation of voice activity data based on the voice signals, and transmits the voice activity data to the generation AI engine; The voice activity data received from the terminal device is input to the generation AI engine, and the generation AI engine generates the care information based on the voice activity data; a control unit of the terminal device monitoring speech energy using energy-based voice activity detection (VAD) and determining that the speech activity data occurs when the speech energy threshold is within a predetermined range; the microphone of the terminal device is a wired microphone or a wireless microphone, A care information generating system, wherein the control unit of the terminal device is provided with a wired threshold of sound energy for the wired microphone and a wireless threshold of sound energy for the wireless microphone.
Citation Information
Patent Citations
Voice interval detection device and method
JP2015207002A
Care system
JP2022056253A
Program, information processing method, and information processing apparatus
JP2023059601A