Communication Assist System
The communication assist system objectively evaluates communication between children and adults using speech and gaze analysis, enhancing interaction quality through device and cloud-based processing.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-27
- Publication Date
- 2026-03-11
AI Technical Summary
Existing communication software lacks the ability to objectively evaluate the communication situation between children with developmental disabilities and their guardians, hindering smoother interactions.
A communication assist system that includes application software with integrated voice and image analysis units to objectively assess communication through speech and gaze analysis, and facial expression detection, utilizing both local device and cloud processing to reduce load.
Provides an objective evaluation of communication effectiveness, facilitating smoother interactions by offering actionable insights for improvement.
Smart Images

Figure 2026042605000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a communication assist system including application software for use by children who are users and adults, including their guardians. [Background technology]
[0002] Known application software for this type of communication is application software that can be used for digital therapy for language development in children, including those with developmental disorders such as autism, as disclosed in Patent Document 1 below, by the inventor of the present application.
[0003] This application software is intended for use by children, including those with developmental disabilities, and adults, including their guardians, and the application software displays objects of interest to children, including those with developmental disabilities, on the screen of a communication terminal device, allowing them to operate freely.
[0004] This allows the object of interest to move in response to live commentary by adults, including the parents, and words spoken by children, including children with developmental disabilities, and can be used by children, including children with developmental disabilities, and adults, including their parents. [Prior art documents] [Patent documents]
[0005] [Patent Document 1] Patent No. 7383758 Summary of the Invention [Problem to be solved by the invention]
[0006] The inventors of the present application have now come to the realization that such application software can objectively evaluate the communication situation not only between children with developmental disabilities but also between children in general and adults, including their guardians, and contribute to smoother communication.
[0007] In view of the above circumstances, an object of the present invention is to provide a communication assist system that objectively evaluates the communication situation between children and adults, including their guardians, and contributes to smoother communication. [Means for solving the problem]
[0008] The communication assist system of the first invention is a communication assist system including application software for use by a child user and an adult, including a guardian, in a device having a processor, a voice recording unit for recording a conversation between the child and the adult while using the application software for a predetermined period of time; an image recording unit that records captured images of the child and the adult while using the application software for the predetermined period of time; a voice analysis unit that analyzes the number of utterances made by the child in the conversation from the conversation recorded in the voice recording unit; an image analysis unit that analyzes the gaze time of the child in the captured image from the captured image recorded in the image recording unit; an analysis result output unit that outputs the analysis results of either one or both of the audio analysis unit and the image analysis unit; The present invention is characterized by comprising:
[0009] According to the communication assistance system of the first invention, either or both of the following analysis results are output from the conversation recorded in the audio recording unit: an audio analysis that analyzes at least the number of times the child speaks in the conversation, i.e., the content of the child's speech, such as the number of times the child speaks; and an image analysis that analyzes at least the gaze time of the child in the captured image, i.e., the gaze time of the child, from the captured image recorded in the image recording unit.
[0010] Furthermore, such output results can serve as an evaluation index that objectively assesses the communication situation between children and adults, including their guardians, and the output evaluation index can be used as a reference to facilitate communication.
[0011] In this way, the communication assist system of the first aspect of the invention can objectively evaluate the communication situation between a wide range of people, including children and adults, including their guardians, and contribute to smoother communication.
[0012] The communication assist system of the second invention is the communication assist system of the first invention, a facial expression detection unit that detects a smiling face of the child in a captured image of the child and the adult during the use of the application software for the predetermined time; The facial expression detection unit outputs the number of smiles detected in the predetermined time period and the captured image.
[0013] According to the communication assist system of the second invention, the facial expression detection unit detects facial expressions, including smiles, of children, and outputs the number of facial expressions, including at least the number of smiles, detected over a predetermined time period, and captured images. This makes it possible to objectively evaluate the communication situation between children and adults, including their guardians, based on the number of smiles and other facial expressions, and also to output captured images of actual smiles (and other facial expressions, if necessary), allowing users to get a sense of the practice of communication and contributing to smoother communication.
[0014] In this way, the communication assist system of the second invention can objectively and multilaterally evaluate the communication situation between children and adults, including their guardians, and contribute to smoother communication.
[0015] A communication assist system according to a third aspect of the present invention is the communication assist system according to the second aspect of the present invention, a cloud server connected to the device via a network; the voice recording unit, the image recording unit, and the facial expression detection unit are configured in the device by the application software, The audio analysis unit and the image analysis unit are configured in the cloud server.
[0016] According to the communication assistance system of the third invention, through distributed processing between a device executed by application software and a cloud server, an audio recording unit, an image recording unit, and a facial expression detection unit are configured on the device, and an audio analysis unit and an image analysis unit are configured on the cloud server.
[0017] In this way, according to the communication assist system of the third invention, it is possible to actually provide a communication assist system that reduces the processing load on devices and cloud servers through distributed processing, objectively evaluates the communication situation between children and adults, including their guardians, and contributes to smoother communication. [Brief explanation of the drawings]
[0018] [Figure 1] 1 is a system configuration diagram showing the overall configuration of a communication assist system according to an embodiment of the present invention. [Figure 2] FIG. 2 is an explanatory diagram showing a display example on the device of FIG. 1. [Figure 3] 2 is a flowchart showing the processing content of the voice analysis unit in FIG. 1; [Figure 4]3 is a flowchart showing the processing content of the image analysis unit in FIG. 1; DETAILED DESCRIPTION OF THE INVENTION
[0019] A communication assist system according to an embodiment of the present invention will be described below with reference to FIG.
[0020] As shown in Figure 1, the communication assist system is a system comprising a device 1 having a processor, a screen, a camera for capturing images of the user, and a cloud server 2 connected to device 1 via a network, and is provided with an audio recording unit 11, an image recording unit 12, and a facial expression detection unit 23 configured in device 1, and an audio analysis unit 21 and an image analysis unit 22 configured in cloud server 2.
[0021] Furthermore, the device 1 and the cloud server 2 are capable of sharing analysis results with an analysis result output device 30 of a linked facility 3 such as a hospital or nursery school that has been assigned a facility link code in advance.
[0022] First, the voice recording unit 11, the image recording unit 12, and the facial expression detection unit 23 configured in the device 1 will be described.
[0023] The device 1 may be any device that runs browser application software, such as a smartphone or tablet, and functions as each processing unit when the application software program is installed on the device 1.
[0024] The application software here is described in detail in the above-mentioned patent document (Japanese Patent No. 7383758) by the inventor of the present application, and therefore will not be described here, but it uses augmented reality (AR) or holography technology to generate a composite image in which an object of interest to a child, such as a child with a developmental disability, is superimposed on an image captured by the external camera (not shown) of the device 1, and the composite image is displayed on the screen, allowing the object to move freely. The application software also allows the object of interest to move in response to live commentary by adults, including parents, and to words spoken by the child.
[0025] The voice recording unit 11 records, via a microphone (not shown) of the device 1, a conversation between a child and an adult while using application software for a predetermined period of time (for example, 5 minutes).
[0026] The image recording unit 12 records, via the in-camera (not shown) of the device 1, images of children and adults captured while using the application software for a predetermined time (for example, 5 minutes).
[0027] The facial expression detection unit 23 uses sequential processing to detect the number of facial expressions, including smiles, of children in images captured of children and adults using application software for a predetermined period of time (e.g., 5 minutes), and outputs the number of smiles and other facial expressions detected in the predetermined period of time (e.g., 5 minutes) and the captured images at that time.
[0028] Next, the audio analysis unit 21 and the image analysis unit 22 configured in the cloud server 2 will be described.
[0029] The cloud server 2 is a device that functions as a so-called virtual machine (VM), and executes processing according to the device 1 and outputs the processing results to the device 1 and the analysis output device 30 of the affiliated facility 3.
[0030] The voice analysis unit 21 acquires the conversation recorded in the voice recording unit 11, and performs an analysis of the acquired conversation between a child and an adult, primarily including the number of times the child speaks in the conversation, i.e., an analysis of the content of the child's speech, such as the number of times the child speaks.
[0031] The image analysis unit 22 acquires the captured images recorded in the image recording unit 12, and performs gaze analysis, primarily on the gaze duration of the child, i.e., the gaze duration of the child, from the acquired captured images of the conversation between the child and the adult.
[0032] On the other hand, the analysis result output device 30 of the affiliated facility 3 is a personal computer, smartphone, tablet, etc., and is capable of outputting the analysis results by the cloud server 2 as well as the usage status of the application software of the device 1.
[0033] The above is the configuration of the communication assist system of this embodiment. In the above configuration, each of the processing units 11, 12, and 13 of the device 1 and each of the processing units 21 and 22 of the cloud server 2 is configured with hardware such as a CPU (Central Processing Unit), a ROM (Read Only Memory), and a RAM (Random Access Memory), and stores and holds programs for executing various processes described below in memory (not shown), and functions as a computing device (sequencer) for executing various processes by executing the programs.
[0034] Next, the details of the analysis process performed by the communication assist system will be described.
[0035] First, as a premise, the communication assistance system generates a composite image in which the child's object of interest is superimposed on an image taken by the external camera (not shown) of device 1 when the application software of device 1 is launched, and displays it on the screen and allows it to operate freely.
[0036] The conversation between the child and the adult for a predetermined time (5 minutes) while using the application software is recorded by the audio recording unit 11, and the video recording unit 12 records captured images of the child and the adult using the application software.
[0037] Furthermore, during a predetermined period of use (5 minutes), the facial expression detection unit 13 detects the child's smiles (using existing facial recognition technology, etc.) and temporarily records the number of smiles detected and the captured image at the time of detection.
[0038] As shown in Figure 2, the number of smiles temporarily recorded and the captured images of the smiles are displayed on the screen of device 1 when use ends, i.e., after a predetermined time (5 minutes) has elapsed, allowing the user (child or adult) to recognize the number of smiles and to select and save the smile they like best as a best shot.
[0039] Furthermore, on this screen, it is possible to self-evaluate the status of practice of the items (tips) that should be practiced as "tips for drawing." For example, from the items (tips) that should be practiced, such as "make eye contact with your child," "smile," "talk (words)," and "put your child's actions into words," an adult guardian can select one item (tips) in advance for each session, which will help raise awareness of that item (tips), and at the end of the session, the status of practice can be self-evaluated and the evaluation results can be accumulated.
[0040] In this embodiment, the case where the number of smiles is detected and the detected number of smiles and the captured image at that time are displayed has been described, but this is not limited to this, and facial expressions other than smiles may be detected and the number of those facial expressions and the captured image at that time may be displayed.
[0041] Under the assumption that application software is being used on device 1, as shown in Figure 3, the audio analysis unit 21 of cloud server 2 acquires the conversation, which is the recorded audio recorded by the audio recording unit 11 of device 1 during the use of the application for a predetermined period of time (STEP 211 / Figure 3).
[0042] Next, the voice analysis unit 21 separates the speakers of the conversation from the acquired recorded voice (STEP 222 / FIG. 3).
[0043] Note that various methods of speaker separation may be employed for the speaker separation process here. For example, when installing an application software program on device 1, user profiles for children and adults who are users of device 1 may be created, and when creating these user profiles, voice data (reading of fixed phrases) of children and adults may be input, and a neural network for speaker separation may be trained by machine learning (deep learning) using this voice data, and the outputs of multiple channels of the neural network may be compared with the correct voice.
[0044] Next, the voice analysis unit 21 executes the process in the next step 213 for children among the speaker-separated recorded voices, and executes the process in the next step 214 for adults.
[0045] First, for a child, the speech analysis unit 21 performs a linguistic analysis of the child's speech content from the speaker-separated recorded speech (STEP 213 / FIG. 3). Here, for example, the number of utterances and the number of speeches are analyzed, that is, the number of utterances and the number of speeches are counted.
[0046] Furthermore, for children, the voice analysis unit 21 analyzes the number of words in a sentence (any one to four word sentence) from the speaker-separated recorded voice, that is, counts the number of words contained in one sentence (one utterance) of the child.
[0047] Here, when counting the number of utterances and speeches, and the number of words and sentences, existing speech analysis technology may be adopted as appropriate, or machine learning (deep learning) may be used to count the number of utterances and speeches, and the number of words and sentences, by comparing them with the number of correct answers.
[0048] On the other hand, for adults, the speech analysis unit 21 performs a linguistic analysis of the speech content of the adults from the speaker-separated recorded speech (STEP 214 / FIG. 3). Here, for example, the number of speeches is analyzed, that is, the number of speeches in the conversation is counted.
[0049] Furthermore, for adults, the voice analysis unit 21 analyzes the content of calls from the speaker-separated recorded voices. The analysis of the call content may involve analyzing and outputting the content itself, or may involve classifying the call content into good and bad based on the analyzed content.
[0050] Here, when counting the number of calls and analyzing the content of the calls, existing voice analysis technology may be adopted as appropriate, or machine learning (deep learning) may be used to count the number of calls and compare the content of the calls with the number of correct answers.
[0051] Next, when the analysis of the conversations, which are the recorded voices of the children and adults, is completed, the voice analysis unit 21 updates and saves the analysis results (STEP 215 / FIG. 3).
[0052] The above is the details of the processing by the voice analysis unit 21.
[0053] Next, as shown in FIG. 4, the image analysis unit 22 of the cloud server 2 acquires the recorded images recorded by the image recording unit 12 of the device 1 when the application was used for a predetermined period of time (STEP 221 / FIG. 4).
[0054] Next, the image analysis unit 22 analyzes the gaze of the child in the acquired recorded image (STEP 222 / FIG. 4). Specifically, the gaze analysis employs existing gaze tracking technology such as eye tracking, and calculates the gaze time based on the results of the gaze analysis.
[0055] Then, when a series of analyses of the child's gaze is completed, the image analysis unit 22 updates and saves the analysis results (STEP 223 / Fig. 4). When the analysis results by the image analysis unit 22 are saved, the results of the analysis are updated and saved in the cloud server 2, including the detection result of the number of smiles (including the best shot of the smile, if necessary) by the facial expression detection unit 13 of the device 1.
[0056] The above is the details of the processing by the image analysis unit 22.
[0057] The analysis results of the voice and images updated and saved in the cloud server 2 in STEP 215 and STEP 223 are output to doctors, childcare workers, caregivers, etc. via analysis result output device 30 at affiliated facilities 3 such as hospitals, nurseries, homes, etc. that are accessible by the affiliated code. Therefore, particularly in cases where a child has autism, etc., it becomes possible to provide continuous and detailed, highly specialized medical care tailored to the child's symptoms even in rural areas where there are no specialists, thereby promoting social participation and reducing parental anxiety and burden.
[0058] In this way, according to the communication assistance system of this embodiment, audio analysis results are output by the audio analysis unit 21 analyzing the number of times the child speaks from the conversation recorded in the audio recording unit 11, and image analysis results are output by the image recording unit 22 analyzing the duration of the child's gaze from the captured image recorded in the image recording unit 12.
[0059] Furthermore, such output results can serve as an evaluation index that objectively assesses the communication situation between children and adults, including their guardians, and the output evaluation index can be used as a reference to facilitate communication.
[0060] In this embodiment, the case where the voice analysis results by the voice analysis unit 21 and the image analysis results by the image recording unit 22 are output to the analysis result output device 30 of the affiliated facility 3 has been described, but this is not limited to this, and the analysis results may be output to the device 1 as well as to the personal computer of a previously registered adult guardian.
[0061] In addition, in this embodiment, we have described the use of application software using the in-camera and out-camera and microphone of device 1, but this is not limited to this, and application software may also be used using an external camera or microphone. [Explanation of symbols]
[0062] 1...device, 2...cloud server, 3...associated facility, 11...audio recording unit, 12...image recording unit, 13...facial expression detection unit, 21...audio analysis unit, 22...image analysis unit, 30...analysis result output device.
Claims
1. A communication assist system including application software for use by a child user and an adult including a guardian, in a device having a processor, a voice recording unit for recording a conversation between the child and the adult while using the application software for a predetermined period of time; an image recording unit that records captured images of the child and the adult while using the application software for the predetermined period of time; a voice analysis unit that analyzes the number of utterances made by the child in the conversation from the conversation recorded in the voice recording unit; an image analysis unit that analyzes the gaze time of the child in the captured image from the captured image recorded in the image recording unit; an analysis result output unit that outputs the analysis results of either one or both of the audio analysis unit and the image analysis unit; A communication assist system comprising:
2. 2. The communication assist system according to claim 1, a facial expression detection unit that detects facial expressions, including a smile, of the child in captured images of the child and the adult during the use of the application software for the predetermined period of time; The communication assist system is characterized in that the facial expression detection unit outputs the number of smiles detected within the predetermined time period and the captured image.
3. 3. The communication assist system according to claim 2, a cloud server connected to the device via a network; the voice recording unit, the image recording unit, and the facial expression detection unit are configured in the device by the application software, A communication assist system characterized in that the voice analysis unit and the image analysis unit are configured in the cloud server.
Citation Information
Patent Citations
Utterance meter, method of measuring quantity of utterance, program and recording medium
JP2004275220A
Autism treatment support system, autism treatment support device, and program
JP2020151092A
Application Software
JP7383758B1