system
The system automates the creation of business records through video recording, facial recognition, and summarization, addressing the manual burden and enhancing efficiency.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-18
- Publication Date
- 2026-05-01
AI Technical Summary
Conventional methods require manual observation and input for creating business records, leading to a significant burden.
A system comprising a recording unit, capture unit, recognition unit, documentation unit, and summarization unit that automates the creation of business records by recording video, performing facial recognition, documenting situations, summarizing information, and creating work records.
The system automates the creation of business records, reducing the workload and improving operational efficiency by generating accurate work records with minimal human intervention.
Smart Images

Figure 2026072840000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In the conventional technology, there is a problem that manual observation and input are required in creating business records, resulting in a large burden.
[0005] The system according to the embodiment aims to automate the creation of business records and reduce the burden.
Means for Solving the Problems
[0006] The system according to this embodiment comprises a recording unit, a capture unit, a recognition unit, a documentation unit, a summarization unit, and a record creation unit. The recording unit records video. The capture unit captures the video recorded by the recording unit at regular intervals. The recognition unit performs face recognition based on the video captured by the capture unit. The documentation unit documents the situation based on the information recognized by the recognition unit. The summarization unit summarizes the information documented by the documentation unit. The record creation unit automatically creates a business record based on the information summarized by the summarization unit. [Effects of the Invention]
[0007] The system according to this embodiment can automate the creation of business records and reduce the burden. [Brief explanation of the drawing]
[0008] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Modes for carrying out the invention]
[0009] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.
[0010] First, let's explain the terminology used in the following explanation.
[0011] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), or TPU (Tensor Processing Unit).
[0012] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.
[0013] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0014] In the following embodiments, the numbered communication I / F (Interface) is an interface including a communication processor, an antenna, etc. The communication I / F manages communication between multiple computers. Examples of communication standards applicable to the communication I / F include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0015] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it may be only A, only B, or a combination of A and B. Also, in this specification, when expressing three or more matters connected by "and / or", the same concept as "A and / or B" is applied.
[0016] [First Embodiment] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0017] As shown in FIG. 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0018] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. Also, the database 24 and the communication I / F 26 are connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0019] The smart device 14 comprises a computer 36, a receiving device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The receiving device 38, output device 40, and camera 42 are also connected to the bus 52.
[0020] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, and accepts user input. The touch panel 38A accepts user input via touch by detecting contact with an object (e.g., a pen or finger). The microphone 38B accepts user input via voice by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 (see Figure 2) acquires the data indicating the user input.
[0021] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user by outputting the data in a form perceptible to the user (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0022] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0023] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0024] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0025] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0026] In the smart device 14, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The specific processing program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart device 14 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.
[0027] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device having the data generation model 58. The data processing device 12 may also be a server device or a terminal device owned by a user (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of processing by the data processing system 10 according to the first embodiment will be described.
[0028] (Example of form 1) The automated system according to an embodiment of the present invention is a system for reducing the burden of creating work records in places such as daycare centers and nursing homes. This system captures video recorded by a camera at set intervals, performs facial recognition on each person, and a generating AI documents the situation. Furthermore, once a day's worth of documents for each interval have been accumulated, the system summarizes them and automatically creates a day's worth of work records. With this system, the generating AI can take over all the tasks of "observation," "document creation," and "recording" that were previously performed by humans. For example, the system captures video recorded by a camera at set intervals and performs facial recognition on each person. In a daycare center, the camera records the children's behavior upon arrival and during playtime, capturing the footage at set intervals. Next, the generating AI analyzes the captured video and documents the situation. For example, the generating AI documents information such as what kind of play the children were doing and who they were playing with. Furthermore, once a day's worth of documents for each interval have been accumulated, the generating AI summarizes them and automatically creates a day's worth of work records. In a daycare center, the system summarizes the children's behavior from arrival to nap time and compiles it into a day's worth of work records. This work record only requires human review and supplementation as needed. This system significantly reduces the workload of childcare workers and caregivers. Furthermore, by utilizing the data generated by the AI as behavioral data in healthcare businesses with the patient's consent, it becomes possible to provide more accurate medical option recommendations. For example, the AI can suggest the optimal treatment method based on the patient's behavioral data. In this way, the automated system is a groundbreaking system that automates work record keeping, reducing workload and providing highly accurate medical support. As a result, the automated system can reduce the burden of creating work records in childcare centers, nursing homes, and other facilities, and improve operational efficiency.
[0029] The automated system according to this embodiment comprises a recording unit, a capture unit, a recognition unit, a documentation unit, a summarization unit, and a record creation unit. The recording unit records video. The recording unit records video for the purpose of creating business records, for example, in a nursery school or a nursing home. For example, in a nursery school, the recording unit can record the behavior of children when they arrive and the situation during playtime. The capture unit captures the video recorded by the recording unit at regular intervals. For example, the capture unit captures video at regular intervals. For example, in a nursery school, the capture unit can capture the behavior of children when they arrive and the situation during playtime at regular intervals. The recognition unit performs face recognition based on the video captured by the capture unit. For example, the recognition unit performs face recognition based on the captured video. For example, in a nursery school, the recognition unit can recognize the faces of children when they arrive. The documentation unit documents the situation based on the information recognized by the recognition unit. For example, the documentation unit documents the situation based on the recognized information. The documentation unit can, for example, in a nursery school, document information such as what kind of games the children were playing and who they were playing with. The summarization unit summarizes the information documented by the documentation unit. The summarization unit can, for example, summarize the documented information. In a nursery school, for example, the summarization unit can summarize the children's activities from arrival to nap time. The record creation unit automatically creates work records based on the information summarized by the summarization unit. The record creation unit can, for example, automatically create work records based on the summarized information. In a nursery school, for example, the record creation unit can automatically create a day's worth of work records. As a result, the automated system according to this embodiment can automate the creation of work records and reduce the workload.
[0030] The recording unit records video. For example, it records video for creating operational records in places like daycare centers and nursing homes. In a daycare center, for instance, it can record children's behavior upon arrival and their playtime. Specifically, the recording unit uses high-resolution cameras to clearly record a wide range of images. Cameras are selected according to the application, including fixed cameras and pan-tilt-zoom (PTZ) cameras. Fixed cameras constantly monitor a specific area, while PTZ cameras allow for remotely changing the viewpoint, making them suitable for scenes with movement. The recording unit streams video data in real time and saves it to a cloud server or local storage. The saved video data is managed with timestamps and metadata so that it can be used later for playback and analysis. Furthermore, the recording unit has functions to adjust video quality and frame rate, allowing for optimal settings based on network bandwidth and storage capacity. This enables the recording unit to accurately record various situations occurring in places like daycare centers and nursing homes, which can then be used for later analysis and reporting.
[0031] The capture unit captures video recorded by the recording unit at regular intervals. For example, in a nursery school, the capture unit can capture the children's behavior upon arrival and their playtime at regular intervals. Specifically, the capture unit extracts still images from the recorded video data at specific time intervals. For example, it can capture every minute or every five minutes to ensure that important moments are not missed. The captured still images are saved with a timestamp and used for later analysis and reporting. The capture unit can use image processing technology to remove noise and adjust contrast in order to generate high-resolution still images suitable for video analysis. Furthermore, the capture unit has a trigger function that automatically captures when a specific event or movement is detected, ensuring that important scenes are recorded. This allows the capture unit to efficiently extract necessary information from the recorded video and use it for later analysis and reporting.
[0032] The recognition unit performs face recognition based on the video captured by the capture unit. For example, in a nursery school, the recognition unit can recognize the faces of children arriving at the school. Specifically, the recognition unit implements an AI-based face recognition algorithm that detects and identifies people's faces from captured still images. The face recognition algorithm utilizes deep learning technology and achieves high-precision recognition by learning from a large amount of face image data. The recognition unit extracts facial feature points and identifies individual people by comparing them with a pre-registered face database. For example, in a nursery school, it can recognize children's faces and be used for attendance confirmation and safety management upon arrival. To improve the accuracy of face recognition, the recognition unit is equipped with correction functions for lighting conditions and face orientation, and exhibits stable recognition performance in various environments. As a result, the recognition unit can accurately identify people from captured video, which is useful for subsequent documentation and reporting.
[0033] The documentation unit documents the situation based on the information recognized by the recognition unit. For example, in a nursery school, the documentation unit can document information such as what kind of play the children were doing and who they were playing with. Specifically, the documentation unit generates text-based reports using natural language processing (NLP) technology based on facial recognition results and other data provided by the recognition unit. The documentation unit uses AI to automatically organize the recognized information and convert it into easy-to-understand text. For example, in a nursery school, it can record the children's play and activities in detail and provide this information to parents and staff. The documentation unit stores the generated documents in cloud storage or on local servers and manages them so that they can be searched and viewed as needed. Furthermore, the documentation unit has the ability to customize the format and content of documents, allowing it to create reports tailored to specific needs. This enables the documentation unit to efficiently document recognized information and support the creation of work records.
[0034] The summarization unit summarizes information documented by the documentation unit. For example, the summarization unit can summarize documented information. For instance, in a daycare center, the summarization unit can summarize a child's activities from arrival to nap time. Specifically, the summarization unit analyzes documented information, extracts key points and keywords, and generates a concise summary. The summarization unit utilizes AI-based natural language processing technology to extract necessary information from lengthy documents and create summaries quickly. For example, in a daycare center, it can concisely summarize a child's daily activities and provide them to parents. To improve the accuracy of summaries, the summarization unit has the ability to understand the content and context of documents, ensuring that important information is not omitted. Furthermore, the summarization unit has the ability to customize the format and length of summaries, providing summaries tailored to specific needs. This allows the summarization unit to efficiently summarize documented information and support the creation of work records.
[0035] The record creation unit automatically creates work records based on information summarized by the summarization unit. For example, the record creation unit can automatically create work records based on summarized information. In a nursery school, for example, the record creation unit can automatically create a day's worth of work records. Specifically, the record creation unit generates work records according to a standard format based on the summarized information provided by the summarization unit. The record creation unit uses AI to convert the summarized information into an appropriate format and automatically fills in the necessary items. For example, in a nursery school, it can automatically create work records including the children's daily activities, attendance, and special notes, and provide them to parents and staff. The record creation unit saves the generated work records to cloud storage or a local server and manages them so that they can be searched and viewed as needed. Furthermore, the record creation unit has a function to customize the format and content of work records, so that records can be created to meet specific needs. As a result, the record creation unit can efficiently create work records based on summarized information and reduce its workload.
[0036] The record creation unit includes a data utilization unit for using the data created by the generation AI in healthcare operations. The data utilization unit, for example, utilizes the data created by the generation AI in healthcare operations. For example, the data utilization unit can use AI to suggest optimal treatment methods based on patient behavior data. For example, the data utilization unit can utilize data collected at daycare centers and nursing homes in healthcare operations. This improves the accuracy of medical treatment options by utilizing the data created by the generation AI in healthcare operations.
[0037] The recording unit analyzes ambient and background sounds during recording and highlights important audio events. For example, the recording unit can detect and highlight children's laughter or crying during recording. For example, the recording unit can detect and highlight emergency call sounds in a nursing home during recording. For example, the recording unit can detect and highlight specific conversations or instructions during recording. By highlighting important audio events, important information is not missed. Some or all of the above processing in the recording unit may be performed using AI, for example, or without AI.
[0038] The recording unit detects specific actions or behaviors during recording and automatically highlights those parts. For example, the recording unit detects a scene of a child playing during recording and highlights that part. For example, the recording unit detects a scene of a caregiver performing a specific type of care during recording and highlights that part. For example, the recording unit detects a specific event (for example, mealtime) during recording and highlights that part. This allows important scenes to be emphasized by highlighting specific actions or behaviors. Some or all of the above processing in the recording unit may be performed using AI, for example, or without AI.
[0039] The recording unit synchronizes multiple cameras during recording to simultaneously record footage from different angles. For example, the recording unit can record playtime at a nursery school with multiple cameras, simultaneously recording footage from different angles. For example, the recording unit can record care activities at a nursing home with multiple cameras, simultaneously recording footage from different angles. For example, the recording unit can record a specific event (e.g., a sports day) with multiple cameras, simultaneously recording footage from different angles. This allows for a more multifaceted perspective by simultaneously recording footage from multiple angles. Some or all of the above-described processes in the recording unit may be performed using AI, for example, or without AI.
[0040] The recording unit tracks the movements of a specific person during recording and focuses the recording on those movements. For example, the recording unit tracks the movements of a specific child in a nursery school and focuses the recording on those movements. For example, the recording unit tracks the movements of a specific user in a nursing home and focuses the recording on those movements. For example, the recording unit tracks the movements of a specific person at a specific event (e.g., a presentation) and focuses the recording on those movements. This ensures that important scenes are recorded without being missed by tracking the movements of a specific person. Some or all of the above processing in the recording unit may be performed using AI, for example, or without using AI.
[0041] The capture unit automatically adjusts the brightness and contrast of the video during capture to obtain the optimal image. For example, the capture unit automatically adjusts the brightness and contrast to obtain the optimal image in a scene of indoor play at a nursery school. For example, the capture unit automatically adjusts the brightness and contrast to obtain the optimal image in a scene of care at a nursing home. For example, the capture unit automatically adjusts the brightness and contrast to obtain the optimal image in a specific event (for example, a sports day). In this way, the optimal image can be obtained by automatically adjusting the brightness and contrast of the video. Some or all of the above processing in the capture unit may be performed using AI, for example, or without using AI.
[0042] The capture unit detects specific events or actions during capture and prioritizes capturing those moments. For example, the capture unit detects a specific event (e.g., a birthday party) during playtime at a nursery school and prioritizes capturing that moment. For example, the capture unit detects a specific action (e.g., rehabilitation) during care time at a nursing home and prioritizes capturing that moment. For example, the capture unit detects a specific action (e.g., a speech) at a specific event (e.g., a recital) and prioritizes capturing that moment. By prioritizing the capture of specific events or actions, important moments can be recorded without being missed. Some or all of the above processing in the capture unit may be performed using AI, for example, or without using AI.
[0043] The capture unit improves visibility by applying color tones and filters to the video during capture. For example, the capture unit improves visibility by applying color tones and filters to indoor play scenes in a nursery school. For example, the capture unit improves visibility by applying color tones and filters to care scenes in a nursing home. For example, the capture unit improves visibility by applying color tones and filters to specific events (for example, a sports day). In this way, visibility can be improved by applying color tones and filters to the video. Some or all of the above processing in the capture unit may be performed using AI, for example, or without using AI.
[0044] The capture unit zooms in on a specific area during capture to obtain a detailed image. For example, the capture unit zooms in on a specific child's face during playtime at a nursery school to obtain a detailed image. For example, the capture unit zooms in on a specific user's movements during care time at a nursing home to obtain a detailed image. For example, the capture unit zooms in on a specific person's facial expression at a specific event (e.g., a presentation) to obtain a detailed image. In this way, a detailed image can be obtained by zooming in on a specific area. Some or all of the above processing in the capture unit may be performed using AI, for example, or without using AI.
[0045] The recognition unit analyzes facial expressions and movements during recognition to estimate emotions and actions. For example, the recognition unit analyzes a child's facial expressions during playtime at a nursery school to estimate their emotions. For example, the recognition unit analyzes a user's movements during care time at a nursing home to estimate their actions. For example, the recognition unit analyzes a specific person's facial expressions at a specific event (e.g., a recital) to estimate their emotions. In this way, emotions and actions can be estimated by analyzing facial expressions and movements. Some or all of the above processing in the recognition unit may be performed using AI, for example, or without using AI.
[0046] The recognition unit recognizes multiple people simultaneously during recognition and tracks their actions. For example, the recognition unit recognizes multiple children simultaneously during playtime at a nursery school and tracks their actions. For example, the recognition unit recognizes multiple users simultaneously during care time at a nursing home and tracks their actions. For example, the recognition unit recognizes multiple people simultaneously at a specific event (e.g., a presentation) and tracks their actions. This allows for tracking the actions of multiple people by recognizing them simultaneously. Some or all of the above processing in the recognition unit may be performed using AI, for example, or without AI.
[0047] The recognition unit identifies individuals using features other than faces during recognition. For example, the recognition unit identifies individuals using children's clothing during playtime at a nursery school. For example, the recognition unit identifies individuals using users' belongings during care time at a nursing home. For example, the recognition unit identifies individuals using the clothing of specific individuals at specific events (e.g., recitals). This allows for more accurate identification of individuals by using features other than faces. Some or all of the above-described processes in the recognition unit may be performed using AI, for example, or without AI.
[0048] The recognition unit improves recognition accuracy by referring to past recognition data during recognition. For example, the recognition unit improves recognition accuracy by referring to past recognition data during playtime at a nursery school. For example, the recognition unit improves recognition accuracy by referring to past recognition data during care time at a nursing home. For example, the recognition unit improves recognition accuracy by referring to past recognition data at a specific event (for example, a recital). In this way, recognition accuracy can be improved by referring to past recognition data. Some or all of the above processing in the recognition unit may be performed using AI, for example, or without using AI.
[0049] The documentation unit analyzes the audio data of the video during the documentation process and reflects the audio information in the document. For example, the documentation unit analyzes the conversations of children during playtime at a nursery school and reflects them in the document. For example, the documentation unit analyzes the conversations of users during care time at a nursing home and reflects them in the document. For example, the documentation unit analyzes speeches at a specific event (e.g., a presentation) and reflects them in the document. By reflecting audio information in the document, more detailed records can be created. Some or all of the above processing in the documentation unit may be performed using AI, for example, or without using AI.
[0050] The documentation unit emphasizes specific keywords or phrases during the documentation process. For example, the documentation unit emphasizes specific keywords (e.g., types of play) during playtime at a nursery school. For example, the documentation unit emphasizes specific phrases (e.g., types of care) during care time at a nursing home. For example, the documentation unit emphasizes specific keywords (e.g., themes) at a specific event (e.g., a presentation). This allows important information to be highlighted by emphasizing specific keywords or phrases. Some or all of the above processing in the documentation unit may be performed using AI, for example, or without AI.
[0051] The documentation unit adds time information from the video during the documentation process to create a chronological document. For example, the documentation unit adds time information from the video to record playtime at a nursery school to create a chronological document. For example, the documentation unit adds time information from the video to record care time at a nursing home to create a chronological document. For example, the documentation unit adds time information from the video to record a specific event (for example, a presentation) to create a chronological document. In this way, by adding time information from the video, a chronological document can be created. Some or all of the above processing in the documentation unit may be performed using AI, for example, or without using AI.
[0052] The documentation unit integrates multiple video sources to create a single document during the documentation process. For example, the documentation unit integrates video footage from multiple cameras during playtime at a nursery school to create a single document. For example, the documentation unit integrates video footage from multiple cameras during care time at a nursing home to create a single document. For example, the documentation unit integrates video footage from multiple cameras at a specific event (e.g., a presentation) to create a single document. This allows for the creation of more comprehensive documents by integrating multiple video sources. Some or all of the above-described processes in the documentation unit may be performed using AI, for example, or without AI.
[0053] The summarization unit prioritizes summarizing important events and actions. For example, it might prioritize summarizing important events during playtime at a nursery school (e.g., birthday parties). For example, it might prioritize summarizing important actions during care time at a nursing home (e.g., rehabilitation). For example, it might prioritize summarizing important actions (e.g., speeches) at a specific event (e.g., a presentation). This ensures that important information is not overlooked by prioritizing the summarization of important events and actions. Some or all of the processing described above in the summarization unit may be performed using AI, for example, or not.
[0054] The summarization unit improves the accuracy of summarization by referring to past summarization data during the summarization process. For example, the summarization unit improves the accuracy of summarization by referring to past summarization data during playtime at a nursery school. For example, the summarization unit improves the accuracy of summarization by referring to past summarization data during care time at a nursing home. For example, the summarization unit improves the accuracy of summarization by referring to past summarization data at a specific event (e.g., a presentation). In this way, the accuracy of summarization can be improved by referring to past summarization data. Some or all of the above processing in the summarization unit may be performed using AI, for example, or without using AI.
[0055] The summarization unit incorporates video metadata (e.g., filming location and time) during the summarization process. For example, the summarization unit creates a summary by incorporating video metadata for playtime at a nursery school. For example, the summarization unit creates a summary by incorporating video metadata for care time at a nursing home. For example, the summarization unit creates a summary by incorporating video metadata for a specific event (e.g., a presentation). This allows for the creation of more detailed summaries by incorporating video metadata. Some or all of the above processing in the summarization unit may be performed using AI, for example, or without AI.
[0056] The summarization unit combines multiple summaries to create a single summary. For example, the summarization unit combines multiple summaries during playtime at a nursery school to create a single summary. For example, the summarization unit combines multiple summaries during care time at a nursing home to create a single summary. For example, the summarization unit combines multiple summaries at a specific event (e.g., a presentation) to create a single summary. This allows for the creation of a more comprehensive summary by combining multiple summaries. Some or all of the above processing in the summarization unit may be performed using AI, for example, or without AI.
[0057] The record-creation unit improves the accuracy of records by referring to past record data when creating records. For example, the record-creation unit improves the accuracy of records by referring to past record data during playtime at a nursery school. For example, the record-creation unit improves the accuracy of records by referring to past record data during care time at a nursing home. For example, the record-creation unit improves the accuracy of records by referring to past record data at a specific event (for example, a presentation). In this way, the accuracy of records can be improved by referring to past record data. Some or all of the above processing in the record-creation unit may be performed using AI, for example, or without using AI.
[0058] The recording unit highlights specific events or actions when creating records. For example, the recording unit highlights specific events (e.g., birthday parties) during playtime at a nursery school. For example, the recording unit highlights specific actions (e.g., rehabilitation) during care time at a nursing home. For example, the recording unit highlights specific actions (e.g., speeches) at a specific event (e.g., a presentation). By highlighting specific events or actions, important information is not overlooked. Some or all of the above processing in the recording unit may be performed using AI, for example, or without AI.
[0059] The recording unit adds time information to the video footage during recording to create a chronological record. For example, the recording unit adds time information to the video footage during playtime at a nursery school to create a chronological record. For example, the recording unit adds time information to the video footage during care time at a nursing home to create a chronological record. For example, the recording unit adds time information to the video footage during a specific event (for example, a recital) to create a chronological record. In this way, by adding time information to the video footage, a chronological record can be created. Some or all of the above processing in the recording unit may be performed using AI, for example, or without using AI.
[0060] The record creation unit integrates multiple records to create a single record when creating a record. For example, the record creation unit integrates multiple records of playtime at a nursery school to create a single record. For example, the record creation unit integrates multiple records of care time at a nursing home to create a single record. For example, the record creation unit integrates multiple records of a specific event (for example, a presentation) to create a single record. In this way, a more comprehensive record can be created by integrating multiple records. Some or all of the above processing in the record creation unit may be performed using AI, for example, or without using AI.
[0061] The data utilization unit improves the accuracy of data utilization by referring to past data when utilizing data. For example, the data utilization unit improves the accuracy of data utilization by referring to past data during playtime at a nursery school. For example, the data utilization unit improves the accuracy of data utilization by referring to past data during care time at a nursing home. For example, the data utilization unit improves the accuracy of data utilization by referring to past data at a specific event (for example, a presentation). In this way, the accuracy of data utilization can be improved by referring to past data. Some or all of the above processing in the data utilization unit may be performed using AI, for example, or without using AI.
[0062] The data utilization unit filters data according to specific purposes when utilizing it. For example, the data utilization unit filters data during playtime at a nursery school according to specific purposes (e.g., analyzing children's behavior). For example, the data utilization unit filters data during care time at a nursing home according to specific purposes (e.g., analyzing users' health status). For example, the data utilization unit filters data at a specific event (e.g., a presentation) according to specific purposes (e.g., analyzing participants' performance). By filtering data according to specific purposes, more appropriate data utilization becomes possible. Some or all of the above processing in the data utilization unit may be performed using AI, for example, or without using AI.
[0063] The Data Utilization Department incorporates metadata (e.g., location and time of acquisition) into the data when it is used. For example, the Data Utilization Department might incorporate metadata into data during playtime at a nursery school. For example, the Data Utilization Department might incorporate metadata into data during care time at a nursing home. For example, the Data Utilization Department might incorporate metadata into data for a specific event (e.g., a presentation). By incorporating metadata into the data, more detailed data utilization becomes possible. Some or all of the above-described processes in the Data Utilization Department may be performed using AI, for example, or without AI.
[0064] The Data Utilization Department integrates multiple data sources to create a single dataset when utilizing data. For example, the Data Utilization Department integrates multiple data sources to create a single dataset during playtime at a nursery school. For example, the Data Utilization Department integrates multiple data sources to create a single dataset during care time at a nursing home. For example, the Data Utilization Department integrates multiple data sources to create a single dataset at a specific event (e.g., a presentation). In this way, a more comprehensive dataset can be created by integrating multiple data sources. Some or all of the above processing in the Data Utilization Department may be performed using AI, for example, or without using AI.
[0065] The Data Utilization Department visualizes data during data utilization and presents it to users in an easy-to-understand manner. For example, the Data Utilization Department visualizes data during playtime at a nursery school to clearly show children's behavioral patterns. For example, the Data Utilization Department visualizes data during care time at a nursing home to clearly show users' health status. For example, the Data Utilization Department visualizes data at a specific event (e.g., a presentation) to clearly show participants' performance. In this way, data visualization makes it possible to present data to users in an easy-to-understand manner. Some or all of the above processing in the Data Utilization Department may be performed using AI, for example, or without using AI.
[0066] The Data Utilization Department ensures data security and protects privacy when utilizing data. For example, the Data Utilization Department ensures data security and protects children's privacy during playtime at a nursery school. For example, the Data Utilization Department ensures data security and protects user privacy during care time at a nursing home. For example, the Data Utilization Department ensures data security and protects participants' privacy at specific events (e.g., presentations). This ensures data security and protects privacy, allowing for safe data utilization. Some or all of the above-described processes in the Data Utilization Department may be performed using AI, for example, or without AI.
[0067] The Data Utilization Department backs up data when it is used to prevent data loss. For example, the Data Utilization Department backs up data during playtime at a nursery school to prevent data loss. For example, the Data Utilization Department backs up data during care time at a nursing home to prevent data loss. For example, the Data Utilization Department backs up data at a specific event (for example, a presentation) to prevent data loss. In this way, data loss can be prevented by backing up data. Some or all of the above processes in the Data Utilization Department may be performed using AI, for example, or without using AI.
[0068] The Data Utilization Department outputs the results of data analysis in report format when utilizing data. For example, the Data Utilization Department outputs the results of data analysis in report format during playtime at a nursery school, reporting on the children's behavioral patterns. For example, the Data Utilization Department outputs the results of data analysis in report format during care time at a nursing home, reporting on the users' health status. For example, the Data Utilization Department outputs the results of data analysis in report format at a specific event (e.g., a presentation), reporting on the performance of the participants. By outputting the data analysis results in report format, the analysis results can be presented in an easy-to-understand manner. Some or all of the above processing in the Data Utilization Department may be performed using AI, for example, or without using AI.
[0069] The Data Utilization Department analyzes data trends and makes future predictions when utilizing data. For example, the Data Utilization Department analyzes data trends during playtime at a nursery school to predict future behavioral patterns of children. For example, the Data Utilization Department analyzes data trends during care time at a nursing home to predict future health conditions of users. For example, the Data Utilization Department analyzes data trends at a specific event (e.g., a presentation) to predict future performance of participants. By analyzing data trends and making future predictions, it is possible to understand future trends. Some or all of the above processing in the Data Utilization Department may be performed using AI, for example, or without AI.
[0070] The data utilization unit detects anomalies in the data during data utilization and issues alerts. For example, the data utilization unit can detect anomalies in the data during playtime at a nursery school and issue an alert if there are abnormalities in the children's behavior. For example, the data utilization unit can detect anomalies in the data during care time at a nursing home and issue an alert if there are abnormalities in the user's health condition. For example, the data utilization unit can detect anomalies in the data at a specific event (e.g., a presentation) and issue an alert if there are abnormalities in the participants' performance. In this way, anomalies can be detected early by detecting anomalies in the data and issuing alerts. Some or all of the above processing in the data utilization unit may be performed using AI, for example, or without using AI.
[0071] The Data Utilization Department visualizes data during data utilization and presents it to users in an easy-to-understand manner. For example, the Data Utilization Department visualizes data during playtime at a nursery school to clearly show children's behavioral patterns. For example, the Data Utilization Department visualizes data during care time at a nursing home to clearly show users' health status. For example, the Data Utilization Department visualizes data at a specific event (e.g., a presentation) to clearly show participants' performance. In this way, data visualization makes it possible to present data to users in an easy-to-understand manner. Some or all of the above processing in the Data Utilization Department may be performed using AI, for example, or without using AI.
[0072] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.
[0073] The recording unit can detect specific environmental conditions (e.g., lighting and temperature) during recording and automatically adjust the recording settings based on those conditions. For example, if the lighting is dim during playtime at a nursery school, the recording unit can adjust the recording settings to compensate for the brightness. For example, if the room temperature is high during care time at a nursing home, the recording unit can adjust the recording settings to record video suitable for the temperature. For example, if the weather is bad at a specific event (e.g., a sports day), the recording unit can adjust the recording settings to record optimal video. In this way, by automatically adjusting the recording settings according to specific environmental conditions, more appropriate video can be recorded.
[0074] The capture unit can detect specific colors or patterns during capture and highlight those areas during the capture process. For example, the capture unit can detect a specific color (e.g., a red toy) during playtime at a nursery school and highlight that area during the capture. For example, the capture unit can detect a specific pattern (e.g., checkered clothing) during care time at a nursing home and highlight that area during the capture. For example, the capture unit can detect specific colors or patterns at a specific event (e.g., a recital) and highlight that area during the capture. This allows important information to be recorded without being missed by highlighting specific colors or patterns.
[0075] The recognition unit can detect specific gestures or actions during recognition and trigger specific actions based on those actions. For example, the recognition unit can detect a child raising their hand during playtime at a nursery school and trigger a specific action (e.g., starting recording) based on that action. For example, the recognition unit can detect a user waving their hand during care time at a nursing home and trigger a specific action (e.g., issuing an alert) based on that action. For example, the recognition unit can detect specific gestures or actions at a specific event (e.g., a presentation) and trigger specific actions based on those actions. This allows important actions to be automatically triggered by detecting specific gestures or actions.
[0076] The documentation unit can emphasize specific keywords or phrases during the documentation process. For example, it can emphasize specific keywords (e.g., types of play) during playtime at a nursery school. For example, it can emphasize specific phrases (e.g., types of care) during care time at a nursing home. For example, it can emphasize specific keywords (e.g., themes) during a specific event (e.g., a presentation). This allows important information to be highlighted by emphasizing specific keywords or phrases.
[0077] The summarization section can prioritize summarizing important events and actions. For example, it can prioritize summarizing important events (e.g., birthday parties) during playtime at a nursery school. For example, it can prioritize summarizing important actions (e.g., rehabilitation) during care time at a nursing home. For example, it can prioritize summarizing important actions (e.g., speeches) at a specific event (e.g., a presentation). By prioritizing the summarization of important events and actions, important information is not overlooked.
[0078] The following briefly describes the processing flow for example form 1.
[0079] Step 1: The recording unit records video. For example, it can record video for creating operational records in places like daycare centers and nursing homes. In a daycare center, it can record the children's behavior when they arrive and their playtime. Step 2: The capture unit captures the video recorded by the recording unit at scheduled intervals. For example, in a nursery school, it can capture the children's behavior upon arrival and their playtime at scheduled intervals. Step 3: The recognition unit performs facial recognition based on the video captured by the capture unit. For example, in a nursery school, it can recognize the faces of children as they arrive. Step 4: The documentation unit documents the situation based on the information recognized by the recognition unit. For example, in a nursery school, information such as what kind of games the children were playing and who they were playing with can be documented. Step 5: The summarization section summarizes the information documented by the documentation section. For example, in a daycare center, this could summarize the children's activities from arrival time to nap time. Step 6: The record-keeping unit automatically creates work records based on the information summarized by the summarization unit. For example, a daycare center can automatically create a work record for one day.
[0080] (Example of form 2) The automated system according to an embodiment of the present invention is a system for reducing the burden of creating work records in places such as daycare centers and nursing homes. This system captures video recorded by a camera at set intervals, performs facial recognition on each person, and a generating AI documents the situation. Furthermore, once a day's worth of documents for each interval have been accumulated, the system summarizes them and automatically creates a day's worth of work records. With this system, the generating AI can take over all the tasks of "observation," "document creation," and "recording" that were previously performed by humans. For example, the system captures video recorded by a camera at set intervals and performs facial recognition on each person. In a daycare center, the camera records the children's behavior upon arrival and during playtime, capturing the footage at set intervals. Next, the generating AI analyzes the captured video and documents the situation. For example, the generating AI documents information such as what kind of play the children were doing and who they were playing with. Furthermore, once a day's worth of documents for each interval have been accumulated, the generating AI summarizes them and automatically creates a day's worth of work records. In a daycare center, the system summarizes the children's behavior from arrival to nap time and compiles it into a day's worth of work records. This work record only requires human review and supplementation as needed. This system significantly reduces the workload of childcare workers and caregivers. Furthermore, by utilizing the data generated by the AI as behavioral data in healthcare businesses with the patient's consent, it becomes possible to provide more accurate medical option recommendations. For example, the AI can suggest the optimal treatment method based on the patient's behavioral data. In this way, the automated system is a groundbreaking system that automates work record keeping, reducing workload and providing highly accurate medical support. As a result, the automated system can reduce the burden of creating work records in childcare centers, nursing homes, and other facilities, and improve operational efficiency.
[0081] The automated system according to this embodiment comprises a recording unit, a capture unit, a recognition unit, a documentation unit, a summarization unit, and a record creation unit. The recording unit records video. The recording unit records video for the purpose of creating business records, for example, in a nursery school or a nursing home. For example, in a nursery school, the recording unit can record the behavior of children when they arrive and the situation during playtime. The capture unit captures the video recorded by the recording unit at regular intervals. For example, the capture unit captures video at regular intervals. For example, in a nursery school, the capture unit can capture the behavior of children when they arrive and the situation during playtime at regular intervals. The recognition unit performs face recognition based on the video captured by the capture unit. For example, the recognition unit performs face recognition based on the captured video. For example, in a nursery school, the recognition unit can recognize the faces of children when they arrive. The documentation unit documents the situation based on the information recognized by the recognition unit. For example, the documentation unit documents the situation based on the recognized information. The documentation unit can, for example, in a nursery school, document information such as what kind of games the children were playing and who they were playing with. The summarization unit summarizes the information documented by the documentation unit. The summarization unit can, for example, summarize the documented information. In a nursery school, for example, the summarization unit can summarize the children's activities from arrival to nap time. The record creation unit automatically creates work records based on the information summarized by the summarization unit. The record creation unit can, for example, automatically create work records based on the summarized information. In a nursery school, for example, the record creation unit can automatically create a day's worth of work records. As a result, the automated system according to this embodiment can automate the creation of work records and reduce the workload.
[0082] The recording unit records video. For example, it records video for creating operational records in places like daycare centers and nursing homes. In a daycare center, for instance, it can record children's behavior upon arrival and their playtime. Specifically, the recording unit uses high-resolution cameras to clearly record a wide range of images. Cameras are selected according to the application, including fixed cameras and pan-tilt-zoom (PTZ) cameras. Fixed cameras constantly monitor a specific area, while PTZ cameras allow for remotely changing the viewpoint, making them suitable for scenes with movement. The recording unit streams video data in real time and saves it to a cloud server or local storage. The saved video data is managed with timestamps and metadata so that it can be used later for playback and analysis. Furthermore, the recording unit has functions to adjust video quality and frame rate, allowing for optimal settings based on network bandwidth and storage capacity. This enables the recording unit to accurately record various situations occurring in places like daycare centers and nursing homes, which can then be used for later analysis and reporting.
[0083] The capture unit captures video recorded by the recording unit at regular intervals. For example, in a nursery school, the capture unit can capture the children's behavior upon arrival and their playtime at regular intervals. Specifically, the capture unit extracts still images from the recorded video data at specific time intervals. For example, it can capture every minute or every five minutes to ensure that important moments are not missed. The captured still images are saved with a timestamp and used for later analysis and reporting. The capture unit can use image processing technology to remove noise and adjust contrast in order to generate high-resolution still images suitable for video analysis. Furthermore, the capture unit has a trigger function that automatically captures when a specific event or movement is detected, ensuring that important scenes are recorded. This allows the capture unit to efficiently extract necessary information from the recorded video and use it for later analysis and reporting.
[0084] The recognition unit performs face recognition based on the video captured by the capture unit. For example, in a nursery school, the recognition unit can recognize the faces of children arriving at the school. Specifically, the recognition unit implements an AI-based face recognition algorithm that detects and identifies people's faces from captured still images. The face recognition algorithm utilizes deep learning technology and achieves high-precision recognition by learning from a large amount of face image data. The recognition unit extracts facial feature points and identifies individual people by comparing them with a pre-registered face database. For example, in a nursery school, it can recognize children's faces and be used for attendance confirmation and safety management upon arrival. To improve the accuracy of face recognition, the recognition unit is equipped with correction functions for lighting conditions and face orientation, and exhibits stable recognition performance in various environments. As a result, the recognition unit can accurately identify people from captured video, which is useful for subsequent documentation and reporting.
[0085] The documentation unit documents the situation based on the information recognized by the recognition unit. For example, in a nursery school, the documentation unit can document information such as what kind of play the children were doing and who they were playing with. Specifically, the documentation unit generates text-based reports using natural language processing (NLP) technology based on facial recognition results and other data provided by the recognition unit. The documentation unit uses AI to automatically organize the recognized information and convert it into easy-to-understand text. For example, in a nursery school, it can record the children's play and activities in detail and provide this information to parents and staff. The documentation unit stores the generated documents in cloud storage or on local servers and manages them so that they can be searched and viewed as needed. Furthermore, the documentation unit has the ability to customize the format and content of documents, allowing it to create reports tailored to specific needs. This enables the documentation unit to efficiently document recognized information and support the creation of work records.
[0086] The summarization unit summarizes information documented by the documentation unit. For example, the summarization unit can summarize documented information. For instance, in a daycare center, the summarization unit can summarize a child's activities from arrival to nap time. Specifically, the summarization unit analyzes documented information, extracts key points and keywords, and generates a concise summary. The summarization unit utilizes AI-based natural language processing technology to extract necessary information from lengthy documents and create summaries quickly. For example, in a daycare center, it can concisely summarize a child's daily activities and provide them to parents. To improve the accuracy of summaries, the summarization unit has the ability to understand the content and context of documents, ensuring that important information is not omitted. Furthermore, the summarization unit has the ability to customize the format and length of summaries, providing summaries tailored to specific needs. This allows the summarization unit to efficiently summarize documented information and support the creation of work records.
[0087] The record creation unit automatically creates work records based on information summarized by the summarization unit. For example, the record creation unit can automatically create work records based on summarized information. In a nursery school, for example, the record creation unit can automatically create a day's worth of work records. Specifically, the record creation unit generates work records according to a standard format based on the summarized information provided by the summarization unit. The record creation unit uses AI to convert the summarized information into an appropriate format and automatically fills in the necessary items. For example, in a nursery school, it can automatically create work records including the children's daily activities, attendance, and special notes, and provide them to parents and staff. The record creation unit saves the generated work records to cloud storage or a local server and manages them so that they can be searched and viewed as needed. Furthermore, the record creation unit has a function to customize the format and content of work records, so that records can be created to meet specific needs. As a result, the record creation unit can efficiently create work records based on summarized information and reduce its workload.
[0088] The record creation unit includes a data utilization unit for using the data created by the generation AI in healthcare operations. The data utilization unit, for example, utilizes the data created by the generation AI in healthcare operations. For example, the data utilization unit can use AI to suggest optimal treatment methods based on patient behavior data. For example, the data utilization unit can utilize data collected at daycare centers and nursing homes in healthcare operations. This improves the accuracy of medical treatment options by utilizing the data created by the generation AI in healthcare operations.
[0089] The recording unit estimates the user's emotions and adjusts the recording start time based on the estimated emotions. For example, if the user is stressed, the recording unit delays the start of recording to give the user time to relax. For example, if the user is excited, the recording unit starts recording immediately to capture that moment. For example, if the user is tired, the recording unit adjusts the start of recording to allow for a break. In this way, by adjusting the recording start time according to the user's emotions, more appropriate footage can be recorded. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI.
[0090] The recording unit analyzes ambient and background sounds during recording and highlights important audio events. For example, the recording unit can detect and highlight children's laughter or crying during recording. For example, the recording unit can detect and highlight emergency call sounds in a nursing home during recording. For example, the recording unit can detect and highlight specific conversations or instructions during recording. By highlighting important audio events, important information is not missed. Some or all of the above processing in the recording unit may be performed using AI, for example, or without AI.
[0091] The recording unit detects specific actions or behaviors during recording and automatically highlights those parts. For example, the recording unit detects a scene of a child playing during recording and highlights that part. For example, the recording unit detects a scene of a caregiver performing a specific type of care during recording and highlights that part. For example, the recording unit detects a specific event (for example, mealtime) during recording and highlights that part. This allows important scenes to be emphasized by highlighting specific actions or behaviors. Some or all of the above processing in the recording unit may be performed using AI, for example, or without AI.
[0092] The recording unit estimates the user's emotions and adjusts the recording end time based on the estimated emotions. For example, if the user is relaxed, the recording unit delays the recording end time to capture the relaxed state. For example, if the user is in a hurry, the recording unit shortens the recording end time to efficiently end the recording. For example, if the user is excited, the recording unit adjusts the recording end time to capture the peak of excitement. In this way, by adjusting the recording end time according to the user's emotions, more appropriate footage can be recorded. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI.
[0093] The recording unit synchronizes multiple cameras during recording to simultaneously record footage from different angles. For example, the recording unit can record playtime at a nursery school with multiple cameras, simultaneously recording footage from different angles. For example, the recording unit can record care activities at a nursing home with multiple cameras, simultaneously recording footage from different angles. For example, the recording unit can record a specific event (e.g., a sports day) with multiple cameras, simultaneously recording footage from different angles. This allows for a more multifaceted perspective by simultaneously recording footage from multiple angles. Some or all of the above-described processes in the recording unit may be performed using AI, for example, or without AI.
[0094] The recording unit tracks the movements of a specific person during recording and focuses the recording on those movements. For example, the recording unit tracks the movements of a specific child in a nursery school and focuses the recording on those movements. For example, the recording unit tracks the movements of a specific user in a nursing home and focuses the recording on those movements. For example, the recording unit tracks the movements of a specific person at a specific event (e.g., a presentation) and focuses the recording on those movements. This ensures that important scenes are recorded without being missed by tracking the movements of a specific person. Some or all of the above processing in the recording unit may be performed using AI, for example, or without using AI.
[0095] The capture unit estimates the user's emotions and adjusts the capture frequency based on the estimated emotions. For example, if the user is relaxed, the capture unit lowers the capture frequency to record the relaxed state. For example, if the user is excited, the capture unit increases the capture frequency to capture moments of excitement. For example, if the user is tired, the capture unit adjusts the capture frequency to allow for rest time. In this way, by adjusting the capture frequency according to the user's emotions, more appropriate video can be recorded. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI.
[0096] The capture unit automatically adjusts the brightness and contrast of the video during capture to obtain the optimal image. For example, the capture unit automatically adjusts the brightness and contrast to obtain the optimal image in a scene of indoor play at a nursery school. For example, the capture unit automatically adjusts the brightness and contrast to obtain the optimal image in a scene of care at a nursing home. For example, the capture unit automatically adjusts the brightness and contrast to obtain the optimal image in a specific event (for example, a sports day). In this way, the optimal image can be obtained by automatically adjusting the brightness and contrast of the video. Some or all of the above processing in the capture unit may be performed using AI, for example, or without using AI.
[0097] The capture unit detects specific events or actions during capture and prioritizes capturing those moments. For example, the capture unit detects a specific event (e.g., a birthday party) during playtime at a nursery school and prioritizes capturing that moment. For example, the capture unit detects a specific action (e.g., rehabilitation) during care time at a nursing home and prioritizes capturing that moment. For example, the capture unit detects a specific action (e.g., a speech) at a specific event (e.g., a recital) and prioritizes capturing that moment. By prioritizing the capture of specific events or actions, important moments can be recorded without being missed. Some or all of the above processing in the capture unit may be performed using AI, for example, or without using AI.
[0098] The capture unit estimates the user's emotions and adjusts the capture resolution based on the estimated emotions. For example, if the user is relaxed, the capture unit captures at a low resolution to record the relaxed state. For example, if the user is excited, the capture unit captures at a high resolution to capture the moment of excitement. For example, if the user is tired, the capture unit adjusts the resolution to allow for rest time. In this way, by adjusting the capture resolution according to the user's emotions, more appropriate video can be recorded. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI.
[0099] The capture unit improves visibility by applying color tones and filters to the video during capture. For example, the capture unit improves visibility by applying color tones and filters to indoor play scenes in a nursery school. For example, the capture unit improves visibility by applying color tones and filters to care scenes in a nursing home. For example, the capture unit improves visibility by applying color tones and filters to specific events (for example, a sports day). In this way, visibility can be improved by applying color tones and filters to the video. Some or all of the above processing in the capture unit may be performed using AI, for example, or without using AI.
[0100] The capture unit zooms in on a specific area during capture to obtain a detailed image. For example, the capture unit zooms in on a specific child's face during playtime at a nursery school to obtain a detailed image. For example, the capture unit zooms in on a specific user's movements during care time at a nursing home to obtain a detailed image. For example, the capture unit zooms in on a specific person's facial expression at a specific event (e.g., a presentation) to obtain a detailed image. In this way, a detailed image can be obtained by zooming in on a specific area. Some or all of the above processing in the capture unit may be performed using AI, for example, or without using AI.
[0101] The recognition unit estimates the user's emotions and adjusts the accuracy of face recognition based on the estimated emotions. For example, if the user is relaxed, the recognition unit lowers the accuracy of face recognition to record the relaxed state. For example, if the user is excited, the recognition unit increases the accuracy of face recognition to capture moments of excitement. For example, if the user is tired, the recognition unit adjusts the accuracy of face recognition to ensure rest time. In this way, by adjusting the accuracy of face recognition according to the user's emotions, more appropriate recognition results can be obtained. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI.
[0102] The recognition unit analyzes facial expressions and movements during recognition to estimate emotions and actions. For example, the recognition unit analyzes a child's facial expressions during playtime at a nursery school to estimate their emotions. For example, the recognition unit analyzes a user's movements during care time at a nursing home to estimate their actions. For example, the recognition unit analyzes a specific person's facial expressions at a specific event (e.g., a recital) to estimate their emotions. In this way, emotions and actions can be estimated by analyzing facial expressions and movements. Some or all of the above processing in the recognition unit may be performed using AI, for example, or without using AI.
[0103] The recognition unit recognizes multiple people simultaneously during recognition and tracks their actions. For example, the recognition unit recognizes multiple children simultaneously during playtime at a nursery school and tracks their actions. For example, the recognition unit recognizes multiple users simultaneously during care time at a nursing home and tracks their actions. For example, the recognition unit recognizes multiple people simultaneously at a specific event (e.g., a presentation) and tracks their actions. This allows for tracking the actions of multiple people by recognizing them simultaneously. Some or all of the above processing in the recognition unit may be performed using AI, for example, or without AI.
[0104] The recognition unit estimates the user's emotions and adjusts the display method of the recognition results based on the estimated emotions. For example, if the user is relaxed, the recognition unit displays detailed recognition results. For example, if the user is excited, the recognition unit displays concise recognition results. For example, if the user is tired, the recognition unit displays highly visible recognition results. In this way, by adjusting the display method of recognition results according to the user's emotions, more appropriate information can be provided. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI.
[0105] The recognition unit identifies individuals using features other than faces during recognition. For example, the recognition unit identifies individuals using children's clothing during playtime at a nursery school. For example, the recognition unit identifies individuals using users' belongings during care time at a nursing home. For example, the recognition unit identifies individuals using the clothing of specific individuals at specific events (e.g., recitals). This allows for more accurate identification of individuals by using features other than faces. Some or all of the above-described processes in the recognition unit may be performed using AI, for example, or without AI.
[0106] The recognition unit improves recognition accuracy by referring to past recognition data during recognition. For example, the recognition unit improves recognition accuracy by referring to past recognition data during playtime at a nursery school. For example, the recognition unit improves recognition accuracy by referring to past recognition data during care time at a nursing home. For example, the recognition unit improves recognition accuracy by referring to past recognition data at a specific event (for example, a recital). In this way, recognition accuracy can be improved by referring to past recognition data. Some or all of the above processing in the recognition unit may be performed using AI, for example, or without using AI.
[0107] The documentation unit estimates the user's emotions and adjusts the document's expression based on the estimated emotions. For example, if the user is relaxed, the documentation unit creates a document using soft language. If the user is excited, the documentation unit creates a document using strong language. If the user is tired, the documentation unit creates a document using concise and easy-to-understand language. This allows for the creation of more appropriate documents by adjusting the document's expression according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI.
[0108] The documentation unit analyzes the audio data of the video during the documentation process and reflects the audio information in the document. For example, the documentation unit analyzes the conversations of children during playtime at a nursery school and reflects them in the document. For example, the documentation unit analyzes the conversations of users during care time at a nursing home and reflects them in the document. For example, the documentation unit analyzes speeches at a specific event (e.g., a presentation) and reflects them in the document. By reflecting audio information in the document, more detailed records can be created. Some or all of the above processing in the documentation unit may be performed using AI, for example, or without using AI.
[0109] The documentation unit emphasizes specific keywords or phrases during the documentation process. For example, the documentation unit emphasizes specific keywords (e.g., types of play) during playtime at a nursery school. For example, the documentation unit emphasizes specific phrases (e.g., types of care) during care time at a nursing home. For example, the documentation unit emphasizes specific keywords (e.g., themes) at a specific event (e.g., a presentation). This allows important information to be highlighted by emphasizing specific keywords or phrases. Some or all of the above processing in the documentation unit may be performed using AI, for example, or without AI.
[0110] The documentation unit estimates the user's emotions and adjusts the length of the document based on the estimated emotions. For example, if the user is relaxed, the documentation unit creates a detailed document. If the user is excited, the documentation unit creates a concise document. If the user is tired, the documentation unit creates a short, to-the-point document. This allows for the creation of more appropriate documents by adjusting the length of the document according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI.
[0111] The documentation unit adds time information from the video during the documentation process to create a chronological document. For example, the documentation unit adds time information from the video to record playtime at a nursery school to create a chronological document. For example, the documentation unit adds time information from the video to record care time at a nursing home to create a chronological document. For example, the documentation unit adds time information from the video to record a specific event (for example, a presentation) to create a chronological document. In this way, by adding time information from the video, a chronological document can be created. Some or all of the above processing in the documentation unit may be performed using AI, for example, or without using AI.
[0112] The documentation unit integrates multiple video sources to create a single document during the documentation process. For example, the documentation unit integrates video footage from multiple cameras during playtime at a nursery school to create a single document. For example, the documentation unit integrates video footage from multiple cameras during care time at a nursing home to create a single document. For example, the documentation unit integrates video footage from multiple cameras at a specific event (e.g., a presentation) to create a single document. This allows for the creation of more comprehensive documents by integrating multiple video sources. Some or all of the above-described processes in the documentation unit may be performed using AI, for example, or without AI.
[0113] The summarization unit estimates the user's emotions and adjusts the level of detail in the summary based on the estimated emotions. For example, if the user is relaxed, the summarization unit will create a detailed summary. If the user is excited, the summarization unit will create a concise summary. If the user is tired, the summarization unit will create a short, to-the-point summary. This allows for the creation of more appropriate summaries by adjusting the level of detail in the summary according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI.
[0114] The summarization unit prioritizes summarizing important events and actions. For example, it might prioritize summarizing important events during playtime at a nursery school (e.g., birthday parties). For example, it might prioritize summarizing important actions during care time at a nursing home (e.g., rehabilitation). For example, it might prioritize summarizing important actions (e.g., speeches) at a specific event (e.g., a presentation). This ensures that important information is not overlooked by prioritizing the summarization of important events and actions. Some or all of the processing described above in the summarization unit may be performed using AI, for example, or not.
[0115] The summarization unit improves the accuracy of summarization by referring to past summarization data during the summarization process. For example, the summarization unit improves the accuracy of summarization by referring to past summarization data during playtime at a nursery school. For example, the summarization unit improves the accuracy of summarization by referring to past summarization data during care time at a nursing home. For example, the summarization unit improves the accuracy of summarization by referring to past summarization data at a specific event (e.g., a presentation). In this way, the accuracy of summarization can be improved by referring to past summarization data. Some or all of the above processing in the summarization unit may be performed using AI, for example, or without using AI.
[0116] The summarization unit estimates the user's emotions and adjusts the order of the summary based on the estimated emotions. For example, if the user is relaxed, the summarization unit creates a detailed summary. If the user is excited, the summarization unit creates a concise summary. If the user is tired, the summarization unit creates a short, to-the-point summary. This allows for the creation of a more appropriate summary by adjusting the order of the summary according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI.
[0117] The summarization unit incorporates video metadata (e.g., filming location and time) during the summarization process. For example, the summarization unit creates a summary by incorporating video metadata for playtime at a nursery school. For example, the summarization unit creates a summary by incorporating video metadata for care time at a nursing home. For example, the summarization unit creates a summary by incorporating video metadata for a specific event (e.g., a presentation). This allows for the creation of more detailed summaries by incorporating video metadata. Some or all of the above processing in the summarization unit may be performed using AI, for example, or without AI.
[0118] The summarization unit combines multiple summaries to create a single summary. For example, the summarization unit combines multiple summaries during playtime at a nursery school to create a single summary. For example, the summarization unit combines multiple summaries during care time at a nursing home to create a single summary. For example, the summarization unit combines multiple summaries at a specific event (e.g., a presentation) to create a single summary. This allows for the creation of a more comprehensive summary by combining multiple summaries. Some or all of the above processing in the summarization unit may be performed using AI, for example, or without AI.
[0119] The recording unit estimates the user's emotions and adjusts the recording format based on the estimated emotions. For example, if the user is relaxed, the recording unit creates a detailed recording. If the user is excited, the recording unit creates a concise recording. If the user is tired, the recording unit creates a short, to-the-point recording. By adjusting the recording format according to the user's emotions, a more appropriate recording can be created. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI.
[0120] The record-creation unit improves the accuracy of records by referring to past record data when creating records. For example, the record-creation unit improves the accuracy of records by referring to past record data during playtime at a nursery school. For example, the record-creation unit improves the accuracy of records by referring to past record data during care time at a nursing home. For example, the record-creation unit improves the accuracy of records by referring to past record data at a specific event (for example, a presentation). In this way, the accuracy of records can be improved by referring to past record data. Some or all of the above processing in the record-creation unit may be performed using AI, for example, or without using AI.
[0121] The recording unit highlights specific events or actions when creating records. For example, the recording unit highlights specific events (e.g., birthday parties) during playtime at a nursery school. For example, the recording unit highlights specific actions (e.g., rehabilitation) during care time at a nursing home. For example, the recording unit highlights specific actions (e.g., speeches) at a specific event (e.g., a presentation). By highlighting specific events or actions, important information is not overlooked. Some or all of the above processing in the recording unit may be performed using AI, for example, or without AI.
[0122] The recording unit estimates the user's emotions and determines the priority of recordings based on the estimated emotions. For example, if the user is relaxed, the recording unit prioritizes detailed recordings. If the user is excited, for example, the recording unit prioritizes concise recordings. If the user is tired, for example, the recording unit prioritizes short, to-the-point recordings. This allows for the creation of more appropriate recordings by determining the priority of recordings according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI.
[0123] The recording unit adds time information to the video footage during recording to create a chronological record. For example, the recording unit adds time information to the video footage during playtime at a nursery school to create a chronological record. For example, the recording unit adds time information to the video footage during care time at a nursing home to create a chronological record. For example, the recording unit adds time information to the video footage during a specific event (for example, a recital) to create a chronological record. In this way, by adding time information to the video footage, a chronological record can be created. Some or all of the above processing in the recording unit may be performed using AI, for example, or without using AI.
[0124] The record creation unit integrates multiple records to create a single record when creating a record. For example, the record creation unit integrates multiple records of playtime at a nursery school to create a single record. For example, the record creation unit integrates multiple records of care time at a nursing home to create a single record. For example, the record creation unit integrates multiple records of a specific event (for example, a presentation) to create a single record. In this way, a more comprehensive record can be created by integrating multiple records. Some or all of the above processing in the record creation unit may be performed using AI, for example, or without using AI.
[0125] The data utilization unit estimates the user's emotions and adjusts how the data is used based on the estimated emotions. For example, if the user is relaxed, the data utilization unit performs a detailed data analysis. For example, if the user is excited, the data utilization unit performs a concise data analysis. For example, if the user is tired, the data utilization unit performs a short, to-the-point data analysis. By adjusting how the data is used according to the user's emotions, more appropriate data utilization can be achieved. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI.
[0126] The data utilization unit improves the accuracy of data utilization by referring to past data when utilizing data. For example, the data utilization unit improves the accuracy of data utilization by referring to past data during playtime at a nursery school. For example, the data utilization unit improves the accuracy of data utilization by referring to past data during care time at a nursing home. For example, the data utilization unit improves the accuracy of data utilization by referring to past data at a specific event (for example, a presentation). In this way, the accuracy of data utilization can be improved by referring to past data. Some or all of the above processing in the data utilization unit may be performed using AI, for example, or without using AI.
[0127] The data utilization unit filters data according to specific purposes when utilizing it. For example, the data utilization unit filters data during playtime at a nursery school according to specific purposes (e.g., analyzing children's behavior). For example, the data utilization unit filters data during care time at a nursing home according to specific purposes (e.g., analyzing users' health status). For example, the data utilization unit filters data at a specific event (e.g., a presentation) according to specific purposes (e.g., analyzing participants' performance). By filtering data according to specific purposes, more appropriate data utilization becomes possible. Some or all of the above processing in the data utilization unit may be performed using AI, for example, or without using AI.
[0128] The data utilization unit estimates the user's emotions and prioritizes data based on the estimated emotions. For example, if the user is relaxed, the data utilization unit prioritizes detailed data. If the user is excited, for example, the data utilization unit prioritizes concise data. If the user is tired, for example, the data utilization unit prioritizes short, to the point data. This allows for more appropriate data utilization by prioritizing data according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI.
[0129] The Data Utilization Department incorporates metadata (e.g., location and time of acquisition) into the data when it is used. For example, the Data Utilization Department might incorporate metadata into data during playtime at a nursery school. For example, the Data Utilization Department might incorporate metadata into data during care time at a nursing home. For example, the Data Utilization Department might incorporate metadata into data for a specific event (e.g., a presentation). By incorporating metadata into the data, more detailed data utilization becomes possible. Some or all of the above-described processes in the Data Utilization Department may be performed using AI, for example, or without AI.
[0130] The Data Utilization Department integrates multiple data sources to create a single dataset when utilizing data. For example, the Data Utilization Department integrates multiple data sources to create a single dataset during playtime at a nursery school. For example, the Data Utilization Department integrates multiple data sources to create a single dataset during care time at a nursing home. For example, the Data Utilization Department integrates multiple data sources to create a single dataset at a specific event (e.g., a presentation). In this way, a more comprehensive dataset can be created by integrating multiple data sources. Some or all of the above processing in the Data Utilization Department may be performed using AI, for example, or without using AI.
[0131] The Data Utilization Department visualizes data during data utilization and presents it to users in an easy-to-understand manner. For example, the Data Utilization Department visualizes data during playtime at a nursery school to clearly show children's behavioral patterns. For example, the Data Utilization Department visualizes data during care time at a nursing home to clearly show users' health status. For example, the Data Utilization Department visualizes data at a specific event (e.g., a presentation) to clearly show participants' performance. In this way, data visualization makes it possible to present data to users in an easy-to-understand manner. Some or all of the above processing in the Data Utilization Department may be performed using AI, for example, or without using AI.
[0132] The Data Utilization Department ensures data security and protects privacy when utilizing data. For example, the Data Utilization Department ensures data security and protects children's privacy during playtime at a nursery school. For example, the Data Utilization Department ensures data security and protects user privacy during care time at a nursing home. For example, the Data Utilization Department ensures data security and protects participants' privacy at specific events (e.g., presentations). This ensures data security and protects privacy, allowing for safe data utilization. Some or all of the above-described processes in the Data Utilization Department may be performed using AI, for example, or without AI.
[0133] The Data Utilization Department backs up data when it is used to prevent data loss. For example, the Data Utilization Department backs up data during playtime at a nursery school to prevent data loss. For example, the Data Utilization Department backs up data during care time at a nursing home to prevent data loss. For example, the Data Utilization Department backs up data at a specific event (for example, a presentation) to prevent data loss. In this way, data loss can be prevented by backing up data. Some or all of the above processes in the Data Utilization Department may be performed using AI, for example, or without using AI.
[0134] The Data Utilization Department outputs the results of data analysis in report format when utilizing data. For example, the Data Utilization Department outputs the results of data analysis in report format during playtime at a nursery school, reporting on the children's behavioral patterns. For example, the Data Utilization Department outputs the results of data analysis in report format during care time at a nursing home, reporting on the users' health status. For example, the Data Utilization Department outputs the results of data analysis in report format at a specific event (e.g., a presentation), reporting on the performance of the participants. By outputting the data analysis results in report format, the analysis results can be presented in an easy-to-understand manner. Some or all of the above processing in the Data Utilization Department may be performed using AI, for example, or without using AI.
[0135] The Data Utilization Department analyzes data trends and makes future predictions when utilizing data. For example, the Data Utilization Department analyzes data trends during playtime at a nursery school to predict future behavioral patterns of children. For example, the Data Utilization Department analyzes data trends during care time at a nursing home to predict future health conditions of users. For example, the Data Utilization Department analyzes data trends at a specific event (e.g., a presentation) to predict future performance of participants. By analyzing data trends and making future predictions, it is possible to understand future trends. Some or all of the above processing in the Data Utilization Department may be performed using AI, for example, or without AI.
[0136] The data utilization unit detects anomalies in the data during data utilization and issues alerts. For example, the data utilization unit can detect anomalies in the data during playtime at a nursery school and issue an alert if there are abnormalities in the children's behavior. For example, the data utilization unit can detect anomalies in the data during care time at a nursing home and issue an alert if there are abnormalities in the user's health condition. For example, the data utilization unit can detect anomalies in the data at a specific event (e.g., a presentation) and issue an alert if there are abnormalities in the participants' performance. In this way, anomalies can be detected early by detecting anomalies in the data and issuing alerts. Some or all of the above processing in the data utilization unit may be performed using AI, for example, or without using AI.
[0137] The Data Utilization Department visualizes data during data utilization and presents it to users in an easy-to-understand manner. For example, the Data Utilization Department visualizes data during playtime at a nursery school to clearly show children's behavioral patterns. For example, the Data Utilization Department visualizes data during care time at a nursing home to clearly show users' health status. For example, the Data Utilization Department visualizes data at a specific event (e.g., a presentation) to clearly show participants' performance. In this way, data visualization makes it possible to present data to users in an easy-to-understand manner. Some or all of the above processing in the Data Utilization Department may be performed using AI, for example, or without using AI.
[0138] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.
[0139] The recording unit can detect specific environmental conditions (e.g., lighting and temperature) during recording and automatically adjust the recording settings based on those conditions. For example, if the lighting is dim during playtime at a nursery school, the recording unit can adjust the recording settings to compensate for the brightness. For example, if the room temperature is high during care time at a nursing home, the recording unit can adjust the recording settings to record video suitable for the temperature. For example, if the weather is bad at a specific event (e.g., a sports day), the recording unit can adjust the recording settings to record optimal video. In this way, by automatically adjusting the recording settings according to specific environmental conditions, more appropriate video can be recorded.
[0140] The recording unit can estimate the user's emotions and adjust the recording frame rate based on the estimated emotions. For example, if the user is relaxed, the recording unit can record at a low frame rate to capture the relaxed state. For example, if the user is excited, the recording unit can record at a high frame rate to capture the moment of excitement. For example, if the user is tired, the recording unit can adjust the frame rate to allow for rest time. In this way, by adjusting the recording frame rate according to the user's emotions, more appropriate footage can be recorded. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI.
[0141] The capture unit can detect specific colors or patterns during capture and highlight those areas during the capture process. For example, the capture unit can detect a specific color (e.g., a red toy) during playtime at a nursery school and highlight that area during the capture. For example, the capture unit can detect a specific pattern (e.g., checkered clothing) during care time at a nursing home and highlight that area during the capture. For example, the capture unit can detect specific colors or patterns at a specific event (e.g., a recital) and highlight that area during the capture. This allows important information to be recorded without being missed by highlighting specific colors or patterns.
[0142] The capture unit can estimate the user's emotions and adjust the capture resolution based on the estimated emotions. For example, if the user is relaxed, the capture unit can capture at a low resolution to record the relaxed state. For example, if the user is excited, the capture unit can capture at a high resolution to capture the moment of excitement. For example, if the user is tired, the capture unit can adjust the resolution to allow for rest time. In this way, by adjusting the capture resolution according to the user's emotions, more appropriate video can be recorded. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI.
[0143] The recognition unit can detect specific gestures or actions during recognition and trigger specific actions based on those actions. For example, the recognition unit can detect a child raising their hand during playtime at a nursery school and trigger a specific action (e.g., starting recording) based on that action. For example, the recognition unit can detect a user waving their hand during care time at a nursing home and trigger a specific action (e.g., issuing an alert) based on that action. For example, the recognition unit can detect specific gestures or actions at a specific event (e.g., a presentation) and trigger specific actions based on those actions. This allows important actions to be automatically triggered by detecting specific gestures or actions.
[0144] The recognition unit can estimate the user's emotions and adjust the accuracy of face recognition based on the estimated emotions. For example, if the user is relaxed, the recognition unit can lower the accuracy of face recognition to record the relaxed state. For example, if the user is excited, the recognition unit can increase the accuracy of face recognition to capture the moment of excitement. For example, if the user is tired, the recognition unit can adjust the accuracy of face recognition to ensure a rest period. By adjusting the accuracy of face recognition according to the user's emotions, more appropriate recognition results can be obtained. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI.
[0145] The documentation unit can emphasize specific keywords or phrases during the documentation process. For example, it can emphasize specific keywords (e.g., types of play) during playtime at a nursery school. For example, it can emphasize specific phrases (e.g., types of care) during care time at a nursing home. For example, it can emphasize specific keywords (e.g., themes) during a specific event (e.g., a presentation). This allows important information to be highlighted by emphasizing specific keywords or phrases.
[0146] The documentation unit can estimate the user's emotions and adjust the document's expression based on those emotions. For example, if the user is relaxed, the documentation unit can create a document using soft language. If the user is excited, the documentation unit can create a document using strong language. If the user is tired, the documentation unit can create a document using concise and easy-to-understand language. By adjusting the document's expression according to the user's emotions, a more appropriate document can be created. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI.
[0147] The summarization section can prioritize summarizing important events and actions. For example, it can prioritize summarizing important events (e.g., birthday parties) during playtime at a nursery school. For example, it can prioritize summarizing important actions (e.g., rehabilitation) during care time at a nursing home. For example, it can prioritize summarizing important actions (e.g., speeches) at a specific event (e.g., a presentation). By prioritizing the summarization of important events and actions, important information is not overlooked.
[0148] The summarization unit can estimate the user's emotions and adjust the level of detail in the summary based on the estimated emotions. For example, if the user is relaxed, the summarization unit can create a detailed summary. If the user is excited, for example, the summarization unit can create a concise summary. If the user is tired, for example, the summarization unit can create a short, to-the-point summary. In this way, by adjusting the level of detail in the summary according to the user's emotions, a more appropriate summary can be created. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI.
[0149] The following briefly describes the processing flow for example form 2.
[0150] Step 1: The recording unit records video. For example, it can record video for creating operational records in places like daycare centers and nursing homes. In a daycare center, it can record the children's behavior when they arrive and their playtime. Step 2: The capture unit captures the video recorded by the recording unit at scheduled intervals. For example, in a nursery school, it can capture the children's behavior upon arrival and their playtime at scheduled intervals. Step 3: The recognition unit performs facial recognition based on the video captured by the capture unit. For example, in a nursery school, it can recognize the faces of children as they arrive. Step 4: The documentation unit documents the situation based on the information recognized by the recognition unit. For example, in a nursery school, information such as what kind of games the children were playing and who they were playing with can be documented. Step 5: The summarization section summarizes the information documented by the documentation section. For example, in a daycare center, this could summarize the children's activities from arrival time to nap time. Step 6: The record-keeping unit automatically creates work records based on the information summarized by the summarization unit. For example, a daycare center can automatically create a work record for one day.
[0151] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0152] Data generation model 58 is a form of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> Examples of generative AI include text generation AI, image generation AI, and multimodal generation AI. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats from audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVMs), k-means clustering, convolutional neural networks (CNNs), recurrent neural networks (RNNs), generative adversarial networks (GANs), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each of the above parts is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example.Furthermore, processing performed by AI, including generative AI, may be replaced with rule-based processing, and rule-based processing may be replaced with processing performed by AI, including generative AI.
[0153] Furthermore, the processing performed by the data processing system 10 described above is carried out by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it may also be carried out by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart device 14 or an external device, and the smart device 14 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0154] Each of the multiple elements described above, including the recording unit, capture unit, recognition unit, documentation unit, summarization unit, record creation unit, and data utilization unit, is implemented in at least one of the smart device 14 and the data processing unit 12. For example, the recording unit records video using the camera 42 of the smart device 14 and processes the recorded video using the specific processing unit 290 of the data processing unit 12. The capture unit captures video at regular intervals using the control unit 46A of the smart device 14 and processes the captured video using the specific processing unit 290 of the data processing unit 12. The recognition unit performs face recognition based on the captured video using the specific processing unit 290 of the data processing unit 12. The documentation unit documents the situation based on the information recognized by the specific processing unit 290 of the data processing unit 12. The summarization unit summarizes the information documented by the specific processing unit 290 of the data processing unit 12. The record creation unit automatically creates a business record based on the information summarized by the specific processing unit 290 of the data processing unit 12. The data utilization unit utilizes the data generated by the AI generated by the specific processing unit 290 of the data processing device 12 for healthcare business. The correspondence between each unit and the device or control unit is not limited to the example described above and can be changed in various ways.
[0155] [Second Embodiment] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0156] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0157] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.
[0158] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0159] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0160] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).
[0161] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0162] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing by the processor 28. The storage 32 stores the specific processing program 56.
[0163] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0164] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0165] In the smart glasses 214, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart glasses 214 also have a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.
[0166] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).
[0167] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0168] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.
[0169] The data processing system 210 according to the second embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 210 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart glasses 214 or an external device, and the smart glasses 214 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0170] Each of the multiple elements described above, including the recording unit, capture unit, recognition unit, documentation unit, summarization unit, record creation unit, and data utilization unit, is implemented, for example, in at least one of the smart glasses 214 and the data processing unit 12. For example, the recording unit records video using the camera 42 of the smart glasses 214 and processes the recorded video using the specific processing unit 290 of the data processing unit 12. The capture unit captures video at regular intervals using the control unit 46A of the smart glasses 214 and processes the captured video using the specific processing unit 290 of the data processing unit 12. The recognition unit performs face recognition based on the captured video using the specific processing unit 290 of the data processing unit 12. The documentation unit documents the situation based on the information recognized by the specific processing unit 290 of the data processing unit 12. The summarization unit summarizes the information documented by the specific processing unit 290 of the data processing unit 12. The record creation unit automatically creates a business record based on the information summarized by the specific processing unit 290 of the data processing unit 12. The data utilization unit utilizes the data generated by the AI generated by the specific processing unit 290 of the data processing device 12 for healthcare business. The correspondence between each unit and the device or control unit is not limited to the example described above and can be changed in various ways.
[0171] [Third Embodiment] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0172] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0173] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.
[0174] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0175] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0176] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).
[0177] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0178] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0179] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0180] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0181] In the headset terminal 314, specific processing is performed by the processor 46. The storage 50 stores a specific program 60. The processor 46 reads the specific program 60 from the storage 50 and executes the read specific program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific program 60 executed on the RAM 48. The headset terminal 314 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.
[0182] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).
[0183] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0184] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.
[0185] The data processing system 310 according to the third embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 310 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the headset terminal 314, but may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the headset terminal 314. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the headset terminal 314 or an external device, and the headset terminal 314 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0186] Each of the multiple elements described above, including the recording unit, capture unit, recognition unit, documentation unit, summarization unit, record creation unit, and data utilization unit, is implemented in at least one of the headset terminal 314 and the data processing unit 12. For example, the recording unit records video using the camera 42 of the headset terminal 314 and processes the recorded video using the specific processing unit 290 of the data processing unit 12. The capture unit captures video at regular intervals using the control unit 46A of the headset terminal 314 and processes the captured video using the specific processing unit 290 of the data processing unit 12. The recognition unit performs face recognition based on the captured video using the specific processing unit 290 of the data processing unit 12. The documentation unit documents the situation based on the information recognized by the specific processing unit 290 of the data processing unit 12. The summarization unit summarizes the documented information using the specific processing unit 290 of the data processing unit 12. The record creation unit automatically creates a business record based on the information summarized by the specific processing unit 290 of the data processing unit 12. The data utilization unit utilizes the data generated by the AI generated by the specific processing unit 290 of the data processing device 12 for healthcare business. The correspondence between each unit and the device or control unit is not limited to the example described above and can be changed in various ways.
[0187] [Fourth Embodiment] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0188] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0189] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.
[0190] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0191] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0192] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS image sensor or CCD image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).
[0193] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0194] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. The robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0195] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0196] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0197] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0198] In robot 414, specific processing is performed by processor 46. A specific program 60 is stored in storage 50. Processor 46 reads the specific program 60 from storage 50 and executes it on RAM 48. The specific processing is achieved by processor 46 acting as a control unit 46A according to the specific program 60 executed on RAM 48. Robot 414 also has data generation model 58 and emotion identification model 59, similar to those of the robot, and can perform processing similar to that of the specific processing unit 290 using these models.
[0199] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).
[0200] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0201] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.
[0202] The data processing system 410 according to the fourth embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 410 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the robot 414 or an external device, and the robot 414 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0203] Each of the multiple elements described above, including the recording unit, capture unit, recognition unit, documentation unit, summarization unit, record creation unit, and data utilization unit, is implemented, for example, in at least one of the robot 414 and the data processing unit 12. For example, the recording unit records video using the camera 42 of the robot 414 and processes the recorded video using the specific processing unit 290 of the data processing unit 12. The capture unit captures video at regular intervals using the control unit 46A of the robot 414 and processes the captured video using the specific processing unit 290 of the data processing unit 12. The recognition unit performs face recognition based on the captured video using the specific processing unit 290 of the data processing unit 12. The documentation unit documents the situation based on the information recognized by the specific processing unit 290 of the data processing unit 12. The summarization unit summarizes the information documented by the specific processing unit 290 of the data processing unit 12. The record creation unit automatically creates a business record based on the information summarized by the specific processing unit 290 of the data processing unit 12. The data utilization unit utilizes the data generated by the AI generated by the specific processing unit 290 of the data processing device 12 for healthcare business. The correspondence between each unit and the device or control unit is not limited to the example described above and can be changed in various ways.
[0204] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0205] Figure 9 shows the emotion map 400, in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0206] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0207] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0208] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, and motorcycles, emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated based, for example, on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0209] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0210] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0211] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing method for the specific process may be used, which includes computer 22 and multiple other computers.
[0212] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0213] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0214] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0215] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0216] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0217] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0218] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0219] Furthermore, although the above-described examples were divided into four embodiments, some or all of these embodiments may be combined. Also, the smart device 14, smart glasses 214, headset terminal 314, and robot 414 are just examples, and they may be combined, or other devices may be used. Also, although the above-described examples were divided into two embodiments, Embodiment 1 and Embodiment 2, these may be combined.
[0220] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and other things that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0221] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0222] (Note 1) A recording unit that records video, A capture unit that captures video recorded by the aforementioned recording unit at regular intervals, A recognition unit that performs face recognition based on the video captured by the capture unit, A documentation unit that documents the situation based on the information recognized by the recognition unit, A summarization unit that summarizes the information documented by the aforementioned documentation unit, A record creation unit that automatically creates business records based on the information summarized by the summarization unit, Equipped with A system characterized by the following features. (Note 2) The company has a data utilization department to use the data generated by the AI for healthcare businesses. The system described in Appendix 1, characterized by the features described herein. (Note 3) The aforementioned recording unit is It estimates the user's emotions and adjusts the recording start time based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 4) The aforementioned recording unit is During recording, ambient and background sounds are analyzed to highlight important audio events. The system described in Appendix 1, characterized by the features described herein. (Note 5) The aforementioned recording unit is During recording, it detects specific actions or movements and automatically highlights those parts. The system described in Appendix 1, characterized by the features described herein. (Note 6) The aforementioned recording unit is It estimates the user's emotions and adjusts the recording end time based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 7) The aforementioned recording unit is During recording, multiple cameras are synchronized to simultaneously record footage from different angles. The system described in Appendix 1, characterized by the features described herein. (Note 8) The aforementioned recording unit is During recording, the system tracks the movements of a specific person and focuses the recording on those movements. The system described in Appendix 1, characterized by the features described herein. (Note 9) The aforementioned capture unit is It estimates the user's emotions and adjusts the capture frequency based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 10) The aforementioned capture unit is During capture, the brightness and contrast of the video are automatically adjusted to obtain the optimal image. The system described in Appendix 1, characterized by the features described herein. (Note 11) The aforementioned capture unit is During capture, it detects specific events or actions and prioritizes capturing those moments. The system described in Appendix 1, characterized by the features described herein. (Note 12) The aforementioned capture unit is It estimates the user's emotions and adjusts the capture resolution based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 13) The aforementioned capture unit is During capture, apply color tones and filters to the video to improve visibility. The system described in Appendix 1, characterized by the features described herein. (Note 14) The aforementioned capture unit is During capture, zoom in on a specific area to obtain a detailed image. The system described in Appendix 1, characterized by the features described herein. (Note 15) The recognition unit, It estimates the user's emotions and adjusts the accuracy of facial recognition based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 16) The recognition unit, During recognition, facial expressions and movements are analyzed to estimate emotions and actions. The system described in Appendix 1, characterized by the features described herein. (Note 17) The recognition unit, During recognition, it recognizes multiple people simultaneously and tracks their actions. The system described in Appendix 1, characterized by the features described herein. (Note 18) The recognition unit, It estimates the user's emotions and adjusts how the recognition results are displayed based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 19) The recognition unit, During recognition, a person is identified using features other than their face. The system described in Appendix 1, characterized by the features described herein. (Note 20) The recognition unit, During recognition, past recognition data is referenced to improve recognition accuracy. The system described in Appendix 1, characterized by the features described herein. (Note 21) The aforementioned documentation unit, It estimates the user's emotions and adjusts the way the document is written based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 22) The aforementioned documentation unit, During documentation, the audio data from the video is analyzed, and the audio information is reflected in the document. The system described in Appendix 1, characterized by the features described herein. (Note 23) The aforementioned documentation unit, When documenting, emphasize specific keywords or phrases. The system described in Appendix 1, characterized by the features described herein. (Note 24) The aforementioned documentation unit, It estimates the user's emotions and adjusts the document length based on those emotions. The system described in Appendix 1, characterized by the features described herein. (Note 25) The aforementioned documentation unit, When documenting, time information from the video is added to create a document that follows a chronological order. The system described in Appendix 1, characterized by the features described herein. (Note 26) The aforementioned documentation unit, When documenting, multiple video sources are integrated to create a single document. The system described in Appendix 1, characterized by the features described herein. (Note 27) The summary section above is, It estimates the user's sentiment and adjusts the level of detail in the summary based on the estimated user sentiment. The system described in Appendix 1, characterized by the features described herein. (Note 28) The summary section above is, When summarizing, prioritize summarizing important events and actions. The system described in Appendix 1, characterized by the features described herein. (Note 29) The summary section above is, When summarizing, we refer to past summarization data to improve the accuracy of the summary. The system described in Appendix 1, characterized by the features described herein. (Note 30) The summary section above is, It estimates the user's sentiment and adjusts the order of summaries based on the estimated user sentiment. The system described in Appendix 1, characterized by the features described herein. (Note 31) The summary section above is, When summarizing, reflect the video metadata. The system described in Appendix 1, characterized by the features described herein. (Note 32) The summary section above is, When summarizing, combine multiple summaries to create a single summary. The system described in Appendix 1, characterized by the features described herein. (Note 33) The aforementioned record creation unit, It estimates the user's emotions and adjusts the recording format based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 34) The aforementioned record creation unit, When creating records, refer to past record data to improve the accuracy of the records. The system described in Appendix 1, characterized by the features described herein. (Note 35) The aforementioned record creation unit, When creating a record, highlight specific events or actions in your recording. The system described in Appendix 1, characterized by the features described herein. (Note 36) The aforementioned record creation unit, The system estimates the user's emotions and prioritizes recordings based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 37) The aforementioned record creation unit, When creating a record, time information is added to the video to create a record that follows a chronological order. The system described in Appendix 1, characterized by the features described herein. (Note 38) The aforementioned record creation unit, When creating a record, multiple records are combined into a single record. The system described in Appendix 1, characterized by the features described herein. (Note 39) The aforementioned data utilization unit is We estimate user emotions and adjust how data is used based on those estimated emotions. The system described in Appendix 2, characterized by the features described herein. (Note 40) The aforementioned data utilization unit is When utilizing data, referencing past data improves the accuracy of the utilization. The system described in Appendix 2, characterized by the features described herein. (Note 41) The aforementioned data utilization unit is When using data, filter the data according to a specific purpose. The system according to Appendix 2, characterized in that... (Appendix 42) The data utilization unit estimates the user's emotions and determines the priority of data based on the estimated user emotions The system according to Appendix 2, characterized in that... (Appendix 43) The data utilization unit reflects the metadata of the data during data utilization The system according to Appendix 2, characterized in that... (Appendix 44) The data utilization unit integrates multiple data sources to create one data set during data utilization The system according to Appendix 2, characterized in that... (Appendix 45) The data utilization unit visualizes the data and presents it in an easy-to-understand manner to the user during data utilization The system according to Appendix 2, characterized in that... (Appendix 46) The data utilization unit ensures the security of the data and protects privacy during data utilization The system according to Appendix 2, characterized in that... (Appendix 47) The data utilization unit performs a backup of the data to prevent data loss during data utilization The system according to Appendix 2, characterized in that... (Appendix 48) The data utilization unit outputs the analysis results of the data in a report format during data utilization The system according to Appendix 2, characterized in that... (Appendix 49) The data utilization unit analyzes the trend of the data and makes future predictions during data utilization The system according to Appendix 2, characterized in that... (Appendix 50) The aforementioned data utilization unit is When using data, detect anomalies in the data and issue alerts. The system described in Appendix 2, characterized by the features described herein. (Note 51) The aforementioned data utilization unit is When utilizing data, visualize the data and present it in an easy-to-understand manner for the user. The system described in Appendix 2, characterized by the features described herein. [Explanation of Symbols]
[0223] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots
Claims
1. A recording unit that records video, A capture unit that captures video recorded by the aforementioned recording unit at regular intervals, A recognition unit that performs face recognition based on the video captured by the capture unit, A documentation unit that documents the situation based on the information recognized by the recognition unit, A summarization unit that summarizes the information documented by the aforementioned documentation unit, A record creation unit that automatically creates business records based on the information summarized by the summarization unit, Equipped with A system characterized by the following features.
2. It includes a data utilization department for using data generated by AI in healthcare businesses. The system according to feature 1.
3. The aforementioned recording unit is It estimates the user's emotions and adjusts the recording start time based on the estimated emotions. The system according to feature 1.
4. The aforementioned recording unit is During recording, ambient and background sounds are analyzed to highlight important audio events. The system according to feature 1.
5. The aforementioned recording unit is During recording, it detects specific actions or movements and automatically highlights those parts. The system according to feature 1.
6. The aforementioned recording unit is It estimates the user's emotions and adjusts the recording end time based on the estimated emotions. The system according to feature 1.
7. The aforementioned recording unit is During recording, multiple cameras are synchronized to simultaneously record footage from different angles. The system according to feature 1.
8. The aforementioned recording unit is During recording, the system tracks the movements of a specific person and focuses the recording on those movements. The system according to feature 1.
9. The aforementioned capture unit is It estimates the user's emotions and adjusts the capture frequency based on the estimated user emotions. The system according to feature 1.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A