System
A system using high-resolution cameras and AI analysis automates childcare diary generation, monitoring, and provides safety alerts, addressing the challenges of workload and safety in childcare settings.
Patent Information
- Application Number
- JP2024129436
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-05
- Publication Date
- 2026-02-18
AI Technical Summary
Creating daily childcare diaries for multiple children is time-consuming and labor-intensive, and monitoring children's behavior for safety is burdensome for childcare workers, making it difficult to provide high-quality and safe care.
A system using high-resolution cameras, facial recognition, and AI analysis to automatically generate childcare diaries, distribute them to parents, monitor risky behavior, and provide growth analysis and advice to childcare workers.
Reduces the workload of childcare workers, ensures the safety and quality of childcare by automating diary generation, monitoring, and providing real-time alerts and growth insights.
Smart Images

Figure 2026027015000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] In today's childcare settings, creating daily childcare diaries for many children is extremely time-consuming and labor-intensive. Furthermore, a shortage of childcare workers and a lack of time make it difficult for them to record the detailed behavior of each child and share it with parents. Additionally, preventing accidents within the nursery is an important issue, but monitoring children's risky behavior in real time places a heavy burden on personnel. In these circumstances, a system that reduces the burden on childcare workers while ensuring safe, high-quality childcare is needed. [Means for solving the problem]
[0005] In order to solve the above problems, the present invention provides the following means. A means is provided for capturing images of classrooms and playgrounds with a high-resolution camera and receiving the video data. A means is provided for applying facial recognition technology using the received video data to identify multiple individual subjects (children). A means is included for tracking the behavior of the identified individual subjects and automatically generating a childcare diary. Furthermore, a means is provided for automatically distributing the generated childcare diary to parents. A means is also provided for analyzing the video data and issuing an alert if dangerous behavior is detected, thereby ensuring the safety of the children. Finally, a system is proposed that includes a means for analyzing the growth of individual subjects based on the childcare diary and providing advice to childcare workers.
[0006] A "high-resolution camera" is a high-resolution camera that can capture detailed images of childcare environments such as classrooms and playgrounds.
[0007] "Video data" refers to information about videos and images taken with high-resolution cameras.
[0008] "Facial recognition technology" refers to algorithms and software technology that can identify faces captured in video data and identify specific individuals.
[0009] "Individual targets" refers to children in nursery schools, kindergartens, etc.
[0010] "Tracking" is the process of following and recording the movements of individual objects based on video data.
[0011] A "childcare diary" is a document that records an individual's daily activities, health condition, interactions, etc.
[0012] "Automatic generation means" refers to a method or device that allows a system to automatically create data without human intervention.
[0013] The "distribution means" refers to a method or device for automatically transmitting the created childcare diary to parents.
[0014] "Dangerous behavior" refers to behavior that could lead to an accident or injury to a child.
[0015] An "alert means" is a method or device for issuing a warning signal when dangerous behavior is detected.
[0016] "Means for analyzing growth" are methods or devices for evaluating the growth and development of individual subjects based on data from childcare diaries.
[0017] The "means for providing advice" refers to a method or device for providing specific advice or suggestions to childcare workers based on the results of the growth analysis. [Brief explanation of the drawings]
[0018] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0019] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0020] First, the terms used in the following description will be explained.
[0021] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0022] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0023] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0024] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0026] [First embodiment]
[0027] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0028] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0029] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0030] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0031] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0033] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0034] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0035] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0036] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0037] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0038] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0039] This invention is a system that reduces the burden on childcare workers and ensures safe, high-quality childcare. Using high-resolution cameras, servers, and terminals, the system utilizes face recognition technology and AI analysis technology to automatically generate childcare diaries, distribute them to parents, monitor and alert for risky behavior, analyze child development, and provide advice to childcare workers.
[0040] 1. Camera video capture
[0041] The server receives real-time video data from high-resolution cameras installed in classrooms and on the playground, which have a wide field of view and can accurately capture detailed activity within the school.
[0042] 2. Face recognition and individual object identification
[0043] The server analyzes the received video data using facial recognition AI technology to identify the children in the video. Facial recognition technology identifies the ID of each individual child by comparing the facial image in the video with a pre-registered database.
[0044] 3. Behavior tracking and automatic generation of childcare diary
[0045] The server tracks the behavior of identified children and records their behavioral patterns (e.g., types of play, interpersonal relationships, and health status). Based on this, a daycare diary is automatically generated. The diary details each child's daily activities, interactions, and health status.
[0046] 4. Distribution of the childcare diary to parents
[0047] The server automatically distributes the generated childcare diary to the parents of the corresponding child via their email address or a dedicated app, allowing parents to keep track of their child's daily activities in real time.
[0048] 5. Monitoring risky behavior and issuing alerts
[0049] The server analyzes the video data in real time and monitors the children's risky behavior. If any risky behavior (e.g., jumping from a high place or rough play) is detected, an alert is sent immediately.
[0050] The terminal (childcare worker) receives this alert and can respond quickly.
[0051] 6. Growth analysis and advice for childcare workers
[0052] The server periodically compiles data from the childcare diary and analyzes the growth of each child. Based on the results of this analysis, it identifies individual growth trends and provides advice to childcare workers regarding specific childcare policies and activities. The advice is automatically sent to the childcare workers' devices and can be used in practice.
[0053] For example, when Taro is playing in the playground, a camera captures his movements, and the server identifies him using facial recognition technology and records his actions. His behavior as he plays with building blocks with his friends is tracked, and a diary entry is made stating, "Taro played with building blocks with his friends and worked together to build a big tower." This diary entry is automatically sent to Taro's mother's email address that afternoon.
[0054] In addition, if Taro attempts dangerous behavior in the playground, such as standing up high on a swing, the server will detect this and send an alert to the nursery teacher's device, allowing the nursery teacher to immediately rush over to Taro and ensure his safety.
[0055] The above is a specific embodiment for carrying out the present invention. This system allows childcare workers to save time and effort, ensure the safety of children, and provide high-quality childcare.
[0056] The processing flow will be explained below.
[0057] Processing for automatic childcare diary generation system using face recognition
[0058] From video capture to automatic generation of childcare diary
[0059] Step 1: Receiving video data
[0060] The server receives video data in real time from high-resolution cameras installed in classrooms and playgrounds, and the video data is temporarily stored in dedicated storage.
[0061] Step 2: Face recognition and individual subject identification
[0062] The server inputs the received video data into a facial recognition AI model to identify the child in the video. The facial recognition technology compares the facial images in the video with a pre-registered database to identify the ID of each individual subject (child).
[0063] Step 3: Tracking behavior
[0064] The movements of identified individuals are tracked in real time, with the server recording the start and end of specific actions (e.g., playing, crafting, talking) with timestamps.
[0065] Step 4: Creating a childcare diary
[0066] The server automatically generates a daily diary for each child based on the tracking data, detailing each child's daily activities, interactions, and health status.
[0067] Step 5: Distributing the journal
[0068] The server automatically distributes the generated childcare diary to the parents of the corresponding children via their email address or a dedicated app.
[0069] Monitoring and alerting for risky behavior
[0070] Step 1: Real-time monitoring of video data
[0071] The server analyzes the video data in real time and monitors the behavior of the children.
[0072] Step 2: Detect risky behavior
[0073] The server uses a vision AI model to detect dangerous behavior (e.g., jumping from a high place, rough play).
[0074] Step 3: Trigger an alert
[0075] If dangerous behavior is detected, the server immediately generates an alert and sends it to the childcare worker's device.
[0076] Step 4: Receive and respond to alerts
[0077] The device (childcare worker) receives the alert and takes action to respond in real time.
[0078] Growth analysis and advice for childcare workers
[0079] Step 1: Collecting childcare diary data
[0080] The server periodically compiles data from the daycare diary and analyzes growth and behavioral patterns over a set period of time.
[0081] Step 2: Conduct a growth analysis
[0082] The server generates and evaluates the growth trends of each child based on the analysis data.
[0083] Step 3: Generate Advice
[0084] Based on growth trends, the server generates specific advice for nursery teachers, including specific activities and methods for promoting children's growth.
[0085] Step 4: Delivering advice
[0086] The server delivers the generated advice to the childcare worker's device, who receives it and puts it into practice.
[0087] This is the specific processing flow of the system, which reduces the burden on nursery teachers and ensures the safety and growth of children.
[0088] Example 1
[0089] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0090] In conventional childcare settings, childcare workers have to manually create childcare records, monitor children's behavior, and regularly provide information to parents. This increases the burden on childcare workers and can lead to a decline in the quality of childcare. It is also difficult to immediately detect risky behavior or provide specific advice based on growth analysis. Therefore, there was a need for a system that could reduce the burden on childcare workers and ensure safe, high-quality childcare.
[0091] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0092] In this invention, the server includes: means for capturing images of classrooms and playgrounds with a high-resolution camera; means for receiving the image data and identifying multiple individual objects using image authentication technology; means for tracking the behavior of the identified individual objects using computer vision technology and automatically generating a childcare record; means for automatically distributing the generated childcare record to parents; means for issuing a warning when dangerous behavior is detected; means for analyzing growth based on the childcare record and providing advice to childcare workers; means for processing the image data frame by frame and comparing it with multiple databases; means for tracking the positions and movements of children using computer vision technology; means for saving the tracking data in chronological order and performing activity mapping in real time; means for creating a childcare record in natural language using a generative AI model; means for automatically allocating contact information for each parent; and means for generating advice regarding childcare policies and activities based on the activity analysis results. This reduces the burden on childcare workers, improves safety, and enables the provision of high-quality childcare.
[0093] A "high-definition camera" is a camera that can capture wide-area images in high resolution.
[0094] "Video data" refers to information recorded in digital format from video captured by a camera.
[0095] "Image recognition technology" refers to algorithms and techniques used to identify individual subjects from video data.
[0096] "Individual subjects" refer to individual people, such as children, who participate in activities within the nursery school.
[0097] "Computer vision technology" is a technology that uses computers to analyze images and videos to recognize objects and track behavior.
[0098] A "childcare record" is a childcare diary that records in detail the daily activities and behavior of the children.
[0099] A "warning" is an alert that is issued when risky behavior is detected.
[0100] "Guardian" refers to the parent or legal guardian who is the caretaker of the child.
[0101] A "server" is a computer system that receives, analyzes, stores, and distributes video data.
[0102] "Tracking" means to continuously follow the movement of an object.
[0103] A "generative AI model" is an artificial intelligence model that generates sentences in natural language based on accumulated data.
[0104] "Contact information" refers to the parent's email address and account information for the dedicated application.
[0105] "Activity mapping" means visualizing the behavior and activities of kindergarten children along a timeline.
[0106] "Advice" means specific instructions or advice given to improve the quality of childcare.
[0107] "Database" means an electronic record containing information used to identify or match individuals.
[0108] This invention is a system that reduces the workload of childcare workers and provides safe, high-quality childcare. This system uses high-resolution image capture devices, servers, and terminals, and makes full use of image authentication and computer vision technologies to provide the following integrated functions.
[0109] The server acquires video data in real time from high-resolution camera devices installed in classrooms and playgrounds. These cameras are capable of covering a wide area in high resolution, and the video is streamed to the server using the RTSP protocol. The server stores this video data in a buffer frame by frame and begins processing.
[0110] The server then analyzes the received video data using image recognition technology. Specifically, it uses OpenCV and the Dlib library to detect the faces of the children and compare them with a pre-registered database to identify each individual child. This allows it to assign a unique ID to each child.
[0111] The behavior of identified children is tracked using computer vision technology on the server. Specifically, the OpenPose library is used to analyze the children's positions and movements. This data is saved in chronological order, and activity mapping is performed in real time. This allows for a detailed record of each child's activities.
[0112] Based on this tracking data, the server uses a generative AI model to automatically generate childcare records. The generated childcare records are detailed in natural language and include the child's daily activities, interactions, health status, etc. For example, tracking Taro playing with his friends using building blocks might result in a note in the childcare record saying, "Taro worked together with his friends to build a big tower."
[0113] The generated childcare records are automatically sent to the parents of the corresponding children via a server. The records are sent via email addresses registered in advance by the parents or via a dedicated childcare app. This allows parents to keep track of their child's daily activities in real time.
[0114] The server also analyzes video data in real time and monitors for dangerous behavior. For example, if the server detects a child engaging in dangerous behavior, such as jumping from a high place, it immediately issues an alert and notifies the childcare worker's device. This alert is sent as a push notification, allowing the childcare worker to respond quickly.
[0115] Furthermore, the server periodically compiles the childcare records and analyzes the growth of each child. Based on the results of this analysis, the generative AI model generates specific advice on childcare policies and activities, which are automatically distributed to the childcare worker's device. The childcare worker can use this advice to adjust and improve the childcare plan.
[0116] As a concrete example, when Taro is playing in the playground, a camera captures his movements, and the server uses facial recognition technology to identify him and record his actions. Computer vision technology tracks Taro playing with his friends using building blocks, and a generative AI model is used to create a childcare record stating, "Taro worked with his friends to build a big tower." This childcare record is automatically sent to Taro's mother. Furthermore, if the server detects Taro standing up high on the swing, it immediately sends an alert to the childcare worker's device, allowing the worker to respond quickly.
[0117] Prompt Sentence Examples
[0118] "How do you analyze real-time footage of children playing outside and send alerts to childcare workers if a particular child is exhibiting risky behavior?"
[0119] By utilizing this system, the workload of childcare workers will be significantly reduced, the safety of children will be improved, and high-quality childcare will be provided.
[0120] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0121] Step 1:
[0122] The server acquires video data in real time from high-resolution camera devices installed in classrooms and playgrounds. The input is the video stream from the camera, and the output is frame data stored on the server. Specifically, the server acquires data from the camera using the RTSP protocol and adds each frame to the processing queue.
[0123] Step 2:
[0124] The server analyzes the received video data using image recognition technology to detect the faces of the children. The input is the frame data from step 1, and the output is a feature vector of the detected face image. Specifically, it uses OpenCV and the Dlib library to detect faces in the frame and extract facial features.
[0125] Step 3:
[0126] The server uses facial recognition technology to compare the facial features detected with a pre-registered database to identify individual subjects (children). The input is the facial feature vector from step 2, and the output is the child's ID. Specifically, the server compares the facial feature vector with existing features in the database and identifies matching IDs based on cosine similarity.
[0127] Step 4:
[0128] The server uses computer vision technology to track the behavior of the identified children. The input is the child's ID from step 3 and the frame data from step 1, and the output is time-series data of the child's position and movement. Specific movements are tracked by detecting the child's skeleton using the OpenPose library.
[0129] Step 5:
[0130] The server automatically generates childcare records based on the tracking data using a generative AI model. The input is the time series data from step 4, and the output is a childcare record written in natural language. Specifically, the tracking data is input into the generative AI model, and a sentence describing the daily activities of the children is generated.
[0131] Step 6:
[0132] The server automatically distributes the generated childcare record to the corresponding guardian. The input is the childcare record from step 5, and the output is the information sent to the guardian's device. Specifically, the server refers to the guardian's contact database and sends the childcare record via email or a dedicated app.
[0133] Step 7:
[0134] The server analyzes the video data in real time and monitors for risky behavior. The input is the frame data from step 1, and the output is a warning notification when risky behavior is detected. Specifically, it uses computer vision technology to detect risky behavior patterns and sends a warning to the nursery teacher's device.
[0135] Step 8:
[0136] The server periodically compiles the childcare records and analyzes the children's growth. The input is the childcare record data from Step 5, and the output is the growth analysis results and advice. Specific operations include analyzing the compiled data and using a generative AI model to generate specific advice regarding childcare policies and activities.
[0137] Prompt Sentence Examples
[0138] "How do you analyze real-time footage of children playing outside and send alerts to childcare workers if a particular child is exhibiting risky behavior?"
[0139] This series of processes enables nursery school teachers to ensure the safety of children while saving time and effort, and to provide high-quality childcare.
[0140] (Application example 1)
[0141] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0142] In recent years, improving work efficiency and ensuring safety in factories have become extremely important issues. In particular, in manufacturing sites where many workers work simultaneously, there is a high risk of human error and dangerous behavior, and there is a growing demand for systems that can quickly detect and address such errors. Furthermore, manually creating work logs requires time and effort, making it difficult to immediately improve work efficiency. The present invention addresses these issues by providing advanced monitoring and automated work log generation to improve work efficiency and ensure safety in factories.
[0143] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0144] In this invention, the server includes means for capturing images of the work environment with a high-resolution camera, means for receiving the image data and identifying multiple individual objects using facial recognition technology, means for tracking the behavior of the identified individual objects and automatically generating a work log, means for automatically distributing the generated work log to relevant parties, means for issuing an alert when dangerous behavior is detected, and means for analyzing work efficiency based on the work log and providing relevant parties with points for improvement. This makes it possible to immediately grasp and analyze work efficiency within a factory and improve safety.
[0145] A "high-resolution camera" is a photographic device with high resolution that can record the work environment in detail.
[0146] "Work environment" refers to the place or area within a factory where work is carried out.
[0147] "Video data" is digital data containing visual information captured by a high-resolution camera.
[0148] "Facial recognition technology" is a biometric authentication technology for identifying people in video data.
[0149] "Individual Subject" refers to each identified worker within a factory.
[0150] "Behavior tracking" refers to following and recording the movements of identified individual subjects.
[0151] A "work log" is a digital record that records the actions and work of an individual.
[0152] "Automatically generating a work log" means that the server automatically creates a work log from the behavioral data of an individual subject.
[0153] "Stakeholders" refers to factory managers and supervisors who require work logs.
[0154] "Distribution" means sending the generated work log to the relevant parties.
[0155] "Issuing an alert" means sending a warning message when dangerous behavior is detected.
[0156] "Analyzing work efficiency" means analyzing the data in the work log to evaluate the efficiency of the work.
[0157] "Providing improvements" means presenting specific suggestions to stakeholders to improve work efficiency.
[0158] This invention is a system that improves work efficiency and ensures safety in factories by utilizing high-resolution cameras, facial recognition technology, and AI analysis technology. The system consists of a server, a camera, smart glasses, and a dedicated app.
[0159] System configuration
[0160] 1. High-quality camera:
[0161] The working environment within the factory is photographed in high resolution, and detailed video data is sent to a server in real time.
[0162] 2. Server:
[0163] The server processes the received video data and identifies the worker using facial recognition technology. Azure Face API, for example, is used for facial recognition. The actions of identified workers are tracked, and a work log is automatically generated based on that data. The work log contains detailed records of each worker's movements and tasks.
[0164] 3. Risky behavior monitoring and alerting:
[0165] The server analyzes video data and work logs in real time and immediately issues an alert if any dangerous behavior is detected. The alert is then sent to the smart glasses to alert the worker, enabling a rapid response.
[0166] 4. Work log distribution:
[0167] The generated work logs are automatically distributed to relevant parties (factory managers and supervisors) via email addresses or a dedicated app.
[0168] 5. Analyzing work efficiency and providing improvements:
[0169] The server analyzes work efficiency based on the work logs and provides specific improvements to the relevant parties, thereby optimizing work efficiency throughout the factory.
[0170] Hardware and software used
[0171] Hardware: high-resolution cameras, smart glasses (e.g., Microsoft HoloLens)
[0172] Software: Facial recognition AI (e.g., Azure Face API), behavioral analysis AI (e.g., TensorFlow, OpenCV), database (e.g., Azure SQL Database)
[0173] Example of operation
[0174] For example, if a worker is operating heavy machinery in an inappropriate manner, a high-resolution camera captures this, the server uses facial recognition technology to identify the worker, and AI determines that the operation is inappropriate. An alert is immediately sent to the smart glasses, prompting the worker to operate in a safe manner. At the same time, details of the operation are automatically generated as a work log and sent to the factory manager. Based on this work log, areas for improving work efficiency are analyzed and provided to relevant parties.
[0175] Example of input prompt for generative AI model
[0176] prompt:
[0177] Design a smart glasses application for factory monitoring with self-diagnosis function. Please pay attention to the following points:
[0178] 1. Tracking worker behavior in real time using cameras and smart glasses.
[0179] 2. Use facial recognition technology to identify individual workers.
[0180] 3. If risky behavior is detected, an alert is sent immediately.
[0181] 4. Automatically generate work logs and analyze areas for improvement.
[0182] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0183] Step 1:
[0184] High-resolution cameras capture the factory environment. The input is real-time footage from within the factory, and the output is high-resolution video data. Specifically, the cameras have a field of view that covers the entire factory, capturing detailed images of workers and machine operations.
[0185] Step 2:
[0186] The server receives the video data sent from the camera. The input is high-definition video data that is temporarily stored in the server. The output is video data stored in the server's storage.
[0187] Step 3:
[0188] The server analyzes the video data using facial recognition AI (Azure Face API) to identify multiple workers. The input is the saved video data, and the output is the identification information of each worker. Specifically, the AI analyzes the facial images in the video and identifies the worker's ID by comparing them with a pre-registered database.
[0189] Step 4:
[0190] The server tracks the actions of each identified worker. The input is identification information and real-time video data, and the output is individual behavioral patterns. Specifically, the server continuously tracks the movements of workers and records what tasks they are performing.
[0191] Step 5:
[0192] The server automatically generates work logs based on the behavioral data it tracks. The input is behavioral pattern data, and the output is a digital work log. Specifically, the AI organizes the recorded information and provides a detailed record of each worker's daily activities and their content.
[0193] Step 6:
[0194] The server automatically distributes the generated work log to the relevant parties. The input is the generated work log, and the output is the work log sent to the relevant parties' email addresses or a dedicated app. Specifically, the log file is sent to the specified recipient via a mail server or application.
[0195] Step 7:
[0196] The server analyzes real-time video data and issues an alert if it detects dangerous behavior. The input is real-time video data, and the output is an alert message sent to the smart glasses. Specifically, when the AI detects an abnormal behavior pattern, it immediately generates an alert and sends it to the worker's smart glasses.
[0197] Step 8:
[0198] The server analyzes work efficiency based on work logs and provides suggestions for improvement to the relevant parties. The input is daily work log data, and the output is an analysis report including suggestions for improvement. Specifically, the server compares the data with past work data, proposes specific improvements for improving work efficiency and safety, and sends them to the relevant parties.
[0199] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0200] This invention is a system that reduces the burden on childcare workers and ensures safe, high-quality childcare. Using high-resolution cameras, servers, and terminals, the system utilizes face recognition technology, emotion engines, and AI analysis technology to automatically generate childcare diaries, distribute them to parents, monitor and alert on risky behavior, analyze child development, and provide advice to childcare workers.
[0201] 1. Camera video capture
[0202] The server receives real-time video data from high-resolution cameras installed in classrooms and on the playground, which have a wide field of view and can accurately capture detailed activity within the school.
[0203] 2. Face recognition and individual object identification
[0204] The server analyzes the received video data using facial recognition AI technology to identify the children in the video. The facial recognition technology compares the facial images in the video with a pre-registered database to identify the ID of each individual subject (child).
[0205] 3. Behavior tracking and automatic generation of childcare diary
[0206] The server tracks the behavior of identified children and records their behavioral patterns (e.g., types of play, interpersonal relationships, and health status). Based on this, a daycare diary is automatically generated. The diary details each child's daily activities, interactions, and health status.
[0207] 4. Distribution of the childcare diary to parents
[0208] The server automatically distributes the generated childcare diary to the parents of the corresponding child via their email address or a dedicated app, allowing parents to keep track of their child's daily activities in real time.
[0209] 5. Monitoring risky behavior and issuing alerts
[0210] The server analyzes the video data in real time and monitors the children's risky behavior. If any risky behavior (e.g., jumping from a high place or rough play) is detected, an alert is sent immediately.
[0211] The terminal (childcare worker) receives this alert and can respond quickly.
[0212] 6. Growth analysis and advice for childcare workers
[0213] The server periodically compiles data from the childcare diary and analyzes the growth of each child. Based on the results of this analysis, it identifies individual growth trends and provides advice to childcare workers regarding specific childcare policies and activities. The advice is automatically sent to the childcare workers' devices and can be used in practice.
[0214] 7. Introducing the Emotion Engine
[0215] The server inputs the video and audio data into the emotion engine, which analyzes the child's emotional state based on their facial expressions and tone of voice. The emotion engine then identifies the emotion the child is feeling (e.g., joy, sadness, anger, anxiety).
[0216] Based on the emotional states identified by the emotion engine, the server adds emotional data to the childcare diary and development analysis, for example, recording how children felt during play and how specific activities affected their emotions.
[0217] The server generates and delivers advice to childcare workers and parents based on their emotional state, suggesting appropriate ways to respond. For example, if a child is feeling anxious, the server identifies the root cause and provides specific advice on how to support the child.
[0218] For example, when Taro is playing in the playground, a camera captures his movements, and the server identifies him using facial recognition technology and records his actions. His behavior as he plays with building blocks with his friends is tracked, and a diary entry is made stating, "Taro played with building blocks with his friends and worked together to build a big tower." This diary entry is automatically sent to Taro's mother's email address that afternoon.
[0219] Furthermore, if Taro attempts dangerous behavior such as standing up too high on the swing, the server will detect this and immediately send an alert to the nursery teacher's device, allowing the nursery teacher to immediately rush over to Taro and ensure his safety.
[0220] Furthermore, the emotion engine analyzes Taro's facial expressions and tone of voice and detects that he looks anxious. Based on this, the server sends advice to the nursery teacher's device, such as, "Taro is likely feeling anxious, so please talk to him a bit to reassure him."
[0221] The above is a specific embodiment for carrying out the present invention. This system allows childcare workers to reduce their workload, ensure the safety and emotional growth of children, and provide high-quality childcare.
[0222] The processing flow will be explained below.
[0223] Processing for an automatic childcare diary generation system that combines face recognition and emotion engine
[0224] From video capture to emotion analysis
[0225] Step 1: Receiving video data
[0226] The server receives video data in real time from high-resolution cameras installed in classrooms and playgrounds, and the video data is temporarily stored in dedicated storage.
[0227] Step 2: Face recognition and individual subject identification
[0228] The server inputs the received video data into a facial recognition AI model to identify the children in the video. The facial recognition technology compares the facial images in the video with a pre-registered database to identify the ID of each individual subject (child).
[0229] Step 3: Tracking behavior
[0230] The movements of identified individuals are tracked in real time, with the server recording the start and end of specific actions (e.g., playing, crafting, talking) with timestamps.
[0231] Step 4: Analysis by Emotion Engine
[0232] The server inputs video and audio data into an emotion engine, which analyzes the children's facial expressions and tone of voice to determine their emotional state (e.g., joy, sadness, anger, anxiety).
[0233] Step 5: Creating a childcare diary
[0234] The server automatically generates a daily diary for each child based on tracking data and emotion analysis data, detailing each child's daily activities, interactions, health status, and emotional state.
[0235] Delivery of childcare diary to parents
[0236] Step 1: Generate a journal
[0237] The server automatically distributes the generated childcare diary to the parents of the corresponding children via their email addresses or a dedicated app.
[0238] Monitoring and alerting for risky behavior
[0239] Step 1: Real-time monitoring of video data
[0240] The server analyzes the video data in real time and monitors the behavior of the children.
[0241] Step 2: Detect risky behavior
[0242] The server uses a vision AI model to detect dangerous behavior (e.g., jumping from a high place, rough play).
[0243] Step 3: Trigger an alert
[0244] If dangerous behavior is detected, the server immediately generates an alert and sends it to the childcare worker's device.
[0245] Step 4: Receive and respond to alerts
[0246] The device (childcare worker) receives the alert and takes action to respond in real time.
[0247] Growth analysis and advice for childcare workers
[0248] Step 1: Collecting childcare diary data
[0249] The server periodically compiles data from the daycare diary and analyzes growth and behavioral patterns over a set period of time.
[0250] Step 2: Conduct a growth analysis
[0251] The server generates and evaluates the growth trends of each child based on the analysis data.
[0252] Step 3: Generate Advice
[0253] Based on growth trends, the server generates specific advice for nursery teachers, including specific activities and methods for promoting children's growth.
[0254] Step 4: Delivering advice
[0255] The server delivers the generated advice to the childcare worker's device, who receives it and puts it into practice.
[0256] Further processing by the emotion engine
[0257] Step 1: Collecting emotion data
[0258] The server integrates the analytical data from the emotion engine into the childcare diary, allowing for an understanding of emotional fluctuations and trends.
[0259] Step 2: Generating sentiment-based advice
[0260] Based on the emotional data, the server generates advice suggesting appropriate responses for childcare workers and parents.
[0261] Step 3: Delivering emotional advice
[0262] The server delivers advice based on the generated emotions to the childcare worker's device, who then receives it and puts it into practice.
[0263] For example, when Taro is playing in the playground, a camera captures his movements, and the server identifies him using facial recognition technology and records his actions. His behavior as he plays with building blocks with his friends is tracked, and a diary entry is made stating, "Taro played with building blocks with his friends and worked together to build a big tower." This diary entry is automatically sent to Taro's mother's email address that afternoon.
[0264] Furthermore, if Taro attempts dangerous behavior such as standing up too high on the swing, the server will detect this and immediately send an alert to the nursery teacher's device, allowing the nursery teacher to immediately rush over to Taro and ensure his safety.
[0265] Furthermore, the emotion engine analyzes Taro's facial expressions and tone of voice and detects that he looks anxious. Based on this, the server sends advice to the nursery teacher's device, such as, "Taro is likely feeling anxious, so please talk to him a bit to reassure him."
[0266] The above is a specific embodiment for carrying out the present invention. This system allows childcare workers to reduce their workload, ensure the safety and emotional growth of children, and provide high-quality childcare.
[0267] Example 2
[0268] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0269] In childcare facilities, the workload of childcare workers is increasing, and it is often the case that they are not doing enough to ensure the safety of children, monitor their growth, or report to parents. It is also difficult to accurately grasp the emotional state of children and respond appropriately. This puts the quality of childcare at risk of declining.
[0270] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0271] In this invention, the server includes means for capturing images of classrooms and playgrounds with a high-resolution camera, means for receiving the image data and identifying multiple individual objects using facial recognition technology, means for tracking the behavior of the identified individual objects and automatically generating a childcare diary, means for automatically distributing the generated childcare diary to parents, means for issuing an alert when dangerous behavior is detected, means for analyzing growth based on the childcare diary and providing advice to childcare workers, means for analyzing emotional states from video data and audio data using an emotion engine, and means for suggesting appropriate response methods to childcare workers and parents based on the analyzed emotional states. This reduces the burden on childcare workers, enables management of the safety and growth of children, and enables understanding of emotional states and appropriate responses.
[0272] A "high-resolution camera" is a camera that has a wide field of view and can capture detailed video data in high resolution.
[0273] "Video data" refers to video information acquired from cameras installed in classrooms and playgrounds.
[0274] "Facial recognition technology" is a technology that identifies individuals by identifying faces in video data and comparing them with a pre-registered database.
[0275] "Individual object" is a term that refers to a specific person, especially a kindergartener.
[0276] "Behavioral tracking" is a method of monitoring and recording the movements and behavioral patterns of individual subjects in real time.
[0277] A "childcare diary" is a report that details the behavior, interactions, and health status of the children during a day's childcare activities.
[0278] "Guardian" is a term that refers to the parents or guardians of kindergarten children.
[0279] An "alert" is a warning message that is sent immediately when dangerous behavior or an abnormality is detected in real time.
[0280] "Growth analysis" is a method of analyzing children's growth patterns based on data from the nursery school diary and evaluating developmental trends.
[0281] "Advice" refers to specific methods of response and guidelines provided to childcare workers and parents based on growth analysis and emotional analysis.
[0282] An "emotion engine" is a technology that analyzes video and audio data to identify the emotional state of an individual subject.
[0283] "Response methods" are specific measures that, based on the analyzed data, indicate appropriate approaches in caring for and guiding children.
[0284] This invention is a system that reduces the burden on childcare workers and ensures safe, high-quality childcare. This system uses high-resolution cameras, servers, and terminals, and makes full use of face recognition technology, emotion engines, and AI analysis technology to automatically generate childcare diaries, distribute them to parents, monitor and alert on risky behavior, analyze child development, and provide advice to childcare workers.
[0285] Hardware and software used
[0286] server
[0287] Data servers with high processing power (e.g., Dell PowerEdge series)
[0288] Facial recognition AI technology (e.g., FaceNet)
[0289] Emotion engine (e.g. Microsoft Azure Emotion API)
[0290] Machine learning algorithms (e.g., Python's scikit-learn library)
[0291] camera
[0292] High-resolution network camera (e.g. AXIS P1368-E)
[0293] Terminal
[0294] Tablet or smartphone for childcare workers (e.g. iPad, Android tablet)
[0295] Explanation of program processing
[0296] High-definition camera video capture
[0297] The server receives real-time video data from high-resolution cameras installed in classrooms and playgrounds. These cameras have a wide field of view and can accurately capture detailed activity within the school. The video data is stored in the server's storage and used for analysis.
[0298] Identifying individual subjects using facial recognition technology
[0299] The server analyzes the received video data using facial recognition AI technology to identify the children in the video. The facial recognition technology compares the facial images in the video with a pre-registered database to identify the ID of each individual subject (child). This allows for accurate tracking of each child's activities.
[0300] Behavior tracking and automatic generation of childcare diary
[0301] The server tracks the identified children's behavior in real time and records their behavioral patterns (e.g., types of play, interpersonal relationships, and health status). Based on this behavioral data, a daycare diary is automatically generated, detailing each child's daily activities, interactions, and health status.
[0302] Delivery of childcare diary to parents
[0303] The generated daycare diary is automatically sent to the parents of the children via their email address or a dedicated app (e.g., Custom Nursery App). This allows parents to keep track of their children's daily activities in real time.
[0304] Monitoring risky behavior and issuing alerts
[0305] The server analyzes the video data in real time and monitors the children's risky behavior. For example, if dangerous behavior such as jumping from a high place or rough play is detected, an alert is immediately issued. A push notification is sent to the childcare worker's device (e.g., tablet), allowing the childcare worker to respond quickly.
[0306] Growth analysis and advice for childcare workers
[0307] The server periodically compiles data from the childcare diary and analyzes the growth of each child. Based on the results of this analysis, it identifies individual growth trends and provides advice to childcare workers regarding specific childcare policies and activities. This advice is automatically sent to devices and can be used by childcare workers in their daily work.
[0308] Introducing the Emotion Engine
[0309] The server inputs video and audio data into an emotion engine, which analyzes the emotional state of the child based on facial expressions, tone of voice, etc. The emotion engine identifies the child's emotions (e.g., joy, sadness, anger, anxiety). This emotional state is reflected in the childcare diary and growth analysis, and is used as part of advice for childcare workers and parents. Specifically, the server calls an emotion analysis API to obtain the analysis results, and a recommendation algorithm then suggests appropriate responses based on these results.
[0310] Specific examples
[0311] While Taro is playing in the playground, a high-definition camera captures his movements, and the server uses facial recognition technology to identify him and record his actions. His behavior as he plays with building blocks with his friends is tracked, and an entry is made in the daycare diary, stating, "Taro played with building blocks with his friends and worked together to build a big tower." This diary is automatically sent to Taro's mother's email address that afternoon. Furthermore, if Taro attempts the dangerous act of standing up high on the swing, the server detects this and immediately sends an alert to the nursery teacher's device, allowing the nursery teacher to quickly rush to Taro's aid and ensure his safety.
[0312] Prompt Sentence Examples
[0313] "High-resolution cameras and AI technology are used in daycare centers to track the behavior and emotions of children in real time, and facial recognition technology is used to identify each child individually. Please explain in detail how the AI will automatically generate daycare diaries, provide information to parents, monitor and alert for risky behavior, analyze growth, and provide advice to childcare workers."
[0314] The above is a specific embodiment for carrying out the present invention. This system allows childcare workers to reduce their workload, ensure the safety and emotional growth of children, and provide high-quality childcare.
[0315] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0316] Step 1:
[0317] Receiving video data
[0318] The server receives real-time video data of classrooms and playgrounds from high-resolution cameras. The input is the video stream from each camera, and the output is the received video data. Specifically, the server specifies the IP address of the camera and obtains the video stream using HTTP or RTSP protocol. This allows the server to monitor the current activities of the children in real time.
[0319] Step 2:
[0320] Identifying individual subjects through facial recognition
[0321] The server analyzes the received video data using facial recognition technology. The input is the video data and a database of pre-registered faces of the children, and the output is the ID of the identified child. Specifically, the server uses a facial recognition algorithm (e.g., FaceNet) to detect faces in the video data and compare them with the database. Based on the results of this comparison, the server identifies individual subjects (children).
[0322] Step 3:
[0323] Behavior tracking and data recording
[0324] The server tracks the behavior of identified children and records the data. The input is the child's ID and video data, and the output is a log of behavioral data. Specifically, the server uses a behavioral analysis algorithm to analyze the children's movements and location information and records it in chronological order. For example, a detailed behavioral log is generated, such as "Taro started playing in the sandbox at 9:15 and was talking to his classmate Hanako at 9:30."
[0325] Step 4:
[0326] Automatic generation of childcare diary
[0327] The server automatically generates a childcare diary based on the recorded behavioral data. The input is a log of behavioral data, and the output is a childcare diary. Specifically, the server organizes the behavioral logs based on a template and creates a diary-style report. This report details the children's daily activities, interactions, and health status.
[0328] Step 5:
[0329] Distribution of childcare journals
[0330] The server automatically distributes the generated childcare diary to the children's parents. The input is the childcare diary and the parents' contact information, and the output is the sent email or app notification. Specifically, the server uses an email sending API or push notification API to send the diary via the parents' email address or a dedicated app. This allows parents to keep track of their children's daily activities in real time.
[0331] Step 6:
[0332] Monitoring and alerting for risky behavior
[0333] The server analyzes video data in real time and monitors children's risky behavior. The input is video data, and the output is an alert notification. Specifically, the server uses a motion analysis algorithm to detect risky behavior (e.g., jumping from high places, rough play). If any risky behavior is detected, an alert notification is immediately pushed to the nursery teacher's device.
[0334] Step 7:
[0335] Providing growth analysis and advice
[0336] The server periodically compiles the data from the childcare diary and analyzes the growth of each child. The input is the childcare diary data, and the output is a growth analysis report and advice. Specifically, the server uses a machine learning algorithm to analyze past behavioral data and predict the child's growth trends. Based on this, it provides childcare workers with advice on appropriate childcare policies and activities.
[0337] Step 8:
[0338] Emotion analysis using an emotion engine
[0339] The server inputs video and audio data into an emotion engine to analyze the emotional state of the children. The input is video and audio data, and the output is the emotion analysis results. Specifically, the server calls an emotion analysis API (e.g., Microsoft Azure Emotion API) to analyze the children's facial expressions and tone of voice to identify their emotional state. The results of this analysis are reflected in the childcare diary and growth analysis, and are used as part of advice for childcare workers and parents.
[0340] (Application example 2)
[0341] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0342] There are challenges in monitoring worker safety and managing work efficiently in factories. In particular, it is difficult to grasp workers' behavior and emotional state in real time, quickly detect dangerous behavior, and take countermeasures. Furthermore, there are insufficient systems to provide appropriate advice based on workers' emotional state, which hinders workers' stress reduction and efficient work performance.
[0343] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0344] In this invention, the server includes means for capturing images of an area with a high-resolution camera, means for receiving the image data and identifying multiple individual subjects using facial recognition technology, means for tracking the behavior of the identified individuals and automatically generating records, means for automatically distributing the generated records to relevant parties, means for issuing warnings when dangerous behavior is detected, means for analyzing growth or progress based on the records and providing advice to the person in charge, and means for analyzing the emotional state of workers in real time using a portable display device worn by the workers and providing appropriate advice, thereby enabling safety monitoring of workers in factories and efficient work management.
[0345] A "high-resolution imaging device" is a high-resolution camera device that has a wide field of view and can capture detailed images.
[0346] "Area" refers to the space or location where a particular activity takes place, including, for example, a work area in a factory or an area requiring safety monitoring.
[0347] "Facial recognition technology" is a technology that can identify individual subjects by analyzing facial images in video and comparing them with a pre-registered database.
[0348] "Distinct objects" refer to specific people or objects that the system needs to identify, including, for example, factory workers.
[0349] "Behavioral tracking" refers to tracking the movements and actions of individual subjects in real time and recording that data.
[0350] "Automatic generation of records" refers to automatically creating diaries and reports based on tracked behavioral data.
[0351] "Stakeholders" refers to people who need the system's output data, including, for example, factory managers and the workers' families.
[0352] "Sending a warning" means that the system will send an alert in real time when it detects dangerous behavior or an unexpected situation.
[0353] "Analyzing growth or progress" refers to analyzing the progress or areas for improvement of an individual subject based on recorded data.
[0354] "Providing advice to the person in charge" refers to notifying the person in charge of specific instructions or suggestions based on the analysis results.
[0355] A "portable display device" is a device with a small display that can be worn and used by a worker.
[0356] "Analyzing emotional states" is a technology that identifies emotions from video and audio data and understands those states.
[0357] The system of this invention uses high-resolution imaging devices, servers, facial recognition technology, emotion analysis engines, and portable display devices (hereinafter referred to as smart glasses) to monitor the safety of workers within factories and achieve efficient business management.
[0358] Hardware and software used:
[0359] High-resolution camera equipment: A high-resolution camera equipment with a wide field of view that can capture detailed images.
[0360] Server: A cloud or local server that receives and analyzes data in real time.
[0361] Facial recognition technology: Technology that analyzes facial images in video and identifies individual subjects (e.g., Amazon Rekognition, Microsoft Azure Face API).
[0362] Emotion analysis engine: Technology that identifies emotions from video and audio data and understands their state (e.g., Affectiva, IBM Watson Tone Analyzer).
[0363] Portable display devices (smart glasses): Devices worn by workers that display information in real time (e.g., Google Glass, Microsoft HoloLens).
[0364] Real-time data analysis engine: Technology that performs real-time analysis of data (e.g., Apache Kafka, Apache Flink).
[0365] The server uses high-resolution camera equipment to capture images of the factory work area. The captured image data is sent to the server in real time. The server's facial recognition technology analyzes the image data and identifies multiple individual subjects (factory workers). The actions of the identified individuals (workers) are tracked by the server, and an activity record is automatically generated.
[0366] The recorded behavioral data is automatically distributed to relevant parties (such as factory managers) via email or a dedicated application. If the system detects any dangerous behavior, it will immediately issue a warning. Workers wearing the smart glasses will receive a warning message in real time, allowing them to respond quickly.
[0367] Furthermore, the server analyzes the progress of workers based on the recorded data and provides advice to managers, thereby improving work efficiency and safety. The emotion analysis engine analyzes the emotional state of workers in real time and provides appropriate advice. For example, if a worker is in a high-stress state, the smart glasses will display advice such as "You need to take a break."
[0368] As a concrete example, if a factory worker is wearing smart glasses while working at height, a camera will capture his movements and send the data to a server, where facial recognition technology will identify the worker and begin tracking his behavior and analyzing his emotions. If the worker makes an unstable movement, the server will determine this as a dangerous behavior and display a warning on the smart glasses saying, "Your current movement is dangerous. Please move to a safe location."
[0369] Prompt Sentence Examples
[0370] Design a system to monitor the safety of factory workers using facial recognition technology, emotion engines, and AI analysis technology. Specifically, include the following features:
[0371] 1. Facial recognition function to identify workers using smart glasses
[0372] 2. Tracking worker behavior and detecting dangerous behavior
[0373] 3. Monitor emotional states using a sentiment analysis engine
[0374] 4. Alert workers when they are in a dangerous situation
[0375] 5. Providing advice on safety and stress reduction
[0376] Specific examples of technologies to consider using include Amazon Rekognition, Affectiva, and Apache Flink.
[0377] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0378] Step 1:
[0379] High-resolution camera equipment captures video data of the factory work area, including the overall situation within the factory and the actions of workers.
[0380] Input: Factory area video
[0381] Output: High-resolution video data
[0382] Step 2:
[0383] The server receives the video data in real time, and then transmits the video data to the server, where the analysis process begins.
[0384] Input: High-resolution video data
[0385] Output: Video data required for analysis
[0386] Step 3:
[0387] The server uses facial recognition technology to identify workers in the video data by comparing facial images in the video with a pre-registered database to identify individual targets.
[0388] Input: Video data, face database
[0389] Output: ID of the identified individual (worker)
[0390] Step 4:
[0391] The server tracks the behavior of identified individuals, recording the movements and actions of workers in real time.
[0392] Input: ID of identified individual, video data
[0393] Output: Tracked behavior data
[0394] Step 5:
[0395] The server automatically generates records based on the collected behavioral data, including the worker's daily activities and tasks.
[0396] Input: Tracked behavioral data
[0397] Output: Automatically generated activity log
[0398] Step 6:
[0399] The server automatically distributes the generated records to the relevant parties via email or a dedicated application.
[0400] Input: Auto-generated event log, stakeholder list
[0401] Output: Records sent to interested parties
[0402] Step 7:
[0403] The server analyzes the video data in real time to detect dangerous behavior and issues a warning if any.
[0404] Input: Real-time video data
[0405] Output: Risky behavior warning
[0406] Step 8:
[0407] Workers wearing smart glasses receive real-time warning messages that are immediately displayed to workers, enabling them to take prompt action.
[0408] Input: Risky behavior warning message
[0409] Output: Real-time displayed warnings
[0410] Step 9:
[0411] The server analyzes progress based on the activity records and provides advice to the administrator, thereby improving work efficiency and safety.
[0412] Input: Action log
[0413] Output: Progress analysis and advice
[0414] Step 10:
[0415] The emotion analysis engine analyzes the emotional state of the worker in real time, and based on the analysis results, appropriate advice is generated and displayed on the smart glasses.
[0416] Input: Video data, audio data
[0417] Output: Emotional state analysis and advice
[0418] For example, if the server detects unstable movements of a worker working at height, it will display a warning on the smart glasses saying, "The current movement is dangerous. Please move to a safe location." This warning allows the worker to take immediate action to ensure safety.
[0419] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0420] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0421] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0422] [Second embodiment]
[0423] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0424] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0425] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0426] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0427] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0428] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0429] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0430] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0431] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0432] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0433] In the smart glasses 214, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0434] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0435] This invention is a system that reduces the burden on childcare workers and ensures safe, high-quality childcare. Using high-resolution cameras, servers, and terminals, the system utilizes face recognition technology and AI analysis technology to automatically generate childcare diaries, distribute them to parents, monitor and alert for risky behavior, analyze child development, and provide advice to childcare workers.
[0436] 1. Camera video capture
[0437] The server receives real-time video data from high-resolution cameras installed in classrooms and on the playground, which have a wide field of view and can accurately capture detailed activity within the school.
[0438] 2. Face recognition and individual object identification
[0439] The server analyzes the received video data using facial recognition AI technology to identify the children in the video. Facial recognition technology identifies the ID of each individual child by comparing the facial image in the video with a pre-registered database.
[0440] 3. Behavior tracking and automatic generation of childcare diary
[0441] The server tracks the behavior of identified children and records their behavioral patterns (e.g., types of play, interpersonal relationships, and health status). Based on this, a daycare diary is automatically generated. The diary details each child's daily activities, interactions, and health status.
[0442] 4. Distribution of the childcare diary to parents
[0443] The server automatically distributes the generated childcare diary to the parents of the corresponding child via their email address or a dedicated app, allowing parents to keep track of their child's daily activities in real time.
[0444] 5. Monitoring risky behavior and issuing alerts
[0445] The server analyzes the video data in real time and monitors the children's risky behavior. If any risky behavior (e.g., jumping from a high place or rough play) is detected, an alert is sent immediately.
[0446] The terminal (childcare worker) receives this alert and can respond quickly.
[0447] 6. Growth analysis and advice for childcare workers
[0448] The server periodically compiles data from the childcare diary and analyzes the growth of each child. Based on the results of this analysis, it identifies individual growth trends and provides advice to childcare workers regarding specific childcare policies and activities. The advice is automatically sent to the childcare workers' devices and can be used in practice.
[0449] For example, when Taro is playing in the playground, a camera captures his movements, and the server identifies him using facial recognition technology and records his actions. His behavior as he plays with building blocks with his friends is tracked, and a diary entry is made stating, "Taro played with building blocks with his friends and worked together to build a big tower." This diary entry is automatically sent to Taro's mother's email address that afternoon.
[0450] In addition, if Taro attempts dangerous behavior in the playground, such as standing up high on a swing, the server will detect this and send an alert to the nursery teacher's device, allowing the nursery teacher to immediately rush over to Taro and ensure his safety.
[0451] The above is a specific embodiment for carrying out the present invention. This system allows childcare workers to save time and effort, ensure the safety of children, and provide high-quality childcare.
[0452] The processing flow will be explained below.
[0453] Processing for automatic childcare diary generation system using face recognition
[0454] From video capture to automatic generation of childcare diary
[0455] Step 1: Receiving video data
[0456] The server receives video data in real time from high-resolution cameras installed in classrooms and playgrounds, and the video data is temporarily stored in dedicated storage.
[0457] Step 2: Face recognition and individual subject identification
[0458] The server inputs the received video data into a facial recognition AI model to identify the child in the video. The facial recognition technology compares the facial images in the video with a pre-registered database to identify the ID of each individual subject (child).
[0459] Step 3: Tracking behavior
[0460] The movements of identified individuals are tracked in real time, with the server recording the start and end of specific actions (e.g., playing, crafting, talking) with timestamps.
[0461] Step 4: Creating a childcare diary
[0462] The server automatically generates a daily diary for each child based on the tracking data, detailing each child's daily activities, interactions, and health status.
[0463] Step 5: Distributing the journal
[0464] The server automatically distributes the generated childcare diary to the parents of the corresponding children via their email address or a dedicated app.
[0465] Monitoring and alerting for risky behavior
[0466] Step 1: Real-time monitoring of video data
[0467] The server analyzes the video data in real time and monitors the behavior of the children.
[0468] Step 2: Detect risky behavior
[0469] The server uses a vision AI model to detect dangerous behavior (e.g., jumping from a high place, rough play).
[0470] Step 3: Trigger an alert
[0471] If dangerous behavior is detected, the server immediately generates an alert and sends it to the childcare worker's device.
[0472] Step 4: Receive and respond to alerts
[0473] The device (childcare worker) receives the alert and takes action to respond in real time.
[0474] Growth analysis and advice for childcare workers
[0475] Step 1: Collecting childcare diary data
[0476] The server periodically compiles data from the daycare diary and analyzes growth and behavioral patterns over a set period of time.
[0477] Step 2: Conduct a growth analysis
[0478] The server generates and evaluates the growth trends of each child based on the analysis data.
[0479] Step 3: Generate Advice
[0480] Based on growth trends, the server generates specific advice for nursery teachers, including specific activities and methods for promoting children's growth.
[0481] Step 4: Delivering advice
[0482] The server delivers the generated advice to the childcare worker's device, who receives it and puts it into practice.
[0483] This is the specific processing flow of the system, which reduces the burden on nursery teachers and ensures the safety and growth of children.
[0484] Example 1
[0485] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0486] In conventional childcare settings, childcare workers have to manually create childcare records, monitor children's behavior, and regularly provide information to parents. This increases the burden on childcare workers and can lead to a decline in the quality of childcare. It is also difficult to immediately detect risky behavior or provide specific advice based on growth analysis. Therefore, there was a need for a system that could reduce the burden on childcare workers and ensure safe, high-quality childcare.
[0487] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0488] In this invention, the server includes: means for capturing images of classrooms and playgrounds with a high-resolution camera; means for receiving the image data and identifying multiple individual objects using image authentication technology; means for tracking the behavior of the identified individual objects using computer vision technology and automatically generating a childcare record; means for automatically distributing the generated childcare record to parents; means for issuing a warning when dangerous behavior is detected; means for analyzing growth based on the childcare record and providing advice to childcare workers; means for processing the image data frame by frame and comparing it with multiple databases; means for tracking the positions and movements of children using computer vision technology; means for saving the tracking data in chronological order and performing activity mapping in real time; means for creating a childcare record in natural language using a generative AI model; means for automatically allocating contact information for each parent; and means for generating advice regarding childcare policies and activities based on the activity analysis results. This reduces the burden on childcare workers, improves safety, and enables the provision of high-quality childcare.
[0489] A "high-definition camera" is a camera that can capture wide-area images in high resolution.
[0490] "Video data" refers to information recorded in digital format from video captured by a camera.
[0491] "Image recognition technology" refers to algorithms and techniques used to identify individual subjects from video data.
[0492] "Individual subjects" refer to individual people, such as children, who participate in activities within the nursery school.
[0493] "Computer vision technology" is a technology that uses computers to analyze images and videos to recognize objects and track behavior.
[0494] A "childcare record" is a childcare diary that records in detail the daily activities and behavior of the children.
[0495] A "warning" is an alert that is issued when risky behavior is detected.
[0496] "Guardian" refers to the parent or legal guardian who is the caretaker of the child.
[0497] A "server" is a computer system that receives, analyzes, stores, and distributes video data.
[0498] "Tracking" means to continuously follow the movement of an object.
[0499] A "generative AI model" is an artificial intelligence model that generates sentences in natural language based on accumulated data.
[0500] "Contact information" refers to the parent's email address and account information for the dedicated application.
[0501] "Activity mapping" means visualizing the behavior and activities of kindergarten children along a timeline.
[0502] "Advice" means specific instructions or advice given to improve the quality of childcare.
[0503] "Database" means an electronic record containing information used to identify or match individuals.
[0504] This invention is a system that reduces the workload of childcare workers and provides safe, high-quality childcare. This system uses high-resolution image capture devices, servers, and terminals, and makes full use of image authentication and computer vision technologies to provide the following integrated functions.
[0505] The server acquires video data in real time from high-resolution camera devices installed in classrooms and playgrounds. These cameras are capable of covering a wide area in high resolution, and the video is streamed to the server using the RTSP protocol. The server stores this video data in a buffer frame by frame and begins processing.
[0506] The server then analyzes the received video data using image recognition technology. Specifically, it uses OpenCV and the Dlib library to detect the faces of the children and compare them with a pre-registered database to identify each individual child. This allows it to assign a unique ID to each child.
[0507] The behavior of identified children is tracked using computer vision technology on the server. Specifically, the OpenPose library is used to analyze the children's positions and movements. This data is saved in chronological order, and activity mapping is performed in real time. This allows for a detailed record of each child's activities.
[0508] Based on this tracking data, the server uses a generative AI model to automatically generate childcare records. The generated childcare records are detailed in natural language and include the child's daily activities, interactions, health status, etc. For example, tracking Taro playing with his friends using building blocks might result in a note in the childcare record saying, "Taro worked together with his friends to build a big tower."
[0509] The generated childcare records are automatically sent to the parents of the corresponding children via a server. The records are sent via email addresses registered in advance by the parents or via a dedicated childcare app. This allows parents to keep track of their child's daily activities in real time.
[0510] The server also analyzes video data in real time and monitors for dangerous behavior. For example, if the server detects a child engaging in dangerous behavior, such as jumping from a high place, it immediately issues an alert and notifies the childcare worker's device. This alert is sent as a push notification, allowing the childcare worker to respond quickly.
[0511] Furthermore, the server periodically compiles the childcare records and analyzes the growth of each child. Based on the results of this analysis, the generative AI model generates specific advice on childcare policies and activities, which are automatically distributed to the childcare worker's device. The childcare worker can use this advice to adjust and improve the childcare plan.
[0512] As a concrete example, when Taro is playing in the playground, a camera captures his movements, and the server uses facial recognition technology to identify him and record his actions. Computer vision technology tracks Taro playing with his friends using building blocks, and a generative AI model is used to create a childcare record stating, "Taro worked with his friends to build a big tower." This childcare record is automatically sent to Taro's mother. Furthermore, if the server detects Taro standing up high on the swing, it immediately sends an alert to the childcare worker's device, allowing the worker to respond quickly.
[0513] Prompt Sentence Examples
[0514] "How do you analyze real-time footage of children playing outside and send alerts to childcare workers if a particular child is exhibiting risky behavior?"
[0515] By utilizing this system, the workload of childcare workers will be significantly reduced, the safety of children will be improved, and high-quality childcare will be provided.
[0516] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0517] Step 1:
[0518] The server acquires video data in real time from high-resolution camera devices installed in classrooms and playgrounds. The input is the video stream from the camera, and the output is frame data stored on the server. Specifically, the server acquires data from the camera using the RTSP protocol and adds each frame to the processing queue.
[0519] Step 2:
[0520] The server analyzes the received video data using image recognition technology to detect the faces of the children. The input is the frame data from step 1, and the output is a feature vector of the detected face image. Specifically, it uses OpenCV and the Dlib library to detect faces in the frame and extract facial features.
[0521] Step 3:
[0522] The server uses facial recognition technology to compare the facial features detected with a pre-registered database to identify individual subjects (children). The input is the facial feature vector from step 2, and the output is the child's ID. Specifically, the server compares the facial feature vector with existing features in the database and identifies matching IDs based on cosine similarity.
[0523] Step 4:
[0524] The server uses computer vision technology to track the behavior of the identified children. The input is the child's ID from step 3 and the frame data from step 1, and the output is time-series data of the child's position and movement. Specific movements are tracked by detecting the child's skeleton using the OpenPose library.
[0525] Step 5:
[0526] The server automatically generates childcare records based on the tracking data using a generative AI model. The input is the time series data from step 4, and the output is a childcare record written in natural language. Specifically, the tracking data is input into the generative AI model, and a sentence describing the daily activities of the children is generated.
[0527] Step 6:
[0528] The server automatically distributes the generated childcare record to the corresponding guardian. The input is the childcare record from step 5, and the output is the information sent to the guardian's device. Specifically, the server refers to the guardian's contact database and sends the childcare record via email or a dedicated app.
[0529] Step 7:
[0530] The server analyzes the video data in real time and monitors for risky behavior. The input is the frame data from step 1, and the output is a warning notification when risky behavior is detected. Specifically, it uses computer vision technology to detect risky behavior patterns and sends a warning to the nursery teacher's device.
[0531] Step 8:
[0532] The server periodically compiles the childcare records and analyzes the children's growth. The input is the childcare record data from Step 5, and the output is the growth analysis results and advice. Specific operations include analyzing the compiled data and using a generative AI model to generate specific advice regarding childcare policies and activities.
[0533] Prompt Sentence Examples
[0534] "How do you analyze real-time footage of children playing outside and send alerts to childcare workers if a particular child is exhibiting risky behavior?"
[0535] This series of processes enables nursery school teachers to ensure the safety of children while saving time and effort, and to provide high-quality childcare.
[0536] (Application example 1)
[0537] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0538] In recent years, improving work efficiency and ensuring safety in factories have become extremely important issues. In particular, in manufacturing sites where many workers work simultaneously, there is a high risk of human error and dangerous behavior, and there is a growing demand for systems that can quickly detect and address such errors. Furthermore, manually creating work logs requires time and effort, making it difficult to immediately improve work efficiency. The present invention addresses these issues by providing advanced monitoring and automated work log generation to improve work efficiency and ensure safety in factories.
[0539] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0540] In this invention, the server includes means for capturing images of the work environment with a high-resolution camera, means for receiving the image data and identifying multiple individual objects using facial recognition technology, means for tracking the behavior of the identified individual objects and automatically generating a work log, means for automatically distributing the generated work log to relevant parties, means for issuing an alert when dangerous behavior is detected, and means for analyzing work efficiency based on the work log and providing relevant parties with points for improvement. This makes it possible to immediately grasp and analyze work efficiency within a factory and improve safety.
[0541] A "high-resolution camera" is a photographic device with high resolution that can record the work environment in detail.
[0542] "Work environment" refers to the place or area within a factory where work is carried out.
[0543] "Video data" is digital data containing visual information captured by a high-resolution camera.
[0544] "Facial recognition technology" is a biometric authentication technology for identifying people in video data.
[0545] "Individual Subject" refers to each identified worker within a factory.
[0546] "Behavior tracking" refers to following and recording the movements of identified individual subjects.
[0547] A "work log" is a digital record that records the actions and work of an individual.
[0548] "Automatically generating a work log" means that the server automatically creates a work log from the behavioral data of an individual subject.
[0549] "Stakeholders" refers to factory managers and supervisors who require work logs.
[0550] "Distribution" means sending the generated work log to the relevant parties.
[0551] "Issuing an alert" means sending a warning message when dangerous behavior is detected.
[0552] "Analyzing work efficiency" means analyzing the data in the work log to evaluate the efficiency of the work.
[0553] "Providing improvements" means presenting specific suggestions to stakeholders to improve work efficiency.
[0554] This invention is a system that improves work efficiency and ensures safety in factories by utilizing high-resolution cameras, facial recognition technology, and AI analysis technology. The system consists of a server, a camera, smart glasses, and a dedicated app.
[0555] System configuration
[0556] 1. High-quality camera:
[0557] The working environment within the factory is photographed in high resolution, and detailed video data is sent to a server in real time.
[0558] 2. Server:
[0559] The server processes the received video data and identifies the worker using facial recognition technology. Azure Face API, for example, is used for facial recognition. The actions of identified workers are tracked, and a work log is automatically generated based on that data. The work log contains detailed records of each worker's movements and tasks.
[0560] 3. Risky behavior monitoring and alerting:
[0561] The server analyzes video data and work logs in real time and immediately issues an alert if any dangerous behavior is detected. The alert is then sent to the smart glasses to alert the worker, enabling a rapid response.
[0562] 4. Work log distribution:
[0563] The generated work logs are automatically distributed to relevant parties (factory managers and supervisors) via email addresses or a dedicated app.
[0564] 5. Analyzing work efficiency and providing improvements:
[0565] The server analyzes work efficiency based on the work logs and provides specific improvements to the relevant parties, thereby optimizing work efficiency throughout the factory.
[0566] Hardware and software used
[0567] Hardware: high-resolution cameras, smart glasses (e.g., Microsoft HoloLens)
[0568] Software: Facial recognition AI (e.g., Azure Face API), behavioral analysis AI (e.g., TensorFlow, OpenCV), database (e.g., Azure SQL Database)
[0569] Example of operation
[0570] For example, if a worker is operating heavy machinery in an inappropriate manner, a high-resolution camera captures this, the server uses facial recognition technology to identify the worker, and AI determines that the operation is inappropriate. An alert is immediately sent to the smart glasses, prompting the worker to operate in a safe manner. At the same time, details of the operation are automatically generated as a work log and sent to the factory manager. Based on this work log, areas for improving work efficiency are analyzed and provided to relevant parties.
[0571] Example of input prompt for generative AI model
[0572] prompt:
[0573] Design a smart glasses application for factory monitoring with self-diagnosis function. Please pay attention to the following points:
[0574] 1. Tracking worker behavior in real time using cameras and smart glasses.
[0575] 2. Use facial recognition technology to identify individual workers.
[0576] 3. If risky behavior is detected, an alert is sent immediately.
[0577] 4. Automatically generate work logs and analyze areas for improvement.
[0578] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0579] Step 1:
[0580] High-resolution cameras capture the factory environment. The input is real-time footage from within the factory, and the output is high-resolution video data. Specifically, the cameras have a field of view that covers the entire factory, capturing detailed images of workers and machine operations.
[0581] Step 2:
[0582] The server receives the video data sent from the camera. The input is high-definition video data that is temporarily stored in the server. The output is video data stored in the server's storage.
[0583] Step 3:
[0584] The server analyzes the video data using facial recognition AI (Azure Face API) to identify multiple workers. The input is the saved video data, and the output is the identification information of each worker. Specifically, the AI analyzes the facial images in the video and identifies the worker's ID by comparing them with a pre-registered database.
[0585] Step 4:
[0586] The server tracks the actions of each identified worker. The input is identification information and real-time video data, and the output is individual behavioral patterns. Specifically, the server continuously tracks the movements of workers and records what tasks they are performing.
[0587] Step 5:
[0588] The server automatically generates work logs based on the behavioral data it tracks. The input is behavioral pattern data, and the output is a digital work log. Specifically, the AI organizes the recorded information and provides a detailed record of each worker's daily activities and their content.
[0589] Step 6:
[0590] The server automatically distributes the generated work log to the relevant parties. The input is the generated work log, and the output is the work log sent to the relevant parties' email addresses or a dedicated app. Specifically, the log file is sent to the specified recipient via a mail server or application.
[0591] Step 7:
[0592] The server analyzes real-time video data and issues an alert if it detects dangerous behavior. The input is real-time video data, and the output is an alert message sent to the smart glasses. Specifically, when the AI detects an abnormal behavior pattern, it immediately generates an alert and sends it to the worker's smart glasses.
[0593] Step 8:
[0594] The server analyzes work efficiency based on work logs and provides suggestions for improvement to the relevant parties. The input is daily work log data, and the output is an analysis report including suggestions for improvement. Specifically, the server compares the data with past work data, proposes specific improvements for improving work efficiency and safety, and sends them to the relevant parties.
[0595] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0596] This invention is a system that reduces the burden on childcare workers and ensures safe, high-quality childcare. Using high-resolution cameras, servers, and terminals, the system utilizes face recognition technology, emotion engines, and AI analysis technology to automatically generate childcare diaries, distribute them to parents, monitor and alert on risky behavior, analyze child development, and provide advice to childcare workers.
[0597] 1. Camera video capture
[0598] The server receives real-time video data from high-resolution cameras installed in classrooms and on the playground, which have a wide field of view and can accurately capture detailed activity within the school.
[0599] 2. Face recognition and individual object identification
[0600] The server analyzes the received video data using facial recognition AI technology to identify the children in the video. The facial recognition technology compares the facial images in the video with a pre-registered database to identify the ID of each individual subject (child).
[0601] 3. Behavior tracking and automatic generation of childcare diary
[0602] The server tracks the behavior of identified children and records their behavioral patterns (e.g., types of play, interpersonal relationships, and health status). Based on this, a daycare diary is automatically generated. The diary details each child's daily activities, interactions, and health status.
[0603] 4. Distribution of the childcare diary to parents
[0604] The server automatically distributes the generated childcare diary to the parents of the corresponding child via their email address or a dedicated app, allowing parents to keep track of their child's daily activities in real time.
[0605] 5. Monitoring risky behavior and issuing alerts
[0606] The server analyzes the video data in real time and monitors the children's risky behavior. If any risky behavior (e.g., jumping from a high place or rough play) is detected, an alert is sent immediately.
[0607] The terminal (childcare worker) receives this alert and can respond quickly.
[0608] 6. Growth analysis and advice for childcare workers
[0609] The server periodically compiles data from the childcare diary and analyzes the growth of each child. Based on the results of this analysis, it identifies individual growth trends and provides advice to childcare workers regarding specific childcare policies and activities. The advice is automatically sent to the childcare workers' devices and can be used in practice.
[0610] 7. Introducing the Emotion Engine
[0611] The server inputs the video and audio data into the emotion engine, which analyzes the child's emotional state based on their facial expressions and tone of voice. The emotion engine then identifies the emotion the child is feeling (e.g., joy, sadness, anger, anxiety).
[0612] Based on the emotional states identified by the emotion engine, the server adds emotional data to the childcare diary and development analysis, for example, recording how children felt during play and how specific activities affected their emotions.
[0613] The server generates and delivers advice to childcare workers and parents based on their emotional state, suggesting appropriate ways to respond. For example, if a child is feeling anxious, the server identifies the root cause and provides specific advice on how to support the child.
[0614] For example, when Taro is playing in the playground, a camera captures his movements, and the server identifies him using facial recognition technology and records his actions. His behavior as he plays with building blocks with his friends is tracked, and a diary entry is made stating, "Taro played with building blocks with his friends and worked together to build a big tower." This diary entry is automatically sent to Taro's mother's email address that afternoon.
[0615] Furthermore, if Taro attempts dangerous behavior such as standing up too high on the swing, the server will detect this and immediately send an alert to the nursery teacher's device, allowing the nursery teacher to immediately rush over to Taro and ensure his safety.
[0616] Furthermore, the emotion engine analyzes Taro's facial expressions and tone of voice and detects that he looks anxious. Based on this, the server sends advice to the nursery teacher's device, such as, "Taro is likely feeling anxious, so please talk to him a bit to reassure him."
[0617] The above is a specific embodiment for carrying out the present invention. This system allows childcare workers to reduce their workload, ensure the safety and emotional growth of children, and provide high-quality childcare.
[0618] The processing flow will be explained below.
[0619] Processing for an automatic childcare diary generation system that combines face recognition and emotion engine
[0620] From video capture to emotion analysis
[0621] Step 1: Receiving video data
[0622] The server receives video data in real time from high-resolution cameras installed in classrooms and playgrounds, and the video data is temporarily stored in dedicated storage.
[0623] Step 2: Face recognition and individual subject identification
[0624] The server inputs the received video data into a facial recognition AI model to identify the children in the video. The facial recognition technology compares the facial images in the video with a pre-registered database to identify the ID of each individual subject (child).
[0625] Step 3: Tracking behavior
[0626] The movements of identified individuals are tracked in real time, with the server recording the start and end of specific actions (e.g., playing, crafting, talking) with timestamps.
[0627] Step 4: Analysis by Emotion Engine
[0628] The server inputs video and audio data into an emotion engine, which analyzes the children's facial expressions and tone of voice to determine their emotional state (e.g., joy, sadness, anger, anxiety).
[0629] Step 5: Creating a childcare diary
[0630] The server automatically generates a daily diary for each child based on tracking data and emotion analysis data, detailing each child's daily activities, interactions, health status, and emotional state.
[0631] Delivery of childcare diary to parents
[0632] Step 1: Generate a journal
[0633] The server automatically distributes the generated childcare diary to the parents of the corresponding children via their email addresses or a dedicated app.
[0634] Monitoring and alerting for risky behavior
[0635] Step 1: Real-time monitoring of video data
[0636] The server analyzes the video data in real time and monitors the behavior of the children.
[0637] Step 2: Detect risky behavior
[0638] The server uses a vision AI model to detect dangerous behavior (e.g., jumping from a high place, rough play).
[0639] Step 3: Trigger an alert
[0640] If dangerous behavior is detected, the server immediately generates an alert and sends it to the childcare worker's device.
[0641] Step 4: Receive and respond to alerts
[0642] The device (childcare worker) receives the alert and takes action to respond in real time.
[0643] Growth analysis and advice for childcare workers
[0644] Step 1: Collecting childcare diary data
[0645] The server periodically compiles data from the daycare diary and analyzes growth and behavioral patterns over a set period of time.
[0646] Step 2: Conduct a growth analysis
[0647] The server generates and evaluates the growth trends of each child based on the analysis data.
[0648] Step 3: Generate Advice
[0649] Based on growth trends, the server generates specific advice for nursery teachers, including specific activities and methods for promoting children's growth.
[0650] Step 4: Delivering advice
[0651] The server delivers the generated advice to the childcare worker's device, who receives it and puts it into practice.
[0652] Further processing by the emotion engine
[0653] Step 1: Collecting emotion data
[0654] The server integrates the analytical data from the emotion engine into the childcare diary, allowing for an understanding of emotional fluctuations and trends.
[0655] Step 2: Generating sentiment-based advice
[0656] Based on the emotional data, the server generates advice suggesting appropriate responses for childcare workers and parents.
[0657] Step 3: Delivering emotional advice
[0658] The server delivers advice based on the generated emotions to the childcare worker's device, who then receives it and puts it into practice.
[0659] For example, when Taro is playing in the playground, a camera captures his movements, and the server identifies him using facial recognition technology and records his actions. His behavior as he plays with building blocks with his friends is tracked, and a diary entry is made stating, "Taro played with building blocks with his friends and worked together to build a big tower." This diary entry is automatically sent to Taro's mother's email address that afternoon.
[0660] Furthermore, if Taro attempts dangerous behavior such as standing up too high on the swing, the server will detect this and immediately send an alert to the nursery teacher's device, allowing the nursery teacher to immediately rush over to Taro and ensure his safety.
[0661] Furthermore, the emotion engine analyzes Taro's facial expressions and tone of voice and detects that he looks anxious. Based on this, the server sends advice to the nursery teacher's device, such as, "Taro is likely feeling anxious, so please talk to him a bit to reassure him."
[0662] The above is a specific embodiment for carrying out the present invention. This system allows childcare workers to reduce their workload, ensure the safety and emotional growth of children, and provide high-quality childcare.
[0663] Example 2
[0664] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0665] In childcare facilities, the workload of childcare workers is increasing, and it is often the case that they are not doing enough to ensure the safety of children, monitor their growth, or report to parents. It is also difficult to accurately grasp the emotional state of children and respond appropriately. This puts the quality of childcare at risk of declining.
[0666] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0667] In this invention, the server includes means for capturing images of classrooms and playgrounds with a high-resolution camera, means for receiving the image data and identifying multiple individual objects using facial recognition technology, means for tracking the behavior of the identified individual objects and automatically generating a childcare diary, means for automatically distributing the generated childcare diary to parents, means for issuing an alert when dangerous behavior is detected, means for analyzing growth based on the childcare diary and providing advice to childcare workers, means for analyzing emotional states from video data and audio data using an emotion engine, and means for suggesting appropriate response methods to childcare workers and parents based on the analyzed emotional states. This reduces the burden on childcare workers, enables management of the safety and growth of children, and enables understanding of emotional states and appropriate responses.
[0668] A "high-resolution camera" is a camera that has a wide field of view and can capture detailed video data in high resolution.
[0669] "Video data" refers to video information acquired from cameras installed in classrooms and playgrounds.
[0670] "Facial recognition technology" is a technology that identifies individuals by identifying faces in video data and comparing them with a pre-registered database.
[0671] "Individual object" is a term that refers to a specific person, especially a kindergartener.
[0672] "Behavioral tracking" is a method of monitoring and recording the movements and behavioral patterns of individual subjects in real time.
[0673] A "childcare diary" is a report that details the behavior, interactions, and health status of the children during a day's childcare activities.
[0674] "Guardian" is a term that refers to the parents or guardians of kindergarten children.
[0675] An "alert" is a warning message that is sent immediately when dangerous behavior or an abnormality is detected in real time.
[0676] "Growth analysis" is a method of analyzing children's growth patterns based on data from the nursery school diary and evaluating developmental trends.
[0677] "Advice" refers to specific methods of response and guidelines provided to childcare workers and parents based on growth analysis and emotional analysis.
[0678] An "emotion engine" is a technology that analyzes video and audio data to identify the emotional state of an individual subject.
[0679] "Response methods" are specific measures that, based on the analyzed data, indicate appropriate approaches in caring for and guiding children.
[0680] This invention is a system that reduces the burden on childcare workers and ensures safe, high-quality childcare. This system uses high-resolution cameras, servers, and terminals, and makes full use of face recognition technology, emotion engines, and AI analysis technology to automatically generate childcare diaries, distribute them to parents, monitor and alert on risky behavior, analyze child development, and provide advice to childcare workers.
[0681] Hardware and software used
[0682] server
[0683] Data servers with high processing power (e.g., Dell PowerEdge series)
[0684] Facial recognition AI technology (e.g., FaceNet)
[0685] Emotion engine (e.g. Microsoft Azure Emotion API)
[0686] Machine learning algorithms (e.g., Python's scikit-learn library)
[0687] camera
[0688] High-resolution network camera (e.g. AXIS P1368-E)
[0689] Terminal
[0690] Tablet or smartphone for childcare workers (e.g. iPad, Android tablet)
[0691] Explanation of program processing
[0692] High-definition camera video capture
[0693] The server receives real-time video data from high-resolution cameras installed in classrooms and playgrounds. These cameras have a wide field of view and can accurately capture detailed activity within the school. The video data is stored in the server's storage and used for analysis.
[0694] Identifying individual subjects using facial recognition technology
[0695] The server analyzes the received video data using facial recognition AI technology to identify the children in the video. The facial recognition technology compares the facial images in the video with a pre-registered database to identify the ID of each individual subject (child). This allows for accurate tracking of each child's activities.
[0696] Behavior tracking and automatic generation of childcare diary
[0697] The server tracks the identified children's behavior in real time and records their behavioral patterns (e.g., types of play, interpersonal relationships, and health status). Based on this behavioral data, a daycare diary is automatically generated, detailing each child's daily activities, interactions, and health status.
[0698] Delivery of childcare diary to parents
[0699] The generated daycare diary is automatically sent to the parents of the children via their email address or a dedicated app (e.g., Custom Nursery App). This allows parents to keep track of their children's daily activities in real time.
[0700] Monitoring risky behavior and issuing alerts
[0701] The server analyzes the video data in real time and monitors the children's risky behavior. For example, if dangerous behavior such as jumping from a high place or rough play is detected, an alert is immediately issued. A push notification is sent to the childcare worker's device (e.g., tablet), allowing the childcare worker to respond quickly.
[0702] Growth analysis and advice for childcare workers
[0703] The server periodically compiles data from the childcare diary and analyzes the growth of each child. Based on the results of this analysis, it identifies individual growth trends and provides advice to childcare workers regarding specific childcare policies and activities. This advice is automatically sent to devices and can be used by childcare workers in their daily work.
[0704] Introducing the Emotion Engine
[0705] The server inputs video and audio data into an emotion engine, which analyzes the emotional state of the child based on facial expressions, tone of voice, etc. The emotion engine identifies the child's emotions (e.g., joy, sadness, anger, anxiety). This emotional state is reflected in the childcare diary and growth analysis, and is used as part of advice for childcare workers and parents. Specifically, the server calls an emotion analysis API to obtain the analysis results, and a recommendation algorithm then suggests appropriate responses based on these results.
[0706] Specific examples
[0707] While Taro is playing in the playground, a high-definition camera captures his movements, and the server uses facial recognition technology to identify him and record his actions. His behavior as he plays with building blocks with his friends is tracked, and an entry is made in the daycare diary, stating, "Taro played with building blocks with his friends and worked together to build a big tower." This diary is automatically sent to Taro's mother's email address that afternoon. Furthermore, if Taro attempts the dangerous act of standing up high on the swing, the server detects this and immediately sends an alert to the nursery teacher's device, allowing the nursery teacher to quickly rush to Taro's aid and ensure his safety.
[0708] Prompt Sentence Examples
[0709] "High-resolution cameras and AI technology are used in daycare centers to track the behavior and emotions of children in real time, and facial recognition technology is used to identify each child individually. Please explain in detail how the AI will automatically generate daycare diaries, provide information to parents, monitor and alert for risky behavior, analyze growth, and provide advice to childcare workers."
[0710] The above is a specific embodiment for carrying out the present invention. This system allows childcare workers to reduce their workload, ensure the safety and emotional growth of children, and provide high-quality childcare.
[0711] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0712] Step 1:
[0713] Receiving video data
[0714] The server receives real-time video data of classrooms and playgrounds from high-resolution cameras. The input is the video stream from each camera, and the output is the received video data. Specifically, the server specifies the IP address of the camera and obtains the video stream using HTTP or RTSP protocol. This allows the server to monitor the current activities of the children in real time.
[0715] Step 2:
[0716] Identifying individual subjects through facial recognition
[0717] The server analyzes the received video data using facial recognition technology. The input is the video data and a database of pre-registered faces of the children, and the output is the ID of the identified child. Specifically, the server uses a facial recognition algorithm (e.g., FaceNet) to detect faces in the video data and compare them with the database. Based on the results of this comparison, the server identifies individual subjects (children).
[0718] Step 3:
[0719] Behavior tracking and data recording
[0720] The server tracks the behavior of identified children and records the data. The input is the child's ID and video data, and the output is a log of behavioral data. Specifically, the server uses a behavioral analysis algorithm to analyze the children's movements and location information and records it in chronological order. For example, a detailed behavioral log is generated, such as "Taro started playing in the sandbox at 9:15 and was talking to his classmate Hanako at 9:30."
[0721] Step 4:
[0722] Automatic generation of childcare diary
[0723] The server automatically generates a childcare diary based on the recorded behavioral data. The input is a log of behavioral data, and the output is a childcare diary. Specifically, the server organizes the behavioral logs based on a template and creates a diary-style report. This report details the children's daily activities, interactions, and health status.
[0724] Step 5:
[0725] Distribution of childcare journals
[0726] The server automatically distributes the generated childcare diary to the children's parents. The input is the childcare diary and the parents' contact information, and the output is the sent email or app notification. Specifically, the server uses an email sending API or push notification API to send the diary via the parents' email address or a dedicated app. This allows parents to keep track of their children's daily activities in real time.
[0727] Step 6:
[0728] Monitoring and alerting for risky behavior
[0729] The server analyzes video data in real time and monitors children's risky behavior. The input is video data, and the output is an alert notification. Specifically, the server uses a motion analysis algorithm to detect risky behavior (e.g., jumping from high places, rough play). If any risky behavior is detected, an alert notification is immediately pushed to the nursery teacher's device.
[0730] Step 7:
[0731] Providing growth analysis and advice
[0732] The server periodically compiles the data from the childcare diary and analyzes the growth of each child. The input is the childcare diary data, and the output is a growth analysis report and advice. Specifically, the server uses a machine learning algorithm to analyze past behavioral data and predict the child's growth trends. Based on this, it provides childcare workers with advice on appropriate childcare policies and activities.
[0733] Step 8:
[0734] Emotion analysis using an emotion engine
[0735] The server inputs video and audio data into an emotion engine to analyze the emotional state of the children. The input is video and audio data, and the output is the emotion analysis results. Specifically, the server calls an emotion analysis API (e.g., Microsoft Azure Emotion API) to analyze the children's facial expressions and tone of voice to identify their emotional state. The results of this analysis are reflected in the childcare diary and growth analysis, and are used as part of advice for childcare workers and parents.
[0736] (Application example 2)
[0737] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0738] There are challenges in monitoring worker safety and managing work efficiently in factories. In particular, it is difficult to grasp workers' behavior and emotional state in real time, quickly detect dangerous behavior, and take countermeasures. Furthermore, there are insufficient systems to provide appropriate advice based on workers' emotional state, which hinders workers' stress reduction and efficient work performance.
[0739] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0740] In this invention, the server includes means for capturing images of an area with a high-resolution camera, means for receiving the image data and identifying multiple individual subjects using facial recognition technology, means for tracking the behavior of the identified individuals and automatically generating records, means for automatically distributing the generated records to relevant parties, means for issuing warnings when dangerous behavior is detected, means for analyzing growth or progress based on the records and providing advice to the person in charge, and means for analyzing the emotional state of workers in real time using a portable display device worn by the workers and providing appropriate advice, thereby enabling safety monitoring of workers in factories and efficient work management.
[0741] A "high-resolution imaging device" is a high-resolution camera device that has a wide field of view and can capture detailed images.
[0742] "Area" refers to the space or location where a particular activity takes place, including, for example, a work area in a factory or an area requiring safety monitoring.
[0743] "Facial recognition technology" is a technology that can identify individual subjects by analyzing facial images in video and comparing them with a pre-registered database.
[0744] "Distinct objects" refer to specific people or objects that the system needs to identify, including, for example, factory workers.
[0745] "Behavioral tracking" refers to tracking the movements and actions of individual subjects in real time and recording that data.
[0746] "Automatic generation of records" refers to automatically creating diaries and reports based on tracked behavioral data.
[0747] "Stakeholders" refers to people who need the system's output data, including, for example, factory managers and the workers' families.
[0748] "Sending a warning" means that the system will send an alert in real time when it detects dangerous behavior or an unexpected situation.
[0749] "Analyzing growth or progress" refers to analyzing the progress or areas for improvement of an individual subject based on recorded data.
[0750] "Providing advice to the person in charge" refers to notifying the person in charge of specific instructions or suggestions based on the analysis results.
[0751] A "portable display device" is a device with a small display that can be worn and used by a worker.
[0752] "Analyzing emotional states" is a technology that identifies emotions from video and audio data and understands those states.
[0753] The system of this invention uses high-resolution imaging devices, servers, facial recognition technology, emotion analysis engines, and portable display devices (hereinafter referred to as smart glasses) to monitor the safety of workers within factories and achieve efficient business management.
[0754] Hardware and software used:
[0755] High-resolution camera equipment: A high-resolution camera equipment with a wide field of view that can capture detailed images.
[0756] Server: A cloud or local server that receives and analyzes data in real time.
[0757] Facial recognition technology: Technology that analyzes facial images in video and identifies individual subjects (e.g., Amazon Rekognition, Microsoft Azure Face API).
[0758] Emotion analysis engine: Technology that identifies emotions from video and audio data and understands their state (e.g., Affectiva, IBM Watson Tone Analyzer).
[0759] Portable display devices (smart glasses): Devices worn by workers that display information in real time (e.g., Google Glass, Microsoft HoloLens).
[0760] Real-time data analysis engine: Technology that performs real-time analysis of data (e.g., Apache Kafka, Apache Flink).
[0761] The server uses high-resolution camera equipment to capture images of the factory work area. The captured image data is sent to the server in real time. The server's facial recognition technology analyzes the image data and identifies multiple individual subjects (factory workers). The actions of the identified individuals (workers) are tracked by the server, and an activity record is automatically generated.
[0762] The recorded behavioral data is automatically distributed to relevant parties (such as factory managers) via email or a dedicated application. If the system detects any dangerous behavior, it will immediately issue a warning. Workers wearing the smart glasses will receive a warning message in real time, allowing them to respond quickly.
[0763] Furthermore, the server analyzes the progress of workers based on the recorded data and provides advice to managers, thereby improving work efficiency and safety. The emotion analysis engine analyzes the emotional state of workers in real time and provides appropriate advice. For example, if a worker is in a high-stress state, the smart glasses will display advice such as "You need to take a break."
[0764] As a concrete example, if a factory worker is wearing smart glasses while working at height, a camera will capture his movements and send the data to a server, where facial recognition technology will identify the worker and begin tracking his behavior and analyzing his emotions. If the worker makes an unstable movement, the server will determine this as a dangerous behavior and display a warning on the smart glasses saying, "Your current movement is dangerous. Please move to a safe location."
[0765] Prompt Sentence Examples
[0766] Design a system to monitor the safety of factory workers using facial recognition technology, emotion engines, and AI analysis technology. Specifically, include the following features:
[0767] 1. Facial recognition function to identify workers using smart glasses
[0768] 2. Tracking worker behavior and detecting dangerous behavior
[0769] 3. Monitor emotional states using a sentiment analysis engine
[0770] 4. Alert workers when they are in a dangerous situation
[0771] 5. Providing advice on safety and stress reduction
[0772] Specific examples of technologies to consider using include Amazon Rekognition, Affectiva, and Apache Flink.
[0773] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0774] Step 1:
[0775] High-resolution camera equipment captures video data of the factory work area, including the overall situation within the factory and the actions of workers.
[0776] Input: Factory area video
[0777] Output: High-resolution video data
[0778] Step 2:
[0779] The server receives the video data in real time, and then transmits the video data to the server, where the analysis process begins.
[0780] Input: High-resolution video data
[0781] Output: Video data required for analysis
[0782] Step 3:
[0783] The server uses facial recognition technology to identify workers in the video data by comparing facial images in the video with a pre-registered database to identify individual targets.
[0784] Input: Video data, face database
[0785] Output: ID of the identified individual (worker)
[0786] Step 4:
[0787] The server tracks the behavior of identified individuals, recording the movements and actions of workers in real time.
[0788] Input: ID of identified individual, video data
[0789] Output: Tracked behavior data
[0790] Step 5:
[0791] The server automatically generates records based on the collected behavioral data, including the worker's daily activities and tasks.
[0792] Input: Tracked behavioral data
[0793] Output: Automatically generated activity log
[0794] Step 6:
[0795] The server automatically distributes the generated records to the relevant parties via email or a dedicated application.
[0796] Input: Auto-generated event log, stakeholder list
[0797] Output: Records sent to interested parties
[0798] Step 7:
[0799] The server analyzes the video data in real time to detect dangerous behavior and issues a warning if any.
[0800] Input: Real-time video data
[0801] Output: Risky behavior warning
[0802] Step 8:
[0803] Workers wearing smart glasses receive real-time warning messages that are immediately displayed to workers, enabling them to take prompt action.
[0804] Input: Risky behavior warning message
[0805] Output: Real-time displayed warnings
[0806] Step 9:
[0807] The server analyzes progress based on the activity records and provides advice to the administrator, thereby improving work efficiency and safety.
[0808] Input: Action log
[0809] Output: Progress analysis and advice
[0810] Step 10:
[0811] The emotion analysis engine analyzes the emotional state of the worker in real time, and based on the analysis results, appropriate advice is generated and displayed on the smart glasses.
[0812] Input: Video data, audio data
[0813] Output: Emotional state analysis and advice
[0814] For example, if the server detects unstable movements of a worker working at height, it will display a warning on the smart glasses saying, "The current movement is dangerous. Please move to a safe location." This warning allows the worker to take immediate action to ensure safety.
[0815] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0816] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0817] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0818] [Third embodiment]
[0819] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0820] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0821] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0822] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0823] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0824] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0825] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0826] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0827] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0828] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0829] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0830] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0831] This invention is a system that reduces the burden on childcare workers and ensures safe, high-quality childcare. Using high-resolution cameras, servers, and terminals, the system utilizes face recognition technology and AI analysis technology to automatically generate childcare diaries, distribute them to parents, monitor and alert for risky behavior, analyze child development, and provide advice to childcare workers.
[0832] 1. Camera video capture
[0833] The server receives real-time video data from high-resolution cameras installed in classrooms and on the playground, which have a wide field of view and can accurately capture detailed activity within the school.
[0834] 2. Face recognition and individual object identification
[0835] The server analyzes the received video data using facial recognition AI technology to identify the children in the video. Facial recognition technology identifies the ID of each individual child by comparing the facial image in the video with a pre-registered database.
[0836] 3. Behavior tracking and automatic generation of childcare diary
[0837] The server tracks the behavior of identified children and records their behavioral patterns (e.g., types of play, interpersonal relationships, and health status). Based on this, a daycare diary is automatically generated. The diary details each child's daily activities, interactions, and health status.
[0838] 4. Distribution of the childcare diary to parents
[0839] The server automatically distributes the generated childcare diary to the parents of the corresponding child via their email address or a dedicated app, allowing parents to keep track of their child's daily activities in real time.
[0840] 5. Monitoring risky behavior and issuing alerts
[0841] The server analyzes the video data in real time and monitors the children's risky behavior. If any risky behavior (e.g., jumping from a high place or rough play) is detected, an alert is sent immediately.
[0842] The terminal (childcare worker) receives this alert and can respond quickly.
[0843] 6. Growth analysis and advice for childcare workers
[0844] The server periodically compiles data from the childcare diary and analyzes the growth of each child. Based on the results of this analysis, it identifies individual growth trends and provides advice to childcare workers regarding specific childcare policies and activities. The advice is automatically sent to the childcare workers' devices and can be used in practice.
[0845] For example, when Taro is playing in the playground, a camera captures his movements, and the server identifies him using facial recognition technology and records his actions. His behavior as he plays with building blocks with his friends is tracked, and a diary entry is made stating, "Taro played with building blocks with his friends and worked together to build a big tower." This diary entry is automatically sent to Taro's mother's email address that afternoon.
[0846] In addition, if Taro attempts dangerous behavior in the playground, such as standing up high on a swing, the server will detect this and send an alert to the nursery teacher's device, allowing the nursery teacher to immediately rush over to Taro and ensure his safety.
[0847] The above is a specific embodiment for carrying out the present invention. This system allows childcare workers to save time and effort, ensure the safety of children, and provide high-quality childcare.
[0848] The processing flow will be explained below.
[0849] Processing for automatic childcare diary generation system using face recognition
[0850] From video capture to automatic generation of childcare diary
[0851] Step 1: Receiving video data
[0852] The server receives video data in real time from high-resolution cameras installed in classrooms and playgrounds, and the video data is temporarily stored in dedicated storage.
[0853] Step 2: Face recognition and individual subject identification
[0854] The server inputs the received video data into a facial recognition AI model to identify the child in the video. The facial recognition technology compares the facial images in the video with a pre-registered database to identify the ID of each individual subject (child).
[0855] Step 3: Tracking behavior
[0856] The movements of identified individuals are tracked in real time, with the server recording the start and end of specific actions (e.g., playing, crafting, talking) with timestamps.
[0857] Step 4: Creating a childcare diary
[0858] The server automatically generates a daily diary for each child based on the tracking data, detailing each child's daily activities, interactions, and health status.
[0859] Step 5: Distributing the journal
[0860] The server automatically distributes the generated childcare diary to the parents of the corresponding children via their email address or a dedicated app.
[0861] Monitoring and alerting for risky behavior
[0862] Step 1: Real-time monitoring of video data
[0863] The server analyzes the video data in real time and monitors the behavior of the children.
[0864] Step 2: Detect risky behavior
[0865] The server uses a vision AI model to detect dangerous behavior (e.g., jumping from a high place, rough play).
[0866] Step 3: Trigger an alert
[0867] If dangerous behavior is detected, the server immediately generates an alert and sends it to the childcare worker's device.
[0868] Step 4: Receive and respond to alerts
[0869] The device (childcare worker) receives the alert and takes action to respond in real time.
[0870] Growth analysis and advice for childcare workers
[0871] Step 1: Collecting childcare diary data
[0872] The server periodically compiles data from the daycare diary and analyzes growth and behavioral patterns over a set period of time.
[0873] Step 2: Conduct a growth analysis
[0874] The server generates and evaluates the growth trends of each child based on the analysis data.
[0875] Step 3: Generate Advice
[0876] Based on growth trends, the server generates specific advice for nursery teachers, including specific activities and methods for promoting children's growth.
[0877] Step 4: Delivering advice
[0878] The server delivers the generated advice to the childcare worker's device, who receives it and puts it into practice.
[0879] This is the specific processing flow of the system, which reduces the burden on nursery teachers and ensures the safety and growth of children.
[0880] Example 1
[0881] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0882] In conventional childcare settings, childcare workers have to manually create childcare records, monitor children's behavior, and regularly provide information to parents. This increases the burden on childcare workers and can lead to a decline in the quality of childcare. It is also difficult to immediately detect risky behavior or provide specific advice based on growth analysis. Therefore, there was a need for a system that could reduce the burden on childcare workers and ensure safe, high-quality childcare.
[0883] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0884] In this invention, the server includes: means for capturing images of classrooms and playgrounds with a high-resolution camera; means for receiving the image data and identifying multiple individual objects using image authentication technology; means for tracking the behavior of the identified individual objects using computer vision technology and automatically generating a childcare record; means for automatically distributing the generated childcare record to parents; means for issuing a warning when dangerous behavior is detected; means for analyzing growth based on the childcare record and providing advice to childcare workers; means for processing the image data frame by frame and comparing it with multiple databases; means for tracking the positions and movements of children using computer vision technology; means for saving the tracking data in chronological order and performing activity mapping in real time; means for creating a childcare record in natural language using a generative AI model; means for automatically allocating contact information for each parent; and means for generating advice regarding childcare policies and activities based on the activity analysis results. This reduces the burden on childcare workers, improves safety, and enables the provision of high-quality childcare.
[0885] A "high-definition camera" is a camera that can capture wide-area images in high resolution.
[0886] "Video data" refers to information recorded in digital format from video captured by a camera.
[0887] "Image recognition technology" refers to algorithms and techniques used to identify individual subjects from video data.
[0888] "Individual subjects" refer to individual people, such as children, who participate in activities within the nursery school.
[0889] "Computer vision technology" is a technology that uses computers to analyze images and videos to recognize objects and track behavior.
[0890] A "childcare record" is a childcare diary that records in detail the daily activities and behavior of the children.
[0891] A "warning" is an alert that is issued when risky behavior is detected.
[0892] "Guardian" refers to the parent or legal guardian who is the caretaker of the child.
[0893] A "server" is a computer system that receives, analyzes, stores, and distributes video data.
[0894] "Tracking" means to continuously follow the movement of an object.
[0895] A "generative AI model" is an artificial intelligence model that generates sentences in natural language based on accumulated data.
[0896] "Contact information" refers to the parent's email address and account information for the dedicated application.
[0897] "Activity mapping" means visualizing the behavior and activities of kindergarten children along a timeline.
[0898] "Advice" means specific instructions or advice given to improve the quality of childcare.
[0899] "Database" means an electronic record containing information used to identify or match individuals.
[0900] This invention is a system that reduces the workload of childcare workers and provides safe, high-quality childcare. This system uses high-resolution image capture devices, servers, and terminals, and makes full use of image authentication and computer vision technologies to provide the following integrated functions.
[0901] The server acquires video data in real time from high-resolution camera devices installed in classrooms and playgrounds. These cameras are capable of covering a wide area in high resolution, and the video is streamed to the server using the RTSP protocol. The server stores this video data in a buffer frame by frame and begins processing.
[0902] The server then analyzes the received video data using image recognition technology. Specifically, it uses OpenCV and the Dlib library to detect the faces of the children and compare them with a pre-registered database to identify each individual child. This allows it to assign a unique ID to each child.
[0903] The behavior of identified children is tracked using computer vision technology on the server. Specifically, the OpenPose library is used to analyze the children's positions and movements. This data is saved in chronological order, and activity mapping is performed in real time. This allows for a detailed record of each child's activities.
[0904] Based on this tracking data, the server uses a generative AI model to automatically generate childcare records. The generated childcare records are detailed in natural language and include the child's daily activities, interactions, health status, etc. For example, tracking Taro playing with his friends using building blocks might result in a note in the childcare record saying, "Taro worked together with his friends to build a big tower."
[0905] The generated childcare records are automatically sent to the parents of the corresponding children via a server. The records are sent via email addresses registered in advance by the parents or via a dedicated childcare app. This allows parents to keep track of their child's daily activities in real time.
[0906] The server also analyzes video data in real time and monitors for dangerous behavior. For example, if the server detects a child engaging in dangerous behavior, such as jumping from a high place, it immediately issues an alert and notifies the childcare worker's device. This alert is sent as a push notification, allowing the childcare worker to respond quickly.
[0907] Furthermore, the server periodically compiles the childcare records and analyzes the growth of each child. Based on the results of this analysis, the generative AI model generates specific advice on childcare policies and activities, which are automatically distributed to the childcare worker's device. The childcare worker can use this advice to adjust and improve the childcare plan.
[0908] As a concrete example, when Taro is playing in the playground, a camera captures his movements, and the server uses facial recognition technology to identify him and record his actions. Computer vision technology tracks Taro playing with his friends using building blocks, and a generative AI model is used to create a childcare record stating, "Taro worked with his friends to build a big tower." This childcare record is automatically sent to Taro's mother. Furthermore, if the server detects Taro standing up high on the swing, it immediately sends an alert to the childcare worker's device, allowing the worker to respond quickly.
[0909] Prompt Sentence Examples
[0910] "How do you analyze real-time footage of children playing outside and send alerts to childcare workers if a particular child is exhibiting risky behavior?"
[0911] By utilizing this system, the workload of childcare workers will be significantly reduced, the safety of children will be improved, and high-quality childcare will be provided.
[0912] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0913] Step 1:
[0914] The server acquires video data in real time from high-resolution camera devices installed in classrooms and playgrounds. The input is the video stream from the camera, and the output is frame data stored on the server. Specifically, the server acquires data from the camera using the RTSP protocol and adds each frame to the processing queue.
[0915] Step 2:
[0916] The server analyzes the received video data using image recognition technology to detect the faces of the children. The input is the frame data from step 1, and the output is a feature vector of the detected face image. Specifically, it uses OpenCV and the Dlib library to detect faces in the frame and extract facial features.
[0917] Step 3:
[0918] The server uses facial recognition technology to compare the facial features detected with a pre-registered database to identify individual subjects (children). The input is the facial feature vector from step 2, and the output is the child's ID. Specifically, the server compares the facial feature vector with existing features in the database and identifies matching IDs based on cosine similarity.
[0919] Step 4:
[0920] The server uses computer vision technology to track the behavior of the identified children. The input is the child's ID from step 3 and the frame data from step 1, and the output is time-series data of the child's position and movement. Specific movements are tracked by detecting the child's skeleton using the OpenPose library.
[0921] Step 5:
[0922] The server automatically generates childcare records based on the tracking data using a generative AI model. The input is the time series data from step 4, and the output is a childcare record written in natural language. Specifically, the tracking data is input into the generative AI model, and a sentence describing the daily activities of the children is generated.
[0923] Step 6:
[0924] The server automatically distributes the generated childcare record to the corresponding guardian. The input is the childcare record from step 5, and the output is the information sent to the guardian's device. Specifically, the server refers to the guardian's contact database and sends the childcare record via email or a dedicated app.
[0925] Step 7:
[0926] The server analyzes the video data in real time and monitors for risky behavior. The input is the frame data from step 1, and the output is a warning notification when risky behavior is detected. Specifically, it uses computer vision technology to detect risky behavior patterns and sends a warning to the nursery teacher's device.
[0927] Step 8:
[0928] The server periodically compiles the childcare records and analyzes the children's growth. The input is the childcare record data from Step 5, and the output is the growth analysis results and advice. Specific operations include analyzing the compiled data and using a generative AI model to generate specific advice regarding childcare policies and activities.
[0929] Prompt Sentence Examples
[0930] "How do you analyze real-time footage of children playing outside and send alerts to childcare workers if a particular child is exhibiting risky behavior?"
[0931] This series of processes enables nursery school teachers to ensure the safety of children while saving time and effort, and to provide high-quality childcare.
[0932] (Application example 1)
[0933] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0934] In recent years, improving work efficiency and ensuring safety in factories have become extremely important issues. In particular, in manufacturing sites where many workers work simultaneously, there is a high risk of human error and dangerous behavior, and there is a growing demand for systems that can quickly detect and address such errors. Furthermore, manually creating work logs requires time and effort, making it difficult to immediately improve work efficiency. The present invention addresses these issues by providing advanced monitoring and automated work log generation to improve work efficiency and ensure safety in factories.
[0935] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0936] In this invention, the server includes means for capturing images of the work environment with a high-resolution camera, means for receiving the image data and identifying multiple individual objects using facial recognition technology, means for tracking the behavior of the identified individual objects and automatically generating a work log, means for automatically distributing the generated work log to relevant parties, means for issuing an alert when dangerous behavior is detected, and means for analyzing work efficiency based on the work log and providing relevant parties with points for improvement. This makes it possible to immediately grasp and analyze work efficiency within a factory and improve safety.
[0937] A "high-resolution camera" is a photographic device with high resolution that can record the work environment in detail.
[0938] "Work environment" refers to the place or area within a factory where work is carried out.
[0939] "Video data" is digital data containing visual information captured by a high-resolution camera.
[0940] "Facial recognition technology" is a biometric authentication technology for identifying people in video data.
[0941] "Individual Subject" refers to each identified worker within a factory.
[0942] "Behavior tracking" refers to following and recording the movements of identified individual subjects.
[0943] A "work log" is a digital record that records the actions and work of an individual.
[0944] "Automatically generating a work log" means that the server automatically creates a work log from the behavioral data of an individual subject.
[0945] "Stakeholders" refers to factory managers and supervisors who require work logs.
[0946] "Distribution" means sending the generated work log to the relevant parties.
[0947] "Issuing an alert" means sending a warning message when dangerous behavior is detected.
[0948] "Analyzing work efficiency" means analyzing the data in the work log to evaluate the efficiency of the work.
[0949] "Providing improvements" means presenting specific suggestions to stakeholders to improve work efficiency.
[0950] This invention is a system that improves work efficiency and ensures safety in factories by utilizing high-resolution cameras, facial recognition technology, and AI analysis technology. The system consists of a server, a camera, smart glasses, and a dedicated app.
[0951] System configuration
[0952] 1. High-quality camera:
[0953] The working environment within the factory is photographed in high resolution, and detailed video data is sent to a server in real time.
[0954] 2. Server:
[0955] The server processes the received video data and identifies the worker using facial recognition technology. Azure Face API, for example, is used for facial recognition. The actions of identified workers are tracked, and a work log is automatically generated based on that data. The work log contains detailed records of each worker's movements and tasks.
[0956] 3. Risky behavior monitoring and alerting:
[0957] The server analyzes video data and work logs in real time and immediately issues an alert if any dangerous behavior is detected. The alert is then sent to the smart glasses to alert the worker, enabling a rapid response.
[0958] 4. Work log distribution:
[0959] The generated work logs are automatically distributed to relevant parties (factory managers and supervisors) via email addresses or a dedicated app.
[0960] 5. Analyzing work efficiency and providing improvements:
[0961] The server analyzes work efficiency based on the work logs and provides specific improvements to the relevant parties, thereby optimizing work efficiency throughout the factory.
[0962] Hardware and software used
[0963] Hardware: high-resolution cameras, smart glasses (e.g., Microsoft HoloLens)
[0964] Software: Facial recognition AI (e.g., Azure Face API), behavioral analysis AI (e.g., TensorFlow, OpenCV), database (e.g., Azure SQL Database)
[0965] Example of operation
[0966] For example, if a worker is operating heavy machinery in an inappropriate manner, a high-resolution camera captures this, the server uses facial recognition technology to identify the worker, and AI determines that the operation is inappropriate. An alert is immediately sent to the smart glasses, prompting the worker to operate in a safe manner. At the same time, details of the operation are automatically generated as a work log and sent to the factory manager. Based on this work log, areas for improving work efficiency are analyzed and provided to relevant parties.
[0967] Example of input prompt for generative AI model
[0968] prompt:
[0969] Design a smart glasses application for factory monitoring with self-diagnosis function. Please pay attention to the following points:
[0970] 1. Tracking worker behavior in real time using cameras and smart glasses.
[0971] 2. Use facial recognition technology to identify individual workers.
[0972] 3. If risky behavior is detected, an alert is sent immediately.
[0973] 4. Automatically generate work logs and analyze areas for improvement.
[0974] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0975] Step 1:
[0976] High-resolution cameras capture the factory environment. The input is real-time footage from within the factory, and the output is high-resolution video data. Specifically, the cameras have a field of view that covers the entire factory, capturing detailed images of workers and machine operations.
[0977] Step 2:
[0978] The server receives the video data sent from the camera. The input is high-definition video data that is temporarily stored in the server. The output is video data stored in the server's storage.
[0979] Step 3:
[0980] The server analyzes the video data using facial recognition AI (Azure Face API) to identify multiple workers. The input is the saved video data, and the output is the identification information of each worker. Specifically, the AI analyzes the facial images in the video and identifies the worker's ID by comparing them with a pre-registered database.
[0981] Step 4:
[0982] The server tracks the actions of each identified worker. The input is identification information and real-time video data, and the output is individual behavioral patterns. Specifically, the server continuously tracks the movements of workers and records what tasks they are performing.
[0983] Step 5:
[0984] The server automatically generates work logs based on the behavioral data it tracks. The input is behavioral pattern data, and the output is a digital work log. Specifically, the AI organizes the recorded information and provides a detailed record of each worker's daily activities and their content.
[0985] Step 6:
[0986] The server automatically distributes the generated work log to the relevant parties. The input is the generated work log, and the output is the work log sent to the relevant parties' email addresses or a dedicated app. Specifically, the log file is sent to the specified recipient via a mail server or application.
[0987] Step 7:
[0988] The server analyzes real-time video data and issues an alert if it detects dangerous behavior. The input is real-time video data, and the output is an alert message sent to the smart glasses. Specifically, when the AI detects an abnormal behavior pattern, it immediately generates an alert and sends it to the worker's smart glasses.
[0989] Step 8:
[0990] The server analyzes work efficiency based on work logs and provides suggestions for improvement to the relevant parties. The input is daily work log data, and the output is an analysis report including suggestions for improvement. Specifically, the server compares the data with past work data, proposes specific improvements for improving work efficiency and safety, and sends them to the relevant parties.
[0991] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0992] This invention is a system that reduces the burden on childcare workers and ensures safe, high-quality childcare. Using high-resolution cameras, servers, and terminals, the system utilizes face recognition technology, emotion engines, and AI analysis technology to automatically generate childcare diaries, distribute them to parents, monitor and alert on risky behavior, analyze child development, and provide advice to childcare workers.
[0993] 1. Camera video capture
[0994] The server receives real-time video data from high-resolution cameras installed in classrooms and on the playground, which have a wide field of view and can accurately capture detailed activity within the school.
[0995] 2. Face recognition and individual object identification
[0996] The server analyzes the received video data using facial recognition AI technology to identify the children in the video. The facial recognition technology compares the facial images in the video with a pre-registered database to identify the ID of each individual subject (child).
[0997] 3. Behavior tracking and automatic generation of childcare diary
[0998] The server tracks the behavior of identified children and records their behavioral patterns (e.g., types of play, interpersonal relationships, and health status). Based on this, a daycare diary is automatically generated. The diary details each child's daily activities, interactions, and health status.
[0999] 4. Distribution of the childcare diary to parents
[1000] The server automatically distributes the generated childcare diary to the parents of the corresponding child via their email address or a dedicated app, allowing parents to keep track of their child's daily activities in real time.
[1001] 5. Monitoring risky behavior and issuing alerts
[1002] The server analyzes the video data in real time and monitors the children's risky behavior. If any risky behavior (e.g., jumping from a high place or rough play) is detected, an alert is sent immediately.
[1003] The terminal (childcare worker) receives this alert and can respond quickly.
[1004] 6. Growth analysis and advice for childcare workers
[1005] The server periodically compiles data from the childcare diary and analyzes the growth of each child. Based on the results of this analysis, it identifies individual growth trends and provides advice to childcare workers regarding specific childcare policies and activities. The advice is automatically sent to the childcare workers' devices and can be used in practice.
[1006] 7. Introducing the Emotion Engine
[1007] The server inputs the video and audio data into the emotion engine, which analyzes the child's emotional state based on their facial expressions and tone of voice. The emotion engine then identifies the emotion the child is feeling (e.g., joy, sadness, anger, anxiety).
[1008] Based on the emotional states identified by the emotion engine, the server adds emotional data to the childcare diary and development analysis, for example, recording how children felt during play and how specific activities affected their emotions.
[1009] The server generates and delivers advice to childcare workers and parents based on their emotional state, suggesting appropriate ways to respond. For example, if a child is feeling anxious, the server identifies the root cause and provides specific advice on how to support the child.
[1010] For example, when Taro is playing in the playground, a camera captures his movements, and the server identifies him using facial recognition technology and records his actions. His behavior as he plays with building blocks with his friends is tracked, and a diary entry is made stating, "Taro played with building blocks with his friends and worked together to build a big tower." This diary entry is automatically sent to Taro's mother's email address that afternoon.
[1011] Furthermore, if Taro attempts dangerous behavior such as standing up too high on the swing, the server will detect this and immediately send an alert to the nursery teacher's device, allowing the nursery teacher to immediately rush over to Taro and ensure his safety.
[1012] Furthermore, the emotion engine analyzes Taro's facial expressions and tone of voice and detects that he looks anxious. Based on this, the server sends advice to the nursery teacher's device, such as, "Taro is likely feeling anxious, so please talk to him a bit to reassure him."
[1013] The above is a specific embodiment for carrying out the present invention. This system allows childcare workers to reduce their workload, ensure the safety and emotional growth of children, and provide high-quality childcare.
[1014] The processing flow will be explained below.
[1015] Processing for an automatic childcare diary generation system that combines face recognition and emotion engine
[1016] From video capture to emotion analysis
[1017] Step 1: Receiving video data
[1018] The server receives video data in real time from high-resolution cameras installed in classrooms and playgrounds, and the video data is temporarily stored in dedicated storage.
[1019] Step 2: Face recognition and individual subject identification
[1020] The server inputs the received video data into a facial recognition AI model to identify the children in the video. The facial recognition technology compares the facial images in the video with a pre-registered database to identify the ID of each individual subject (child).
[1021] Step 3: Tracking behavior
[1022] The movements of identified individuals are tracked in real time, with the server recording the start and end of specific actions (e.g., playing, crafting, talking) with timestamps.
[1023] Step 4: Analysis by Emotion Engine
[1024] The server inputs video and audio data into an emotion engine, which analyzes the children's facial expressions and tone of voice to determine their emotional state (e.g., joy, sadness, anger, anxiety).
[1025] Step 5: Creating a childcare diary
[1026] The server automatically generates a daily diary for each child based on tracking data and emotion analysis data, detailing each child's daily activities, interactions, health status, and emotional state.
[1027] Delivery of childcare diary to parents
[1028] Step 1: Generate a journal
[1029] The server automatically distributes the generated childcare diary to the parents of the corresponding children via their email addresses or a dedicated app.
[1030] Monitoring and alerting for risky behavior
[1031] Step 1: Real-time monitoring of video data
[1032] The server analyzes the video data in real time and monitors the behavior of the children.
[1033] Step 2: Detect risky behavior
[1034] The server uses a vision AI model to detect dangerous behavior (e.g., jumping from a high place, rough play).
[1035] Step 3: Trigger an alert
[1036] If dangerous behavior is detected, the server immediately generates an alert and sends it to the childcare worker's device.
[1037] Step 4: Receive and respond to alerts
[1038] The device (childcare worker) receives the alert and takes action to respond in real time.
[1039] Growth analysis and advice for childcare workers
[1040] Step 1: Collecting childcare diary data
[1041] The server periodically compiles data from the daycare diary and analyzes growth and behavioral patterns over a set period of time.
[1042] Step 2: Conduct a growth analysis
[1043] The server generates and evaluates the growth trends of each child based on the analysis data.
[1044] Step 3: Generate Advice
[1045] Based on growth trends, the server generates specific advice for nursery teachers, including specific activities and methods for promoting children's growth.
[1046] Step 4: Delivering advice
[1047] The server delivers the generated advice to the childcare worker's device, who receives it and puts it into practice.
[1048] Further processing by the emotion engine
[1049] Step 1: Collecting emotion data
[1050] The server integrates the analytical data from the emotion engine into the childcare diary, allowing for an understanding of emotional fluctuations and trends.
[1051] Step 2: Generating sentiment-based advice
[1052] Based on the emotional data, the server generates advice suggesting appropriate responses for childcare workers and parents.
[1053] Step 3: Delivering emotional advice
[1054] The server delivers advice based on the generated emotions to the childcare worker's device, who then receives it and puts it into practice.
[1055] For example, when Taro is playing in the playground, a camera captures his movements, and the server identifies him using facial recognition technology and records his actions. His behavior as he plays with building blocks with his friends is tracked, and a diary entry is made stating, "Taro played with building blocks with his friends and worked together to build a big tower." This diary entry is automatically sent to Taro's mother's email address that afternoon.
[1056] Furthermore, if Taro attempts dangerous behavior such as standing up too high on the swing, the server will detect this and immediately send an alert to the nursery teacher's device, allowing the nursery teacher to immediately rush over to Taro and ensure his safety.
[1057] Furthermore, the emotion engine analyzes Taro's facial expressions and tone of voice and detects that he looks anxious. Based on this, the server sends advice to the nursery teacher's device, such as, "Taro is likely feeling anxious, so please talk to him a bit to reassure him."
[1058] The above is a specific embodiment for carrying out the present invention. This system allows childcare workers to reduce their workload, ensure the safety and emotional growth of children, and provide high-quality childcare.
[1059] Example 2
[1060] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1061] In childcare facilities, the workload of childcare workers is increasing, and it is often the case that they are not doing enough to ensure the safety of children, monitor their growth, or report to parents. It is also difficult to accurately grasp the emotional state of children and respond appropriately. This puts the quality of childcare at risk of declining.
[1062] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1063] In this invention, the server includes means for capturing images of classrooms and playgrounds with a high-resolution camera, means for receiving the image data and identifying multiple individual objects using facial recognition technology, means for tracking the behavior of the identified individual objects and automatically generating a childcare diary, means for automatically distributing the generated childcare diary to parents, means for issuing an alert when dangerous behavior is detected, means for analyzing growth based on the childcare diary and providing advice to childcare workers, means for analyzing emotional states from video data and audio data using an emotion engine, and means for suggesting appropriate response methods to childcare workers and parents based on the analyzed emotional states. This reduces the burden on childcare workers, enables management of the safety and growth of children, and enables understanding of emotional states and appropriate responses.
[1064] A "high-resolution camera" is a camera that has a wide field of view and can capture detailed video data in high resolution.
[1065] "Video data" refers to video information acquired from cameras installed in classrooms and playgrounds.
[1066] "Facial recognition technology" is a technology that identifies individuals by identifying faces in video data and comparing them with a pre-registered database.
[1067] "Individual object" is a term that refers to a specific person, especially a kindergartener.
[1068] "Behavioral tracking" is a method of monitoring and recording the movements and behavioral patterns of individual subjects in real time.
[1069] A "childcare diary" is a report that details the behavior, interactions, and health status of the children during a day's childcare activities.
[1070] "Guardian" is a term that refers to the parents or guardians of kindergarten children.
[1071] An "alert" is a warning message that is sent immediately when dangerous behavior or an abnormality is detected in real time.
[1072] "Growth analysis" is a method of analyzing children's growth patterns based on data from the nursery school diary and evaluating developmental trends.
[1073] "Advice" refers to specific methods of response and guidelines provided to childcare workers and parents based on growth analysis and emotional analysis.
[1074] An "emotion engine" is a technology that analyzes video and audio data to identify the emotional state of an individual subject.
[1075] "Response methods" are specific measures that, based on the analyzed data, indicate appropriate approaches in caring for and guiding children.
[1076] This invention is a system that reduces the burden on childcare workers and ensures safe, high-quality childcare. This system uses high-resolution cameras, servers, and terminals, and makes full use of face recognition technology, emotion engines, and AI analysis technology to automatically generate childcare diaries, distribute them to parents, monitor and alert on risky behavior, analyze child development, and provide advice to childcare workers.
[1077] Hardware and software used
[1078] server
[1079] Data servers with high processing power (e.g., Dell PowerEdge series)
[1080] Facial recognition AI technology (e.g., FaceNet)
[1081] Emotion engine (e.g. Microsoft Azure Emotion API)
[1082] Machine learning algorithms (e.g., Python's scikit-learn library)
[1083] camera
[1084] High-resolution network camera (e.g. AXIS P1368-E)
[1085] Terminal
[1086] Tablet or smartphone for childcare workers (e.g. iPad, Android tablet)
[1087] Explanation of program processing
[1088] High-definition camera video capture
[1089] The server receives real-time video data from high-resolution cameras installed in classrooms and playgrounds. These cameras have a wide field of view and can accurately capture detailed activity within the school. The video data is stored in the server's storage and used for analysis.
[1090] Identifying individual subjects using facial recognition technology
[1091] The server analyzes the received video data using facial recognition AI technology to identify the children in the video. The facial recognition technology compares the facial images in the video with a pre-registered database to identify the ID of each individual subject (child). This allows for accurate tracking of each child's activities.
[1092] Behavior tracking and automatic generation of childcare diary
[1093] The server tracks the identified children's behavior in real time and records their behavioral patterns (e.g., types of play, interpersonal relationships, and health status). Based on this behavioral data, a daycare diary is automatically generated, detailing each child's daily activities, interactions, and health status.
[1094] Delivery of childcare diary to parents
[1095] The generated daycare diary is automatically sent to the parents of the children via their email address or a dedicated app (e.g., Custom Nursery App). This allows parents to keep track of their children's daily activities in real time.
[1096] Monitoring risky behavior and issuing alerts
[1097] The server analyzes the video data in real time and monitors the children's risky behavior. For example, if dangerous behavior such as jumping from a high place or rough play is detected, an alert is immediately issued. A push notification is sent to the childcare worker's device (e.g., tablet), allowing the childcare worker to respond quickly.
[1098] Growth analysis and advice for childcare workers
[1099] The server periodically compiles data from the childcare diary and analyzes the growth of each child. Based on the results of this analysis, it identifies individual growth trends and provides advice to childcare workers regarding specific childcare policies and activities. This advice is automatically sent to devices and can be used by childcare workers in their daily work.
[1100] Introducing the Emotion Engine
[1101] The server inputs video and audio data into an emotion engine, which analyzes the emotional state of the child based on facial expressions, tone of voice, etc. The emotion engine identifies the child's emotions (e.g., joy, sadness, anger, anxiety). This emotional state is reflected in the childcare diary and growth analysis, and is used as part of advice for childcare workers and parents. Specifically, the server calls an emotion analysis API to obtain the analysis results, and a recommendation algorithm then suggests appropriate responses based on these results.
[1102] Specific examples
[1103] While Taro is playing in the playground, a high-definition camera captures his movements, and the server uses facial recognition technology to identify him and record his actions. His behavior as he plays with building blocks with his friends is tracked, and an entry is made in the daycare diary, stating, "Taro played with building blocks with his friends and worked together to build a big tower." This diary is automatically sent to Taro's mother's email address that afternoon. Furthermore, if Taro attempts the dangerous act of standing up high on the swing, the server detects this and immediately sends an alert to the nursery teacher's device, allowing the nursery teacher to quickly rush to Taro's aid and ensure his safety.
[1104] Prompt Sentence Examples
[1105] "High-resolution cameras and AI technology are used in daycare centers to track the behavior and emotions of children in real time, and facial recognition technology is used to identify each child individually. Please explain in detail how the AI will automatically generate daycare diaries, provide information to parents, monitor and alert for risky behavior, analyze growth, and provide advice to childcare workers."
[1106] The above is a specific embodiment for carrying out the present invention. This system allows childcare workers to reduce their workload, ensure the safety and emotional growth of children, and provide high-quality childcare.
[1107] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1108] Step 1:
[1109] Receiving video data
[1110] The server receives real-time video data of classrooms and playgrounds from high-resolution cameras. The input is the video stream from each camera, and the output is the received video data. Specifically, the server specifies the IP address of the camera and obtains the video stream using HTTP or RTSP protocol. This allows the server to monitor the current activities of the children in real time.
[1111] Step 2:
[1112] Identifying individual subjects through facial recognition
[1113] The server analyzes the received video data using facial recognition technology. The input is the video data and a database of pre-registered faces of the children, and the output is the ID of the identified child. Specifically, the server uses a facial recognition algorithm (e.g., FaceNet) to detect faces in the video data and compare them with the database. Based on the results of this comparison, the server identifies individual subjects (children).
[1114] Step 3:
[1115] Behavior tracking and data recording
[1116] The server tracks the behavior of identified children and records the data. The input is the child's ID and video data, and the output is a log of behavioral data. Specifically, the server uses a behavioral analysis algorithm to analyze the children's movements and location information and records it in chronological order. For example, a detailed behavioral log is generated, such as "Taro started playing in the sandbox at 9:15 and was talking to his classmate Hanako at 9:30."
[1117] Step 4:
[1118] Automatic generation of childcare diary
[1119] The server automatically generates a childcare diary based on the recorded behavioral data. The input is a log of behavioral data, and the output is a childcare diary. Specifically, the server organizes the behavioral logs based on a template and creates a diary-style report. This report details the children's daily activities, interactions, and health status.
[1120] Step 5:
[1121] Distribution of childcare journals
[1122] The server automatically distributes the generated childcare diary to the children's parents. The input is the childcare diary and the parents' contact information, and the output is the sent email or app notification. Specifically, the server uses an email sending API or push notification API to send the diary via the parents' email address or a dedicated app. This allows parents to keep track of their children's daily activities in real time.
[1123] Step 6:
[1124] Monitoring and alerting for risky behavior
[1125] The server analyzes video data in real time and monitors children's risky behavior. The input is video data, and the output is an alert notification. Specifically, the server uses a motion analysis algorithm to detect risky behavior (e.g., jumping from high places, rough play). If any risky behavior is detected, an alert notification is immediately pushed to the nursery teacher's device.
[1126] Step 7:
[1127] Providing growth analysis and advice
[1128] The server periodically compiles the data from the childcare diary and analyzes the growth of each child. The input is the childcare diary data, and the output is a growth analysis report and advice. Specifically, the server uses a machine learning algorithm to analyze past behavioral data and predict the child's growth trends. Based on this, it provides childcare workers with advice on appropriate childcare policies and activities.
[1129] Step 8:
[1130] Emotion analysis using an emotion engine
[1131] The server inputs video and audio data into an emotion engine to analyze the emotional state of the children. The input is video and audio data, and the output is the emotion analysis results. Specifically, the server calls an emotion analysis API (e.g., Microsoft Azure Emotion API) to analyze the children's facial expressions and tone of voice to identify their emotional state. The results of this analysis are reflected in the childcare diary and growth analysis, and are used as part of advice for childcare workers and parents.
[1132] (Application example 2)
[1133] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1134] There are challenges in monitoring worker safety and managing work efficiently in factories. In particular, it is difficult to grasp workers' behavior and emotional state in real time, quickly detect dangerous behavior, and take countermeasures. Furthermore, there are insufficient systems to provide appropriate advice based on workers' emotional state, which hinders workers' stress reduction and efficient work performance.
[1135] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1136] In this invention, the server includes means for capturing images of an area with a high-resolution camera, means for receiving the image data and identifying multiple individual subjects using facial recognition technology, means for tracking the behavior of the identified individuals and automatically generating records, means for automatically distributing the generated records to relevant parties, means for issuing warnings when dangerous behavior is detected, means for analyzing growth or progress based on the records and providing advice to the person in charge, and means for analyzing the emotional state of workers in real time using a portable display device worn by the workers and providing appropriate advice, thereby enabling safety monitoring of workers in factories and efficient work management.
[1137] A "high-resolution imaging device" is a high-resolution camera device that has a wide field of view and can capture detailed images.
[1138] "Area" refers to the space or location where a particular activity takes place, including, for example, a work area in a factory or an area requiring safety monitoring.
[1139] "Facial recognition technology" is a technology that can identify individual subjects by analyzing facial images in video and comparing them with a pre-registered database.
[1140] "Distinct objects" refer to specific people or objects that the system needs to identify, including, for example, factory workers.
[1141] "Behavioral tracking" refers to tracking the movements and actions of individual subjects in real time and recording that data.
[1142] "Automatic generation of records" refers to automatically creating diaries and reports based on tracked behavioral data.
[1143] "Stakeholders" refers to people who need the system's output data, including, for example, factory managers and the workers' families.
[1144] "Sending a warning" means that the system will send an alert in real time when it detects dangerous behavior or an unexpected situation.
[1145] "Analyzing growth or progress" refers to analyzing the progress or areas for improvement of an individual subject based on recorded data.
[1146] "Providing advice to the person in charge" refers to notifying the person in charge of specific instructions or suggestions based on the analysis results.
[1147] A "portable display device" is a device with a small display that can be worn and used by a worker.
[1148] "Analyzing emotional states" is a technology that identifies emotions from video and audio data and understands those states.
[1149] The system of this invention uses high-resolution imaging devices, servers, facial recognition technology, emotion analysis engines, and portable display devices (hereinafter referred to as smart glasses) to monitor the safety of workers within factories and achieve efficient business management.
[1150] Hardware and software used:
[1151] High-resolution camera equipment: A high-resolution camera equipment with a wide field of view that can capture detailed images.
[1152] Server: A cloud or local server that receives and analyzes data in real time.
[1153] Facial recognition technology: Technology that analyzes facial images in video and identifies individual subjects (e.g., Amazon Rekognition, Microsoft Azure Face API).
[1154] Emotion analysis engine: Technology that identifies emotions from video and audio data and understands their state (e.g., Affectiva, IBM Watson Tone Analyzer).
[1155] Portable display devices (smart glasses): Devices worn by workers that display information in real time (e.g., Google Glass, Microsoft HoloLens).
[1156] Real-time data analysis engine: Technology that performs real-time analysis of data (e.g., Apache Kafka, Apache Flink).
[1157] The server uses high-resolution camera equipment to capture images of the factory work area. The captured image data is sent to the server in real time. The server's facial recognition technology analyzes the image data and identifies multiple individual subjects (factory workers). The actions of the identified individuals (workers) are tracked by the server, and an activity record is automatically generated.
[1158] The recorded behavioral data is automatically distributed to relevant parties (such as factory managers) via email or a dedicated application. If the system detects any dangerous behavior, it will immediately issue a warning. Workers wearing the smart glasses will receive a warning message in real time, allowing them to respond quickly.
[1159] Furthermore, the server analyzes the progress of workers based on the recorded data and provides advice to managers, thereby improving work efficiency and safety. The emotion analysis engine analyzes the emotional state of workers in real time and provides appropriate advice. For example, if a worker is in a high-stress state, the smart glasses will display advice such as "You need to take a break."
[1160] As a concrete example, if a factory worker is wearing smart glasses while working at height, a camera will capture his movements and send the data to a server, where facial recognition technology will identify the worker and begin tracking his behavior and analyzing his emotions. If the worker makes an unstable movement, the server will determine this as a dangerous behavior and display a warning on the smart glasses saying, "Your current movement is dangerous. Please move to a safe location."
[1161] Prompt Sentence Examples
[1162] Design a system to monitor the safety of factory workers using facial recognition technology, emotion engines, and AI analysis technology. Specifically, include the following features:
[1163] 1. Facial recognition function to identify workers using smart glasses
[1164] 2. Tracking worker behavior and detecting dangerous behavior
[1165] 3. Monitor emotional states using a sentiment analysis engine
[1166] 4. Alert workers when they are in a dangerous situation
[1167] 5. Providing advice on safety and stress reduction
[1168] Specific examples of technologies to consider using include Amazon Rekognition, Affectiva, and Apache Flink.
[1169] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1170] Step 1:
[1171] High-resolution camera equipment captures video data of the factory work area, including the overall situation within the factory and the actions of workers.
[1172] Input: Factory area video
[1173] Output: High-resolution video data
[1174] Step 2:
[1175] The server receives the video data in real time, and then transmits the video data to the server, where the analysis process begins.
[1176] Input: High-resolution video data
[1177] Output: Video data required for analysis
[1178] Step 3:
[1179] The server uses facial recognition technology to identify workers in the video data by comparing facial images in the video with a pre-registered database to identify individual targets.
[1180] Input: Video data, face database
[1181] Output: ID of the identified individual (worker)
[1182] Step 4:
[1183] The server tracks the behavior of identified individuals, recording the movements and actions of workers in real time.
[1184] Input: ID of identified individual, video data
[1185] Output: Tracked behavior data
[1186] Step 5:
[1187] The server automatically generates records based on the collected behavioral data, including the worker's daily activities and tasks.
[1188] Input: Tracked behavioral data
[1189] Output: Automatically generated activity log
[1190] Step 6:
[1191] The server automatically distributes the generated records to the relevant parties via email or a dedicated application.
[1192] Input: Auto-generated event log, stakeholder list
[1193] Output: Records sent to interested parties
[1194] Step 7:
[1195] The server analyzes the video data in real time to detect dangerous behavior and issues a warning if any.
[1196] Input: Real-time video data
[1197] Output: Risky behavior warning
[1198] Step 8:
[1199] Workers wearing smart glasses receive real-time warning messages that are immediately displayed to workers, enabling them to take prompt action.
[1200] Input: Risky behavior warning message
[1201] Output: Real-time displayed warnings
[1202] Step 9:
[1203] The server analyzes progress based on the activity records and provides advice to the administrator, thereby improving work efficiency and safety.
[1204] Input: Action log
[1205] Output: Progress analysis and advice
[1206] Step 10:
[1207] The emotion analysis engine analyzes the emotional state of the worker in real time, and based on the analysis results, appropriate advice is generated and displayed on the smart glasses.
[1208] Input: Video data, audio data
[1209] Output: Emotional state analysis and advice
[1210] For example, if the server detects unstable movements of a worker working at height, it will display a warning on the smart glasses saying, "The current movement is dangerous. Please move to a safe location." This warning allows the worker to take immediate action to ensure safety.
[1211] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1212] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1213] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1214] [Fourth embodiment]
[1215] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1216] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1217] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1218] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1219] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1220] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1221] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1222] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1223] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1224] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1225] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1226] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1227] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1228] This invention is a system that reduces the burden on childcare workers and ensures safe, high-quality childcare. Using high-resolution cameras, servers, and terminals, the system utilizes face recognition technology and AI analysis technology to automatically generate childcare diaries, distribute them to parents, monitor and alert for risky behavior, analyze child development, and provide advice to childcare workers.
[1229] 1. Camera video capture
[1230] The server receives real-time video data from high-resolution cameras installed in classrooms and on the playground, which have a wide field of view and can accurately capture detailed activity within the school.
[1231] 2. Face recognition and individual object identification
[1232] The server analyzes the received video data using facial recognition AI technology to identify the children in the video. Facial recognition technology identifies the ID of each individual child by comparing the facial image in the video with a pre-registered database.
[1233] 3. Behavior tracking and automatic generation of childcare diary
[1234] The server tracks the behavior of identified children and records their behavioral patterns (e.g., types of play, interpersonal relationships, and health status). Based on this, a daycare diary is automatically generated. The diary details each child's daily activities, interactions, and health status.
[1235] 4. Distribution of the childcare diary to parents
[1236] The server automatically distributes the generated childcare diary to the parents of the corresponding child via their email address or a dedicated app, allowing parents to keep track of their child's daily activities in real time.
[1237] 5. Monitoring risky behavior and issuing alerts
[1238] The server analyzes the video data in real time and monitors the children's risky behavior. If any risky behavior (e.g., jumping from a high place or rough play) is detected, an alert is sent immediately.
[1239] The terminal (childcare worker) receives this alert and can respond quickly.
[1240] 6. Growth analysis and advice for childcare workers
[1241] The server periodically compiles data from the childcare diary and analyzes the growth of each child. Based on the results of this analysis, it identifies individual growth trends and provides advice to childcare workers regarding specific childcare policies and activities. The advice is automatically sent to the childcare workers' devices and can be used in practice.
[1242] For example, when Taro is playing in the playground, a camera captures his movements, and the server identifies him using facial recognition technology and records his actions. His behavior as he plays with building blocks with his friends is tracked, and a diary entry is made stating, "Taro played with building blocks with his friends and worked together to build a big tower." This diary entry is automatically sent to Taro's mother's email address that afternoon.
[1243] In addition, if Taro attempts dangerous behavior in the playground, such as standing up high on a swing, the server will detect this and send an alert to the nursery teacher's device, allowing the nursery teacher to immediately rush over to Taro and ensure his safety.
[1244] The above is a specific embodiment for carrying out the present invention. This system allows childcare workers to save time and effort, ensure the safety of children, and provide high-quality childcare.
[1245] The processing flow will be explained below.
[1246] Processing for automatic childcare diary generation system using face recognition
[1247] From video capture to automatic generation of childcare diary
[1248] Step 1: Receiving video data
[1249] The server receives video data in real time from high-resolution cameras installed in classrooms and playgrounds, and the video data is temporarily stored in dedicated storage.
[1250] Step 2: Face recognition and individual subject identification
[1251] The server inputs the received video data into a facial recognition AI model to identify the child in the video. The facial recognition technology compares the facial images in the video with a pre-registered database to identify the ID of each individual subject (child).
[1252] Step 3: Tracking behavior
[1253] The movements of identified individuals are tracked in real time, with the server recording the start and end of specific actions (e.g., playing, crafting, talking) with timestamps.
[1254] Step 4: Creating a childcare diary
[1255] The server automatically generates a daily diary for each child based on the tracking data, detailing each child's daily activities, interactions, and health status.
[1256] Step 5: Distributing the journal
[1257] The server automatically distributes the generated childcare diary to the parents of the corresponding children via their email address or a dedicated app.
[1258] Monitoring and alerting for risky behavior
[1259] Step 1: Real-time monitoring of video data
[1260] The server analyzes the video data in real time and monitors the behavior of the children.
[1261] Step 2: Detect risky behavior
[1262] The server uses a vision AI model to detect dangerous behavior (e.g., jumping from a high place, rough play).
[1263] Step 3: Trigger an alert
[1264] If dangerous behavior is detected, the server immediately generates an alert and sends it to the childcare worker's device.
[1265] Step 4: Receive and respond to alerts
[1266] The device (childcare worker) receives the alert and takes action to respond in real time.
[1267] Growth analysis and advice for childcare workers
[1268] Step 1: Collecting childcare diary data
[1269] The server periodically compiles data from the daycare diary and analyzes growth and behavioral patterns over a set period of time.
[1270] Step 2: Conduct a growth analysis
[1271] The server generates and evaluates the growth trends of each child based on the analysis data.
[1272] Step 3: Generate Advice
[1273] Based on growth trends, the server generates specific advice for nursery teachers, including specific activities and methods for promoting children's growth.
[1274] Step 4: Delivering advice
[1275] The server delivers the generated advice to the childcare worker's device, who receives it and puts it into practice.
[1276] This is the specific processing flow of the system, which reduces the burden on nursery teachers and ensures the safety and growth of children.
[1277] Example 1
[1278] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1279] In conventional childcare settings, childcare workers have to manually create childcare records, monitor children's behavior, and regularly provide information to parents. This increases the burden on childcare workers and can lead to a decline in the quality of childcare. It is also difficult to immediately detect risky behavior or provide specific advice based on growth analysis. Therefore, there was a need for a system that could reduce the burden on childcare workers and ensure safe, high-quality childcare.
[1280] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1281] In this invention, the server includes: means for capturing images of classrooms and playgrounds with a high-resolution camera; means for receiving the image data and identifying multiple individual objects using image authentication technology; means for tracking the behavior of the identified individual objects using computer vision technology and automatically generating a childcare record; means for automatically distributing the generated childcare record to parents; means for issuing a warning when dangerous behavior is detected; means for analyzing growth based on the childcare record and providing advice to childcare workers; means for processing the image data frame by frame and comparing it with multiple databases; means for tracking the positions and movements of children using computer vision technology; means for saving the tracking data in chronological order and performing activity mapping in real time; means for creating a childcare record in natural language using a generative AI model; means for automatically allocating contact information for each parent; and means for generating advice regarding childcare policies and activities based on the activity analysis results. This reduces the burden on childcare workers, improves safety, and enables the provision of high-quality childcare.
[1282] A "high-definition camera" is a camera that can capture wide-area images in high resolution.
[1283] "Video data" refers to information recorded in digital format from video captured by a camera.
[1284] "Image recognition technology" refers to algorithms and techniques used to identify individual subjects from video data.
[1285] "Individual subjects" refer to individual people, such as children, who participate in activities within the nursery school.
[1286] "Computer vision technology" is a technology that uses computers to analyze images and videos to recognize objects and track behavior.
[1287] A "childcare record" is a childcare diary that records in detail the daily activities and behavior of the children.
[1288] A "warning" is an alert that is issued when risky behavior is detected.
[1289] "Guardian" refers to the parent or legal guardian who is the caretaker of the child.
[1290] A "server" is a computer system that receives, analyzes, stores, and distributes video data.
[1291] "Tracking" means to continuously follow the movement of an object.
[1292] A "generative AI model" is an artificial intelligence model that generates sentences in natural language based on accumulated data.
[1293] "Contact information" refers to the parent's email address and account information for the dedicated application.
[1294] "Activity mapping" means visualizing the behavior and activities of kindergarten children along a timeline.
[1295] "Advice" means specific instructions or advice given to improve the quality of childcare.
[1296] "Database" means an electronic record containing information used to identify or match individuals.
[1297] This invention is a system that reduces the workload of childcare workers and provides safe, high-quality childcare. This system uses high-resolution image capture devices, servers, and terminals, and makes full use of image authentication and computer vision technologies to provide the following integrated functions.
[1298] The server acquires video data in real time from high-resolution camera devices installed in classrooms and playgrounds. These cameras are capable of covering a wide area in high resolution, and the video is streamed to the server using the RTSP protocol. The server stores this video data in a buffer frame by frame and begins processing.
[1299] The server then analyzes the received video data using image recognition technology. Specifically, it uses OpenCV and the Dlib library to detect the faces of the children and compare them with a pre-registered database to identify each individual child. This allows it to assign a unique ID to each child.
[1300] The behavior of identified children is tracked using computer vision technology on the server. Specifically, the OpenPose library is used to analyze the children's positions and movements. This data is saved in chronological order, and activity mapping is performed in real time. This allows for a detailed record of each child's activities.
[1301] Based on this tracking data, the server uses a generative AI model to automatically generate childcare records. The generated childcare records are detailed in natural language and include the child's daily activities, interactions, health status, etc. For example, tracking Taro playing with his friends using building blocks might result in a note in the childcare record saying, "Taro worked together with his friends to build a big tower."
[1302] The generated childcare records are automatically sent to the parents of the corresponding children via a server. The records are sent via email addresses registered in advance by the parents or via a dedicated childcare app. This allows parents to keep track of their child's daily activities in real time.
[1303] The server also analyzes video data in real time and monitors for dangerous behavior. For example, if the server detects a child engaging in dangerous behavior, such as jumping from a high place, it immediately issues an alert and notifies the childcare worker's device. This alert is sent as a push notification, allowing the childcare worker to respond quickly.
[1304] Furthermore, the server periodically compiles the childcare records and analyzes the growth of each child. Based on the results of this analysis, the generative AI model generates specific advice on childcare policies and activities, which are automatically distributed to the childcare worker's device. The childcare worker can use this advice to adjust and improve the childcare plan.
[1305] As a concrete example, when Taro is playing in the playground, a camera captures his movements, and the server uses facial recognition technology to identify him and record his actions. Computer vision technology tracks Taro playing with his friends using building blocks, and a generative AI model is used to create a childcare record stating, "Taro worked with his friends to build a big tower." This childcare record is automatically sent to Taro's mother. Furthermore, if the server detects Taro standing up high on the swing, it immediately sends an alert to the childcare worker's device, allowing the worker to respond quickly.
[1306] Prompt Sentence Examples
[1307] "How do you analyze real-time footage of children playing outside and send alerts to childcare workers if a particular child is exhibiting risky behavior?"
[1308] By utilizing this system, the workload of childcare workers will be significantly reduced, the safety of children will be improved, and high-quality childcare will be provided.
[1309] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1310] Step 1:
[1311] The server acquires video data in real time from high-resolution camera devices installed in classrooms and playgrounds. The input is the video stream from the camera, and the output is frame data stored on the server. Specifically, the server acquires data from the camera using the RTSP protocol and adds each frame to the processing queue.
[1312] Step 2:
[1313] The server analyzes the received video data using image recognition technology to detect the faces of the children. The input is the frame data from step 1, and the output is a feature vector of the detected face image. Specifically, it uses OpenCV and the Dlib library to detect faces in the frame and extract facial features.
[1314] Step 3:
[1315] The server uses facial recognition technology to compare the facial features detected with a pre-registered database to identify individual subjects (children). The input is the facial feature vector from step 2, and the output is the child's ID. Specifically, the server compares the facial feature vector with existing features in the database and identifies matching IDs based on cosine similarity.
[1316] Step 4:
[1317] The server uses computer vision technology to track the behavior of the identified children. The input is the child's ID from step 3 and the frame data from step 1, and the output is time-series data of the child's position and movement. Specific movements are tracked by detecting the child's skeleton using the OpenPose library.
[1318] Step 5:
[1319] The server automatically generates childcare records based on the tracking data using a generative AI model. The input is the time series data from step 4, and the output is a childcare record written in natural language. Specifically, the tracking data is input into the generative AI model, and a sentence describing the daily activities of the children is generated.
[1320] Step 6:
[1321] The server automatically distributes the generated childcare record to the corresponding guardian. The input is the childcare record from step 5, and the output is the information sent to the guardian's device. Specifically, the server refers to the guardian's contact database and sends the childcare record via email or a dedicated app.
[1322] Step 7:
[1323] The server analyzes the video data in real time and monitors for risky behavior. The input is the frame data from step 1, and the output is a warning notification when risky behavior is detected. Specifically, it uses computer vision technology to detect risky behavior patterns and sends a warning to the nursery teacher's device.
[1324] Step 8:
[1325] The server periodically compiles the childcare records and analyzes the children's growth. The input is the childcare record data from Step 5, and the output is the growth analysis results and advice. Specific operations include analyzing the compiled data and using a generative AI model to generate specific advice regarding childcare policies and activities.
[1326] Prompt Sentence Examples
[1327] "How do you analyze real-time footage of children playing outside and send alerts to childcare workers if a particular child is exhibiting risky behavior?"
[1328] This series of processes enables nursery school teachers to ensure the safety of children while saving time and effort, and to provide high-quality childcare.
[1329] (Application example 1)
[1330] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1331] In recent years, improving work efficiency and ensuring safety in factories have become extremely important issues. In particular, in manufacturing sites where many workers work simultaneously, there is a high risk of human error and dangerous behavior, and there is a growing demand for systems that can quickly detect and address such errors. Furthermore, manually creating work logs requires time and effort, making it difficult to immediately improve work efficiency. The present invention addresses these issues by providing advanced monitoring and automated work log generation to improve work efficiency and ensure safety in factories.
[1332] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1333] In this invention, the server includes means for capturing images of the work environment with a high-resolution camera, means for receiving the image data and identifying multiple individual objects using facial recognition technology, means for tracking the behavior of the identified individual objects and automatically generating a work log, means for automatically distributing the generated work log to relevant parties, means for issuing an alert when dangerous behavior is detected, and means for analyzing work efficiency based on the work log and providing relevant parties with points for improvement. This makes it possible to immediately grasp and analyze work efficiency within a factory and improve safety.
[1334] A "high-resolution camera" is a photographic device with high resolution that can record the work environment in detail.
[1335] "Work environment" refers to the place or area within a factory where work is carried out.
[1336] "Video data" is digital data containing visual information captured by a high-resolution camera.
[1337] "Facial recognition technology" is a biometric authentication technology for identifying people in video data.
[1338] "Individual Subject" refers to each identified worker within a factory.
[1339] "Behavior tracking" refers to following and recording the movements of identified individual subjects.
[1340] A "work log" is a digital record that records the actions and work of an individual.
[1341] "Automatically generating a work log" means that the server automatically creates a work log from the behavioral data of an individual subject.
[1342] "Stakeholders" refers to factory managers and supervisors who require work logs.
[1343] "Distribution" means sending the generated work log to the relevant parties.
[1344] "Issuing an alert" means sending a warning message when dangerous behavior is detected.
[1345] "Analyzing work efficiency" means analyzing the data in the work log to evaluate the efficiency of the work.
[1346] "Providing improvements" means presenting specific suggestions to stakeholders to improve work efficiency.
[1347] This invention is a system that improves work efficiency and ensures safety in factories by utilizing high-resolution cameras, facial recognition technology, and AI analysis technology. The system consists of a server, a camera, smart glasses, and a dedicated app.
[1348] System configuration
[1349] 1. High-quality camera:
[1350] The working environment within the factory is photographed in high resolution, and detailed video data is sent to a server in real time.
[1351] 2. Server:
[1352] The server processes the received video data and identifies the worker using facial recognition technology. Azure Face API, for example, is used for facial recognition. The actions of identified workers are tracked, and a work log is automatically generated based on that data. The work log contains detailed records of each worker's movements and tasks.
[1353] 3. Risky behavior monitoring and alerting:
[1354] The server analyzes video data and work logs in real time and immediately issues an alert if any dangerous behavior is detected. The alert is then sent to the smart glasses to alert the worker, enabling a rapid response.
[1355] 4. Work log distribution:
[1356] The generated work logs are automatically distributed to relevant parties (factory managers and supervisors) via email addresses or a dedicated app.
[1357] 5. Analyzing work efficiency and providing improvements:
[1358] The server analyzes work efficiency based on the work logs and provides specific improvements to the relevant parties, thereby optimizing work efficiency throughout the factory.
[1359] Hardware and software used
[1360] Hardware: high-resolution cameras, smart glasses (e.g., Microsoft HoloLens)
[1361] Software: Facial recognition AI (e.g., Azure Face API), behavioral analysis AI (e.g., TensorFlow, OpenCV), database (e.g., Azure SQL Database)
[1362] Example of operation
[1363] For example, if a worker is operating heavy machinery in an inappropriate manner, a high-resolution camera captures this, the server uses facial recognition technology to identify the worker, and AI determines that the operation is inappropriate. An alert is immediately sent to the smart glasses, prompting the worker to operate in a safe manner. At the same time, details of the operation are automatically generated as a work log and sent to the factory manager. Based on this work log, areas for improving work efficiency are analyzed and provided to relevant parties.
[1364] Example of input prompt for generative AI model
[1365] prompt:
[1366] Design a smart glasses application for factory monitoring with self-diagnosis function. Please pay attention to the following points:
[1367] 1. Tracking worker behavior in real time using cameras and smart glasses.
[1368] 2. Use facial recognition technology to identify individual workers.
[1369] 3. If risky behavior is detected, an alert is sent immediately.
[1370] 4. Automatically generate work logs and analyze areas for improvement.
[1371] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1372] Step 1:
[1373] High-resolution cameras capture the factory environment. The input is real-time footage from within the factory, and the output is high-resolution video data. Specifically, the cameras have a field of view that covers the entire factory, capturing detailed images of workers and machine operations.
[1374] Step 2:
[1375] The server receives the video data sent from the camera. The input is high-definition video data that is temporarily stored in the server. The output is video data stored in the server's storage.
[1376] Step 3:
[1377] The server analyzes the video data using facial recognition AI (Azure Face API) to identify multiple workers. The input is the saved video data, and the output is the identification information of each worker. Specifically, the AI analyzes the facial images in the video and identifies the worker's ID by comparing them with a pre-registered database.
[1378] Step 4:
[1379] The server tracks the actions of each identified worker. The input is identification information and real-time video data, and the output is individual behavioral patterns. Specifically, the server continuously tracks the movements of workers and records what tasks they are performing.
[1380] Step 5:
[1381] The server automatically generates work logs based on the behavioral data it tracks. The input is behavioral pattern data, and the output is a digital work log. Specifically, the AI organizes the recorded information and provides a detailed record of each worker's daily activities and their content.
[1382] Step 6:
[1383] The server automatically distributes the generated work log to the relevant parties. The input is the generated work log, and the output is the work log sent to the relevant parties' email addresses or a dedicated app. Specifically, the log file is sent to the specified recipient via a mail server or application.
[1384] Step 7:
[1385] The server analyzes real-time video data and issues an alert if it detects dangerous behavior. The input is real-time video data, and the output is an alert message sent to the smart glasses. Specifically, when the AI detects an abnormal behavior pattern, it immediately generates an alert and sends it to the worker's smart glasses.
[1386] Step 8:
[1387] The server analyzes work efficiency based on work logs and provides suggestions for improvement to the relevant parties. The input is daily work log data, and the output is an analysis report including suggestions for improvement. Specifically, the server compares the data with past work data, proposes specific improvements for improving work efficiency and safety, and sends them to the relevant parties.
[1388] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1389] This invention is a system that reduces the burden on childcare workers and ensures safe, high-quality childcare. Using high-resolution cameras, servers, and terminals, the system utilizes face recognition technology, emotion engines, and AI analysis technology to automatically generate childcare diaries, distribute them to parents, monitor and alert on risky behavior, analyze child development, and provide advice to childcare workers.
[1390] 1. Camera video capture
[1391] The server receives real-time video data from high-resolution cameras installed in classrooms and on the playground, which have a wide field of view and can accurately capture detailed activity within the school.
[1392] 2. Face recognition and individual object identification
[1393] The server analyzes the received video data using facial recognition AI technology to identify the children in the video. The facial recognition technology compares the facial images in the video with a pre-registered database to identify the ID of each individual subject (child).
[1394] 3. Behavior tracking and automatic generation of childcare diary
[1395] The server tracks the behavior of identified children and records their behavioral patterns (e.g., types of play, interpersonal relationships, and health status). Based on this, a daycare diary is automatically generated. The diary details each child's daily activities, interactions, and health status.
[1396] 4. Distribution of the childcare diary to parents
[1397] The server automatically distributes the generated childcare diary to the parents of the corresponding child via their email address or a dedicated app, allowing parents to keep track of their child's daily activities in real time.
[1398] 5. Monitoring risky behavior and issuing alerts
[1399] The server analyzes the video data in real time and monitors the children's risky behavior. If any risky behavior (e.g., jumping from a high place or rough play) is detected, an alert is sent immediately.
[1400] The terminal (childcare worker) receives this alert and can respond quickly.
[1401] 6. Growth analysis and advice for childcare workers
[1402] The server periodically compiles data from the childcare diary and analyzes the growth of each child. Based on the results of this analysis, it identifies individual growth trends and provides advice to childcare workers regarding specific childcare policies and activities. The advice is automatically sent to the childcare workers' devices and can be used in practice.
[1403] 7. Introducing the Emotion Engine
[1404] The server inputs the video and audio data into the emotion engine, which analyzes the child's emotional state based on their facial expressions and tone of voice. The emotion engine then identifies the emotion the child is feeling (e.g., joy, sadness, anger, anxiety).
[1405] Based on the emotional states identified by the emotion engine, the server adds emotional data to the childcare diary and development analysis, for example, recording how children felt during play and how specific activities affected their emotions.
[1406] The server generates and delivers advice to childcare workers and parents based on their emotional state, suggesting appropriate ways to respond. For example, if a child is feeling anxious, the server identifies the root cause and provides specific advice on how to support the child.
[1407] For example, when Taro is playing in the playground, a camera captures his movements, and the server identifies him using facial recognition technology and records his actions. His behavior as he plays with building blocks with his friends is tracked, and a diary entry is made stating, "Taro played with building blocks with his friends and worked together to build a big tower." This diary entry is automatically sent to Taro's mother's email address that afternoon.
[1408] Furthermore, if Taro attempts dangerous behavior such as standing up too high on the swing, the server will detect this and immediately send an alert to the nursery teacher's device, allowing the nursery teacher to immediately rush over to Taro and ensure his safety.
[1409] Furthermore, the emotion engine analyzes Taro's facial expressions and tone of voice and detects that he looks anxious. Based on this, the server sends advice to the nursery teacher's device, such as, "Taro is likely feeling anxious, so please talk to him a bit to reassure him."
[1410] The above is a specific embodiment for carrying out the present invention. This system allows childcare workers to reduce their workload, ensure the safety and emotional growth of children, and provide high-quality childcare.
[1411] The processing flow will be explained below.
[1412] Processing for an automatic childcare diary generation system that combines face recognition and emotion engine
[1413] From video capture to emotion analysis
[1414] Step 1: Receiving video data
[1415] The server receives video data in real time from high-resolution cameras installed in classrooms and playgrounds, and the video data is temporarily stored in dedicated storage.
[1416] Step 2: Face recognition and individual subject identification
[1417] The server inputs the received video data into a facial recognition AI model to identify the children in the video. The facial recognition technology compares the facial images in the video with a pre-registered database to identify the ID of each individual subject (child).
[1418] Step 3: Tracking behavior
[1419] The movements of identified individuals are tracked in real time, with the server recording the start and end of specific actions (e.g., playing, crafting, talking) with timestamps.
[1420] Step 4: Analysis by Emotion Engine
[1421] The server inputs video and audio data into an emotion engine, which analyzes the children's facial expressions and tone of voice to determine their emotional state (e.g., joy, sadness, anger, anxiety).
[1422] Step 5: Creating a childcare diary
[1423] The server automatically generates a daily diary for each child based on tracking data and emotion analysis data, detailing each child's daily activities, interactions, health status, and emotional state.
[1424] Delivery of childcare diary to parents
[1425] Step 1: Generate a journal
[1426] The server automatically distributes the generated childcare diary to the parents of the corresponding children via their email addresses or a dedicated app.
[1427] Monitoring and alerting for risky behavior
[1428] Step 1: Real-time monitoring of video data
[1429] The server analyzes the video data in real time and monitors the behavior of the children.
[1430] Step 2: Detect risky behavior
[1431] The server uses a vision AI model to detect dangerous behavior (e.g., jumping from a high place, rough play).
[1432] Step 3: Trigger an alert
[1433] If dangerous behavior is detected, the server immediately generates an alert and sends it to the childcare worker's device.
[1434] Step 4: Receive and respond to alerts
[1435] The device (childcare worker) receives the alert and takes action to respond in real time.
[1436] Growth analysis and advice for childcare workers
[1437] Step 1: Collecting childcare diary data
[1438] The server periodically compiles data from the daycare diary and analyzes growth and behavioral patterns over a set period of time.
[1439] Step 2: Conduct a growth analysis
[1440] The server generates and evaluates the growth trends of each child based on the analysis data.
[1441] Step 3: Generate Advice
[1442] Based on growth trends, the server generates specific advice for nursery teachers, including specific activities and methods for promoting children's growth.
[1443] Step 4: Delivering advice
[1444] The server delivers the generated advice to the childcare worker's device, who receives it and puts it into practice.
[1445] Further processing by the emotion engine
[1446] Step 1: Collecting emotion data
[1447] The server integrates the analytical data from the emotion engine into the childcare diary, allowing for an understanding of emotional fluctuations and trends.
[1448] Step 2: Generating sentiment-based advice
[1449] Based on the emotional data, the server generates advice suggesting appropriate responses for childcare workers and parents.
[1450] Step 3: Delivering emotional advice
[1451] The server delivers advice based on the generated emotions to the childcare worker's device, who then receives it and puts it into practice.
[1452] For example, when Taro is playing in the playground, a camera captures his movements, and the server identifies him using facial recognition technology and records his actions. His behavior as he plays with building blocks with his friends is tracked, and a diary entry is made stating, "Taro played with building blocks with his friends and worked together to build a big tower." This diary entry is automatically sent to Taro's mother's email address that afternoon.
[1453] Furthermore, if Taro attempts dangerous behavior such as standing up too high on the swing, the server will detect this and immediately send an alert to the nursery teacher's device, allowing the nursery teacher to immediately rush over to Taro and ensure his safety.
[1454] Furthermore, the emotion engine analyzes Taro's facial expressions and tone of voice and detects that he looks anxious. Based on this, the server sends advice to the nursery teacher's device, such as, "Taro is likely feeling anxious, so please talk to him a bit to reassure him."
[1455] The above is a specific embodiment for carrying out the present invention. This system allows childcare workers to reduce their workload, ensure the safety and emotional growth of children, and provide high-quality childcare.
[1456] Example 2
[1457] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1458] In childcare facilities, the workload of childcare workers is increasing, and it is often the case that they are not doing enough to ensure the safety of children, monitor their growth, or report to parents. It is also difficult to accurately grasp the emotional state of children and respond appropriately. This puts the quality of childcare at risk of declining.
[1459] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1460] In this invention, the server includes means for capturing images of classrooms and playgrounds with a high-resolution camera, means for receiving the image data and identifying multiple individual objects using facial recognition technology, means for tracking the behavior of the identified individual objects and automatically generating a childcare diary, means for automatically distributing the generated childcare diary to parents, means for issuing an alert when dangerous behavior is detected, means for analyzing growth based on the childcare diary and providing advice to childcare workers, means for analyzing emotional states from video data and audio data using an emotion engine, and means for suggesting appropriate response methods to childcare workers and parents based on the analyzed emotional states. This reduces the burden on childcare workers, enables management of the safety and growth of children, and enables understanding of emotional states and appropriate responses.
[1461] A "high-resolution camera" is a camera that has a wide field of view and can capture detailed video data in high resolution.
[1462] "Video data" refers to video information acquired from cameras installed in classrooms and playgrounds.
[1463] "Facial recognition technology" is a technology that identifies individuals by identifying faces in video data and comparing them with a pre-registered database.
[1464] "Individual object" is a term that refers to a specific person, especially a kindergartener.
[1465] "Behavioral tracking" is a method of monitoring and recording the movements and behavioral patterns of individual subjects in real time.
[1466] A "childcare diary" is a report that details the behavior, interactions, and health status of the children during a day's childcare activities.
[1467] "Guardian" is a term that refers to the parents or guardians of kindergarten children.
[1468] An "alert" is a warning message that is sent immediately when dangerous behavior or an abnormality is detected in real time.
[1469] "Growth analysis" is a method of analyzing children's growth patterns based on data from the nursery school diary and evaluating developmental trends.
[1470] "Advice" refers to specific methods of response and guidelines provided to childcare workers and parents based on growth analysis and emotional analysis.
[1471] An "emotion engine" is a technology that analyzes video and audio data to identify the emotional state of an individual subject.
[1472] "Response methods" are specific measures that, based on the analyzed data, indicate appropriate approaches in caring for and guiding children.
[1473] This invention is a system that reduces the burden on childcare workers and ensures safe, high-quality childcare. This system uses high-resolution cameras, servers, and terminals, and makes full use of face recognition technology, emotion engines, and AI analysis technology to automatically generate childcare diaries, distribute them to parents, monitor and alert on risky behavior, analyze child development, and provide advice to childcare workers.
[1474] Hardware and software used
[1475] server
[1476] Data servers with high processing power (e.g., Dell PowerEdge series)
[1477] Facial recognition AI technology (e.g., FaceNet)
[1478] Emotion engine (e.g. Microsoft Azure Emotion API)
[1479] Machine learning algorithms (e.g., Python's scikit-learn library)
[1480] camera
[1481] High-resolution network camera (e.g. AXIS P1368-E)
[1482] Terminal
[1483] Tablet or smartphone for childcare workers (e.g. iPad, Android tablet)
[1484] Explanation of program processing
[1485] High-definition camera video capture
[1486] The server receives real-time video data from high-resolution cameras installed in classrooms and playgrounds. These cameras have a wide field of view and can accurately capture detailed activity within the school. The video data is stored in the server's storage and used for analysis.
[1487] Identifying individual subjects using facial recognition technology
[1488] The server analyzes the received video data using facial recognition AI technology to identify the children in the video. The facial recognition technology compares the facial images in the video with a pre-registered database to identify the ID of each individual subject (child). This allows for accurate tracking of each child's activities.
[1489] Behavior tracking and automatic generation of childcare diary
[1490] The server tracks the identified children's behavior in real time and records their behavioral patterns (e.g., types of play, interpersonal relationships, and health status). Based on this behavioral data, a daycare diary is automatically generated, detailing each child's daily activities, interactions, and health status.
[1491] Delivery of childcare diary to parents
[1492] The generated daycare diary is automatically sent to the parents of the children via their email address or a dedicated app (e.g., Custom Nursery App). This allows parents to keep track of their children's daily activities in real time.
[1493] Monitoring risky behavior and issuing alerts
[1494] The server analyzes the video data in real time and monitors the children's risky behavior. For example, if dangerous behavior such as jumping from a high place or rough play is detected, an alert is immediately issued. A push notification is sent to the childcare worker's device (e.g., tablet), allowing the childcare worker to respond quickly.
[1495] Growth analysis and advice for childcare workers
[1496] The server periodically compiles data from the childcare diary and analyzes the growth of each child. Based on the results of this analysis, it identifies individual growth trends and provides advice to childcare workers regarding specific childcare policies and activities. This advice is automatically sent to devices and can be used by childcare workers in their daily work.
[1497] Introducing the Emotion Engine
[1498] The server inputs video and audio data into an emotion engine, which analyzes the emotional state of the child based on facial expressions, tone of voice, etc. The emotion engine identifies the child's emotions (e.g., joy, sadness, anger, anxiety). This emotional state is reflected in the childcare diary and growth analysis, and is used as part of advice for childcare workers and parents. Specifically, the server calls an emotion analysis API to obtain the analysis results, and a recommendation algorithm then suggests appropriate responses based on these results.
[1499] Specific examples
[1500] While Taro is playing in the playground, a high-definition camera captures his movements, and the server uses facial recognition technology to identify him and record his actions. His behavior as he plays with building blocks with his friends is tracked, and an entry is made in the daycare diary, stating, "Taro played with building blocks with his friends and worked together to build a big tower." This diary is automatically sent to Taro's mother's email address that afternoon. Furthermore, if Taro attempts the dangerous act of standing up high on the swing, the server detects this and immediately sends an alert to the nursery teacher's device, allowing the nursery teacher to quickly rush to Taro's aid and ensure his safety.
[1501] Prompt Sentence Examples
[1502] "High-resolution cameras and AI technology are used in daycare centers to track the behavior and emotions of children in real time, and facial recognition technology is used to identify each child individually. Please explain in detail how the AI will automatically generate daycare diaries, provide information to parents, monitor and alert for risky behavior, analyze growth, and provide advice to childcare workers."
[1503] The above is a specific embodiment for carrying out the present invention. This system allows childcare workers to reduce their workload, ensure the safety and emotional growth of children, and provide high-quality childcare.
[1504] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1505] Step 1:
[1506] Receiving video data
[1507] The server receives real-time video data of classrooms and playgrounds from high-resolution cameras. The input is the video stream from each camera, and the output is the received video data. Specifically, the server specifies the IP address of the camera and obtains the video stream using HTTP or RTSP protocol. This allows the server to monitor the current activities of the children in real time.
[1508] Step 2:
[1509] Identifying individual subjects through facial recognition
[1510] The server analyzes the received video data using facial recognition technology. The input is the video data and a database of pre-registered faces of the children, and the output is the ID of the identified child. Specifically, the server uses a facial recognition algorithm (e.g., FaceNet) to detect faces in the video data and compare them with the database. Based on the results of this comparison, the server identifies individual subjects (children).
[1511] Step 3:
[1512] Behavior tracking and data recording
[1513] The server tracks the behavior of identified children and records the data. The input is the child's ID and video data, and the output is a log of behavioral data. Specifically, the server uses a behavioral analysis algorithm to analyze the children's movements and location information and records it in chronological order. For example, a detailed behavioral log is generated, such as "Taro started playing in the sandbox at 9:15 and was talking to his classmate Hanako at 9:30."
[1514] Step 4:
[1515] Automatic generation of childcare diary
[1516] The server automatically generates a childcare diary based on the recorded behavioral data. The input is a log of behavioral data, and the output is a childcare diary. Specifically, the server organizes the behavioral logs based on a template and creates a diary-style report. This report details the children's daily activities, interactions, and health status.
[1517] Step 5:
[1518] Distribution of childcare journals
[1519] The server automatically distributes the generated childcare diary to the children's parents. The input is the childcare diary and the parents' contact information, and the output is the sent email or app notification. Specifically, the server uses an email sending API or push notification API to send the diary via the parents' email address or a dedicated app. This allows parents to keep track of their children's daily activities in real time.
[1520] Step 6:
[1521] Monitoring and alerting for risky behavior
[1522] The server analyzes video data in real time and monitors children's risky behavior. The input is video data, and the output is an alert notification. Specifically, the server uses a motion analysis algorithm to detect risky behavior (e.g., jumping from high places, rough play). If any risky behavior is detected, an alert notification is immediately pushed to the nursery teacher's device.
[1523] Step 7:
[1524] Providing growth analysis and advice
[1525] The server periodically compiles the data from the childcare diary and analyzes the growth of each child. The input is the childcare diary data, and the output is a growth analysis report and advice. Specifically, the server uses a machine learning algorithm to analyze past behavioral data and predict the child's growth trends. Based on this, it provides childcare workers with advice on appropriate childcare policies and activities.
[1526] Step 8:
[1527] Emotion analysis using an emotion engine
[1528] The server inputs video and audio data into an emotion engine to analyze the emotional state of the children. The input is video and audio data, and the output is the emotion analysis results. Specifically, the server calls an emotion analysis API (e.g., Microsoft Azure Emotion API) to analyze the children's facial expressions and tone of voice to identify their emotional state. The results of this analysis are reflected in the childcare diary and growth analysis, and are used as part of advice for childcare workers and parents.
[1529] (Application example 2)
[1530] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1531] There are challenges in monitoring worker safety and managing work efficiently in factories. In particular, it is difficult to grasp workers' behavior and emotional state in real time, quickly detect dangerous behavior, and take countermeasures. Furthermore, there are insufficient systems to provide appropriate advice based on workers' emotional state, which hinders workers' stress reduction and efficient work performance.
[1532] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1533] In this invention, the server includes means for capturing images of an area with a high-resolution camera, means for receiving the image data and identifying multiple individual subjects using facial recognition technology, means for tracking the behavior of the identified individuals and automatically generating records, means for automatically distributing the generated records to relevant parties, means for issuing warnings when dangerous behavior is detected, means for analyzing growth or progress based on the records and providing advice to the person in charge, and means for analyzing the emotional state of workers in real time using a portable display device worn by the workers and providing appropriate advice, thereby enabling safety monitoring of workers in factories and efficient work management.
[1534] A "high-resolution imaging device" is a high-resolution camera device that has a wide field of view and can capture detailed images.
[1535] "Area" refers to the space or location where a particular activity takes place, including, for example, a work area in a factory or an area requiring safety monitoring.
[1536] "Facial recognition technology" is a technology that can identify individual subjects by analyzing facial images in video and comparing them with a pre-registered database.
[1537] "Distinct objects" refer to specific people or objects that the system needs to identify, including, for example, factory workers.
[1538] "Behavioral tracking" refers to tracking the movements and actions of individual subjects in real time and recording that data.
[1539] "Automatic generation of records" refers to automatically creating diaries and reports based on tracked behavioral data.
[1540] "Stakeholders" refers to people who need the system's output data, including, for example, factory managers and the workers' families.
[1541] "Sending a warning" means that the system will send an alert in real time when it detects dangerous behavior or an unexpected situation.
[1542] "Analyzing growth or progress" refers to analyzing the progress or areas for improvement of an individual subject based on recorded data.
[1543] "Providing advice to the person in charge" refers to notifying the person in charge of specific instructions or suggestions based on the analysis results.
[1544] A "portable display device" is a device with a small display that can be worn and used by a worker.
[1545] "Analyzing emotional states" is a technology that identifies emotions from video and audio data and understands those states.
[1546] The system of this invention uses high-resolution imaging devices, servers, facial recognition technology, emotion analysis engines, and portable display devices (hereinafter referred to as smart glasses) to monitor the safety of workers within factories and achieve efficient business management.
[1547] Hardware and software used:
[1548] High-resolution camera equipment: A high-resolution camera equipment with a wide field of view that can capture detailed images.
[1549] Server: A cloud or local server that receives and analyzes data in real time.
[1550] Facial recognition technology: Technology that analyzes facial images in video and identifies individual subjects (e.g., Amazon Rekognition, Microsoft Azure Face API).
[1551] Emotion analysis engine: Technology that identifies emotions from video and audio data and understands their state (e.g., Affectiva, IBM Watson Tone Analyzer).
[1552] Portable display devices (smart glasses): Devices worn by workers that display information in real time (e.g., Google Glass, Microsoft HoloLens).
[1553] Real-time data analysis engine: Technology that performs real-time analysis of data (e.g., Apache Kafka, Apache Flink).
[1554] The server uses high-resolution camera equipment to capture images of the factory work area. The captured image data is sent to the server in real time. The server's facial recognition technology analyzes the image data and identifies multiple individual subjects (factory workers). The actions of the identified individuals (workers) are tracked by the server, and an activity record is automatically generated.
[1555] The recorded behavioral data is automatically distributed to relevant parties (such as factory managers) via email or a dedicated application. If the system detects any dangerous behavior, it will immediately issue a warning. Workers wearing the smart glasses will receive a warning message in real time, allowing them to respond quickly.
[1556] Furthermore, the server analyzes the progress of workers based on the recorded data and provides advice to managers, thereby improving work efficiency and safety. The emotion analysis engine analyzes the emotional state of workers in real time and provides appropriate advice. For example, if a worker is in a high-stress state, the smart glasses will display advice such as "You need to take a break."
[1557] As a concrete example, if a factory worker is wearing smart glasses while working at height, a camera will capture his movements and send the data to a server, where facial recognition technology will identify the worker and begin tracking his behavior and analyzing his emotions. If the worker makes an unstable movement, the server will determine this as a dangerous behavior and display a warning on the smart glasses saying, "Your current movement is dangerous. Please move to a safe location."
[1558] Prompt Sentence Examples
[1559] Design a system to monitor the safety of factory workers using facial recognition technology, emotion engines, and AI analysis technology. Specifically, include the following features:
[1560] 1. Facial recognition function to identify workers using smart glasses
[1561] 2. Tracking worker behavior and detecting dangerous behavior
[1562] 3. Monitor emotional states using a sentiment analysis engine
[1563] 4. Alert workers when they are in a dangerous situation
[1564] 5. Providing advice on safety and stress reduction
[1565] Specific examples of technologies to consider using include Amazon Rekognition, Affectiva, and Apache Flink.
[1566] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1567] Step 1:
[1568] High-resolution camera equipment captures video data of the factory work area, including the overall situation within the factory and the actions of workers.
[1569] Input: Factory area video
[1570] Output: High-resolution video data
[1571] Step 2:
[1572] The server receives the video data in real time, and then transmits the video data to the server, where the analysis process begins.
[1573] Input: High-resolution video data
[1574] Output: Video data required for analysis
[1575] Step 3:
[1576] The server uses facial recognition technology to identify workers in the video data by comparing facial images in the video with a pre-registered database to identify individual targets.
[1577] Input: Video data, face database
[1578] Output: ID of the identified individual (worker)
[1579] Step 4:
[1580] The server tracks the behavior of identified individuals, recording the movements and actions of workers in real time.
[1581] Input: ID of identified individual, video data
[1582] Output: Tracked behavior data
[1583] Step 5:
[1584] The server automatically generates records based on the collected behavioral data, including the worker's daily activities and tasks.
[1585] Input: Tracked behavioral data
[1586] Output: Automatically generated activity log
[1587] Step 6:
[1588] The server automatically distributes the generated records to the relevant parties via email or a dedicated application.
[1589] Input: Auto-generated event log, stakeholder list
[1590] Output: Records sent to interested parties
[1591] Step 7:
[1592] The server analyzes the video data in real time to detect dangerous behavior and issues a warning if any.
[1593] Input: Real-time video data
[1594] Output: Risky behavior warning
[1595] Step 8:
[1596] Workers wearing smart glasses receive real-time warning messages that are immediately displayed to workers, enabling them to take prompt action.
[1597] Input: Risky behavior warning message
[1598] Output: Real-time displayed warnings
[1599] Step 9:
[1600] The server analyzes progress based on the activity records and provides advice to the administrator, thereby improving work efficiency and safety.
[1601] Input: Action log
[1602] Output: Progress analysis and advice
[1603] Step 10:
[1604] The emotion analysis engine analyzes the emotional state of the worker in real time, and based on the analysis results, appropriate advice is generated and displayed on the smart glasses.
[1605] Input: Video data, audio data
[1606] Output: Emotional state analysis and advice
[1607] For example, if the server detects unstable movements of a worker working at height, it will display a warning on the smart glasses saying, "The current movement is dangerous. Please move to a safe location." This warning allows the worker to take immediate action to ensure safety.
[1608] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1609] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1610] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1611] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1612] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1613] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1614] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1615] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1616] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1617] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1618] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1619] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1620] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1621] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1622] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1623] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1624] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1625] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1626] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1627] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1628] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1629] The following is further disclosed regarding the above embodiment.
[1630] (Claim 1)
[1631] A means to take pictures of classrooms and playgrounds with high-resolution cameras,
[1632] means for receiving the video data and identifying a plurality of distinct subjects using facial recognition technology;
[1633] A means for tracking the behavior of the identified individual subjects and automatically generating a childcare diary;
[1634] A means for automatically distributing the generated childcare diary to parents,
[1635] a means for issuing an alert when a risky behavior is detected;
[1636] A method to analyze growth based on the childcare diary and provide advice to childcare workers.
[1637] A system including:
[1638] (Claim 2)
[1639] 10. The system of claim 1, further comprising means for processing the video data in real time to immediately detect risky behavior.
[1640] (Claim 3)
[1641] The system according to claim 1, further comprising means for distributing the childcare diary via an email address or a dedicated app.
[1642] "Example 1"
[1643] (Claim 1)
[1644] A means to photograph classrooms and playgrounds with high-resolution photography equipment,
[1645] means for receiving the video data and identifying a plurality of individual objects using image recognition technology;
[1646] A means for tracking the behavior of the identified individual subjects using computer vision technology and automatically generating a childcare record;
[1647] A means for automatically distributing the generated childcare records to parents;
[1648] means for issuing a warning when a dangerous behavior is detected;
[1649] A method for analyzing growth based on childcare records and providing advice to caregivers.
[1650] a means for processing the video data frame by frame and comparing it with multiple databases;
[1651] a means for tracking the position and movement of the child using computer vision technology;
[1652] A means of storing tracking data chronologically and performing real-time activity mapping;
[1653] A means for creating childcare records in natural language using a generative AI model;
[1654] A means to automatically sort contact information for each parent,
[1655] A means for generating advice regarding childcare policies and activities based on the activity analysis results;
[1656] A system including:
[1657] (Claim 2)
[1658] 10. The system of claim 1, further comprising means for processing the video data in real time to immediately detect risky behavior.
[1659] (Claim 3)
[1660] 10. The system of claim 1, further comprising means for distributing the childcare records via an email address or a dedicated application.
[1661] "Application Example 1"
[1662] (Claim 1)
[1663] A means of capturing images of the work environment with a high-resolution camera,
[1664] means for receiving the video data and identifying a plurality of distinct subjects using facial recognition technology;
[1665] a means for tracking the actions of the identified individuals and automatically generating a work log;
[1666] A means to automatically distribute the generated work log to relevant parties;
[1667] a means for issuing an alert when a risky behavior is detected;
[1668] A method to analyze work efficiency based on work logs and provide relevant parties with points for improvement.
[1669] A system including:
[1670] (Claim 2)
[1671] 10. The system of claim 1, further comprising means for processing the video data in real time to immediately detect risky behavior.
[1672] (Claim 3)
[1673] 2. The system according to claim 1, further comprising means for distributing the work log via an email address or a dedicated app.
[1674] "Example 2: Combining Emotion Engines"
[1675] (Claim 1)
[1676] A means to take pictures of classrooms and playgrounds with high-resolution cameras,
[1677] means for receiving the video data and identifying a plurality of distinct subjects using facial recognition technology;
[1678] A means for tracking the behavior of the identified individual subjects and automatically generating a childcare diary;
[1679] A means for automatically distributing the generated childcare diary to parents,
[1680] a means for issuing an alert when a risky behavior is detected;
[1681] A method to analyze growth based on the childcare diary and provide advice to childcare workers.
[1682] A means for analyzing emotional states from video data and audio data using an emotion engine;
[1683] A means for suggesting appropriate ways of responding to childcare workers and parents based on the analyzed emotional state;
[1684] A system including:
[1685] (Claim 2)
[1686] 10. The system of claim 1, further comprising means for processing the video data in real time to immediately detect risky behavior.
[1687] (Claim 3)
[1688] The system according to claim 1, further comprising means for distributing the childcare diary via an email address or a dedicated app.
[1689] "Application example 2 when combining emotion engines"
[1690] (Claim 1)
[1691] means for photographing the area with a high-quality photographing device;
[1692] means for receiving the video data and identifying a plurality of distinct subjects using facial recognition technology;
[1693] means for tracking the behavior of identified individuals and automatically generating records;
[1694] A means for automatically distributing the generated records to interested parties;
[1695] a means for issuing an alert when dangerous behavior is detected;
[1696] A means of analyzing growth or progress based on the records and providing advice to the person in charge;
[1697] A means for analyzing the emotional state of workers in real time and providing appropriate advice using a portable display device worn by the workers;
[1698] A system including:
[1699] (Claim 2)
[1700] 10. The system of claim 1, further comprising means for processing the video data in real time to immediately detect risky behavior.
[1701] (Claim 3)
[1702] 10. The system of claim 1, further comprising means for distributing the records via email or a dedicated application. [Explanation of symbols]
[1703] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. A means to take pictures of classrooms and playgrounds with high-resolution cameras, means for receiving the video data and identifying a plurality of distinct subjects using facial recognition technology; A means for tracking the behavior of the identified individual subjects and automatically generating a childcare diary; A means for automatically distributing the generated childcare diary to parents, a means for issuing an alert when a risky behavior is detected; A method to analyze growth based on the childcare diary and provide advice to childcare workers. A system including:
2. 10. The system of claim 1, further comprising means for processing the video data in real time to immediately detect risky behavior.
3. The system according to claim 1, further comprising means for distributing the childcare diary via an email address or a dedicated app.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A