system

A system using facial recognition and real-time video analysis efficiently locates lost children in large facilities, ensuring quick reunions and alleviating guardian anxiety.

JP2026069011APending Publication Date: 2026-04-23SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-11
Publication Date
2026-04-23

AI Technical Summary

Technical Problem

Large facilities face challenges in quickly and efficiently identifying lost children, which causes anxiety for guardians and affects facility reputation, due to insufficient facial recognition and location accuracy in existing systems.

Method used

A system combining facial recognition technology with real-time video analysis to track and locate lost children, utilizing surveillance devices for data acquisition, matching, and notifying nearby staff through communication devices.

Benefits of technology

Enables rapid identification and safe reunion of lost children, reducing psychological burden on guardians and improving facility safety and reputation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026069011000001_ABST
    Figure 2026069011000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A means of acquiring and storing facial recognition data, A means of comparing the provided photo data with the stored face recognition data to perform matching, A means for receiving real-time video data from a monitoring device and determining the location by comparing it with the matching result, A means of notifying the communication device of a nearby person in charge of the identified location information, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In a large facility, a child getting lost is a major source of anxiety for guardians and also a safety management issue for facility operators. The occurrence of such lost children not only increases the psychological burden on guardians and children but may also affect the reputation of the facility. Therefore, there is a need to provide a system that can efficiently and quickly identify lost children and enable them to reunite safely.

Means for Solving the Problems

[0005] This invention provides a system that combines multiple means to quickly identify a lost child and safely reunite them. Specifically, it includes means for acquiring and storing facial recognition data, means for comparing provided photographic data with stored facial recognition data to perform matching, means for receiving real-time video data from a monitoring device and determining the location by comparing it with the matching result, and means for notifying the communication device of a nearby person in charge of the identified location. This makes it possible to efficiently identify a lost child and notify their guardian, thereby reducing their psychological burden.

[0006] "Facial recognition data" refers to information that quantifies the characteristics of an individual's face and stores it as identifiable data.

[0007] "Provided photo data" refers to still images or video images provided by the user, which are used in the facial recognition system.

[0008] "Matching" is the process of comparing the provided photo data with stored facial recognition data to determine if they match.

[0009] A "surveillance device" refers to a camera device that is placed within a facility and acquires video footage in real time.

[0010] "Real-time video data" refers to video data that is captured by surveillance equipment and processed immediately.

[0011] "Means of determining location" refers to methods or systems that use matched facial recognition data and real-time video data to determine the current location of a specific object.

[0012] A "communication device" is a device that can send and receive information and is used to transmit notifications and communications. [Brief explanation of the drawing]

[0013] [Figure 1]This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]

[0014] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.

[0015] First, the terms used in the following description will be explained.

[0016] In the following embodiments, the labeled processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0017] In the following embodiments, the labeled RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0018] In the following embodiments, the labeled storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.

[0019] In the following embodiments, the labeled communication I / F (Interface) is an interface including a communication processor and an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark), and the like.

[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0021] [First Embodiment]

[0022] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0023] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0024] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0025] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0026] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0028] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0029] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0030] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0031] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0032] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0033] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0034] This invention is a system for quickly identifying lost children in large facilities and safely reuniting them. This system combines facial recognition technology with real-time video analysis to efficiently track lost children and support appropriate responses.

[0035] server

[0036] The server stores visitor facial recognition data acquired at each entrance gate in a database. In the event of a lost child, it receives a photo of the child provided by a parent or guardian and matches it against the existing database using a facial recognition algorithm. Next, it analyzes real-time video data received from surveillance devices to search for a match with the facial recognition data. Once the location of the target is identified, it generates location information and notifies nearby employees' communication devices.

[0037] terminal

[0038] The terminal is a device that receives notifications sent from the server. These notifications include the location information of a lost child. Based on this information, the terminal provides navigation instructions to help employees quickly locate the lost child. This process allows employees to quickly find and safely protect the lost child.

[0039] User

[0040] When a child is found missing, the user (parent) immediately provides a recent photo of their child through the local support desk or support app. This data enables rapid matching in the system's facial recognition process. In addition, parents can check the latest location information provided through the employee's device, giving them peace of mind.

[0041] As a concrete example, consider a case at a theme park. When a parent realizes their child is lost, they immediately register the child's photo at the support desk. The server performs facial recognition and detects a match from real-time video. The server locates the child's position, a notification is sent to an employee's terminal, and the employee quickly goes to the lost child, ensuring a safe reunion between parent and child.

[0042] Thus, the system of the present invention enables a rapid response in the event of a child getting lost, providing a safe and secure environment for both facility operators and users.

[0043] The following describes the processing flow.

[0044] Step 1:

[0045] When a user enters the facility, a facial recognition camera captures the visitor's face and generates facial recognition data. The server receives this data and stores it in a database.

[0046] Step 2:

[0047] The user (parent) reports the child getting lost to the support desk and provides a recent photo of the child's face. The server receives this photo data and starts the matching process using a facial recognition algorithm against existing data in the database.

[0048] Step 3:

[0049] The server receives video data from the monitoring device in real time and searches for matching faces through a facial recognition process. This involves algorithmic image analysis and comparison of facial data.

[0050] Step 4:

[0051] When the server identifies a matching target, it generates the child's current location data. This location information is then formatted as map data.

[0052] Step 5:

[0053] The server sends location information as a notification to the terminals of nearby employees. The notification includes location details and instructions on how to get there.

[0054] Step 6:

[0055] The device receives a notification, and the employee quickly moves to the child's location based on the instructions. The device provides navigation information to support efficient movement.

[0056] Step 7:

[0057] An employee finds the lost child, safely takes them into custody, and then guides them back to their parents. This successful event is then fed back from the terminal to the server, completing the lost child tracking process.

[0058] (Example 1)

[0059] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0060] For facility managers, quickly identifying lost children within the facility and safely reuniting them with their guardians is a critical challenge. Current systems lack sufficient facial recognition and location accuracy, making rapid response difficult in some cases. Furthermore, there is a need for more efficient information notification and navigation.

[0061] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0062] In this invention, the server includes means for acquiring and storing visitor facial data, means for comparing provided image data with stored facial data to confirm a match, and means for receiving real-time video data from a monitoring device and determining the location by comparing it with the matching result. This makes it possible to efficiently identify lost children and quickly transmit their location information.

[0063] "Visitor facial data" refers to a collection of identifying information, including facial images and characteristics of individuals entering the facility.

[0064] "Image data" refers to visual information that is stored electronically and can be processed by a system.

[0065] "Verification" is the process of comparing the provided data with the stored data to confirm whether they match.

[0066] "Surveillance equipment" refers to devices installed inside and outside a facility that capture and transmit video in real time.

[0067] "Real-time video data" refers to video information captured by cameras and sensing devices at a given moment.

[0068] A "match result" is the result of the matching process confirming that the provided data is identical or similar to existing data.

[0069] "Means of location determination" refers to a mechanism within a system for determining the precise location of a person or object.

[0070] "Location information" refers to detailed data about a specific location, and may include maps and coordinate information.

[0071] A "communication device" is an electronic device used to send and receive information, and mainly includes mobile terminals and computers.

[0072] "Navigation instructions" are guidelines that show the route and steps necessary to reach a specific location.

[0073] "Re-authentication to confirm safe reunion" is a facial recognition process performed when a parent and child meet face-to-face again, in order to prevent misidentification.

[0074] The present invention will now be described in terms of embodiments for carrying it out. This system is designed to quickly identify lost children in facilities and safely reunite them with their guardians. The process primarily utilizes facial recognition technology, and real-time location tracking and notification are performed based on that information.

[0075] Hardware and software to be used

[0076] The server acquires visitor facial data via cameras installed at each entrance gate within the facility. For this purpose, it uses image processing software such as OpenCV and TENSORFLOW®. This software extracts features from the acquired facial data and stores them in a database. The server also receives video data in real time from surveillance equipment and performs comparison and matching by analyzing it with a facial recognition algorithm. This process requires high system computing power and applies advanced image processing technology.

[0077] A terminal is a device that receives notifications sent from the server. Specifically, this includes smartphones and tablets, and is operated through a dedicated application compatible with the ANDROID® OS and iOS. The terminal uses the Google® Maps API and other tools to visually display location information and provide navigation instructions to staff.

[0078] When a user (parent) realizes their child is lost, they provide the latest image data of their child through the support desk or support app. This data is a crucial criterion for the facial recognition process on the server. Furthermore, users can confidently cooperate in the search for their child based on the location information received through their device.

[0079] Specific example

[0080] As a concrete example, consider its use in a theme park. If a parent gets separated from their child, a photo of the child is immediately provided at the support desk. This photo is cross-referenced with a server database, and the child's current location is determined through real-time video analysis from surveillance cameras. Based on this information, the location information is sent to employee terminals, enabling a quick reunion between parent and child.

[0081] Example of a prompt

[0082] An example of a prompt to input into a generative AI model is, "Please describe the process of rapidly locating and reuniting lost children using a facial recognition system in a large facility." This prompt is expected to prompt the generative AI model to generate a detailed description of the system's processes and functions.

[0083] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0084] Step 1:

[0085] The server uses cameras installed at each entrance gate of the facility to acquire visitor facial data. The input is real-time video data. The server extracts facial information from this video using a face detection algorithm and stores the features in a database. The output is the facial feature data stored in the database. This process allows for the temporary storage of visitor facial data, preparing it for later matching.

[0086] Step 2:

[0087] When a user (parent) realizes their child is missing, they provide the latest photo of the child to the server via the support desk or a dedicated app. The input to this process is the child's photo data, which the server receives and immediately analyzes using a facial recognition algorithm. The output is a matching result, which confirms that the image matches existing data in the database.

[0088] Step 3:

[0089] The server receives video data in real time from the surveillance camera. The input is a video stream from the surveillance equipment. The server analyzes this video data and compares it with the matching results obtained in step 2 to perform face recognition. This process confirms the match between the face information in the video and the existing data, and identifies the location. The output is the identified location information.

[0090] Step 4:

[0091] The server formats the identified location information in JSON format or another suitable format and sends a notification to the terminals of nearby staff members. The input is the identified location data. The notification includes location details, allowing employees to quickly locate the lost child. The output is navigation instructions displayed on the terminal.

[0092] Step 5:

[0093] The terminal receives notifications from the server and provides visual navigation to employees through its built-in application. The input is location information from the server, and the location is displayed on a map using the Google Maps API. This allows employees to efficiently navigate within the facility. The output consists of a map and a guidance screen.

[0094] Step 6:

[0095] The user (parent) is guided by an employee to a designated location where they are reunited with their child. After the reunion, a re-authentication process using a terminal is performed, and facial recognition is used to confirm safety. The input consists of facial data of both the child and the parent. The server re-authenticates this data and confirms that the parent and child match, preventing misidentification and providing a secure environment. The output is the authentication confirmation result.

[0096] (Application Example 1)

[0097] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0098] In large facilities, there is a need to quickly and reliably locate lost children and safely reunite them with their guardians. However, conventional systems take a long time to pinpoint the location of lost children, making it difficult to reassure guardians. This invention aims to solve this problem and provide peace of mind and safety to facility users.

[0099] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0100] In this invention, the server includes means for acquiring and storing facial recognition data, means for comparing provided visual data with stored facial recognition data to determine a match, means for receiving real-time video information from a monitoring device and determining its location by comparing it with the match determination result, means for notifying the determined location information to the mobile terminal of a nearby staff member, and means for providing a route to the determined location. This enables the rapid identification of a lost child's location and a quick response by staff members.

[0101] "Facial recognition data" refers to information recorded digitally from an individual's face and used for identification purposes.

[0102] "Visual data" refers to visual information, such as images and photographs, recorded as digital data.

[0103] "Matching" is the process of comparing the provided visual data with stored facial recognition data to determine whether they are the same person.

[0104] "Monitoring equipment" is a general term for devices that can capture video in real time and process it as data.

[0105] "Real-time video information" refers to data that can instantly capture and process the current situation as video.

[0106] "Location information" refers to data that indicates the current geographical location of a specific object.

[0107] A "mobile terminal" is a portable device capable of sending, receiving, and displaying information.

[0108] A "travel path" is information that indicates the optimal or designated path to a specific location.

[0109] To implement this invention, it is necessary to build a system that can quickly and reliably find lost children in large facilities. This system works effectively by combining facial recognition and real-time video analysis technologies.

[0110] The server acquires and stores facial recognition data in a database when visitors enter. Existing facial recognition software is used for the facial recognition technology. If a child gets lost, the parent provides the latest photo via a device, and the server uses this visual data to determine if it matches. The facial recognition algorithm compares the stored facial recognition data with the provided visual data to identify the match.

[0111] Meanwhile, monitoring devices are installed in multiple locations within the facility and constantly transmit real-time video information to a server. The server analyzes this video information and identifies videos containing matching facial recognition data. The identified location information is then notified to a mobile terminal. This allows staff to efficiently understand the route to the identified location and quickly guide the lost person to their location.

[0112] As a concrete example, consider a case where a child gets lost in a shopping mall. The parent sends a photo of their face via a smartphone app. The server performs facial recognition, determines the child's location from the video data in real time, and sends the location information, such as "The child is near the food court on the 3rd floor," to the staff member's smartphone. This allows the staff member to quickly go to the scene and guide the lost child back to their parents.

[0113] A concrete example of a prompt for a generated AI model would be: "Explain the steps and techniques necessary to safely reunite a lost child with their parents in a shopping mall. Include specific facial recognition algorithms, video analysis methods, and location information notification processes." Based on this prompt, it is possible to explain the system's detailed operation flow and technical background.

[0114] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0115] Step 1:

[0116] The server stores visitor facial recognition data acquired at the entrance gate into a database. The input is video from surveillance cameras, and the output is digital facial data obtained using a facial recognition algorithm. Faces are detected from the video, feature points are extracted, and stored.

[0117] Step 2:

[0118] When a child goes missing, the user sends the latest visual data of the child from their device to the server. The input is a photograph of the child, and the output is the visual data of that photograph. The data is securely transferred from the device to the server.

[0119] Step 3:

[0120] The server determines whether the received visual data matches the existing face recognition data in the database. The input is the provided visual data and the existing face recognition data, and the output is the matching result. The server then uses a face recognition algorithm to compare each set of data.

[0121] Step 4:

[0122] Real-time video information is transmitted from the monitoring device to the server. The input is live video from the monitoring device, and the output is real-time video information. The server constantly receives and analyzes this real-time video data.

[0123] Step 5:

[0124] The server compares the matching results with real-time video information to pinpoint the lost child's location. The input is the matching results and video information, and the output is the identified location. The entire face visible in the video is analyzed, and the location of the matched face is displayed on a map.

[0125] Step 6:

[0126] The server notifies the assigned person's mobile device of the identified location. The input is location information, and the output is a notification message sent to the mobile device. A message containing navigation information to that location is sent to the assigned person.

[0127] Step 7:

[0128] The terminal guides the assigned person along a route based on the location information it receives. The input is a notification message from the server, and the output is a route suggestion for the assigned person. The route displayed will help the assigned person reach the designated location quickly.

[0129] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0130] This invention is a system that further reduces the psychological burden on parents and facility staff and enables a quick and effective response by incorporating an emotion recognition engine into a lost child tracking system. This system is based on real-time video analysis using facial recognition technology and has the function of dynamically adjusting notification content through emotion recognition.

[0131] server

[0132] The server stores and manages facial recognition data collected at the facility's entrance gates. In the event of a lost child, the server compares the provided child's photo with information in the database using a facial recognition algorithm. Furthermore, the server analyzes real-time video transmitted by monitoring devices and identifies the child's location through facial recognition.

[0133] The emotion recognition engine based on this invention determines the current emotional state based on voice and facial expression data received from parents or facility staff. The server analyzes this emotional information and adjusts the content and format of notifications. For example, if a parent is showing strong anxiety, more detailed location information and reassuring messages are added.

[0134] terminal

[0135] The device receives notifications from the server and provides navigation information to help facility staff quickly locate the lost child. Furthermore, based on the results of the emotion recognition engine, it can present encouraging messages and specific action plans to staff, enabling them to respond calmly.

[0136] User

[0137] When a child goes missing, the user (parent) provides a photo of their child through the support desk or support app and appropriately communicates their emotional state. This allows the system to respond in a way that takes the parent's psychological burden into consideration.

[0138] As a concrete example, if a parent loses sight of their child at a theme park, the emotion recognition engine analyzes the parent's level of anxiety from their tone of voice and facial expression. The server takes this into account and includes additional information in the notification sent to the facility staff's terminal, providing reassurance. This aims to facilitate a smooth reunion between parent and child and create a safe and stress-free environment throughout the facility.

[0139] Thus, the system of the present invention, by combining facial recognition and emotion recognition, realizes a psychologically reassuring approach to lost children, bringing benefits to both facility management and users.

[0140] The following describes the processing flow.

[0141] Step 1:

[0142] When a user enters the facility, a facial recognition camera automatically captures a photograph of the visitor's face. The server receives this facial recognition data and securely stores it in a database.

[0143] Step 2:

[0144] If a user (parent) finds their lost child, they provide a recent photo of the child's face via the support desk or mobile app. The server receives this photo and begins matching it with existing facial recognition data in its database.

[0145] Step 3:

[0146] The server receives video data in real time from the monitoring device and uses a facial recognition algorithm to detect matches between photos and videos. If a match is detected, the server identifies the child's current location and compiles that information.

[0147] Step 4:

[0148] The emotion recognition engine analyzes the user's emotions through voice input and a real-time video interface. Based on these results, the server understands the user's state and adjusts the response accordingly.

[0149] Step 5:

[0150] The server generates an appropriate response message along with the child's precise location information and sends a notification to the terminal of nearby facility staff. The notification includes the parent's emotional state and provides instructions and messages appropriate to the situation.

[0151] Step 6:

[0152] The terminal (staff member's device) receives the notification and acts quickly according to the displayed details and navigation. The terminal also provides messages to staff members to encourage them to act calmly.

[0153] Step 7:

[0154] The user (parent) is reassured after receiving a direct update from facility staff and confirming that their child is safe. The emotion recognition engine confirms the parent's state of reassurance and feeds this information back to the server. The server records the completion of the process within the system.

[0155] (Example 2)

[0156] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0157] When a child gets lost in a public facility, it places a significant psychological burden not only on the parents but also on the facility staff. Conventional lost-child tracking systems often struggle to provide a quick response or alleviate parents' anxiety, which can result in a decrease in confidence in the safety of the facility. This invention aims to improve this situation and enable the rapid discovery of lost children while reducing the psychological burden.

[0158] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0159] In this invention, the server includes means for acquiring and storing facial recognition data, means for comparing and matching provided image data with stored facial recognition data, means for receiving real-time image data from a monitoring device and determining the location by comparing it with the matching result, means for analyzing emotional data from parents and staff and adjusting the notification content, means for transmitting the adjusted notification content to the communication device of a nearby staff member, and means for guiding staff members based on geographical information. This makes it possible to quickly locate children and provide psychological support to parents and staff.

[0160] "Facial recognition data" refers to data in which the characteristics of an individual's face are extracted as digital information and stored in an identifiable format.

[0161] "Image data" refers to still or dynamic visual information stored in digital format.

[0162] "Matching" is the process of comparing the provided data with the stored data and evaluating the degree of similarity.

[0163] A "monitoring device" is a device that monitors a designated environment in real time and acquires necessary image and audio data.

[0164] "Real-time image data" refers to digital data that instantly captures and transmits video footage occurring under certain circumstances.

[0165] "Emotional data" refers to information about emotional states extracted from voice, facial expressions, and actions.

[0166] "Notification content" refers to a message structured to convey specific information to a recipient.

[0167] A "communication device" is a device that combines hardware and software for sending and receiving information.

[0168] "Geographic information" refers to data related to location or place, and is usually presented visually on a map.

[0169] This invention incorporates emotion recognition technology into a lost child tracking system, aiming to quickly and effectively locate lost children within a facility. The system utilizes facial recognition and emotion recognition capabilities to provide a series of processes for processing information and taking appropriate action.

[0170] server

[0171] The server collects facial recognition data in real time from cameras placed throughout the facility and stores it in a management database. This process can utilize OpenCV, a widely used facial recognition algorithm, or general image analysis APIs, which are cloud-based services. When a missing child is reported, the server matches the child's photo provided by the parents or staff with the facial data in the database, and combines this with video data from surveillance devices to determine the child's current location.

[0172] Furthermore, the server analyzes voice and facial expression data from parents and staff using an emotion recognition engine, dynamically adjusting notification content according to each individual's emotional state. Based on these results, it generates a message to provide reassurance and notifies staff members of their mobile devices.

[0173] terminal

[0174] The terminal receives notifications from the server and provides instructions for facility staff to quickly locate the child. Furthermore, it displays advice messages and response procedures that reflect the emotion recognition results, supporting staff in dealing with the situation in the most appropriate way.

[0175] User

[0176] When a child goes missing, the user (parent) provides the child's latest photo and their own emotional state through the support center or a dedicated application. This allows for appropriate responses that take into account the user's psychological state.

[0177] As a concrete example, if a parent loses sight of their child in a theme park, the system will sense the parent's anxiety and provide reassuring guidance. It analyzes data from the parent's voice tone and facial expressions and uses this to include reassuring messages in notifications sent to staff.

[0178] Example of a prompt:

[0179] "Describe a system for finding lost children within a theme park, and provide specific examples of notification features that utilize emotion recognition to alleviate parental anxiety."

[0180] This system enables early detection of children and provides a sense of psychological security to facility users, thereby improving the overall safety and reliability of the facility.

[0181] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0182] Step 1:

[0183] The server collects facial recognition data in real time from cameras within the facility and stores it in a database. The input is raw video data acquired by surveillance cameras. The server processes this video data using a facial recognition algorithm, extracts facial features, and outputs the facial recognition data to the database. This data is retained for subsequent matching processes.

[0184] Step 2:

[0185] The user (parent) provides a photo of their lost child and information about their own emotional state through the facility's support desk or app. The input consists of still image data and emotional information in text / voice format provided by the parent. The server receives the photo data and compares it with face recognition data in a stored database. This process calculates the degree of face matching and outputs the data with the best match.

[0186] Step 3:

[0187] The server acquires real-time video data from the surveillance camera and identifies the child's location by comparing it with the matching results obtained in step 2. The inputs are the real-time video data and the matching results. The server analyzes the video data using a similar facial recognition algorithm and generates the child's location coordinates as output. The location information is represented using GPS data or location coordinates within the facility.

[0188] Step 4:

[0189] The server uses an emotion recognition engine to determine the emotional state based on voice and facial expression data collected from parents and staff. Input can be text, voice data, or image data. The server performs a series of data calculations to analyze this data and outputs an emotional state (e.g., reassured, anxious, tense). This emotional information, along with the child's location information, is incorporated into the notification message.

[0190] Step 5:

[0191] The terminal distributes notification messages received from the server to facility staff and provides navigation information to the child's current location. Input consists of location information and a coordinated notification message from the server. The terminal converts this information into a format understandable to staff and displays it on the screen. This allows staff to quickly and appropriately reach the scene.

[0192] Step 6:

[0193] The user (parent) can receive feedback from the server and check the lost child's status in real time. The input is a status report from the server, which is output as text or audio to the parent's mobile device. This information includes reassuring messages and specific instructions for reunion, and is designed to alleviate the parent's anxiety.

[0194] (Application Example 2)

[0195] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0196] There is a need to alleviate the psychological burden and anxiety felt by parents and caregivers when they lose their children, and to provide support for quickly and effectively finding lost children. However, conventional systems lack the flexibility to respond flexibly to emotional states, and thus have limitations in reducing psychological burden. Therefore, this invention aims to provide a more reassuring lost child response by analyzing the emotional state of parents and caregivers and dynamically adjusting the notification content.

[0197] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0198] In this invention, the server includes means for acquiring and storing facial recognition information, means for comparing provided image information with stored facial recognition information and performing matching, means for receiving current video information from a monitoring device and determining the location by comparing it with the matching result, means for analyzing the emotional state of the parent or administrator using an emotion recognition engine, and means for dynamically adjusting the notification content based on the analyzed emotional information. This makes it possible to alleviate the anxiety of the parent or administrator and to find the lost child safely and quickly.

[0199] "Facial recognition information" refers to data that extracts the facial features of an individual and stores them as digital data.

[0200] "Image information" refers to digital data of photographs and videos acquired through vision.

[0201] "Matching" is the process of comparing points of agreement between two sets of data to determine if they are identical or similar.

[0202] "Surveillance equipment" refers to devices used to record or monitor the surrounding environment, and which can acquire video and audio.

[0203] "Current video information" refers to real-time video data acquired continuously over time.

[0204] An "emotion recognition engine" is a system equipped with algorithms for analyzing and estimating an individual's emotional state from data such as voice and facial expressions.

[0205] "Dynamically adjusting notification content" means changing the content of messages and alerts based on the information acquired, according to the situation at hand and the individual's status.

[0206] To implement this invention, hardware and software for teaching face recognition and emotion recognition technologies are required. The server first acquires face recognition information from cameras installed at entrance gates and various facilities and stores it in a database. This makes it possible to retain the characteristics of the subjects as digital data. Image processing libraries such as OpenCV are used for face recognition.

[0207] If a user loses sight of their child, the parent uses a smartphone application to provide a photo of the child to the server. This information is compared with facial recognition data stored on the server to ensure a proper match. This is then compared with current video footage received from surveillance equipment to pinpoint the child's location. The server then uses this location information to send navigation information to nearby staff members' terminals.

[0208] Furthermore, to determine the user's emotional state, an emotion recognition engine analyzes voice and facial expression data. This process utilizes emotion recognition tools such as EmotionAnalyzer. The emotional information obtained through this analysis is used to dynamically adjust the content of messages sent to staff. For example, if a user is feeling anxious, the notification displayed on the staff terminal will include not only detailed location information but also a message of encouragement.

[0209] A concrete example of its implementation is a case where a parent loses sight of their child at a theme park. Based on the information provided by the parent, an emotion recognition engine analyzes the parent's anxiety, and the server takes appropriate action. The notification includes specific instructions such as, "Your child is currently in the playground on the north side. Please check immediately." By quickly locating the child, the system helps facilitate the reunion of parent and child.

[0210] By using generative AI models, the system can further enhance its ability to analyze parental emotions and provide real-time responses. An example of a prompt message would be, "A child is lost. Analyze the parent's emotions and provide staff with quick and appropriate information." This allows the system to accurately support parents and facility staff, ensuring a safe and secure environment throughout the entire facility.

[0211] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0212] Step 1:

[0213] The server acquires facial recognition information from cameras installed at entrance gates and within the facility. The input is real-time video from the cameras, which is converted into facial recognition data using the OpenCV library and stored in a database. The output is the stored facial recognition information.

[0214] Step 2:

[0215] The user provides photo data of their child to the server via a smartphone app. The input is the image data of the child provided by the user, which the server receives and compares with pre-stored facial recognition information. The output is the matching result.

[0216] Step 3:

[0217] The server uses the current video information received from the monitoring equipment and compares it with the matching results to determine the child's location. The input consists of the matching results and real-time video data, and by comparing these, the server outputs the child's position coordinates.

[0218] Step 4:

[0219] The server notifies nearby staff terminals of the identified location information. The input is location coordinates, and the output is navigation information displayed on the staff terminals. This information is visually represented on a map.

[0220] Step 5:

[0221] The server uses an emotion recognition engine to analyze the user's emotional state from voice and facial expression data. The input is voice and facial expression data, which is analyzed using the EmotionAnalyzer tool. The output is data indicating the user's emotional state.

[0222] Step 6:

[0223] The server dynamically adjusts the content of notifications sent to staff terminals based on the analyzed emotional information. The input is the user's emotional data, and the output is a customized notification message displayed on the terminal. Specifically, if a parent is in a highly anxious state, the notification will include a message to calm them down.

[0224] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0225] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0226] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0227] [Second Embodiment]

[0228] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0229] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0230] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0231] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0232] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0233] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0234] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0235] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0236] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0237] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0238] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0239] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0240] This invention is a system for quickly identifying lost children in large facilities and safely reuniting them. This system combines facial recognition technology with real-time video analysis to efficiently track lost children and support appropriate responses.

[0241] server

[0242] The server stores visitor facial recognition data acquired at each entrance gate in a database. In the event of a lost child, it receives a photo of the child provided by a parent or guardian and matches it against the existing database using a facial recognition algorithm. Next, it analyzes real-time video data received from surveillance devices to search for a match with the facial recognition data. Once the location of the target is identified, it generates location information and notifies nearby employees' communication devices.

[0243] terminal

[0244] The terminal is a device that receives notifications sent from the server. These notifications include the location information of a lost child. Based on this information, the terminal provides navigation instructions to help employees quickly locate the lost child. This process allows employees to quickly find and safely protect the lost child.

[0245] User

[0246] When a child is found missing, the user (parent) immediately provides a recent photo of their child through the local support desk or support app. This data enables rapid matching in the system's facial recognition process. In addition, parents can check the latest location information provided through the employee's device, giving them peace of mind.

[0247] As a concrete example, consider a case at a theme park. When a parent realizes their child is lost, they immediately register the child's photo at the support desk. The server performs facial recognition and detects a match from real-time video. The server locates the child's position, a notification is sent to an employee's terminal, and the employee quickly goes to the lost child, ensuring a safe reunion between parent and child.

[0248] Thus, the system of the present invention enables a rapid response in the event of a child getting lost, providing a safe and secure environment for both facility operators and users.

[0249] The following describes the processing flow.

[0250] Step 1:

[0251] When a user enters the facility, a facial recognition camera captures the visitor's face and generates facial recognition data. The server receives this data and stores it in a database.

[0252] Step 2:

[0253] The user (parent) reports the child getting lost to the support desk and provides a recent photo of the child's face. The server receives this photo data and starts the matching process using a facial recognition algorithm against existing data in the database.

[0254] Step 3:

[0255] The server receives video data from the monitoring device in real time and searches for matching faces through a facial recognition process. This involves algorithmic image analysis and comparison of facial data.

[0256] Step 4:

[0257] When the server identifies a matching target, it generates the child's current location data. This location information is then formatted as map data.

[0258] Step 5:

[0259] The server sends location information as a notification to the terminals of nearby employees. The notification includes location details and instructions on how to get there.

[0260] Step 6:

[0261] The device receives a notification, and the employee quickly moves to the child's location based on the instructions. The device provides navigation information to support efficient movement.

[0262] Step 7:

[0263] An employee finds the lost child, safely takes them into custody, and then guides them back to their parents. This successful event is then fed back from the terminal to the server, completing the lost child tracking process.

[0264] (Example 1)

[0265] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0266] For facility managers, quickly identifying lost children within the facility and safely reuniting them with their guardians is a critical challenge. Current systems lack sufficient facial recognition and location accuracy, making rapid response difficult in some cases. Furthermore, there is a need for more efficient information notification and navigation.

[0267] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0268] In this invention, the server includes means for acquiring and storing visitor facial data, means for comparing provided image data with stored facial data to confirm a match, and means for receiving real-time video data from a monitoring device and determining the location by comparing it with the matching result. This makes it possible to efficiently identify lost children and quickly transmit their location information.

[0269] "Visitor facial data" refers to a collection of identifying information, including facial images and characteristics of individuals entering the facility.

[0270] "Image data" refers to visual information that is stored electronically and can be processed by a system.

[0271] "Verification" is the process of comparing the provided data with the stored data to confirm whether they match.

[0272] "Surveillance equipment" refers to devices installed inside and outside a facility that capture and transmit video in real time.

[0273] "Real-time video data" refers to video information captured by cameras and sensing devices at a given moment.

[0274] A "match result" is the result of the matching process confirming that the provided data is identical or similar to existing data.

[0275] "Means of location determination" refers to a mechanism within a system for determining the precise location of a person or object.

[0276] "Location information" refers to detailed data about a specific location, and may include maps and coordinate information.

[0277] A "communication device" is an electronic device used to send and receive information, and mainly includes mobile terminals and computers.

[0278] "Navigation instructions" are guidelines that show the route and steps necessary to reach a specific location.

[0279] "Re-authentication to confirm safe reunion" is a facial recognition process performed when a parent and child meet face-to-face again, in order to prevent misidentification.

[0280] The present invention will now be described in terms of embodiments for carrying it out. This system is designed to quickly identify lost children in facilities and safely reunite them with their guardians. The process primarily utilizes facial recognition technology, and real-time location tracking and notification are performed based on that information.

[0281] Hardware and software to be used

[0282] The server acquires the face data of visitors via cameras installed at each entrance gate within the facility. For this purpose, OpenCV or TensorFlow is used as image processing software. These software extract the acquired face data as feature quantities and save them in the database. Also, the server receives video data in real time from monitoring devices and performs comparison and matching by analyzing it with a face recognition algorithm. In this process, the computing power of the system is high and advanced image processing technologies are applied.

[0283] The terminal is a device that receives notifications sent from the server. Specifically, smartphones and tablets fall under this category and are operated through dedicated applications compatible with Android OS and iOS. The terminal visually displays location information using, for example, the Google Maps API and provides navigation instructions to the staff.

[0284] When the user (parent) realizes that their child is lost, they provide the latest image data of the child through the support desk or support app. This data serves as an important criterion in the face authentication process on the server. Furthermore, based on the location information received through the terminal, the user can cooperate with confidence in the search for the child.

[0285] Specific example

[0286] As a specific example, consider its use in a theme park. When a parent loses their child, they immediately provide a photo of the child at the support desk. This photo is compared with the server's database, and the current location of the child is identified through real-time video analysis of the surveillance cameras. Based on this information, the location information is sent to the employee's terminal, enabling a quick reunion of the parent and child.

[0287] Examples of prompt sentences

[0288] An example of a prompt to input into a generative AI model is, "Please describe the process of rapidly locating and reuniting lost children using a facial recognition system in a large facility." This prompt is expected to prompt the generative AI model to generate a detailed description of the system's processes and functions.

[0289] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0290] Step 1:

[0291] The server uses cameras installed at each entrance gate of the facility to acquire visitor facial data. The input is real-time video data. The server extracts facial information from this video using a face detection algorithm and stores the features in a database. The output is the facial feature data stored in the database. This process allows for the temporary storage of visitor facial data, preparing it for later matching.

[0292] Step 2:

[0293] When a user (parent) realizes their child is missing, they provide the latest photo of the child to the server via the support desk or a dedicated app. The input to this process is the child's photo data, which the server receives and immediately analyzes using a facial recognition algorithm. The output is a matching result, which confirms that the image matches existing data in the database.

[0294] Step 3:

[0295] The server receives video data in real time from the surveillance camera. The input is a video stream from the surveillance equipment. The server analyzes this video data and compares it with the matching results obtained in step 2 to perform face recognition. This process confirms the match between the face information in the video and the existing data, and identifies the location. The output is the identified location information.

[0296] Step 4:

[0297] The server formats the identified location information in JSON format or another suitable format and sends a notification to the terminals of nearby staff members. The input is the identified location data. The notification includes location details, allowing employees to quickly locate the lost child. The output is navigation instructions displayed on the terminal.

[0298] Step 5:

[0299] The terminal receives notifications from the server and provides visual navigation to employees through its built-in application. The input is location information from the server, and the location is displayed on a map using the Google Maps API. This allows employees to efficiently navigate within the facility. The output consists of a map and a guidance screen.

[0300] Step 6:

[0301] The user (parent) is guided by an employee to a designated location where they are reunited with their child. After the reunion, a re-authentication process using a terminal is performed, and facial recognition is used to confirm safety. The input consists of facial data of both the child and the parent. The server re-authenticates this data and confirms that the parent and child match, preventing misidentification and providing a secure environment. The output is the authentication confirmation result.

[0302] (Application Example 1)

[0303] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0304] In large facilities, there is a need to quickly and reliably locate lost children and safely reunite them with their guardians. However, conventional systems take a long time to pinpoint the location of lost children, making it difficult to reassure guardians. This invention aims to solve this problem and provide peace of mind and safety to facility users.

[0305] The specific processing by the specific processing unit 290 of the data processing apparatus 12 in Application Example 1 is realized by the following means.

[0306] In this invention, the server includes means for acquiring and storing face recognition data, means for comparing the provided visual data with the stored face recognition data to perform a coincidence determination, means for receiving real-time video information from a monitoring device and specifying a position by comparing with the coincidence determination result, means for notifying the specified position information to a mobile terminal of a nearby person in charge, and means for providing a movement route to the specified position. Thereby, it becomes possible to quickly specify the position of a lost person and for the person in charge to quickly respond.

[0307] "Face recognition data" is information that records an individual's face in digital form and is used for identification.

[0308] "Visual data" is what records visual information such as images and photos as digital data.

[0309] "Coincidence determination" is a process of comparing the provided visual data with the stored face recognition data to determine whether they are of the same person.

[0310] "Monitoring device" is a general term for devices that can shoot video in real time and process it as data.

[0311] "Real-time video information" is data that can immediately capture and process the current situation as video.

[0312] "Position information" is data indicating the current geographical position of a specific target.

[0313] "Mobile terminal" refers to a device that is portable and can send and receive information and display it.

[0314] "Movement route" refers to information indicating an optimal or specified route for a specific position.

[0315] To implement this invention, it is necessary to build a system that can quickly and reliably find lost children in large facilities. This system works effectively by combining facial recognition and real-time video analysis technologies.

[0316] The server acquires and stores facial recognition data in a database when visitors enter. Existing facial recognition software is used for the facial recognition technology. If a child gets lost, the parent provides the latest photo via a device, and the server uses this visual data to determine if it matches. The facial recognition algorithm compares the stored facial recognition data with the provided visual data to identify the match.

[0317] Meanwhile, monitoring devices are installed in multiple locations within the facility and constantly transmit real-time video information to a server. The server analyzes this video information and identifies videos containing matching facial recognition data. The identified location information is then notified to a mobile terminal. This allows staff to efficiently understand the route to the identified location and quickly guide the lost person to their location.

[0318] As a concrete example, consider a case where a child gets lost in a shopping mall. The parent sends a photo of their face via a smartphone app. The server performs facial recognition, determines the child's location from the video data in real time, and sends the location information, such as "The child is near the food court on the 3rd floor," to the staff member's smartphone. This allows the staff member to quickly go to the scene and guide the lost child back to their parents.

[0319] A concrete example of a prompt for a generated AI model would be: "Explain the steps and techniques necessary to safely reunite a lost child with their parents in a shopping mall. Include specific facial recognition algorithms, video analysis methods, and location information notification processes." Based on this prompt, it is possible to explain the system's detailed operation flow and technical background.

[0320] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0321] Step 1:

[0322] The server stores visitor facial recognition data acquired at the entrance gate into a database. The input is video from surveillance cameras, and the output is digital facial data obtained using a facial recognition algorithm. Faces are detected from the video, feature points are extracted, and stored.

[0323] Step 2:

[0324] When a child goes missing, the user sends the latest visual data of the child from their device to the server. The input is a photograph of the child, and the output is the visual data of that photograph. The data is securely transferred from the device to the server.

[0325] Step 3:

[0326] The server determines whether the received visual data matches the existing face recognition data in the database. The input is the provided visual data and the existing face recognition data, and the output is the matching result. The server then uses a face recognition algorithm to compare each set of data.

[0327] Step 4:

[0328] Real-time video information is transmitted from the monitoring device to the server. The input is live video from the monitoring device, and the output is real-time video information. The server constantly receives and analyzes this real-time video data.

[0329] Step 5:

[0330] The server compares the matching results with real-time video information to pinpoint the lost child's location. The input is the matching results and video information, and the output is the identified location. The entire face visible in the video is analyzed, and the location of the matched face is displayed on a map.

[0331] Step 6:

[0332] The server notifies the assigned person's mobile device of the identified location. The input is location information, and the output is a notification message sent to the mobile device. A message containing navigation information to that location is sent to the assigned person.

[0333] Step 7:

[0334] The terminal guides the assigned person along a route based on the location information it receives. The input is a notification message from the server, and the output is a route suggestion for the assigned person. The route displayed will help the assigned person reach the designated location quickly.

[0335] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0336] This invention is a system that further reduces the psychological burden on parents and facility staff and enables a quick and effective response by incorporating an emotion recognition engine into a lost child tracking system. This system is based on real-time video analysis using facial recognition technology and has the function of dynamically adjusting notification content through emotion recognition.

[0337] server

[0338] The server stores and manages facial recognition data collected at the facility's entrance gates. In the event of a lost child, the server compares the provided child's photo with information in the database using a facial recognition algorithm. Furthermore, the server analyzes real-time video transmitted by monitoring devices and identifies the child's location through facial recognition.

[0339] The emotion recognition engine based on this invention determines the current emotional state based on voice and facial expression data received from parents or facility staff. The server analyzes this emotional information and adjusts the content and format of notifications. For example, if a parent is showing strong anxiety, more detailed location information and reassuring messages are added.

[0340] terminal

[0341] The device receives notifications from the server and provides navigation information to help facility staff quickly locate the lost child. Furthermore, based on the results of the emotion recognition engine, it can present encouraging messages and specific action plans to staff, enabling them to respond calmly.

[0342] User

[0343] When a child goes missing, the user (parent) provides a photo of their child through the support desk or support app and appropriately communicates their emotional state. This allows the system to respond in a way that takes the parent's psychological burden into consideration.

[0344] As a concrete example, if a parent loses sight of their child at a theme park, the emotion recognition engine analyzes the parent's level of anxiety from their tone of voice and facial expression. The server takes this into account and includes additional information in the notification sent to the facility staff's terminal, providing reassurance. This aims to facilitate a smooth reunion between parent and child and create a safe and stress-free environment throughout the facility.

[0345] Thus, the system of the present invention, by combining facial recognition and emotion recognition, realizes a psychologically reassuring approach to lost children, bringing benefits to both facility management and users.

[0346] The following describes the processing flow.

[0347] Step 1:

[0348] When a user enters the facility, a facial recognition camera automatically captures a photograph of the visitor's face. The server receives this facial recognition data and securely stores it in a database.

[0349] Step 2:

[0350] If a user (parent) finds their lost child, they provide a recent photo of the child's face via the support desk or mobile app. The server receives this photo and begins matching it with existing facial recognition data in its database.

[0351] Step 3:

[0352] The server receives video data in real time from the monitoring device and uses a facial recognition algorithm to detect matches between photos and videos. If a match is detected, the server identifies the child's current location and compiles that information.

[0353] Step 4:

[0354] The emotion recognition engine analyzes the user's emotions through voice input and a real-time video interface. Based on these results, the server understands the user's state and adjusts the response accordingly.

[0355] Step 5:

[0356] The server generates an appropriate response message along with the child's precise location information and sends a notification to the terminal of nearby facility staff. The notification includes the parent's emotional state and provides instructions and messages appropriate to the situation.

[0357] Step 6:

[0358] The terminal (staff member's device) receives the notification and acts quickly according to the displayed details and navigation. The terminal also provides messages to staff members to encourage them to act calmly.

[0359] Step 7:

[0360] The user (parent) is reassured after receiving a direct update from facility staff and confirming that their child is safe. The emotion recognition engine confirms the parent's state of reassurance and feeds this information back to the server. The server records the completion of the process within the system.

[0361] (Example 2)

[0362] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0363] When a child gets lost in a public facility, it places a significant psychological burden not only on the parents but also on the facility staff. Conventional lost-child tracking systems often struggle to provide a quick response or alleviate parents' anxiety, which can result in a decrease in confidence in the safety of the facility. This invention aims to improve this situation and enable the rapid discovery of lost children while reducing the psychological burden.

[0364] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0365] In this invention, the server includes means for acquiring and storing facial recognition data, means for comparing and matching provided image data with stored facial recognition data, means for receiving real-time image data from a monitoring device and determining the location by comparing it with the matching result, means for analyzing emotional data from parents and staff and adjusting the notification content, means for transmitting the adjusted notification content to the communication device of a nearby staff member, and means for guiding staff members based on geographical information. This makes it possible to quickly locate children and provide psychological support to parents and staff.

[0366] "Facial recognition data" refers to data in which the characteristics of an individual's face are extracted as digital information and stored in an identifiable format.

[0367] "Image data" refers to still or dynamic visual information stored in digital format.

[0368] "Matching" is the process of comparing the provided data with the stored data and evaluating the degree of similarity.

[0369] A "monitoring device" is a device that monitors a designated environment in real time and acquires necessary image and audio data.

[0370] "Real-time image data" refers to digital data that instantly captures and transmits video footage occurring under certain circumstances.

[0371] "Emotional data" refers to information about emotional states extracted from voice, facial expressions, and actions.

[0372] "Notification content" refers to a message structured to convey specific information to a recipient.

[0373] A "communication device" is a device that combines hardware and software for sending and receiving information.

[0374] "Geographic information" refers to data related to location or place, and is usually presented visually on a map.

[0375] This invention incorporates emotion recognition technology into a lost child tracking system, aiming to quickly and effectively locate lost children within a facility. The system utilizes facial recognition and emotion recognition capabilities to provide a series of processes for processing information and taking appropriate action.

[0376] server

[0377] The server collects facial recognition data in real time from cameras placed throughout the facility and stores it in a management database. This process can utilize OpenCV, a widely used facial recognition algorithm, or general image analysis APIs, which are cloud-based services. When a missing child is reported, the server matches the child's photo provided by the parents or staff with the facial data in the database, and combines this with video data from surveillance devices to determine the child's current location.

[0378] Furthermore, the server analyzes voice and facial expression data from parents and staff using an emotion recognition engine, dynamically adjusting notification content according to each individual's emotional state. Based on these results, it generates a message to provide reassurance and notifies staff members of their mobile devices.

[0379] terminal

[0380] The terminal receives notifications from the server and provides instructions for facility staff to quickly locate the child. Furthermore, it displays advice messages and response procedures that reflect the emotion recognition results, supporting staff in dealing with the situation in the most appropriate way.

[0381] User

[0382] When a child goes missing, the user (parent) provides the child's latest photo and their own emotional state through the support center or a dedicated application. This allows for appropriate responses that take into account the user's psychological state.

[0383] As a concrete example, if a parent loses sight of their child in a theme park, the system will sense the parent's anxiety and provide reassuring guidance. It analyzes data from the parent's voice tone and facial expressions and uses this to include reassuring messages in notifications sent to staff.

[0384] Example of a prompt:

[0385] "Describe a system for finding lost children within a theme park, and provide specific examples of notification features that utilize emotion recognition to alleviate parental anxiety."

[0386] This system enables early detection of children and provides a sense of psychological security to facility users, thereby improving the overall safety and reliability of the facility.

[0387] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0388] Step 1:

[0389] The server collects facial recognition data in real time from cameras within the facility and stores it in a database. The input is raw video data acquired by surveillance cameras. The server processes this video data using a facial recognition algorithm, extracts facial features, and outputs the facial recognition data to the database. This data is retained for subsequent matching processes.

[0390] Step 2:

[0391] The user (parent) provides a photo of their lost child and information about their own emotional state through the facility's support desk or app. The input consists of still image data and emotional information in text / voice format provided by the parent. The server receives the photo data and compares it with face recognition data in a stored database. This process calculates the degree of face matching and outputs the data with the best match.

[0392] Step 3:

[0393] The server acquires real-time video data from the surveillance camera and identifies the child's location by comparing it with the matching results obtained in step 2. The inputs are the real-time video data and the matching results. The server analyzes the video data using a similar facial recognition algorithm and generates the child's location coordinates as output. The location information is represented using GPS data or location coordinates within the facility.

[0394] Step 4:

[0395] The server uses an emotion recognition engine to determine the emotional state based on voice and facial expression data collected from parents and staff. Input can be text, voice data, or image data. The server performs a series of data calculations to analyze this data and outputs an emotional state (e.g., reassured, anxious, tense). This emotional information, along with the child's location information, is incorporated into the notification message.

[0396] Step 5:

[0397] The terminal distributes notification messages received from the server to facility staff and provides navigation information to the child's current location. Input consists of location information and a coordinated notification message from the server. The terminal converts this information into a format understandable to staff and displays it on the screen. This allows staff to quickly and appropriately reach the scene.

[0398] Step 6:

[0399] The user (parent) can receive feedback from the server and check the lost child's status in real time. The input is a status report from the server, which is output as text or audio to the parent's mobile device. This information includes reassuring messages and specific instructions for reunion, and is designed to alleviate the parent's anxiety.

[0400] (Application Example 2)

[0401] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0402] There is a need to alleviate the psychological burden and anxiety felt by parents and caregivers when they lose their children, and to provide support for quickly and effectively finding lost children. However, conventional systems lack the flexibility to respond flexibly to emotional states, and thus have limitations in reducing psychological burden. Therefore, this invention aims to provide a more reassuring lost child response by analyzing the emotional state of parents and caregivers and dynamically adjusting the notification content.

[0403] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0404] In this invention, the server includes means for acquiring and storing facial recognition information, means for comparing provided image information with stored facial recognition information and performing matching, means for receiving current video information from a monitoring device and determining the location by comparing it with the matching result, means for analyzing the emotional state of the parent or administrator using an emotion recognition engine, and means for dynamically adjusting the notification content based on the analyzed emotional information. This makes it possible to alleviate the anxiety of the parent or administrator and to find the lost child safely and quickly.

[0405] "Facial recognition information" refers to data that extracts the facial features of an individual and stores them as digital data.

[0406] "Image information" refers to digital data of photographs and videos acquired through vision.

[0407] "Matching" is the process of comparing points of agreement between two sets of data to determine if they are identical or similar.

[0408] "Surveillance equipment" refers to devices used to record or monitor the surrounding environment, and which can acquire video and audio.

[0409] "Current video information" refers to real-time video data acquired continuously over time.

[0410] An "emotion recognition engine" is a system equipped with algorithms for analyzing and estimating an individual's emotional state from data such as voice and facial expressions.

[0411] "Dynamically adjusting notification content" means changing the content of messages and alerts based on the information acquired, according to the situation at hand and the individual's status.

[0412] To implement this invention, hardware and software for teaching face recognition and emotion recognition technologies are required. The server first acquires face recognition information from cameras installed at entrance gates and various facilities and stores it in a database. This makes it possible to retain the characteristics of the subjects as digital data. Image processing libraries such as OpenCV are used for face recognition.

[0413] If a user loses sight of their child, the parent uses a smartphone application to provide a photo of the child to the server. This information is compared with facial recognition data stored on the server to ensure a proper match. This is then compared with current video footage received from surveillance equipment to pinpoint the child's location. The server then uses this location information to send navigation information to nearby staff members' terminals.

[0414] Furthermore, to determine the user's emotional state, an emotion recognition engine analyzes voice and facial expression data. This process utilizes emotion recognition tools such as EmotionAnalyzer. The emotional information obtained through this analysis is used to dynamically adjust the content of messages sent to staff. For example, if a user is feeling anxious, the notification displayed on the staff terminal will include not only detailed location information but also a message of encouragement.

[0415] A concrete example of its implementation is a case where a parent loses sight of their child at a theme park. Based on the information provided by the parent, an emotion recognition engine analyzes the parent's anxiety, and the server takes appropriate action. The notification includes specific instructions such as, "Your child is currently in the playground on the north side. Please check immediately." By quickly locating the child, the system helps facilitate the reunion of parent and child.

[0416] By using generative AI models, the system can further enhance its ability to analyze parental emotions and provide real-time responses. An example of a prompt message would be, "A child is lost. Analyze the parent's emotions and provide staff with quick and appropriate information." This allows the system to accurately support parents and facility staff, ensuring a safe and secure environment throughout the entire facility.

[0417] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0418] Step 1:

[0419] The server acquires facial recognition information from cameras installed at entrance gates and within the facility. The input is real-time video from the cameras, which is converted into facial recognition data using the OpenCV library and stored in a database. The output is the stored facial recognition information.

[0420] Step 2:

[0421] The user provides photo data of their child to the server via a smartphone app. The input is the image data of the child provided by the user, which the server receives and compares with pre-stored facial recognition information. The output is the matching result.

[0422] Step 3:

[0423] The server uses the current video information received from the monitoring equipment and compares it with the matching results to determine the child's location. The input consists of the matching results and real-time video data, and by comparing these, the server outputs the child's position coordinates.

[0424] Step 4:

[0425] The server notifies nearby staff terminals of the identified location information. The input is location coordinates, and the output is navigation information displayed on the staff terminals. This information is visually represented on a map.

[0426] Step 5:

[0427] The server uses an emotion recognition engine to analyze the user's emotional state from voice and facial expression data. The input is voice and facial expression data, which is analyzed using the EmotionAnalyzer tool. The output is data indicating the user's emotional state.

[0428] Step 6:

[0429] The server dynamically adjusts the content of notifications sent to staff terminals based on the analyzed emotional information. The input is the user's emotional data, and the output is a customized notification message displayed on the terminal. Specifically, if a parent is in a highly anxious state, the notification will include a message to calm them down.

[0430] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0431] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0432] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0433] [Third Embodiment]

[0434] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0435] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0436] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0437] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0438] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0439] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0440] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0441] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0442] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0443] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0444] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0445] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0446] This invention is a system for quickly identifying lost children in large facilities and safely reuniting them. This system combines facial recognition technology with real-time video analysis to efficiently track lost children and support appropriate responses.

[0447] server

[0448] The server stores visitor facial recognition data acquired at each entrance gate in a database. In the event of a lost child, it receives a photo of the child provided by a parent or guardian and matches it against the existing database using a facial recognition algorithm. Next, it analyzes real-time video data received from surveillance devices to search for a match with the facial recognition data. Once the location of the target is identified, it generates location information and notifies nearby employees' communication devices.

[0449] terminal

[0450] The terminal is a device that receives notifications sent from the server. These notifications include the location information of a lost child. Based on this information, the terminal provides navigation instructions to help employees quickly locate the lost child. This process allows employees to quickly find and safely protect the lost child.

[0451] User

[0452] When a child is found missing, the user (parent) immediately provides a recent photo of their child through the local support desk or support app. This data enables rapid matching in the system's facial recognition process. In addition, parents can check the latest location information provided through the employee's device, giving them peace of mind.

[0453] As a concrete example, consider a case at a theme park. When a parent realizes their child is lost, they immediately register the child's photo at the support desk. The server performs facial recognition and detects a match from real-time video. The server locates the child's position, a notification is sent to an employee's terminal, and the employee quickly goes to the lost child, ensuring a safe reunion between parent and child.

[0454] Thus, the system of the present invention enables a rapid response in the event of a child getting lost, providing a safe and secure environment for both facility operators and users.

[0455] The following describes the processing flow.

[0456] Step 1:

[0457] When a user enters the facility, a facial recognition camera captures the visitor's face and generates facial recognition data. The server receives this data and stores it in a database.

[0458] Step 2:

[0459] The user (parent) reports the child getting lost to the support desk and provides a recent photo of the child's face. The server receives this photo data and starts the matching process using a facial recognition algorithm against existing data in the database.

[0460] Step 3:

[0461] The server receives video data from the monitoring device in real time and searches for matching faces through a facial recognition process. This involves algorithmic image analysis and comparison of facial data.

[0462] Step 4:

[0463] When the server identifies a matching target, it generates the child's current location data. This location information is then formatted as map data.

[0464] Step 5:

[0465] The server sends location information as a notification to the terminals of nearby employees. The notification includes location details and instructions on how to get there.

[0466] Step 6:

[0467] The device receives a notification, and the employee quickly moves to the child's location based on the instructions. The device provides navigation information to support efficient movement.

[0468] Step 7:

[0469] An employee finds the lost child, safely takes them into custody, and then guides them back to their parents. This successful event is then fed back from the terminal to the server, completing the lost child tracking process.

[0470] (Example 1)

[0471] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0472] For facility managers, quickly identifying lost children within the facility and safely reuniting them with their guardians is a critical challenge. Current systems lack sufficient facial recognition and location accuracy, making rapid response difficult in some cases. Furthermore, there is a need for more efficient information notification and navigation.

[0473] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0474] In this invention, the server includes means for acquiring and storing visitor facial data, means for comparing provided image data with stored facial data to confirm a match, and means for receiving real-time video data from a monitoring device and determining the location by comparing it with the matching result. This makes it possible to efficiently identify lost children and quickly transmit their location information.

[0475] "Visitor facial data" refers to a collection of identifying information, including facial images and characteristics of individuals entering the facility.

[0476] "Image data" refers to visual information that is stored electronically and can be processed by a system.

[0477] "Verification" is the process of comparing the provided data with the stored data to confirm whether they match.

[0478] "Surveillance equipment" refers to devices installed inside and outside a facility that capture and transmit video in real time.

[0479] "Real-time video data" refers to video information captured by cameras and sensing devices at a given moment.

[0480] A "match result" is the result of the matching process confirming that the provided data is identical or similar to existing data.

[0481] "Means of location determination" refers to a mechanism within a system for determining the precise location of a person or object.

[0482] "Location information" refers to detailed data about a specific location, and may include maps and coordinate information.

[0483] A "communication device" is an electronic device used to send and receive information, and mainly includes mobile terminals and computers.

[0484] "Navigation instructions" are guidelines that show the route and steps necessary to reach a specific location.

[0485] "Re-authentication to confirm safe reunion" is a facial recognition process performed when a parent and child meet face-to-face again, in order to prevent misidentification.

[0486] The present invention will now be described in terms of embodiments for carrying it out. This system is designed to quickly identify lost children in facilities and safely reunite them with their guardians. The process primarily utilizes facial recognition technology, and real-time location tracking and notification are performed based on that information.

[0487] Hardware and software to be used

[0488] The server acquires visitor facial data via cameras installed at each entrance gate within the facility. For this purpose, it uses image processing software such as OpenCV and TensorFlow. This software extracts features from the acquired facial data and stores them in a database. The server also receives video data in real time from surveillance equipment and performs comparison and matching by analyzing it with a facial recognition algorithm. This process requires high system computing power and applies advanced image processing techniques.

[0489] The terminal is a device that receives notifications sent from the server. Specifically, this includes smartphones and tablets, and is operated through a dedicated application compatible with Android OS and iOS. The terminal visually displays location information using the Google Maps API and other tools, and provides navigation instructions to staff.

[0490] When a user (parent) realizes their child is lost, they provide the latest image data of their child through the support desk or support app. This data is a crucial criterion for the facial recognition process on the server. Furthermore, users can confidently cooperate in the search for their child based on the location information received through their device.

[0491] Specific example

[0492] As a concrete example, consider its use in a theme park. If a parent gets separated from their child, a photo of the child is immediately provided at the support desk. This photo is cross-referenced with a server database, and the child's current location is determined through real-time video analysis from surveillance cameras. Based on this information, the location information is sent to employee terminals, enabling a quick reunion between parent and child.

[0493] Example of a prompt

[0494] An example of a prompt to input into a generative AI model is, "Please describe the process of rapidly locating and reuniting lost children using a facial recognition system in a large facility." This prompt is expected to prompt the generative AI model to generate a detailed description of the system's processes and functions.

[0495] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0496] Step 1:

[0497] The server uses cameras installed at each entrance gate of the facility to acquire visitor facial data. The input is real-time video data. The server extracts facial information from this video using a face detection algorithm and stores the features in a database. The output is the facial feature data stored in the database. This process allows for the temporary storage of visitor facial data, preparing it for later matching.

[0498] Step 2:

[0499] When a user (parent) realizes their child is missing, they provide the latest photo of the child to the server via the support desk or a dedicated app. The input to this process is the child's photo data, which the server receives and immediately analyzes using a facial recognition algorithm. The output is a matching result, which confirms that the image matches existing data in the database.

[0500] Step 3:

[0501] The server receives video data in real time from the surveillance camera. The input is a video stream from the surveillance equipment. The server analyzes this video data and compares it with the matching results obtained in step 2 to perform face recognition. This process confirms the match between the face information in the video and the existing data, and identifies the location. The output is the identified location information.

[0502] Step 4:

[0503] The server formats the identified location information in JSON format or another suitable format and sends a notification to the terminals of nearby staff members. The input is the identified location data. The notification includes location details, allowing employees to quickly locate the lost child. The output is navigation instructions displayed on the terminal.

[0504] Step 5:

[0505] The terminal receives notifications from the server and provides visual navigation to employees through its built-in application. The input is location information from the server, and the location is displayed on a map using the Google Maps API. This allows employees to efficiently navigate within the facility. The output consists of a map and a guidance screen.

[0506] Step 6:

[0507] The user (parent) is guided by an employee to a designated location where they are reunited with their child. After the reunion, a re-authentication process using a terminal is performed, and facial recognition is used to confirm safety. The input consists of facial data of both the child and the parent. The server re-authenticates this data and confirms that the parent and child match, preventing misidentification and providing a secure environment. The output is the authentication confirmation result.

[0508] (Application Example 1)

[0509] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0510] In large facilities, there is a need to quickly and reliably locate lost children and safely reunite them with their guardians. However, conventional systems take a long time to pinpoint the location of lost children, making it difficult to reassure guardians. This invention aims to solve this problem and provide peace of mind and safety to facility users.

[0511] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0512] In this invention, the server includes means for acquiring and storing facial recognition data, means for comparing provided visual data with stored facial recognition data to determine a match, means for receiving real-time video information from a monitoring device and determining its location by comparing it with the match determination result, means for notifying the determined location information to the mobile terminal of a nearby staff member, and means for providing a route to the determined location. This enables the rapid identification of a lost child's location and a quick response by staff members.

[0513] "Facial recognition data" refers to information recorded digitally from an individual's face and used for identification purposes.

[0514] "Visual data" refers to visual information, such as images and photographs, recorded as digital data.

[0515] "Matching" is the process of comparing the provided visual data with stored facial recognition data to determine whether they are the same person.

[0516] "Monitoring equipment" is a general term for devices that can capture video in real time and process it as data.

[0517] "Real-time video information" refers to data that can instantly capture and process the current situation as video.

[0518] "Location information" refers to data that indicates the current geographical location of a specific object.

[0519] A "mobile terminal" is a portable device capable of sending, receiving, and displaying information.

[0520] A "travel path" is information that indicates the optimal or designated path to a specific location.

[0521] To implement this invention, it is necessary to build a system that can quickly and reliably find lost children in large facilities. This system works effectively by combining facial recognition and real-time video analysis technologies.

[0522] The server acquires and stores facial recognition data in a database when visitors enter. Existing facial recognition software is used for the facial recognition technology. If a child gets lost, the parent provides the latest photo via a device, and the server uses this visual data to determine if it matches. The facial recognition algorithm compares the stored facial recognition data with the provided visual data to identify the match.

[0523] Meanwhile, monitoring devices are installed in multiple locations within the facility and constantly transmit real-time video information to a server. The server analyzes this video information and identifies videos containing matching facial recognition data. The identified location information is then notified to a mobile terminal. This allows staff to efficiently understand the route to the identified location and quickly guide the lost person to their location.

[0524] As a concrete example, consider a case where a child gets lost in a shopping mall. The parent sends a photo of their face via a smartphone app. The server performs facial recognition, determines the child's location from the video data in real time, and sends the location information, such as "The child is near the food court on the 3rd floor," to the staff member's smartphone. This allows the staff member to quickly go to the scene and guide the lost child back to their parents.

[0525] A concrete example of a prompt for a generated AI model would be: "Explain the steps and techniques necessary to safely reunite a lost child with their parents in a shopping mall. Include specific facial recognition algorithms, video analysis methods, and location information notification processes." Based on this prompt, it is possible to explain the system's detailed operation flow and technical background.

[0526] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0527] Step 1:

[0528] The server stores visitor facial recognition data acquired at the entrance gate into a database. The input is video from surveillance cameras, and the output is digital facial data obtained using a facial recognition algorithm. Faces are detected from the video, feature points are extracted, and stored.

[0529] Step 2:

[0530] When a child goes missing, the user sends the latest visual data of the child from their device to the server. The input is a photograph of the child, and the output is the visual data of that photograph. The data is securely transferred from the device to the server.

[0531] Step 3:

[0532] The server determines whether the received visual data matches the existing face recognition data in the database. The input is the provided visual data and the existing face recognition data, and the output is the matching result. The server then uses a face recognition algorithm to compare each set of data.

[0533] Step 4:

[0534] Real-time video information is transmitted from the monitoring device to the server. The input is live video from the monitoring device, and the output is real-time video information. The server constantly receives and analyzes this real-time video data.

[0535] Step 5:

[0536] The server compares the matching results with real-time video information to pinpoint the lost child's location. The input is the matching results and video information, and the output is the identified location. The entire face visible in the video is analyzed, and the location of the matched face is displayed on a map.

[0537] Step 6:

[0538] The server notifies the assigned person's mobile device of the identified location. The input is location information, and the output is a notification message sent to the mobile device. A message containing navigation information to that location is sent to the assigned person.

[0539] Step 7:

[0540] The terminal guides the assigned person along a route based on the location information it receives. The input is a notification message from the server, and the output is a route suggestion for the assigned person. The route displayed will help the assigned person reach the designated location quickly.

[0541] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0542] This invention is a system that further reduces the psychological burden on parents and facility staff and enables a quick and effective response by incorporating an emotion recognition engine into a lost child tracking system. This system is based on real-time video analysis using facial recognition technology and has the function of dynamically adjusting notification content through emotion recognition.

[0543] server

[0544] The server stores and manages facial recognition data collected at the facility's entrance gates. In the event of a lost child, the server compares the provided child's photo with information in the database using a facial recognition algorithm. Furthermore, the server analyzes real-time video transmitted by monitoring devices and identifies the child's location through facial recognition.

[0545] The emotion recognition engine based on this invention determines the current emotional state based on voice and facial expression data received from parents or facility staff. The server analyzes this emotional information and adjusts the content and format of notifications. For example, if a parent is showing strong anxiety, more detailed location information and reassuring messages are added.

[0546] terminal

[0547] The device receives notifications from the server and provides navigation information to help facility staff quickly locate the lost child. Furthermore, based on the results of the emotion recognition engine, it can present encouraging messages and specific action plans to staff, enabling them to respond calmly.

[0548] User

[0549] When a child goes missing, the user (parent) provides a photo of their child through the support desk or support app and appropriately communicates their emotional state. This allows the system to respond in a way that takes the parent's psychological burden into consideration.

[0550] As a concrete example, if a parent loses sight of their child at a theme park, the emotion recognition engine analyzes the parent's level of anxiety from their tone of voice and facial expression. The server takes this into account and includes additional information in the notification sent to the facility staff's terminal, providing reassurance. This aims to facilitate a smooth reunion between parent and child and create a safe and stress-free environment throughout the facility.

[0551] Thus, the system of the present invention, by combining facial recognition and emotion recognition, realizes a psychologically reassuring approach to lost children, bringing benefits to both facility management and users.

[0552] The following describes the processing flow.

[0553] Step 1:

[0554] When a user enters the facility, a facial recognition camera automatically captures a photograph of the visitor's face. The server receives this facial recognition data and securely stores it in a database.

[0555] Step 2:

[0556] If a user (parent) finds their lost child, they provide a recent photo of the child's face via the support desk or mobile app. The server receives this photo and begins matching it with existing facial recognition data in its database.

[0557] Step 3:

[0558] The server receives video data in real time from the monitoring device and uses a facial recognition algorithm to detect matches between photos and videos. If a match is detected, the server identifies the child's current location and compiles that information.

[0559] Step 4:

[0560] The emotion recognition engine analyzes the user's emotions through voice input and a real-time video interface. Based on these results, the server understands the user's state and adjusts the response accordingly.

[0561] Step 5:

[0562] The server generates an appropriate response message along with the child's precise location information and sends a notification to the terminal of nearby facility staff. The notification includes the parent's emotional state and provides instructions and messages appropriate to the situation.

[0563] Step 6:

[0564] The terminal (staff member's device) receives the notification and acts quickly according to the displayed details and navigation. The terminal also provides messages to staff members to encourage them to act calmly.

[0565] Step 7:

[0566] The user (parent) is reassured after receiving a direct update from facility staff and confirming that their child is safe. The emotion recognition engine confirms the parent's state of reassurance and feeds this information back to the server. The server records the completion of the process within the system.

[0567] (Example 2)

[0568] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0569] When a child gets lost in a public facility, it places a significant psychological burden not only on the parents but also on the facility staff. Conventional lost-child tracking systems often struggle to provide a quick response or alleviate parents' anxiety, which can result in a decrease in confidence in the safety of the facility. This invention aims to improve this situation and enable the rapid discovery of lost children while reducing the psychological burden.

[0570] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0571] In this invention, the server includes means for acquiring and storing facial recognition data, means for comparing and matching provided image data with stored facial recognition data, means for receiving real-time image data from a monitoring device and determining the location by comparing it with the matching result, means for analyzing emotional data from parents and staff and adjusting the notification content, means for transmitting the adjusted notification content to the communication device of a nearby staff member, and means for guiding staff members based on geographical information. This makes it possible to quickly locate children and provide psychological support to parents and staff.

[0572] "Facial recognition data" refers to data in which the characteristics of an individual's face are extracted as digital information and stored in an identifiable format.

[0573] "Image data" refers to still or dynamic visual information stored in digital format.

[0574] "Matching" is the process of comparing the provided data with the stored data and evaluating the degree of similarity.

[0575] A "monitoring device" is a device that monitors a designated environment in real time and acquires necessary image and audio data.

[0576] "Real-time image data" refers to digital data that instantly captures and transmits video footage occurring under certain circumstances.

[0577] "Emotional data" refers to information about emotional states extracted from voice, facial expressions, and actions.

[0578] "Notification content" refers to a message structured to convey specific information to a recipient.

[0579] A "communication device" is a device that combines hardware and software for sending and receiving information.

[0580] "Geographic information" refers to data related to location or place, and is usually presented visually on a map.

[0581] This invention incorporates emotion recognition technology into a lost child tracking system, aiming to quickly and effectively locate lost children within a facility. The system utilizes facial recognition and emotion recognition capabilities to provide a series of processes for processing information and taking appropriate action.

[0582] server

[0583] The server collects facial recognition data in real time from cameras placed throughout the facility and stores it in a management database. This process can utilize OpenCV, a widely used facial recognition algorithm, or general image analysis APIs, which are cloud-based services. When a missing child is reported, the server matches the child's photo provided by the parents or staff with the facial data in the database, and combines this with video data from surveillance devices to determine the child's current location.

[0584] Furthermore, the server analyzes voice and facial expression data from parents and staff using an emotion recognition engine, dynamically adjusting notification content according to each individual's emotional state. Based on these results, it generates a message to provide reassurance and notifies staff members of their mobile devices.

[0585] terminal

[0586] The terminal receives notifications from the server and provides instructions for facility staff to quickly locate the child. Furthermore, it displays advice messages and response procedures that reflect the emotion recognition results, supporting staff in dealing with the situation in the most appropriate way.

[0587] User

[0588] When a child goes missing, the user (parent) provides the child's latest photo and their own emotional state through the support center or a dedicated application. This allows for appropriate responses that take into account the user's psychological state.

[0589] As a concrete example, if a parent loses sight of their child in a theme park, the system will sense the parent's anxiety and provide reassuring guidance. It analyzes data from the parent's voice tone and facial expressions and uses this to include reassuring messages in notifications sent to staff.

[0590] Example of a prompt:

[0591] "Describe a system for finding lost children within a theme park, and provide specific examples of notification features that utilize emotion recognition to alleviate parental anxiety."

[0592] This system enables early detection of children and provides a sense of psychological security to facility users, thereby improving the overall safety and reliability of the facility.

[0593] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0594] Step 1:

[0595] The server collects facial recognition data in real time from cameras within the facility and stores it in a database. The input is raw video data acquired by surveillance cameras. The server processes this video data using a facial recognition algorithm, extracts facial features, and outputs the facial recognition data to the database. This data is retained for subsequent matching processes.

[0596] Step 2:

[0597] The user (parent) provides a photo of their lost child and information about their own emotional state through the facility's support desk or app. The input consists of still image data and emotional information in text / voice format provided by the parent. The server receives the photo data and compares it with face recognition data in a stored database. This process calculates the degree of face matching and outputs the data with the best match.

[0598] Step 3:

[0599] The server acquires real-time video data from the surveillance camera and identifies the child's location by comparing it with the matching results obtained in step 2. The inputs are the real-time video data and the matching results. The server analyzes the video data using a similar facial recognition algorithm and generates the child's location coordinates as output. The location information is represented using GPS data or location coordinates within the facility.

[0600] Step 4:

[0601] The server uses an emotion recognition engine to determine the emotional state based on voice and facial expression data collected from parents and staff. Input can be text, voice data, or image data. The server performs a series of data calculations to analyze this data and outputs an emotional state (e.g., reassured, anxious, tense). This emotional information, along with the child's location information, is incorporated into the notification message.

[0602] Step 5:

[0603] The terminal distributes notification messages received from the server to facility staff and provides navigation information to the child's current location. Input consists of location information and a coordinated notification message from the server. The terminal converts this information into a format understandable to staff and displays it on the screen. This allows staff to quickly and appropriately reach the scene.

[0604] Step 6:

[0605] The user (parent) can receive feedback from the server and check the lost child's status in real time. The input is a status report from the server, which is output as text or audio to the parent's mobile device. This information includes reassuring messages and specific instructions for reunion, and is designed to alleviate the parent's anxiety.

[0606] (Application Example 2)

[0607] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0608] There is a need to alleviate the psychological burden and anxiety felt by parents and caregivers when they lose their children, and to provide support for quickly and effectively finding lost children. However, conventional systems lack the flexibility to respond flexibly to emotional states, and thus have limitations in reducing psychological burden. Therefore, this invention aims to provide a more reassuring lost child response by analyzing the emotional state of parents and caregivers and dynamically adjusting the notification content.

[0609] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0610] In this invention, the server includes means for acquiring and storing facial recognition information, means for comparing provided image information with stored facial recognition information and performing matching, means for receiving current video information from a monitoring device and determining the location by comparing it with the matching result, means for analyzing the emotional state of the parent or administrator using an emotion recognition engine, and means for dynamically adjusting the notification content based on the analyzed emotional information. This makes it possible to alleviate the anxiety of the parent or administrator and to find the lost child safely and quickly.

[0611] "Facial recognition information" refers to data that extracts the facial features of an individual and stores them as digital data.

[0612] "Image information" refers to digital data of photographs and videos acquired through vision.

[0613] "Matching" is the process of comparing points of agreement between two sets of data to determine if they are identical or similar.

[0614] "Surveillance equipment" refers to devices used to record or monitor the surrounding environment, and which can acquire video and audio.

[0615] "Current video information" refers to real-time video data acquired continuously over time.

[0616] An "emotion recognition engine" is a system equipped with algorithms for analyzing and estimating an individual's emotional state from data such as voice and facial expressions.

[0617] "Dynamically adjusting notification content" means changing the content of messages and alerts based on the information acquired, according to the situation at hand and the individual's status.

[0618] To implement this invention, hardware and software for teaching face recognition and emotion recognition technologies are required. The server first acquires face recognition information from cameras installed at entrance gates and various facilities and stores it in a database. This makes it possible to retain the characteristics of the subjects as digital data. Image processing libraries such as OpenCV are used for face recognition.

[0619] If a user loses sight of their child, the parent uses a smartphone application to provide a photo of the child to the server. This information is compared with facial recognition data stored on the server to ensure a proper match. This is then compared with current video footage received from surveillance equipment to pinpoint the child's location. The server then uses this location information to send navigation information to nearby staff members' terminals.

[0620] Furthermore, to determine the user's emotional state, an emotion recognition engine analyzes voice and facial expression data. This process utilizes emotion recognition tools such as EmotionAnalyzer. The emotional information obtained through this analysis is used to dynamically adjust the content of messages sent to staff. For example, if a user is feeling anxious, the notification displayed on the staff terminal will include not only detailed location information but also a message of encouragement.

[0621] A concrete example of its implementation is a case where a parent loses sight of their child at a theme park. Based on the information provided by the parent, an emotion recognition engine analyzes the parent's anxiety, and the server takes appropriate action. The notification includes specific instructions such as, "Your child is currently in the playground on the north side. Please check immediately." By quickly locating the child, the system helps facilitate the reunion of parent and child.

[0622] By using generative AI models, the system can further enhance its ability to analyze parental emotions and provide real-time responses. An example of a prompt message would be, "A child is lost. Analyze the parent's emotions and provide staff with quick and appropriate information." This allows the system to accurately support parents and facility staff, ensuring a safe and secure environment throughout the entire facility.

[0623] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0624] Step 1:

[0625] The server acquires facial recognition information from cameras installed at entrance gates and within the facility. The input is real-time video from the cameras, which is converted into facial recognition data using the OpenCV library and stored in a database. The output is the stored facial recognition information.

[0626] Step 2:

[0627] The user provides photo data of their child to the server via a smartphone app. The input is the image data of the child provided by the user, which the server receives and compares with pre-stored facial recognition information. The output is the matching result.

[0628] Step 3:

[0629] The server uses the current video information received from the monitoring equipment and compares it with the matching results to determine the child's location. The input consists of the matching results and real-time video data, and by comparing these, the server outputs the child's position coordinates.

[0630] Step 4:

[0631] The server notifies nearby staff terminals of the identified location information. The input is location coordinates, and the output is navigation information displayed on the staff terminals. This information is visually represented on a map.

[0632] Step 5:

[0633] The server uses an emotion recognition engine to analyze the user's emotional state from voice and facial expression data. The input is voice and facial expression data, which is analyzed using the EmotionAnalyzer tool. The output is data indicating the user's emotional state.

[0634] Step 6:

[0635] The server dynamically adjusts the content of notifications sent to staff terminals based on the analyzed emotional information. The input is the user's emotional data, and the output is a customized notification message displayed on the terminal. Specifically, if a parent is in a highly anxious state, the notification will include a message to calm them down.

[0636] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0637] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0638] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0639] [Fourth Embodiment]

[0640] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0641] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0642] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0643] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0644] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0645] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0646] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0647] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0648] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0649] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0650] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0651] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0652] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0653] This invention is a system for quickly identifying lost children in large facilities and safely reuniting them. This system combines facial recognition technology with real-time video analysis to efficiently track lost children and support appropriate responses.

[0654] server

[0655] The server stores visitor facial recognition data acquired at each entrance gate in a database. In the event of a lost child, it receives a photo of the child provided by a parent or guardian and matches it against the existing database using a facial recognition algorithm. Next, it analyzes real-time video data received from surveillance devices to search for a match with the facial recognition data. Once the location of the target is identified, it generates location information and notifies nearby employees' communication devices.

[0656] terminal

[0657] The terminal is a device that receives notifications sent from the server. These notifications include the location information of a lost child. Based on this information, the terminal provides navigation instructions to help employees quickly locate the lost child. This process allows employees to quickly find and safely protect the lost child.

[0658] User

[0659] When a child is found missing, the user (parent) immediately provides a recent photo of their child through the local support desk or support app. This data enables rapid matching in the system's facial recognition process. In addition, parents can check the latest location information provided through the employee's device, giving them peace of mind.

[0660] As a concrete example, consider a case at a theme park. When a parent realizes their child is lost, they immediately register the child's photo at the support desk. The server performs facial recognition and detects a match from real-time video. The server locates the child's position, a notification is sent to an employee's terminal, and the employee quickly goes to the lost child, ensuring a safe reunion between parent and child.

[0661] Thus, the system of the present invention enables a rapid response in the event of a child getting lost, providing a safe and secure environment for both facility operators and users.

[0662] The following describes the processing flow.

[0663] Step 1:

[0664] When a user enters the facility, a facial recognition camera captures the visitor's face and generates facial recognition data. The server receives this data and stores it in a database.

[0665] Step 2:

[0666] The user (parent) reports the child getting lost to the support desk and provides a recent photo of the child's face. The server receives this photo data and starts the matching process using a facial recognition algorithm against existing data in the database.

[0667] Step 3:

[0668] The server receives video data from the monitoring device in real time and searches for matching faces through a facial recognition process. This involves algorithmic image analysis and comparison of facial data.

[0669] Step 4:

[0670] When the server identifies a matching target, it generates the child's current location data. This location information is then formatted as map data.

[0671] Step 5:

[0672] The server sends location information as a notification to the terminals of nearby employees. The notification includes location details and instructions on how to get there.

[0673] Step 6:

[0674] The device receives a notification, and the employee quickly moves to the child's location based on the instructions. The device provides navigation information to support efficient movement.

[0675] Step 7:

[0676] An employee finds the lost child, safely takes them into custody, and then guides them back to their parents. This successful event is then fed back from the terminal to the server, completing the lost child tracking process.

[0677] (Example 1)

[0678] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0679] For facility managers, quickly identifying lost children within the facility and safely reuniting them with their guardians is a critical challenge. Current systems lack sufficient facial recognition and location accuracy, making rapid response difficult in some cases. Furthermore, there is a need for more efficient information notification and navigation.

[0680] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0681] In this invention, the server includes means for acquiring and storing visitor facial data, means for comparing provided image data with stored facial data to confirm a match, and means for receiving real-time video data from a monitoring device and determining the location by comparing it with the matching result. This makes it possible to efficiently identify lost children and quickly transmit their location information.

[0682] "Visitor facial data" refers to a collection of identifying information, including facial images and characteristics of individuals entering the facility.

[0683] "Image data" refers to visual information that is stored electronically and can be processed by a system.

[0684] "Verification" is the process of comparing the provided data with the stored data to confirm whether they match.

[0685] "Surveillance equipment" refers to devices installed inside and outside a facility that capture and transmit video in real time.

[0686] "Real-time video data" refers to video information captured by cameras and sensing devices at a given moment.

[0687] A "match result" is the result of the matching process confirming that the provided data is identical or similar to existing data.

[0688] "Means of location determination" refers to a mechanism within a system for determining the precise location of a person or object.

[0689] "Location information" refers to detailed data about a specific location, and may include maps and coordinate information.

[0690] A "communication device" is an electronic device used to send and receive information, and mainly includes mobile terminals and computers.

[0691] "Navigation instructions" are guidelines that show the route and steps necessary to reach a specific location.

[0692] "Re-authentication to confirm safe reunion" is a facial recognition process performed when a parent and child meet face-to-face again, in order to prevent misidentification.

[0693] The present invention will now be described in terms of embodiments for carrying it out. This system is designed to quickly identify lost children in facilities and safely reunite them with their guardians. The process primarily utilizes facial recognition technology, and real-time location tracking and notification are performed based on that information.

[0694] Hardware and software to be used

[0695] The server acquires visitor facial data via cameras installed at each entrance gate within the facility. For this purpose, it uses image processing software such as OpenCV and TensorFlow. This software extracts features from the acquired facial data and stores them in a database. The server also receives video data in real time from surveillance equipment and performs comparison and matching by analyzing it with a facial recognition algorithm. This process requires high system computing power and applies advanced image processing techniques.

[0696] The terminal is a device that receives notifications sent from the server. Specifically, this includes smartphones and tablets, and is operated through a dedicated application compatible with Android OS and iOS. The terminal visually displays location information using the Google Maps API and other tools, and provides navigation instructions to staff.

[0697] When a user (parent) realizes their child is lost, they provide the latest image data of their child through the support desk or support app. This data is a crucial criterion for the facial recognition process on the server. Furthermore, users can confidently cooperate in the search for their child based on the location information received through their device.

[0698] Specific example

[0699] As a concrete example, consider its use in a theme park. If a parent gets separated from their child, a photo of the child is immediately provided at the support desk. This photo is cross-referenced with a server database, and the child's current location is determined through real-time video analysis from surveillance cameras. Based on this information, the location information is sent to employee terminals, enabling a quick reunion between parent and child.

[0700] Example of a prompt

[0701] An example of a prompt to input into a generative AI model is, "Please describe the process of rapidly locating and reuniting lost children using a facial recognition system in a large facility." This prompt is expected to prompt the generative AI model to generate a detailed description of the system's processes and functions.

[0702] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0703] Step 1:

[0704] The server uses cameras installed at each entrance gate of the facility to acquire visitor facial data. The input is real-time video data. The server extracts facial information from this video using a face detection algorithm and stores the features in a database. The output is the facial feature data stored in the database. This process allows for the temporary storage of visitor facial data, preparing it for later matching.

[0705] Step 2:

[0706] When a user (parent) realizes their child is missing, they provide the latest photo of the child to the server via the support desk or a dedicated app. The input to this process is the child's photo data, which the server receives and immediately analyzes using a facial recognition algorithm. The output is a matching result, which confirms that the image matches existing data in the database.

[0707] Step 3:

[0708] The server receives video data in real time from the surveillance camera. The input is a video stream from the surveillance equipment. The server analyzes this video data and compares it with the matching results obtained in step 2 to perform face recognition. This process confirms the match between the face information in the video and the existing data, and identifies the location. The output is the identified location information.

[0709] Step 4:

[0710] The server formats the identified location information in JSON format or another suitable format and sends a notification to the terminals of nearby staff members. The input is the identified location data. The notification includes location details, allowing employees to quickly locate the lost child. The output is navigation instructions displayed on the terminal.

[0711] Step 5:

[0712] The terminal receives notifications from the server and provides visual navigation to employees through its built-in application. The input is location information from the server, and the location is displayed on a map using the Google Maps API. This allows employees to efficiently navigate within the facility. The output consists of a map and a guidance screen.

[0713] Step 6:

[0714] The user (parent) is guided by an employee to a designated location where they are reunited with their child. After the reunion, a re-authentication process using a terminal is performed, and facial recognition is used to confirm safety. The input consists of facial data of both the child and the parent. The server re-authenticates this data and confirms that the parent and child match, preventing misidentification and providing a secure environment. The output is the authentication confirmation result.

[0715] (Application Example 1)

[0716] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0717] In large facilities, there is a need to quickly and reliably locate lost children and safely reunite them with their guardians. However, conventional systems take a long time to pinpoint the location of lost children, making it difficult to reassure guardians. This invention aims to solve this problem and provide peace of mind and safety to facility users.

[0718] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0719] In this invention, the server includes means for acquiring and storing facial recognition data, means for comparing provided visual data with stored facial recognition data to determine a match, means for receiving real-time video information from a monitoring device and determining its location by comparing it with the match determination result, means for notifying the determined location information to the mobile terminal of a nearby staff member, and means for providing a route to the determined location. This enables the rapid identification of a lost child's location and a quick response by staff members.

[0720] "Facial recognition data" refers to information recorded digitally from an individual's face and used for identification purposes.

[0721] "Visual data" refers to visual information, such as images and photographs, recorded as digital data.

[0722] "Matching" is the process of comparing the provided visual data with stored facial recognition data to determine whether they are the same person.

[0723] "Monitoring equipment" is a general term for devices that can capture video in real time and process it as data.

[0724] "Real-time video information" refers to data that can instantly capture and process the current situation as video.

[0725] "Location information" refers to data that indicates the current geographical location of a specific object.

[0726] A "mobile terminal" is a portable device capable of sending, receiving, and displaying information.

[0727] A "travel path" is information that indicates the optimal or designated path to a specific location.

[0728] To implement this invention, it is necessary to build a system that can quickly and reliably find lost children in large facilities. This system works effectively by combining facial recognition and real-time video analysis technologies.

[0729] The server acquires and stores facial recognition data in a database when visitors enter. Existing facial recognition software is used for the facial recognition technology. If a child gets lost, the parent provides the latest photo via a device, and the server uses this visual data to determine if it matches. The facial recognition algorithm compares the stored facial recognition data with the provided visual data to identify the match.

[0730] Meanwhile, monitoring devices are installed in multiple locations within the facility and constantly transmit real-time video information to a server. The server analyzes this video information and identifies videos containing matching facial recognition data. The identified location information is then notified to a mobile terminal. This allows staff to efficiently understand the route to the identified location and quickly guide the lost person to their location.

[0731] As a concrete example, consider a case where a child gets lost in a shopping mall. The parent sends a photo of their face via a smartphone app. The server performs facial recognition, determines the child's location from the video data in real time, and sends the location information, such as "The child is near the food court on the 3rd floor," to the staff member's smartphone. This allows the staff member to quickly go to the scene and guide the lost child back to their parents.

[0732] A concrete example of a prompt for a generated AI model would be: "Explain the steps and techniques necessary to safely reunite a lost child with their parents in a shopping mall. Include specific facial recognition algorithms, video analysis methods, and location information notification processes." Based on this prompt, it is possible to explain the system's detailed operation flow and technical background.

[0733] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0734] Step 1:

[0735] The server stores visitor facial recognition data acquired at the entrance gate into a database. The input is video from surveillance cameras, and the output is digital facial data obtained using a facial recognition algorithm. Faces are detected from the video, feature points are extracted, and stored.

[0736] Step 2:

[0737] When a child goes missing, the user sends the latest visual data of the child from their device to the server. The input is a photograph of the child, and the output is the visual data of that photograph. The data is securely transferred from the device to the server.

[0738] Step 3:

[0739] The server determines whether the received visual data matches the existing face recognition data in the database. The input is the provided visual data and the existing face recognition data, and the output is the matching result. The server then uses a face recognition algorithm to compare each set of data.

[0740] Step 4:

[0741] Real-time video information is transmitted from the monitoring device to the server. The input is live video from the monitoring device, and the output is real-time video information. The server constantly receives and analyzes this real-time video data.

[0742] Step 5:

[0743] The server compares the matching results with real-time video information to pinpoint the lost child's location. The input is the matching results and video information, and the output is the identified location. The entire face visible in the video is analyzed, and the location of the matched face is displayed on a map.

[0744] Step 6:

[0745] The server notifies the assigned person's mobile device of the identified location. The input is location information, and the output is a notification message sent to the mobile device. A message containing navigation information to that location is sent to the assigned person.

[0746] Step 7:

[0747] The terminal guides the assigned person along a route based on the location information it receives. The input is a notification message from the server, and the output is a route suggestion for the assigned person. The route displayed will help the assigned person reach the designated location quickly.

[0748] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0749] This invention is a system that further reduces the psychological burden on parents and facility staff and enables a quick and effective response by incorporating an emotion recognition engine into a lost child tracking system. This system is based on real-time video analysis using facial recognition technology and has the function of dynamically adjusting notification content through emotion recognition.

[0750] server

[0751] The server stores and manages facial recognition data collected at the facility's entrance gates. In the event of a lost child, the server compares the provided child's photo with information in the database using a facial recognition algorithm. Furthermore, the server analyzes real-time video transmitted by monitoring devices and identifies the child's location through facial recognition.

[0752] The emotion recognition engine based on this invention determines the current emotional state based on voice and facial expression data received from parents or facility staff. The server analyzes this emotional information and adjusts the content and format of notifications. For example, if a parent is showing strong anxiety, more detailed location information and reassuring messages are added.

[0753] terminal

[0754] The device receives notifications from the server and provides navigation information to help facility staff quickly locate the lost child. Furthermore, based on the results of the emotion recognition engine, it can present encouraging messages and specific action plans to staff, enabling them to respond calmly.

[0755] User

[0756] When a child goes missing, the user (parent) provides a photo of their child through the support desk or support app and appropriately communicates their emotional state. This allows the system to respond in a way that takes the parent's psychological burden into consideration.

[0757] As a concrete example, if a parent loses sight of their child at a theme park, the emotion recognition engine analyzes the parent's level of anxiety from their tone of voice and facial expression. The server takes this into account and includes additional information in the notification sent to the facility staff's terminal, providing reassurance. This aims to facilitate a smooth reunion between parent and child and create a safe and stress-free environment throughout the facility.

[0758] Thus, the system of the present invention, by combining facial recognition and emotion recognition, realizes a psychologically reassuring approach to lost children, bringing benefits to both facility management and users.

[0759] The following describes the processing flow.

[0760] Step 1:

[0761] When a user enters the facility, a facial recognition camera automatically captures a photograph of the visitor's face. The server receives this facial recognition data and securely stores it in a database.

[0762] Step 2:

[0763] If a user (parent) finds their lost child, they provide a recent photo of the child's face via the support desk or mobile app. The server receives this photo and begins matching it with existing facial recognition data in its database.

[0764] Step 3:

[0765] The server receives video data in real time from the monitoring device and uses a facial recognition algorithm to detect matches between photos and videos. If a match is detected, the server identifies the child's current location and compiles that information.

[0766] Step 4:

[0767] The emotion recognition engine analyzes the user's emotions through voice input and a real-time video interface. Based on these results, the server understands the user's state and adjusts the response accordingly.

[0768] Step 5:

[0769] The server generates an appropriate response message along with the child's precise location information and sends a notification to the terminal of nearby facility staff. The notification includes the parent's emotional state and provides instructions and messages appropriate to the situation.

[0770] Step 6:

[0771] The terminal (staff member's device) receives the notification and acts quickly according to the displayed details and navigation. The terminal also provides messages to staff members to encourage them to act calmly.

[0772] Step 7:

[0773] The user (parent) is reassured after receiving a direct update from facility staff and confirming that their child is safe. The emotion recognition engine confirms the parent's state of reassurance and feeds this information back to the server. The server records the completion of the process within the system.

[0774] (Example 2)

[0775] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0776] When a child gets lost in a public facility, it places a significant psychological burden not only on the parents but also on the facility staff. Conventional lost-child tracking systems often struggle to provide a quick response or alleviate parents' anxiety, which can result in a decrease in confidence in the safety of the facility. This invention aims to improve this situation and enable the rapid discovery of lost children while reducing the psychological burden.

[0777] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0778] In this invention, the server includes means for acquiring and storing facial recognition data, means for comparing and matching provided image data with stored facial recognition data, means for receiving real-time image data from a monitoring device and determining the location by comparing it with the matching result, means for analyzing emotional data from parents and staff and adjusting the notification content, means for transmitting the adjusted notification content to the communication device of a nearby staff member, and means for guiding staff members based on geographical information. This makes it possible to quickly locate children and provide psychological support to parents and staff.

[0779] "Facial recognition data" refers to data in which the characteristics of an individual's face are extracted as digital information and stored in an identifiable format.

[0780] "Image data" refers to still or dynamic visual information stored in digital format.

[0781] "Matching" is the process of comparing the provided data with the stored data and evaluating the degree of similarity.

[0782] A "monitoring device" is a device that monitors a designated environment in real time and acquires necessary image and audio data.

[0783] "Real-time image data" refers to digital data that instantly captures and transmits video footage occurring under certain circumstances.

[0784] "Emotional data" refers to information about emotional states extracted from voice, facial expressions, and actions.

[0785] "Notification content" refers to a message structured to convey specific information to a recipient.

[0786] A "communication device" is a device that combines hardware and software for sending and receiving information.

[0787] "Geographic information" refers to data related to location or place, and is usually presented visually on a map.

[0788] This invention incorporates emotion recognition technology into a lost child tracking system, aiming to quickly and effectively locate lost children within a facility. The system utilizes facial recognition and emotion recognition capabilities to provide a series of processes for processing information and taking appropriate action.

[0789] server

[0790] The server collects facial recognition data in real time from cameras placed throughout the facility and stores it in a management database. This process can utilize OpenCV, a widely used facial recognition algorithm, or general image analysis APIs, which are cloud-based services. When a missing child is reported, the server matches the child's photo provided by the parents or staff with the facial data in the database, and combines this with video data from surveillance devices to determine the child's current location.

[0791] Furthermore, the server analyzes voice and facial expression data from parents and staff using an emotion recognition engine, dynamically adjusting notification content according to each individual's emotional state. Based on these results, it generates a message to provide reassurance and notifies staff members of their mobile devices.

[0792] terminal

[0793] The terminal receives notifications from the server and provides instructions for facility staff to quickly locate the child. Furthermore, it displays advice messages and response procedures that reflect the emotion recognition results, supporting staff in dealing with the situation in the most appropriate way.

[0794] User

[0795] When a child goes missing, the user (parent) provides the child's latest photo and their own emotional state through the support center or a dedicated application. This allows for appropriate responses that take into account the user's psychological state.

[0796] As a concrete example, if a parent loses sight of their child in a theme park, the system will sense the parent's anxiety and provide reassuring guidance. It analyzes data from the parent's voice tone and facial expressions and uses this to include reassuring messages in notifications sent to staff.

[0797] Example of a prompt:

[0798] "Describe a system for finding lost children within a theme park, and provide specific examples of notification features that utilize emotion recognition to alleviate parental anxiety."

[0799] This system enables early detection of children and provides a sense of psychological security to facility users, thereby improving the overall safety and reliability of the facility.

[0800] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0801] Step 1:

[0802] The server collects facial recognition data in real time from cameras within the facility and stores it in a database. The input is raw video data acquired by surveillance cameras. The server processes this video data using a facial recognition algorithm, extracts facial features, and outputs the facial recognition data to the database. This data is retained for subsequent matching processes.

[0803] Step 2:

[0804] The user (parent) provides a photo of their lost child and information about their own emotional state through the facility's support desk or app. The input consists of still image data and emotional information in text / voice format provided by the parent. The server receives the photo data and compares it with face recognition data in a stored database. This process calculates the degree of face matching and outputs the data with the best match.

[0805] Step 3:

[0806] The server acquires real-time video data from the surveillance camera and identifies the child's location by comparing it with the matching results obtained in step 2. The inputs are the real-time video data and the matching results. The server analyzes the video data using a similar facial recognition algorithm and generates the child's location coordinates as output. The location information is represented using GPS data or location coordinates within the facility.

[0807] Step 4:

[0808] The server uses an emotion recognition engine to determine the emotional state based on voice and facial expression data collected from parents and staff. Input can be text, voice data, or image data. The server performs a series of data calculations to analyze this data and outputs an emotional state (e.g., reassured, anxious, tense). This emotional information, along with the child's location information, is incorporated into the notification message.

[0809] Step 5:

[0810] The terminal distributes notification messages received from the server to facility staff and provides navigation information to the child's current location. Input consists of location information and a coordinated notification message from the server. The terminal converts this information into a format understandable to staff and displays it on the screen. This allows staff to quickly and appropriately reach the scene.

[0811] Step 6:

[0812] The user (parent) can receive feedback from the server and check the lost child's status in real time. The input is a status report from the server, which is output as text or audio to the parent's mobile device. This information includes reassuring messages and specific instructions for reunion, and is designed to alleviate the parent's anxiety.

[0813] (Application Example 2)

[0814] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0815] There is a need to alleviate the psychological burden and anxiety felt by parents and caregivers when they lose their children, and to provide support for quickly and effectively finding lost children. However, conventional systems lack the flexibility to respond flexibly to emotional states, and thus have limitations in reducing psychological burden. Therefore, this invention aims to provide a more reassuring lost child response by analyzing the emotional state of parents and caregivers and dynamically adjusting the notification content.

[0816] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0817] In this invention, the server includes means for acquiring and storing facial recognition information, means for comparing provided image information with stored facial recognition information and performing matching, means for receiving current video information from a monitoring device and determining the location by comparing it with the matching result, means for analyzing the emotional state of the parent or administrator using an emotion recognition engine, and means for dynamically adjusting the notification content based on the analyzed emotional information. This makes it possible to alleviate the anxiety of the parent or administrator and to find the lost child safely and quickly.

[0818] "Facial recognition information" refers to data that extracts the facial features of an individual and stores them as digital data.

[0819] "Image information" refers to digital data of photographs and videos acquired through vision.

[0820] "Matching" is the process of comparing points of agreement between two sets of data to determine if they are identical or similar.

[0821] "Surveillance equipment" refers to devices used to record or monitor the surrounding environment, and which can acquire video and audio.

[0822] "Current video information" refers to real-time video data acquired continuously over time.

[0823] An "emotion recognition engine" is a system equipped with algorithms for analyzing and estimating an individual's emotional state from data such as voice and facial expressions.

[0824] "Dynamically adjusting notification content" means changing the content of messages and alerts based on the information acquired, according to the situation at hand and the individual's status.

[0825] To implement this invention, hardware and software for teaching face recognition and emotion recognition technologies are required. The server first acquires face recognition information from cameras installed at entrance gates and various facilities and stores it in a database. This makes it possible to retain the characteristics of the subjects as digital data. Image processing libraries such as OpenCV are used for face recognition.

[0826] If a user loses sight of their child, the parent uses a smartphone application to provide a photo of the child to the server. This information is compared with facial recognition data stored on the server to ensure a proper match. This is then compared with current video footage received from surveillance equipment to pinpoint the child's location. The server then uses this location information to send navigation information to nearby staff members' terminals.

[0827] Furthermore, to determine the user's emotional state, an emotion recognition engine analyzes voice and facial expression data. This process utilizes emotion recognition tools such as EmotionAnalyzer. The emotional information obtained through this analysis is used to dynamically adjust the content of messages sent to staff. For example, if a user is feeling anxious, the notification displayed on the staff terminal will include not only detailed location information but also a message of encouragement.

[0828] A concrete example of its implementation is a case where a parent loses sight of their child at a theme park. Based on the information provided by the parent, an emotion recognition engine analyzes the parent's anxiety, and the server takes appropriate action. The notification includes specific instructions such as, "Your child is currently in the playground on the north side. Please check immediately." By quickly locating the child, the system helps facilitate the reunion of parent and child.

[0829] By using generative AI models, the system can further enhance its ability to analyze parental emotions and provide real-time responses. An example of a prompt message would be, "A child is lost. Analyze the parent's emotions and provide staff with quick and appropriate information." This allows the system to accurately support parents and facility staff, ensuring a safe and secure environment throughout the entire facility.

[0830] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0831] Step 1:

[0832] The server acquires facial recognition information from cameras installed at entrance gates and within the facility. The input is real-time video from the cameras, which is converted into facial recognition data using the OpenCV library and stored in a database. The output is the stored facial recognition information.

[0833] Step 2:

[0834] The user provides photo data of their child to the server via a smartphone app. The input is the image data of the child provided by the user, which the server receives and compares with pre-stored facial recognition information. The output is the matching result.

[0835] Step 3:

[0836] The server uses the current video information received from the monitoring equipment and compares it with the matching results to determine the child's location. The input consists of the matching results and real-time video data, and by comparing these, the server outputs the child's position coordinates.

[0837] Step 4:

[0838] The server notifies nearby staff terminals of the identified location information. The input is location coordinates, and the output is navigation information displayed on the staff terminals. This information is visually represented on a map.

[0839] Step 5:

[0840] The server uses an emotion recognition engine to analyze the user's emotional state from voice and facial expression data. The input is voice and facial expression data, which is analyzed using the EmotionAnalyzer tool. The output is data indicating the user's emotional state.

[0841] Step 6:

[0842] The server dynamically adjusts the content of notifications sent to staff terminals based on the analyzed emotional information. The input is the user's emotional data, and the output is a customized notification message displayed on the terminal. Specifically, if a parent is in a highly anxious state, the notification will include a message to calm them down.

[0843] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0844] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0845] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0846] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0847] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. In the upper and lower directions of the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. Also, the upper side of the concentric circles is where "pleasant" emotions are located, and the lower side is where "unpleasant" emotions are located. In this way, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0848] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0849] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0850] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0851] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0852] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0853] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0854] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0855] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0856] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0857] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0858] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0859] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0860] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0861] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0862] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0863] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0864] The following is further disclosed regarding the embodiments described above.

[0865] (Claim 1)

[0866] A means of acquiring and storing facial recognition data,

[0867] A means of comparing the provided photo data with the stored face recognition data to perform matching,

[0868] A means for receiving real-time video data from a monitoring device and determining the location by comparing it with the matching result,

[0869] A means of notifying the communication device of a nearby person in charge of the identified location information,

[0870] A system that includes this.

[0871] (Claim 2)

[0872] The system according to claim 1, wherein the provided photo data is received from the administrator communication device.

[0873] (Claim 3)

[0874] The system according to claim 1, wherein the location information is provided as map information.

[0875] "Example 1"

[0876] (Claim 1)

[0877] A means of acquiring and storing visitor facial data,

[0878] A means of comparing the provided image data with the stored facial data to confirm a match,

[0879] A means for receiving real-time video data from a monitoring device and determining the location by comparing it with the matching result,

[0880] A means of notifying the communication devices of relevant parties in the vicinity of information about the identified location,

[0881] A means of providing directions and instructions based on location information within the facility,

[0882] A means of re-authentication to confirm safe resumption,

[0883] A system that includes this.

[0884] (Claim 2)

[0885] The system according to claim 1, wherein the provided image data is received from the management communication device.

[0886] (Claim 3)

[0887] The system according to claim 1, wherein the location data is provided as map data.

[0888] "Application Example 1"

[0889] (Claim 1)

[0890] A means of acquiring and storing facial recognition data,

[0891] A means for comparing the provided visual data with stored face recognition data to determine if they match,

[0892] A means for receiving real-time video information from a monitoring device and determining the location by comparing it with the matching determination result,

[0893] A means of notifying the identified location information to the mobile terminal of a nearby staff member,

[0894] Means for providing a path to a specified location,

[0895] A system that includes this.

[0896] (Claim 2)

[0897] The system according to claim 1, wherein the provided visual data is received from an information terminal.

[0898] (Claim 3)

[0899] The system according to claim 1, wherein the aforementioned location information is provided as spatial information.

[0900] "Example 2 of combining an emotion engine"

[0901] (Claim 1)

[0902] A means of acquiring and storing facial recognition data,

[0903] A means for comparing and matching the provided image data with stored face recognition data,

[0904] A means for receiving real-time image data from a monitoring device and determining the position by comparing it with the matching result,

[0905] A means of analyzing emotional data from parents and staff and adjusting the content of notifications,

[0906] A means of transmitting the adjusted notification content to the communication device of a nearby employee,

[0907] A means of guiding staff based on geographical information,

[0908] A system that includes this.

[0909] (Claim 2)

[0910] The system according to claim 1, wherein the provided image data is received from the user's communication device.

[0911] (Claim 3)

[0912] The system according to claim 1, wherein the location information and the adjusted notification content are provided as geographical indication information.

[0913] "Application example 2 when combining with an emotional engine"

[0914] (Claim 1)

[0915] A means of acquiring and storing facial recognition information,

[0916] A means of performing matching by comparing the provided image information with stored face recognition information,

[0917] A means for receiving current video information from a monitoring device and determining the location by comparing it with the matching result,

[0918] A means of notifying the communication device of a nearby person in charge of the identified location information,

[0919] A means for analyzing the emotional state of a parent or caregiver using an emotion recognition engine,

[0920] A means for dynamically adjusting notification content based on analyzed emotional information,

[0921] A system that includes this.

[0922] (Claim 2)

[0923] The system according to claim 1, wherein the provided image information is received from the administrator communication device.

[0924] (Claim 3)

[0925] The system according to claim 1, wherein the location information is provided as map information. [Explanation of Symbols]

[0926] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means of acquiring and storing facial recognition data, A means of comparing the provided photo data with the stored face recognition data to perform matching, A means for receiving real-time video data from a monitoring device and determining the location by comparing it with the matching result, A means of notifying the communication device of a nearby person in charge of the identified location information, A system that includes this.

2. The system according to claim 1, wherein the provided photo data is received from the administrator communication device.

3. The system according to claim 1, wherein the location information is provided as map information.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A