system
A biometric tracking system for large facilities uses facial recognition and communication terminals to quickly locate and protect lost children by comparing visitor data with surveillance footage, enhancing safety and efficiency.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-12-16
- Publication Date
- 2026-06-26
AI Technical Summary
Existing methods for locating lost children in large-scale facilities are inefficient, taking a long time and requiring a large number of personnel, and there is a need for a system that can quickly and accurately find and protect lost children.
A system that acquires biometric information of visitors upon entry, compares it with image data from video equipment, and notifies employees when a match is detected, providing the current location of the found visitor, utilizing facial recognition technology and communication terminals for efficient and accurate tracking.
Enables rapid location and protection of lost children by improving the efficiency and accuracy of the tracking process, allowing for quick reunions with minimal personnel involvement.
Smart Images

Figure 2026105435000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] The problem of lost children occurring in large-scale facilities is a significant mental burden for parents and children. Existing methods often take a long time to find lost children, and there is a need for means to quickly and accurately find children. Also, in order to monitor a wide area within a facility, a large number of personnel are required, so there is a need for a system that can search for lost children efficiently and automatically.
Means for Solving the Problems
[0005] This invention provides a system to solve the problem of lost children in large-scale facilities by acquiring biometric information of visitors upon entry and comparing it with image data acquired from video equipment installed within the facility. Based on the matching results, when matching biometric information is detected, a warning is quickly notified to employees within the facility, and the current location information of the found visitor is provided, thereby realizing a means to quickly locate and protect lost children. Furthermore, by utilizing facial recognition technology and efficiently acquiring and using biometric information via communication terminals, the overall efficiency and accuracy of the system are improved.
[0006] A "large-scale facility" refers to a place where many people gather and which covers a wide area that cannot be fully covered by normal surveillance methods.
[0007] "Biometric information" refers to characteristic information that enables the individual identification of a person, and in particular, information that processes facial features as digital data.
[0008] "Video equipment" refers to devices used to acquire real-time monitorable image data, and includes equipment such as cameras installed within a facility.
[0009] "Image data" refers to visual information acquired from a video device, and is data of still images or videos processed in digital format.
[0010] "Facial recognition technology" refers to technology that extracts specific identifying information from an individual's face and compares and matches that information with other facial information.
[0011] A "communication terminal" is an electronic device capable of sending and receiving digital data, and specifically refers to smartphones and tablets.
[0012] A "warning" refers to an informational notification sent from a system to an employee to draw attention or prompt action under specific circumstances. [Brief explanation of the drawing]
[0013] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]
[0014] An example of an embodiment of the system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0015] First, the terms used in the following description will be explained.
[0016] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0017] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0018] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs, various parameters, and the like. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.
[0019] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0021] [First Embodiment]
[0022] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0023] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0024] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0025] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0026] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0028] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0029] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0030] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0031] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0032] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0033] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0034] This invention is a tracking system for quickly finding and protecting visitors who get lost in large-scale facilities. This system acquires facial data of individual visitors upon entry and then compares it with video data acquired by surveillance cameras within the facility to enable the discovery of lost persons.
[0035] System Configuration
[0036] 1. Server
[0037] The server manages a database containing biometric information and stores facial data acquired upon entry.
[0038] The server processes the video data transmitted in real time from the surveillance cameras and compares it with the stored facial data.
[0039] The server runs a facial recognition algorithm, and when it detects a match in the facial data, it sends an alert to employees in the corresponding area.
[0040] 2. Terminals (employee devices)
[0041] The terminal receives a warning from the server and checks the video information for the relevant area.
[0042] Employees equipped with the device will quickly head to the scene based on the notification, physically confirm the lost child's whereabouts, and take them into their care.
[0043] 3. User (Parent)
[0044] Users provide their child's facial data upon entry. This can be done using a communication device owned by the user.
[0045] When a lost child is found, the user receives a notification from the server and is given directions to the location where the child was found.
[0046] Characteristic behavior
[0047] Acquisition and analysis of facial data
[0048] The server registers facial images of parents and children in a database during the inspection process at the entrance gate. This facial data can also be transmitted to the server via the parent's communication device.
[0049] Real-time monitoring
[0050] The server periodically processes video feeds from cameras installed throughout the facility and compares them with stored facial data. This allows for the instantaneous identification of which area a lost child is in.
[0051] Warnings and notifications
[0052] Based on the facial recognition results, the server generates a warning message if a match is found and sends it to the employee terminal in the relevant area. This message includes a snapshot of the face that matches the location information of the camera in question.
[0053] Specific example
[0054] For example, in the case of a family visiting a theme park, the parents provide a photo of their child's face upon entry. If the parents and child become separated, the parents inform a staff member, and the system immediately begins monitoring. After about two minutes, the server detects a match between the child's face data and security camera footage and sends an alert to the nearest staff member. Upon receiving the notification, the staff member rushes to the designated area and safely retrieves the child. The parents then follow the instructions and can be safely reunited with their child at the designated meeting point.
[0055] This process enables the system to quickly locate lost children and facilitates the rapid reunion of parents and their offspring.
[0056] The following describes the processing flow.
[0057] Step 1:
[0058] Users provide their child's facial data upon entry. This is done using the user's own communication device or a dedicated terminal at the facility. The facial data is sent to a server and registered in a database.
[0059] Step 2:
[0060] The server assigns an identification ID to the facial data registered in the database and stores it securely along with other visitor data.
[0061] Step 3:
[0062] A user reports a lost child to a facility employee. The employee then uses a dedicated application to notify the server of the lost child.
[0063] Step 4:
[0064] Upon receiving a report of a lost child, the server begins acquiring real-time video data from surveillance cameras.
[0065] Step 5:
[0066] The server inputs video data from surveillance cameras into a facial recognition algorithm and performs a comparison with registered facial data.
[0067] Step 6:
[0068] When the server finds a match in facial data, it identifies the current location of the matched person and sends that information to a terminal assigned to the nearby area.
[0069] Step 7:
[0070] Employees with terminals receive notifications from the server, rush to the designated camera location, and identify and protect the designated person.
[0071] Step 8:
[0072] The server receives information from an employee that the lost child has been found and sends a notification to the parents. This notification includes the reunion location and any necessary information.
[0073] Step 9:
[0074] The user follows the server's instructions to the reunion point and is safely reunited with the protected child.
[0075] This series of steps makes it possible to quickly locate lost children within large facilities, providing peace of mind to parents and children.
[0076] (Example 1)
[0077] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0078] In large-scale facilities, the challenge lies in providing effective tracking and protection methods to address the problem of visitors getting lost, which reduces safety and convenience. In particular, there is a need to improve the situation where conventional tracking methods make rapid retrieval difficult, and to streamline the response of facility staff.
[0079] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0080] In this invention, the server includes means for acquiring visitor characteristic information upon entry, means for comparing the acquired characteristic information with video information acquired from a monitoring device, and means for generating an alarm when matching characteristic information is detected based on the comparison. This enables the rapid discovery and protection of lost visitors, thereby improving safety and convenience within the facility.
[0081] A "large-scale facility" refers to a commercial complex or theme park that covers a wide area and can be used by a large number of visitors simultaneously.
[0082] "At the time of entry" refers to the moment when a visitor passes through the entrance of the facility and begins using the facility.
[0083] "Characteristic information" refers to biometric or physical information used to identify visitors, and specifically includes facial information.
[0084] "Means of acquisition" refers to the methods and technologies used to gather necessary information, such as collecting data through entrance gates and communication devices.
[0085] "Monitoring equipment" refers to devices such as cameras and sensors that are placed to record or observe the situation within a facility.
[0086] "Video information" refers to image and video data captured by surveillance devices.
[0087] "Means of comparison" refers to the techniques and processes used to compare acquired information and determine its similarities and differences.
[0088] An "alarm" refers to a notification or alert that is issued when a match is confirmed, and is intended to quickly communicate information to staff within the facility.
[0089] "Means of triggering" refers to methods or systems for initiating notifications or actions when a specific event occurs.
[0090] This invention is a tracking system for quickly locating and safely protecting visitors who get lost in large-scale facilities. The system primarily uses visitor facial information as its main data source and performs information processing based on facial recognition technology. Specifically, it utilizes the following hardware and software.
[0091] The server acquires facial information from cameras installed at the facility's entrance gates and from communication devices owned by users. Standard facial recognition software is used for facial recognition. The acquired facial information is stored in a secure cloud environment. The server also receives real-time video information from surveillance cameras installed within the facility. This video information is compared with the stored facial information using a comparison algorithm, and an alarm is generated each time a match is detected.
[0092] The terminal receives alarm messages sent from the server, providing a means for facility staff to respond quickly. This allows staff to determine the lost child's current location and provide necessary protection. The terminals are provided to facility staff as personal digital assistants or wearable devices.
[0093] Users can participate in the system by sending their child's facial information to the server upon entry. This information is provided using the user's smartphone or other communication device. If a lost child is identified, the user will receive a notification from the server and be guided to the location where the child was found.
[0094] This system can be used, for example, when a family enters a theme park, where parents register their child's photo on their smartphone. If a child gets lost, the server quickly begins monitoring and automatically notifies staff of their location, enabling early discovery and reunion of the child.
[0095] An example of a prompt would be: "Design a facial recognition tracking system for families to find a lost child in a large theme park. Describe in detail how the system should function and the process involved."
[0096] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0097] Step 1:
[0098] Users take photos of themselves and their children at the facility entrance using their smartphones and send them to the server. The input is facial information captured by the smartphone, and the output is registered in the server's database as securely encrypted facial data. Users send data through a communication application, and the server stores the received data on a cloud server.
[0099] Step 2:
[0100] The server collects video data in real time from cameras within the facility. The input is video information acquired from surveillance cameras, and the output is image frames to be supplied to a face recognition algorithm. To process this quickly, the server captures the video information at an appropriate resolution and extracts the face portion.
[0101] Step 3:
[0102] The server compares the collected video data with pre-registered face data. This process uses face recognition technology to detect faces in the input video frames and outputs the results of the match confirmation. The server runs a face recognition algorithm and determines in real time which video frames contain registered faces.
[0103] Step 4:
[0104] If a matching face is detected, the server sends an alert to employee terminals in the corresponding area. The input is the matching face information and its location data, and the output is an alert message. The server attaches a snapshot of the detected face and area information to the notification message and sends it quickly.
[0105] Step 5:
[0106] The terminal receives warning messages sent from the server. The input is the warning message, and the output is an alert that notifies the employee. The terminal displays the received notification on the screen and prompts the staff to take the necessary action.
[0107] Step 6:
[0108] Employees equipped with the device rush to the scene to physically locate and secure the lost visitor. The input is location information obtained from the notification, and the output is the visitor's safe custody. This process allows staff to quickly and efficiently locate the lost child and return them to their parents or guardians.
[0109] Step 7:
[0110] Once the server confirms that a lost child has been found, it notifies the user and provides directions to the location where the child was found. The input is an employee's report of finding a lost child, and the output is a notification message to the parent. The server then helps parents to safely reunite with their children.
[0111] (Application Example 1)
[0112] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0113] In large commercial facilities, it is necessary to quickly locate lost visitors, especially minors, and minimize the time it takes for them to be reunited with their guardians. However, currently, there is a lack of effective means to quickly find lost children, and a system is needed that enables rapid discovery and safe protection.
[0114] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0115] In this invention, the server includes means for acquiring visitor biometric information upon entry, means for comparing the collected biometric information with image data acquired from a video device, and means for identifying and protecting the visitor using real-time video from multiple cameras. This enables the rapid detection and secure protection of visitors.
[0116] A "large-scale facility" refers to a building or site that attracts many visitors and has a large area with multiple sections or floors.
[0117] "Means of acquiring visitor biometric information upon entry" refers to devices or methods that record a visitor's face and other biometric characteristics when they enter a facility.
[0118] "Means for comparing collected biometric information with image data acquired from a video device" refers to a technology or algorithm for comparing and verifying a visitor's biometric information with image data.
[0119] "Means of issuing warnings when matching biometric information is detected" refers to devices or programs that issue alerts when they detect faces or characteristics that match pre-registered biometric information.
[0120] "Means of notifying facility staff of the current location of a discovered visitor" refers to a means of communication that informs facility staff of the identified location of a visitor in real time.
[0121] "Means of identifying and protecting visitors using real-time video from multiple cameras" refers to a process of tracking visitors in real time and ensuring their safety by utilizing a camera network within the facility.
[0122] "Means of notifying the visitor's guardian of the location of a protected visitor" refers to a means of sending a message to inform the guardian of the safe location of the visitor.
[0123] To implement this invention, the server first acquires the visitor's biometric information, specifically facial data, at the facility entrance upon entry. This allows the server to use a facial recognition algorithm to compare and match the visitor's facial data with real-time video data transmitted from multiple surveillance cameras installed within the facility. Existing facial recognition software is used for comparing the facial data. If the discovered visitor matches pre-registered information, the server notifies facility employees of the visitor's current location. The visitor's location is displayed on employee terminals along with a warning message, allowing them to quickly go to a specific area. Furthermore, the server notifies the visitor's guardian of the protected visitor's location information via a communication terminal. This notification is sent using SMS or a dedicated application.
[0124] For example, this system would be useful if a parent and child get separated while visiting a shopping mall. When the parent notifies the facility staff that they and their child are missing, the server immediately begins searching for a matching face in the security camera footage. Once the server identifies the child and detects a match, it sends an alert to the terminal of the nearest staff member. Based on this notification, the staff member rushes to the designated location to find the child. The parent is then reunited with their child at a designated meeting point after receiving a notification.
[0125] An example of a prompt for a generative AI model is: "Consider a specific application example of a lost child prevention system using facial recognition in a shopping mall. How does this system work, and how does it handle situations where a family member gets separated? Please explain, including specific technical means and procedures."
[0126] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0127] Step 1:
[0128] The server acquires biometric information, specifically facial data, from visitors at the facility's entrance. It uses camera footage as input and employs facial recognition software to extract facial features from this footage. The output is registered in a database as individual visitor facial data.
[0129] Step 2:
[0130] The server receives real-time video from surveillance cameras installed within the facility. Using the video feed from the surveillance cameras as input, it detects and extracts facial features from each frame of the video. The resulting output is the facial data contained in the current video frame.
[0131] Step 3:
[0132] The server compares and matches registered face data with face data extracted from camera footage. The input consists of face data in the database and face data obtained from surveillance cameras. This is then matched using a face recognition algorithm, and the output is a numerical representation of the degree of match.
[0133] Step 4:
[0134] The server generates a warning message when the matching score exceeds a certain threshold. It uses the matching score as input and creates a warning message based on the visitor information where a match was detected. The output is the warning message and the visitor's location information.
[0135] Step 5:
[0136] The terminal receives warning messages from the server and notifies facility employees. The input is the warning message from the server, and the output is a notification screen that employees can immediately check.
[0137] Step 6:
[0138] The user receives information from facility staff and notifies lost visitors and their guardians of their location. The input is confirmation information from facility staff, and the output is a rendezvous point guide provided to guardians after the visitor's safety has been confirmed.
[0139] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0140] This invention is a tracking system for quickly locating lost visitors in large-scale facilities and understanding their emotional state to enable appropriate action. The system acquires individual visitor facial data and biometric information upon entry, and then compares this data with video data obtained from surveillance cameras within the facility to locate lost visitors. Furthermore, by using an emotion recognition engine that recognizes the visitor's emotional state from the acquired video data, the system can understand the lost person's psychological state and enable appropriate responses.
[0141] System Configuration
[0142] 1. Server
[0143] The server registers visitors' facial data and biometric information in a database and processes video data acquired from surveillance cameras based on that data.
[0144] The server executes a facial recognition algorithm, compares it with stored facial data, and contributes to finding lost children.
[0145] An emotion recognition engine is used to analyze the emotional state of visitors from video data. Based on the analysis results, the priority of warnings to facility staff is determined.
[0146] 2. Terminals (employee devices)
[0147] The device receives an alert from the server that the lost child has been found, along with the emotional state determined by the emotion recognition engine.
[0148] Employees equipped with devices can use information about visitors' emotional states to interact with them in an appropriate manner, thereby creating a sense of security.
[0149] 3. User (Parent)
[0150] Users can provide their child's facial data upon entry. This is done via the user's own communication device.
[0151] When a lost child is found, users can receive information about the child's emotional state along with the child's location.
[0152] Characteristic behavior
[0153] Acquisition of facial data and analysis of emotions
[0154] The server acquires visitor facial data upon entry and uses this data to compare with surveillance camera footage. Furthermore, an emotion recognition engine analyzes the visitor's emotional state based on the video data.
[0155] Notifications and responses based on emotional state
[0156] In addition to locating lost children, the server analyzes visitors' emotional states and, based on that analysis, notifies facility staff of appropriate responses. For example, if a found child appears anxious, employees will be instructed to approach them with greater empathy.
[0157] Specific example
[0158] For example, if a child who came to a theme park with their mother gets lost, the facial data provided by the mother upon entry is stored on the server. After a report of the lost child is received, the server matches the facial data with surveillance camera footage to locate the child. Simultaneously, an emotion recognition engine analyzes the child's emotional state. If the child's face is determined to show anxiety or fear, the server sends this information to a nearby employee, who then approaches the child cautiously and gently. This information is also sent to the mother, facilitating the child's safety check and a swift reunion. This provides reassurance to the lost child and their family and enables rapid problem resolution.
[0159] The following describes the processing flow.
[0160] Step 1:
[0161] Users provide their child's facial data upon entry, and the server stores this data in a database. Users can also upload facial data using their own communication devices.
[0162] Step 2:
[0163] The server assigns an identification ID to the registered facial data and begins linking with the facility's surveillance camera system.
[0164] Step 3:
[0165] When a user reports a lost child, the server acquires video data from surveillance cameras in real time and attempts to detect the lost child using a facial recognition algorithm.
[0166] Step 4:
[0167] The server detects faces from the video data, compares them with a biometric information database, and finds matching face data.
[0168] Step 5:
[0169] The server analyzes the visitor's emotional state using an emotion recognition engine, along with matching facial data. Emotion recognition identifies feelings such as anxiety, fear, and joy from facial expressions.
[0170] Step 6:
[0171] Based on matching faces and analyzed emotional information, the server sends an alert to a nearby facility employee's terminal. This information includes the current location and emotional state of the discovered visitor.
[0172] Step 7:
[0173] Employees with terminals will check notifications from the server and move to the designated area. Depending on the child's emotional state, they will speak to them gently or take swift action to protect them.
[0174] Step 8:
[0175] The employee reports to the server that the lost child has been found, and the server notifies the parents of this information. The parents are also provided with information about the child's emotional state and instructions on where to meet the child.
[0176] Step 9:
[0177] Based on instructions received from the server, the user travels to the reunion point and is successfully reunited with their child. The approach is tailored to the child's feelings, providing a sense of security.
[0178] This series of steps ensures that lost children within the facility are found and psychological care is provided quickly.
[0179] (Example 2)
[0180] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0181] In large facilities, visitors can get lost, causing anxiety and confusion for both visitors and their families. Furthermore, responding to a lost visitor, especially a child, without understanding their psychological state makes appropriate support difficult. Therefore, a system is needed that can quickly identify visitors and provide the most appropriate response, taking their psychological state into consideration.
[0182] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0183] In this invention, the server includes means for processing visitor biometric data acquired upon entry, means for comparing the acquired biometric data with image data obtained from video equipment, and emotion analysis means for analyzing the visitor's facial expressions to identify their emotional state. This enables rapid identification of visitors and the provision of appropriate support based on their emotional state.
[0184] A "large-scale facility" refers to a building or complex that has a wide area and is likely to attract a large number of people.
[0185] "Visitors" refer to customers or guests who come to large-scale facilities, and are the individuals from whom facial information and biometric data are collected.
[0186] "Biometric data" refers to information representing physical characteristics, such as facial information, that is used to identify individual visitors.
[0187] "Video equipment" refers to devices such as cameras installed within a facility that capture video and images of visitors.
[0188] A "facial recognition method" refers to an algorithm or technology that uses acquired facial information to identify and authenticate an individual.
[0189] An "alarm" is a warning signal or alert that notifies staff who need to respond when the system detects a lost child.
[0190] "Emotional analysis methods" refer to technologies and engines used to analyze and identify the emotional state of visitors from video data.
[0191] "Staff" refers to employees assigned to operate this system, and their role is to handle lost children and assist visitors.
[0192] A "communication device" refers to a device such as a smartphone owned by a visitor or their parent, which is used to send and receive information.
[0193] The system of this invention functions by combining various hardware and software to ensure the safety and comfort of visitors in large-scale facilities. The main components of this system include servers, terminals (employee devices), and user-owned communication devices.
[0194] The server is responsible for acquiring visitors' biometric data, specifically facial information, at the facility's entrance and storing it in a secure cloud database. Based on this information, the server compares it with real-time video data from surveillance cameras installed within the facility. By executing a facial recognition algorithm, lost visitors can be quickly located. The server also uses an emotion recognition engine to analyze the visitor's emotional state from the video data. Based on this analysis, it generates detailed warnings, including emotional states, and sends them to staff terminals.
[0195] The terminal (employee device) can receive information transmitted from the server. This terminal allows for real-time verification of information regarding the identification of lost children and their emotional state, supporting a rapid response at the scene. For example, the displayed map information allows for the precise location of a lost child, enabling the fastest possible arrival at the scene. The terminal also provides instructions on how to respond based on the emotional state; for example, if a visitor is showing strong signs of anxiety, the terminal provides instructions on how to calm them down.
[0196] Users (parents) can use their own communication devices to provide their child's facial data to the server upon entry. This allows parents to receive information on the same communication device if a lost child is identified early. This is to enable parents to quickly reunite with their children by providing immediate notification of the lost child's location and emotional state.
[0197] As a concrete example, consider its use in a theme park. In this case, visitor facial information is registered on a server at the entrance, and surveillance cameras compare the images against that information. If a child gets lost, an emotion recognition engine detects an anxious expression and sends the relevant information to a terminal. Staff members with terminals then take the instructed actions and quickly head to the child's location. This information is also sent to the parents' communication devices, allowing them to understand the situation in real time.
[0198] Examples of prompts that utilize generative AI models to improve the accuracy of emotional state recognition include, "What is the best approach when a child appears anxious?" and "Please explain the benefits of a lost child tracking system using facial recognition technology in large facilities." These prompts can be used to improve emotion analysis and response strategies.
[0199] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0200] Step 1:
[0201] The server acquires facial data through a facial recognition device when a visitor enters the facility. This becomes the input data. This facial data undergoes a cleansing process using image processing technology and is stored in a cloud-based database. The output is the cleaned facial data that has been stored. Specifically, a camera module captures the visitor's face, an AI-based algorithm detects the face, and transfers it to the server in digital format.
[0202] Step 2:
[0203] The server receives video data in real time from surveillance cameras installed within the facility. The input is the video data transmitted from the surveillance cameras. The server then applies a face recognition algorithm and compares it with face data in a database. As part of the data processing, still images are extracted from the video stream and features for face recognition are generated. The output is the ID information of the identified visitor.
[0204] Step 3:
[0205] The server begins analyzing the emotional state of the visitor identified from the facial recognition results. The input is the facial data of the identified visitor, and the emotion recognition engine analyzes the facial expressions based on this data to determine the visitor's emotional state. A generative AI model may be used at this stage. The output is data indicating the visitor's emotional state. Specifically, the emotion model is matched based on the facial feature points to obtain classification results such as "anxiety" or "joy."
[0206] Step 4:
[0207] The server generates an alert and sends it to an employee's terminal when it determines that a identified visitor is lost. The input is the result of facial recognition and emotion analysis, and the output is a lost visitor alert with details about their emotional state. This alert message includes the visitor's current location and suggested actions based on their emotion. Specifically, the server determines the priority and generates a corresponding text message, which is then sent to the terminal.
[0208] Step 5:
[0209] The terminal receives alarms sent from the server and displays them on its interface so that staff can respond. The input is the alarm message from the server, and the output is the staff's prompt on-site response. Specifically, the terminal displays the visitor's location on a map and guides employees using voice guidance and visual alerts.
[0210] Step 6:
[0211] The user receives information notified from the server on their communication device. The input is information about the location and emotional state from the server, and the output is a quick response and reassurance for the lost child. Specifically, a push notification is sent to the parent's smartphone, allowing them to check a real-time map and the child's status.
[0212] (Application Example 2)
[0213] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0214] In large-scale facilities, visitors sometimes get lost, but conventional technology makes it difficult to find them quickly. Furthermore, it is difficult to understand the emotional state of visitors once they are found, making it challenging to provide appropriate support.
[0215] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0216] In this invention, the server includes means for acquiring biometric information of visitors upon entry, means for comparing the collected biometric information with image data acquired from a video device, means for issuing a warning when matching biometric information is detected based on the comparison, and means for notifying facility staff who receive the warning of the current location and emotional state of the discovered visitor, and for facility staff to obtain visual information using a visualization device. This enables the rapid detection of visitors and appropriate responses according to their emotional state.
[0217] A "large-scale facility" is a place where many people gather, and typically includes multiple areas or spaces with different uses.
[0218] "Visitors" refer to people who use or visit a facility.
[0219] "Biometric information" refers to identifiable personal information such as facial data and physical characteristics.
[0220] "Image equipment" refers to cameras or similar devices installed to acquire image data.
[0221] "Image data" refers to visual information acquired through video equipment.
[0222] "Matching" is the process of comparing acquired biometric information with image data to confirm a match.
[0223] A "warning" is an alert that notifies you that a specific event has occurred.
[0224] "Facility staff" refers to the staff and managers who work within the facility.
[0225] "Current location" refers to information that indicates where a visitor is located within the facility.
[0226] "Emotional state" refers to information that indicates the visitor's psychological state, and is usually judged from facial expressions and other such observations.
[0227] A "visualization device" is a pair of glasses or display devices used to visually display information.
[0228] "Visual information" refers to visual data and notifications provided through visualization devices.
[0229] To implement this invention, a server, terminals belonging to facility employees, and communication terminals owned by visitors are utilized.
[0230] Server Role
[0231] The server acquires biometric information from visitors via communication terminals when they enter the facility and registers it in a database. Specifically, it collects and manages biometric information, primarily facial data. This information is compared with image data sent from multiple video devices installed throughout the facility. Facial recognition and emotion analysis technologies are used for the comparison, utilizing open-source libraries such as OpenCV and TENSORFLOW®. When a facial match is confirmed, the server generates a warning and sends this information to the terminals of facility employees. Furthermore, the emotion recognition engine analyzes the visitor's emotional state, and the results are also notified simultaneously.
[0232] Terminal role
[0233] Terminals carried by facility staff receive warnings and emotional status information transmitted from a server. Furthermore, smart glasses and head-mounted displays, acting as visualization devices, display visual information, enabling efficient instruction from supervisors and guidance for lost individuals. This allows staff to respond flexibly to different situations.
[0234] User roles
[0235] Parents of visitors, who are also users, can provide their child's facial data via a communication device. If a lost child is found, they can receive information about the child's location and emotional state in real time. This enables a quick reunion and provides peace of mind to visitors.
[0236] Specific example
[0237] For example, when searching for a lost child in a large theme park, employees can wear smart glasses to check the child's latest location and emotional state, enabling them to take an appropriate approach. If the lost child is found and it is determined that the child is in an anxious state, employees can gently interact with the child and take steps to quickly reunite them with their guardians. This information is also simultaneously transmitted to the parents' communication devices, providing reassurance and enabling an efficient response.
[0238] Examples of prompts for the generating AI:
[0239] "Please provide an overview and specific use cases of a smart glasses application for locating lost children and performing emotion analysis in theme parks."
[0240] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0241] Step 1:
[0242] The server acquires biometric information from visitors via a communication terminal when they enter the facility. This biometric information includes facial data and is registered in a database. The input is biometric information from the communication terminal, and the output is the information registered in the database. Specifically, facial data is received in image format and stored as an identifier.
[0243] Step 2:
[0244] The server receives image data in real time from video equipment installed within the facility. This input image data is then used to perform a matching process against registered face data. The output is either a match or a mismatch. Specifically, the system uses the OpenCV library to perform image analysis and compare facial feature points.
[0245] Step 3:
[0246] The server generates an alert and sends it to the facility employee's terminal when a match is found. This alert includes the visitor's current location information. The input is the matching information from the matching result, and the output is the alert message. Specifically, the alert message is generated based on the location obtained from the facial recognition API and sent over the network.
[0247] Step 4:
[0248] The server analyzes the visitor's emotional state using emotion analysis technology based on video data. The input is video data acquired in real time, and the output is information indicating the visitor's emotional state. Specifically, it uses a TensorFlow model to estimate emotions from facial expressions and categorize them.
[0249] Step 5:
[0250] The terminals used by facility staff display warnings and emotional state information received from the server on a visualization device. Inputs are warning messages and emotional information, while output is visual information displayed on smart glasses. As a specific example, a visitor's photo and location information are overlaid on the display.
[0251] Step 6:
[0252] The user receives lost child discovery information on their communication terminal and checks the child's location and emotional state in real time. Input is a notification from the server, and output is information displayed on the terminal screen. Specifically, the application receives a push notification and displays it as a pop-up on the screen.
[0253] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0254] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0255] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0256] [Second Embodiment]
[0257] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0258] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0259] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0260] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0261] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0262] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0263] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0264] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0265] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0266] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0267] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0268] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0269] This invention is a tracking system for quickly finding and protecting visitors who get lost in large-scale facilities. This system acquires facial data of individual visitors upon entry and then compares it with video data acquired by surveillance cameras within the facility to enable the discovery of lost persons.
[0270] System Configuration
[0271] 1. Server
[0272] The server manages a database containing biometric information and stores facial data acquired upon entry.
[0273] The server processes the video data transmitted in real time from the surveillance cameras and performs matching with the stored face data.
[0274] The server executes a face authentication algorithm and, when detecting a match in the face data, sends a warning to the employees in the corresponding area.
[0275] 2. Terminal (Employee Device)
[0276] The terminal receives a warning from the server and checks the video information in the corresponding area.
[0277] The employee with the terminal promptly heads to the scene based on the notification, physically checks for the missing child, and protects them.
[0278] 3. User (Parent)
[0279] The user provides the face data of the child upon entry. For this, the user can use the communication terminal they own.
[0280] When a missing child is discovered, the user receives a notification from the server and gets guidance on the location where the missing child is protected.
[0281] Characteristic Operations
[0282] Acquisition and Analysis of Face Data
[0283] The server registers the face photos of the parent and child in the database during the inspection process at the entrance gate. This face data can also be transmitted to the server through the communication terminal owned by the parent.
[0284] Real - Time Monitoring [[ID=四十七]]
[0285] The server periodically processes the video feed from the cameras installed throughout the facility and performs matching with the stored face data. Thereby, it can instantly identify in which area the missing child is.
[0286] Warnings and Notifications
[0287] Based on the face recognition result, if there is a match, the server generates a warning message and sends it to the employee terminals in the corresponding area. This message is attached with the location information of the corresponding camera and a snapshot of the matching face.
[0288] Specific Example
[0289] For example, in the case of a family visiting a theme park, the parent provides a face photo of the child upon entry. When the parent and child get separated and the parent informs the staff, the system immediately starts monitoring. After about 2 minutes, the server detects a match between the child's face data and the video of the monitoring camera and sends a warning to the nearby staff. The staff receives the notification and rushes to the designated area to safely protect the child. The parent can then follow the subsequent instructions and reunite with the child safely at the designated meeting point.
[0290] Through this process, the system enables the early detection of lost children and the rapid reunion of parents and children.
[0291] The following describes the processing flow.
[0292] Step 1:
[0293] The user provides the face data of the child upon entry. This is done using the communication terminal owned by the user or the dedicated terminal of the facility. The face data is sent to the server and registered in the database.
[0294] Step 2:
[0295] The server assigns an identification ID to the face data registered in the database and securely stores it together with other visitor data.
[0296] Step 3:
[0297] The user reports the occurrence of a lost child to the facility staff. The staff notifies the server of the lost child using a dedicated application.
[0298] Step 4:
[0299] Upon receiving the report of the lost child, the server starts real-time acquisition of video data from the surveillance cameras.
[0300] Step 5:
[0301] The server inputs the video data from the surveillance cameras into a face recognition algorithm and performs a comparison with the registered face data.
[0302] Step 6:
[0303] When the server finds a match in the face data, it identifies the current location of the matching person and transmits that information to the terminal of the staff in charge of the adjacent area.
[0304] Step 7:
[0305] The staff member with the terminal receives the notification from the server, rushes to the location of the designated camera, and confirms and protects the designated person.
[0306] Step 8:
[0307] The server receives information from the staff that the protection of the lost child is complete and sends a notification to the parent. This notification includes the reunion location and the necessary information.
[0308] Step 9:
[0309] The user goes to the reunion location as instructed by the server and safely reunites with the protected child.
[0310] Through this series of steps, it becomes possible to quickly discover a lost child within a large-scale facility and provide peace of mind to parents and children.
[0311] (Example 1)
[0312] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0313] In large-scale facilities, the challenge lies in providing effective tracking and protection methods to address the problem of visitors getting lost, which reduces safety and convenience. In particular, there is a need to improve the situation where conventional tracking methods make rapid retrieval difficult, and to streamline the response of facility staff.
[0314] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0315] In this invention, the server includes means for acquiring visitor characteristic information upon entry, means for comparing the acquired characteristic information with video information acquired from a monitoring device, and means for generating an alarm when matching characteristic information is detected based on the comparison. This enables the rapid discovery and protection of lost visitors, thereby improving safety and convenience within the facility.
[0316] A "large-scale facility" refers to a commercial complex or theme park that covers a wide area and can be used by a large number of visitors simultaneously.
[0317] "At the time of entry" refers to the moment when a visitor passes through the entrance of the facility and begins using the facility.
[0318] "Characteristic information" refers to biometric or physical information used to identify visitors, and specifically includes facial information.
[0319] "Means of acquisition" refers to the methods and technologies used to gather necessary information, such as collecting data through entrance gates and communication devices.
[0320] "Monitoring equipment" refers to devices such as cameras and sensors that are placed to record or observe the situation within a facility.
[0321] "Video information" refers to image and video data captured by surveillance devices.
[0322] "Means of comparison" refers to the techniques and processes used to compare acquired information and determine its similarities and differences.
[0323] An "alarm" refers to a notification or alert that is issued when a match is confirmed, and is intended to quickly communicate information to staff within the facility.
[0324] "Means of triggering" refers to methods or systems for initiating notifications or actions when a specific event occurs.
[0325] This invention is a tracking system for quickly locating and safely protecting visitors who get lost in large-scale facilities. The system primarily uses visitor facial information as its main data source and performs information processing based on facial recognition technology. Specifically, it utilizes the following hardware and software.
[0326] The server acquires facial information from cameras installed at the facility's entrance gates and from communication devices owned by users. Standard facial recognition software is used for facial recognition. The acquired facial information is stored in a secure cloud environment. The server also receives real-time video information from surveillance cameras installed within the facility. This video information is compared with the stored facial information using a comparison algorithm, and an alarm is generated each time a match is detected.
[0327] The terminal receives alarm messages sent from the server, providing a means for facility staff to respond quickly. This allows staff to determine the lost child's current location and provide necessary protection. The terminals are provided to facility staff as personal digital assistants or wearable devices.
[0328] Users can participate in the system by sending their child's facial information to the server upon entry. This information is provided using the user's smartphone or other communication device. If a lost child is identified, the user will receive a notification from the server and be guided to the location where the child was found.
[0329] This system can be used, for example, when a family enters a theme park, where parents register their child's photo on their smartphone. If a child gets lost, the server quickly begins monitoring and automatically notifies staff of their location, enabling early discovery and reunion of the child.
[0330] An example of a prompt would be: "Design a facial recognition tracking system for families to find a lost child in a large theme park. Describe in detail how the system should function and the process involved."
[0331] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0332] Step 1:
[0333] Users take photos of themselves and their children at the facility entrance using their smartphones and send them to the server. The input is facial information captured by the smartphone, and the output is registered in the server's database as securely encrypted facial data. Users send data through a communication application, and the server stores the received data on a cloud server.
[0334] Step 2:
[0335] The server collects video data in real time from cameras within the facility. The input is video information acquired from surveillance cameras, and the output is image frames to be supplied to a face recognition algorithm. To process this quickly, the server captures the video information at an appropriate resolution and extracts the face portion.
[0336] Step 3:
[0337] The server compares the collected video data with pre-registered face data. This process uses face recognition technology to detect faces in the input video frames and outputs the results of the match confirmation. The server runs a face recognition algorithm and determines in real time which video frames contain registered faces.
[0338] Step 4:
[0339] If a matching face is detected, the server sends an alert to employee terminals in the corresponding area. The input is the matching face information and its location data, and the output is an alert message. The server attaches a snapshot of the detected face and area information to the notification message and sends it quickly.
[0340] Step 5:
[0341] The terminal receives warning messages sent from the server. The input is the warning message, and the output is an alert that notifies the employee. The terminal displays the received notification on the screen and prompts the staff to take the necessary action.
[0342] Step 6:
[0343] Employees equipped with the device rush to the scene to physically locate and secure the lost visitor. The input is location information obtained from the notification, and the output is the visitor's safe custody. This process allows staff to quickly and efficiently locate the lost child and return them to their parents or guardians.
[0344] Step 7:
[0345] Once the server confirms that a lost child has been found, it notifies the user and provides directions to the location where the child was found. The input is an employee's report of finding a lost child, and the output is a notification message to the parent. The server then helps parents to safely reunite with their children.
[0346] (Application Example 1)
[0347] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0348] In large commercial facilities, it is necessary to quickly locate lost visitors, especially minors, and minimize the time it takes for them to be reunited with their guardians. However, currently, there is a lack of effective means to quickly find lost children, and a system is needed that enables rapid discovery and safe protection.
[0349] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0350] In this invention, the server includes means for acquiring visitor biometric information upon entry, means for comparing the collected biometric information with image data acquired from a video device, and means for identifying and protecting the visitor using real-time video from multiple cameras. This enables the rapid detection and secure protection of visitors.
[0351] A "large-scale facility" refers to a building or site that attracts many visitors and has a large area with multiple sections or floors.
[0352] "Means of acquiring visitor biometric information upon entry" refers to devices or methods that record a visitor's face and other biometric characteristics when they enter a facility.
[0353] "Means for comparing collected biometric information with image data acquired from a video device" refers to a technology or algorithm for comparing and verifying a visitor's biometric information with image data.
[0354] "Means of issuing warnings when matching biometric information is detected" refers to devices or programs that issue alerts when they detect faces or characteristics that match pre-registered biometric information.
[0355] "Means of notifying facility staff of the current location of a discovered visitor" refers to a means of communication that informs facility staff of the identified location of a visitor in real time.
[0356] "Means of identifying and protecting visitors using real-time video from multiple cameras" refers to a process of tracking visitors in real time and ensuring their safety by utilizing a camera network within the facility.
[0357] "Means of notifying the visitor's guardian of the location of a protected visitor" refers to a means of sending a message to inform the guardian of the safe location of the visitor.
[0358] To implement this invention, the server first acquires the visitor's biometric information, specifically facial data, at the facility entrance upon entry. This allows the server to use a facial recognition algorithm to compare and match the visitor's facial data with real-time video data transmitted from multiple surveillance cameras installed within the facility. Existing facial recognition software is used for comparing the facial data. If the discovered visitor matches pre-registered information, the server notifies facility employees of the visitor's current location. The visitor's location is displayed on employee terminals along with a warning message, allowing them to quickly go to a specific area. Furthermore, the server notifies the visitor's guardian of the protected visitor's location information via a communication terminal. This notification is sent using SMS or a dedicated application.
[0359] For example, this system would be useful if a parent and child get separated while visiting a shopping mall. When the parent notifies the facility staff that they and their child are missing, the server immediately begins searching for a matching face in the security camera footage. Once the server identifies the child and detects a match, it sends an alert to the terminal of the nearest staff member. Based on this notification, the staff member rushes to the designated location to find the child. The parent is then reunited with their child at a designated meeting point after receiving a notification.
[0360] An example of a prompt for a generative AI model is: "Consider a specific application example of a lost child prevention system using facial recognition in a shopping mall. How does this system work, and how does it handle situations where a family member gets separated? Please explain, including specific technical means and procedures."
[0361] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0362] Step 1:
[0363] The server acquires biometric information, specifically facial data, from visitors at the facility's entrance. It uses camera footage as input and employs facial recognition software to extract facial features from this footage. The output is registered in a database as individual visitor facial data.
[0364] Step 2:
[0365] The server receives real-time video from surveillance cameras installed within the facility. Using the video feed from the surveillance cameras as input, it detects and extracts facial features from each frame of the video. The resulting output is the facial data contained in the current video frame.
[0366] Step 3:
[0367] The server compares and matches registered face data with face data extracted from camera footage. The input consists of face data in the database and face data obtained from surveillance cameras. This is then matched using a face recognition algorithm, and the output is a numerical representation of the degree of match.
[0368] Step 4:
[0369] The server generates a warning message when the matching score exceeds a certain threshold. It uses the matching score as input and creates a warning message based on the visitor information where a match was detected. The output is the warning message and the visitor's location information.
[0370] Step 5:
[0371] The terminal receives warning messages from the server and notifies facility employees. The input is the warning message from the server, and the output is a notification screen that employees can immediately check.
[0372] Step 6:
[0373] The user receives information from facility staff and notifies lost visitors and their guardians of their location. The input is confirmation information from facility staff, and the output is a rendezvous point guide provided to guardians after the visitor's safety has been confirmed.
[0374] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0375] This invention is a tracking system for quickly locating lost visitors in large-scale facilities and understanding their emotional state to enable appropriate action. The system acquires individual visitor facial data and biometric information upon entry, and then compares this data with video data obtained from surveillance cameras within the facility to locate lost visitors. Furthermore, by using an emotion recognition engine that recognizes the visitor's emotional state from the acquired video data, the system can understand the lost person's psychological state and enable appropriate responses.
[0376] System Configuration
[0377] 1. Server
[0378] The server registers visitors' facial data and biometric information in a database and processes video data acquired from surveillance cameras based on that data.
[0379] The server executes a facial recognition algorithm, compares it with stored facial data, and contributes to finding lost children.
[0380] An emotion recognition engine is used to analyze the emotional state of visitors from video data. Based on the analysis results, the priority of warnings to facility staff is determined.
[0381] 2. Terminals (employee devices)
[0382] The device receives an alert from the server that the lost child has been found, along with the emotional state determined by the emotion recognition engine.
[0383] Employees equipped with devices can use information about visitors' emotional states to interact with them in an appropriate manner, thereby creating a sense of security.
[0384] 3. User (Parent)
[0385] Users can provide their child's facial data upon entry. This is done via the user's own communication device.
[0386] When a lost child is found, users can receive information about the child's emotional state along with the child's location.
[0387] Characteristic behavior
[0388] Acquisition of facial data and analysis of emotions
[0389] The server acquires visitor facial data upon entry and uses this data to compare with surveillance camera footage. Furthermore, an emotion recognition engine analyzes the visitor's emotional state based on the video data.
[0390] Notifications and responses based on emotional state
[0391] In addition to locating lost children, the server analyzes visitors' emotional states and, based on that analysis, notifies facility staff of appropriate responses. For example, if a found child appears anxious, employees will be instructed to approach them with greater empathy.
[0392] Specific example
[0393] For example, if a child who came to a theme park with their mother gets lost, the facial data provided by the mother upon entry is stored on the server. After a report of the lost child is received, the server matches the facial data with surveillance camera footage to locate the child. Simultaneously, an emotion recognition engine analyzes the child's emotional state. If the child's face is determined to show anxiety or fear, the server sends this information to a nearby employee, who then approaches the child cautiously and gently. This information is also sent to the mother, facilitating the child's safety check and a swift reunion. This provides reassurance to the lost child and their family and enables rapid problem resolution.
[0394] The following describes the processing flow.
[0395] Step 1:
[0396] Users provide their child's facial data upon entry, and the server stores this data in a database. Users can also upload facial data using their own communication devices.
[0397] Step 2:
[0398] The server assigns an identification ID to the registered facial data and begins linking with the facility's surveillance camera system.
[0399] Step 3:
[0400] When a user reports a lost child, the server acquires video data from surveillance cameras in real time and attempts to detect the lost child using a facial recognition algorithm.
[0401] Step 4:
[0402] The server detects faces from the video data, compares them with a biometric information database, and finds matching face data.
[0403] Step 5:
[0404] The server analyzes the visitor's emotional state using an emotion recognition engine, along with matching facial data. Emotion recognition identifies feelings such as anxiety, fear, and joy from facial expressions.
[0405] Step 6:
[0406] Based on matching faces and analyzed emotional information, the server sends an alert to a nearby facility employee's terminal. This information includes the current location and emotional state of the discovered visitor.
[0407] Step 7:
[0408] Employees with terminals will check notifications from the server and move to the designated area. Depending on the child's emotional state, they will speak to them gently or take swift action to protect them.
[0409] Step 8:
[0410] The employee reports to the server that the lost child has been found, and the server notifies the parents of this information. The parents are also provided with information about the child's emotional state and instructions on where to meet the child.
[0411] Step 9:
[0412] Based on instructions received from the server, the user travels to the reunion point and is successfully reunited with their child. The approach is tailored to the child's feelings, providing a sense of security.
[0413] This series of steps ensures that lost children within the facility are found and psychological care is provided quickly.
[0414] (Example 2)
[0415] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0416] In large facilities, visitors can get lost, causing anxiety and confusion for both visitors and their families. Furthermore, responding to a lost visitor, especially a child, without understanding their psychological state makes appropriate support difficult. Therefore, a system is needed that can quickly identify visitors and provide the most appropriate response, taking their psychological state into consideration.
[0417] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0418] In this invention, the server includes means for processing visitor biometric data acquired upon entry, means for comparing the acquired biometric data with image data obtained from video equipment, and emotion analysis means for analyzing the visitor's facial expressions to identify their emotional state. This enables rapid identification of visitors and the provision of appropriate support based on their emotional state.
[0419] A "large-scale facility" refers to a building or complex that has a wide area and is likely to attract a large number of people.
[0420] "Visitors" refer to customers or guests who come to large-scale facilities, and are the individuals from whom facial information and biometric data are collected.
[0421] "Biometric data" refers to information representing physical characteristics, such as facial information, that is used to identify individual visitors.
[0422] "Video equipment" refers to devices such as cameras installed within a facility that capture video and images of visitors.
[0423] A "facial recognition method" refers to an algorithm or technology that uses acquired facial information to identify and authenticate an individual.
[0424] An "alarm" is a warning signal or alert that notifies staff who need to respond when the system detects a lost child.
[0425] "Emotional analysis methods" refer to technologies and engines used to analyze and identify the emotional state of visitors from video data.
[0426] "Staff" refers to employees assigned to operate this system, and their role is to handle lost children and assist visitors.
[0427] A "communication device" refers to a device such as a smartphone owned by a visitor or their parent, which is used to send and receive information.
[0428] The system of this invention functions by combining various hardware and software to ensure the safety and comfort of visitors in large-scale facilities. The main components of this system include servers, terminals (employee devices), and user-owned communication devices.
[0429] The server is responsible for acquiring visitors' biometric data, specifically facial information, at the facility's entrance and storing it in a secure cloud database. Based on this information, the server compares it with real-time video data from surveillance cameras installed within the facility. By executing a facial recognition algorithm, lost visitors can be quickly located. The server also uses an emotion recognition engine to analyze the visitor's emotional state from the video data. Based on this analysis, it generates detailed warnings, including emotional states, and sends them to staff terminals.
[0430] The terminal (employee device) can receive information transmitted from the server. This terminal allows for real-time verification of information regarding the identification of lost children and their emotional state, supporting a rapid response at the scene. For example, the displayed map information allows for the precise location of a lost child, enabling the fastest possible arrival at the scene. The terminal also provides instructions on how to respond based on the emotional state; for example, if a visitor is showing strong signs of anxiety, the terminal provides instructions on how to calm them down.
[0431] Users (parents) can use their own communication devices to provide their child's facial data to the server upon entry. This allows parents to receive information on the same communication device if a lost child is identified early. This is to enable parents to quickly reunite with their children by providing immediate notification of the lost child's location and emotional state.
[0432] As a concrete example, consider its use in a theme park. In this case, visitor facial information is registered on a server at the entrance, and surveillance cameras compare the images against that information. If a child gets lost, an emotion recognition engine detects an anxious expression and sends the relevant information to a terminal. Staff members with terminals then take the instructed actions and quickly head to the child's location. This information is also sent to the parents' communication devices, allowing them to understand the situation in real time.
[0433] Examples of prompts that utilize generative AI models to improve the accuracy of emotional state recognition include, "What is the best approach when a child appears anxious?" and "Please explain the benefits of a lost child tracking system using facial recognition technology in large facilities." These prompts can be used to improve emotion analysis and response strategies.
[0434] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0435] Step 1:
[0436] The server acquires facial data through a facial recognition device when a visitor enters the facility. This becomes the input data. This facial data undergoes a cleansing process using image processing technology and is stored in a cloud-based database. The output is the cleaned facial data that has been stored. Specifically, a camera module captures the visitor's face, an AI-based algorithm detects the face, and transfers it to the server in digital format.
[0437] Step 2:
[0438] The server receives video data in real time from surveillance cameras installed within the facility. The input is the video data transmitted from the surveillance cameras. The server then applies a face recognition algorithm and compares it with face data in a database. As part of the data processing, still images are extracted from the video stream and features for face recognition are generated. The output is the ID information of the identified visitor.
[0439] Step 3:
[0440] The server begins analyzing the emotional state of the visitor identified from the facial recognition results. The input is the facial data of the identified visitor, and the emotion recognition engine analyzes the facial expressions based on this data to determine the visitor's emotional state. A generative AI model may be used at this stage. The output is data indicating the visitor's emotional state. Specifically, the emotion model is matched based on the facial feature points to obtain classification results such as "anxiety" or "joy."
[0441] Step 4:
[0442] The server generates an alert and sends it to an employee's terminal when it determines that a identified visitor is lost. The input is the result of facial recognition and emotion analysis, and the output is a lost visitor alert with details about their emotional state. This alert message includes the visitor's current location and suggested actions based on their emotion. Specifically, the server determines the priority and generates a corresponding text message, which is then sent to the terminal.
[0443] Step 5:
[0444] The terminal receives alarms sent from the server and displays them on its interface so that staff can respond. The input is the alarm message from the server, and the output is the staff's prompt on-site response. Specifically, the terminal displays the visitor's location on a map and guides employees using voice guidance and visual alerts.
[0445] Step 6:
[0446] The user receives information notified from the server on their communication device. The input is information about the location and emotional state from the server, and the output is a quick response and reassurance for the lost child. Specifically, a push notification is sent to the parent's smartphone, allowing them to check a real-time map and the child's status.
[0447] (Application Example 2)
[0448] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0449] In large-scale facilities, visitors sometimes get lost, but conventional technology makes it difficult to find them quickly. Furthermore, it is difficult to understand the emotional state of visitors once they are found, making it challenging to provide appropriate support.
[0450] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0451] In this invention, the server includes means for acquiring biometric information of visitors upon entry, means for comparing the collected biometric information with image data acquired from a video device, means for issuing a warning when matching biometric information is detected based on the comparison, and means for notifying facility staff who receive the warning of the current location and emotional state of the discovered visitor, and for facility staff to obtain visual information using a visualization device. This enables the rapid detection of visitors and appropriate responses according to their emotional state.
[0452] A "large-scale facility" is a place where many people gather, and typically includes multiple areas or spaces with different uses.
[0453] "Visitors" refer to people who use or visit a facility.
[0454] "Biometric information" refers to identifiable personal information such as facial data and physical characteristics.
[0455] "Image equipment" refers to cameras or similar devices installed to acquire image data.
[0456] "Image data" refers to visual information acquired through video equipment.
[0457] "Matching" is the process of comparing acquired biometric information with image data to confirm a match.
[0458] A "warning" is an alert that notifies you that a specific event has occurred.
[0459] "Facility staff" refers to the staff and managers who work within the facility.
[0460] "Current location" refers to information that indicates where a visitor is located within the facility.
[0461] "Emotional state" refers to information that indicates the visitor's psychological state, and is usually judged from facial expressions and other such observations.
[0462] A "visualization device" is a pair of glasses or display devices used to visually display information.
[0463] "Visual information" refers to visual data and notifications provided through visualization devices.
[0464] To implement this invention, a server, terminals belonging to facility employees, and communication terminals owned by visitors are utilized.
[0465] Server Role
[0466] The server acquires biometric information from visitors via communication terminals when they enter the facility and registers it in a database. Specifically, it collects and manages biometric information, primarily facial data. This information is compared with image data sent from multiple video devices installed throughout the facility. Facial recognition and emotion analysis technologies are used for the comparison, utilizing open-source libraries such as OpenCV and TensorFlow. When a facial match is confirmed, the server generates a warning and sends this information to the terminals of facility employees. Furthermore, the emotion recognition engine analyzes the visitor's emotional state, and the results are also notified simultaneously.
[0467] Terminal role
[0468] Terminals carried by facility staff receive warnings and emotional status information transmitted from a server. Furthermore, smart glasses and head-mounted displays, acting as visualization devices, display visual information, enabling efficient instruction from supervisors and guidance for lost individuals. This allows staff to respond flexibly to different situations.
[0469] User roles
[0470] Parents of visitors, who are also users, can provide their child's facial data via a communication device. If a lost child is found, they can receive information about the child's location and emotional state in real time. This enables a quick reunion and provides peace of mind to visitors.
[0471] Specific example
[0472] For example, when searching for a lost child in a large theme park, employees can wear smart glasses to check the child's latest location and emotional state, enabling them to take an appropriate approach. If the lost child is found and it is determined that the child is in an anxious state, employees can gently interact with the child and take steps to quickly reunite them with their guardians. This information is also simultaneously transmitted to the parents' communication devices, providing reassurance and enabling an efficient response.
[0473] Examples of prompts for the generating AI:
[0474] "Please provide an overview and specific use cases of a smart glasses application for locating lost children and performing emotion analysis in theme parks."
[0475] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0476] Step 1:
[0477] The server acquires biometric information from visitors via a communication terminal when they enter the facility. This biometric information includes facial data and is registered in a database. The input is biometric information from the communication terminal, and the output is the information registered in the database. Specifically, facial data is received in image format and stored as an identifier.
[0478] Step 2:
[0479] The server receives image data in real time from video equipment installed within the facility. This input image data is then used to perform a matching process against registered face data. The output is either a match or a mismatch. Specifically, the system uses the OpenCV library to perform image analysis and compare facial feature points.
[0480] Step 3:
[0481] The server generates an alert and sends it to the facility employee's terminal when a match is found. This alert includes the visitor's current location information. The input is the matching information from the matching result, and the output is the alert message. Specifically, the alert message is generated based on the location obtained from the facial recognition API and sent over the network.
[0482] Step 4:
[0483] The server analyzes the visitor's emotional state using emotion analysis technology based on video data. The input is video data acquired in real time, and the output is information indicating the visitor's emotional state. Specifically, it uses a TensorFlow model to estimate emotions from facial expressions and categorize them.
[0484] Step 5:
[0485] The terminals used by facility staff display warnings and emotional state information received from the server on a visualization device. Inputs are warning messages and emotional information, while output is visual information displayed on smart glasses. As a specific example, a visitor's photo and location information are overlaid on the display.
[0486] Step 6:
[0487] The user receives lost child discovery information on their communication terminal and checks the child's location and emotional state in real time. Input is a notification from the server, and output is information displayed on the terminal screen. Specifically, the application receives a push notification and displays it as a pop-up on the screen.
[0488] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0489] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0490] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0491] [Third Embodiment]
[0492] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0493] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0494] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0495] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0496] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0497] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0498] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0499] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0500] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0501] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0502] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0503] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0504] This invention is a tracking system for quickly finding and protecting visitors who get lost in large-scale facilities. This system acquires facial data of individual visitors upon entry and then compares it with video data acquired by surveillance cameras within the facility to enable the discovery of lost persons.
[0505] System Configuration
[0506] 1. Server
[0507] The server manages a database containing biometric information and stores facial data acquired upon entry.
[0508] The server processes the video data transmitted in real time from the surveillance cameras and compares it with the stored facial data.
[0509] The server runs a facial recognition algorithm, and when it detects a match in the facial data, it sends an alert to employees in the corresponding area.
[0510] 2. Terminals (employee devices)
[0511] The terminal receives a warning from the server and checks the video information for the relevant area.
[0512] Employees equipped with the device will quickly head to the scene based on the notification, physically confirm the lost child's whereabouts, and take them into their care.
[0513] 3. User (Parent)
[0514] Users provide their child's facial data upon entry. This can be done using a communication device owned by the user.
[0515] When a lost child is found, the user receives a notification from the server and is given directions to the location where the child was found.
[0516] Characteristic behavior
[0517] Acquisition and analysis of facial data
[0518] The server registers facial images of parents and children in a database during the inspection process at the entrance gate. This facial data can also be transmitted to the server via the parent's communication device.
[0519] Real-time monitoring
[0520] The server periodically processes video feeds from cameras installed throughout the facility and compares them with stored facial data. This allows for the instantaneous identification of which area a lost child is in.
[0521] Warnings and notifications
[0522] Based on the facial recognition results, the server generates a warning message if a match is found and sends it to the employee terminal in the relevant area. This message includes a snapshot of the face that matches the location information of the camera in question.
[0523] Specific example
[0524] For example, in the case of a family visiting a theme park, the parents provide a photo of their child's face upon entry. If the parents and child become separated, the parents inform a staff member, and the system immediately begins monitoring. After about two minutes, the server detects a match between the child's face data and security camera footage and sends an alert to the nearest staff member. Upon receiving the notification, the staff member rushes to the designated area and safely retrieves the child. The parents then follow the instructions and can be safely reunited with their child at the designated meeting point.
[0525] This process enables the system to quickly locate lost children and facilitates the rapid reunion of parents and their offspring.
[0526] The following describes the processing flow.
[0527] Step 1:
[0528] Users provide their child's facial data upon entry. This is done using the user's own communication device or a dedicated terminal at the facility. The facial data is sent to a server and registered in a database.
[0529] Step 2:
[0530] The server assigns an identification ID to the facial data registered in the database and stores it securely along with other visitor data.
[0531] Step 3:
[0532] A user reports a lost child to a facility employee. The employee then uses a dedicated application to notify the server of the lost child.
[0533] Step 4:
[0534] Upon receiving a report of a lost child, the server begins acquiring real-time video data from surveillance cameras.
[0535] Step 5:
[0536] The server inputs video data from surveillance cameras into a facial recognition algorithm and performs a comparison with registered facial data.
[0537] Step 6:
[0538] When the server finds a match in facial data, it identifies the current location of the matched person and sends that information to a terminal assigned to the nearby area.
[0539] Step 7:
[0540] Employees with terminals receive notifications from the server, rush to the designated camera location, and identify and protect the designated person.
[0541] Step 8:
[0542] The server receives information from an employee that the lost child has been found and sends a notification to the parents. This notification includes the reunion location and any necessary information.
[0543] Step 9:
[0544] The user follows the server's instructions to the reunion point and is safely reunited with the protected child.
[0545] This series of steps makes it possible to quickly locate lost children within large facilities, providing peace of mind to parents and children.
[0546] (Example 1)
[0547] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0548] In large-scale facilities, the challenge lies in providing effective tracking and protection methods to address the problem of visitors getting lost, which reduces safety and convenience. In particular, there is a need to improve the situation where conventional tracking methods make rapid retrieval difficult, and to streamline the response of facility staff.
[0549] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0550] In this invention, the server includes means for acquiring visitor characteristic information upon entry, means for comparing the acquired characteristic information with video information acquired from a monitoring device, and means for generating an alarm when matching characteristic information is detected based on the comparison. This enables the rapid discovery and protection of lost visitors, thereby improving safety and convenience within the facility.
[0551] A "large-scale facility" refers to a commercial complex or theme park that covers a wide area and can be used by a large number of visitors simultaneously.
[0552] "At the time of entry" refers to the moment when a visitor passes through the entrance of the facility and begins using the facility.
[0553] "Characteristic information" refers to biometric or physical information used to identify visitors, and specifically includes facial information.
[0554] "Means of acquisition" refers to the methods and technologies used to gather necessary information, such as collecting data through entrance gates and communication devices.
[0555] "Monitoring equipment" refers to devices such as cameras and sensors that are placed to record or observe the situation within a facility.
[0556] "Video information" refers to image and video data captured by surveillance devices.
[0557] "Means of comparison" refers to the techniques and processes used to compare acquired information and determine its similarities and differences.
[0558] An "alarm" refers to a notification or alert that is issued when a match is confirmed, and is intended to quickly communicate information to staff within the facility.
[0559] "Means of triggering" refers to methods or systems for initiating notifications or actions when a specific event occurs.
[0560] This invention is a tracking system for quickly locating and safely protecting visitors who get lost in large-scale facilities. The system primarily uses visitor facial information as its main data source and performs information processing based on facial recognition technology. Specifically, it utilizes the following hardware and software.
[0561] The server acquires facial information from cameras installed at the facility's entrance gates and from communication devices owned by users. Standard facial recognition software is used for facial recognition. The acquired facial information is stored in a secure cloud environment. The server also receives real-time video information from surveillance cameras installed within the facility. This video information is compared with the stored facial information using a comparison algorithm, and an alarm is generated each time a match is detected.
[0562] The terminal receives alarm messages sent from the server, providing a means for facility staff to respond quickly. This allows staff to determine the lost child's current location and provide necessary protection. The terminals are provided to facility staff as personal digital assistants or wearable devices.
[0563] Users can participate in the system by sending their child's facial information to the server upon entry. This information is provided using the user's smartphone or other communication device. If a lost child is identified, the user will receive a notification from the server and be guided to the location where the child was found.
[0564] This system can be used, for example, when a family enters a theme park, where parents register their child's photo on their smartphone. If a child gets lost, the server quickly begins monitoring and automatically notifies staff of their location, enabling early discovery and reunion of the child.
[0565] An example of a prompt would be: "Design a facial recognition tracking system for families to find a lost child in a large theme park. Describe in detail how the system should function and the process involved."
[0566] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0567] Step 1:
[0568] Users take photos of themselves and their children at the facility entrance using their smartphones and send them to the server. The input is facial information captured by the smartphone, and the output is registered in the server's database as securely encrypted facial data. Users send data through a communication application, and the server stores the received data on a cloud server.
[0569] Step 2:
[0570] The server collects video data in real time from cameras within the facility. The input is video information acquired from surveillance cameras, and the output is image frames to be supplied to a face recognition algorithm. To process this quickly, the server captures the video information at an appropriate resolution and extracts the face portion.
[0571] Step 3:
[0572] The server compares the collected video data with pre-registered face data. This process uses face recognition technology to detect faces in the input video frames and outputs the results of the match confirmation. The server runs a face recognition algorithm and determines in real time which video frames contain registered faces.
[0573] Step 4:
[0574] If a matching face is detected, the server sends an alert to employee terminals in the corresponding area. The input is the matching face information and its location data, and the output is an alert message. The server attaches a snapshot of the detected face and area information to the notification message and sends it quickly.
[0575] Step 5:
[0576] The terminal receives warning messages sent from the server. The input is the warning message, and the output is an alert that notifies the employee. The terminal displays the received notification on the screen and prompts the staff to take the necessary action.
[0577] Step 6:
[0578] Employees equipped with the device rush to the scene to physically locate and secure the lost visitor. The input is location information obtained from the notification, and the output is the visitor's safe custody. This process allows staff to quickly and efficiently locate the lost child and return them to their parents or guardians.
[0579] Step 7:
[0580] Once the server confirms that a lost child has been found, it notifies the user and provides directions to the location where the child was found. The input is an employee's report of finding a lost child, and the output is a notification message to the parent. The server then helps parents to safely reunite with their children.
[0581] (Application Example 1)
[0582] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0583] In large commercial facilities, it is necessary to quickly locate lost visitors, especially minors, and minimize the time it takes for them to be reunited with their guardians. However, currently, there is a lack of effective means to quickly find lost children, and a system is needed that enables rapid discovery and safe protection.
[0584] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0585] In this invention, the server includes means for acquiring visitor biometric information upon entry, means for comparing the collected biometric information with image data acquired from a video device, and means for identifying and protecting the visitor using real-time video from multiple cameras. This enables the rapid detection and secure protection of visitors.
[0586] A "large-scale facility" refers to a building or site that attracts many visitors and has a large area with multiple sections or floors.
[0587] "Means of acquiring visitor biometric information upon entry" refers to devices or methods that record a visitor's face and other biometric characteristics when they enter a facility.
[0588] "Means for comparing collected biometric information with image data acquired from a video device" refers to a technology or algorithm for comparing and verifying a visitor's biometric information with image data.
[0589] "Means of issuing warnings when matching biometric information is detected" refers to devices or programs that issue alerts when they detect faces or characteristics that match pre-registered biometric information.
[0590] "Means of notifying facility staff of the current location of a discovered visitor" refers to a means of communication that informs facility staff of the identified location of a visitor in real time.
[0591] "Means of identifying and protecting visitors using real-time video from multiple cameras" refers to a process of tracking visitors in real time and ensuring their safety by utilizing a camera network within the facility.
[0592] "Means of notifying the visitor's guardian of the location of a protected visitor" refers to a means of sending a message to inform the guardian of the safe location of the visitor.
[0593] To implement this invention, the server first acquires the visitor's biometric information, specifically facial data, at the facility entrance upon entry. This allows the server to use a facial recognition algorithm to compare and match the visitor's facial data with real-time video data transmitted from multiple surveillance cameras installed within the facility. Existing facial recognition software is used for comparing the facial data. If the discovered visitor matches pre-registered information, the server notifies facility employees of the visitor's current location. The visitor's location is displayed on employee terminals along with a warning message, allowing them to quickly go to a specific area. Furthermore, the server notifies the visitor's guardian of the protected visitor's location information via a communication terminal. This notification is sent using SMS or a dedicated application.
[0594] For example, this system would be useful if a parent and child get separated while visiting a shopping mall. When the parent notifies the facility staff that they and their child are missing, the server immediately begins searching for a matching face in the security camera footage. Once the server identifies the child and detects a match, it sends an alert to the terminal of the nearest staff member. Based on this notification, the staff member rushes to the designated location to find the child. The parent is then reunited with their child at a designated meeting point after receiving a notification.
[0595] An example of a prompt for a generative AI model is: "Consider a specific application example of a lost child prevention system using facial recognition in a shopping mall. How does this system work, and how does it handle situations where a family member gets separated? Please explain, including specific technical means and procedures."
[0596] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0597] Step 1:
[0598] The server acquires biometric information, specifically facial data, from visitors at the facility's entrance. It uses camera footage as input and employs facial recognition software to extract facial features from this footage. The output is registered in a database as individual visitor facial data.
[0599] Step 2:
[0600] The server receives real-time video from surveillance cameras installed within the facility. Using the video feed from the surveillance cameras as input, it detects and extracts facial features from each frame of the video. The resulting output is the facial data contained in the current video frame.
[0601] Step 3:
[0602] The server compares and matches registered face data with face data extracted from camera footage. The input consists of face data in the database and face data obtained from surveillance cameras. This is then matched using a face recognition algorithm, and the output is a numerical representation of the degree of match.
[0603] Step 4:
[0604] The server generates a warning message when the matching score exceeds a certain threshold. It uses the matching score as input and creates a warning message based on the visitor information where a match was detected. The output is the warning message and the visitor's location information.
[0605] Step 5:
[0606] The terminal receives warning messages from the server and notifies facility employees. The input is the warning message from the server, and the output is a notification screen that employees can immediately check.
[0607] Step 6:
[0608] The user receives information from facility staff and notifies lost visitors and their guardians of their location. The input is confirmation information from facility staff, and the output is a rendezvous point guide provided to guardians after the visitor's safety has been confirmed.
[0609] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0610] This invention is a tracking system for quickly locating lost visitors in large-scale facilities and understanding their emotional state to enable appropriate action. The system acquires individual visitor facial data and biometric information upon entry, and then compares this data with video data obtained from surveillance cameras within the facility to locate lost visitors. Furthermore, by using an emotion recognition engine that recognizes the visitor's emotional state from the acquired video data, the system can understand the lost person's psychological state and enable appropriate responses.
[0611] System Configuration
[0612] 1. Server
[0613] The server registers visitors' facial data and biometric information in a database and processes video data acquired from surveillance cameras based on that data.
[0614] The server executes a facial recognition algorithm, compares it with stored facial data, and contributes to finding lost children.
[0615] An emotion recognition engine is used to analyze the emotional state of visitors from video data. Based on the analysis results, the priority of warnings to facility staff is determined.
[0616] 2. Terminals (employee devices)
[0617] The device receives an alert from the server that the lost child has been found, along with the emotional state determined by the emotion recognition engine.
[0618] Employees equipped with devices can use information about visitors' emotional states to interact with them in an appropriate manner, thereby creating a sense of security.
[0619] 3. User (Parent)
[0620] Users can provide their child's facial data upon entry. This is done via the user's own communication device.
[0621] When a lost child is found, users can receive information about the child's emotional state along with the child's location.
[0622] Characteristic behavior
[0623] Acquisition of facial data and analysis of emotions
[0624] The server acquires visitor facial data upon entry and uses this data to compare with surveillance camera footage. Furthermore, an emotion recognition engine analyzes the visitor's emotional state based on the video data.
[0625] Notifications and responses based on emotional state
[0626] In addition to locating lost children, the server analyzes visitors' emotional states and, based on that analysis, notifies facility staff of appropriate responses. For example, if a found child appears anxious, employees will be instructed to approach them with greater empathy.
[0627] Specific example
[0628] For example, if a child who came to a theme park with their mother gets lost, the facial data provided by the mother upon entry is stored on the server. After a report of the lost child is received, the server matches the facial data with surveillance camera footage to locate the child. Simultaneously, an emotion recognition engine analyzes the child's emotional state. If the child's face is determined to show anxiety or fear, the server sends this information to a nearby employee, who then approaches the child cautiously and gently. This information is also sent to the mother, facilitating the child's safety check and a swift reunion. This provides reassurance to the lost child and their family and enables rapid problem resolution.
[0629] The following describes the processing flow.
[0630] Step 1:
[0631] Users provide their child's facial data upon entry, and the server stores this data in a database. Users can also upload facial data using their own communication devices.
[0632] Step 2:
[0633] The server assigns an identification ID to the registered facial data and begins linking with the facility's surveillance camera system.
[0634] Step 3:
[0635] When a user reports a lost child, the server acquires video data from surveillance cameras in real time and attempts to detect the lost child using a facial recognition algorithm.
[0636] Step 4:
[0637] The server detects faces from the video data, compares them with a biometric information database, and finds matching face data.
[0638] Step 5:
[0639] The server analyzes the visitor's emotional state using an emotion recognition engine, along with matching facial data. Emotion recognition identifies feelings such as anxiety, fear, and joy from facial expressions.
[0640] Step 6:
[0641] Based on matching faces and analyzed emotional information, the server sends an alert to a nearby facility employee's terminal. This information includes the current location and emotional state of the discovered visitor.
[0642] Step 7:
[0643] Employees with terminals will check notifications from the server and move to the designated area. Depending on the child's emotional state, they will speak to them gently or take swift action to protect them.
[0644] Step 8:
[0645] The employee reports to the server that the lost child has been found, and the server notifies the parents of this information. The parents are also provided with information about the child's emotional state and instructions on where to meet the child.
[0646] Step 9:
[0647] Based on instructions received from the server, the user travels to the reunion point and is successfully reunited with their child. The approach is tailored to the child's feelings, providing a sense of security.
[0648] This series of steps ensures that lost children within the facility are found and psychological care is provided quickly.
[0649] (Example 2)
[0650] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0651] In large facilities, visitors can get lost, causing anxiety and confusion for both visitors and their families. Furthermore, responding to a lost visitor, especially a child, without understanding their psychological state makes appropriate support difficult. Therefore, a system is needed that can quickly identify visitors and provide the most appropriate response, taking their psychological state into consideration.
[0652] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0653] In this invention, the server includes means for processing visitor biometric data acquired upon entry, means for comparing the acquired biometric data with image data obtained from video equipment, and emotion analysis means for analyzing the visitor's facial expressions to identify their emotional state. This enables rapid identification of visitors and the provision of appropriate support based on their emotional state.
[0654] A "large-scale facility" refers to a building or complex that has a wide area and is likely to attract a large number of people.
[0655] "Visitors" refer to customers or guests who come to large-scale facilities, and are the individuals from whom facial information and biometric data are collected.
[0656] "Biometric data" refers to information representing physical characteristics, such as facial information, that is used to identify individual visitors.
[0657] "Video equipment" refers to devices such as cameras installed within a facility that capture video and images of visitors.
[0658] A "facial recognition method" refers to an algorithm or technology that uses acquired facial information to identify and authenticate an individual.
[0659] An "alarm" is a warning signal or alert that notifies staff who need to respond when the system detects a lost child.
[0660] "Emotional analysis methods" refer to technologies and engines used to analyze and identify the emotional state of visitors from video data.
[0661] "Staff" refers to employees assigned to operate this system, and their role is to handle lost children and assist visitors.
[0662] A "communication device" refers to a device such as a smartphone owned by a visitor or their parent, which is used to send and receive information.
[0663] The system of this invention functions by combining various hardware and software to ensure the safety and comfort of visitors in large-scale facilities. The main components of this system include servers, terminals (employee devices), and user-owned communication devices.
[0664] The server is responsible for acquiring visitors' biometric data, specifically facial information, at the facility's entrance and storing it in a secure cloud database. Based on this information, the server compares it with real-time video data from surveillance cameras installed within the facility. By executing a facial recognition algorithm, lost visitors can be quickly located. The server also uses an emotion recognition engine to analyze the visitor's emotional state from the video data. Based on this analysis, it generates detailed warnings, including emotional states, and sends them to staff terminals.
[0665] The terminal (employee device) can receive information transmitted from the server. This terminal allows for real-time verification of information regarding the identification of lost children and their emotional state, supporting a rapid response at the scene. For example, the displayed map information allows for the precise location of a lost child, enabling the fastest possible arrival at the scene. The terminal also provides instructions on how to respond based on the emotional state; for example, if a visitor is showing strong signs of anxiety, the terminal provides instructions on how to calm them down.
[0666] Users (parents) can use their own communication devices to provide their child's facial data to the server upon entry. This allows parents to receive information on the same communication device if a lost child is identified early. This is to enable parents to quickly reunite with their children by providing immediate notification of the lost child's location and emotional state.
[0667] As a concrete example, consider its use in a theme park. In this case, visitor facial information is registered on a server at the entrance, and surveillance cameras compare the images against that information. If a child gets lost, an emotion recognition engine detects an anxious expression and sends the relevant information to a terminal. Staff members with terminals then take the instructed actions and quickly head to the child's location. This information is also sent to the parents' communication devices, allowing them to understand the situation in real time.
[0668] Examples of prompts that utilize generative AI models to improve the accuracy of emotional state recognition include, "What is the best approach when a child appears anxious?" and "Please explain the benefits of a lost child tracking system using facial recognition technology in large facilities." These prompts can be used to improve emotion analysis and response strategies.
[0669] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0670] Step 1:
[0671] The server acquires facial data through a facial recognition device when a visitor enters the facility. This becomes the input data. This facial data undergoes a cleansing process using image processing technology and is stored in a cloud-based database. The output is the cleaned facial data that has been stored. Specifically, a camera module captures the visitor's face, an AI-based algorithm detects the face, and transfers it to the server in digital format.
[0672] Step 2:
[0673] The server receives video data in real time from surveillance cameras installed within the facility. The input is the video data transmitted from the surveillance cameras. The server then applies a face recognition algorithm and compares it with face data in a database. As part of the data processing, still images are extracted from the video stream and features for face recognition are generated. The output is the ID information of the identified visitor.
[0674] Step 3:
[0675] The server begins analyzing the emotional state of the visitor identified from the facial recognition results. The input is the facial data of the identified visitor, and the emotion recognition engine analyzes the facial expressions based on this data to determine the visitor's emotional state. A generative AI model may be used at this stage. The output is data indicating the visitor's emotional state. Specifically, the emotion model is matched based on the facial feature points to obtain classification results such as "anxiety" or "joy."
[0676] Step 4:
[0677] The server generates an alert and sends it to an employee's terminal when it determines that a identified visitor is lost. The input is the result of facial recognition and emotion analysis, and the output is a lost visitor alert with details about their emotional state. This alert message includes the visitor's current location and suggested actions based on their emotion. Specifically, the server determines the priority and generates a corresponding text message, which is then sent to the terminal.
[0678] Step 5:
[0679] The terminal receives alarms sent from the server and displays them on its interface so that staff can respond. The input is the alarm message from the server, and the output is the staff's prompt on-site response. Specifically, the terminal displays the visitor's location on a map and guides employees using voice guidance and visual alerts.
[0680] Step 6:
[0681] The user receives information notified from the server on their communication device. The input is information about the location and emotional state from the server, and the output is a quick response and reassurance for the lost child. Specifically, a push notification is sent to the parent's smartphone, allowing them to check a real-time map and the child's status.
[0682] (Application Example 2)
[0683] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0684] In large-scale facilities, visitors sometimes get lost, but conventional technology makes it difficult to find them quickly. Furthermore, it is difficult to understand the emotional state of visitors once they are found, making it challenging to provide appropriate support.
[0685] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0686] In this invention, the server includes means for acquiring biometric information of visitors upon entry, means for comparing the collected biometric information with image data acquired from a video device, means for issuing a warning when matching biometric information is detected based on the comparison, and means for notifying facility staff who receive the warning of the current location and emotional state of the discovered visitor, and for facility staff to obtain visual information using a visualization device. This enables the rapid detection of visitors and appropriate responses according to their emotional state.
[0687] A "large-scale facility" is a place where many people gather, and typically includes multiple areas or spaces with different uses.
[0688] "Visitors" refer to people who use or visit a facility.
[0689] "Biometric information" refers to identifiable personal information such as facial data and physical characteristics.
[0690] "Image equipment" refers to cameras or similar devices installed to acquire image data.
[0691] "Image data" refers to visual information acquired through video equipment.
[0692] "Matching" is the process of comparing acquired biometric information with image data to confirm a match.
[0693] A "warning" is an alert that notifies you that a specific event has occurred.
[0694] "Facility staff" refers to the staff and managers who work within the facility.
[0695] "Current location" refers to information that indicates where a visitor is located within the facility.
[0696] "Emotional state" refers to information that indicates the visitor's psychological state, and is usually judged from facial expressions and other such observations.
[0697] A "visualization device" is a pair of glasses or display devices used to visually display information.
[0698] "Visual information" refers to visual data and notifications provided through visualization devices.
[0699] To implement this invention, a server, terminals belonging to facility employees, and communication terminals owned by visitors are utilized.
[0700] Server Role
[0701] The server acquires biometric information from visitors via communication terminals when they enter the facility and registers it in a database. Specifically, it collects and manages biometric information, primarily facial data. This information is compared with image data sent from multiple video devices installed throughout the facility. Facial recognition and emotion analysis technologies are used for the comparison, utilizing open-source libraries such as OpenCV and TensorFlow. When a facial match is confirmed, the server generates a warning and sends this information to the terminals of facility employees. Furthermore, the emotion recognition engine analyzes the visitor's emotional state, and the results are also notified simultaneously.
[0702] Terminal role
[0703] Terminals carried by facility staff receive warnings and emotional status information transmitted from a server. Furthermore, smart glasses and head-mounted displays, acting as visualization devices, display visual information, enabling efficient instruction from supervisors and guidance for lost individuals. This allows staff to respond flexibly to different situations.
[0704] User roles
[0705] Parents of visitors, who are also users, can provide their child's facial data via a communication device. If a lost child is found, they can receive information about the child's location and emotional state in real time. This enables a quick reunion and provides peace of mind to visitors.
[0706] Specific example
[0707] For example, when searching for a lost child in a large theme park, employees can wear smart glasses to check the child's latest location and emotional state, enabling them to take an appropriate approach. If the lost child is found and it is determined that the child is in an anxious state, employees can gently interact with the child and take steps to quickly reunite them with their guardians. This information is also simultaneously transmitted to the parents' communication devices, providing reassurance and enabling an efficient response.
[0708] Examples of prompts for the generating AI:
[0709] "Please provide an overview and specific use cases of a smart glasses application for locating lost children and performing emotion analysis in theme parks."
[0710] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0711] Step 1:
[0712] The server acquires biometric information from visitors via a communication terminal when they enter the facility. This biometric information includes facial data and is registered in a database. The input is biometric information from the communication terminal, and the output is the information registered in the database. Specifically, facial data is received in image format and stored as an identifier.
[0713] Step 2:
[0714] The server receives image data in real time from video equipment installed within the facility. This input image data is then used to perform a matching process against registered face data. The output is either a match or a mismatch. Specifically, the system uses the OpenCV library to perform image analysis and compare facial feature points.
[0715] Step 3:
[0716] The server generates an alert and sends it to the facility employee's terminal when a match is found. This alert includes the visitor's current location information. The input is the matching information from the matching result, and the output is the alert message. Specifically, the alert message is generated based on the location obtained from the facial recognition API and sent over the network.
[0717] Step 4:
[0718] The server analyzes the visitor's emotional state using emotion analysis technology based on video data. The input is video data acquired in real time, and the output is information indicating the visitor's emotional state. Specifically, it uses a TensorFlow model to estimate emotions from facial expressions and categorize them.
[0719] Step 5:
[0720] The terminals used by facility staff display warnings and emotional state information received from the server on a visualization device. Inputs are warning messages and emotional information, while output is visual information displayed on smart glasses. As a specific example, a visitor's photo and location information are overlaid on the display.
[0721] Step 6:
[0722] The user receives lost child discovery information on their communication terminal and checks the child's location and emotional state in real time. Input is a notification from the server, and output is information displayed on the terminal screen. Specifically, the application receives a push notification and displays it as a pop-up on the screen.
[0723] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0724] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0725] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0726] [Fourth Embodiment]
[0727] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0728] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0729] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0730] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0731] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0732] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0733] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0734] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0735] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0736] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0737] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0738] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0739] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0740] This invention is a tracking system for quickly finding and protecting visitors who get lost in large-scale facilities. This system acquires facial data of individual visitors upon entry and then compares it with video data acquired by surveillance cameras within the facility to enable the discovery of lost persons.
[0741] System Configuration
[0742] 1. Server
[0743] The server manages a database containing biometric information and stores facial data acquired upon entry.
[0744] The server processes the video data transmitted in real time from the surveillance cameras and compares it with the stored facial data.
[0745] The server runs a facial recognition algorithm, and when it detects a match in the facial data, it sends an alert to employees in the corresponding area.
[0746] 2. Terminals (employee devices)
[0747] The terminal receives a warning from the server and checks the video information for the relevant area.
[0748] Employees equipped with the device will quickly head to the scene based on the notification, physically confirm the lost child's whereabouts, and take them into their care.
[0749] 3. User (Parent)
[0750] Users provide their child's facial data upon entry. This can be done using a communication device owned by the user.
[0751] When a lost child is found, the user receives a notification from the server and is given directions to the location where the child was found.
[0752] Characteristic behavior
[0753] Acquisition and analysis of facial data
[0754] The server registers facial images of parents and children in a database during the inspection process at the entrance gate. This facial data can also be transmitted to the server via the parent's communication device.
[0755] Real-time monitoring
[0756] The server periodically processes video feeds from cameras installed throughout the facility and compares them with stored facial data. This allows for the instantaneous identification of which area a lost child is in.
[0757] Warnings and notifications
[0758] Based on the facial recognition results, the server generates a warning message if a match is found and sends it to the employee terminal in the relevant area. This message includes a snapshot of the face that matches the location information of the camera in question.
[0759] Specific example
[0760] For example, in the case of a family visiting a theme park, the parents provide a photo of their child's face upon entry. If the parents and child become separated, the parents inform a staff member, and the system immediately begins monitoring. After about two minutes, the server detects a match between the child's face data and security camera footage and sends an alert to the nearest staff member. Upon receiving the notification, the staff member rushes to the designated area and safely retrieves the child. The parents then follow the instructions and can be safely reunited with their child at the designated meeting point.
[0761] This process enables the system to quickly locate lost children and facilitates the rapid reunion of parents and their offspring.
[0762] The following describes the processing flow.
[0763] Step 1:
[0764] Users provide their child's facial data upon entry. This is done using the user's own communication device or a dedicated terminal at the facility. The facial data is sent to a server and registered in a database.
[0765] Step 2:
[0766] The server assigns an identification ID to the facial data registered in the database and stores it securely along with other visitor data.
[0767] Step 3:
[0768] A user reports a lost child to a facility employee. The employee then uses a dedicated application to notify the server of the lost child.
[0769] Step 4:
[0770] Upon receiving a report of a lost child, the server begins acquiring real-time video data from surveillance cameras.
[0771] Step 5:
[0772] The server inputs video data from surveillance cameras into a facial recognition algorithm and performs a comparison with registered facial data.
[0773] Step 6:
[0774] When the server finds a match in facial data, it identifies the current location of the matched person and sends that information to a terminal assigned to the nearby area.
[0775] Step 7:
[0776] Employees with terminals receive notifications from the server, rush to the designated camera location, and identify and protect the designated person.
[0777] Step 8:
[0778] The server receives information from an employee that the lost child has been found and sends a notification to the parents. This notification includes the reunion location and any necessary information.
[0779] Step 9:
[0780] The user follows the server's instructions to the reunion point and is safely reunited with the protected child.
[0781] This series of steps makes it possible to quickly locate lost children within large facilities, providing peace of mind to parents and children.
[0782] (Example 1)
[0783] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0784] In large-scale facilities, the challenge lies in providing effective tracking and protection methods to address the problem of visitors getting lost, which reduces safety and convenience. In particular, there is a need to improve the situation where conventional tracking methods make rapid retrieval difficult, and to streamline the response of facility staff.
[0785] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0786] In this invention, the server includes means for acquiring visitor characteristic information upon entry, means for comparing the acquired characteristic information with video information acquired from a monitoring device, and means for generating an alarm when matching characteristic information is detected based on the comparison. This enables the rapid discovery and protection of lost visitors, thereby improving safety and convenience within the facility.
[0787] A "large-scale facility" refers to a commercial complex or theme park that covers a wide area and can be used by a large number of visitors simultaneously.
[0788] "At the time of entry" refers to the moment when a visitor passes through the entrance of the facility and begins using the facility.
[0789] "Characteristic information" refers to biometric or physical information used to identify visitors, and specifically includes facial information.
[0790] "Means of acquisition" refers to the methods and technologies used to gather necessary information, such as collecting data through entrance gates and communication devices.
[0791] "Monitoring equipment" refers to devices such as cameras and sensors that are placed to record or observe the situation within a facility.
[0792] "Video information" refers to image and video data captured by surveillance devices.
[0793] "Means of comparison" refers to the techniques and processes used to compare acquired information and determine its similarities and differences.
[0794] An "alarm" refers to a notification or alert that is issued when a match is confirmed, and is intended to quickly communicate information to staff within the facility.
[0795] "Means of triggering" refers to methods or systems for initiating notifications or actions when a specific event occurs.
[0796] This invention is a tracking system for quickly locating and safely protecting visitors who get lost in large-scale facilities. The system primarily uses visitor facial information as its main data source and performs information processing based on facial recognition technology. Specifically, it utilizes the following hardware and software.
[0797] The server acquires facial information from cameras installed at the facility's entrance gates and from communication devices owned by users. Standard facial recognition software is used for facial recognition. The acquired facial information is stored in a secure cloud environment. The server also receives real-time video information from surveillance cameras installed within the facility. This video information is compared with the stored facial information using a comparison algorithm, and an alarm is generated each time a match is detected.
[0798] The terminal receives alarm messages sent from the server, providing a means for facility staff to respond quickly. This allows staff to determine the lost child's current location and provide necessary protection. The terminals are provided to facility staff as personal digital assistants or wearable devices.
[0799] Users can participate in the system by sending their child's facial information to the server upon entry. This information is provided using the user's smartphone or other communication device. If a lost child is identified, the user will receive a notification from the server and be guided to the location where the child was found.
[0800] This system can be used, for example, when a family enters a theme park, where parents register their child's photo on their smartphone. If a child gets lost, the server quickly begins monitoring and automatically notifies staff of their location, enabling early discovery and reunion of the child.
[0801] An example of a prompt would be: "Design a facial recognition tracking system for families to find a lost child in a large theme park. Describe in detail how the system should function and the process involved."
[0802] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0803] Step 1:
[0804] Users take photos of themselves and their children at the facility entrance using their smartphones and send them to the server. The input is facial information captured by the smartphone, and the output is registered in the server's database as securely encrypted facial data. Users send data through a communication application, and the server stores the received data on a cloud server.
[0805] Step 2:
[0806] The server collects video data in real time from cameras within the facility. The input is video information acquired from surveillance cameras, and the output is image frames to be supplied to a face recognition algorithm. To process this quickly, the server captures the video information at an appropriate resolution and extracts the face portion.
[0807] Step 3:
[0808] The server compares the collected video data with pre-registered face data. This process uses face recognition technology to detect faces in the input video frames and outputs the results of the match confirmation. The server runs a face recognition algorithm and determines in real time which video frames contain registered faces.
[0809] Step 4:
[0810] If a matching face is detected, the server sends an alert to employee terminals in the corresponding area. The input is the matching face information and its location data, and the output is an alert message. The server attaches a snapshot of the detected face and area information to the notification message and sends it quickly.
[0811] Step 5:
[0812] The terminal receives warning messages sent from the server. The input is the warning message, and the output is an alert that notifies the employee. The terminal displays the received notification on the screen and prompts the staff to take the necessary action.
[0813] Step 6:
[0814] Employees equipped with the device rush to the scene to physically locate and secure the lost visitor. The input is location information obtained from the notification, and the output is the visitor's safe custody. This process allows staff to quickly and efficiently locate the lost child and return them to their parents or guardians.
[0815] Step 7:
[0816] Once the server confirms that a lost child has been found, it notifies the user and provides directions to the location where the child was found. The input is an employee's report of finding a lost child, and the output is a notification message to the parent. The server then helps parents to safely reunite with their children.
[0817] (Application Example 1)
[0818] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0819] In large commercial facilities, it is necessary to quickly locate lost visitors, especially minors, and minimize the time it takes for them to be reunited with their guardians. However, currently, there is a lack of effective means to quickly find lost children, and a system is needed that enables rapid discovery and safe protection.
[0820] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0821] In this invention, the server includes means for acquiring visitor biometric information upon entry, means for comparing the collected biometric information with image data acquired from a video device, and means for identifying and protecting the visitor using real-time video from multiple cameras. This enables the rapid detection and secure protection of visitors.
[0822] A "large-scale facility" refers to a building or site that attracts many visitors and has a large area with multiple sections or floors.
[0823] "Means of acquiring visitor biometric information upon entry" refers to devices or methods that record a visitor's face and other biometric characteristics when they enter a facility.
[0824] "Means for comparing collected biometric information with image data acquired from a video device" refers to a technology or algorithm for comparing and verifying a visitor's biometric information with image data.
[0825] "Means of issuing warnings when matching biometric information is detected" refers to devices or programs that issue alerts when they detect faces or characteristics that match pre-registered biometric information.
[0826] "Means of notifying facility staff of the current location of a discovered visitor" refers to a means of communication that informs facility staff of the identified location of a visitor in real time.
[0827] "Means of identifying and protecting visitors using real-time video from multiple cameras" refers to a process of tracking visitors in real time and ensuring their safety by utilizing a camera network within the facility.
[0828] "Means of notifying the visitor's guardian of the location of a protected visitor" refers to a means of sending a message to inform the guardian of the safe location of the visitor.
[0829] To implement this invention, the server first acquires the visitor's biometric information, specifically facial data, at the facility entrance upon entry. This allows the server to use a facial recognition algorithm to compare and match the visitor's facial data with real-time video data transmitted from multiple surveillance cameras installed within the facility. Existing facial recognition software is used for comparing the facial data. If the discovered visitor matches pre-registered information, the server notifies facility employees of the visitor's current location. The visitor's location is displayed on employee terminals along with a warning message, allowing them to quickly go to a specific area. Furthermore, the server notifies the visitor's guardian of the protected visitor's location information via a communication terminal. This notification is sent using SMS or a dedicated application.
[0830] For example, this system would be useful if a parent and child get separated while visiting a shopping mall. When the parent notifies the facility staff that they and their child are missing, the server immediately begins searching for a matching face in the security camera footage. Once the server identifies the child and detects a match, it sends an alert to the terminal of the nearest staff member. Based on this notification, the staff member rushes to the designated location to find the child. The parent is then reunited with their child at a designated meeting point after receiving a notification.
[0831] An example of a prompt for a generative AI model is: "Consider a specific application example of a lost child prevention system using facial recognition in a shopping mall. How does this system work, and how does it handle situations where a family member gets separated? Please explain, including specific technical means and procedures."
[0832] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0833] Step 1:
[0834] The server acquires biometric information, specifically facial data, from visitors at the facility's entrance. It uses camera footage as input and employs facial recognition software to extract facial features from this footage. The output is registered in a database as individual visitor facial data.
[0835] Step 2:
[0836] The server receives real-time video from surveillance cameras installed within the facility. Using the video feed from the surveillance cameras as input, it detects and extracts facial features from each frame of the video. The resulting output is the facial data contained in the current video frame.
[0837] Step 3:
[0838] The server compares and matches registered face data with face data extracted from camera footage. The input consists of face data in the database and face data obtained from surveillance cameras. This is then matched using a face recognition algorithm, and the output is a numerical representation of the degree of match.
[0839] Step 4:
[0840] The server generates a warning message when the matching score exceeds a certain threshold. It uses the matching score as input and creates a warning message based on the visitor information where a match was detected. The output is the warning message and the visitor's location information.
[0841] Step 5:
[0842] The terminal receives warning messages from the server and notifies facility employees. The input is the warning message from the server, and the output is a notification screen that employees can immediately check.
[0843] Step 6:
[0844] The user receives information from facility staff and notifies lost visitors and their guardians of their location. The input is confirmation information from facility staff, and the output is a rendezvous point guide provided to guardians after the visitor's safety has been confirmed.
[0845] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0846] This invention is a tracking system for quickly locating lost visitors in large-scale facilities and understanding their emotional state to enable appropriate action. The system acquires individual visitor facial data and biometric information upon entry, and then compares this data with video data obtained from surveillance cameras within the facility to locate lost visitors. Furthermore, by using an emotion recognition engine that recognizes the visitor's emotional state from the acquired video data, the system can understand the lost person's psychological state and enable appropriate responses.
[0847] System Configuration
[0848] 1. Server
[0849] The server registers visitors' facial data and biometric information in a database and processes video data acquired from surveillance cameras based on that data.
[0850] The server executes a facial recognition algorithm, compares it with stored facial data, and contributes to finding lost children.
[0851] An emotion recognition engine is used to analyze the emotional state of visitors from video data. Based on the analysis results, the priority of warnings to facility staff is determined.
[0852] 2. Terminals (employee devices)
[0853] The device receives an alert from the server that the lost child has been found, along with the emotional state determined by the emotion recognition engine.
[0854] Employees equipped with devices can use information about visitors' emotional states to interact with them in an appropriate manner, thereby creating a sense of security.
[0855] 3. User (Parent)
[0856] Users can provide their child's facial data upon entry. This is done via the user's own communication device.
[0857] When a lost child is found, users can receive information about the child's emotional state along with the child's location.
[0858] Characteristic behavior
[0859] Acquisition of facial data and analysis of emotions
[0860] The server acquires visitor facial data upon entry and uses this data to compare with surveillance camera footage. Furthermore, an emotion recognition engine analyzes the visitor's emotional state based on the video data.
[0861] Notifications and responses based on emotional state
[0862] In addition to locating lost children, the server analyzes visitors' emotional states and, based on that analysis, notifies facility staff of appropriate responses. For example, if a found child appears anxious, employees will be instructed to approach them with greater empathy.
[0863] Specific example
[0864] For example, if a child who came to a theme park with their mother gets lost, the facial data provided by the mother upon entry is stored on the server. After a report of the lost child is received, the server matches the facial data with surveillance camera footage to locate the child. Simultaneously, an emotion recognition engine analyzes the child's emotional state. If the child's face is determined to show anxiety or fear, the server sends this information to a nearby employee, who then approaches the child cautiously and gently. This information is also sent to the mother, facilitating the child's safety check and a swift reunion. This provides reassurance to the lost child and their family and enables rapid problem resolution.
[0865] The following describes the processing flow.
[0866] Step 1:
[0867] Users provide their child's facial data upon entry, and the server stores this data in a database. Users can also upload facial data using their own communication devices.
[0868] Step 2:
[0869] The server assigns an identification ID to the registered facial data and begins linking with the facility's surveillance camera system.
[0870] Step 3:
[0871] When a user reports a lost child, the server acquires video data from surveillance cameras in real time and attempts to detect the lost child using a facial recognition algorithm.
[0872] Step 4:
[0873] The server detects faces from the video data, compares them with a biometric information database, and finds matching face data.
[0874] Step 5:
[0875] The server analyzes the visitor's emotional state using an emotion recognition engine, along with matching facial data. Emotion recognition identifies feelings such as anxiety, fear, and joy from facial expressions.
[0876] Step 6:
[0877] Based on matching faces and analyzed emotional information, the server sends an alert to a nearby facility employee's terminal. This information includes the current location and emotional state of the discovered visitor.
[0878] Step 7:
[0879] Employees with terminals will check notifications from the server and move to the designated area. Depending on the child's emotional state, they will speak to them gently or take swift action to protect them.
[0880] Step 8:
[0881] The employee reports to the server that the lost child has been found, and the server notifies the parents of this information. The parents are also provided with information about the child's emotional state and instructions on where to meet the child.
[0882] Step 9:
[0883] Based on instructions received from the server, the user travels to the reunion point and is successfully reunited with their child. The approach is tailored to the child's feelings, providing a sense of security.
[0884] This series of steps ensures that lost children within the facility are found and psychological care is provided quickly.
[0885] (Example 2)
[0886] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0887] In large facilities, visitors can get lost, causing anxiety and confusion for both visitors and their families. Furthermore, responding to a lost visitor, especially a child, without understanding their psychological state makes appropriate support difficult. Therefore, a system is needed that can quickly identify visitors and provide the most appropriate response, taking their psychological state into consideration.
[0888] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0889] In this invention, the server includes means for processing visitor biometric data acquired upon entry, means for comparing the acquired biometric data with image data obtained from video equipment, and emotion analysis means for analyzing the visitor's facial expressions to identify their emotional state. This enables rapid identification of visitors and the provision of appropriate support based on their emotional state.
[0890] A "large-scale facility" refers to a building or complex that has a wide area and is likely to attract a large number of people.
[0891] "Visitors" refer to customers or guests who come to large-scale facilities, and are the individuals from whom facial information and biometric data are collected.
[0892] "Biometric data" refers to information representing physical characteristics, such as facial information, that is used to identify individual visitors.
[0893] "Video equipment" refers to devices such as cameras installed within a facility that capture video and images of visitors.
[0894] A "facial recognition method" refers to an algorithm or technology that uses acquired facial information to identify and authenticate an individual.
[0895] An "alarm" is a warning signal or alert that notifies staff who need to respond when the system detects a lost child.
[0896] "Emotional analysis methods" refer to technologies and engines used to analyze and identify the emotional state of visitors from video data.
[0897] "Staff" refers to employees assigned to operate this system, and their role is to handle lost children and assist visitors.
[0898] A "communication device" refers to a device such as a smartphone owned by a visitor or their parent, which is used to send and receive information.
[0899] The system of this invention functions by combining various hardware and software to ensure the safety and comfort of visitors in large-scale facilities. The main components of this system include servers, terminals (employee devices), and user-owned communication devices.
[0900] The server is responsible for acquiring visitors' biometric data, specifically facial information, at the facility's entrance and storing it in a secure cloud database. Based on this information, the server compares it with real-time video data from surveillance cameras installed within the facility. By executing a facial recognition algorithm, lost visitors can be quickly located. The server also uses an emotion recognition engine to analyze the visitor's emotional state from the video data. Based on this analysis, it generates detailed warnings, including emotional states, and sends them to staff terminals.
[0901] The terminal (employee device) can receive information transmitted from the server. This terminal allows for real-time verification of information regarding the identification of lost children and their emotional state, supporting a rapid response at the scene. For example, the displayed map information allows for the precise location of a lost child, enabling the fastest possible arrival at the scene. The terminal also provides instructions on how to respond based on the emotional state; for example, if a visitor is showing strong signs of anxiety, the terminal provides instructions on how to calm them down.
[0902] Users (parents) can use their own communication devices to provide their child's facial data to the server upon entry. This allows parents to receive information on the same communication device if a lost child is identified early. This is to enable parents to quickly reunite with their children by providing immediate notification of the lost child's location and emotional state.
[0903] As a concrete example, consider its use in a theme park. In this case, visitor facial information is registered on a server at the entrance, and surveillance cameras compare the images against that information. If a child gets lost, an emotion recognition engine detects an anxious expression and sends the relevant information to a terminal. Staff members with terminals then take the instructed actions and quickly head to the child's location. This information is also sent to the parents' communication devices, allowing them to understand the situation in real time.
[0904] Examples of prompts that utilize generative AI models to improve the accuracy of emotional state recognition include, "What is the best approach when a child appears anxious?" and "Please explain the benefits of a lost child tracking system using facial recognition technology in large facilities." These prompts can be used to improve emotion analysis and response strategies.
[0905] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0906] Step 1:
[0907] The server acquires facial data through a facial recognition device when a visitor enters the facility. This becomes the input data. This facial data undergoes a cleansing process using image processing technology and is stored in a cloud-based database. The output is the cleaned facial data that has been stored. Specifically, a camera module captures the visitor's face, an AI-based algorithm detects the face, and transfers it to the server in digital format.
[0908] Step 2:
[0909] The server receives video data in real time from surveillance cameras installed within the facility. The input is the video data transmitted from the surveillance cameras. The server then applies a face recognition algorithm and compares it with face data in a database. As part of the data processing, still images are extracted from the video stream and features for face recognition are generated. The output is the ID information of the identified visitor.
[0910] Step 3:
[0911] The server begins analyzing the emotional state of the visitor identified from the facial recognition results. The input is the facial data of the identified visitor, and the emotion recognition engine analyzes the facial expressions based on this data to determine the visitor's emotional state. A generative AI model may be used at this stage. The output is data indicating the visitor's emotional state. Specifically, the emotion model is matched based on the facial feature points to obtain classification results such as "anxiety" or "joy."
[0912] Step 4:
[0913] The server generates an alert and sends it to an employee's terminal when it determines that a identified visitor is lost. The input is the result of facial recognition and emotion analysis, and the output is a lost visitor alert with details about their emotional state. This alert message includes the visitor's current location and suggested actions based on their emotion. Specifically, the server determines the priority and generates a corresponding text message, which is then sent to the terminal.
[0914] Step 5:
[0915] The terminal receives alarms sent from the server and displays them on its interface so that staff can respond. The input is the alarm message from the server, and the output is the staff's prompt on-site response. Specifically, the terminal displays the visitor's location on a map and guides employees using voice guidance and visual alerts.
[0916] Step 6:
[0917] The user receives information notified from the server on their communication device. The input is information about the location and emotional state from the server, and the output is a quick response and reassurance for the lost child. Specifically, a push notification is sent to the parent's smartphone, allowing them to check a real-time map and the child's status.
[0918] (Application Example 2)
[0919] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0920] In large-scale facilities, visitors sometimes get lost, but conventional technology makes it difficult to find them quickly. Furthermore, it is difficult to understand the emotional state of visitors once they are found, making it challenging to provide appropriate support.
[0921] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0922] In this invention, the server includes means for acquiring biometric information of visitors upon entry, means for comparing the collected biometric information with image data acquired from a video device, means for issuing a warning when matching biometric information is detected based on the comparison, and means for notifying facility staff who receive the warning of the current location and emotional state of the discovered visitor, and for facility staff to obtain visual information using a visualization device. This enables the rapid detection of visitors and appropriate responses according to their emotional state.
[0923] A "large-scale facility" is a place where many people gather, and typically includes multiple areas or spaces with different uses.
[0924] "Visitors" refer to people who use or visit a facility.
[0925] "Biometric information" refers to identifiable personal information such as facial data and physical characteristics.
[0926] "Image equipment" refers to cameras or similar devices installed to acquire image data.
[0927] "Image data" refers to visual information acquired through video equipment.
[0928] "Matching" is the process of comparing acquired biometric information with image data to confirm a match.
[0929] A "warning" is an alert that notifies you that a specific event has occurred.
[0930] "Facility staff" refers to the staff and managers who work within the facility.
[0931] "Current location" refers to information that indicates where a visitor is located within the facility.
[0932] "Emotional state" refers to information that indicates the visitor's psychological state, and is usually judged from facial expressions and other such observations.
[0933] A "visualization device" is a pair of glasses or display devices used to visually display information.
[0934] "Visual information" refers to visual data and notifications provided through visualization devices.
[0935] To implement this invention, a server, terminals belonging to facility employees, and communication terminals owned by visitors are utilized.
[0936] Server Role
[0937] The server acquires biometric information from visitors via communication terminals when they enter the facility and registers it in a database. Specifically, it collects and manages biometric information, primarily facial data. This information is compared with image data sent from multiple video devices installed throughout the facility. Facial recognition and emotion analysis technologies are used for the comparison, utilizing open-source libraries such as OpenCV and TensorFlow. When a facial match is confirmed, the server generates a warning and sends this information to the terminals of facility employees. Furthermore, the emotion recognition engine analyzes the visitor's emotional state, and the results are also notified simultaneously.
[0938] Terminal role
[0939] Terminals carried by facility staff receive warnings and emotional status information transmitted from a server. Furthermore, smart glasses and head-mounted displays, acting as visualization devices, display visual information, enabling efficient instruction from supervisors and guidance for lost individuals. This allows staff to respond flexibly to different situations.
[0940] User roles
[0941] Parents of visitors, who are also users, can provide their child's facial data via a communication device. If a lost child is found, they can receive information about the child's location and emotional state in real time. This enables a quick reunion and provides peace of mind to visitors.
[0942] Specific example
[0943] For example, when searching for a lost child in a large theme park, employees can wear smart glasses to check the child's latest location and emotional state, enabling them to take an appropriate approach. If the lost child is found and it is determined that the child is in an anxious state, employees can gently interact with the child and take steps to quickly reunite them with their guardians. This information is also simultaneously transmitted to the parents' communication devices, providing reassurance and enabling an efficient response.
[0944] Examples of prompts for the generating AI:
[0945] "Please provide an overview and specific use cases of a smart glasses application for locating lost children and performing emotion analysis in theme parks."
[0946] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0947] Step 1:
[0948] The server acquires biometric information from visitors via a communication terminal when they enter the facility. This biometric information includes facial data and is registered in a database. The input is biometric information from the communication terminal, and the output is the information registered in the database. Specifically, facial data is received in image format and stored as an identifier.
[0949] Step 2:
[0950] The server receives image data in real time from video equipment installed within the facility. This input image data is then used to perform a matching process against registered face data. The output is either a match or a mismatch. Specifically, the system uses the OpenCV library to perform image analysis and compare facial feature points.
[0951] Step 3:
[0952] The server generates an alert and sends it to the facility employee's terminal when a match is found. This alert includes the visitor's current location information. The input is the matching information from the matching result, and the output is the alert message. Specifically, the alert message is generated based on the location obtained from the facial recognition API and sent over the network.
[0953] Step 4:
[0954] The server analyzes the visitor's emotional state using emotion analysis technology based on video data. The input is video data acquired in real time, and the output is information indicating the visitor's emotional state. Specifically, it uses a TensorFlow model to estimate emotions from facial expressions and categorize them.
[0955] Step 5:
[0956] The terminals used by facility staff display warnings and emotional state information received from the server on a visualization device. Inputs are warning messages and emotional information, while output is visual information displayed on smart glasses. As a specific example, a visitor's photo and location information are overlaid on the display.
[0957] Step 6:
[0958] The user receives lost child discovery information on their communication terminal and checks the child's location and emotional state in real time. Input is a notification from the server, and output is information displayed on the terminal screen. Specifically, the application receives a push notification and displays it as a pop-up on the screen.
[0959] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0960] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0961] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0962] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0963] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0964] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0965] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0966] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0967] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0968] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0969] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0970] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0971] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0972] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0973] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0974] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0975] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0976] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0977] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0978] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0979] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0980] The following is further disclosed regarding the embodiments described above.
[0981] (Claim 1)
[0982] In large-scale facilities, a means of acquiring visitors' biometric information upon entry,
[0983] A means for comparing collected biometric information with image data acquired from a video device,
[0984] A means of issuing a warning when matching biometric information is detected based on the matching process,
[0985] A means of notifying facility staff who receive a warning of the current location of the discovered visitor,
[0986] A system that includes this.
[0987] (Claim 2)
[0988] The system according to claim 1, which uses facial data as biometric information and matches it with image data using facial recognition technology.
[0989] (Claim 3)
[0990] The system according to claim 1, comprising means of using a communication terminal owned by the visitor to provide the visitor's biometric information.
[0991] "Example 1"
[0992] (Claim 1)
[0993] In large-scale facilities, a means of acquiring visitor characteristic information upon entry,
[0994] A means of comparing acquired feature information with video information acquired from a monitoring device,
[0995] A means for generating an alarm when matching feature information is detected based on comparison,
[0996] A means of notifying facility staff who received the alarm of the confirmed visitor's current location,
[0997] A system that includes this.
[0998] (Claim 2)
[0999] The system according to claim 1, which uses facial information as feature information and compares it with video information using facial recognition technology.
[1000] (Claim 3)
[1001] The system according to claim 1, comprising means of using a communication device owned by the visitor to provide visitor characteristic information.
[1002] "Application Example 1"
[1003] (Claim 1)
[1004] In large-scale facilities, a means of acquiring visitors' biometric information upon entry,
[1005] A means for comparing collected biometric information with image data acquired from a video device,
[1006] A means of issuing a warning when matching biometric information is detected based on the matching process,
[1007] A means of notifying facility staff who receive a warning of the current location of the discovered visitor,
[1008] A means of identifying and protecting visitors using real-time video from multiple cameras,
[1009] A means of notifying the visitor's guardian of the location information of a protected visitor,
[1010] A system that includes this.
[1011] (Claim 2)
[1012] The system according to claim 1, which uses facial data as biometric information and matches it with image data using facial recognition technology.
[1013] (Claim 3)
[1014] The system according to claim 1, comprising means of using a communication terminal owned by the visitor to provide the visitor's biometric information.
[1015] "Example 2 of combining an emotion engine"
[1016] (Claim 1)
[1017] In large-scale facilities, a means of acquiring visitor biometric data upon entry,
[1018] A means for comparing collected biometric data with image data obtained from video equipment,
[1019] A means of issuing an alarm when matching biometric data is detected based on the matching process,
[1020] A means of notifying the staff member who received the alarm of the location information of the discovered visitor,
[1021] An emotion analysis method that analyzes the facial expressions of visitors to identify their emotional state,
[1022] A means of instructing staff on the most appropriate response method based on their emotional state,
[1023] A means of communication to share the visitor's location and emotional state with related persons such as parents,
[1024] A system that includes this.
[1025] (Claim 2)
[1026] The system according to claim 1, which uses facial information as biometric data and matches it with image data using a facial recognition method.
[1027] (Claim 3)
[1028] The system according to claim 1, comprising means of using a mobile communication device owned by the visitor to provide the visitor's biometric data.
[1029] "Application example 2 when combining with an emotional engine"
[1030] (Claim 1)
[1031] In large-scale facilities, a means of acquiring visitors' biometric information upon entry,
[1032] A means for comparing collected biometric information with image data acquired from a video device,
[1033] A means of issuing a warning when matching biometric information is detected based on the matching process,
[1034] A means of notifying facility staff who receive a warning of the current location and emotional state of the discovered visitor,
[1035] A means for facility employees to obtain visual information using visualization devices,
[1036] A system that includes this.
[1037] (Claim 2)
[1038] The system according to claim 1, which uses facial data as biometric information and matches it with image data using facial recognition technology and emotion analysis technology.
[1039] (Claim 3)
[1040] The system according to claim 1, comprising means of using a communication terminal owned by the visitor to provide the visitor's biometric information, and means of providing the information via a visualization device worn by an employee. [Explanation of Symbols]
[1041] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. In large-scale facilities, a means of acquiring visitors' biometric information upon entry, A means for comparing collected biometric information with image data acquired from a video device, A means of issuing a warning when matching biometric information is detected based on the matching process, A means of notifying facility staff who receive a warning of the current location of the discovered visitor, A means of identifying and protecting visitors using real-time video from multiple cameras, A means of notifying the visitor's guardian of the location information of a protected visitor, A system that includes this.
2. The system according to claim 1, which uses facial data as biometric information and matches it with image data using facial recognition technology.
3. The system according to claim 1, further comprising means of using a communication terminal owned by the visitor to provide the visitor's biometric information.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A